<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Sig. Proc.</journal-id>
<journal-title>Frontiers in Signal Processing</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Sig. Proc.</abbrev-journal-title>
<issn pub-type="epub">2673-8198</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">874200</article-id>
<article-id pub-id-type="doi">10.3389/frsip.2022.874200</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Signal Processing</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>BVI-CC: A Dataset for Research on Video Compression and Quality Assessment</article-title>
<alt-title alt-title-type="left-running-head">Katsenou et al.</alt-title>
<alt-title alt-title-type="right-running-head">Introducing BVI-CC Dataset</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Katsenou</surname>
<given-names>Angeliki</given-names>
</name>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1168997/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Zhang</surname>
<given-names>Fan</given-names>
</name>
<uri xlink:href="https://loop.frontiersin.org/people/1445585/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Afonso</surname>
<given-names>Mariana</given-names>
</name>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Dimitrov</surname>
<given-names>Goce</given-names>
</name>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Bull&#x2009;</surname>
<given-names>David R.</given-names>
</name>
</contrib>
</contrib-group>
<aff>
<institution>Visual Information Laboratory</institution>, <institution>Department of Electrical and Electronic Engineering</institution>, <institution>University of Bristol</institution>, <addr-line>Bristol</addr-line>, <country>United Kingdom</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1126837/overview">Raouf Hamzaoui</ext-link>, De Montfort University, United Kingdom</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1686922/overview">Xin Lu</ext-link>, De Montfort University, United Kingdom</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1687497/overview">Hadi Amirpour</ext-link>, University of Klagenfurt, Austria</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Angeliki Katsenou, <email>angeliki.katsenou@bristol.ac.uk</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Image Processing, a section of the journal Frontiers in Signal Processing</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>11</day>
<month>04</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>2</volume>
<elocation-id>874200</elocation-id>
<history>
<date date-type="received">
<day>11</day>
<month>02</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>09</day>
<month>03</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2022 Katsenou, Zhang, Afonso, Dimitrov and Bull&#x2009;.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Katsenou, Zhang, Afonso, Dimitrov and Bull&#x2009;</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>The video technology scenery has been very vivid over the past years, with novel video coding technologies introduced that promise improved compression performance over state-of-the-art technologies. Despite the fact that a lot of video datasets are available, representative content of the wide parameter space along with subjective evaluations of variations of encoded content from an unpartial end is required. In response to this requirement, this paper features a dataset, the BVI-CC. Three video codecs were deployed to create the variations of the encoded sequences: High Efficiency Video Coding Test Model (HM), AOMedia Video 1 (AV1), and Versatile Video Coding Test Model (VTM). Nine source video sequences were carefully selected to offer both diversity and representativeness in the spatio-temporal domain. Different spatial resolution versions of the sequences were created and encoded by all three codecs at pre-defined target bit rates. The compression efficiency of the codecs was evaluated with commonly used objective quality metrics, and the subjective quality of their reconstructed content was also evaluated through psychophysical experiments. Additionally, an adaptive bit rate (convex hull rate-distortion optimization across spatial resolutions) test case was assessed using both objective and subjective evaluations. Finally, the computational complexities of the tested codecs were examined. All data have been made publicly available as part of the dataset, which can be used for coding performance evaluation and video quality metric development.</p>
</abstract>
<kwd-group>
<kwd>video codec dataset</kwd>
<kwd>codec comparison</kwd>
<kwd>HEVC</kwd>
<kwd>AV1</kwd>
<kwd>VVC</kwd>
<kwd>objective quality assessment</kwd>
<kwd>subjective quality assessment</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>Video technology is ubiquitous in modern life, with wired and wireless video streaming, terrestrial and satellite TV, Blu-ray players, digital cameras, video conferencing and surveillance all underpinned by efficient signal representations. It was predicted that, by 2022, 82% (approximately 4.0&#xa0;ZB) of all global internet traffic per year will be video content <xref ref-type="bibr" rid="B12">CISCO (2018)</xref>. This projected figure was probably hit earlier due to the increased use of video technologies during the Covid-19 pandemic. It is therefore a very challenging time for compression, which must efficiently encode these increased quantities of video at higher spatial and temporal resolutions, dynamic resolutions and qualities.</p>
<p>The last 3&#xa0;decades have witnessed significant advances in video compression technology, from the first international video coding standard H.120 [<xref ref-type="bibr" rid="B27">ITU-T Rec. H.120 (1993)</xref>], to the widely adopted.</p>
<p>MPEG-2/H.262 [<xref ref-type="bibr" rid="B28">ITU-T Rec. H.262 (2012)</xref>], and H.264 Advanced Video Coding (H.264/AVC) [<xref ref-type="bibr" rid="B25">ITU-T Rec H.264 (2005)</xref>] standards. Recently, ISO/IEC Moving Picture Experts Group (MPEG) and ITU-T Video Coding Experts Group (VCEG) have released a new video coding standard, Versatile Video Coding (VVC) [<xref ref-type="bibr" rid="B8">Bross et al. (2019)</xref>], with the aim of reducing bit rates by 30&#x2013;50% compared to the current High Efficiency Video Coding (HEVC) standard [<xref ref-type="bibr" rid="B26">ITU-T Rec H.265 (2015)</xref>]. In parallel, the Alliance for Open Media (AOMedia) have developed royalty-free open-source video codecs to compete with MPEG standards. The first AOMedia Video 1 (AV1) codec [<xref ref-type="bibr" rid="B4">AOM (2019)</xref>; <xref ref-type="bibr" rid="B11">Chen et al. (2020)</xref>] has been reported to outperform its predecessor VP9, developed by <xref ref-type="bibr" rid="B19">Google (2017)</xref>. In order to benchmark these coding algorithms, their rate quality performance can be evaluated using objective and/or subjective assessment methods. Existing works, by <xref ref-type="bibr" rid="B3">Akyazi and Ebrahimi (2018)</xref>; <xref ref-type="bibr" rid="B20">Grois et al. (2016)</xref>; <xref ref-type="bibr" rid="B16">Dias et al. (2018)</xref>; <xref ref-type="bibr" rid="B21">Guo et al. (2018)</xref>; <xref ref-type="bibr" rid="B62">Zabrovskiy et al. (2018)</xref>; <xref ref-type="bibr" rid="B30">Katsavounidis and Guo (2018)</xref>; <xref ref-type="bibr" rid="B43">Nguyen and Marpe (2021)</xref>, have reported comparisons for contemporary codecs, with perplexing results and conclusions, mainly due to the use of different coding configurations. Also, most of these studies are solely based on objective quality assessment. Finally, the majority of these works do not publicly release the produced data.</p>
<p>The significant impact of data availability has always been important in video technology research and has become even more crucial over the past years due to the deployment of machine learning and deep learning methods [see <xref ref-type="bibr" rid="B40">Ma et al. (2021)</xref>]. Furthermore, existing datasets lack of variety in the encoded versions of the raw sequences, as the majority offers only H.264 e.g., VQEGHD3 [<xref ref-type="bibr" rid="B55">Video Quality Experts Group (2010)</xref>], LIVE [<xref ref-type="bibr" rid="B50">Seshadrinathan et al. (2010)</xref>], NFLX-P and VMAF&#x2b; [<xref ref-type="bibr" rid="B38">Li et al. (2016)</xref>] or HEVC e.g., BVI-HD [<xref ref-type="bibr" rid="B65">Zhang et al. (2018)</xref>], BVI-Texture [<xref ref-type="bibr" rid="B48">Papadopoulos et al. (2015)</xref>], BVI-SynTex [<xref ref-type="bibr" rid="B31">Katsenou A. V. et al. (2021)</xref>] encodings.</p>
<p>In this context, this paper presents a video dataset, referred to as BVI-CC, which comprises a complete set of 306 encodings using VVC, HEVC, and AV1, on nine representative source sequences typically used by the standardisation bodies. To this end, a controlled set of three experiments were designed taking into consideration the codecs corresponding common test conditions. The source sequences native spatial resolution is Ultra High Definition (UHD), 3840&#x00D7;2160. Additionally to the UHD resolution, the sequences were spatially downscaled to 1920&#x00D7;1080, 1280&#x00D7;720, and 960&#x00D7;540 resolution and three experiments were designed. Two experiments were traditionally configured using constant resolution, one at UHD and one at HD. The third experiment was designed on an adaptive bit rate use case implementing the Dynamic Optimizer (DO) approach<xref ref-type="fn" rid="fn1">
<sup>1</sup>
</xref> as described by <xref ref-type="bibr" rid="B29">Katsavounidis (2018)</xref> and implemented by <xref ref-type="bibr" rid="B33">Katsenou et al. (2019b)</xref> (up to FHD resolution only). This work provides a comprehensive extension of our previous work [<xref ref-type="bibr" rid="B33">Katsenou et al. (2019b)</xref>], where only AV1 and HEVC results were presented based on the DO approach. BVI-CC is complemented by data collected during the quality assessment: 1) anonymised opinion scores from psychovisual experiments in labs and 2) values of six commonly used objective quality metrics. BVI-CC dataset is publicly available after request [see <xref ref-type="bibr" rid="B34">Katsenou A. et al. (2021a)</xref>] and can be utilised either for research on video compression and research on image/video quality assessment.</p>
<p>The rest of this paper is organised as follows. <xref ref-type="sec" rid="s2">Section 2</xref> briefly reviews the history of video coding and related work on codec comparison. <xref ref-type="sec" rid="s3">Section 3</xref> presents the selected source sequences and the coding configurations employed in generating various compressed content. In <xref ref-type="sec" rid="s4">Section 4</xref>, the conducted subjective experiments are described in detail, while the evaluation results through both objective and subjective assessment are reported and discussed in <xref ref-type="sec" rid="s5">Section 5</xref>. Finally, <xref ref-type="sec" rid="s6">Section 6</xref> outlines the conclusion and future work.</p>
</sec>
<sec id="s2">
<title>2 Background</title>
<p>This section provides a brief overview of video coding standards and reports on existing datasets for research on video compression and quality assessment.</p>
<sec id="s2-1">
<title>2.1 Video Coding Standards and Technologies</title>
<p>Video coding standards normally define the syntax of bitstream and the decoding process, while encoders generate standard-compliant bitstream and thus determine compression performance. Each generation of video coding standard comes with a reference test model, such as HM (HEVC Test Model) for HEVC, which can be used to provide a performance benchmark. H.264/MPEG-4-AVC [ITU-T Rec H.264 (2005)] was launched in 2004, and is still the most prolific video coding standard, despite the fact that its successor, H.265/HEVC [ITU-T Rec H.265 (2015)] finalised in 2013, provides enhanced coding performance. In 2020, the first version of the latest video coding standard, Versatile Video Coding (VVC) [<xref ref-type="bibr" rid="B8">Bross et al. (2019)</xref>], has been finalised, which can achieve 30&#x2013;50% coding gain over H.265/HEVC, supporting immersive formats (360&#xb0; videos) and higher spatial resolutions, up to 16&#xa0;K.</p>
<p>Alongside recent MPEG standardisation, there has been increasing activity in the development of open-source royalty-free video codecs, particularly by the Alliance for Open Media (AOMedia), a consortium of video-related companies. VP9 [<xref ref-type="bibr" rid="B19">Google (2017)</xref>], which was earlier developed by Google to compete with MPEG, provided a basis for AV1 (AOMedia Video 1) [<xref ref-type="bibr" rid="B4">AOM (2019)</xref>; <xref ref-type="bibr" rid="B11">Chen et al. (2020)</xref>], which was released in 2018. <xref ref-type="bibr" rid="B11">Chen et al. (2020)</xref> set as AV1&#x2019;s main goal to provide an open source and royalty-free video coding format that substantially outperforms its primary market competitors in compression efficiency and concurrently achieves a practical decoding complexity, is optimized for hardware feasibility and scalability on modern devices. Since its release, its pool of contributors has expanded and they have provided significant updates towards its goal. The AOM contributors are currently working towards developing the next generation, AV2 [<xref ref-type="bibr" rid="B53">Shostak et al. (2021)</xref>]. For further details on existing video coding standards and formats, the readers are referred to <xref ref-type="bibr" rid="B9">Bull and Zhang (2021)</xref>; <xref ref-type="bibr" rid="B59">Wien (2015)</xref>; <xref ref-type="bibr" rid="B45">Ohm (2015)</xref>.</p>
<p>The performance of video coding algorithms is usually assessed by comparing their rate-distortion (RD) or rate-quality (RQ) performance on various test sequences. According to Recommendation ITU-R BT.500-12 (2012), the selection of test content is important and should provide a diverse and representative coverage of the video parameter space. Objective quality metrics or/and subjective opinion measurements are normally employed to assess compressed video quality. The overall RD or RQ performance difference between codecs can be then calculated using Bj&#xf8;ntegaard delta metrics for objective quality metrics according to <xref ref-type="bibr" rid="B5">Bj&#xf8;ntegaard (2001)</xref> or SCENIC developed by <xref ref-type="bibr" rid="B22">Hanhart and Ebrahimi (2014)</xref> for subjective assessments.</p>
<p>Most recent literature on video codec comparisons has focused on performance evaluations between MPEG codecs (H.264/AVC, HEVC, and VVC) and royalty-free (VP9 and AV1) codecs [<xref ref-type="bibr" rid="B3">Akyazi and Ebrahimi (2018)</xref>; <xref ref-type="bibr" rid="B20">Grois et al. (2016)</xref>; <xref ref-type="bibr" rid="B35">Lee et al. (2011)</xref>; <xref ref-type="bibr" rid="B16">Dias et al. (2018)</xref>; <xref ref-type="bibr" rid="B43">Nguyen and Marpe (2021)</xref>] and on their application in adaptive video steaming services <xref ref-type="bibr" rid="B21">Guo et al. (2018)</xref>; <xref ref-type="bibr" rid="B62">Zabrovskiy et al. (2018)</xref>; <xref ref-type="bibr" rid="B30">Katsavounidis and Guo (2018)</xref>. However, the results presented are acknowledged to be highly inconsistent [<xref ref-type="bibr" rid="B43">Nguyen and Marpe (2021)</xref>], mainly due to the different configurations employed across codecs.</p>
<p>Our work contributes towards a fair codec comparison by releasing publicly a dataset that includes objective and subjective codec comparisons in three different use cases: encoding at FHD and UHD resolution and encoding within the framework of adaptive streaming.</p>
</sec>
<sec id="s2-2">
<title>2.2 Datasets for Video Compression Purposes</title>
<p>The literature is rich of video datasets developed for many different purposes, mainly for computer vision related tasks such as object detection, action recognition, summarization, etc. These datasets, however, are not suitable for research in video compression. The main reason is that those have been designed for training deep network or models to infer very specific information. For example, the EPIC-Kitchens dataset [<xref ref-type="bibr" rid="B15">Damen et al. (2021)</xref>], contains scenes of daily activities in the kitchen, e.g., slicing bread, peeling carrots, stirring soup, etc. Therefore, the content bears significant similarities by repeating specific patterns in similar although diverse set-ups. Not providing a wide range of scenes with spatial and temporal information, this type of datasets cannot adequately represent video content and, thus, form a basis for a fair comparison of video compression algorithms and/or video quality metrics. Furthermore, the datasets published for computer vision tasks include already encoded versions (usually H.264-based encodings) as exported automatically from the video recording device. Although this type of content resembles the features of user generated content (UGC), a large portion of the videos streamed are coming from the creative industry. Thus research on video technology additionally requires pristine videos as exported from post-production.</p>
<p>In <xref ref-type="table" rid="T1">Table 1</xref>, a selection of commonly use datasets for video compression research is listed along with some of their basic features. Some of these datasets consist of pristine sequences [those that include raw (uncompressed) video sequences in the dataset] and others of UGC content [no raw sequences available as in KonViD-1K (<xref ref-type="bibr" rid="B24">Hosu et al. (2017)</xref>], YouTube-UGC [<xref ref-type="bibr" rid="B56">Wang et al. (2019)</xref>], DVL 2021 [<xref ref-type="bibr" rid="B61">Xing et al. (2022)</xref>]. Most of the existing datasets with pristine content, only include encoded sequences with one codec, usually H.264 {e.g., VQEGHD3 [<xref ref-type="bibr" rid="B55">Video Quality Experts Group (2010)</xref>], LIVE [<xref ref-type="bibr" rid="B50">Seshadrinathan et al. (2010)</xref>], NFLX-P and VMAF&#x2b; [<xref ref-type="bibr" rid="B38">Li et al. (2016)</xref>]} or HEVC {e.g., BVI-HD [<xref ref-type="bibr" rid="B65">Zhang et al. (2018)</xref>], BVI-Texture [<xref ref-type="bibr" rid="B48">Papadopoulos et al. (2015)</xref>], BVI-SynTex [<xref ref-type="bibr" rid="B31">Katsenou A. V. et al. (2021)</xref>]}.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Selection of state-of-the-art datasets for video research purposes.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Dataset</th>
<th align="center">Resolution</th>
<th align="center">Raw</th>
<th align="center">Encoded</th>
<th align="center">Codec</th>
<th align="center">QMs</th>
<th align="center">OS</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">VQEGHD3 [<xref ref-type="bibr" rid="B55">Video Quality Experts Group (2010)</xref>]</td>
<td align="center">1080p</td>
<td align="char" char=".">13</td>
<td align="center">72</td>
<td align="center">MPEG2/H264</td>
<td align="center">&#x2014;</td>
<td align="center">
<italic>&#x2713;</italic>
</td>
</tr>
<tr>
<td align="left">LIVE [<xref ref-type="bibr" rid="B50">Seshadrinathan et al. (2010)</xref>]</td>
<td align="center">1080p</td>
<td align="char" char=".">10</td>
<td align="center">150</td>
<td align="center">H.264</td>
<td align="center">&#x2014;</td>
<td align="center">
<italic>&#x2713;</italic>
</td>
</tr>
<tr>
<td align="left">MCL-V [<xref ref-type="bibr" rid="B39">Lin et al. (2015)</xref>]</td>
<td align="center">1080p</td>
<td align="char" char=".">12</td>
<td align="center">96</td>
<td align="center">H.264</td>
<td align="center">&#x2014;</td>
<td align="center">
<italic>&#x2713;</italic>
</td>
</tr>
<tr>
<td align="left">BVI-HD [<xref ref-type="bibr" rid="B65">Zhang et al. (2018)</xref>]</td>
<td align="center">1080p</td>
<td align="char" char=".">32</td>
<td align="center">192</td>
<td align="center">HEVC</td>
<td align="center">&#x2014;</td>
<td align="center">
<italic>&#x2713;</italic>
</td>
</tr>
<tr>
<td align="left">BVI-Texture [<xref ref-type="bibr" rid="B48">Papadopoulos et al. (2015)</xref>]</td>
<td align="center">1080p</td>
<td align="char" char=".">20</td>
<td align="center">80</td>
<td align="center">HEVC</td>
<td align="center">&#x2014;</td>
<td align="center">
<italic>&#x2713;</italic>
</td>
</tr>
<tr>
<td align="left">BVI-SynTex [<xref ref-type="bibr" rid="B31">Katsenou et al. (2021b)</xref>]</td>
<td align="center">1080p</td>
<td align="char" char=".">49</td>
<td align="center">196</td>
<td align="center">HEVC</td>
<td align="center">
<italic>&#x2713;</italic>
</td>
<td align="center">
<italic>&#x2713;</italic>
</td>
</tr>
<tr>
<td align="left">BVI-DVC [<xref ref-type="bibr" rid="B40">Ma et al. (2021)</xref>]</td>
<td align="center">up to 2160p</td>
<td align="char" char=".">800</td>
<td align="center">&#x2014;</td>
<td align="center">&#x2014;</td>
<td align="center">&#x2014;</td>
<td align="center">&#x2014;</td>
</tr>
<tr>
<td align="left">100-4K [<xref ref-type="bibr" rid="B32">Katsenou et al. (2019a)</xref>]</td>
<td align="center">2160p</td>
<td align="char" char=".">100</td>
<td align="center">&#x2014;</td>
<td align="center">&#x2014;</td>
<td align="center">&#x2014;</td>
<td align="center">&#x2014;</td>
</tr>
<tr>
<td align="left">VMAF&#x2b; [<xref ref-type="bibr" rid="B38">Li et al. (2016)</xref>]</td>
<td align="center">up to 1080p</td>
<td align="char" char=".">23</td>
<td align="center">230</td>
<td align="center">H.264</td>
<td align="center">&#x2014;</td>
<td align="center">
<italic>&#x2713;</italic>
</td>
</tr>
<tr>
<td align="left">NFLX-P [<xref ref-type="bibr" rid="B38">Li et al. (2016)</xref>]</td>
<td align="center">up to 1080p</td>
<td align="char" char=".">9</td>
<td align="center">70</td>
<td align="center">H.264</td>
<td align="center">&#x2014;</td>
<td align="center">
<italic>&#x2713;</italic>
</td>
</tr>
<tr>
<td align="left">YouTube-UGC [<xref ref-type="bibr" rid="B56">Wang et al. (2019)</xref>]</td>
<td align="center">up to 2160p</td>
<td align="center">&#x2014;</td>
<td align="center">&#x223c; 1,500</td>
<td align="center">H.264, VP92</td>
<td align="center">
<italic>&#x2713;</italic>
</td>
<td align="center">
<italic>&#x2713;</italic>
</td>
</tr>
<tr>
<td align="left">KonViD-1K [<xref ref-type="bibr" rid="B24">Hosu et al. (2017)</xref>]</td>
<td align="center">540p</td>
<td align="center">&#x2014;</td>
<td align="center">1,200</td>
<td align="center">H.264</td>
<td align="center">&#x2014;</td>
<td align="center">
<italic>&#x2713;</italic>
</td>
</tr>
<tr>
<td align="left">DVL 2021 [<xref ref-type="bibr" rid="B61">Xing et al. (2022)</xref>]</td>
<td align="center">2160p</td>
<td align="center">&#x2014;</td>
<td align="center">206</td>
<td align="center">&#x2014;</td>
<td align="center">&#x2014;</td>
<td align="center">
<italic>&#x2713;</italic>
</td>
</tr>
<tr>
<td align="left">BVI-CC</td>
<td align="center">up to 2160p</td>
<td align="char" char=".">9</td>
<td align="center">306</td>
<td align="center">HEVC, VVC, AV1</td>
<td align="center">
<italic>&#x2713;</italic>
</td>
<td align="center">
<italic>&#x2713;</italic>
</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>From <xref ref-type="table" rid="T1">Table 1</xref>, it is evident that there is no dataset available that offers variations of encoded sequences based on different state-of-the-art codecs. Furthermore, to the best of our knowledge, there is no other dataset available that offers encoded sequences based on VVC and AV1. This is a very important contribution of the introduced dataset, BVI-CC, as it could facilitate research on video compression and comparison across these three different codecs. It is also noticeable that although almost all datasets provide opinion scores (OS) data, most of the datasets do not provide computed values of objective quality metrics (QMs).</p>
</sec>
</sec>
<sec id="s3">
<title>3 Test Content and Codec Configurations</title>
<p>This section describes the selection of source sequences and the different codec configurations used to generate their various compressed versions.</p>
<sec id="s3-1">
<title>3.1 Source Sequence Selection</title>
<p>Nine source sequences were selected from <xref ref-type="bibr" rid="B23">Harmonic (2019)</xref>, BVI-Texture [<xref ref-type="bibr" rid="B48">Papadopoulos et al. (2015)</xref>] and JVET (Joint Video Exploration Team) CTC (Common Test Conditions) datasets. Each sequence is progressively scanned, at UHD resolution, with a frame rate of 60 frames per second (fps) and without scene-cuts. All were truncated from their original lengths to 5&#xa0;seconds [rather than the recommended 10&#xa0;s in ITU standard Recommendation ITU-R BT.500-12 (2012)]. This reflects the recommendations of a recent study on optimal video duration for subjective quality assessment by <xref ref-type="bibr" rid="B41">Mercer Moss et al. (2016a</xref>, <xref ref-type="bibr" rid="B42">2016b)</xref>. Sample frames from the selected nine source sequences alongside clip names and indices are shown in <xref ref-type="fig" rid="F1">Figure 1</xref>. The dataset includes three sequences with only local motion (without any camera motion, V1-V3), three sequences with dynamic textures [for definitions see <xref ref-type="bibr" rid="B64">Zhang and Bull (2011)</xref>, V4-V6], and three with complex camera movements (V7-V9). The coverage of the video parameter space is confirmed in <xref ref-type="fig" rid="F2">Figure 2</xref>, where the Spatial and Temporal Information (SI and TI, respectively), the colourfulness (CF), and average contrast, as defined by <xref ref-type="bibr" rid="B60">Winkler (2012)</xref> are plotted.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Sample video frames from the selected source sequences.</p>
</caption>
<graphic xlink:href="frsip-02-874200-g001.tif"/>
</fig>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Scatter plots of low level content characteristics: for the selected source sequences: <bold>(A)</bold> SI vs. TI and <bold>(B)</bold> CF vs. Contrast.</p>
</caption>
<graphic xlink:href="frsip-02-874200-g002.tif"/>
</fig>
<p>In order to investigate coding performance across different resolutions and within an adaptive streaming framework, three spatial resolution groups were generated from the source sequences: (A) UHD (3840&#x00D7;2160) only, (B) HD (1920 601,080) only, and (C) HD-Dynamic Optimizer (HD-DO). For group C, coding results for three different resolutions (1920&#x00D7;1080, 1,280&#x00D7;720, and 960&#x00D7;540) and with various quantisation parameters (QPs) were firstly generated. The reconstructed videos were then up-sampled to HD resolution (in order to provide a basis for comparison with the original HD sequences). Here, spatial resolution re-sampling was implemented using Lanczos-3 filters designed by <xref ref-type="bibr" rid="B18">Duchon (1979)</xref>. The rate points with optimal rate-quality performance based on Video Multimethod Assessment Fusion (VMAF) [<xref ref-type="bibr" rid="B38">Li et al. (2016)</xref>] were selected across the three tested resolutions for each target bit rate and codec. This process is repeated to create the entire convex hull in the DO approach [<xref ref-type="bibr" rid="B29">Katsavounidis (2018)</xref>]. The resulting selected resolutions across the set of target bit rates are reported in <xref ref-type="table" rid="T2">Table 2</xref>.</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Resolution selections per sequence after applying the DO methodology on AV1 and HM for resolution group C.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="left">Sequences</th>
<th colspan="5" align="center">HM</th>
<th colspan="5" align="center">AV1</th>
</tr>
<tr>
<th align="center">R1</th>
<th align="center">R2</th>
<th align="center">R3</th>
<th align="center">R4</th>
<th align="center">R5</th>
<th align="center">R1</th>
<th align="center">R2</th>
<th align="center">R3</th>
<th align="center">R4</th>
<th align="center">R5</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">V1: AirAcrobatic</td>
<td align="char" char=".">720p</td>
<td align="char" char=".">1080p</td>
<td align="char" char=".">1080p</td>
<td align="char" char=".">1080p</td>
<td align="char" char=".">1080p</td>
<td align="char" char=".">720p</td>
<td align="char" char=".">720p</td>
<td align="char" char=".">720p</td>
<td align="char" char=".">720p</td>
<td align="char" char=".">720p</td>
</tr>
<tr>
<td align="left">V2: CatRobot</td>
<td align="char" char=".">1080p</td>
<td align="char" char=".">1080p</td>
<td align="char" char=".">1080p</td>
<td align="char" char=".">1080p</td>
<td align="char" char=".">1080p</td>
<td align="char" char=".">720p</td>
<td align="char" char=".">720p</td>
<td align="char" char=".">1080p</td>
<td align="char" char=".">1080p</td>
<td align="char" char=".">1080p</td>
</tr>
<tr>
<td align="left">V3: Myanmar</td>
<td align="char" char=".">1080p</td>
<td align="char" char=".">1080p</td>
<td align="char" char=".">1080p</td>
<td align="char" char=".">1080p</td>
<td align="char" char=".">1080p</td>
<td align="char" char=".">720p</td>
<td align="char" char=".">1080p</td>
<td align="char" char=".">1080p</td>
<td align="char" char=".">1080p</td>
<td align="char" char=".">1080p</td>
</tr>
<tr>
<td align="left">V4: CalmingWater</td>
<td align="char" char=".">544p</td>
<td align="char" char=".">544p</td>
<td align="char" char=".">544p</td>
<td align="char" char=".">544p</td>
<td align="char" char=".">720p</td>
<td align="char" char=".">544p</td>
<td align="char" char=".">544p</td>
<td align="char" char=".">544p</td>
<td align="char" char=".">544p</td>
<td align="char" char=".">720p</td>
</tr>
<tr>
<td align="left">V5: ToddlerFountain</td>
<td align="char" char=".">544p</td>
<td align="char" char=".">544p</td>
<td align="char" char=".">544p</td>
<td align="char" char=".">544p</td>
<td align="char" char=".">720p</td>
<td align="char" char=".">544p</td>
<td align="char" char=".">544p</td>
<td align="char" char=".">544p</td>
<td align="char" char=".">544p</td>
<td align="char" char=".">720p</td>
</tr>
<tr>
<td align="left">V6: LampLeaves</td>
<td align="char" char=".">720p</td>
<td align="char" char=".">720p</td>
<td align="char" char=".">1080p</td>
<td align="char" char=".">1080p</td>
<td align="char" char=".">1080p</td>
<td align="char" char=".">544p</td>
<td align="char" char=".">544p</td>
<td align="char" char=".">720p</td>
<td align="char" char=".">720p</td>
<td align="char" char=".">720p</td>
</tr>
<tr>
<td align="left">V7: DaylightRoad</td>
<td align="char" char=".">720p</td>
<td align="char" char=".">1080p</td>
<td align="char" char=".">1080p</td>
<td align="char" char=".">1080p</td>
<td align="char" char=".">1080p</td>
<td align="char" char=".">720p</td>
<td align="char" char=".">720p</td>
<td align="char" char=".">1080p</td>
<td align="char" char=".">1080p</td>
<td align="char" char=".">1080p</td>
</tr>
<tr>
<td align="left">V8: RedRock</td>
<td align="char" char=".">720p</td>
<td align="char" char=".">720p</td>
<td align="char" char=".">1080p</td>
<td align="char" char=".">1080p</td>
<td align="char" char=".">1080p</td>
<td align="char" char=".">544p</td>
<td align="char" char=".">720p</td>
<td align="char" char=".">720p</td>
<td align="char" char=".">720p</td>
<td align="char" char=".">720p</td>
</tr>
<tr>
<td align="left">V9: RollerCoaster</td>
<td align="char" char=".">544p</td>
<td align="char" char=".">544p</td>
<td align="char" char=".">720p</td>
<td align="char" char=".">720p</td>
<td align="char" char=".">1080p</td>
<td align="char" char=".">544p</td>
<td align="char" char=".">544p</td>
<td align="char" char=".">720p</td>
<td align="char" char=".">720p</td>
<td align="char" char=".">720p</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3-2">
<title>3.2 Coding Configurations</title>
<p>The reference test models of HEVC and VVC, and their major competitor, AV1 have been evaluated in this study. Each codec was configured using the coding parameters defined in their common test conditions [see <xref ref-type="bibr" rid="B51">Sharman and Suehring (2018)</xref>; <xref ref-type="bibr" rid="B6">Bossen et al. (2019)</xref>; <xref ref-type="bibr" rid="B14">Daede et al. (2019)</xref>], with fixed quantisation parameters (rate control disabled), the same structural delay (e.g., defined as GOP size in the HEVC HM software) of 16 frames and the same random access intervals (e.g., defined as IntraPeriod in the HEVC HM software) of 64 frames. The actual codec versions and configuration parameters are provided in <xref ref-type="table" rid="T3">Table 3</xref>.</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>The software versions and configurations of the evaluated video codecs.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Codec</th>
<th align="center">Version</th>
<th align="center">Configuration parameters</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">HEVC HM</td>
<td align="center">16.18</td>
<td align="left">Random access configuration for Main10 profile <xref ref-type="bibr" rid="B51">Sharman and Suehring (2018)</xref>. IntraPeriod &#x3d; 64 and GOPSize &#x3d; 16</td>
</tr>
<tr>
<td align="left">AOM AV1</td>
<td align="center">0.1.0-9647-ga6fa0877f</td>
<td align="left">Common settings with high latency CQP configuration <xref ref-type="bibr" rid="B14">Daede et al. (2019)</xref>. Other coding parameters: passes &#x3d; 2, cpu-used &#x3d; 1, kf-max-dist &#x3d; 64, kf-min-dist &#x3d; 64, arnr-maxframes &#x3d; 7, arnr-strength &#x3d; 5, lag-in-frames &#x3d; 16, aq-mode &#x3d; 0, bias-pct &#x3d; 100, minsection-pct &#x3d; 1, maxsection-pct &#x3d; 10,000, auto-alt-ref &#x3d; 1, min-q &#x3d; 0, max-q &#x3d; 63, max-gf-interval &#x3d; 16, min-gf-interval &#x3d; 4 and color-primaries &#x3d; bt709</td>
</tr>
<tr>
<td rowspan="2" align="left">VVC VTM</td>
<td rowspan="2" align="center">4.01</td>
<td align="left">Random access configuration <xref ref-type="bibr" rid="B6">Bossen et al. (2019)</xref>
</td>
</tr>
<tr>
<td align="left">IntraPeriod &#x3d; 64 and GOPSize &#x3d; 16</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Different target bit rates were pre-determined for each test sequence and for each resolution group (four points for resolution group A and B, and five for HD-DO group), and their values are shown in <xref ref-type="table" rid="T4">Table 4</xref>. These were determined based on the preliminary encoding results of the test sequences for each resolution group using AV1. This decision was made because the version of AV1 employed restricted production of bitstreams at pre-defined bit rates, as only integer quantisation parameters could be used. On the other hand, for HEVC HM and VVC VTM this was easier to achieve by enabling the &#x201c;QPIncrementFrame&#x201d; parameter. In order to achieve these target bit rates, the quantisation parameter values were iteratively adjusted to ensure the output bit rates were sufficiently close to the targets (within a range of &#xb1;3<italic>%</italic>).</p>
<table-wrap id="T4" position="float">
<label>TABLE 4</label>
<caption>
<p>Pre-determined target bit rates for all test sequences in three resolution groups.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="3" align="left">Sequence</th>
<th colspan="13" align="center">Target bit rates (kbps)</th>
</tr>
<tr>
<th colspan="4" align="center">Resolution group A (UHD only)</th>
<th colspan="4" align="center">Resolution group B (HD only)</th>
<th colspan="5" align="center">Resolution group C (HD-DO)</th>
</tr>
<tr>
<th align="center">R1</th>
<th align="center">R2</th>
<th align="center">R3</th>
<th align="center">R4</th>
<th align="center">R1</th>
<th align="center">R2</th>
<th align="center">R3</th>
<th align="center">R4</th>
<th align="center">R1</th>
<th align="center">R2</th>
<th align="center">R3</th>
<th align="center">R4</th>
<th align="center">R5</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">V1: AirAcrobatic</td>
<td align="right">1,300</td>
<td align="right">2,250</td>
<td align="right">4,700</td>
<td align="right">9,270</td>
<td align="right">550</td>
<td align="right">920</td>
<td align="right">1850</td>
<td align="right">3,400</td>
<td align="right">305</td>
<td align="right">575</td>
<td align="right">940</td>
<td align="right">1770</td>
<td align="right">3,350</td>
</tr>
<tr>
<td align="left">V2: CatRobot</td>
<td align="right">3,170</td>
<td align="right">5,450</td>
<td align="right">8,450</td>
<td align="right">14,500</td>
<td align="right">1,480</td>
<td align="right">2,200</td>
<td align="right">3,250</td>
<td align="right">5,440</td>
<td align="right">910</td>
<td align="right">1,500</td>
<td align="right">2,200</td>
<td align="right">3,250</td>
<td align="right">5,500</td>
</tr>
<tr>
<td align="left">V3: Myanmar</td>
<td align="right">11,000</td>
<td align="right">21,100</td>
<td align="right">33,500</td>
<td align="right">46,000</td>
<td align="right">3,450</td>
<td align="right">5,500</td>
<td align="right">8,200</td>
<td align="right">10,800</td>
<td align="right">2,100</td>
<td align="right">3,450</td>
<td align="right">5,500</td>
<td align="right">8,150</td>
<td align="right">10,800</td>
</tr>
<tr>
<td align="left">V4: CalmingWater</td>
<td align="right">10,100</td>
<td align="right">19,250</td>
<td align="right">30,000</td>
<td align="right">50,000</td>
<td align="right">3,100</td>
<td align="right">6,400</td>
<td align="right">12,000</td>
<td align="right">21,000</td>
<td align="right">1,140</td>
<td align="right">3,050</td>
<td align="right">6,550</td>
<td align="right">12,200</td>
<td align="right">20,500</td>
</tr>
<tr>
<td align="left">V5: ToddlerFountain</td>
<td align="right">13,180</td>
<td align="right">27,000</td>
<td align="right">38,350</td>
<td align="right">69,800</td>
<td align="right">6,150</td>
<td align="right">13,500</td>
<td align="right">21,500</td>
<td align="right">34,500</td>
<td align="right">2,900</td>
<td align="right">5,900</td>
<td align="right">13,200</td>
<td align="right">20,500</td>
<td align="right">34,900</td>
</tr>
<tr>
<td align="left">V6: LampLeaves</td>
<td align="right">14,550</td>
<td align="right">26,460</td>
<td align="right">43,900</td>
<td align="right">69,800</td>
<td align="right">8,100</td>
<td align="right">14,200</td>
<td align="right">20,500</td>
<td align="right">33,000</td>
<td align="right">5,030</td>
<td align="right">8,100</td>
<td align="right">14,000</td>
<td align="right">20,500</td>
<td align="right">33,500</td>
</tr>
<tr>
<td align="left">V7: DaylightRoad</td>
<td align="right">2,650</td>
<td align="right">4,450</td>
<td align="right">7,050</td>
<td align="right">12,170</td>
<td align="right">1,220</td>
<td align="right">1800</td>
<td align="right">3,000</td>
<td align="right">5,300</td>
<td align="right">810</td>
<td align="right">1,220</td>
<td align="right">1800</td>
<td align="right">3,000</td>
<td align="right">5,300</td>
</tr>
<tr>
<td align="left">V8: RedRock</td>
<td align="right">1,500</td>
<td align="right">2,600</td>
<td align="right">4,000</td>
<td align="right">6,380</td>
<td align="right">680</td>
<td align="right">1,000</td>
<td align="right">1,650</td>
<td align="right">2,500</td>
<td align="right">460</td>
<td align="right">650</td>
<td align="right">1,020</td>
<td align="right">1,600</td>
<td align="right">2,450</td>
</tr>
<tr>
<td align="left">V9: RollerCoaster</td>
<td align="right">1750</td>
<td align="right">2,880</td>
<td align="right">4,600</td>
<td align="right">7,350</td>
<td align="right">880</td>
<td align="right">1,480</td>
<td align="right">2,270</td>
<td align="right">3,564</td>
<td align="right">550</td>
<td align="right">850</td>
<td align="right">1,480</td>
<td align="right">2,280</td>
<td align="right">3,580</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3-3">
<title>3.3 Summary</title>
<p>In summary, a total number of 306 distorted sequences were produced: there are 108 (9 source sequences &#xd7; 4 rate points &#xd7; 3 codecs) for Resolution Group A (UHD only), 108 (9 3 4 co3) for Resolution Group B (HD only), and 90 (9 ) for 2) for Resolution Group C (HD-DO)<xref ref-type="fn" rid="fn2">
<sup>2,</sup>
</xref>
<xref ref-type="fn" rid="fn3">
<sup>3</sup>
</xref>.</p>
</sec>
</sec>
<sec id="s4">
<title>4 Subjective Experiments</title>
<p>Three subjective experiment sessions were conducted separately on the test sequences in the three resolution groups. The experimental setup, procedure, test methodology and data processing approach are reported in this section.</p>
<sec id="s4-1">
<title>4.1 Environmental Setup</title>
<p>All three experiment sessions were conducted in a darkened, living room-style environment. The background luminance level was set to 15% of the peak luminance of the monitor used (62.5 lux) Recommendation ITU-R BT.500-12 (2012). All test sequences were shown at their native spatial resolution and frame rates, on a consumer display, a SONY KD65Z9D LCD TV, which measures 1,429 29804&#xa0;mm, with a peak luminance of 410 lux. The viewing distance was set to 121&#xa0;cm (1.5 times the screen height) for Resolution Group A (UHD) and 241&#xa0;cm (three times the screen height) for Resolution Group B (HD) and C (HD-DO), following the <xref ref-type="bibr" rid="B49">Recommendation ITU-R BT.500-12 (2012)</xref> and <xref ref-type="bibr" rid="B46">P.910 (1999)</xref>. The presentation of video sequences was controlled by a Windows PC running an open source software, BVI-SVQA, developed by <xref ref-type="bibr" rid="B17">Dimitrov et al. (2019)</xref> at the University of Bristol for psychophysical experiments.</p>
</sec>
<sec id="s4-2">
<title>4.2 Experimental Procedure</title>
<p>In all three experiments, the Double Stimulus Continuous Quality Scale (DSCQS) [Recommendation ITU-R BT.500-12 (2012)] methodology was used. In each trial, participants were shown a pair of sequences twice, including original and encoded versions. The presentation order was randomised in each trial and was unknown to each participant. Participants had unlimited time to respond to the following question (presented on the video monitor): &#x201c;Please rate the quality (0&#x2013;100) of the first/second video. Excellent&#x2013;90, Good&#x2013;70, Fair&#x2013;50, Poor&#x2013;30 and Bad&#x2013;10&#x201d;. Participants then used a mouse to scroll through the vertical scale and score (0&#x2013;100) for these two videos. The total duration of each experimental session was approximately 50 (Resolution Group A and B) or 60 (Resolution Group C) minutes, and each was split into two sub-sessions with a 10&#xa0;min break in between. Before the formal test, a training session was conducted, under the supervision of the instructor, consisting of three trials (different from those used in the formal test) to allow the participants time to familiarize.</p>
</sec>
<sec id="s4-3">
<title>4.3 Participants and Data Processing</title>
<p>A total of 62 subjects, with an average age of 27 (age range 20&#x2013;45), from the University of Bristol (students and staff), were compensated for their participation in the experiments. All of them were tested for normal or corrected-to normal vision. Consent forms were signed by each participant and the data were anonymized. Responses from the subjects were first recorded as quality scores in the range 0&#x2013;100, as explained earlier. Difference scores were then calculated for each tested sequence and each subject <italic>i</italic> by subtracting the quality score of the distorted sequence <italic>OS</italic>
<sub>
<italic>dis</italic>
</sub> from its corresponding reference <italic>OS</italic>
<sub>
<italic>ref</italic>
</sub>. Difference Mean Opinion Scores (DMOS) and the respective statistics (standard error, confidence intervals, etc) were then obtained for each trial by taking the mean of the difference scores among participants <italic>N</italic>. Particularly, DMOS values for a video <italic>v</italic> were calculated as follows:<disp-formula id="e1">
<mml:math id="m1">
<mml:mi>D</mml:mi>
<mml:mi>M</mml:mi>
<mml:mi>O</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>O</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>ref</mml:mtext>
<mml:mo>,</mml:mo>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>O</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>dis</mml:mtext>
<mml:mo>,</mml:mo>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>.</mml:mo>
</mml:math>
<label>(1)</label>
</disp-formula>
</p>
<p>In order to reflect the quality taking into account the relative difference from the score of the hidden reference (in this case the original) for each participant <italic>i</italic>, we subtracted the <italic>DMOS</italic>
<sub>
<italic>v</italic>
</sub> from the maximum quality score (100).</p>
<p>Further to the score collection, we performed post-screening of the subjects. Following the Recommendation ITU-R BT.500-12 (2012) protocols, for each resolution group, we performed outlier rejection on all participant scores. No participants were rejected. Moreover, we calculated the subject bias according the Recommendation <xref ref-type="bibr" rid="B47">P.913 (2021)</xref> and removed it before performing any statistical analysis. As defined in the recommendation, subject bias is the difference between the average of one subject&#x2019;s ratings and the average of all subjects&#x2019; ratings for each processed video sequence. To remove subject bias, the recommendation proposed to subtract that its value from each one of subject&#x2019;s ratings.</p>
</sec>
</sec>
<sec sec-type="results|discussion" id="s5">
<title>5 Results and Discussion</title>
<p>This section presents the codec comparison results based on objective and subjective quality assessments of BVI-CC, alongside encoder and decoder complexity assessments. For the objective evaluation, two video quality metrics have been employed: the commonly used Peak-Signal-to-Noise-Ratio (PSNR) and Video Multi-method Assessment Fusion (VMAF) <xref ref-type="bibr" rid="B38">Li et al. (2016)</xref>. The latter is a machine learning-based video quality metric, which predicts subjective quality by combining multiple quality metrics and video features, including the Detail Loss Metric (DLM) [<xref ref-type="bibr" rid="B37">Li et al. (2011)</xref>], Visual Information Fidelity measure (VIF) [<xref ref-type="bibr" rid="B52">Sheikh et al. (2005)</xref>], and averaged temporal frame difference [<xref ref-type="bibr" rid="B38">Li et al. (2016)</xref>]. The fusion process employs a &#x3bd;-Support Vector machine (&#x3bd;-SVM) regressor [<xref ref-type="bibr" rid="B13">Cortes and Vapnik (1995)</xref>]. VMAF has been evaluated on various video quality databases, and shows improved correlation with subjective scores [<xref ref-type="bibr" rid="B38">Li et al. (2016)</xref>; <xref ref-type="bibr" rid="B65">Zhang et al. (2018)</xref>; <xref ref-type="bibr" rid="B36">Li et al. (2019)</xref>]. In this work, VMAF has also been employed to determine optimum resolution for each test rate point and sequence, following the procedure described in <xref ref-type="sec" rid="s3-2">Section 3.2</xref>. The difference between test video codecs in terms of coding efficiency was calculated using the <xref ref-type="bibr" rid="B5">Bj&#xf8;ntegaard (2001)</xref> Delta (BD) measurements benchmarked against HEVC HM.</p>
<p>For the subjective assessment of the BVI-CC dataset, after following the experimental procedure defined in <xref ref-type="sec" rid="s4-2">Section 4.2</xref>, the raw opinion scores were collected for each trial in confidentiality and anonymized. The rate-quality curves have been plotted for each test sequence in all three resolution groups (see <xref ref-type="sec" rid="s5-2">Section 5.2</xref>), where the subjective quality is defined as 100-DMOS (see <xref ref-type="disp-formula" rid="e1">Eq. 1</xref>). Before computing the DMOS values and performing a codec comparison, the subject bias was estimated and removed according to the Recommendation P.913 (2021). A significance test was then conducted using one-way Analysis of Variance (ANOVA) between each paired of codecs on all rate points and sequences.</p>
<p>The subjective data collected was also used to evaluate six popular objective video quality metrics (see <xref ref-type="sec" rid="s5-3">Section 5.3</xref>), including PSNR, Structural Similarity Index (SSIM) [<xref ref-type="bibr" rid="B57">Wang et al. (2004)</xref>], multi-scale SSIM (MS-SSIM) [<xref ref-type="bibr" rid="B58">Wang et al. (2003)</xref>], VIF [<xref ref-type="bibr" rid="B52">Sheikh et al. (2005)</xref>], Visual Signal-to-Noise Ratio (VSNR) [<xref ref-type="bibr" rid="B10">Chandler and Hemami (2007)</xref>], and VMAF [<xref ref-type="bibr" rid="B38">Li et al. (2016)</xref>]. According to the recommendation the bias was only removed from the opinion scores for the subjective comparison of the codecs. For the objective comparison, the raw opinion scores were utilized. Following the procedure in <xref ref-type="bibr" rid="B54">Video Quality Experts Group (2000)</xref>, their quality indices and the subjective DMOS were fitted based on a weighted least-squares approach using a logistic fitting function for three different resolution groups. The correlation performance of these quality metrics was assessed using four correlation statistics, the Spearman Rank Order Correlation Coefficient (SROCC), the Linear Correlation Coefficient (LCC), the Outlier Ratio (OR) and the Root Mean Squared Error (RMSE). The definitions of these parameters can be found in <xref ref-type="bibr" rid="B54">Video Quality Experts Group (2000)</xref>.</p>
<p>Finally, the computational complexity of the three tested encoders was calculated and normalised to HEVC HM for Resolution Group A and B (see <xref ref-type="sec" rid="s5-4">Section 5.4</xref>). They were executed on the CPU nodes of a shared cluster, Blue Crystal Phase 3, of the <xref ref-type="bibr" rid="B1">Advanced Computing Research Centre, University of Bristol (2021)</xref>. Each node has 16 6E2.6&#xa0;GHz SandyBridge cores and 64&#xa0;GB RAM.</p>
<sec id="s5-1">
<title>5.1 Results Based on Objective Quality Assessment</title>
<p>
<xref ref-type="table" rid="T5">Table 5</xref> summarises the Bj&#xf8;ntegaard Delta measurements (BD-rate) of AOM AV1 (for three resolution groups) and VVC VTM (for Resolution Group A and B only) compared with HEVC HM, based on both PSNR and VMAF. For the tested codec versions and configurations, it can be observed that AV1 achieves an average bit rate saving of 7.3% against HEVC HM for the UHD test content assessed by PSNR, and this figure reduces (3.8%) at HD resolution. When VMAF is employed for quality assessment, the coding gains of AV1 over HM are slightly higher, averaging 8.6 and 5.0% for UHD and HD respectively. Comparing to AV1, VTM provides significant bit rate savings for both HD and UHD test content, with average BD-rate values between &#x2212;27% and &#x2212;30% for PSNR and VMAF. For resolution group C, where VMAF-based DO was applied for HM and AV1, the coding gain achieved by AV1 is 6.3% (over HM) assessed by VMAF, while there is a BD-rate (1.8%) loss when PSNR is employed. In overall conclusion, the performance of AV1 makes a small improvement over HM on the test content, and both AV1 and HM perform (significantly) worse than VTM.</p>
<table-wrap id="T5" position="float">
<label>TABLE 5</label>
<caption>
<p>Codec comparison results based on PSNR and VMAF quality metrics. Here Bj&#xf8;ntegaard Delta <xref ref-type="bibr" rid="B5">Bj&#xf8;ntegaard (2001)</xref> measurements (BD-rate) were employed, and HEVC HM was used as benchmark.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Resolution group</th>
<th colspan="4" align="center">A (UHD)</th>
<th colspan="4" align="center">B (HD)</th>
<th colspan="2" align="center">C (HD-DO)</th>
</tr>
<tr>
<th align="left">Codec</th>
<th colspan="2" align="center">PSNR</th>
<th colspan="2" align="center">VMAF</th>
<th colspan="2" align="center">PSNR</th>
<th colspan="2" align="center">VMAF</th>
<th align="center">PSNR</th>
<th align="center">VMAF</th>
</tr>
<tr>
<th align="left">Sequence&#x2216;BD-rate</th>
<th align="center">AV1</th>
<th align="center">VTM</th>
<th align="center">AV1</th>
<th align="center">VTM</th>
<th align="center">AV1</th>
<th align="center">VTM</th>
<th align="center">AV1</th>
<th align="center">VTM</th>
<th align="center">AV1</th>
<th align="center">AV1</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">AirAcrobatic</td>
<td align="center">&#x2212;12.1%</td>
<td align="center">&#x2212;25.5%</td>
<td align="center">&#x2212;12.0%</td>
<td align="center">&#x2212;28.6%</td>
<td align="center">&#x2212;2.6%</td>
<td align="center">&#x2212;21.7%</td>
<td align="center">4.2%</td>
<td align="center">&#x2212;20.0%</td>
<td align="center">13.3%</td>
<td align="center">&#x2212;0.1%</td>
</tr>
<tr>
<td align="left">CatRobot</td>
<td align="center">&#x2212;6.2%</td>
<td align="center">&#x2212;38.0%</td>
<td align="center">&#x2212;12.8%</td>
<td align="center">&#x2212;39.6%</td>
<td align="center">&#x2212;4.0%</td>
<td align="center">&#x2212;37.7%</td>
<td align="center">&#x2212;10.4%</td>
<td align="center">&#x2212;41.2%</td>
<td align="center">&#x2212;2.1%</td>
<td align="center">&#x2212;11.3%</td>
</tr>
<tr>
<td align="left">Myanmar</td>
<td align="center">4.3%</td>
<td align="center">&#x2212;17.2%</td>
<td align="center">1.3%</td>
<td align="center">&#x2212;21.3%</td>
<td align="center">6.5%</td>
<td align="center">&#x2212;15.5%</td>
<td align="center">3.5%</td>
<td align="center">&#x2212;18.6%</td>
<td align="center">8.4%</td>
<td align="center">5.1%</td>
</tr>
<tr>
<td align="left">CalmingWater</td>
<td align="center">&#x2212;15.5%</td>
<td align="center">&#x2212;21.5%</td>
<td align="center">&#x2212;9.6%</td>
<td align="center">&#x2212;18.9%</td>
<td align="center">&#x2212;15.7%</td>
<td align="center">&#x2212;22.6%</td>
<td align="center">&#x2212;10.2%</td>
<td align="center">&#x2212;19.6%</td>
<td align="center">&#x2212;13.0%</td>
<td align="center">&#x2212;10.4%</td>
</tr>
<tr>
<td align="left">ToddlerFountain</td>
<td align="center">&#x2212;6.6%</td>
<td align="center">&#x2212;18.7%</td>
<td align="center">&#x2212;2.0%</td>
<td align="center">&#x2212;17.4%</td>
<td align="center">&#x2212;8.1%</td>
<td align="center">&#x2212;18.2%</td>
<td align="center">&#x2212;3.7%</td>
<td align="center">&#x2212;16.4%</td>
<td align="center">&#x2212;7.8%</td>
<td align="center">&#x2212;7.3%</td>
</tr>
<tr>
<td align="left">LampLeaves</td>
<td align="center">&#x2212;6.8%</td>
<td align="center">&#x2212;26.2%</td>
<td align="center">&#x2212;6.2%</td>
<td align="center">&#x2212;26.1%</td>
<td align="center">&#x2212;2.8%</td>
<td align="center">&#x2212;23.7%</td>
<td align="center">&#x2212;0.4%</td>
<td align="center">&#x2212;24.8%</td>
<td align="center">6.7%</td>
<td align="center">&#x2212;1.2%</td>
</tr>
<tr>
<td align="left">DaylightRoad</td>
<td align="center">&#x2212;3.8%</td>
<td align="center">&#x2212;38.0%</td>
<td align="center">&#x2212;12.4%</td>
<td align="center">&#x2212;40.3%</td>
<td align="center">&#x2212;0.9%</td>
<td align="center">&#x2212;37.6%</td>
<td align="center">&#x2212;10.4%</td>
<td align="center">&#x2212;42.4%</td>
<td align="center">0.9%</td>
<td align="center">&#x2212;9.8%</td>
</tr>
<tr>
<td align="left">RedRock</td>
<td align="center">&#x2212;3.5%</td>
<td align="center">&#x2212;32.5%</td>
<td align="center">&#x2212;9.0%</td>
<td align="center">&#x2212;37.9%</td>
<td align="center">0.8%</td>
<td align="center">&#x2212;31.5%</td>
<td align="center">&#x2212;6.4%</td>
<td align="center">&#x2212;37.7%</td>
<td align="center">16.1%</td>
<td align="center">&#x2212;8.0%</td>
</tr>
<tr>
<td align="left">RollerCoaster</td>
<td align="center">&#x2212;15.3%</td>
<td align="center">&#x2212;39.9%</td>
<td align="center">&#x2212;14.5%</td>
<td align="center">&#x2212;41.7%</td>
<td align="center">&#x2212;7.9%</td>
<td align="center">&#x2212;38.9%</td>
<td align="center">&#x2212;11.5%</td>
<td align="center">&#x2212;39.8%</td>
<td align="center">&#x2212;6.6%</td>
<td align="center">&#x2212;13.5%</td>
</tr>
<tr>
<td align="left">Average</td>
<td align="center">&#x2212;7.3%</td>
<td align="center">&#x2212;28.5%</td>
<td align="center">&#x2212;8.6%</td>
<td align="center">&#x2212;30.2%</td>
<td align="center">&#x2212;3.8%</td>
<td align="center">&#x2212;27.5%</td>
<td align="center">&#x2212;5.0%</td>
<td align="center">&#x2212;28.9%</td>
<td align="center">1.8%</td>
<td align="center">&#x2212;6.3%</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In order to further compare performance across different spatial resolutions and within the context of DO, the average rate-VMAF curves of the nine test sequences (HD resolution only) for HM and AV1 with and without DO are shown in <xref ref-type="fig" rid="F3">Figure 3</xref>. It can be observed that DO has achieved slightly higher overall coding gains for AV1 (BD-rate is -8.2%) on the tested content compared to HM (BD-rate is -6.1%). For both codecs, the savings become lower for higher bit rates (low QP and high quality). It should be noted that the DO approach employed was based on efficient up-sampling using simple spatial filters. <xref ref-type="bibr" rid="B2">Afonso et al. (2019)</xref>, <xref ref-type="bibr" rid="B63">Zhang et al. (2019)</xref> and others have reported significant improvement when advanced up-sampling approaches are applied, such as deep learning based super-resolution.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>The average rate-VMAF curves of the nine test sequences for <bold>(A)</bold> HM and <bold>(B)</bold> AV1 with and without applying DO.</p>
</caption>
<graphic xlink:href="frsip-02-874200-g003.tif"/>
</fig>
</sec>
<sec id="s5-2">
<title>5.2 Results Based on Subjective Quality Assessment</title>
<p>As outlined in the introduction of <xref ref-type="sec" rid="s5">Section 5</xref>, following the ITU-BT.500 protocols, for each resolution group, we performed outlier rejection on all participant scores. No participants were rejected. Next, following the ITU-T P.913 recommendation P.913 (2021), we estimated the subject bias for the different trials and compensated the raw scores based on that. <xref ref-type="fig" rid="F4">Figures 4</xref>, <xref ref-type="fig" rid="F5">5</xref> illustrate the 100-DMOS against the achieved bit rate for all tested codecs and sequences. Then, we performed one-way ANOVA analysis between pairs of the tested codecs to assess the significance of the differences. <xref ref-type="table" rid="T6">Table 6</xref> summarises this comparison.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>The DMOS-Rate curves for all sequences <bold>(A&#x2013;I)</bold> of Resolution Group A and all three codecs along with the standard error bars after compensating for subject bias, as described in Recommendation P.913 (2021).</p>
</caption>
<graphic xlink:href="frsip-02-874200-g004.tif"/>
</fig>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>The DMOS-Rate curves for all sequences <bold>(A&#x2013;I)</bold> of Resolution Group B and all three codecs along with the standard error bars after compensating for subject bias, as described in Recommendation P.913 (2021).</p>
</caption>
<graphic xlink:href="frsip-02-874200-g005.tif"/>
</fig>
<table-wrap id="T6" position="float">
<label>TABLE 6</label>
<caption>
<p>Aggregated significant difference of perceived quality among the tested codecs based on the ANOVA. The first ratio represents the number of sequences where statistically significant difference has been recorded. The ratio in the parentheses show which codec is significantly better (&#x2b;) or worse (&#x2212;) in the pairwise comparison of horizontal/vertical codec.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="left">Codecs</th>
<th colspan="3" align="center">Resolution group A</th>
<th colspan="3" align="center">Resolution group B</th>
<th colspan="2" align="center">Resolution group C</th>
</tr>
<tr>
<th align="center">AV1</th>
<th align="center">HM</th>
<th align="center">VTM</th>
<th align="center">AV1</th>
<th align="center">HM</th>
<th align="center">VTM</th>
<th align="center">AV1</th>
<th align="center">HM</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">AV1</td>
<td align="center">&#x2014;</td>
<td align="center">4/36, (2/-2)</td>
<td align="center">9/36, (0/-9)</td>
<td align="center">&#x2014;</td>
<td align="center">3/36, (0/-3)</td>
<td align="center">17/36, (0/-17)</td>
<td align="center">-</td>
<td align="center">5/45, (0/5)</td>
</tr>
<tr>
<td align="left">HM</td>
<td align="center">4/36, (2/-2)</td>
<td align="center">&#x2014;</td>
<td align="center">10/36, (0/-10)</td>
<td align="center">3/36, (3/0)</td>
<td align="center">&#x2014;</td>
<td align="center">15/36, (0/-15)</td>
<td align="center">5/45, (-5/0)</td>
<td align="center">&#x2014;</td>
</tr>
<tr>
<td align="left">VTM</td>
<td align="center">9/36, (9/0)</td>
<td align="center">10/36, (10/0)</td>
<td align="center">&#x2014;</td>
<td align="center">17/36, (17/0)</td>
<td align="center">15/36, (15/0)</td>
<td align="center">&#x2014;</td>
<td align="center">&#x2014;</td>
<td align="center">&#x2014;</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Results on Resolution Group A: A first impression from <xref ref-type="fig" rid="F4">Figure 4</xref> is that in most cases the VTM curve is on top of the other curves but also that the confidence intervals are overlapping in most cases. This is confirmed in <xref ref-type="table" rid="T6">Table 6</xref>, where the significance tests indicate only four cases which exhibit significant difference (<italic>p</italic> &#x3c; 0.05) between HM and AV1. In two of these, AV1 is significantly better than HM: in the cases of CalmingWater sequence at R1 and LampLeaves at R2. Both of these sequences consist of dynamic textures that are very challenging to compress. The opposite happens at R2 for AirAcrobatics and at R1 for Myanmar sequence; two sequences with a static background and slow moving objects. Performing significance tests between VTM and HM, a higher number of cases with significant differences were identified mostly at the lower bit rates. Particularly, for CatRobot at R1, R3; for CalmingWater at R1; for LampLeaves at R1-R2; for DaylightRoad at R1-R3; and for RedRock at R1-R2. A similar number of cases where VTM was significantly better than AV1 were identified: for AirAcrobatics at R3; for CatRobot at R1; for Myanmar at R1; for ToddlerFontain at R1; for DaylightRoad at R1-R3; and for RedRock at R1-R2. It is worth mentioning that from the curves in <xref ref-type="fig" rid="F4">Figure 4</xref> VTM achieves very good quality <inline-formula id="inf1">
<mml:math id="m2">
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mo>&#x2265;</mml:mo>
<mml:mn>80</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> even at the lowest bit rates for all sequences except for the ones with dynamic textures (CalmingWater, ToddlerFontain, and LampLeaves). That remains an area for improvement for the future codecs.</p>
<p>Results on Resolution Group B: We performed the same post-screening of the subjects and statistical analysis as in resolution group A. <xref ref-type="fig" rid="F5">Figure 5</xref> demonstrates the subjective quality against the bit rate after removing the subject bias and <xref ref-type="table" rid="T6">Table 6</xref> summarises the results of the one-way ANOVA comparison of the codecs. Generally, the results for the resolution group B align with those from resolution group A, that in most cases AV1 and HM result in equivalent video quality according to the viewers and that VTM in many cases prevails both codecs. Particularly, from <xref ref-type="table" rid="T6">Table 6</xref>, it can be observed that HM is significantly better than AV1 in only three cases: for Myanmar at R1-R2 and for RedRock at R2. VTM is significantly better than HM in 15 cases: for ToddlerFontain at R1-R4; for Myanmar at R1-R2; for CalmingWater at R1; for LampLeaves at R1-R2; for DaylightRoad at R1-R3; at RedRock at R1; and RollerCoaster at R3-R4. Similarly, VTM is significantly better than AV1 in 17 cases (almost have of the test sequences): for AirAcrobatics at R1; for CatRobot at R1-R3; for Myanmar at R1-R4; for LampLeaves at R1; for DaylightRoad at R1-R3; for RedRock at R1-R3; and RollerCoaster at R1-R2.</p>
<p>The reason that significant differences are noticed between AV1 and both HM and VTM in the case of the Myanmar sequence might be associated with the observation that, at lower bit rates, AV1 encoder demonstrates noticeable artifacts on regions of interest, namely in the center of the frame and on the heads of the walking monks in front of a still background with a static camera. An indicative example of these artifacts has been captured in <xref ref-type="fig" rid="F6">Figure 6</xref>. This, however, is probably a rare case as no similar cases have been reported so far in recent literature.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>Example of artifacts experienced on the Myanmar sequence that results in subjectively significantly different opinion scores between <bold>(A)</bold> HM and <bold>(B)</bold> AV1. The particular patches were captured at R1 from the resolution group B.</p>
</caption>
<graphic xlink:href="frsip-02-874200-g006.tif"/>
</fig>
<p>Results on Resolution Group C: The removal of the subject bias leads to slightly different results than the ones presented for this case in our previous work [see <xref ref-type="bibr" rid="B33">Katsenou et al. (2019b)</xref>], as illustrated in <xref ref-type="fig" rid="F7">Figure 7</xref>. After performing the significance test using one-way ANOVA between paired AV1 and HM sequences, the <italic>p</italic>-values of five rate points were indicated as significantly different. In all cases, HM is significantly better than AV1: at R1-R4 for Myanmar; and at R2 and R5 for LampLeaves. Myanmar encoded sequences suffer as mentioned earlier from unique artifacts and LampLeaves is a challenging dynamic texture. Although the findings from this resolution group are generally aligned with the observation from the other two resolution groups, for the LampLeaves sequence we notice a degraded performance of AV1 compared to resolution group B. This is attributed to the selected set of resolutions by the DO algorithm for AV1, which comprises lower resolution sequences than those selected for HM: at R2 540p instead of 720p and at R5 720p instead of 1080p (see <xref ref-type="table" rid="T2">Table 2</xref>).</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption>
<p>The DMOS-Rate curves for all sequences <bold>(A&#x2013;I)</bold> of Resolution Group C and all three codecs along with the standard error bars after compensating for subject bias, as described in Recommendation P.913 (2021).</p>
</caption>
<graphic xlink:href="frsip-02-874200-g007.tif"/>
</fig>
</sec>
<sec id="s5-3">
<title>5.3 Objective Quality Metric Performance Comparison</title>
<p>The correlation performance of six tested objective quality metrics for three resolution groups (in terms of SROCC values) is summarised in <xref ref-type="table" rid="T7">Table 7</xref>. It can be observed that VMAF outperforms the other metrics on all three test databases with the highest SROCC and LCC values, and lowest OR and RMSE. PSNR results in much lower performance, especially for the UHD resolution group. It is also noted that, for all test quality metrics, the SROCC values for three resolution groups are all below 0.9, which indicates that further enhancement is still needed to achieve more accurate prediction.</p>
<table-wrap id="T7" position="float">
<label>TABLE 7</label>
<caption>
<p>The correlation statistics of six popular quality metrics when evaluated on three subject datasets (UHD, HD and HD-DO).</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="left">Database metric</th>
<th colspan="4" align="center">UHD (108)</th>
<th colspan="4" align="center">HD (108)</th>
<th colspan="4" align="center">HD-DO (90)</th>
</tr>
<tr>
<th align="center">SROCC</th>
<th align="center">LCC</th>
<th align="center">OR</th>
<th align="center">RMSE</th>
<th align="center">SROCC</th>
<th align="center">LCC</th>
<th align="center">OR</th>
<th align="center">RMSE</th>
<th align="center">SROCC</th>
<th align="center">LCC</th>
<th align="center">OR</th>
<th align="center">RMSE</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">PSNR</td>
<td align="char" char=".">0.5517</td>
<td align="char" char=".">0.6278</td>
<td align="char" char=".">0.3056</td>
<td align="char" char=".">8.7540</td>
<td align="char" char=".">0.6097</td>
<td align="char" char=".">0.6268</td>
<td align="char" char=".">0.5556</td>
<td align="char" char=".">12.5870</td>
<td align="char" char=".">0.7462</td>
<td align="char" char=".">0.7439</td>
<td align="char" char=".">0.4222</td>
<td align="char" char=".">13.3191</td>
</tr>
<tr>
<td align="left">SSIM</td>
<td align="char" char=".">0.5911</td>
<td align="char" char=".">0.5853</td>
<td align="char" char=".">0.3148</td>
<td align="char" char=".">9.2195</td>
<td align="char" char=".">0.7194</td>
<td align="char" char=".">0.6757</td>
<td align="char" char=".">0.4907</td>
<td align="char" char=".">11.5968</td>
<td align="char" char=".">0.8026</td>
<td align="char" char=".">0.7836</td>
<td align="char" char=".">0.3778</td>
<td align="char" char=".">12.2184</td>
</tr>
<tr>
<td align="left">MSSSIM</td>
<td align="char" char=".">0.7426</td>
<td align="char" char=".">0.7436</td>
<td align="char" char=".">0.2130</td>
<td align="char" char=".">7.4102</td>
<td align="char" char=".">0.7534</td>
<td align="char" char=".">0.7241</td>
<td align="char" char=".">0.4537</td>
<td align="char" char=".">10.7594</td>
<td align="char" char=".">0.8321</td>
<td align="char" char=".">0.8228</td>
<td align="char" char=".">0.3556</td>
<td align="char" char=".">11.1398</td>
</tr>
<tr>
<td align="left">VIF</td>
<td align="char" char=".">0.7464</td>
<td align="char" char=".">0.7749</td>
<td align="char" char=".">0.1852</td>
<td align="char" char=".">6.9273</td>
<td align="char" char=".">0.7459</td>
<td align="char" char=".">0.7592</td>
<td align="char" char=".">0.3796</td>
<td align="char" char=".">10.0815</td>
<td align="char" char=".">0.8232</td>
<td align="char" char=".">0.8321</td>
<td align="char" char=".">0.3778</td>
<td align="char" char=".">10.8851</td>
</tr>
<tr>
<td align="left">VSNR</td>
<td align="char" char=".">0.5961</td>
<td align="char" char=".">0.6580</td>
<td align="char" char=".">0.2500</td>
<td align="char" char=".">8.4062</td>
<td align="char" char=".">0.5763</td>
<td align="char" char=".">0.6587</td>
<td align="char" char=".">0.3889</td>
<td align="char" char=".">12.0502</td>
<td align="char" char=".">0.6581</td>
<td align="char" char=".">0.7039</td>
<td align="char" char=".">0.4778</td>
<td align="char" char=".">14.0736</td>
</tr>
<tr>
<td align="left">ADM</td>
<td align="char" char=".">0.7532</td>
<td align="char" char=".">0.7639</td>
<td align="char" char=".">0.1759</td>
<td align="char" char=".">7.1573</td>
<td align="char" char=".">0.6858</td>
<td align="char" char=".">0.7290</td>
<td align="char" char=".">0.4352</td>
<td align="char" char=".">10.8612</td>
<td align="char" char=".">0.7928</td>
<td align="char" char=".">0.8143</td>
<td align="char" char=".">0.3556</td>
<td align="char" char=".">11.4781</td>
</tr>
<tr>
<td align="left">STVMAF</td>
<td align="char" char=".">0.7386</td>
<td align="char" char=".">0.7471</td>
<td align="char" char=".">0.2130</td>
<td align="char" char=".">7.3607</td>
<td align="char" char=".">0.7727</td>
<td align="char" char=".">0.7743</td>
<td align="char" char=".">0.4167</td>
<td align="char" char=".">9.8542</td>
<td align="char" char=".">0.3147</td>
<td align="char" char=".">0.4884</td>
<td align="char" char=".">0.5556</td>
<td align="char" char=".">17.5840</td>
</tr>
<tr>
<td align="left">VMAF</td>
<td align="char" char=".">
<bold>0.8463</bold>
</td>
<td align="char" char=".">
<bold>0.8375</bold>
</td>
<td align="char" char=".">
<bold>0.1574</bold>
</td>
<td align="char" char=".">
<bold>5.9972</bold>
</td>
<td align="char" char=".">
<bold>0.8723</bold>
</td>
<td align="char" char=".">
<bold>0.8476</bold>
</td>
<td align="char" char=".">
<bold>0.2870</bold>
</td>
<td align="char" char=".">
<bold>7.9969</bold>
</td>
<td align="char" char=".">
<bold>0.8783</bold>
</td>
<td align="char" char=".">
<bold>0.8840</bold>
</td>
<td align="char" char=".">
<bold>0.2556</bold>
</td>
<td align="char" char=".">
<bold>9.1395</bold>
</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>The best goodness-of-fit metric values are in bold.</p>
</fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s5-4">
<title>5.4 Computational Complexity Analysis</title>
<p>The average complexity figures for encoding UHD and HD content are summarised in <xref ref-type="table" rid="T8">Table 8</xref>, where the HM encoder has been used for benchmarking. The average complexity is computed as the average ratio of the execution time of the tested codec for all rate points over the benchmark. As can be seen, for the tested codec versions, AV1 has a higher complexity compared to VTM<xref ref-type="fn" rid="fn4">
<sup>4</sup>
</xref>. Interestingly these figures are higher for the HD than the UHD resolution. The relationship between the relative complexity and encoding performance (in terms of average coding gains for PSNR and VMAF) is also shown in <xref ref-type="fig" rid="F8">Figure 8</xref>.</p>
<table-wrap id="T8" position="float">
<label>TABLE 8</label>
<caption>
<p>Computational complexity comparison.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Resolution group/Codecs</th>
<th align="center">HM</th>
<th align="center">AV1</th>
<th align="center">VTM</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Resolution Group A (UHD)</td>
<td align="char" char=".">1</td>
<td align="char" char=".">9.37&#xd7;</td>
<td align="char" char=".">7.04&#xd7;</td>
</tr>
<tr>
<td align="left">Resolution Group B (HD)</td>
<td align="char" char=".">1</td>
<td align="char" char=".">14.29&#xd7;</td>
<td align="char" char=".">8.84&#xd7;</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="F8" position="float">
<label>FIGURE 8</label>
<caption>
<p>The relationship between the relative codec complexity (benchmarked on HM) and encoding performance (in terms of average coding gains) for different resolution groups and quality metrics: <bold>(A)</bold> UHD/PSNR, <bold>(B)</bold> UHD/VMAF, <bold>(C)</bold> HD/PSNR, and <bold>(D)</bold> HD/VMAF.</p>
</caption>
<graphic xlink:href="frsip-02-874200-g008.tif"/>
</fig>
</sec>
</sec>
<sec id="s6">
<title>6 Conclusion</title>
<p>This paper presents a video database of nine representative UHD source sequences and 306 compressed versions of these along with their associated objective and subjective quality assessment results. The testing configurations include spatial resampling (from 540p to 1080p) and encoding by three major contemporary video codecs, HEVC HM, AV1, and VVC VTM at pre-defined target bit rates. For one of the three test cases, the convex hull rate-distortion optimisation has been employed to compare HEVC HM and AV1 across different resolutions (from 540p to 1080p) and across a wider bit rate range. This test case is particularly useful in the bitrate ladder construction for adaptive video streaming. All the original and compressed video sequences and their corresponding quality scores are available online for public testing [see <xref ref-type="bibr" rid="B34">Katsenou A. et al. (2021)</xref>] with the aim to facilitate research on video compression and video quality across video codecs. To the best of our knowledge, this is the first public dataset that contains encodings from VVC, AV1, and HEVC.</p>
<p>As research on video technologies evolves to data-greedy algorithms, in the near future, we intend to extend this dataset by incorporating more sequences, at higher spatial resolutions, varying framerates, and bitdepths, and by running a large-scale subjective evaluation through crowdsourcing. Furthermore, as the newest VTM releases have shown significant improvement over past versions, we intend to update the VTM version with the latest release. Finally, we will expand the set of codecs by including optimized versions of existing standards, such as SVT-AV1 from <xref ref-type="bibr" rid="B44">Norkin et al. (2020)</xref>, VVenC/VVdeC from <xref ref-type="bibr" rid="B7">Brandenburg et al. (2020)</xref>, etc.</p>
</sec>
</body>
<back>
<sec id="s7">
<title>Data Availability Statement</title>
<p>The datasets presented in this study can be found in online repositories listed in this paper.</p>
</sec>
<sec id="s38">
<title>Ethics statement</title>
<p>Ethical review and approval was not required for the study of human participants in accordance with the local legislation and institutional requirements. Written informed consent from the participants was provided following the ITU-R BT.500 protocol.</p>
</sec>
<sec id="s8">
<title>Author Contributions</title>
<p>AK: Conceptualization, Investigation, Methodology, Subjective Tests, Data Curation, Software, Visualization, Validation, Writing&#x2014;Reviewing and Editing. FZ: Conceptualization, Investigation, Methodology, Data Curation, Software, Visualization, Validation, Writing&#x2014;Reviewing and Editing. MA: Investigation, Methodology, Conceptualization, Software. GD: Subjective Tests, Data Curation, Software, Validation. DB: Supervision, Funding acquisition, Conceptualization, Writing&#x2014;Reviewing and Editing.</p>
</sec>
<sec id="s9">
<title>Funding</title>
<p>The authors acknowledge funding from the UK Engineering and Physical Sciences Research Council (EPSRC, project No. EP/M000885/1) and the Leverhulme Early Career Fellowship awarded to A. Katsenou (ECF-2017-413). This study received donation from NVIDIA Corporation in the form of GPUs.</p>
</sec>
<sec sec-type="COI-statement" id="s10">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s11">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<fn-group>
<fn id="fn1">
<label>1</label>
<p>Here only convex hull rate-distortion optimisation within each shot is employed.</p>
</fn>
<fn id="fn2">
<label>2</label>
<p>Only for a subset, for the categories Gaming, Sports, and Vlog videos.</p>
</fn>
<fn id="fn3">
<label>3</label>
<p>We have not compared VVC with other codecs using the DO approach. This is mainly due to the high computation complexity of VVC and the limited computational resources that we have. Preliminary results for Group A and B have already shown the significant improvement of VVC over the other two.</p>
</fn>
<fn id="fn4">
<label>4</label>
<p>It is noted that the complexity for both AV1 and VTM in more recent versions have been significantly reduced.</p>
</fn>
</fn-group>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="book">
<comment>[Dataset]</comment> <collab>Advanced Computing Research Centre, University of Bristol</collab> (<year>2021</year>). <source>BlueCystal Phase 3</source>. <publisher-loc>Bristol, England</publisher-loc>: <publisher-name>Advanced Computing Research Centre, University of Bristol</publisher-name>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://www.acrc.bris.ac.uk/acrc/phase3.htm">https://www.acrc.bris.ac.uk/acrc/phase3.htm</ext-link>
</comment>. </citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Afonso</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Bull</surname>
<given-names>D. R.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Video Compression Based on Spatio-Temporal Resolution Adaptation</article-title>. <source>IEEE Trans. Circuits Syst. Video Technol.</source> <volume>29</volume>, <fpage>275</fpage>&#x2013;<lpage>280</lpage>. <pub-id pub-id-type="doi">10.1109/tcsvt.2018.2878952</pub-id> </citation>
</ref>
<ref id="B3">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Akyazi</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Ebrahimi</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Comparison of Compression Efficiency between HEVC/H. 265, VP9 and AV1 Based on Subjective Quality Assessments</article-title>,&#x201d; in <conf-name>Tenth International Conference on Quality of Multimedia Experience (QoMEX)</conf-name> (<publisher-loc>Cagliari, Italy</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x2013;<lpage>6</lpage>. <pub-id pub-id-type="doi">10.1109/qomex.2018.8463294</pub-id> </citation>
</ref>
<ref id="B4">
<citation citation-type="book">
<comment>[Dataset]</comment> <collab>AOM</collab> (<year>2019</year>). <source>AOMedia Video 1 (AV1)</source>. <publisher-loc>Briarcliff Manor, NY, USA</publisher-loc>: <publisher-name>AOM</publisher-name>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://github.com/AOMediaCodec">https://github.com/AOMediaCodec</ext-link>
</comment>. </citation>
</ref>
<ref id="B5">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Bj&#xf8;ntegaard</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2001</year>). &#x201c;<article-title>Calculation of Average PSNR Differences between RD-Curves</article-title>,&#x201d; in <conf-name>VCEG-M33: 13th VCEG Meeting</conf-name> (<publisher-loc>Austin, Texas, USA</publisher-loc>: <publisher-name>ITU-T</publisher-name>). </citation>
</ref>
<ref id="B6">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Bossen</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Boyce</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Seregin</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>S&#xfc;hring</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>JVET Common Test Conditions and Software Reference Configurations for SDR Video</article-title>,&#x201d; in <conf-name>The JVET meeting (ITU-T and ISO/IEC), JVET-M1001</conf-name>. </citation>
</ref>
<ref id="B7">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Brandenburg</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Wieckowski</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Hinz</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Henkel</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>George</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Zupancic</surname>
<given-names>I.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). &#x201c;<article-title>Towards Fast and Efficient Vvc Encoding</article-title>,&#x201d; in <conf-name>2020 IEEE 22nd International Workshop on Multimedia Signal Processing (MMSP)</conf-name>, <fpage>1</fpage>&#x2013;<lpage>6</lpage>. <pub-id pub-id-type="doi">10.1109/mmsp48831.2020.9287093</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Bross</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Y.-K.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Versatile Video Coding (Draft 7)</article-title>,&#x201d; in <conf-name>The JVET meeting (ITU-T and ISO/IEC), JVET-P2001</conf-name>. </citation>
</ref>
<ref id="B9">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Bull</surname>
<given-names>D. R.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2021</year>). <source>Intelligent Image and Video Compression: Communicating Pictures</source>. <publisher-loc>Cambridge, MA, USA</publisher-loc>: <publisher-name>Academic Press</publisher-name>. </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chandler</surname>
<given-names>D. M.</given-names>
</name>
<name>
<surname>Hemami</surname>
<given-names>S. S.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>VSNR: A Wavelet-Based Visual Signal-To-Noise Ratio for Natural Images</article-title>. <source>IEEE Trans. Image Process.</source> <volume>16</volume>, <fpage>2284</fpage>&#x2013;<lpage>2298</lpage>. <pub-id pub-id-type="doi">10.1109/tip.2007.901820</pub-id> </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Mukherjee</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Han</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Grange</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Parker</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>An Overview of Coding Tools in AV1: the First Video Codec from the Alliance for Open Media</article-title>. <source>APSIPA Trans. Signal Inf. Process.</source> <volume>9</volume>. <pub-id pub-id-type="doi">10.1017/atsip.2020.2</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="book">
<collab>CISCO</collab> (<year>2018</year>). <source>CISCO Visual Networking index: Forecast and Methodology</source>. <publisher-loc>Indianapolis, Indiana</publisher-loc>: <publisher-name>CISCO</publisher-name>, <fpage>2017</fpage>&#x2013;<lpage>2022</lpage>. </citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cortes</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Vapnik</surname>
<given-names>V.</given-names>
</name>
</person-group> (<year>1995</year>). <article-title>Support-vector Networks</article-title>. <source>Mach Learn.</source> <volume>20</volume>, <fpage>273</fpage>&#x2013;<lpage>297</lpage>. <pub-id pub-id-type="doi">10.1007/bf00994018</pub-id> </citation>
</ref>
<ref id="B14">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Daede</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Norkin</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Brailovskiy</surname>
<given-names>I.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Video Codec Testing and Quality Measurement</article-title>,&#x201d; in <source>Internet-Draft</source> (<publisher-name>Network Working Group</publisher-name>). </citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Damen</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Doughty</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Farinella</surname>
<given-names>G. M.</given-names>
</name>
<name>
<surname>Fidler</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Furnari</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Kazakos</surname>
<given-names>E.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>The Epic-Kitchens Dataset: Collection, Challenges and Baselines</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell.</source> <volume>43</volume>, <fpage>4125</fpage>&#x2013;<lpage>4141</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2020.2991965</pub-id> </citation>
</ref>
<ref id="B16">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Dias</surname>
<given-names>A. S.</given-names>
</name>
<name>
<surname>Blasi</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Rivera</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Izquierdo</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Mrak</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>An Overview of Recent Video Coding Developments in MPEG and AOMedia</article-title>,&#x201d; in <source>International Broadcasting Convention (IBC)</source>. </citation>
</ref>
<ref id="B17">
<citation citation-type="web">
<comment>[Dataset]</comment> <person-group person-group-type="author">
<name>
<surname>Dimitrov</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Katsenou</surname>
<given-names>A. V.</given-names>
</name>
<name>
<surname>Bull</surname>
<given-names>D. R.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>BVI-SVQA: Subjective Video Quality Assessment Software</article-title>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://github.com/goceee/SVQA">https://github.com/goceee/SVQA</ext-link>
</comment>. </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Duchon</surname>
<given-names>C. E.</given-names>
</name>
</person-group> (<year>1979</year>). <article-title>Lanczos Filtering in One and Two Dimensions</article-title>. <source>J. Appl. Meteorol.</source> <volume>18</volume>, <fpage>1016</fpage>&#x2013;<lpage>1022</lpage>. <pub-id pub-id-type="doi">10.1175/1520-0450(1979)018&#x3c;1016:lfioat&#x3e;2.0.co;2</pub-id> </citation>
</ref>
<ref id="B19">
<citation citation-type="web">
<comment>[Dataset]</comment> <collab>Google</collab> (<year>2017</year>). <article-title>VP9 Video Codec</article-title>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://www.webmproject.org/vp9/">https://www.webmproject.org/vp9/</ext-link>
</comment>. </citation>
</ref>
<ref id="B20">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Grois</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Nguyen</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Marpe</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>Coding Efficiency Comparison of AV1, VP9, H.265/MPEG-HEVC, and H. 264/MPEG-AVC Encoders</article-title>,&#x201d; in <conf-name>Picture Coding Symposium (PCS)</conf-name> (<publisher-loc>Nuremberg, Germany</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x2013;<lpage>5</lpage>. </citation>
</ref>
<ref id="B21">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Guo</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>De Cock</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Aaron</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Compression Performance Comparison of X264, X265, Libvpx and Aomenc for On-Demand Adaptive Streaming Applications</article-title>,&#x201d; in <conf-name>Picture Coding Symposium (PCS)</conf-name> (<publisher-loc>San Francisco, CA, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>26</fpage>&#x2013;<lpage>30</lpage>. <pub-id pub-id-type="doi">10.1109/pcs.2018.8456302</pub-id> </citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hanhart</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Ebrahimi</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Calculation of Average Coding Efficiency Based on Subjective Quality Scores</article-title>. <source>J. Vis. Commun. Image Representation</source> <volume>25</volume>, <fpage>555</fpage>&#x2013;<lpage>564</lpage>. <pub-id pub-id-type="doi">10.1016/j.jvcir.2013.11.008</pub-id> </citation>
</ref>
<ref id="B23">
<citation citation-type="web">
<comment>[Dataset]</comment> <collab>Harmonic</collab> (<year>2019</year>). <article-title>Harmonic Free 4K Demo Footage</article-title>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://www.harmonicinc.com/news-insights/blog/4k-in-context/">https://www.harmonicinc.com/news-insights/blog/4k-in-context/</ext-link>
</comment>. </citation>
</ref>
<ref id="B24">
<citation citation-type="book">
<comment>[Dataset]</comment> <person-group person-group-type="author">
<name>
<surname>Hosu</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Hahn</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Jenadeleh</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Lin</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Men</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Szir&#xe1;nyi</surname>
<given-names>T.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). <source>KonVid-1K: The Konstanz Natural Video Database</source>. </citation>
</ref>
<ref id="B25">
<citation citation-type="book">
<comment>[Dataset]</comment> <collab>ITU-T Rec H.264</collab> (<year>2005</year>). <source>Advanced Video Coding for Generic Audiovisual Services</source>. </citation>
</ref>
<ref id="B26">
<citation citation-type="book">
<comment>[Dataset]</comment> <collab>ITU-T Rec H.265</collab> (<year>2015</year>). <source>High Efficiency Video Coding</source>. </citation>
</ref>
<ref id="B27">
<citation citation-type="book">
<comment>[Dataset]</comment> <collab>ITU-T Rec. H.120</collab> (<year>1993</year>). <source>Codecs for Videoconferencing Using Primary Digital Group Transmission</source>. </citation>
</ref>
<ref id="B28">
<citation citation-type="book">
<comment>[Dataset]</comment> <collab>ITU-T Rec. H.262</collab> (<year>2012</year>). <source>Information Technology - Generic Coding of Moving Pictures and Associated Audio Information: Video</source>. </citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Katsavounidis</surname>
<given-names>I.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Dynamic Optimizer &#x2013; a Perceptual Video Encoding Optimization Framework</article-title>. <source>Netflix Tech. Blog</source>. </citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Katsavounidis</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Video Codec Comparison Using the Dynamic Optimizer Framework</article-title>. <source>Appl. Digital Image Process. XLI (International Soc. Opt. Photonics)</source> <volume>10752</volume>, <fpage>107520Q</fpage>. <pub-id pub-id-type="doi">10.1117/12.2322118</pub-id> </citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Katsenou</surname>
<given-names>A. V.</given-names>
</name>
<name>
<surname>Dimitrov</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Ma</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Bull</surname>
<given-names>D. R.</given-names>
</name>
</person-group> (<year>2021b</year>). <article-title>BVI-SynTex: A Synthetic Video Texture Dataset for Video Compression and Quality Assessment</article-title>. <source>IEEE Trans. Multimedia</source> <volume>23</volume>, <fpage>26</fpage>&#x2013;<lpage>38</lpage>. <pub-id pub-id-type="doi">10.1109/tmm.2020.2976591</pub-id> </citation>
</ref>
<ref id="B32">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Katsenou</surname>
<given-names>A. V.</given-names>
</name>
<name>
<surname>Sole</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Bull</surname>
<given-names>D. R.</given-names>
</name>
</person-group> (<year>2019a</year>). &#x201c;<article-title>Content-gnostic Bitrate Ladder Prediction for Adaptive Video Streaming</article-title>,&#x201d; in <conf-name>Picture Coding Symposium (PCS)</conf-name>. <pub-id pub-id-type="doi">10.1109/pcs48520.2019.8954529</pub-id> </citation>
</ref>
<ref id="B33">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Katsenou</surname>
<given-names>A. V.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Afonso</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Bull</surname>
<given-names>D. R.</given-names>
</name>
</person-group> (<year>2019b</year>). &#x201c;<article-title>A Subjective Comparison of AV1 and HEVC for Adaptive Video Streaming</article-title>,&#x201d; in <conf-name>Proc. IEEE Int Conf. on Image Processing</conf-name>. <pub-id pub-id-type="doi">10.1109/icip.2019.8803523</pub-id> </citation>
</ref>
<ref id="B34">
<citation citation-type="web">
<comment>[Dataset]</comment> <person-group person-group-type="author">
<name>
<surname>Katsenou</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Afonso</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Dimitrov</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Bull</surname>
<given-names>D. R.</given-names>
</name>
</person-group> (<year>2021a</year>). <article-title>BVI-CC: BVI - Video Codec Comparison</article-title>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://angkats.github.io/Video-Codec-Comparison/">https://angkats.github.io/Video-Codec-Comparison/</ext-link>
</comment>. </citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lee</surname>
<given-names>J.-S.</given-names>
</name>
<name>
<surname>De Simone</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Ebrahimi</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Subjective Quality Evaluation via Paired Comparison: Application to Scalable Video Coding</article-title>. <source>IEEE Trans. Multimedia</source> <volume>13</volume>, <fpage>882</fpage>&#x2013;<lpage>893</lpage>. <pub-id pub-id-type="doi">10.1109/tmm.2011.2157333</pub-id> </citation>
</ref>
<ref id="B36">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Krasula</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Baveye</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Le Callet</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Accann: A New Subjective Assessment Methodology for Measuring Acceptability and Annoyance of Quality of Experience</article-title>. <source>IEEE Trans. Multimedia</source> <volume>21</volume>, <fpage>2589</fpage>&#x2013;<lpage>2602</lpage>. <pub-id pub-id-type="doi">10.1109/tmm.2019.2903722</pub-id> </citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Ma</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Ngan</surname>
<given-names>K. N.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Image Quality Assessment by Separately Evaluating Detail Losses and Additive Impairments</article-title>. <source>IEEE Trans. Multimedia</source> <volume>13</volume>, <fpage>935</fpage>&#x2013;<lpage>949</lpage>. <pub-id pub-id-type="doi">10.1109/tmm.2011.2152382</pub-id> </citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Aaron</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Katsavounidis</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Moorthy</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Manohara</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Toward a Practical Perceptual Video Quality Metric</article-title>. <source>Netflix Tech. Blog</source>. </citation>
</ref>
<ref id="B39">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lin</surname>
<given-names>J. Y.</given-names>
</name>
<name>
<surname>Song</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>C.-H.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Kuo</surname>
<given-names>C.-C. J.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>MCL-V: A Streaming Video Quality Assessment Database</article-title>. <source>J. Vis. Commun. Image Representation</source> <volume>30</volume>, <fpage>1</fpage>&#x2013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.1016/j.jvcir.2015.02.012</pub-id> </citation>
</ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ma</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Bull</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>BVI-DVC: A Training Database for Deep Video Compression</article-title>. <source>IEEE Trans. Multimedia</source>. <pub-id pub-id-type="doi">10.1109/tmm.2021.3108943</pub-id> </citation>
</ref>
<ref id="B41">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mercer Moss</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Baddeley</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Bull</surname>
<given-names>D. R.</given-names>
</name>
</person-group> (<year>2016a</year>). <article-title>On the Optimal Presentation Duration for Subjective Video Quality Assessment</article-title>. <source>IEEE Trans. Circuits Syst. Video Technol.</source> <volume>26</volume>, <fpage>1977</fpage>&#x2013;<lpage>1987</lpage>. <pub-id pub-id-type="doi">10.1109/TCSVT.2015.2461971</pub-id> </citation>
</ref>
<ref id="B42">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mercer Moss</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Yeh</surname>
<given-names>C.-T.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Baddeley</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Bull</surname>
<given-names>D. R.</given-names>
</name>
</person-group> (<year>2016b</year>). <article-title>Support for Reduced Presentation Durations in Subjective Video Quality Assessment</article-title>. <source>Signal. Processing: Image Commun.</source> <volume>48</volume>, <fpage>38</fpage>&#x2013;<lpage>49</lpage>. <pub-id pub-id-type="doi">10.1016/j.image.2016.08.005</pub-id> </citation>
</ref>
<ref id="B43">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nguyen</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Marpe</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Compression Efficiency Analysis of AV1, VVC, and HEVC for Random Access Applications</article-title>. <source>APSIPA Trans. Signal Inf. Process.</source> <volume>10</volume>, <fpage>e11</fpage>. <pub-id pub-id-type="doi">10.1017/atsip.2021.10</pub-id> </citation>
</ref>
<ref id="B44">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Norkin</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Sole</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Afonso</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Swanson</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Opalach</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Moorthy</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>SVT-AV1: Open-Source AV1 Encoder and Decoder</article-title>. <source>Netflix Technol. Blog</source>. </citation>
</ref>
<ref id="B45">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Ohm</surname>
<given-names>J.-R.</given-names>
</name>
</person-group> (<year>2015</year>). <source>Multimedia Signal Coding and Transmission</source>. <publisher-loc>Berlin, Germany</publisher-loc>: <publisher-name>Springer</publisher-name>. </citation>
</ref>
<ref id="B46">
<citation citation-type="book">
<comment>[Dataset]</comment> <collab>P.910, I.-R.</collab> (<year>1999</year>). <source>Subjective Video Quality Assessment Methods for Multimedia Applications</source>. </citation>
</ref>
<ref id="B47">
<citation citation-type="book">
<comment>[Dataset]</comment> <collab>P.913, I.-R.</collab> (<year>2021</year>). <source>Methods for the Subjective Assessment of Video Quality, Audio Quality and Audiovisual Quality of Internet Video and Distribution Quality Television in Any Environment Applications</source>. </citation>
</ref>
<ref id="B48">
<citation citation-type="web">
<comment>[Dataset]</comment> <person-group person-group-type="author">
<name>
<surname>Papadopoulos</surname>
<given-names>M. A.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Agrafiotis</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Bull</surname>
<given-names>D. R.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>BVI-texture: BVI Video Texture Database</article-title>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="http://data.bris.ac.uk/data/dataset/1if54ya4xpph81fbo1gkpk5kk4">http://data.bris.ac.uk/data/dataset/1if54ya4xpph81fbo1gkpk5kk4</ext-link>
</comment>. </citation>
</ref>
<ref id="B49">
<citation citation-type="book">
<collab>Recommendation ITU-R BT.500-12</collab> (<year>2012</year>). <source>Methodology for the Subjective Assessment of the Quality of Television Pictures</source>. </citation>
</ref>
<ref id="B50">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Seshadrinathan</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Soundararajan</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Bovik</surname>
<given-names>A. C.</given-names>
</name>
<name>
<surname>Cormack</surname>
<given-names>L. K.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Study of Subjective and Objective Quality Assessment of Video</article-title>. <source>IEEE Trans. Image Process.</source> <volume>19</volume>, <fpage>1427</fpage>&#x2013;<lpage>1441</lpage>. <pub-id pub-id-type="doi">10.1109/tip.2010.2042111</pub-id> </citation>
</ref>
<ref id="B51">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Sharman</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Suehring</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Common Test Conditions for Hm Video Coding Experiments</article-title>,&#x201d; in <conf-name>The JCT-VC meeting (ITU-T, ISO/IEC), JCTVC-AF1100</conf-name>. </citation>
</ref>
<ref id="B52">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sheikh</surname>
<given-names>H. R.</given-names>
</name>
<name>
<surname>Bovik</surname>
<given-names>A. C.</given-names>
</name>
<name>
<surname>de Veciana</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>An Information Fidelity Criterion for Image Quality Assessment Using Natural Scene Statistics</article-title>. <source>IEEE Trans. Image Process.</source> <volume>14</volume>, <fpage>2117</fpage>&#x2013;<lpage>2128</lpage>. <pub-id pub-id-type="doi">10.1109/tip.2005.859389</pub-id> </citation>
</ref>
<ref id="B53">
<citation citation-type="web">
<comment>[Dataset]</comment> <person-group person-group-type="author">
<name>
<surname>Shostak</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Fedorov</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Zimiche</surname>
<given-names>V.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>AV2 Video Codec &#x2013; Early Performance Evaluation of the Research</article-title>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://ottverse.com/av2-video-codec-evaluation/">https://ottverse.com/av2-video-codec-evaluation/</ext-link>
</comment>. </citation>
</ref>
<ref id="B54">
<citation citation-type="book">
<collab>Video Quality Experts Group</collab> (<year>2000</year>). <source>Final Report from the Video Quality Experts Group on the Validation of Objective Quailty Metrics for Video Quality Assessment</source>. </citation>
</ref>
<ref id="B55">
<citation citation-type="book">
<collab>Video Quality Experts Group</collab> (<year>2010</year>). <source>Report on the Validation of Video Quality Models for High Definition Video Content</source>. </citation>
</ref>
<ref id="B56">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Inguva</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Adsumilli</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Youtube Ugc Dataset for Video Compression Research</article-title>,&#x201d; in <conf-name>2019 IEEE 21st International Workshop on Multimedia Signal Processing (MMSP)</conf-name>, <fpage>1</fpage>&#x2013;<lpage>5</lpage>. <pub-id pub-id-type="doi">10.1109/mmsp.2019.8901772</pub-id> </citation>
</ref>
<ref id="B57">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Bovik</surname>
<given-names>A. C.</given-names>
</name>
<name>
<surname>Sheikh</surname>
<given-names>H. R.</given-names>
</name>
<name>
<surname>Simoncelli</surname>
<given-names>E. P.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>Image Quality Assessment: from Error Visibility to Structural Similarity</article-title>. <source>IEEE Trans. Image Process.</source> <volume>13</volume>, <fpage>600</fpage>&#x2013;<lpage>612</lpage>. <pub-id pub-id-type="doi">10.1109/tip.2003.819861</pub-id> </citation>
</ref>
<ref id="B58">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Simoncelli</surname>
<given-names>E. P.</given-names>
</name>
<name>
<surname>Bovik</surname>
<given-names>A. C.</given-names>
</name>
</person-group> (<year>2003</year>). &#x201c;<article-title>Multi-scale Structural Similarity for Image Quality Assessment</article-title>,&#x201d; in <conf-name>Proc. Asilomar Conference on Signals, Systems and Computers</conf-name> (<publisher-loc>Pacific Grove, CA, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1398</fpage>. <comment>Vol. 2</comment>. </citation>
</ref>
<ref id="B59">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Wien</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2015</year>). <source>High Efficiency Video Coding</source>. <publisher-loc>Berlin, Germany</publisher-loc>: <publisher-name>Springer</publisher-name>. </citation>
</ref>
<ref id="B60">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Winkler</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Analysis of Public Image and Video Database for Quality Assessment</article-title>. <source>IEEE J. Selected Top. Signal Process.</source> <volume>6</volume>, <fpage>1</fpage>&#x2013;<lpage>10</lpage>. <pub-id pub-id-type="doi">10.1109/jstsp.2012.2215007</pub-id> </citation>
</ref>
<ref id="B61">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xing</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Y.-G.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>He</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Yuan</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>DVL2021: An Ultra High Definition Video Dataset for Perceptual Quality Study</article-title>. <source>J. Vis. Commun. Image Representation</source> <volume>82</volume>, <fpage>103374</fpage>. <pub-id pub-id-type="doi">10.1016/j.jvcir.2021.103374</pub-id> </citation>
</ref>
<ref id="B62">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Zabrovskiy</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Feldmann</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Timmerer</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>A Practical Evaluation of Video Codecs for Large-Scale HTTP Adaptive Streaming Services</article-title>,&#x201d; in <conf-name>25th IEEE International Conference on Image Processing (ICIP)</conf-name> (<publisher-loc>Athens, Greece</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>998</fpage>&#x2013;<lpage>1002</lpage>. <pub-id pub-id-type="doi">10.1109/icip.2018.8451017</pub-id> </citation>
</ref>
<ref id="B63">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Afonso</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Bull</surname>
<given-names>D. R.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Vistra2: Video Coding Using Spatial Resolution and Effective Bit Depth Adaptation</article-title>. <source>arXiv preprint arXiv:1911.02833</source>. </citation>
</ref>
<ref id="B64">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Bull</surname>
<given-names>D. R.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>A Parametric Framework for Video Compression Using Region-Based Texture Models</article-title>. <source>IEEE J. Sel. Top. Signal. Process.</source> <volume>5</volume>, <fpage>1378</fpage>&#x2013;<lpage>1392</lpage>. <pub-id pub-id-type="doi">10.1109/jstsp.2011.2165201</pub-id> </citation>
</ref>
<ref id="B65">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Moss</surname>
<given-names>F. M.</given-names>
</name>
<name>
<surname>Baddeley</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Bull</surname>
<given-names>D. R.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>BVI-HD: A Video Quality Database for HEVC Compressed and Texture Synthesized Content</article-title>. <source>IEEE Trans. Multimedia</source> <volume>20</volume>, <fpage>2620</fpage>&#x2013;<lpage>2630</lpage>. <pub-id pub-id-type="doi">10.1109/tmm.2018.2817070</pub-id> </citation>
</ref>
</ref-list>
</back>
</article>