<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Hacking VMAF with Video Color and Contrast Distortion</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>A. Zvezdakova</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Dubna State University</institution>
          ,
          <addr-line>Dubna</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Lomonosov Moscow State University</institution>
          ,
          <addr-line>Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Video quality measurement takes an important role in many applications. Full-reference quality metrics which are usually used in video codecs comparisons are expected to reflect any changes in videos. In this article, we consider different color corrections of compressed videos which increase the values of full-reference metric VMAF and almost don't decrease other widely-used metric SSIM. The proposed video contrast enhancement approach shows the metric in-applicability in some cases for video codecs comparisons, as it may be used for cheating in the comparisons via tuning to improve this metric values.</p>
      </abstract>
      <kwd-group>
        <kwd>video quality</kwd>
        <kwd>quality measuring</kwd>
        <kwd>video-codec comparison</kwd>
        <kwd>quality tuning</kwd>
        <kwd>reference metrics</kwd>
        <kwd>color correction</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        At the moment, video content takes a significant
part of worldwide network traffic and its share is
expected to grow up to 71% by 2021 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Therefore, the
quality of encoded videos is becoming increasingly
important, which leads to growing of an interest in the
area of new video quality assessment methods
development. As new video codec standards appear, the
existing standards are being improved. In order to
choose one or another video encodingsolution, it is
necessary to have appropriate tools for video quality
assessment. Since the best method of video quality
assessment is a subjective evaluation, which is quite
expensive in terms of time and cost of its
implementation, all other objective methods are improving in an
attempt to approach the ground truth-solution
(subjective evaluation).
      </p>
      <p>
        Methods for evaluating encoded videos quality
can be divided into 3 categories [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]: full-reference,
reduced-reference and no-reference. Full-reference
metrics are the most common, as their results are
easily interpreted — usually as an assessment of the
degree of distortions in thevideo and their visibility to
the observer. The only drawback of this approach
compared to the others is the need to have
theoriginal video for comparison with the encoded, which is
often not available.
      </p>
      <p>
        One of the widely-used full-reference metrics which is
gaining popularity in the area of video quality
assessment is Video Multimethod Fusion Approach
(VMAF)[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], announced by Netflix. It is an
opensource learning-based solution. Its main idea is to
combine multiple elementary video quality features,
such as Visual Information Fidelity (VIF)[12], Detail
Loss Metric (DLM)[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] and temporal information (TI) –
the difference between two neighboring frames, and then
to train support vector machine (SVM) regres-sion on
subjective data. The resulting regressor is used for
estimating per-frame quality scores on new videos. The
scheme [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] of this metric isshown in Fig. 1.
      </p>
      <p>Fig. 1. The scheme of VMAF algorithm.</p>
      <p>
        Despite increasing attention to this metric, many
video quality analysis projects, such as Moscow State
University’s (MSU) Annual Video Codec Comparison
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], still use other common metrics developed many
years ago, such as structural similarity (SSIM) and
even peak signal-to-noise ratio (PSNR), which are
based only on difference characteristics of two
images. At the same time, many readers of the reports of
these comparisons send requests to use new metrics of
VMAF type. The main obstacle for the full tran-sition
to the use of VMAF metrics is non-versatility of this
metric and not fully adequate results on some types of
videos [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>The main goal of our investigation was to prove
the no-universality of the current version of VMAF
algorithm. In this paper, we describe video color
and contrast transformations which increase
VMAFscore with keeping SSIM score the same or better.
The possibility to improve full-reference metric score
after adding any transformations to distorted im-age
means that the metric can be cheated in some cases.
Such transformations may allow competitors, for
example, to cheat in video codecs comparisons, if they
“tune” their codecs for increasing VMAF qual-ity
scores. Types of video distortions that we were
looking for change the visual quality of the video,
which should lead to a decrease in the value of any
full-reference metric. The fact that they lead to an
increase in the value of VMAF, is a significant
obstacle to using VMAF for all types of videos as the main
quality indicator and proves the need of modification of
the original VMAF algorithm.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Study Method</title>
      <p>During testing of VMAF algorithm forvideo
codecs comparisons, we noticed that it reacts on
contrast distortions, so we chose colorand contrast
adjustments as basic types of the searched video
transformations. Two famous and common approaches for
color adjustments were tested to find the beststrat-egy
for VMAF scores increasing. Two cases of
transformations application to the video were tested:
applying transformation before and after video
encoding. In general, there was no significant difference
between these options, because the compression step can be
omitted for increasing VMAF with color enhance-ment.
Therefore, further we will describe only the first case
with adjustment before compression, and we leave the
compression step because in our work VMAF tuning is
considered in case of video-codec compar-isons.</p>
      <p>
        We chose 4 videos which represent different
spatial and temporal complexity [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], content and contrast to
test transformations which may influence VMAF
scores. All videos have FullHD resolution and high bit
rate. Bay time-lapse and Red kayak were filmed in flat
colors, which usually require color post-processing.
Three of the videos (Crowd run, Red kayak and Speed
bag) were taken from open video collection on
media.xiph.org and one was taken from MSU video
collection used for selecting testing video sets for annual
video codecs comparison [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The description (and
sources) of the first three videos can be found on site
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], and the rest Bay time-lapse video sequence
contained a scene with water and grass and the grass and
waves on the water.
      </p>
      <p>
        Three versions of VMAF were tested: 0.6.1, 0.6.2,
0.6.3. The implementations of all three metric
versions from MSU Video Quality Measurement Tool [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
were used. The results did not differ much, so the
following plots are presented for the latest (0.6.3) version of
VMAF.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Proposed Tuning Algorithm</title>
      <p>
        For color and brightness adjustment, two known
and widespread image processing algorithms were
chosen: unsharp mask and histogram equalization. We
used the implementations of these algorithms which
are available in open-source scikit-image [13] library.
In this library, unsharp mask has two parameters
which influence image levels: radius (the radius of
Gaussian blur) and amount (how much contrast was
added at the edges). For histogram equalization, a
parameter of clipping limit was analyzed. In order to
find optimal configurations of equalization parame-ters, a
multi-objective optimization algorithm NSGA-II [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]
was used. Only the limits for the parameters were set
to the genetic algorithm, and it was applied to find the
best parameters for each testing video.
      </p>
      <p>SSIM and VMAF scores were calculated for each
video processed with the considered color
enhancement algorithms with diferent parameters. As it was
mentioned before, after color correction the videos
were compressed with medium preset of x264 encoder
on 3 Mbps. Then, the diference between metric scores
of processed videos and original video were calculated
to compare, how color corrections influenced quality
scores. Fig. 2 shows this diference for SSIM metric
of Bay time-lapse video sequence for diferent
parameter values of unsharp mask algorithm. The
similarity scores for VMAF quality metric are presented in
Fig. 3.</p>
      <p>Fig. 2. SSIM scores for Fig. 3. VMAF scores for
different parameters of different parameters of
unsharp mask on Bay unsharp mask on Bay
time-lapse video sequence. time-lapse video sequence.</p>
      <p>On these plots, higher values mean that the
objective quality of the color-adjusted video was better
according to the metric. VMAF shows better scores for
high radius and a medium amount of unsharp mask,
and SSIM becomes worse for high radius and high
amount. The optimal values of the algorithm
parameters can be estimated on the difference in these plots.
For another color adjustment algorithm (histogram
equalization), one parameter was optimized and the
results are presented on Fig. 4 together with the
results of unsharp mask.</p>
      <p>74
)72
3
.
6
.
0
(F70
A
M
V
-Y68
66
0.82</p>
      <p>Histogram equalization
Unsharp mask
Without correction
0.83 0.84 0.85 0.86</p>
      <p>Y-SSIM
0.87
0.88
0.89</p>
      <p>Fig. 4. Comparison of VMAF and SSIM scores for
different configurations of unsharp mask and histogram
equalization on Bay time-lapse video sequence. The
results in the second quadrant, where SSIM values weren't
changed and VMAF values increased, are interesting for
us.</p>
      <p>According to these results, for some configurations of
histogram equalization VMAF become significantly
better (from 68 to 74) and SSIM doesn’t change a lot
(decrease from 0.88 to 0.86). The results slightly dif-fer
for other videos. On Crowd run video sequence,
VMAF was not increased by unsharp mask (Fig. 5a)
and was increased a little by histogram equalization.
For Red kayak and Speed bag videos, unsharp mask
could significantly increase VMAF and justslightly
decrease SSIM (Fig. 5b and Fig. 5c)</p>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <p>The following examples of frames from the
testing videos demonstrate colorcorrections which
increased VMAF and almost did not influence the
values of SSIM. Unsharp maswkith radius = 2.843
and amount = 0.179 increased VMAF without
significant decrease of SSIM for Bay time-lapse (Fig. 6a
and Fig. 6b). The images before and after masking
look equivalent (a comparison in a checkerboard view is
in Fig. 7) and have similar SSIM score, while VMAF
score is better after thetransformation.
54
)52
3
.
6
.
0
(F50
A
M
V
-Y48
46
44
)43
3
.642
.
0
(
F41
A
M
-V40
Y
39
38
99
)398
.
6
.
0
(F97
A
M
V
-96
Y
95</p>
      <p>Histogram equalization
Unsharp mask</p>
      <p>Without correction</p>
      <p>For Crowd run sequence, histogram equalization
with kernelsize = 8 and cliplimit = 0.00419 also
increased VMAF (Fig. 8a and Fig. 8b). The video is
more contrasted, so the decrease in SSIM was more
significant. However, tho images also look similar
(Fig. 9) and have similar SSIM score, while VMAF
showed better score after contrast transformation.</p>
      <p>Red kayak looked better according to VMAF after
unsharp mask with radius = 9.436, amount = 0:045.</p>
      <p>For Speed bag, the following parameters of unsharp
mask allowed to increase VMAF greatly without
influencing SSIM: radius = 9.429, amount = 0.114.
Fig. 8. Frame 1 from Crowd run video sequence and its
histogram with and without color correction. Two
images and their histograms look almost similar.</p>
      <p>Fig. 9. Checkerboard comparison of frame 1 from
Crowd run video sequence before and after distortions.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>Video quality reference metrics are used to show
the difference between original and distorted streams
and are expected to take worse values when any
transformations were applied to the originalvideo.
However, sometimes it is possible to deceive objective
metrics. In our article, we described the way to increase
the values of popular full-reference metric VMAF. If
the video is not contrasted, VMAF can be increased by
color adjustments without influencing SSIM. In
another case, contrasted video can also be tuned for
VMAF but with little SSIMworsening.</p>
      <p>Although VMAF has become popular and
important, particularly for video codec developers and
customers, there are still a number of issues in its
application. This is why SSIM is used in many competitions,
as well as in MSU Video-Codec Comparisons, as a
main objective quality metric.</p>
      <p>We wanted to pay attention to this problem and
hope to see the progress in thisare, which is likely to
happen since the metric is beingactively developed.
Our further research will involve a subjective
comparison of the proposed color adjustments to the original
videos and the development of novel approaches for
metric tuning.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Acknowledgments</title>
      <p>This work was partially supported by the Russian
Foundation for Basic Research under Grant
19-01-00785a.</p>
    </sec>
    <sec id="sec-7">
      <title>7. References</title>
      <p>[12] H. R. Sheikh and A. C. Bovik, “Image informa-tion
and visual quality,” in IEEE International Conference
on Acoustics, Speech, and Signal Pro-cessing, 2004,
. 3. .– iii-709.
[13] S. van der Walt, J. L. Schonberger, J.
NunezIglesias, F. Boulogne, J. D. Warner, N. Yager, E.
Gouillart, T. Yu, and the scikit-image
contributors. scikit-image: Image processing in Python.
PeerJ 2:e453, 2014.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Cisco</given-names>
            <surname>Visual Networking</surname>
          </string-name>
          <article-title>Index: Forecast and</article-title>
          <string-name>
            <surname>Methodology. 2016-</surname>
          </string-name>
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>HEVC</given-names>
            <surname>Video Codec Comparison 2018 (Thirteen MSU Video Codec Comparison</surname>
          </string-name>
          ) http://compression.ru/video/codec_ comparison/ hevc_2018/
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>MSU</given-names>
            <surname>Quality Measurement</surname>
          </string-name>
          <string-name>
            <surname>Tool</surname>
          </string-name>
          : Download Page http://compression.ru/video/quality_ measure/ vqmt_download.html
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Perceptual</given-names>
            <surname>Video Quality</surname>
          </string-name>
          <article-title>Metrics: Are they Ready for the Real World? Available online: https://www.ittiam.com/perceptual-videoquality-metrics-ready-real-world</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <article-title>[5] VMAF: Perceptual video quality assessment based on multi-method fusion</article-title>
          , Netflix, Inc.,
          <year>2017</year>
          https:// github.com/Netflix/vmaf.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <article-title>[6] Xiph.org Video Test Media [derf's collection] https://media</article-title>
          .xiph.org/video/derf/
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>C. G.</given-names>
            <surname>Bampis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A. C.</given-names>
            <surname>Bovik</surname>
          </string-name>
          , “
          <article-title>Spatiotemporal feature integration and model fusion for full reference video quality assessment,”</article-title>
          <source>in IEEE Transactions on Circuits and Systems for Video Technology</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>C.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Inguva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rankin</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Kokaram</surname>
          </string-name>
          , “
          <article-title>A subjective study for the design of multiresolution ABR video streams with the VP9 codec</article-title>
          ,” in Electronic Imaging,
          <year>2016</year>
          (2), pp.
          <fpage>1</fpage>
          -
          <lpage>5</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Chikkerur</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Sundaram</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Reisslein</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L. J.</given-names>
            <surname>Karam</surname>
          </string-name>
          , “
          <article-title>Objective video quality assessment meth-ods: A classification, review, and performance comparison,” in IEEE Transactions on Broadcast-ing</article-title>
          ,
          <volume>57</volume>
          (
          <issue>2</issue>
          ), pp.
          <fpage>165</fpage>
          -
          <lpage>182</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>K.</given-names>
            <surname>Deb</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Pratap</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Agarwal</surname>
          </string-name>
          , and
          <string-name>
            <surname>T. A. M. T. Meyarivan</surname>
          </string-name>
          , “
          <article-title>A fast and elitist multiobjective genetic algorithm: NSGA-II,” in IEEE transactions on evolutionary computation, 6(2</article-title>
          ), pp.
          <fpage>182</fpage>
          -
          <lpage>197</lpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>S.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , L. Ma, and
          <string-name>
            <given-names>K. N.</given-names>
            <surname>Ngan</surname>
          </string-name>
          , “
          <article-title>Image quality assessment by separately evaluating detail losses and additive impairments”</article-title>
          ,
          <source>in IEEE Transactions on Multimedia</source>
          ,
          <year>2011</year>
          ,
          <volume>13</volume>
          (
          <issue>5</issue>
          ), pp.
          <fpage>935</fpage>
          -
          <lpage>949</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>