<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Improving the Neural Network Algorithm for Assessing the Quality of Facial Images*</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Faculty of Computational Mathematics</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Cybernetics</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lomonosov Moscow State University</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Moscow</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Russia</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>nikita.lisin</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>anton.konushin}@graphics.cs.msu.ru</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Video Analysis Technologies LLC</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Moscow</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Russia</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>alexander.gromov</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>vadim}@tevian.ru</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>NRU Higher School of Economics</institution>
          ,
          <addr-line>Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <fpage>0000</fpage>
      <lpage>0003</lpage>
      <abstract>
        <p>The paper considers the task of obtaining a quality assessment of facial images for usage in various video surveillance systems, video analytics and biometric identification. Accuracy of person recognition and classification depends on the quality of the input images. We consider an approach to obtaining single face image quality assessment using neural network model, which is trained on pairs of images that are split into two possible classes: the quality of the first image is better or worse than the quality of the second one. Two modifications of the selected baseline algorithm are proposed. A face recognition system is applied to change the loss function and image and face quality attributes are used when training the model. Experimental studies of the proposed modifications show their effectiveness. The accuracy of selecting the best and worst frame is increased by 1.3% and 1.9%, respectively.</p>
      </abstract>
      <kwd-group>
        <kwd>Computer Vision</kwd>
        <kwd>Face Quality Assessment</kwd>
        <kwd>Face Recognition</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Computer vision algorithms such as face recognition, algorithms for determining
emotions, demographic characteristics and key points of a human face, are widely used in
video surveillance systems, video analytics and biometric identification. The received
data in these systems is a video stream, which contains a set of several frames for each
person. But most algorithms are built so that they process frames independently of each
other and, as has been shown in many studies, their accuracy depends on the quality of
the input images [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Therefore, these systems use face quality assessment algorithms
*Publication is supported by RFBR grant №19-07-00844.
to select the best frame to improve system performance [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] or reject the worst frames
to improve the overall accuracy [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        The task of face quality assessment is to obtain one scalar value for the input image
that reflects the overall quality and takes into account both image quality attributes
(illumination, blur, noise, etc.) and face quality attributes (head pose, face occlusion, etc.).
Usually, this scalar value is enclosed in the range from 0 to 100, where the values 0 and
100 correspond to the image with the worst and best quality, respectively. Algorithms
for obtaining this value are trained either on pairs of images split into two possible
classes (the quality of the first image is better or worse than the quality of the second
one), as in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], or using regression to obtain a specific quality value [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Obtaining
ground truth labels in these works is carried out with the help of experts.
      </p>
      <p>
        Many of the recent works consider the problem of face quality assessment from a
different point of view: as an indicator that reflects the usefulness of the image for the
specific algorithm being used. Algorithms for obtaining this value use one of the
existing face recognition systems, based on which either the training and test dataset is
marked up [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] or the finished model is obtained directly [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. The main idea is that
the confidence of the face recognition system for a pair of images of the same person
and the difference in the quality of these images are interrelated. The lower the
confidence of the face recognition system that the images in a pair belong to the same person,
the more they differ from each other in quality, and vice versa.
      </p>
      <p>
        In this paper, we use the approach from article [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] to obtain an overall quality
indicator. Two modifications are proposed for the baseline algorithm. The first
modification is the use of face recognition system to change the loss function. Our approach
differs from the previous ones in that the developed algorithm remains universal: it can
be applied together with any other algorithm, and not only with the used face
recognition system. The second proposed modification is to apply image and face attributes.
Algorithms based on this approach use training of the neural network model for several
tasks [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], when one of the tasks is quality assessment, and the remaining tasks are image
and face attributes assessments. For example, in [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], the authors use sharpness, tone
and colorfulness, and their experiments show that this leads to improved algorithm
accuracy. At the same time, there is a small number of works that use images of human
face and take into account new properties. An example of such work is [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], which uses
alignment, visibility, deflection and clarity. The disadvantage is that the markup method
chosen in this article is quite subjective: the ground truth labels are the mean opinion
scores in the range from 0 to 1, obtained with the help of experts. We use image
attributes such as illumination and blur, which are marked up for image pairs, as well as face
attributes such as head rotation angles in the range from −90° to +90° and occlusion
marking for 23 areas of the face into two possible classes. We assume that the
considered approach to applying attributes for face quality assessment is more reliable than
previously proposed.
      </p>
      <p>Improving the Neural Network Algorithm for Assessing the Quality of Facial Images 3
2</p>
    </sec>
    <sec id="sec-2">
      <title>Baseline algorithm</title>
      <p>
        In the article [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], on which our algorithm is based, a neural network model is trained in
two stages. At the first stage the neural network model is a siamese network with two
identical branches with shared weights. The input is a pair of images, where second
image in the pair is obtained from the first using some distortion, which degrades the
image quality. The first and second images are the input of the first and second branches
of the network respectively and the output is the scores Q1 and Q2. These values are
used in Hinge Loss as follows (1):
      </p>
      <p>HingeLoss = max(0, Q2 − Q1 + 1).
(1)
The minimum of this function is achieved when the quality score of the second image
in the pair is less than the quality score of the first image in more than 1. Proposed
approach allows to obtain a ranking model by training on pairs of images without
manual markup. At the second stage, only one branch of the siamese network is used to
evaluate the quality of a single face image. This branch is also fine-tuned on a separate
dataset using regression.</p>
      <p>
        Our algorithm trains in the same way on pairs of images, but we use pairs of different
images containing people's faces. For each pair, the markup into two possible classes
was obtained using experts: the quality of the first image is better or worse than the
quality of the second one. The best image in a pair was considered to be the one that
best matches the combination of the following properties: front projection, normal
illumination, absence of occlusion, noise and blur, etc. This approach for obtaining pairs
was chosen because, from our point of view, it allows us to take into account more
complex cases that occur between pairs of different images, which cannot be obtained
by applying distortion to the original image. Different strategies for selecting pairs for
markup were used to cover more cases, and pairs for which it is not possible to uniquely
define a class were not used later. We don't apply fine-tuning on a separate dataset.
Resnet-10 [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], shown in Fig. 1, is used as a neural network model.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>First modification</title>
      <p>As the first modification of the baseline algorithm, a new loss function is proposed,
which uses the output of the face recognition system. In the baseline algorithm loss
function for quality assessment has the following form (2):</p>
      <p>LFQA = max(0, Q2 − Q1 + margin),
(2)
where margin is a constant value equal to 1. For a part of the training dataset that
consists of pairs of images of the same person, we calculate a set of probabilities using the
face recognition system. A probability is in the range from 0 to 1, where the value 1
means that a pair of images belongs to the same person and the value 0 means that they
belong to different people. We use this probability as an indicator of the similarity of
two images in terms of quality. This set is then normalized so that the expected value
would be 0 and the standard deviation would be 1. In the new loss function, margin has
the following form (3):</p>
      <sec id="sec-3-1">
        <title>1, a pair of images of different people</title>
        <p>
          margin= { max(α, 1 − β ∗ FR), a pair of images of a single person ,
(3)
where α ∈ (0, 1) – new minimum value of the margin; β ∈ ℝ+ – custom parameter;
FR ∈ ℝ – output of the face recognition system after normalization. In our experiments,
we use the following parameter values: α = 0.4 , β = 2. Face recognition system is
developed by Video Analysis Technologies [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. Our approach differs from the
previous ones in that the resulting algorithm remains universal: it can be used in the future
together with any other algorithm, and not only with the face recognition system
selected for training.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Second modification</title>
      <p>As a second modification of the baseline algorithm, we consider multi-task learning of
the neural network to evaluate the face quality and image attributes such as illumination
and blur, as well as face attributes such as head pose and face occlusion. The general
loss function has the following form (4):
Loss = w1 ∗ LFQA + w2 ∗ LIllumination + w3 ∗ LBlur + w4 ∗ LPose + w5 ∗ LOcclusion,
(4)
where {wi}i5=1 − weight coefficients. In our experiments, we use the following
parameter values: w1 = 64, w2 = 48, w3 = 48, w4 = 2, w5 = 4.
4.1</p>
      <sec id="sec-4-1">
        <title>Illumination and Blur</title>
        <p>Evaluation of illumination and blur occurs on pairs of images similar to the face quality
assessment in the baseline algorithm, with the use of Hinge Loss as LIllumination and</p>
        <p>Improving the Neural Network Algorithm for Assessing the Quality of Facial Images 5
LBlur. During training, the value of the loss function for pairs without markup is
assumed to be zero.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Head Pose</title>
        <p>The head pose estimation consists of determining three angles for each image in a pair:
pitch, roll and yaw. Each angle is enclosed in the range from −90° to +90°. Training is
performed using regression and Weighted Mean Absolute Error is used as the loss
function (5):</p>
        <p>
          LPose =
α∗∣P−P̅∣+ βα∗+∣Rβ−+R̅γ∣+ γ∗∣Y−Y̅∣ ,
(5)
where ̅P, R̅, ̅Y ∈ ℕ — ground truth labels; P, R, Y ∈ ℕ — neural network output; α,
β, γ ∈ ℝ — weight coefficients for pitch, roll and yaw respectively. In this work, we
use the following parameter values: α = 1, β = 1, γ = 0.05 . The low value of γ is
explained by the fact that the marking for yaw is less accurate than for pitch and roll.
Fig. 2 shows an example of head rotation angles using three guide vectors.
Face occlusion task is to determine one of two types of visibility for 23 areas of the
face: the visible region and the invisible region due to the rotation of the head,
overlapping by an external object, or going beyond the boundaries of the image. The scheme
of dividing the face into regions is based on the approach proposed in [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ], while some
regions were divided into subdomains, and new ones were added. Fig. 3 shows an
example of areas markup. The loss function has the following form (6):
        </p>
        <p>LOcclusion = ∑i2=31 CE Lossi
(6)
where CE Lossi is the Cross Entropy Loss for the i-th region.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Training Dataset</title>
      <p>
        Since the necessary dataset are not publicly available, we created our own training set.
Images for constructing pairs were provided by Video Analysis Technologies [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
Table 1 describes the methods of obtaining labels for quality assessment and attributes.
Expert markup for face quality assessment and blur was performed with the help of five
people and each pair was labeled only by one. The neural network models used for
markup of illumination, head pose, and face occlusion were trained on separate
datasets:
1. the dataset for training face occlusion classifier is marked up with the help of experts;
2. the dataset for training head poses classifier is marked up by determining the rotation
angles based on 68 key points;
3. the dataset for training illumination classifier is marked up into 13 levels of
illumination as follows:
a. the autoencoder was trained;
b. outputs correlating with the degree of illumination were found in the intermediate
representation of the autoencoder;
c. based on the found outputs, all data was divided into 13 classes;
d. additional expert markup was made to remove false cases.
      </p>
      <p>It should be noted that in the case of illumination, training is carried out on pairs of
images, the markup for which is obtained automatically based on their illumination
levels, because this approach turned out to be more stable. General characteristics of
the obtained dataset, taking into account the transitive closure and the number of pairs
used together with the face recognition system, are given in Table 2.</p>
      <p>Improving the Neural Network Algorithm for Assessing the Quality of Facial Images 7
Since the most common datasets are designed to evaluate the quality of an arbitrary
image that does not necessarily contain a human face, we have created a new dataset
for this purpose. The test dataset consists of tracks – sets of 5 to 12 frames containing
images of faces belonging to one person. The total number of tracks is 7070, and for
each track the markup of the best and worst frames was made in terms of quality
assessment. The best frames are those that have the best matches the combination of the
following properties: front projection, normal illumination, absence of occlusion, noise
and blur, etc. Similarly, the worst frames are those that have the worst correspondence
to these properties. Each track was marked by three experts, and the frame was
considered the best or worst only if the opinions of at least two experts were the same. It was
also required to select as few frames as possible. Fig. 4 shows an example of such a
track. For this dataset, we define three metrics: Best Shot Accuracy, Worst Shot
Accuracy and Pair Accuracy.</p>
      <p>Frame 1</p>
      <sec id="sec-5-1">
        <title>Frame 2</title>
      </sec>
      <sec id="sec-5-2">
        <title>Frame 3</title>
      </sec>
      <sec id="sec-5-3">
        <title>Frame 4</title>
      </sec>
      <sec id="sec-5-4">
        <title>Frame 5 (7) (8) (9)</title>
      </sec>
      <sec id="sec-5-5">
        <title>Best Frame</title>
      </sec>
      <sec id="sec-5-6">
        <title>Worst Frame</title>
        <p>
          • S = {St}tN=1 – set of N tracks;
• Mt ∈ ℕ, Mt ∈ [
          <xref ref-type="bibr" rid="ref12 ref5">5, 12</xref>
          ] – number of frames for each track;
• St = {Fkt}kM=t1 – track with the number t, which consists of frames Ft ;
k
• B = {Bt}tN=1 – set of tracks with the best frames, Bt ⊂ St;
• W = {Wt}tN=1 – set of tracks with the worst frames, Wt ⊂ St.
        </p>
        <p>Notation for the algorithm output:
• Q = {Qt}tN=1 – collection of N sets with quality scores;
• Qt = {QFkt}kM=t1 – set of quality scores for a track with the number t.
We define two indicator functions:</p>
        <p>ItB = {
ItW = {
1, Fmt ∈ Bt , m = argmaxQFkt,
0, Fmt ∉ Bt ∀k∈[1, Mt]
1, Fmt ∈ Wt , m = argminQFkt.</p>
        <p>0, Fmt ∉ Wt ∀k∈[1, Mt]
The indicator function (7) corresponds to the choice of the best frame – it is equal to 1
if and only if the frame with the highest quality score is contained in the set of the best
frames. Similarly, (8) corresponds to the selection of the worst frame. Based on the
indicator functions, we define two metrics:</p>
        <p>N∑ ItB
BestShotAccuracy= t=1 ,</p>
        <p>N
Improving the Neural Network Algorithm for Assessing the Quality of Facial Images 9
WorstShotAccuracy= t=N∑1ItW .</p>
        <p>N
(10)
Since each track consists of frames of three types (best, worst and normal), we can also
create image pairs for each track that consist of two different types of frames. The
resulting pairs are marked up into two classes (the quality of the first image is better or
worse than the quality of the second one), which is uniquely determined based on the
types of images in the pair. Pair Accuracy is defined as the percentage of correctly
classified specified pairs, the total number of which is 173 940.
7</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Experiments</title>
      <p>On the test datasets four algorithms are compared: Baseline, Baseline with Modified
Loss, Baseline with Attributes, Baseline with Modified Loss and Attributes. On Fig. 5
the scheme of the baseline algorithm along with two modifications is given.</p>
      <p>
        During training a polynomial change of the learning rate with the degree of
polynomial 2 is used, and the initial value of the learning rate is 0.001. The total number of
epochs is 25. We also use the Ranger optimizer, which is a combination of two methods
proposed in [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] and [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Augmentation is the same transformation of both images in
a pair (changing saturation, illumination, contrast, additive Gaussian noise, blurring,
cropping the image, etc.). Inference latency of the baseline algorithm is 0.002 second
on a single core of Intel Core CPU i5-9400.
The results obtained on the test dataset are shown in Table 3. Experimental assessment
shows that the use of both modifications achieves the best result and increases the
accuracy of selecting the best frame by 1.3%, the accuracy of selecting the worst frame
by 1.9%, pair accuracy by 0.9%. The advantage of the proposed modifications is that
they do not increase the inference time of the baseline algorithm.
To study the dependence of the face recognition system on the quality of input images
we use a second test dataset consisting of 61 500 pairs of images of a single person and
a face recognition system from the company Video Analysis Technologies [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. The
dataset used has a limited variation in image quality and was not specially designed for
this purpose, so it is insufficient to demonstrate the difference between the
modifications. But we present the results as additional confirmation of the applicability of the
developed algorithms. Fig. 6 shows the dependence of True Positive Rate on the
percentage of lowest quality images removed, with a fixed False Acceptance Rate of
0.0001.
      </p>
      <p>Improving the Neural Network Algorithm for Assessing the Quality of Facial Images 11
8</p>
    </sec>
    <sec id="sec-7">
      <title>Conclusion</title>
      <p>
        In this paper the modified neural network algorithm for face quality assessment is
proposed. An approach from article [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] is used for baseline algorithm. As modifications
the face recognition system is applied to change the loss function and image and face
quality attributes are used when training the model. Changes increase the accuracy of
selecting the best and worst frame by 1.3% and 1.9%, respectively, without affecting
performance.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Grother</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hom</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ngan</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanaoka</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Ongoing Face Recognition Vendor Test (FRVT) Part 5: Face Image Quality Asssessment</article-title>
          .
          <source>Information Access Division Information Technology Laboratory</source>
          ,
          <string-name>
            <surname>NIST</surname>
          </string-name>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Nikitin</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Konushin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Konushin</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Face quality assessment for face verification in video</article-title>
          .
          <source>In: Proceedings of the 24th International Conference on Computer Graphics and Vision GraphiCon'2014</source>
          , pp.
          <fpage>111</fpage>
          -
          <lpage>114</lpage>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bagrov</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Konushin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Konushin</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Face recognition with low false positive error rate</article-title>
          .
          <source>In: ISPRS - International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences</source>
          , pp.
          <fpage>11</fpage>
          -
          <lpage>15</lpage>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          , Van De Weijer,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Bagdanov</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          :
          <article-title>RankIQA: Learning from Rankings for No-Reference Image Quality Assessment</article-title>
          .
          <source>In: 2017 IEEE International Conference on Computer Vision</source>
          (ICCV), pp.
          <fpage>1040</fpage>
          -
          <lpage>1049</lpage>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Best-Rowden</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jain</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Learning Face Image Quality From Human Assessments</article-title>
          .
          <source>In: IEEE Transactions on Information Forensics and Security</source>
          , vol.
          <volume>13</volume>
          , no.
          <issue>12</issue>
          , pp.
          <fpage>3064</fpage>
          -
          <lpage>3077</lpage>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Hernandez-Ortega</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Galbally</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fiérrez</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haraksim</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Beslay</surname>
          </string-name>
          , L.:
          <article-title>FaceQnet: Quality Assessment for Face Recognition based on Deep Learning</article-title>
          .
          <source>In: 2019 International Conference on Biometrics (ICB)</source>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Terhorst</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kolf</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Damer</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kirchbuchner</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuijper</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>SER-FIQ: Unsupervised Estimation of Face Image Quality Based on Stochastic Embedding Robustness (</article-title>
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Nikitin</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Konushin</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Konushin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Neural network model for video-based face recognition with frames quality assessment</article-title>
          .
          <source>In: Computer Optics</source>
          <year>2017</year>
          , vol.
          <volume>41</volume>
          , pp.
          <fpage>732</fpage>
          -
          <lpage>742</lpage>
          (
          <year>2017</year>
          ). (In Russian)
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Kuharenko</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Konushin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Simultaneous facial attribute classification with convolutional neural networks</article-title>
          .
          <source>In: 11th International Conference on Pattern Recognition and Image Analysis: New Information Technologies (PRIA-11-2003)</source>
          ,
          <source>IPSI RAS Samara</source>
          , vol.
          <volume>2</volume>
          , pp.
          <fpage>623</fpage>
          -
          <lpage>626</lpage>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peltoketo</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kämäräinen</surname>
          </string-name>
          , J.:
          <article-title>CNN-Based Cross-Dataset No-Reference Image Quality Assessment</article-title>
          . In: 2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW), Seoul, Korea (South), pp.
          <fpage>3913</fpage>
          -
          <lpage>3921</lpage>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Lijun</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xiaohu</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fei</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pingling</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , Xiang-dong,
          <string-name>
            <given-names>Z.</given-names>
            ,
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.:</surname>
          </string-name>
          <article-title>Multi-branch Face Quality Assessment for Face Recognition</article-title>
          .
          <source>In: 2019 IEEE 19th International Conference on Communication Technology (ICCT)</source>
          ,
          <source>Xi'an, China</source>
          , pp.
          <fpage>1659</fpage>
          -
          <lpage>1664</lpage>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>He</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ren</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
          </string-name>
          , J.:
          <article-title>Deep Residual Learning for Image Recognition</article-title>
          .
          <source>In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          , pp.
          <fpage>770</fpage>
          -
          <lpage>778</lpage>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Maze</surname>
          </string-name>
          , Brianna et al.:
          <string-name>
            <surname>IARPA Janus Benchmark - C:</surname>
          </string-name>
          <article-title>Face Dataset and Protocol</article-title>
          .
          <source>In: 2018 International Conference on Biometrics (ICB)</source>
          , pp.
          <fpage>158</fpage>
          -
          <lpage>165</lpage>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lucas</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hinton</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ba</surname>
          </string-name>
          , J.:
          <article-title>Lookahead Optimizer: k steps forward, 1 step back</article-title>
          .
          <source>NeurIPS</source>
          , (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Liu</surname>
          </string-name>
          , Liyuan et al.:
          <article-title>On the Variance of the Adaptive Learning Rate</article-title>
          and Beyond, (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Video Analysis Technologies Homepage</surname>
          </string-name>
          , https://tevian.ru.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>