<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>HCMUS at MediaEval 2021: Facial Data De-identification with Adversarial Generation and Perturbation Methods</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Minh-Khoi Pham</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thang-Long Nguyen-Ho</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Trong-Thang Pham</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hai-Tuan Ho-Nguyen</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hai-Dang Nguyen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Minh-Triet Tran</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>John von Neumann Institute</institution>
          ,
          <addr-line>VNU-HCM</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Science</institution>
          ,
          <addr-line>VNU-HCM</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Vietnam National University</institution>
          ,
          <addr-line>Ho Chi Minh city</addr-line>
          ,
          <country country="VN">Vietnam</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>The 2021 MediaEval Multimedia Evaluation introduces a new data de-identification task, which goal is to explore methods for obscuring driver identity in driver-facing video recordings while maintaining visible human behavioral information. Interested in the challenge, our HCMUS team participate in searching for diferent ideas to tackle the problem. We propose two novel approaches as our main contribution: one is based on generative adversarial networks and the other is based on adversarial attacks. Moreover, a specific combination of evaluation metrics is also included for the later method for a fair comparison. The source code is available at https://github.com/kaylode/mediaeval21-drsf</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        Car accidents are an urgent problem for many countries today.
Currently, we develop vehicles to become safer, more durable, but
the number of accidents is still worrisome. In [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], it is shown that
driver-related problems (e.g. distraction, emotionally agitated,
fatigue) account for 90% of crashes. These findings will help
governments, driver educators, vehicle companies, and the public better
understand the situation so that appropriate measures can be taken.
      </p>
      <p>
        With the desire to study driver behavior, it is required to collect
more driving data. This leads to concerns about privacy and security.
At the "Driving Road Safety Forward: Video Data Privacy"
competition, we need to perform de-identification on SHRP2 dataset [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
The dataset shows the drivers’ faces, bodies, genders, and behaviors.
We aim to de-identify such that we can keep as much information
as possible for the behavioral experts.
      </p>
      <p>In this work, we want to focus on hiding their full face, which is
the most important identifier in the given dataset. Specifically, we
propose two approaches:
• A simple process that swaps the driver’s face with an
anonymous face to keep the most facial information.
• An adversarial pipeline to perturbate the face identity
while preserving main facial attributes in form of
embedding features. This approach is followed by specific
evaluation functions for appropriate assessment.</p>
      <p>* Equal contribution
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>METHOD</title>
    </sec>
    <sec id="sec-3">
      <title>Run 01 - Face Swapping</title>
      <p>
        In this run, we implement the idea of swapping face to hide the
real face of the driver while keeping all other facial features like
gaze, eyes, mouth, nose, and head pose. The proposed procedure is
shown in Figure 1. We first extract face from the given RetinaFace
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] detection results. Then, we use the swapping method from [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]
to swap between an anonymous face identity and the driver’s face.
Finally, we bring back the swapped face to the original video, by
using an overlay module. The implemented overlay module in our
work is the overlay feature of a third-party tool (e.g. ffmpeg).
      </p>
      <p>Overlay</p>
      <p>FaceSwap
2.2</p>
    </sec>
    <sec id="sec-4">
      <title>Run 02 - Adversarial Attack</title>
      <p>This run focus on hiding a person’s identity in the image from
human view and preventing unauthorized deep vision algorithms from
extracting useful information while ensuring correct prediction for
authorized algorithms only. In our research, we only consider the
position with rotation of the human face and its eye gaze vector as
principal information for studying the person’s action and behavior.</p>
      <p>The specific approach is a process consisting of two main steps:
(1) Safeguard identified information from being inferred by
unauthorized models.
(2) Guarantee that the model with a defined set of weights can
extract information with low error.</p>
      <p>In the first step, we apply a simple identity masking technique
pixelation and blurring to anonymize the driver’s faces. This step
Minh-khoi Pham, Thang-Long Nguyen-Ho, Trong-Thang Pham et al.</p>
      <sec id="sec-4-1">
        <title>Original images</title>
        <p>Prediction on
original images
De-identified images</p>
      </sec>
      <sec id="sec-4-2">
        <title>Unauthorized prediction</title>
        <p>on de-identified images</p>
      </sec>
      <sec id="sec-4-3">
        <title>Authorized prediction on de-identified images</title>
        <p>provides strong perturbation to the original image such that the
faces may not be identified by humans nor any vision models.</p>
        <p>
          Secondarily, to ensure that the hidden attributes, which are the
bounding box, landmarks, and gaze vector of the face, can be
revealed only to model with a defined set of weights, we utilize
Iterative Fast Gradient Sign Method (I-FGSM) [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
        </p>
        <p>In general, I-FGSM works by exploiting the gradient of the cost
function with respect to the input image to create a new image that
maximizes the loss such that it drives the model’s outputs towards
the desired target. In our case, I-FGSM is used in the opposite
trend to the common goal, which is to modify the input image and
minimize the loss, resulting in changing the model prediction on
the adversarial sample from false to true.</p>
        <p>
          The targeted models (and their default model weights) which
we choose to attack are listed below:
• For face detection, we experiment on the Retina Face [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]
and MTCNN [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]
• For facial landmark detection, we explore the 2D-FAN [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]
• For gaze vector, we use a simple ResNet pretrained on
        </p>
        <p>
          ETH-XGaze [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] and MPII-Gaze [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] datasets.
        </p>
        <p>We carry out the backpropagation process and find the
corresponding quantity to change on the image. Our objective is to
minimize function as described follows:</p>
        <p>=  +  +</p>
      </sec>
      <sec id="sec-4-4">
        <title>Where:</title>
        <p>•  is the box proposal loss of the detection model.
•  is the L2 error between the predicted heatmap and
the ground truth heatmap of the landmarks estimator.
•  is the L2 error between the gaze vector predicted by
the model and the true gaze vector.</p>
        <p>With the proposed loss function, we compute the network
gradient and use it as a perturbation to update the current input image.</p>
        <p>
          +1 =  , { +  (∇ )} [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]
        </p>
        <p>Given the  0 =  the raw input image, we iteratively add
the perturbation to X until  becomes smaller than a predefined
threshold or until  meets the maximum number of iterations.</p>
        <p>In the concept described above, the goal of the problem is to
ensure that the facial attributes extracted by authorized models
between the original image and its de-identified version have a
slight deviation. Therefore, we propose the following assessment
method, which indicates whether the attributes are well hidden or
not:
    (,  ) =
1 ∑︁</p>
        <p>(1 −  (  ,   )) +  ( ,   ) + Θ (  ,   )</p>
        <p>Where  ,  is the adversarial and raw video sequence consisting
of  frames.  ,   is the predictions of model  on two adversarial
and original frames at the same timestamp  respectively. The 
is the Intersect-over-Union between the two faces location. The
 is measured based on the euclidean distance between landmark
points of the faces, Θ is the angle between two gaze vectors. The
ifnal     is calculated by summing all diferences between
pairs of frames across the timestamp.</p>
        <p>Consequently, we expect     is low for authorized
algorithms and higher     for unauthorized ones.
3</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>EXPERIMENTS AND RESULTS</title>
      <p>In the second run, we pass consecutive video sequences as a batch
with a size equal to 64 into our de-identification pipeline to generate
perturbated videos whose facial attributes have been hidden. We
perform a full adversarial pipeline on 720 videos from the SHRP2
dataset and demonstrate some visual results as shown in Figure 2.
4</p>
    </sec>
    <sec id="sec-6">
      <title>CONCLUSION AND FUTURE WORKS</title>
      <p>Conclusively, we present diferent strategies to address the data
privacy issues for MediaEval Challenge 2021. In the future, we
aim to study the performance of our adversarial attack for the
information preservation method on several deep vision models
regarding facial attributes. We are intent to analyze our proposed
metrics on these experiments as well.</p>
    </sec>
    <sec id="sec-7">
      <title>ACKNOWLEDGMENTS</title>
      <p>This work was funded by Gia Lam Urban Development and
Investment Company Limited, Vingroup and supported by Vingroup
Innovation Foundation (VINIF) under project code VINIF.2019.DA19.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Adrian</given-names>
            <surname>Bulat</surname>
          </string-name>
          and
          <string-name>
            <given-names>Georgios</given-names>
            <surname>Tzimiropoulos</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>How Far are We from Solving the 2D 3D Face Alignment Problem? (and a Dataset of 230,000 3D Facial Landmarks)</article-title>
          .
          <source>2017 IEEE International Conference on Computer Vision</source>
          (ICCV) (
          <year>Oct 2017</year>
          ). https://doi.org/10.1109/iccv.
          <year>2017</year>
          .116
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Jiankang</given-names>
            <surname>Deng</surname>
          </string-name>
          , Jia Guo, Yuxiang Zhou, Jinke Yu,
          <string-name>
            <given-names>Irene</given-names>
            <surname>Kotsia</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Stefanos</given-names>
            <surname>Zafeiriou</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>RetinaFace: Single-stage Dense Face Localisation in the Wild</article-title>
          .
          <article-title>(</article-title>
          <year>2019</year>
          ).
          <article-title>arXiv:cs</article-title>
          .CV/
          <year>1905</year>
          .00641
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Thomas</surname>
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Dingus</surname>
            , Feng Guo,
            <given-names>Suzie</given-names>
          </string-name>
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>Jonathan F.</given-names>
          </string-name>
          <string-name>
            <surname>Antin</surname>
          </string-name>
          , Miguel Perez, Mindy
          <string-name>
            <surname>Buchanan-King</surname>
            ,
            <given-names>and Jonathan</given-names>
          </string-name>
          <string-name>
            <surname>Hankey</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Driver crash risk factors and prevalence evaluation using naturalistic driving data</article-title>
          .
          <source>Proceedings of the National Academy of Sciences</source>
          <volume>113</volume>
          ,
          <issue>10</issue>
          (
          <year>2016</year>
          ),
          <fpage>2636</fpage>
          -
          <lpage>2641</lpage>
          . https://doi.org/10.1073/pnas.1513271113 arXiv:https://www.pnas.org/content/113/10/2636.full.pdf
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Alexey</given-names>
            <surname>Kurakin</surname>
          </string-name>
          , Ian Goodfellow, and
          <string-name>
            <given-names>Samy</given-names>
            <surname>Bengio</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Adversarial examples in the physical world</article-title>
          . (
          <year>2017</year>
          ).
          <source>arXiv:cs.CV/1607.02533</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <source>[5] Transportation Research Board of the National Academy of Sciences. 2013. The 2nd Strategic Highway Research Program Naturalistic Driving Study Dataset</source>
          .
          <article-title>(</article-title>
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Sefik</given-names>
            <surname>Ilkin</surname>
          </string-name>
          Serengil and
          <string-name>
            <given-names>Alper</given-names>
            <surname>Ozpinar</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>HyperExtended LightFace: A Facial Attribute Analysis Framework</article-title>
          . In 2021 International Conference on Engineering and
          <article-title>Emerging Technologies (ICEET)</article-title>
          . IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Aliaksandr</given-names>
            <surname>Siarohin</surname>
          </string-name>
          , Subhankar Roy, Stéphane Lathuilière, Sergey Tulyakov, Elisa Ricci, and
          <string-name>
            <given-names>Nicu</given-names>
            <surname>Sebe</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Motion Supervised co-part Segmentation</article-title>
          .
          <source>arXiv preprint</source>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Kaipeng</given-names>
            <surname>Zhang</surname>
          </string-name>
          , Zhanpeng Zhang,
          <string-name>
            <given-names>Zhifeng</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Yu</given-names>
            <surname>Qiao</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Joint Face Detection and Alignment Using Multitask Cascaded Convolutional Networks</article-title>
          .
          <source>IEEE Signal Processing Letters</source>
          <volume>23</volume>
          ,
          <issue>10</issue>
          (Oct
          <year>2016</year>
          ),
          <fpage>1499</fpage>
          -
          <lpage>1503</lpage>
          . https://doi.org/10.1109/lsp.
          <year>2016</year>
          .2603342
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Xucong</given-names>
            <surname>Zhang</surname>
          </string-name>
          , Seonwook Park, Thabo Beeler, Derek Bradley,
          <string-name>
            <given-names>Siyu</given-names>
            <surname>Tang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Otmar</given-names>
            <surname>Hilliges</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>ETH-XGaze: A Large Scale Dataset for Gaze Estimation under Extreme Head Pose</article-title>
          and
          <string-name>
            <given-names>Gaze</given-names>
            <surname>Variation</surname>
          </string-name>
          . (
          <year>2020</year>
          ).
          <article-title>arXiv:cs</article-title>
          .CV/
          <year>2007</year>
          .15837
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Xucong</surname>
            <given-names>Zhang</given-names>
          </string-name>
          , Yusuke Sugano, Mario Fritz, and
          <string-name>
            <given-names>Andreas</given-names>
            <surname>Bulling</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>MPIIGaze: Real-World Dataset and Deep Appearance-Based Gaze Estimation</article-title>
          . (
          <year>2017</year>
          ).
          <source>arXiv:cs.CV/1711.09017</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>