<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Driving Road Safety Forward: Video Data Privacy Task at MediaEval 2021</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alex Liu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrew Boka</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Asal Baragchizadeh</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chandini Muthukumar</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Victoria Huang</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Arjun Sarup</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Regina Ferrell</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gerald Friedland</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thomas P. Karnowski</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Meredith M. Lee</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alice J. O'Toole</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Division of Computing, Data Science, and Society, University of California</institution>
          ,
          <addr-line>Berkeley</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Oak Ridge National Laboratory</institution>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>School of Behavioral and Brain Sciences, The University of Texas at Dallas</institution>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>This paper gives an overview of the Driving Road Safety Forward: Video Data Privacy Task organized as part of the Benchmarking Initiative for Multimedia Evaluation (MediaEval) 2021. The goal of this video data task is to explore methods for obscuring driver identity in driver-facing video recordings while preserving human behavioral information.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        The lifetime odds for dying in a car crash are 1 in 107 [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Each year,
vehicle crashes cost hundreds of billions of dollars [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Research
shows that driver behavior is a primary factor in 23 of crashes and
a contributing factor in 90% of crashes [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        Video footage from driver-facing cameras presents a unique
opportunity to study driver behavior. Indeed, in the United States,
the Second Strategic Highway Research Program (SHRP2) worked
with drivers across the country to collect more than 1 million hours
of driver video [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. Moreover, the growth of both sensor
technologies and computational capacity provides new avenues for
exploration. However, video data analysis and interpretation
related to identifiable human subjects bring forward a variety of
multifaceted questions and concerns, spanning privacy, security,
bias, and additional implications [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. The goal of the Task will be to
develop identity masking methods that efectively conceal the
identity of the driver, while simultaneously preserving facial actions
that can be informative for understanding driver behaviors that
contribute to accidents and other driving actions that pose potential
safety hazards. This Task aims to advance the state-of-the-art in
video de-identification, encouraging participants from all sectors
to develop and demonstrate techniques to mask facial identity and
preserve facial action using the provided data. Successful methods
balancing driver privacy with fidelity of relevant information have
the potential to not only broaden researcher access to existing data,
but also inform the trajectory of transportation safety research,
policy, and education initiatives [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>DATA</title>
      <p>
        The dataset consists of both high- and low-resolution driver video
data prepared by Oak Ridge National Laboratory (ORNL) for this
Driver Video Privacy Task. The videos were captured using the
same data acquisition system as the larger SHRP2 dataset mentioned
above, which currently has limited access in a secure enclave. For
the data in this Task, there are drivers in choreographed situations
designed to emulate diferent naturalistic driving environments.
Actions include talking, coughing, singing, dancing, waving,
eating, and various other actions that are typical among drivers [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
Through this unique partnership, annotated data from ORNL will be
available to registered participants, alongside experts from the data
collection and processing team who will be available for mentoring
and any questions.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>EVALUATION OVERVIEW</title>
      <p>To assess the de-identification of faces and measure the
consistency in preserving driver actions and emotions, there will be a
preliminary automated evaluation as well as a human evaluation.
The scores for each of the automated and human evaluations will
be combined for an overall assessment, prioritizing the human
assessment of de-identification and action preservation. This Task
is heavily reliant on human evaluation, and we encourage
participants to include in their submission any ideas, methods, and results
from their own evaluation approaches. This includes any available
data from participants, descriptions of methodology, assumptions,
and results. This information will be shared with reviewers and the
project organizers for additional discussion and opportunities for
seed funding for further research.</p>
      <p>Although we encourage all Task participants to think creatively
and holistically about how the expectations of privacy, the risk
from potential attackers, and various threat models may evolve,
our starting assumptions are that: (1) The drivers are not known to
the potential attacker; there is no relationship between the attacker
and the driver; the driver is not a public figure. (2) Information
from the driver’s surroundings will not influence the attacker’s
ability to identify the driver. (3) Access to the data is limited to
registered users who have signed a Data Use Agreement specifying
they will not attempt to learn the identity of individuals in the
videos. (4) Attackers have access to basic computational resources.
(5) There is a low probability of attackers launching an efective
crowd-sourcing strategy to re-identify the drivers, in part due to the
Data Use Agreement and context in which the data were collected.
4</p>
    </sec>
    <sec id="sec-4">
      <title>DE-IDENTIFICATION TESTING</title>
      <p>
        Human evaluation is adapted from the methodology as described by
Baragchizadeh et al. in Evaluation of Automated Identity Masking
Method (AIM) in Naturalistic Driving Study (NDS) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
4.1
      </p>
    </sec>
    <sec id="sec-5">
      <title>Human participants</title>
      <p>Undergraduate student volunteers will be recruited from the
University of Texas at Dallas (UTD) to participate in the study in exchange
for research credit. All procedures will be approved by the
Institutional Review Board (IRB) of UTD. In all cases, a minimum of 10
students will evaluate each video for identity masking success.
4.2</p>
    </sec>
    <sec id="sec-6">
      <title>Procedure</title>
      <p>Evaluations will be conducted using the masked videos submitted
by the Task participants. For each submission1, a subset of at least
116 masked videos will be evaluated by human participants using
an identity-matching procedure. The selected video subset will
be identical for all Task participants and will be chosen by the
organizers of the evaluation. The particular set of videos to be used
in the evaluation will not be revealed until all submissions have
been processed.</p>
      <p>On each trial, the participant will be asked to match the identity
of the person shown in a masked video to one of 5 high-resolution
static facial images presented simultaneously at the top of the screen.
The participants will be ofered responses to indicate which of the
static images shows the person pictured in the video, or to indicate
that the person pictured in the video does not appear in the set of
the static images. The participant will have the option of replaying
the video as many times as they want before entering a response.</p>
      <p>The static face images will be matched demographically to the
person in the video so that gender, race, and age cannot provide
cues to the identity of the person in the video. Each static face image
will be cropped to show only the internal face, so that identification
cannot be based on peripheral face cues such as hair style.
4.3</p>
    </sec>
    <sec id="sec-7">
      <title>Results Analysis</title>
      <p>Identification of the original (unmasked) videos from this dataset
was assessed in a previous study at the Univ. of Texas at Dallas,
using the methods described here for the Task evaluation. The
identification accuracy results for these unmasked videos will be
used to assess the success of Task participants in masking the
face identities. It is important to note that matching the identity
of the faces between the unmasked videos and the face images
is not perfect. This is due to diferences in the image/appearance
conditions between the static face images (high resolution, cropped
to show only the internal face) and the videos (inside the car,
highand low-resolution, variable expression, etc.). Therefore, masking
success will be measured for each trial as the diference between the
identification of the unmasked video (from the previous evaluation
at UTD) and the identification of the masked video (from the human
evaluation to be conducted on the masked videos submitted by Task
participants).
1subject to the constraint that the algorithm is submitted by the deadline published on
the MediaEval website</p>
      <p>Task participants will be given summary data on the overall
accuracy of their submission, as well as complete data on their
performance for the individual videos. These data should be helpful
for troubleshooting and improving the performance of the
masking algorithm. We will also make available summary data on the
performance of the other participants, so that individual Task
participants can determine how their algorithm performed relative to
other submissions.</p>
      <p>The computational face recognition evaluation consists of face
recognition and face detection steps. In the face recognition step,
masked faces from selected frames are compared with the gallery
face of the driver. We log the number of instances where the
matching metric between the gallery face and the unmasked face indicate
a match. In the face detection step, we attempt to detect faces in the
masked video, and compute the intersection of union (IoU) score.
We compute the cumulative score for IoU across tested frames.
5</p>
    </sec>
    <sec id="sec-8">
      <title>ACTION PRESERVATION TESTING</title>
      <p>The human evaluation procedures and results analysis for action
preservation will be similar to those described for de-identification
testing, with the following changes. On each trial, a masked video
will be presented alongside a list of possible actions. The participant
will be asked to select the “most obvious” action they detect in the
video. Again, the results will be compiled as the diference between
the accuracy of identifying the action in the unmasked video (from
the previous UTD evaluation) and the masked video (from the Task
participant) submission.</p>
      <p>
        The automated approach to measuring action preservation will
use a deep-learning based gaze estimator [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. The action
preservation is estimated by extracting the predicted gaze-vectors from both
the original un-filtered video and de-identified video and
measuring the Euclidean diference between the two unit vectors. Scoring
closer to 0 implies quality of action preservation since the gaze
estimation is relatively unchanged.
6
      </p>
    </sec>
    <sec id="sec-9">
      <title>DISCUSSION AND OUTLOOK</title>
      <p>With the increased availability, prominence, and applicability of
data in our daily lives, multidisciplinary connections and
engagement are critical to harnessing societal benefit from advances in
technology and methodology. This focused video de-identification
task serves as a key example of how data science collaborations
designed to bridge research and practice can simultaneously help
address a pragmatic need, while sparking new lines of inquiry and
research trajectories. The Driving Road Safety Forward: Video Data
Privacy Task strives to raise awareness about transportation
fatalities and how data might enable thoughtful discussion, analysis, and
actions for the betterment of our community safety.</p>
    </sec>
    <sec id="sec-10">
      <title>ACKNOWLEDGMENTS</title>
      <p>Special thanks to our collaborators and advisors, including David
Kuehn, Charles Fay, Daniel Morgan, Natalie Evans Harris, Lauren
Smith, René Bastón, Mae Tanner, David E. Culler, the U.S.
Department of Transportation (USDOT), and the National Science
Foundation (NSF) Big Data Hubs network. This efort is made possible
through community volunteers, NSF Grants 1916573, 1916481, and
1915774, and an inter-agency agreement between NSF and USDOT.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <fpage>2021</fpage>
          .
          <article-title>About safety Data: Strategic Highway research Program 2</article-title>
          . (
          <year>2021</year>
          ). http://www.trb.org/StrategicHighwayResearchProgram2SHRP2/ SHRP2DataSafetyAbout.aspx
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <year>2021</year>
          .
          <article-title>A Brief Look at the History of SHRP2</article-title>
          . (
          <year>2021</year>
          ). http://shrp2. transportation.org/pages/History-of-SHRP2.aspx
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <fpage>2021</fpage>
          . Exploratory Advanced Research Program Video Analytics Research Projects. (
          <year>2021</year>
          ). https://www.fhwa.dot.gov/publications/ research/ear/15025/15025.pdf
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Asal</given-names>
            <surname>Baragchizadeh</surname>
          </string-name>
          , Thomas P Karnowski, David S Bolme, and
          <string-name>
            <surname>Alice J O'Toole</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Evaluation of Automated Identity Masking Method (AIM) in Naturalistic Driving Study (NDS)</article-title>
          .
          <source>In 12th IEEE International Conference on Automatic Face &amp; Gesture Recognition (FG</source>
          <year>2017</year>
          ). IEEE,
          <fpage>378</fpage>
          -
          <lpage>385</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Lawrence</given-names>
            <surname>Blincoe</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ted R Miller</surname>
          </string-name>
          ,
          <source>Eduard Zaloshnja, and Bruce A Lawrence</source>
          .
          <year>2015</year>
          .
          <article-title>The economic and societal impact of motor vehicle crashes, 2010 (Revised)</article-title>
          .
          <source>Technical Report.</source>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Thomas</surname>
            <given-names>A Dingus</given-names>
          </string-name>
          , Feng Guo,
          <string-name>
            <given-names>Suzie</given-names>
            <surname>Lee</surname>
          </string-name>
          , Jonathan F Antin,
          <string-name>
            <surname>Miguel Perez</surname>
          </string-name>
          , Mindy
          <string-name>
            <surname>Buchanan-King</surname>
            ,
            <given-names>and Jonathan</given-names>
          </string-name>
          <string-name>
            <surname>Hankey</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Driver crash risk factors and prevalence evaluation using naturalistic driving data</article-title>
          .
          <source>Proceedings of the National Academy of Sciences</source>
          <volume>113</volume>
          ,
          <issue>10</issue>
          (
          <year>2016</year>
          ),
          <fpage>2636</fpage>
          -
          <lpage>2641</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>NSC</given-names>
            <surname>Safety Facts</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Odds of dying</article-title>
          . (
          <year>2021</year>
          ). https://injuryfacts.nsc. org/all-injuries/
          <article-title>preventable-death-overview/odds-of-dying/</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>R.</given-names>
            <surname>Ferrell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Aykac</surname>
          </string-name>
          , Thomas Karnowski, and
          <string-name>
            <given-names>N.</given-names>
            <surname>Srinivas</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>A Publicly Available, Annotated Data Set for Naturalistic Driving Study and Computer Vision Algorithm Development</article-title>
          . (
          <year>2021</year>
          ). https://info. ornl.gov/sites/publications/Files/Pub122418.pdf
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>K.</given-names>
            <surname>Finch</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>A visual guide to practical data de-identification</article-title>
          . (
          <year>2016</year>
          ). https://fpf.org/blog/ a
          <article-title>-visual-guide-to-practical-</article-title>
          <string-name>
            <surname>data-</surname>
          </string-name>
          de-identification
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Xucong</surname>
            <given-names>Zhang</given-names>
          </string-name>
          , Yusuke Sugano, Mario Fritz, and
          <string-name>
            <given-names>Andreas</given-names>
            <surname>Bulling</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>It's Written All Over Your Face: Full-Face Appearance-Based Gaze Estimation</article-title>
          . (
          <year>2017</year>
          ).
          <source>arXiv:cs.CV/1611.08860</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>