<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview of the MediaEval 2014 Visual Privacy Task</article-title>
      </title-group>
      <pub-date>
        <year>2014</year>
      </pub-date>
      <fpage>16</fpage>
      <lpage>17</lpage>
      <abstract>
        <p>This paper presents an overview of the Visual Privacy Task (VPT) of MediaEval 2014, its objectives, related dataset, and evaluation approaches. Participants in this task were required to implement a privacy filter or a combination of filters to protect various personal information regions in video sequences as provided. The challenge was to achieve an adequate balance between the degree of privacy protection, intelligibility (how much useful information is retained post privacy filtering), and pleasantness (how minimal were the adverse effects of filtering on the appearance of the video frames). The submissions from the eight (8) teams who participated in this task were evaluated subjectively by surveillance experts, practitioners, data protection experts and by naïve viewers using a crowdsourcing approach.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The video data includes various scenarios featuring one or several
human subjects walking or interacting. The actors may also carry
specific items, which could potentially reveal their identity and
may therefore need to be privacy-filtered appropriately. For
example, the actors are featured carrying backpacks, umbrellas,
wearing scarves, and performing various actions, such as fighting,
pickpocketing, dropping-a-bag, or simply walking. Actors may be
at a distance from the camera or near the camera, making their
faces appear with varying pixel size and quality. The ambient
lighting conditions of the videos also varied widely as they
recorded a range of indoors, outdoors, day/night-time scenes. The
ground truth was created manually by the task organisers and
consisted of annotations of the bounding boxes containing the
regions of High (H), Medium (M), or Low (L) Personally
Identifiable Information elements (PIIs) including persons’ faces
and accessories. In order to simulate context–aware privacy
protection solutions [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], unusual events occurring within the
video datatset, such as fighting, stealing and dropping-a-bag were
also annotated. The annotations were provided in XML format
alongside a foreground mask in the form of binary sequences.
These included such annotations that distinguished the relative
privacy sensitivity of PIIs; namely for Skin (M), Face (H), Hair
(L), Accessories (M), and for Person’s body (L). The dataset was
provided in accordance with the European Data Protection and
ethical compliance guidelines including informed consent and
access control as required. Figure 1 depicts a sample frame from
the dataset with annotated regions as rectangles.
      </p>
    </sec>
    <sec id="sec-2">
      <title>3. MOTIVATION AND OBJECTIVES</title>
      <p>The MediaEval 2014 Visual Privacy Task was motivated by
application domains such as video privacy filtering of videos
taken in public spaces, by smart phones, web-cams, surveillance
CCTVs, and, videos stored in social websites. For this task, the
participants were encouraged to implement a combination of
several privacy filters to protect various personal information
regions in videos, by optimising the privacy filtering so as to: i)
obscure such personal information effectively whilst, ii) keeping
as much as possible of the ‘useful’ information that would enable
a human viewer to form some ‘useful’ interpretation of the
obscured video frame at some level of abstraction without
compromising the privacy protection level as required by the
person(s) featured in the video-frame. Personal visual
information is subjective human-perceived information that can
expose a person’s identity to a human viewer. This can include
richly detailed image regions such as distinctive facial features or
personal jewellery as well as less rich uniform regions e.g. skin
regions (that expose racial identity) or body silhouette showing a
person’s gait (that generally helps to differentiate women from
men and in some cases it may even enable a close friend or a
spouse to identify the person). Therefore, to satisfy both of the
above-mentioned criteria, i), and, ii) above, privacy protection
solutions were required to take into account different types of
visual personal information. . The participants were encouraged
to exploit the annotation information to achieve the appropriate
level of privacy filtering for each person, object, and
Low/Medium/High information regions and accordingly select the
best-fit filtering. It was anticipated that a single privacy filter
applied to all parts of an image would result in a sub-optimal
solution and a combination of several privacy filters would
provide more effective filtering.</p>
    </sec>
    <sec id="sec-3">
      <title>4. SUBMISSIONS EVALUATIONS</title>
      <p>
        The submitted video clips were methodologically evaluated using
UI-REF based privacy protection requirements. Accordingly the
evaluations attempted to assess the perceived efficacy, as well as
side-effects and affects arising from a proposed privacy filtering
solution -as described in [
        <xref ref-type="bibr" rid="ref4 ref5">4,5</xref>
        ]. Eight (8) research teams
submitted privacy filtered video sequences for the evaluation. In
the context of surveillance scenarios, three distinct user studies
were conducted to ensure the validity of the evaluation results.
The subjective evaluations comprised:
a)
b)
c)
      </p>
      <p>
        Stream 1: crowdsourcing evaluations by the general
public from online communities (“naïve subjects”) in
accordance with the methodology in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ];
Stream 2: subjective evaluations by security system
manufacturers and video-analytics technology and
privacy protection solutions developers;
Stream 3: online subjective evaluations by a focus group
comprising trained CCTV monitoring professionals, and
law enforcement personnel.
      </p>
      <p>For consistency in the analysis of evaluation results from all
streams for all participants’ solutions, the same six (6) video clips
were pre-selected from each submission and evaluated using the
three (3) evaluation streams. A questionnaire consisting of 12
questions had been carefully designed to examine aspects related
to privacy, intelligibility, and pleasantness; this was used in
stream 2 and 3. The first (5) questions were aimed at eliciting the
opinions of the evaluators re the Contents of the viewed videos.
The responses to these questions were considered with respect to
the ground truth. The rest of the questions were aimed at eliciting
the Subjective Opinions of the evaluators re the viewed videos.
Stream 1 used a shortened version of the questionnaire with (7)
questions in total due to crowdsourcing constraints. Some 290
workers responded to the crowd-sourced evaluations. In the
design of the crowdsourcing campaign, special care was taken so
that a worker would not see the same content with different filters
(only one filter per content) and would not see different contents
with the same filter (only one content per filter). Also, only the
answers from reliable crowd-sourcing workers were taken into
account. The reliability was ensured via honeypots, mean and
deviation metrics of time per response to a question, and total
time per campaign. Out of the total 290 workers, 230 were found
to have provided reliable responses to all the 8 evaluation batches,
which resulted in 230/8=29 sets of workers' evaluations for each
filter submitted by each participant.</p>
      <p>In Stream 2 evaluations, the focus group consisted of (65)
participants, (15) of them were females; staff from Thales, France
took part in this evaluation. The majority of the participants were
from the R&amp;D departments, while the rest were from
Management, Security, and other departments. The submissions
were evaluated via paper-based responses to the questions.
In Stream 3 evaluations, the focus group comprised of (59)
participants including (22) females. This group included some key
stakeholder types such as people from R&amp;D, data protection, and
law enforcement, who took part in this study from around the
world. The participants streamed the videos and answered the
questionnaire using online forms. As as results of the described
evaluations, VPT participants received a set of 3 by 3 matrices
comprising the results of each participant for each tier of
evaluation; quantified in terms of the following criteria:
1) The Privacy Protection Level – an average level of
privacy protection across all testing video clips.
2) The Level of Intelligibility – the amount of ‘useful’
information that was retained in the video frames after
privacy filtering had been applied.
3) The Pleasantness of the resulting privacy filtered video
frames in terms of their ‘aesthetic’ perceptual appeal to
human viewers.</p>
      <p>Figure 2 depicts an overview of the results from the three (3)
evaluation streams represented by the median values of the
submissions for each criterion.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Senior</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pankanti</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hampapur</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brown</surname>
          </string-name>
          , L.,
          <string-name>
            <surname>Tian</surname>
            ,
            <given-names>Y.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ekin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Connell</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shu</surname>
            ,
            <given-names>C.F.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <article-title>Enabling video privacy through computer vision; IEEE Security and Privacy 3</article-title>
          , no.
          <issue>3</issue>
          , pp.
          <fpage>50</fpage>
          -
          <lpage>57</lpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Korshunov</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Ebrahimi</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <article-title>PEViD: privacy evaluation video dataset applications of digital image processing</article-title>
          ; XXXVI, San Diego, USA,
          <year>August 2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Badii</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Einig</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tiemann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thiemert</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Lallah</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <article-title>Visual context identification for privacy-respecting video analytics</article-title>
          ;
          <source>in IEEE 14th International Workshop on Multimedia Signal Processing (MMSP</source>
          <year>2012</year>
          ), pp.
          <fpage>366</fpage>
          -
          <lpage>371</lpage>
          , Banff, Canada,
          <year>September 2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Badii</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          , “
          <article-title>UI-REF Methodology”, articles online</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Badii</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Al-Obaidi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Einig</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <article-title>MediaEval 2013 Visual Privacy Task: Holistic Evaluation Framework for Privacy by Co-Design Impact Assessment</article-title>
          .
          <source>Proceedings of the MediaEval 2013 Workshop</source>
          , Barcelona, Spain,
          <fpage>18</fpage>
          -
          <lpage>19</lpage>
          October 2013
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Korshunov</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nemoto</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Skodras</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Ebrahimi</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <article-title>Crowdsourcing-based Evaluation of Privacy in HDR Images</article-title>
          .
          <source>SPIE Photonics Europe</source>
          <year>2014</year>
          ,
          <article-title>Optics, Photonics and Digital Technologies for Multimedia Applications</article-title>
          , Brussels, Belgium,
          <source>Proceedings of SPIE</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>