<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Workshops, Los Angeles, USA, March</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Explainable Activity Recognition in Videos</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Chiradeep Roy</string-name>
          <email>Chiradeep.Roy@utdallas.edu</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vibhav Gogate</string-name>
          <email>Vibhav.Gogate@utdallas.edu</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mahesh Shanbhag</string-name>
          <email>Mahesh.Shanbhag@utdallas.edu</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nicholas Ruozzi</string-name>
          <email>Nicholas.Ruozzi@utdallas.edu</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tahrima Rahman</string-name>
          <email>Tahrima.Rahman@utdallas.edu</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mahsan Nourani</string-name>
          <email>mahsannourani@ufl.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Eric D. Ragan</string-name>
          <email>eragan@ufl.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Samia Kabir</string-name>
          <email>samia.kabir@tamu.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Texas A&amp;M University, College Station</institution>
          ,
          <addr-line>Texas</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Florida</institution>
          ,
          <addr-line>Gainesville, FL</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Texas at Dallas</institution>
          ,
          <addr-line>Richardson, TX</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <volume>20</volume>
      <issue>2019</issue>
      <abstract>
        <p>In this paper, we consider the following activity recognition task: given a video, infer the set of activities being performed in the video along with an assignment of activities to each frame in the video. Although this task can be solved accurately using existing deep learning systems, their use is problematic in interactive settings. In particular, deep learning models are black boxes: it is difcfiult to understand how and why the system assigned a particular activity to a frame. This reduces the users' trust in the system, especially in the case of end-users who need to use the system on a regular basis. We address this problem by feeding the output of deep learning to a tractable interpretable probabilistic graphical model and then performing joint learning over the two. The key benefit of our proposed approach is that deep learning helps achieve high accuracy while the interpretable probabilistic model makes the system explainable. We demonstrate the power of our approach using a visual interface to provide explanations of model outputs for queries about videos.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>CCS CONCEPTS</title>
      <p>• Computing methodologies → Activity recognition and
understanding; • Human-centered computing → Graphical user
interfaces; User studies.
IUI Workshops ’19, March 20, 2019, Los Angeles, USA
Copyright © 2019 for the individual papers by the papers’ authors. Copying permitted
for private and academic purposes. This volume is published and copyrighted by its
editors.
1</p>
    </sec>
    <sec id="sec-2">
      <title>INTRODUCTION</title>
      <p>
        We propose a two-layer architecture that stacks a tractable,
interpretable probabilistic graphical model, specifically a cutset network
[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], layer on top of a deep learning layer to address the
aforementioned drawback. A possible interpretation of this model is that
the deep learning layer provides noisy sensory inputs to the cutset
network layer which in turn removes the noise and provides
explainability. The interaction graph of the cutset network encodes our prior
knowledge about the relationship between various (human
interpretable) random variables in the network. The rationale is that the
prior knowledge will help correct the errors made by the neural
network and thus help improve accuracy. The cutset network provides
explainability not only because it is interpretable but also because we
can perform tractable (linear time in the size of the model) abductive
inference to compute explanations for the decisions made by the
model. To model temporal aspects in video, we propose a novel
tractable dynamic probabilistic modeling framework called dynamic
cutset networks and show that they greatly improve the estimation
accuracy.
      </p>
      <p>
        We experimentally demonstrate the efcfiacy of our proposed
approach by building an interactive visual interface and a machine
learning system for activity recognition for the Textually Annotated
Cooking Scenes (TaCOS) dataset [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. The purpose of building this
system is two-fold. First, we want to show that we can create a
working prototype explainable AI system that not only performs
accurate activity recognition in videos but can also generate human
understandable explanations and answer queries posed by end-users.
Second, we aim to use the resulting system as the basis for user
studies of how different types of explanations affect user trust and
understanding of machine learning models.
2
      </p>
    </sec>
    <sec id="sec-3">
      <title>RELATED WORK</title>
      <p>
        This effort was inspired by the work of Rohrbach et al. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] on
generating a semantic representation from videos at an activity level.
Instead of generating sentences in natural language however, we
assign a number of pre-defined labels divided into categories. We
do this by using deep-architectures with proven results in order to
generate high accuracies for predicting the activity labels. Related
efforts have considered the task of dense captioning [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], i.e.,
generating summaries of texts from particular segments. Song et al. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]
attempted to create captioning methods that require minimum
supervision on the TaCOS dataset. Duan et al. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] attempted to combine
caption generation and sentence localization to feed off each other
to create a weakly supervised training model. These works focus
on creating text summaries for video segments, and as is typical
of deep learning approaches, they are essentially black boxes. Our
approach, on the other hand, aims to create a semantic representation
for activities in each frame that can both be used to answer queries
easily as well as generate explanations that justify these answers.
      </p>
      <p>
        There have also been a number of studies on how trust influences
interactions between humans and automated systems, e.g., [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ],
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. These studies examine factors that might affect the trust
of the user in the system, such as showing the past performance of the
system and making the working of the system more understandable
(Lee et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]). Hoffman [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] provides a more detailed taxonomy of
such factors and explain how trust is context-specific and dynamic.
In other words, trust might vary with respect to specific contexts of
automation and must also be maintained over time. Our aim is to
be able to control and measure the trust of humans with respect to
these systems in order to better understand what kind of explanations
influence the trust variable.
3
      </p>
    </sec>
    <sec id="sec-4">
      <title>PROBLEM DESCRIPTION</title>
      <p>In this section, we will define the problem in precise terms and also
describe the framework we will be using for question-answering and
generating explanations.
3.1</p>
    </sec>
    <sec id="sec-5">
      <title>Activity Recognition with Explanations</title>
      <p>The objective of our proposed system is two-fold: (a) perform
accurate activity recognition in videos, and (b) compile knowledge
acquired while learning to recognize activities into an explanatory
model. The latter can then be used to explain why a particular activity
was assigned to a frame by the system.</p>
      <p>We define an activity as a ( action, object, location) triple. The
action component forms the core part of the activity. These are
usually verbs such as wash, cut, slice, open, etc. The object component
denotes the entities over which the activity is performed. These are
generally nouns such as apples, refrigerator, cutting board, knife, etc.
Finally, the location component tells us where the activity is taking
place. These are generally location nouns such as kitchen, bathroom,
counter top, sink, etc. but can also overlap with the nouns we use
as objects. For example, when we “kick open a door,” the activity
is “kick” and the object is “door,” but the same entity might play a
different semantic role in a different activity such as if a baby “draws
a picture on the door.” Here “draw” is the activity, “picture” is the
object, and “door” is the location.</p>
      <p>For the purposes of our initial system, we make the following
simplifying assumptions.</p>
      <p>(1) We train our system on a closed-domain. In this study, we use
cooking videos.
(2) We assume that only one major activity is taking place per
frame (minor activities are ignored).
(3) The action must always be present, while the object and the
location are optional. For reflexive actions, such as “walking,”
the object is “None.”</p>
      <p>In future, we plan on making activities more complex (so that we
can pose more interesting queries on them). We also plan on defining
hierarchies on activities to create ‘super’ and ‘sub’ activities. For
instance, taking out an egg from a refrigerator might be a sub-activity
of cooking the egg, which in turn might be a sub-activity of cooking
a full-course meal.
3.2</p>
    </sec>
    <sec id="sec-6">
      <title>Formulating Queries</title>
      <p>Now that we have precisely defined activities, we can define queries
and explanations. A query is similar to an activity in that it is also
a triple of the form (action, object, location). Once this triple is
formulated, we run a filtering query—which seeks to assign an
activity to each frame in the video based on the current and previous
frames (but not future frames)—to check how many frames match
our query. For instance, if we wanted to ask the system if the person
in the video sliced an orange on the cutting board, our query tuple
would look something like: (slice, orange, cutting-board). Once we
have formulated this tuple, we simply ask the system to search for
frames where the probability of this tuple being the actual activity is
above a certain threshold.</p>
      <p>We envision that our system will be used to answer a wide-range
of queries including but not limited to:
(1) Selection queries. Did the person slice an orange on the
counter?
(2) Counting. How many oranges did the person slice on the
counter?
(3) Recipe. Did the person deviate from a pre-defined recipe?
(4) Complex: Combination of all of the above
In this paper, we will focus on selection queries and leave the
remaining for future work.
3.3</p>
    </sec>
    <sec id="sec-7">
      <title>Generating Explanations</title>
      <p>The aim of our system is not only to answer queries but also to
explain the predictions to the end-user. As mentioned in the
previous section, selection queries involve formulating the query into a
(action, object, location) tuple and then filtering on the video to find
frames which have a high probability of containing the activity. Note
that two of the three parameters can be optional. This means that
our queries can be as simple as “Did the person wash something?”
(wash, ?, ?) or “Did anything happen on the kitchen counter?” (?, ?,
counter), etc.</p>
      <p>We seek to build a system that can generate three types of
explanations:
(1) Video Explanations: When the system answers “yes” we
want the system to highlight segments (possibly more than
one) of the video where the activity happened. For “no”
answers, we want the system to highlight segments where a
related activity happened (e.g., carrots were cut in the video
but not oranges). If no related activity is found in case of a
“no answer,” we want the system to output the most likely
activity in the video.</p>
      <p>(2) Ranked (action,object,location) Triples: We want the
system to display top-k predicted activity triples in the video that
are relevant to the query.
(3) Most Probable Entities: We want the system to display the
most probable actions, objects and locations (along with their
likelihood) that are relevant to the query.
4</p>
    </sec>
    <sec id="sec-8">
      <title>SYSTEM DESCRIPTION</title>
      <p>This section explains the architecture and the functioning of our
system in detail. Fig. 1 shows a high-level overview of the components
of the system and the processing pipeline. Roughly speaking, the
system can be categorized into the following two layers:
(1) Video Classification Layer</p>
      <p>
        In this layer, we use a convolutional neural network whose
architecture is based on GoogleNet [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. The network takes as
input a number of video frames, a vocabulary file, and a set of
annotated ground truths and then uses a version of
backpropagation with the Adam algorithm to learn the weights. The
output is a set of labels that correspond to each vocabulary
word. The accuracy of network for the multilabel
classification task is measured using standard information retrieval
metrics.
(2) Explanation Layer
      </p>
      <p>
        The ground labels and the predicted labels from the previous
layer are fed to this layer and are used to train a Conditional
Cutset Network [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Once training is done, we are now in a
position to pose queries to the system. The system uses the
trained model to answer these queries and returns the top
kbest explanations using sampling-based inference techniques.
Next, we will describe each layer in more detail.
4.1
      </p>
    </sec>
    <sec id="sec-9">
      <title>Video Classification Layer</title>
      <p>
        For this layer, we will be using the BAIR/BLVC GoogleNet Model
[
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. This is a pre-trained model that uses the ILSVRC dataset.
There are 22 layers in the network. It uses a key component called
the Inception Model (for details, please refer [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]) that creatively
uses convolutions of size 1x1 to increase the representational power
the network without increasing the number of parameters. These
layers are stacked one on top of the other. The architecture preserves
translational invariance.
      </p>
      <p>The other reason we used GoogleNet is because the computational
load does not increase exponentially with the increase in layers.
This is because the Inception modules use 1x1 convolutions. The
vanishing gradient problem is taken care of by using rectified linear
units in the perceptrons. This also avoids the introduction of sparse
activations in the hidden layers of the network.
Finally, we made a slight modification to this architecture by
replacing the softmax layer on top with a fully-connected layer with
28 nodes that use the sigmoid cross-entropy loss.</p>
      <p>The video is first divided into frames which are then fed to this
layer along with the ground labels for each frame. For instance, if
we have 10,000 frames in our dataset and the 328th frame has the
person taking a carrot out from the fridge, then the 328th row in the
ground labels file will have ( take-out, carrot, fridge) set to true and
all other labels set to false.</p>
      <p>The output layer has a node for every label in the vocabulary. For
the purposes of this experiment, we are using 28 labels. The neural
network takes in the video frame as an input and tries to guess the
correct activity labels for each frame. The accuracy of this
multilabel classification task is measured using standard metrics such as
the Hamming loss and the Jaccard index.
4.2</p>
    </sec>
    <sec id="sec-10">
      <title>Explanation Layer</title>
      <p>At the end of processing the first layer, we have a list of predicted
labels for selected video frames. In the explanation layer, we correct
errors in the predicted labels using a probabilistic model at each
frame. Also, we model the dynamics as well as persistence (activities
don’t change rapidly between frames) using a temporal probabilistic
model.</p>
      <p>
        The explanation model uses a Conditional Cutset Network (CCN),
[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] to correct the errors at each frame. The reason for using this
probabilistic model is two-fold. First it is interpretable in that its
structure and parameters can be explained to an expert user using
concepts from graph theory and probability theory respectively.
Second, it is a tractable model. In particular, the model can answer
queries in time that scales linearly with its size. At a high level, the
CCN treats the output of the neural network as a noisy sensor (see
Fig. 2(a)) and computes a conditional joint probability distribution
over the true labels given the predicted (noisy) labels.
      </p>
      <p>
        To model dynamics and persistence, we propose to use dynamic
conditional cutset networks (see Fig. 2(b)). To control the number of
parameters and learning complexity, we use 1-Markov and
stationarity assumptions, which are widely used in temporal models literature
[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Specifically, we assume that each frame is conditionally
independent of all frames before it given the previous frame (1-Markov) and
all conditional distributions are identical (stationarity). We model
these conditional distributions using conditional cutset networks.
Thus, a dynamic conditional cutset network is a two-tuple. At the
ifrst frame, we have a conditional cutset network which models the
conditional distribution over the labels in the frame given labels
predicted by the neural network. At subsequent frames, we have a
conditional cutset network which models the conditional distribution
over the labels in the frame given labels predicted by the neural
network and the true labels in the previous frame.
      </p>
      <p>
        The three explanation types (video, ranked triples and most
probable entities) mentioned in the previous section can be computed from
the explanation layer by performing abductive inference (cf. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ])
over the dynamic cutset network. Since inference in cutset networks
is linear in the size of the network, once learned from data, our
explanation layer yields real-time query answers and explanations.
4.3
      </p>
    </sec>
    <sec id="sec-11">
      <title>Dataset and Data Processing</title>
      <p>
        The dataset we are using for this experiment is the TACoS
MultiLevel corpus - MPII Cooking 2 dataset by Rohrbach et al. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Each
video is annotated as follows. The annotations are filename (e.g.,
s24d28), startFrame (e.g. 781), endFrame (e.g., 1098), descriptionIdx
(e.g., 3), ignore (e.g., 0), sentenceProcessed (e.g., the person took
out the cutting board from the drawer and placed it on the counter),
activity (e.g., take out), tool (e.g., hand), object (e.g., cutting board),
source (e.g., drawer), target (e.g., counter).
      </p>
      <p>We generate the ground labels from these annotations. We create
a text file where every row corresponds to a frame of some video. In
addition to the name of the video and the frame number, we have
28 0/1 values depending on which labels are off and which are on.
The order of the labels follows the same order as that of the output
nodes in the video classification layer. We extract the activity, object
and location by punching together the source and destination fields.
For instance, the location drawer-counter indicates that the source
of the action is the drawer and the final destination is the counter.
We repeat this process for each video frame in our training set.</p>
      <p>After we have a list of ground labels and predicted labels, we
train the explanation model and pose queries to it. This part will be
discussed in the Experiments section.
4.4</p>
    </sec>
    <sec id="sec-12">
      <title>User Interface</title>
      <p>The prototype uses an interactive visual interface that allows users
to load videos, ask queries, and review the model output along
with explanations. The goal for the interface design was to limit
the amount of model information presented to the user in order to
avoid overwhelming users with information. For this reason, the
system uses simple visual representations in the form of graphical
annotations, textual component lists, and simple bar charts. Figure 3
shows the interface.</p>
      <p>The interface includes a video player that allows users to watch
the selected video to help review and analyze the system’s answers to
the queries. When a query is submitted, the video player highlights
the most relevant segments of the video through visual annotations
added under the video play bar (shown as orange and purple bars
under the video in Fig. 3). The video player will also automatically
jump to the appropriate segment to help users see the video frames
most important for determining the output. In addition, the right
side of the interface summarizes the detected video components
(activities, objects, and locations) for the query as well as detected
combinations of components. To help users to quickly judge
component scores, graphical bars are shown underneath detected
components to visually represent the values of the component scores.
Users can select different video segments to view the corresponding
component scores and combinations from different portions of the
video.
5</p>
    </sec>
    <sec id="sec-13">
      <title>EVALUATION</title>
      <p>In order to evaluate our system, we designed two experiments using
the TACoS video dataset and annotations. We performed a model
evaluation to measure how successful the system is at identifying
activities and providing explanations, and we performed a preliminary
user study to assess system understandability and usability.
5.1</p>
    </sec>
    <sec id="sec-14">
      <title>Model (Machine Learning) Evaluation</title>
      <p>We selected 60313 frames for training and 9355 frames for testing
distributed over 17 videos. For each set, we selected a set of ground
labels and used the video classification layer to generate the predicted
labels. We performed the following ablation study: (1) Our system in
which the explanation layer is removed (GoogleNet); (2) Our system
in which the dynamic model is removed but the sensor model is
kept at each frame (Sensor Model); and (3) the full system (dynamic
CCNs).</p>
      <p>Table 1 outlines the accuracy scores for correct activity
recognition according to various evaluation metrics. Since predicting each
activity correctly is a multilabel classification task, we use K-Group
measures to calculate the overall percentage of instances where K
labels out of the total number of labels were predicted currently.
We use the group heuristics K-1, K-2, and K-3 (since each activity
comprises of action, object and location). In addition, we also use
standard measures such as the Hamming Loss and the Jaccard Index.
We observe that in general (with a few exceptions) dynamic CCNs is
more accurate than the sensor model which in turn is more accurate
than GoogleNet.
5.2</p>
    </sec>
    <sec id="sec-15">
      <title>Human Feedback</title>
      <p>We sought user feedback via a preliminary testing with seven
participants. All participants were experienced with AI but were
unfamiliar with the specifics of the system. Rather than testing the
querying functionality and model accuracy, we were interested in
general user feedback about perception of the system and
explanations. Participants were asked to review a set of four videos with vfie
pre-determined queries per video.</p>
      <p>K-1
K-2
K-3
pants used each section. We observe that participants found
video segments very helpful in completing the task.</p>
      <p>After participants finished the study, they were asked about the
utility of each section of the interface, in particular, whether the
section helped them complete the given task quickly. The results
of this questionnaire are shown in Fig. 4. Our testing demonstrates
that users were able to effectively and easily use the system to
submit queries and review results. The simplicity of the visual design
enabled participants to easily see the corresponding relevant
segments of the video and to quickly assess the accuracy of the model’s
output. However, participants had mixed thoughts on component
scores (most probable entities) and component combinations (ranked
triples).</p>
      <p>We plan to further examine differences in interpretation and study
the utility of the explanation design through more extensive user
testing in the future. In particular, we will run more formal controlled
experiments of how different types of explanation and amount of
explanatory information affect understanding of the model and
perception of its accuracy.
6</p>
    </sec>
    <sec id="sec-16">
      <title>CONCLUSION</title>
      <p>From our preliminary user studies, a strong positive correlation has
been observed between user trust and the goodness of explanations.
As a part of our future work, we plan on improving upon our current
system in the following manner:
(1) Adding support for more vocabulary in the video
classification layer</p>
      <sec id="sec-16-1">
        <title>Evaluation Metric</title>
      </sec>
      <sec id="sec-16-2">
        <title>GoogleNet</title>
      </sec>
      <sec id="sec-16-3">
        <title>Dynamic CCNs</title>
        <p>(2) Adding support for complex models and automatic query
conversion from natural language in the explanation layer
(3) Adding support for a larger variety of queries</p>
        <p>We expect that adding these features will increase the trust of the
users in the system since the range of activities, the precision of the
explanations as well as the types of queries will all increase.
7</p>
      </sec>
    </sec>
    <sec id="sec-17">
      <title>ACKNOWLEDGMENTS</title>
      <p>This work was supported by the DARPA Explainable Artificial
Intelligence (XAI) Program under contract number
N66001-17-24032, and by the National Science Foundation grants IIS-1652835,
IIS-1528037, and IIS-1762268.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Xuguang</given-names>
            <surname>Duan</surname>
          </string-name>
          , Wenbing Huang, Chuang Gan, Jingdong Wang,
          <string-name>
            <surname>Wenwu Zhu</surname>
            , and
            <given-names>Junzhou</given-names>
          </string-name>
          <string-name>
            <surname>Huang</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Weakly Supervised Dense Event Captioning in Videos</article-title>
          .
          <source>In Advances in Neural Information Processing Systems</source>
          .
          <volume>3062</volume>
          -
          <fpage>3072</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Kevin</given-names>
            <surname>Anthony</surname>
          </string-name>
          Hoff and
          <string-name>
            <given-names>Masooda</given-names>
            <surname>Bashir</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Trust in automation: Integrating empirical evidence on factors that influence trust</article-title>
          .
          <source>Human Factors</source>
          <volume>57</volume>
          ,
          <issue>3</issue>
          (
          <year>2015</year>
          ),
          <fpage>407</fpage>
          -
          <lpage>434</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Robert</surname>
            <given-names>R</given-names>
          </string-name>
          <string-name>
            <surname>Hoffman</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>8 A Taxonomy of Emergent Trusting in the HumanMachine Relationship</article-title>
          .
          <article-title>Cognitive Systems Engineering: The Future for a Changing World (</article-title>
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Ranjay</given-names>
            <surname>Krishna</surname>
          </string-name>
          , Kenji Hata, Frederic Ren,
          <source>Li Fei-Fei, and Juan Carlos Niebles</source>
          .
          <year>2017</year>
          .
          <article-title>Dense-Captioning Events in Videos.</article-title>
          .
          <source>In ICCV. 706-715.</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>John D</given-names>
            <surname>Lee and Katrina A See</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Trust in automation: Designing for appropriate reliance</article-title>
          .
          <source>Human factors 46</source>
          ,
          <issue>1</issue>
          (
          <year>2004</year>
          ),
          <fpage>50</fpage>
          -
          <lpage>80</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Bonnie</surname>
            <given-names>M</given-names>
          </string-name>
          <string-name>
            <surname>Muir</surname>
          </string-name>
          .
          <year>1994</year>
          .
          <article-title>Trust in automation: Part I. Theoretical issues in the study of trust and human intervention in automated systems</article-title>
          .
          <source>Ergonomics</source>
          <volume>37</volume>
          ,
          <issue>11</issue>
          (
          <year>1994</year>
          ),
          <fpage>1905</fpage>
          -
          <lpage>1922</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Bonnie</surname>
            <given-names>M Muir</given-names>
          </string-name>
          and
          <string-name>
            <given-names>Neville</given-names>
            <surname>Moray</surname>
          </string-name>
          .
          <year>1996</year>
          .
          <article-title>Trust in automation. Part II. Experimental studies of trust and human intervention in a process control simulation</article-title>
          .
          <source>Ergonomics</source>
          <volume>39</volume>
          ,
          <issue>3</issue>
          (
          <year>1996</year>
          ),
          <fpage>429</fpage>
          -
          <lpage>460</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Judea</given-names>
            <surname>Pearl</surname>
          </string-name>
          .
          <year>1988</year>
          .
          <article-title>Probabilistic reasoning in intelligent systems: networks of plausible inference</article-title>
          . Morgan Kaufmann.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Lawrence</surname>
            <given-names>R</given-names>
          </string-name>
          <string-name>
            <surname>Rabiner</surname>
          </string-name>
          .
          <year>1989</year>
          .
          <article-title>A tutorial on hidden Markov models and selected applications in speech recognition</article-title>
          .
          <source>Proc. IEEE 77</source>
          ,
          <issue>2</issue>
          (
          <year>1989</year>
          ),
          <fpage>257</fpage>
          -
          <lpage>286</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Tahrima</surname>
            <given-names>Rahman</given-names>
          </string-name>
          , Prasanna Kothalkar, and
          <string-name>
            <given-names>Vibhav</given-names>
            <surname>Gogate</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Cutset networks: A simple, tractable, and scalable approach for improving the accuracy of ChowLiu trees</article-title>
          .
          <source>In Joint European Conference on Machine Learning and Knowledge Discovery in Databases</source>
          . Springer,
          <fpage>630</fpage>
          -
          <lpage>645</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>M.</given-names>
            <surname>Regneri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rohrbach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Wetzel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Thater</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Schiele</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Pinkal</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Grounding Action Descriptions in Videos. Transactions for the Association of the Computational Linguistics (TACL) 1 (</article-title>
          <year>2013</year>
          ),
          <fpage>25</fpage>
          -
          <lpage>36</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Anna</surname>
            <given-names>Rohrbach</given-names>
          </string-name>
          , Marcus Rohrbach, Wei Qiu, Annemarie Friedrich, Manfred Pinkal, and
          <string-name>
            <given-names>Bernt</given-names>
            <surname>Schiele</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Coherent multi-sentence video description with variable level of detail</article-title>
          .
          <source>In German conference on pattern recognition</source>
          . Springer,
          <fpage>184</fpage>
          -
          <lpage>195</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Young</given-names>
            <surname>Chol</surname>
          </string-name>
          <string-name>
            <surname>Song</surname>
          </string-name>
          , Iftekhar Naim, Abdullah Al Mamun,
          <string-name>
            <given-names>Kaustubh</given-names>
            <surname>Kulkarni</surname>
          </string-name>
          , Parag Singla, Jiebo Luo, Daniel Gildea, and Henry A Kautz.
          <year>2016</year>
          .
          <article-title>Unsupervised Alignment of Actions in Video with Text Descriptions.</article-title>
          .
          <source>In IJCAI. 2025-2031.</source>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Christian</given-names>
            <surname>Szegedy</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Googlenet pre-trained model</article-title>
          . http://dl.caffe.berkeleyvision.org/bvlc_googlenet.caffemodel.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Christian</surname>
            <given-names>Szegedy</given-names>
          </string-name>
          , Wei Liu, Yangqing Jia,
          <string-name>
            <given-names>Pierre</given-names>
            <surname>Sermanet</surname>
          </string-name>
          , Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Rabinovich</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Going deeper with convolutions</article-title>
          .
          <source>In Proceedings of the IEEE conference on computer vision and pattern recognition. 1-9.</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>