<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Feeding Machine Learning with Knowledge Graphs for Explainable Ob ject Detection?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Tanguy Pommellet</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Freddy Lecue</string-name>
          <email>freddy.lecue@inria.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Inria</institution>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Thales</institution>
          ,
          <addr-line>CortAIx</addr-line>
          ,
          <country country="CA">Canada</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Machine Learning (ML), as one of the key driver of Arti cial Intelligence, has demonstrated disruptive results in numerous industries. However one of the most fundamental problem of applying ML, and particularly Arti cial Neural Network models, in critical systems is its inability to provide a rational of their decisions. For instance a ML system recognizes an object to be a warfare mine through comparison with its similar observations. No human-transposable rationale is given, mainly because common sense knowledge or reasoning is out-of-scope of ML systems. We developed an asset, combining ML and knowledge graphs to expose a human-like explanation when recognizing an object of any class in a knowledge graph of 4,233,000 resources.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The current hype of Arti cial Intelligence (AI) mostly refers to the success of
Machine Learning (ML) and its sub-domain of deep learning. However industries
operating with critical systems are either highly regulated, or require high level
of certi cation and robustness. Therefore, such industry constraints do limit
the adoption of non deterministic and ML systems. Answers to the question of
explainability will be intrinsically connected to the adoption of AI in industry at
scale. Indeed explanation, which could be used for debugging intelligent systems
or deciding to follow a recommendation in real-time, will increase acceptance
and (business) user trust. Explainable AI (XAI) is now referring to the core
backup for industry to apply AI in products at scale, particularly for industries
operating with critical systems.</p>
      <p>This is particular valid for object detection task in ML, as objet detection is
usually performed from a large portfolio of Arti cial Neural Networks (ANNs)
architectures such as YOLO trained on large amount of labelled data. In such
contexts explaining object detections is rather di cult due to the high
complexity (i.e., number of layers, lters, convolutions phases) of the most accurate
ANNs. Therefore explanations of an object detection task are limited to features
? Copyright c 2019 for this paper by its authors. Use permitted under Creative
Commons License Attribution 4.0 International (CC BY 4.0).
involved in the data and model e.g., saliency maps [1] or at best to examples [2],
or prototypes [3]. They are the best state-of-the-art approaches but explanations
are limited by data frames feeding the ANNs.</p>
      <p>We present a system expanding and linking initial (training, validation and
test) data (for a ML object detection task) with entities in knowledge graphs,
in order (i) to encode context in data, (ii) to capture complex relations among
objects, classes and properties and even (iii) to support inference and causation,
rather than only correlation.
2</p>
      <p>Explaining Objet Detection with Knowledge Graphs
Training
Dataset</p>
      <p>Data: Open</p>
      <p>Image</p>
      <p>Dataset v4
Task</p>
      <p>Object
Detection</p>
      <p>Task</p>
      <p>Image
s
Labels</p>
      <p>Training
Process</p>
      <p>CNN: Faster</p>
      <p>RCNN
Pre-trained inception R
model esnet v2</p>
      <p>Object Detection</p>
      <p>Model</p>
      <p>Knowledge
Graph Selection
Knowledge</p>
      <p>Graphs</p>
      <p>First step
detections
Selected KG,</p>
      <p>Labels
Vocabulary: Wikidata, DBpedia, YAGO.</p>
      <p>50%
Dictionnary
of Context</p>
      <p>Augmentati
on of
'Paddle'
Score</p>
      <p>Semantic
Augmentation
of Confidence</p>
      <p>
        Context
Generation
74
%
Augmented
detections
" 'Paddle' confidence is
augmented as class 'Boat' and
'Canoe’. are in both (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) image
and (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) as properties range of
Paddle in knowledge graph"
      </p>
      <p>Explainable Layer
(ML) Training Process: We selected Faster RCNN (Region CNN, designed
for object detection) Inception Resnet v2 [4]. Due to resources required to train a
neural network on this dataset, we used pre-trained detection model. Among the
pre-trained models on OID v4 available online, the faster RCNN with Inception
Resnet v2 is the best tradeo between detection performance and speed.
(ML) Object Detection Model: The con guration of the faster-RCNN is as
following: The region proposal networks suggest 100 regions, with non maximum
suppression IOU (Intersection-over-union) threshold at :7, and no NMS
(Nonmaximum suppression) hyper-parameters score threshold. The second stage of
the RCNN infers detections for these 100 regions, with no additional NMS. These
100 predictions are the baseline of our work.
(ML) Object Detection Task: Our method is tested on a subset of the Open
Image v4 Validation Dataset (described below). We designed our system to
detect objects among the 600 categories of the Open Image challenge. All those
categories are used to drive the search in Knowledge graph, and extract
contextual information for each one of these categories to augment the baseline object
detection approach.</p>
      <p>Knowledge Graph Selection: Knowledge graphs are selected based on
overage of labels of Open Image v4 Validation Dataset and an initial set of knowledge
graphs: Wikidata, DBpedia, YAGO. Fuzzy match are used to optimize coverage,
and the knowledge graph with the highest coverage is used for the further steps.
Context Generation: The task is to detect and localize objects out of a set C
of 600 categories. For each category in c 2 C we determine a list of other
categories C0 that are `close enough' (i.e., through identi cation of direct link among
entities in the knowledge graph) to c so that if they are detected simultaneously,
then we can con dently increase detections scores. The main di culty is to
identify a unique resource in the graph that will correspond to the given category.
But often, several web resources can represent the same category, and categories
can also have several homonym in the knowledge graph. Towards this challenge
DBpedia is the most e cient graph to extract a unique resource associated to
a category. Indeed, duplicates are quasi-null, disambiguate pages enabling to
di erentiate between homonyms, and redirection property enable to deal with
synonyms issues. Moreover, it has a propriety that redirects every dbpedia
resource toward YAGO and Wikidata. This is extremely interesting as it is hard
to obtain a unique resource from a category name using wikidata itself, and it
can be useful to combine several knowledge graph. The output of this process is
a dictionary, where every key is a label of our detection task, and the value is
the subset of the labels that are contextually linked to it.</p>
      <p>
        Semantic Augmentation of Con dence: We obtain: 100 predictions with
bounding boxes and a contextual dictionary extracted from the knowledge graph.
The process is as following: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) We de ne a rst hyper-parameter (optimized
during training) which is a score threshold st. Detections con dence under this
threshold are not augmented, and cannot contribute to con dence augmentation
of another detection. (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) For each prediction with an initial score higher than
st, we derive a trustworthy indicator: it should indicate if the context
(meaning the other detections on the image) is coherent with the detected category
according to the dictionary extracted from the knowledge graph. We look at
the list of linked label in the contextual dictionary. For each linked label, we
look if it has also been detected in the image (with a con dence score higher
than st). If positive then we add its con dence score to the trustworthy
indicator. Then we compare the trustworthy indicator to a prede ned threshold (new
hyper-parameter). (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) If it is less to the threshold, the initial detection score is
unchanged. The context does not bring more con dence about the detection. (
        <xref ref-type="bibr" rid="ref4">4</xref>
        )
If it is higher, we augment the initial score. To derive the score to add, we
compute the same indicator than in the rst step, but we do not take into account
contribution whose did not reach the trustworthy threshold in the rst step (we
do not want bad predictions to infer in the increase of con dence).
Explainable Layer: During the augmentation process, we record the
predictions that contribute to the increase in con dence of each detection. We obtain
concepts that account for an intelligible and semantic explanation of prediction.
      </p>
      <p>Demonstration
We illustrate our demonstration (Figure 2) on an image from the open image
validation dataset. We feature detections with con dence score higher than :4
before and after the semantic augmentation of con dence.</p>
      <p>Evaluation: We evaluate detection performance with the Open Images mean
Average Precision @0.5 score used for the Open Image detection challenge3. The
Average Precision score for a category is derived as the era under the
Precisionrecall curve for a speci c IoU threshold (here 0.5). Precision measures how
accurate is the predictions. i.e. the percentage of the predictions are correct and
recall measures how good the identi cation of positive is. The mean Average
Precision or mAP score is calculated by taking the mean AP over all classes.
Results: mAP 0:5 = 49:5% with baseline, mAP 0:5 = 49:9% with our approach.
Our semantic augmentation con dence slightly improves the average detection
performance of the model, while providing an interpretable layer due to the use
of external information extracted from knowledge graphs.
3 https://storage.googleapis.com/openimages/web/challenge.html</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <issue>1</issue>
          .
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>C.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Creager</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goldenberg</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Duvenaud</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Interpreting neural network classi cations with variational dropout saliency maps</article-title>
          .
          <source>In: Proc. NIPS</source>
          . Volume
          <volume>1</volume>
          . (
          <year>2017</year>
          )
          <fpage>6</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rudin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Deep learning for case-based reasoning through prototypes: A neural network that explains its predictions</article-title>
          .
          <source>In: Proceedings of the Thirty-Second AAAI Conference on Arti cial Intelligence</source>
          ,
          <source>(AAAI-18)</source>
          , New Orleans, Louisiana, USA, February 2-
          <issue>7</issue>
          ,
          <year>2018</year>
          . (
          <year>2018</year>
          )
          <volume>3530</volume>
          {
          <fpage>3537</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koyejo</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khanna</surname>
          </string-name>
          , R.:
          <article-title>Examples are not enough, learn to criticize! criticism for interpretability</article-title>
          .
          <source>In: Advances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Systems</source>
          <year>2016</year>
          , December 5-
          <issue>10</issue>
          ,
          <year>2016</year>
          , Barcelona,
          <string-name>
            <surname>Spain.</surname>
          </string-name>
          (
          <year>2016</year>
          )
          <volume>2280</volume>
          {
          <fpage>2288</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Girshick</surname>
          </string-name>
          , R.:
          <string-name>
            <surname>Fast</surname>
          </string-name>
          r-cnn.
          <source>In: Proceedings of the IEEE international conference on computer vision</source>
          . (
          <year>2015</year>
          )
          <volume>1440</volume>
          {
          <fpage>1448</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>