<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>CSK-SNIFFER: Commonsense Knowledge for Snifing Object Detection Errors</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Anurag Garg</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Niket Tandon</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aparna S. Varde</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Allen Institute for AI</institution>
          ,
          <addr-line>Seattle</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Montclair State University</institution>
          ,
          <addr-line>Montclair</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>PQRS Research</institution>
          ,
          <addr-line>Dehradun</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper showcases the demonstration of a system called CSK-SNIFFER to automatically predict failures of an object detection model on images in big data sets from a target domain by identifying errors based on commonsense knowledge. CSK-SNIFFER can be an assistant to a human (as snifer dogs are assistants to police searching for problems at airports). To cut through the clutter after deployment, this “snifer” identifies where a given model is probably wrong. Alerted thus, users can visually explore within our demo, the model's explanation based on spatial correlations that make no sense. In other words, it is impossible for a human without the help of a snifer to flag false positives in such large data sets without knowing ground truth (unknown earlier since it is found after deployment). CSK-SNIFFER spans human-AI collaboration. The AI role is harnessed via embedding commonsense knowledge in the system; while an important human part is played by domain experts providing labeled data for training (besides human commonsense deployed by AI). Another highly significant aspect is that the human-in-the-loop can improve the AI system by the feedback it receives from visualizing object detection errors, while the AI provides actual assistance to the human in object detection. CSK-SNIFFER exemplifies visualization in big data analytics through spatial commonsense and a visually rich demo with numerous complex images from target domains. This paper provides excerpts of the CSK-SNIFFER system demo with a synopsis of its approach and experiments.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Bbox2 car
Trained
object
detection
model M
Human-AI collaboration, the realm of humans and AI bbox1 bbox3 H
systems working together, typically achieves better per- collaborate
formance than either one working alone [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Big data
visualization and analytics can be used to foster interac- bPbrex1dicbtiboxn2 bbx3
tion [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Such areas receive attention, e.g. NEIL (Never
Ending Image Learner) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], active learning approaches CSK Snif er C Update
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], human-in-the-loop learning [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] etc. To that end, iImnpaugte x flag spatially
we demonstrate a system “CSK-SNIFFER” exemplifying implausible
bhyu mviasnu-aAlizIicnogllpaobtoernattiiaolnervrioaresninhalanrcginegcoombjpelcetxddeatteactsieotns, object pairs Inference rrkeenltoeriwveavlenedtogsbepjaetciatl ㎅
harnessing spatial commonsense. This system “snifs”
errors in object detection using spatial collocation anoma- Figure 1: CSK-SNIFFER and the human-in-the-loop: A car
lies, assisting humans analogous to snifer dogs aiding was detected in the image which was flagged bad by
CSKpolice at airports. The process, (Figure 1), is as follows, SNIFFER based on its spatial knowledge w.r.t KB which the
with the CSK-SNIFFER system (), human-in-the-loop human-in-the-loop can update after visualizing errors
(), and inference model ( ) for object detection.
      </p>
      <p>System  interacts with human  and provides object
detection output visualizing potential errors, over model sion and recall, based on the spatial KB. Hence, the two
 , by deploying commonsense knowledge through a directions in this learning loop are as follows.
spatial knowledge base (). The  is derived by
capturing spatial commonsense, especially as collocation •  to : Feedback-based interactive learning
anomalies. Then  sees the visualized errors (output by •  to : Assistance in object detection
) and can thereby enhance  by increasing its
precivisualized 
bbx1
bbx2
bbx3
Thus, the human and the AI work together with the goal
of enhancing object detection in big data. Further, the
inference model  can potentially improve, as an added
benefit of this adversarial learning via human-AI
collaboration. The obtained information can be used to supply
more examples to  on the misclassified categories to
make it more robust. If certain labels are inappropriate
consistently across examples, it is a valuable insight. As
data sets get bigger in volume and variety, such
automation is even more significant in assisting object detection
errors.</p>
      <sec id="sec-1-1">
        <title>The use case in our work focuses on the smart mo</title>
        <p>
          bility domain [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. It entails autonomous vehicles,
selfoperating trafic signals, energy-saving street lights
dimming / brightening as per pedestrian usage etc. In such
        </p>
      </sec>
      <sec id="sec-1-2">
        <title>AI systems, it is crucial to detect objects accurately, especially due to issues such as safety. CSK-SNIFFER plays an important role here, generating large adversarial training data sets by snifing object detection errors.</title>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. The CSK-SNIFFER Approach</title>
      <p>People crossing city streets on pedestrian crossings
Vehicles coming to a full hault at red signals
Vehicles stopping or slowing down at stop signs
Street lights dimming when occupants are few
Street lights brightening when occupants are many
Buses running on trafic-optimal routes
Service dogs helping blind people
People charging phones at WiFi stations
People reading useful information at roadside kiosks
People parking bikes at share-ride spots
Vehicles flashing turning lights for L/R turns
Bikes riding on bike routes only
Trafic cops making hand signals in regular operations
Vehicles driving beneath an overpass
Dogs on a leash walking with their owners
People jogging on sidewalks
People entering and leaving trains when doors open
People using prams for kids in buses
Trees existing on sidewalks
Ropeways carrying passengers to tourist spots
Bikers wearing smart watches
Maglev trains running between airports and cities
Grass existing on freeway sides and city streets
Solar panels existing on roofs of buildings
People using smartphones for talking anytime anywhere
Canal lights dimming when occupants are few
Canal lights brightening with many occupants
People wheeling shopping carts in grocery stores</p>
      <sec id="sec-2-1">
        <title>We summarize the CSK-SNIFFER approach as per its</title>
        <p>
          design and execution [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. In this approach. we represent
and construct a , along with the function  ()
that generates a triple, i.e. &lt; , ˆ ,  &gt; from the
predicted bounding boxes of objects  and  such that
 is a binary vector over relations ().
        </p>
        <p>Gather  using action-vocab( ): Table 1
presents some examples from action-vocab( ) where
 =smart mobility domain. These entries represent
mostly typical and some unique scenes in this
domain. This list was manually compiled by a domain Table 1
expert. Images  are compiled using Web queries ∈ ∼ 10% examples from action-vocab( ) where  =smart
action-vocab( ) (on an image search engine), and an mobility domain. These are used as queries to compile the
object detector predicts bounding boxes over  ∈  . input to the object detector, and then CSK-SNIFFER can flag</p>
        <p>
          KB construction: While in principle, we can directly images in  where the detector failed to predict the correct
use existing s, these have errors as elaborated in bounding boxes.
some works [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. CSK-SNIFFER isolates the efect of these
errors by instead manually creating a  at a very low
cost. The  is defined over a set of objects  and rela- available at https://tinyurl.com/kb-for-csksnifer
tions (). The relation set () comprises Function f(bbox): Similar to the triples in the ,
5 relations (isAbove, isBelow, isInside, isNear, we define a function  () to construct triples &lt;
overlapsWith). We are inspired by other works in the , ˆ ,  &gt; using the predicted bounding boxes of
imliterature such as [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] in picking these relations, and be- age  . The  () input consists of predicted
boundcause our initial analysis proposed their suitability for ing boxes on an image, and the output is a list of triples
bounding box relative relations. An entry in the  in the format: &lt; , ˆ ,  &gt;, for every pair of objects
comprises of a pair of objects ,  ∈  and a binary ,  ∈ the objects detected in the image. For every such
vector  denoting ,  ’s and their spatial relations pair,  () compares the coordinates of the bounding
over (). These spatial relations are manually an- boxes of  and  (this is a known process e.g. [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]). We
notated by a domain expert, according to general likeli- illustrate this for the isInside relation. Let coordinates
hood e.g., it is more likely that a dog is observed near a of a bounding box be 1, 1, 2, 2, then .1 denotes
human, and much less likely that it is observed near a 1 coordinate of . If .2 ≤  .1 and  .2 &gt; .1
whale. The  most popular objects on MSCOCO training and  .1 &lt; .2 then  is inside  . Similarly,
data make up ; in our experiments  = 10, and this other relations in () are built, compiling which
leads to 2 entries in  that need to be annotated with provides ˆ . For anomaly detection, we compare vectors
 . It is remarkable that our experiments demonstrate ˆ and  for overlapping object-pairs ,  , detected
that even with  = 10, the  allows CSK-SNIFFER to in the image and present in the .
achieve good performance. We can infer that selecting Based on this discussion, the following algorithm
a popular subset helps, even if it is small. An entry in summarizes the execution of CSK-SNIFFER.
the KB is denoted as &lt; ,  ,  &gt; The  is publicly
        </p>
        <p>Algorithm 1: CSK-SNIFFER Approach</p>
      </sec>
      <sec id="sec-2-2">
        <title>Input: Object detector  trained on source domain</title>
      </sec>
      <sec id="sec-2-3">
        <title>Manually compiled action-vocab( ) in target domain</title>
        <p>Images  compiled using Web queries ∈
action-vocab( )
1. Define () comprising 5 relations:
isAbove, isBelow,
isInside, isNear, overlapsWith
2. Define commonsense , each entry &lt;
,  ,  &gt; where  is a binary vector over
()
3. Generate triples &lt; , ˆ ,  &gt; from predicted
bounding boxes of  ∈  using function
 ().
4. For each image  ∈  , compare ˆ and  ,
from bounding box triples &lt; , ˆ ,  &gt; and
 triples &lt; ,  ,  &gt;
5. For each  , if ˆ ̸=  then flag  : ,
add  to ′</p>
        <p>Output: Subset ′ where  failed</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Excerpts from System Demo</title>
      <sec id="sec-3-1">
        <title>We have built a live demo to depict the working of CSK</title>
      </sec>
      <sec id="sec-3-2">
        <title>SNIFFER. This demo illustrates the functioning of CSK</title>
        <p>
          SNIFFER to enhance its actual comprehension and aug- Table 2
ment its usage. In addition, this demo paper presents Distribution of spatial relations in  . Spatial relations are
the principles behind the human-in-the-loop functioning of the form: &lt; , ^ ,  &gt;
of CSK-SNIFFER for snifing object detection errors in
large, complex data sets, thereby being added
contributions over our earlier work [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. While this human-in-the- predicted, the demo moves to the home page. This
conloop functioning is explained in the introduction with tains details on the output files generated. The “Images”
an illustration and theoretical justification, its detailed option displays downloaded images as shown in Figure
empirical validations with respect to interactive  up- 2 herewith.
dates constitute ongoing work, based on CSK-SNIFFER Output files generated by CSK-SNIFFER are illustrated
being actively deployed in real-world settings. In fact, as follows. Table 2 shows the first output file
“Collocathis demo paves the way for such interactive  updates tions Map” with triples predicted by CSK-SNIFFER in
via augmenting the usage of CSK-SNIFFER in suitable images with their respective counts. The final output
applications to provide the human-in-the-loop feedback ifle “Error Set” Table 3 contains names of images with
for the addition of such interactive  updates. some odd visual collocations. It also indicates the triple
        </p>
        <p>
          We present some screenshots illustrating the demo. that actually got predicted versus the expectation from
Many more can be provided in a live setting. The user the model. These files help fathom the functioning of
enters any search query related to smart mobility [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. CSK-SNIFFER.
        </p>
        <p>
          Images are downloaded from Google Images based on
this query. Object detection is then performed on the
images using YOLO [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] to start predicting triples in the
image using the  () function. Once the triples are
person, overlapsWith, person
car, is_near, car
person, overlapsWith, car
car, overlapsWith, person
car, overlapsWith, car
person, is_near, backpack
backpack, is_near, person
car, is_near, backpack
backpack, is_near, car
trafic light, is_near, trafic light
each other, such that the AI (CSK-SNIFFER) provides
Table 3 a visual demo of the object detection errors snifed by
Canonical examples of errors flagged by CSK-SNIFFER If in- spatial CSK, thus generating large adversarial data sets
ferred spatial relations over model-generated bounding boxes to assist object detection, while the human can use this
are not consistent with expected spatial relations between ob- feedback to enhance the performance of CSK-SNIFFER,
jects, then predicted bounding boxes are flagged as erroneous. thereby playing its role in the learning loop.
4.2. Error Analysis
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Experimental Evaluation</title>
      <sec id="sec-4-1">
        <title>We now present the precision and recall shortcomings.</title>
        <p>We present examples from our experiments, showing the Recall issues: Actually bad, flagged good: Figure 5
correct and wrong predictions made by CSK-SNIFFER, depicts examples of these types of images. The reason for
along with the error analysis. Here, “bad” refers to images CSK-SNIFFER predicting these images as good instead
containing object detection errors while “good” refers to of bad is that the objects wrongly detected in the image
correctly identified images with no such errors. are not present in our , hence it does not check for
their locations. Thus, they are not found in any of the
4.1. Appropriate Identifications triples predicted, they are skipped so that they do not
make their way to the error set.</p>
        <p>Actually bad, flagged bad: Figure 3 illustrates examples Precision issues: Actually good, flagged bad: Our
in this category. Experimental evaluation shows that our investigation of the source of these errors (see Figure
model is good at identifying odd bounding boxes. 6), concluded that the  relations are authored with a</p>
        <p>Actually good, flagged good: Figure 4 portrays ex- 3D space in perspective, while the images only contain
amples of this type. Experimental evaluation shows that 2 information. Therefore, relations such as above and
CSK-SNIFFER is able to distinguish good predictions. below may be confused with farther and nearer. For</p>
        <p>Other benefits: Interestingly, while analyzing mis- example, if a car is detected in the background,  ()
takes of CSK-SNIFFER , we find that ∼ 10% of the refer- function make an incorrect interpretation as car is
ence data on which  is trained (MSCOCO, expected to above person. The  will flag this as unlikely and
be a high quality), contains wrong bounding boxes. This hence an erroneous detection, leading to a possibly good
provides insights into potentially improving MSCOCO, prediction flagged as an error.
constituting an added benefit of this work. Addressing 2D vs. 3D errors: We calculate the area</p>
        <p>On the whole, the human and the AI collaborate with covered by a bounding box, such that if the area is less
than an empirically estimated threshold that object is
considered to be detected in the background and
therefore does not predict the triple, e.g. car above person
in that image. This helps to increase accuracy to ∼ 80%.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions and Roadmap</title>
      <p>This paper synopsizes the demo (with approach and
experiments) of a system “CSK-SNIFFER” that “snifs”
object detection errors in a big data on an unseen target
domain using spatial commonsense, with high accuracy
at no additional annotation cost. Based on human-AI
collaboration, the AI angle entails spatial CSK imbibed in
the system deployed via visual analytics to assist humans,
while an important human role comes from the domain
expert perspective in image tagging and task
identification for training the system (in addition to the obvious
human contribution of commonsense knowledge in the
system). More significantly, the human and the AI make
contributions to the learning loop by feedback-based
interactive learning, and assistance in object detection
respectively. It is promising to note that our approach
based on simplicity can automatically discover errors in
data of significant volume and variety, and be potentially
useful in this learning setting. We demonstrate that with
high quality, we can generate large complex adversarial
datasets on target domains such as smart mobility.</p>
      <sec id="sec-5-1">
        <title>Future work includes harnessing existing, poten</title>
        <p>tially noisy and incomplete commonsense  in
CSK</p>
      </sec>
      <sec id="sec-5-2">
        <title>SNIFFER Another direction is to study whether auto</title>
        <p>matic adversarial datasets compiled with assistance from</p>
      </sec>
      <sec id="sec-5-3">
        <title>CSK-SNIFFER help train better models on novel target domains. Our work presents interesting facets from big data visualization and analytics along with human-AI collaboration.</title>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Acknowledgments</title>
      <sec id="sec-6-1">
        <title>A. Varde has NSF grants 2018575 (MRI: Acquisition of</title>
        <p>a High-Performance GPU Cluster for Research &amp;
Education); 2117308 (MRI: Acquisition of a Multimodal</p>
      </sec>
      <sec id="sec-6-2">
        <title>Collaborative Robot System (MCROS) to Support Cross</title>
      </sec>
      <sec id="sec-6-3">
        <title>Disciplinary Human-Centered Research &amp; Education).</title>
      </sec>
      <sec id="sec-6-4">
        <title>She is a visiting researcher at Max Planck Institute for</title>
      </sec>
      <sec id="sec-6-5">
        <title>Informatics, Germany.</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>D.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Churchill</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Maes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Shneiderman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>From human-human collaboration to human-ai collaboration: Designing ai systems that can work together with people</article-title>
          ,
          <source>in: CHI</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>F.</given-names>
            <surname>Zuo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ozbay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X. J.</given-names>
            <surname>Ban</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Iyer</surname>
          </string-name>
          ,
          <article-title>An interactive data visualization and analytics tool to evaluate mobility and sociability trends during covid-19</article-title>
          , arXiv:
          <year>2006</year>
          .
          <volume>14882</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>X.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Shrivastava</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gupta</surname>
          </string-name>
          , Neil:
          <article-title>Extracting visual knowledge from web data</article-title>
          , in: ICCV,
          <year>2013</year>
          , pp.
          <fpage>1409</fpage>
          -
          <lpage>1416</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>K.</given-names>
            <surname>Konyushkova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Sznitman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Fua</surname>
          </string-name>
          ,
          <article-title>Learning active learning from data</article-title>
          ,
          <source>Advances in Neural Information Processing Systems</source>
          <volume>30</volume>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>D.</given-names>
            <surname>Xin</surname>
          </string-name>
          , L. Ma, J. Liu,
          <string-name>
            <given-names>S.</given-names>
            <surname>Macke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Parameswaran</surname>
          </string-name>
          ,
          <article-title>Accelerating human-in-the-loop machine learning: Challenges and opportunities</article-title>
          ,
          <source>in: ACM SIGMOD (DEEM workshop)</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>4</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Orlowski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Romanowska</surname>
          </string-name>
          ,
          <article-title>Smart cities concept: Smart mobility indicator, Cybernetics and Systems (Taylor</article-title>
          &amp; Francis)
          <volume>50</volume>
          (
          <year>2019</year>
          )
          <fpage>118</fpage>
          -
          <lpage>131</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Garg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Tandon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. S.</given-names>
            <surname>Varde</surname>
          </string-name>
          ,
          <article-title>I am guessing you can't recognize this: Generating adversarial images for object detection using spatial commonsense</article-title>
          ,
          <source>in: AAAI</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>13789</fpage>
          -
          <lpage>13790</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>N.</given-names>
            <surname>Tandon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. S.</given-names>
            <surname>Varde</surname>
          </string-name>
          , G. de Melo,
          <article-title>Commonsense knowledge in machine intelligence</article-title>
          ,
          <source>ACM SIGMOD Record</source>
          <volume>46</volume>
          (
          <year>2017</year>
          )
          <fpage>49</fpage>
          -
          <lpage>52</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Yatskar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Ordonez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Farhadi</surname>
          </string-name>
          ,
          <article-title>Stating the obvious: Extracting visual common sense</article-title>
          ,
          <source>NAACL</source>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J.</given-names>
            <surname>Redmon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Farhadi</surname>
          </string-name>
          , Yolo9000: Better, faster, stronger,
          <source>CVPR</source>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>