<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>An Adaptive Semantic Stream Reasoning Framework for Deep Neural Networks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Danh Le-Phuoc</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thomas Eiter</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Technical University Berlin</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Technical University Vienna</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Driven by deep neural networks (DNN), the recent development of computer vision makes visual sensors such as stereo cameras and Lidars ubiquitous in autonomous cars, robotics and trafic monitoring. However, due to operational constraints, a processing pipeline like object tracking has to hard-wire an engineered set of DNN models to a fixed processing logic. To overcome this, we propose a novel semantic reasoning approach that uses stream reasoning programs for in-cooperating DNN models with commonsense and domain knowledge using Answer Set Programming (ASP). This approach enables us to realize a reasoning framework that can adaptively reconfigure the reasoning plan in each execution step of incoming stream data.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Semantic Reasoning</kwd>
        <kwd>Neural-symoblic</kwd>
        <kwd>Stream Reasoning</kwd>
        <kwd>Stream Processing</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Motivation</title>
      <p>ple, the recent report of Uber’s accident in Arizona [4]
says "...The ADS detected the pedestrian 5.6 seconds
beThe recent development of computer vision (CV) driven fore impact. Although the ADS continued to track the
by deep neural networks (DNN) makes visual sensors pedestrian until the crash, it never accurately classified
such as stereo cameras and Lidars ubiquitous in au- her as a pedestrian or predicted her path. By the time the
tonomous cars, robotics and trafic monitoring. In par- ADS determined that a collision was imminent (1 second
ticular, many DNN models for object detection [1] and before impact), the situation exceeded the response
spectracking [2] are available. However making reliably ifications of the ADS braking system..." . This accident
working in a real-life online processing pipeline such could have not happened if the ADS could
reconfigas an Automated Driving System (ADS) or a trafic ure the object tracking pipeline on the fly, e.g
changsurveillance query engine is still very challenging. For ing DNN models or using alternative sensor sources to
example, [2] reports that the most accurate DNN-driven improve the accuracy on detection and tracking.
multi-object tracking (MOT) pipelines can process only This motivates us to propose an approach of
com4-5 frames/second. To make such a system work on- bining stream reasoning with probabilistic inference
line e.g. for ADS, where processing delay must be less to continuously configure such processing pipelines
than 100ms [3], one has to hard-wire a fixed sets of based in semantic information representing
commonDNN models with some sacrifices on accuracy and ro- sense and domain knowledge. The use of semantic
inbustness as the design constraints of an ADS limit how formation together with DNNs has proved to be useful
much hardware can be put into a system [3]. For in- and led to better accuracy in image understanding [5]
stance, an additional 400 W power consumption trans- and in object tracking [6]. Similar to ours, these
aplates to a 3.23% reduction in miles per gallon for a 2017 proaches use declarative approaches to represent the
Audi A4 sedan or similarly, the additional power con- processing pipelines of visual data. However, none of
sumption will reduce the total driving range of electric them have considered how to deal with the
aforemenvehicles. tioned operational constraints in the context of stream</p>
      <p>Such a design-time trade-of often leads to unpre- processing. Our approach represents such constraints
dictable fatal errors in a real deployment. For exam- in an extension of Answer Set Programming (ASP).
This extension is proposed by leveraging LARS
formuProceedings of the CIKM 2020 Workshops, October 19-20, Galway, las [7] for expressing stream reasoning programmes to
Ireland. incorporate uncertainty of probabilistic inference
optehmoamila:sd.eaintehr.@leptuhwuoiecn@.atcu.a-bte(Trl.inE.idteer()D. Le-Phuoc); erations under weighted rules similar to LP [8],
orcid: 0000-0003-2480-9261 (D. Le-Phuoc); 0000-0001-6003-6345 called semantic reasoning rules. As a result, we will
(T. Eiter) be able to dynamically express a visual sensor fusion
pipeline, e.g. MOT over multiple cameras, by
seman© 2020 Copyright for this paper by its authors. Use permitted under Creative
CPWrEooUrckReshdoinpgs IhStpN:/c1e6u1r3-w-0s.o7r3g CCoEmUmoRns WLiceonrsekAsthtriobuptioPnr4o.0cIneteerdnaitniognasl ((CCC EBYU4R.0)-.WS.org)
b1
…
1
Semantically, a multi-sensor fusion data pipeline will to evaluate a formula 
at a time point</p>
      <p>(resp.
ev
work Ontology (SSN) [10] . For instance, the example 
in Figure 1 shows 3 image frames are observed by a that at evaluation time  ,</p>
      <p>(, , 
. For example, the formula ⊞+5◊
following standardized W3C/OGC Semantic Sensor Net- shot (substream)  ′ from
 by applying a function</p>
      <p>on
) holds at some
tic reasoning rules to fuse probabilistic inference
operations with ASP-based solving processes. Moreover,
the expressive power of our approach also enables us
to express operational constraints together with
optimisation goals as a probabilistic planning programme</p>
      <p>+ [9] so that our reasoning framework
similar to 
can reconfigure the reasoning plan adaptively in each
execution step of incoming stream data.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Formalization of Semantic</title>
    </sec>
    <sec id="sec-3">
      <title>Stream</title>
    </sec>
    <sec id="sec-4">
      <title>Reasoning with DNN</title>
    </sec>
    <sec id="sec-5">
      <title>Models</title>
      <p>consume the data that is observed by a Sensor as a
stream of observations (represented as an Observation)</p>
      <p>The symbol FeatureOfInterest (FoI) is used to
represent the domain of physical objects which are subjects
for the sensor observations, e.g. tracking objects and
ifeld of views (FoV) of the camera. The relationship
between a Result generated by probabilistic inference
algorithms (e.g. YOLO detection model or Kalman filter
algorithm) to such object is represented by the
predicate isSampleOf (denoted as iSO). As such algorithms
have output with uncertainty, we will use an
abduction reasoning process to search for explainable
evidences for iSO via rules driven commonsense and
domain knowledge similar to [6].</p>
      <p>To formalise the reasoning process with such a
semantic representation of stream data, we need a
temporal model that allows us to reason about the
properties and features of objects from streams of sensor
observations. This model must account for the laws of
the physical world movement and in particular be able
to fill gaps of incomplete information (e.g., if we do
not see objects appearing in observations, or camera
trafic camera. These observations will then be fed into
a probabilistic inference process such as a DNN model 
or a CV algorithm (represented as a Procedure) to
provide derived stream elements which then are
representing Sampling instances. In this example, we have

(, 
1,</p>
      <p>) representing for the output bounding
box  1 from the YOLO detector 
∈</p>
      <p>where 
for Detector) represents for the set of detectors
supported. Similarly, 
bounding box  2
generated by a tracking algorithm (a</p>
      <p>(39,  2) represents for tracking
the popular object tracking algorithm SORT [11].</p>
      <p>Tracker) which associates  2 with the tracklet 39 via rules under LP
(short
≡ ℎ 
(,  )
reads are missing), based on commonsense principles.</p>
      <p>We thus use the LARS framework [7] to represent a
reasoning programme over our semantic stream data
which can evaluated using an ASP solver.</p>
      <p>The LARS framework provides formulas with Boolean
connectives and temporal operators @  , □ , and ◊
ery, some time point) in the current stream 
dow operators ⊞  take for evaluating  a data
snap</p>
      <p>;
win(, , 
(, , 
) states
) is the
time point  ′ in the window [ − 5, ..,  ] selected by
= +5; in our representation,</p>
      <p />
      <p>≡ ℎ
matching condition for "a car was detected in
bounding box  by the YOLO detector". The formula @ 
is aligned with a fluent  resp. an event  in Event</p>
      <p>( ,  ) and
. This will help us to employ
ASP-based EC axioms for common sense reasoning rules
as proposed in [6].</p>
      <p>To deal with uncertainty, we extend LARS with weighted
semantics as in [8]. In LP
facts generated from feature extractors (DNN models
or OpenCV algorithms) as well as abduction rules on
top can be annotated with uncertainty information given
by a weight, which allows reasoning under certain
levels of uncertainty.
a set of weighted rules  of the form</p>
      <p>For our concerns, a semantic reasoning program Π is
 ∶  ←</p>
      <p>(1)
where  ,</p>
      <p>are LARS formulas and  ∈ ℝ ∪ { } is the
weight of the rule. If  = 
erwise a soft rule; by Π
ℍ and Π
, then  is a hard rule,
oth</p>
      <p>we denote the sets
weights and each
of hard and soft rules of Π, respectively. The
semantics of Π is given by the answer streams [7]  of the
LARS program Π obtained from Π by dropping the

where  violates 
←</p>
      <p>; each
such  gets a probability</p>
      <p>Π( ) assigned calculated
from the weights of the rules retained for Π ; for more
information, we refer to [8]. In Section 3 we will
address how to translate restricted programs Π into ASP
(e.g. pedestrians or vehicles) on a labelled re-identification Figure 2: Stream Reasoning Framework
nected tracklets that have two bounding boxes matched in Figure 2. The key components Reasoner and Planner
pearance descriptors for each track. Based on this gallery modules. The control logic of the framework is
govprograms to be fed into an ASP solver, and how the
weights  can be learned from training data.</p>
      <p>To demonstrate how to build a semantic reasoning
program, we will emulate DeepSORT tracking
algorithm [13] via soft rules that can search for supporting
evidences to re-identify objects associated with
tracklets created by Kalman filter above, by using visual
appearance associations. DeepSORT extends SORT with
a DNN that is trained to discriminate targeted objects
dataset.</p>
      <p>Hence, we will search for pairs of bounding boxes of
two similar tracklets w.r.t. visual appearance. Due to
a large search space of possible matches, we will limit
the search space by filtering the candidates based on
their temporal and spatial properties. Therefore, we
use rules with windows to reason about two
disconwithin a time window of  
aligned with the DeepSORT’s gallery of associated
ap</p>
      <p>time points which are
of previous tracked boxes, the cosine distance is
computed to compare appearance information that are
particularly useful to recover identities after long term
occlusions, when motion is less discriminative. Hence,
for merging two adjacent tracklets that have visual
appearance matches, we use the parametrized soft rule
(15) below. The pair of parameters (  , 
specified for each reasoning step via the probabilistic</p>
      <p>) has to be
planning component of our dynamic reasoning
framework in Section 3; 
matching models that represent the association
metrics to discriminate comparing bounding boxes.</p>
      <p>is one of the available visual
 1 ∶ 
( 1,  ) ← @  ( 1,  1), 
( 1,  2),</p>
      <p>( 2,  ),
 &lt;   + 3, ℎ
+ 
( 
◊</p>
      <p>( 2,  2),
,  1,  2)</p>
      <p>(2)
ated.   
FoV of camera  ".</p>
      <p>(, 
tracklets from two adjacent cameras. We use</p>
      <p>( 1, , 
the candidate camera</p>
      <p>at time point  to start the
search for the matches via the auxiliary predicate 
Also,  is filtered by the auxiliary predicate 
ing  is adjacent to the camera where  1
was
gener</p>
      <p>stat ) to specify the time diference</p>
      <p>from
) represents for "object  left the

( 1, , 
( 2,  ), 
( 1,  1), 

( 2,  2), ℎ
+ 
( 1,  ),
◊  
( 
(,  ),
,  1,  2)
Similarly, we can define rules to trigger the object
matching search based on visual appearances of two that do not occur in rule heads.</p>
      <p>D
N
N
m
o
d
e
l
s</p>
      <p>K
B
genSnapshot s
a
o
genPlanPf</p>
      <p>solve
lars2asp
solve</p>
      <p>A
S
P
S
o
l
v
e
r</p>
    </sec>
    <sec id="sec-6">
      <title>3. Dynamic Reasoning</title>
    </sec>
    <sec id="sec-7">
      <title>Framework</title>
      <p>To the realize our reasoning approach in Secion 2, we
proposes a dynamic reasoning framework illustrated
of the framework are built on top an ASP Solver and
a Stream Processor which are pluggable and generic
erned by Algorithm 1.</p>
      <p>While any ASP Solver supporting weak constraints
can be used in our framework, existing stream
processors such as relational or graph stream processing
engines need to be extended with some prerequisite
features to connect with the rest of the framework. For
instance, in our under-development prototype, we
extend CQELS [14] to enable DNN inference on GPUs as
built-in functions for CQELS-QL, the graph-based
continuous query language of CQELS. Via CQELS-QL, the
auxiliary predicates such as , 
tion 2 are expressed as continuous queries in order to
delegate the processing to the stream processor. This
mechanism also helps us to avoid grounding overhead
in continuous solving via materialised views similar to
the over-grounding approach in [15]. In particular, we
leverage continuous multiway-joins with windows of
CQELS to delegate the processing of LARS formulas</p>
      <p>in
Secand</p>
      <p>Moreover, the visual stream data with diferent
formats (e.g. RGB videos and Lidar PointCloud) together
with the knowledge base (KB) have to be normalized
. to the data model supported by the stream processor.</p>
      <p>For example, ontologies and metadata and extracted
symbols as outputs of DNN inference processes are
represented as temporal graphs of CQELS. With these
features, the stream processor is able to generate data
snapshots and planning profiles in ASP readable
for(3)
spectively.
mat via two methods</p>
      <p>The Planner calls the method   
the input for the first reasoning step of each time point
and ℎ</p>
      <p>reto prepare
point 
Output: Optimal answer set  ∗
1:  ← 0, Π̂ ← ∅
2: for { ∶  ←  } ∈ Π do
 ←  + 1
̂
Π
 ← Π̂</p>
      <p>∪ {
̂
Π
 ← Π</p>
      <p>̂ ∪ { ← ,  
̂
Π</p>
      <p>← Π̂ ∪ {∶∼ 
, whose optimal models (answers sets)  
cor  ∶Π̂ ( )⊩ 
argmax   Π̂ (Π |,  , Π )</p>
      <p>(5)
calls the method ℎ
With a chosen plan embedded in some   , the Reasoner</p>
      <p>to generate the input
data of the reasoning program for the next step in line
(11) to carry out the second reasoning step from line
(12) to (14) to generate the output of the whole pipeline
as the optimal model  ∗. To specify the weights of the
soft rules, we use the weight learning approach of [17]
which fits the weights using the training data via
gradient ascent. Training is done ofline but uses
Algorithm 1 to compute an optimal stable model in each
step of updating weights in the gradient method.
( ) ← ,</p>
      <p>}
( )}
 which finds the optimal reasoning plan to achieve a
certain goal Π under an operational constraint Π ,
following a probabilistic planning approach from [9]
in lines (8) to (10) of Algorithm 1. The reasoning
problem formalised in following is to find a configuration
(</p>
      <p>, 
tracking goal. For example, we can specify the goal of</p>
      <p>,   ) to result the highest probability of our
being able to track the objects that were tracked in the
previous time point  − 1 as
Π
(, _) @ 
(, _</p>
      <p>)
ASP rules. For instance, the example rule below
represents the constraint to limit executable plans
and visual a matching model 
the execution time of a candidate detection model</p>
      <p>together a candidate
and 
window parameter   . The auxiliary predicates</p>
      <p>provide the time estimation for
corresponding DNN operations and the number of objects tracked
in camera  .
( ,   ), @</p>
      <p>(,  , 
(,</p>
      <p>),   +  ∗   ∗   &gt;  
From Π , lines (8) to (10) carry out solving the
reasoning problem formalised by formula (4). To generate the
LARS program from the soft rules Π , the algorithm
rewrites Π into LARS formulas with weak constraints
as Π̂ in lines (1) to (7) by extending the similar
algorithm for LP</p>
      <p>in [8]. Then, we use the incremental
ASP encoding algorithm of Ticker [16] for rewriting</p>
      <p>), (4)
Similarly, an operational constraint Π can be expressed to exploit the code bases of CQELS and Ticker. We

(,  , 
) at time point  based on the estimations as the DNN Inference engine. The solving and
infer</p>
    </sec>
    <sec id="sec-8">
      <title>4. Conclusion</title>
      <p>This position paper presented a novel semantic
reasoning approach that enables probabilistic planning
to adaptively optimize the sensor fusion pipeline
under operational constraints expressed in ASP. The
approach is realised with a dynamic reasoning
mechanism that can integrate the uncertainy of DNN
inference with semantic information, e.g common sense and
domain knowledge in conjunction with runtime
information as inputs for operational constraints.</p>
      <p>We are currently implementing an open sourced
prototype of the proposed reasoning framework in Java
use the Java native interface to wrap C/C++ libaries of
Clingo 5.4.0 as the ASP Solver and NVidia CUDA 10.2
ence tasks are coordinated in an asynchronous
multithreading fashion to exploit the massive parallel
capabilities of CPUs and GPUs.</p>
    </sec>
    <sec id="sec-9">
      <title>Acknowledgments</title>
      <p>This work was funded in part by the German Ministry
for Education and Research as BIFOLD - Berlin
Institute for the Foundations of Learning and Data (refs
01IS18025A and 01IS18037A) and supported by the
Austrian Federal Ministry of Transport, Innovation and
Technology (BMVIT) under the program ICT of the
Future (FFG-PNr.: 861263, project DynaCon).
[13] N. Wojke, A. Bewley, D. Paulus, Simple online
and realtime tracking with a deep association
[1] L. Liu, W. Ouyang, X. Wang, P. Fieguth, J. Chen, metric, in: ICIP, 2017.</p>
      <p>X. Liu, M. Pietikäinen, Deep learning for generic [14] D. Le-Phuoc, M. Dao-Tran, J. X. Parreira,
object detection: A survey, IJCV (2019). M. Hauswirth, A native and adaptive approach
[2] G. Ciaparrone, F. L. Sánchez, S. Tabik, L. Troiano, for unified processing of linked streams and
R. Tagliaferri, F. Herrera, Deep learning in video linked data, in: ISWC, 2011, pp. 370–388.
multi-object tracking: A survey, Neurocom- [15] F. Calimeri, G. Ianni, F. Pacenza, S. Perri, J.
Zanputing (2019). URL: http://www.sciencedirect. gari, Incremental answer set programming with
com/science/article/pii/S0925231219315966. overgrounding, TPLP 19 (2019).
doi:https://doi.org/10.1016/j.neucom. [16] H. Beck, T. Eiter, C. Folie, Ticker:
2019.11.023. A system for incremental asp-based
[3] S.-C. Lin, Y. Zhang, C.-H. Hsu, M. Skach, M. E. stream reasoning, TPLP (2017). URL:
Haque, L. Tang, J. Mars, The architectural impli- https://doi.org/10.1017/S1471068417000370.
cations of autonomous driving: Constraints and doi:10.1017/S1471068417000370.
acceleration, in: ASPLOS ’18, 2018. [17] J. Lee, Y. Wang, Weight learning in a probabilistic
[4] NTSB, Collision between vehicle controlled extension of answer set programs, in: KR, 2018,
by developmental automated driving sys- pp. 22–31.
tem and pedestrian in Tempe, Arizona,
https://www.ntsb.gov/news/events/Documents/
2019-HWY18MH010-BMG-abstract.pdf, 2019.</p>
      <p>Accessed: 2020-01-15.
[5] S. Aditya, Y. Yang, C. Baral, Integrating
knowledge and reasoning in image understanding, in:
IJCAI, 2019. URL: https://doi.org/10.24963/ijcai.</p>
      <p>2019/873. doi:10.24963/ijcai.2019/873.
[6] J. Suchan, M. Bhatt, S. Varadarajan, Out of
sight but not out of mind: An answer set
programming based online abduction framework
for visual sensemaking in autonomous
driving, in: IJCAI-19, 2019. URL: https://doi.org/
10.24963/ijcai.2019/260. doi:10.24963/ijcai.</p>
      <p>2019/260.
[7] H. Beck, M. Dao-Tran, T. Eiter, LARS: A
logic-based framework for analytic reasoning
over streams, Artif. Intell. 261 (2018) 16–70.</p>
      <p>URL: https://doi.org/10.1016/j.artint.2018.04.003.</p>
      <p>doi:10.1016/j.artint.2018.04.003.
[8] J. Lee, Z. Yang, LPMLN, Weak Constraints, and</p>
      <p>P-log, in: AAAI, 2017.
[9] J. Lee, Y. Wang, A probabilistic extension of
action language BC+, TPLP 18 (2018) 607–622. URL:
https://doi.org/10.1017/S1471068418000303.</p>
      <p>doi:10.1017/S1471068418000303.
[10] K. Janowicz, A. Haller, S. J. D. Cox, D. L. Phuoc,</p>
      <p>M. Lefrançois, SOSA: A lightweight ontology for
sensors, observations, samples, and actuators, J.</p>
      <p>Web Semant. 56 (2019) 1–10.
[11] A. Bewley, Z. Ge, L. Ott, F. Ramos, B. Upcroft,</p>
      <p>Simple online and realtime tracking, in: ICIP,
2016, pp. 3464–3468.
[12] E. T. Mueller, Commonsense Reasoning: An</p>
      <p>Event Calculus Based Approach, 2 ed., 2015.</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>