<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>RUBEN: A Rule Engine Benchmarking Framework</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Kevin Angele</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jürgen Angele</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Umutcan Şimşek</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dieter Fensel</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>RUBEN</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rules</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rule Engines</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Benchmark</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Onlim GmbH</institution>
          ,
          <addr-line>Weintraubengasse 22, 1020 Vienna</addr-line>
          ,
          <country country="AT">Austria</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Semantic Technology Institute, University of Innsbruck</institution>
          ,
          <addr-line>Technikerstrasse 21a, 6020 Innsbruck</addr-line>
          ,
          <country country="AT">Austria</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>adesso, Competence Center Artificial Intelligence</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Knowledge graphs have become an essential technology for powering intelligent applications. Enriching the knowledge within knowledge graphs based on use case-specific requirements can be achieved using inference rules. Applying rules on knowledge graphs requires performant and scalable rule engines. Analyzing rule engines based on test cases covering various characteristics is crucial for identifying the optimal rule engine for a given use case. To this end, we present RUBEN: A Rule Engine Benchmarking Framework providing interfaces to benchmark rule engines based on given test cases. Besides a description of RUBEN's interfaces, we present a selection of test cases adopted from the OpenRuleBench, and an evaluation of four rule engines. In the future, we aim to benchmark existing rule engines regularly and encourage the community to propose new test cases and include other rule engines.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Knowledge graphs have become an important technology for powering intelligent applications
integrating data from heterogeneous (often incomplete) sources. Parts of the missing knowledge
can be inferred by using inference rules. Besides, rules can be used for data integration or
information extraction. In recent years, many new rule engines [
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1, 2, 3</xref>
        ] were developed targeting
knowledge graphs. Performance and scalability are eminent requirements for rule engines
operating on knowledge graphs to deliver fast responses for intelligent applications. Analyzing
rule engines based on test cases covering various characteristics is crucial for an overview of
the available engines, their performance, and scalability.
      </p>
      <p>
        In 2009 OpenRuleBench [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] providing a set of performance benchmarks for comparing and
analyzing rule engines was published. OpenRuleBench included systems relying on diferent
technologies, including Prolog-based, deductive databases, production rules, triple engines, and
general knowledge bases. The main issues when comparing various academic and commercial
systems are the diferent syntaxes and supported features. Therefore, manually generating the
rules for the various systems was necessary. At least the data for those test cases were generated
programmatically. Due to the diferences in the capabilities, not all test cases are applicable for
all the systems. OpenRuleBench is freely available and encourages the community to contribute.
Unfortunately, OpenRuleBench is not a benchmarking framework but a collection of rule sets
representing diferent test cases and corresponding datasets. Each tested system has its scripts
for running the test cases. This makes running the evaluation quite cumbersome. Besides, the
last assessment was conducted in 2011, and nothing seemed to have happened since then.
      </p>
      <p>Therefore, we present RUBEN: A Rule Engine Benchmarking Framework providing a simple
interface for including rule engines into a given set of test cases. The main aim of this framework
is to provide an easy way to extend the collection of engines to be evaluated and execute the
test cases without running multiple scripts. In the end, the output of all engines is combined
into a single result file.</p>
      <p>
        As a basis for this framework, we rely on the data provided by the OpenRuleBench [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The
test cases were completely adopted, and the test data was adapted to the new versions of the
engines. For the first version of RUBEN, we used a selection of the engines of the diferent
categories:
• Deductive database - Stardog1, VLog [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
• Production and reactive rule systems - Drools2
• Rule engines for triples - Jena3
      </p>
      <p>In the future, we plan to extend this list of engines and also encourage the community to add
new engines.</p>
      <p>In this paper, we give an overview of RUBEN’s implementation (Section 2), introduce the
test cases (Section 3) and present a small subset of the evaluation results (Section 4). Afterward,
related work in this area is presented. Finally, Section 6 concludes the paper and gives an
outlook on future work.</p>
    </sec>
    <sec id="sec-2">
      <title>2. RUBEN</title>
      <p>RUBEN4 is a rule engine benchmarking framework written in Java, bundling the evaluation of
various engines into a single framework that is easy to configure and execute. This chapter
presents RUBEN’s architecture and the interface to be implemented for rule engines that should
be included in the evaluation.</p>
      <p>Figure 1 presents the components RUBEN is composed of. The main component called
Ruben loads the evaluation configuration and triggers the execution of the test cases via the
BenchmarkExecutor. Rule engines and test cases are configurable for the evaluation.</p>
      <p>Table 1 presents the general configuration options, namely name and testDataPath, for RUBEN.
While the name is used as a label for the result file, the testDataPath specifies the location
of the test data. The test data needs to include the data needed for each test case and the
corresponding rules for each engine. So far, the data and rule files need to be provided in the
format supported by the rule engine. In the future, we aim to represent the test cases in a rule
engine-independent format. Then, the independent format needs to be interpreted by each rule
engine and transformed into their format for loading the data and rules. The properties engines
and testCases embed the rule engine and test case configurations.</p>
      <p>1https://www.stardog.com/
2https://www.drools.org/
3https://jena.apache.org/
4Check https://github.com/kev-ang/RUBEN for the source code.</p>
      <p>ReasoningEngineConfiguration</p>
      <p>TestCaseConfiguration
use</p>
      <p>use</p>
      <p>BenchmarkConfiguration
BenchmarkExecution</p>
      <p>Ruben</p>
      <p>use
BenchmarkExecutor</p>
      <p>RuleEngine
RuleEngines
implements
implements
implements</p>
      <p>implements
Drools</p>
      <p>Jena</p>
      <p>Stardog</p>
      <p>VLog</p>
      <p>Each rule engine can be configured by using the properties name, classpath, and settings (see
Table 2). The name of the rule engine is used to identify the corresponding test data within
the test data folder. Besides, the classpath refers to the class within the evaluation framework
used to execute the evaluation. settings can be used to provide rule engine-specific settings
(optional). Including the same rule engine with a diferent name and settings allows evaluating
multiple configurations for the same rule engine.</p>
      <p>Table 3 presents the configuration of test cases and properties that need to be specified. A
test case consists of a testCategory, testCaseIdentifier , and testName. The testCategory is used to
categorize tests. Within those categories the testName identifies the name of the test. Each test</p>
      <sec id="sec-2-1">
        <title>Specify a name for the configuration. Engines to be used for the evaluation. For further details on how to configure the engines, see Table 2.</title>
        <p>Configure the test cases to be included in the
evaluation. For further details on how to
configure the test cases, see Table 3.</p>
        <p>Path to the folder containing the data required
for the evaluation. The structure of the folder
containing the test data must follow a predefined
pattern. For each rule engine to be evaluated a
folder is needed (the folder name must be equal to
the name field in Table 2. Inside this folder there
must be a folder for each test category (see
testCategory in Table 3). Within the category folder
a folder with the name equal to the testName (see
Table 3) needs to be included. The files within
the test folder need to be named according to the
testCaseIdentifier values (see Table 3).</p>
        <p>Description</p>
      </sec>
      <sec id="sec-2-2">
        <title>Name of the rule engine to be evaluated. This</title>
        <p>name must be used as name for the folder within
the test data folder.</p>
        <p>Refers to the implementation of the rule engine
within the framework.</p>
        <p>Define additional settings for the rule engine.</p>
        <p>Those settings are provided as a map consisting
of key values.
can have multiple test cases identified by the testCaseIdentifier .</p>
        <p>The path for loading the test data for each test case and rule engine is composed of diferent
information in the configuration and has the following structure:
{ t e s t D a t a P a t h } / { e n g i n e _ n a m e } / { t e s t C a s e _ t e s t C a t e g o r y } /
{ t e s t C a s e _ t e s t N a m e } / { t e s t C a s e _ t e s t C a s e I d e n t i f i e r }</p>
        <p>Including a rule engine into the evaluation framework requires the implementation of the
RuleEngine interface. Table 4 presents the methods to be implemented for a rule engine.</p>
        <p>For an overview of the flow through the framework, we will present the steps taken to
load the configuration and execute the test cases. Initially, the main component ( Ruben) loads
the evaluation configuration. Afterward, the framework iterates through the provided rule
engine configurations to execute all test cases for each rule engine. To evaluate the test cases,
Ruben forwards the required information about the test data, the rule engine, and the test
cases to the BenchmarkExecutor component. In the first step, the BenchmarkExecutor calls the
prepare method of the current rule engine to load the relevant test data and rules required for
the given test case. Afterward, the queries of the provided test case are executed using the
executeQuery method, and the results in the form of the number of results are stored in a report
together with the execution time. After executing all queries for a given test case, the cleanUp
method is called to prepare for the next test case. Finally, when all test cases are evaluated,
the BenchmarkExecutor calls the shutDown method of the rule engine to stop all processes and
clean up all temporary files. The results are collected, and the framework continues with the
following rule engine.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Test Cases</title>
      <p>
        For the initial version of RUBEN, we rely on the test cases provided by the OpenRuleBench
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. OpenRuleBench’s main aim is to test several tasks rule engines are known to be good at.
Therefore, the authors in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] used datasets of diferent sizes ranging from 50,000 to 1,000,000
facts. Alongside the generated test cases, OpenRuleBench includes four real-world benchmarks:
DBLP database5, Mondial6, wine ontology7, and WordNet8. The selected tests are representative of
database and knowledge representation problems. Although all the rule engines in the original
OpenRuleBench support actions, such as those in production rule systems and Prolog, they are
not included due to the diferent paradigms of the engines.
      </p>
      <p>
        The authors in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] introduce several test categories with dedicated test cases:
• Large join tests
• Datalog recursion
• Default negation
      </p>
      <p>
        The following briefly introduces the tests for the diferent test categories. The descriptions
are taken from [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <sec id="sec-3-1">
        <title>3.1. Large Join Tests</title>
        <p>This test category contains two database join tests Join1 and Join2, LUBM-derived tests, the
Mondial, and DBLP tests.</p>
        <p>Join1 is about a non-recursive tree of binary joins. Those recursions are represented by the
following rules9:
a(X,Y) :- b1(X,Z), b2(Z,Y).
b1(X,Y) :- c1(X,Z), c2(Z,Y).
b2(X,Y) :- c3(X,Z), c4(Z,Y).</p>
        <p>c1(X,Y) :- d1(X,Z), d2(Z,Y).</p>
        <p>The facts for the base relations c2, c3, c4, d1, and d2 are randomly generated, resulting in two
datasets with 50,000 facts and 250,000 facts. Based on the derived predicates a, b1, and b2 the
5http://www.informatik.uni-trier.de/~ley/db/
6http://www.dbis.informatik.uni-goettingen.de/Mondial/
7http://www.w3.org/TR/2004/REC-owl-guide-20040210/#WinePortal
8http://wordnet.princeton.edu
9The rules within this section are represented using the Prolog syntax.
test queries are formulated using diferent bindings for the variables: free-free, free-bound, and
bound-free.</p>
        <p>Additionally, a test called 5*Join1 was created out of five copies of the previous rule set. Here,
the predicate a(X,Y) was renamed to a1, ..., a5. Then, a new rule unioning the results of the
previous five rule set was introduced.</p>
        <p>a(X,Y) :- a1(X,Y); a2(X,Y); a3(X,Y); a4(X,Y); a5(X,Y).</p>
        <p>
          Join2 consists of patterns of joins borrowed from [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. Those joins produce a large intermediate
result but a small set of answers.
        </p>
        <p>ra(A,B,C,D,E) :- p(A),p(B),p(C),p(D),p(E).
rb(A,B,C,D,E) :- p(A),p(B),p(C),p(D),p(E).
r(A,B,C,D,E) :- ra(A,B,C,D,E),rb(A,B,C,D,E).
q(A) :- r(A,_,_,_,_).
q(B) :- r(_,B,_,_,_).
q(C) :- r(_,_,C,_,_).
q(D) :- r(_,_,_,D,_).</p>
        <p>q(E) :- r(_,_,_,_,E).</p>
        <p>The query in this test case determines all facts for predicate q and the content of the base
relation is p(a0), ..., p(a18).</p>
        <p>
          LUBM-derived Tests include three rule sets adapted from the original Lehigh Benchmark,
LUBM [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. LUBM is a dataset generator for a synthetical university dataset consisting of
universities (the amount can be specified for the generation), courses, departments, professors, and
students. For the OpenRuleBench, the authors generated two datasets: one for ten universities
resulting in more than 1,000,000 tuples; one for 50 universities resulting in over 6,000,000
tuples. The original LUBM benchmark provides 14 queries, of which three (Query1, Query2,
and Query9) were selected and adapted.
query1(X) :- takesCourse(X,graduateCourse0), graduateStudent(X).
        </p>
        <p>Query1 joins two relations on an attribute with high selectivity (i.e., each tuple in one relation
joins with a small number of tuples in the other relation).
query2(X,Y,Z) :- graduateStudent(X), memberOf(X,Z),
↪ undergraduateDegreeFrom(X,Y), university(Y), department(Z),
↪ subOrganizationOf_0(Z,Y).</p>
        <p>Query2 joins three unary and three binary relations with high selectivity. Therefore, the final
answer, even for the most extensive dataset consisting of 50 universities, contains only around
100 tuples.
query9(X,Y,Z) :- advisor(X,Y), teacherOf(Y,Z), takesCourse(X,Z), student(X),
↪ faculty(Y), course(Z).</p>
        <p>Query9 joins three binary and three unary relations with a lower selectivity than Query1 and
Query2. The answer to this query is rather large (for the large dataset, around 10,000 tuples).
Mondial is a database consisting of geographical information derived from the CIA
Factbook10. It contains around 60,000 facts providing information about cities, provinces, and
countries worldwide. Only one query is used for this test case delivering statistical
information about provinces in China. The particularity of the query is the large intermediate result
(1,676,942 facts) compared to the small final result (888 facts).</p>
        <p>DBLP is a database representing publications about databases and logic programming. The
database is derived from the Web-based bibliography DBLP11. A single relation with nearly
2,500,000 facts about more than 200,000 publications builds up the test.
q(Id,T,A,Y,M) :- att(Id,title,T), att(Id,year,Y), att(Id,author,A),
↪ att(Id,month,M).</p>
        <p>The query within this test case is a 4-way join of parts of the same database relation.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Datalog Recursion</title>
        <p>This category includes test cases using recursion, a significant feature for distinguishing
rulebased systems from traditional database management systems. The authors of the
OpenRuleBench intended to evaluate the performance of such queries by providing recursive tests.
The tests in this category include:
• Classical transitive closure
• The same-generation siblings problem
• WordNet (application from natural language processing)
• The wine ontology in a rule-based representation
Transitive Closure of a binary relation (par ) is the smalles transitive relation containing par.
tc(X,Y) :- par(X,Y).
tc(X,Y) :- par(X,Z), tc(Z,Y).</p>
        <p>Deriving the relation tc causes trouble in traditional prologs if the predicates in the second
rule switch places. Also, when the par relation represents a cyclic graph, Prolog might run into
an infinite loop. Four datasets were randomly generated for this test: with cycles with 50,000
facts and 500,000 facts and without cycles (also 50,000 facts and 500,000 facts).
Same-Generation Problem tries to find all siblings in the same generation. The same
generation refers to the equal distance from a common ancestor.
sg(X,Y) :- sib(X,Y).
sg(X,Y) :- par(X,Z), sg(Z,Z1), par(Y,Z1).</p>
        <p>The base relations of par and sib were randomly generated, and again two types of datasets
(cyclic and acyclic) are used for this test. The smaller datasets consist of 6000 and the larger of
24000 facts.</p>
        <p>10https://www.cia.gov/the-world-factbook/
11http://www.informatik.uni-trier.de/~ley/db/
WordNet includes common queries from natural language processing. The queries of the
provided test case seek to find all:
• hypernyms - words more general than the given word
• hyponyms - words more specific than the given word
• meronyms - words related by the part-of-a-whole semantic relation
• holonyms - words related by the composed-of relation
• troponyms - words more precise than the given word
• same-synset - a word that is in the same set of synonyms as the given word
• glosses
• antonyms - a word with the opposite meaning of the given word
• adjective-clusters</p>
        <p>As the basis for the test data, WordNet Version 3.0 was used. The database consists of around
115,000 synsets containing over 150,000 words in total. Most tests deliver more than 400,000
facts, and some cases exceed 2,000,000 facts. In this paper, we only show the hypernyms test.
For a complete description of the tests, check the GitHub-Repository12
hypernyms(W1,W2) :- s(S1,_,W1,_,_,_), hypernymSynsets(S1,S2),</p>
        <p>↪ s(S2,_,W2,_,_,_).
hypernymSynsets(S1,S2) :- hypernym(S1,S2).
hypernymSynsets(S1,S2) :- hypernym(S1,S3), hypernymSynsets(S3,S2).
Wine Ontology is a rule-based representation of the OWL wine ontology, consisting of 815
rules and 654 facts. A characteristic of this test is the recursive dependency of many predicates
on each other through chains of rules, resulting in large groups of predicates connected via the
depends-on relationship. Those large groups are especially critical for top-down engines.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Default Negation</title>
        <p>
          In this category, the tests include default negation in the body of the rules. In contrast to the
OpenRuleBench, we focus only on predicate-stratified negation [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], so far. A modified
samegeneration problem is used for the predicate-stratified negation.
        </p>
        <p>The modified same-generation test is as follows:
nonsg(X,Y) :- tc(X,Y).
nonsg(X,Y) :- tc(Y,X).
sg2(X,Y) :- sg(X,Y), not nonsg(X,Y).</p>
        <p>Again, as for the original same-generation problem, the base relations of par and sib were
randomly generated, with cycles in the data. This test is executed for the two data sets of 6000
and 24000 facts.</p>
        <p>12https://github.com/kev-ang/RUBEN</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Evaluation and Results</title>
      <p>
        This section presents a subset of the evaluation results produced by RUBEN. We focused on the
test cases and a subset of the tools evaluated in OpenRuleBench [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. This section first presents
the experimental setup, then the methodology, and finally, the evaluation results.
      </p>
      <sec id="sec-4-1">
        <title>4.1. Experimental Setup</title>
        <p>The evaluation framework was hosted on a server with an Intel XEON E5-v3, 6 Cores / 12 Threads,
3.50 GHz Base Frequency, 3.80 GHz Max Turbo Frequency processor and 256GB RAM. Out of the
256GB RAM, only 32GB were made available to the framework and, therefore, to the engines.
The operating system on the machine is Debian GNU / Linux 11.</p>
        <p>
          Similar to the OpenRuleBench, we evaluate engines from diferent categories. In total, we
evaluate four engines out of three categories. The engines to be evaluated are assigned to the
following categories:
• Deductive Databases
• Production rule system
• Rule engines for triples
– Stardog - is implemented in Java providing extensive reasoning capabilities, including
datalog evaluation.
– VLog - is implemented in C++ and provides an eficient Datalog engine for large
knowledge graphs supporting RDF, OWL, and SPARQL. The eficiency (memory usage and
speed) is a result of the combination of a column-based layout with novel optimization
methods. VLog’s code is open source and freely available.
– Drools - a bottom-up engine based on the Rete algorithm [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. The Rete algorithm
combines semi-naive bottom-up computation [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] with a certain heuristic for common
expression elimination [10].
– Apache Jena - a Java-based framework including two rule engines (bottom-up and
top-down).
        </p>
        <p>Apache Jena, Drools, and Stardog are rule engines not materializing the rules. Materializing
implies calculating all the implicit facts on load time and storing them. In contrast, VLog
materializes the rules, and for evaluating a query, VLog can directly access the materialized
facts13. Therefore, the results are not directly comparable.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Methodology</title>
        <p>The used data sets ranging from 50,000 to 1,000,000 facts are generated with the help of data
generators. For running the various tests OpenRuleBench provides, we adopted the test cases.
We provided a data file containing the facts, a rule file, and a file containing the queries to be
evaluated. This allows providing the files in a format optimal for the rule engine. Additionally,
13The paper mentions a query-driven reasoning mode, which we did not find so far in the examples.
rules can be optimized for specific test cases and for each engine. Each query is evaluated two
times, and the response time of the second value is taken as the result. The first execution is
used to initialize all indices and fill the caches. The second query then represents the optimized
query response time. The query response time includes the rule evaluation and the result
counting. The results for each test case are collected and returned in a file.</p>
        <p>Due to the limited space, we present a small subset of the evaluation results:
• Large join tests - include database joins (50,000 facts).
• Datalog recursion - include classical transitive closure and the well-known
samegeneration siblings problem.</p>
        <p>The selection of those test cases was based on the ability of the given rule engines. Since
Drools and Jena participated in the OpenRuleBench, we knew upfront the intersection between
those. Those test case candidates were also checked for the other rule engines.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Results</title>
        <p>Focusing on the number of test cases the engines evaluated, Drools and VLog supported the
large join and datalog recursion tests. Jena and Stardog could not evaluate tests like Join2
and Same Generation for negation. The first requires predicates like p("abcd0") and Same
Generation requires negation which is not supported by both engines in their rule language.
Tables 5, 6, and 7 present the results of those tests that were supported by the selected engines.</p>
        <p>Special values in cells specify an unexpected behavior: error implies that the query evaluation
did not finish as expected, timeout indicates that the evaluation was not finished within 15
minutes. All times within the table are given in seconds and contain only the rule evaluation and
result counting. Loading time is excluded from those numbers. For VLog, we discuss the results
separately, as it materializes all the rules before evaluating the queries. The materialization
time for the diferent test cases for VLog is given in brackets in the table cell.</p>
        <p>The results in Table 5 show that Stardog is the fastest rule-based engine for the large join
test that does not materialize the rules upfront. The large join test uses the same dataset for
all three queries. Therefore, the materialization time shown for VLog is relevant only once for
the whole test case. As you can see, the materialization time for this test case for VLog is quite
high, at nearly 6 minutes. However, the query times are then a tiny fraction of the query time
of Stardog.</p>
        <p>Jena is the fastest engine for the Datalog recursion same generation test in Table 6, as Drools
only throws errors. Here, VLog would be faster than Jena even if the materialization time is
added to the query times. Stardog was not able to execute this test case and was therefore left
out.</p>
        <p>Stardog is the fastest rule engine for the Datalog recursion transitive closure test in Table 7
for the engines without materialization. As for the test case before, VLog would be faster than
Stardog even if each query’s materialization times are included.</p>
        <p>Drools threw an exception in all test cases with an OutOfMemoryError. When evaluating
the rules in Drools, Java objects are generated. This number of objects proliferates, causing
an out-of-memory error. For the next run (next year), optimizing and adapting the rules is
necessary to avoid out-of-memory errors.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Related Work</title>
      <p>This section presents related work about rule engine benchmarking frameworks. Besides, in
the second part of this section, we present some benchmarking datasets used by various rule
engine implementations for the evaluation.</p>
      <sec id="sec-5-1">
        <title>5.1. Benchmarking Frameworks</title>
        <p>
          One framework already known is the OpenRuleBench [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. OpenRuleBench analyzes the
performance and scalability of various rule engines by providing a set of test cases covering diferent
functionalities. OpenRuleBench is open source and welcomes contributions from the
community. However, the last evaluation report dates 2011. Besides, running OpenRuleBench is
cumbersome due to customized scripts for each supported rule engine. Integrating a new rule
engine into the framework requires writing shell scripts and adopting the given test cases to
introduce the rule engine.
        </p>
        <p>A more up-to-date framework for benchmarking rule-based inference engines is [11]. The
authors provide a fully automated framework for benchmarking rule-based reasoning engines.
Therefore, a configuration for a generation tool for generating a meta-model needs to be
provided. Then, the generated meta-model is translated for the diferent rule models of the
rule engines to be benchmarked. The translated model is then fed into the respective rule
engine, and reasoning time and memory consumption is measured. The diferent parts of the
framework are implemented in Java. However, when integrating a new rule engine into the
framework, a translation class needs to be implemented to translate the meta-model into the
corresponding rule model. Afterward, a shell script needs to be provided to run the test cases
on the new rule engine. In the end, there is an overall shell script responsible for running the
whole benchmark. This framework supports only generated test cases which is a drawback
compared to OpenRuleBench, for example.</p>
        <p>In contrast to OpenRuleBench and the other framework, RUBEN is not based on shell scripts
and is fully implemented in Java. RUBEN provides interfaces implemented by rule engines that
should be included in the benchmark, allowing seamless integration.</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Benchmarking Datasets</title>
        <p>
          Two popular rule engines published in recent years are RDFox [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] and Vlog [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. The authors did
not use one of the previously introduced frameworks for their evaluation. They used existing
datasets and adapted them to measure the performance. In the following, we first present
datasets of rules followed by datasets for benchmarking RDF and OWL systems.
        </p>
        <p>Datasets including rules are, e.g., Manners or WaltzDB14. These datasets focus on evaluating
pure matching algorithms (selecting rules to be evaluated) or testing a subset of available
reasoning engines. However, they can not be used for other systems than those for which they
are defined.</p>
        <p>
          RDFox and Vlog rely more or less on the same set of test data15. One of those datasets is
Claros, a cultural database catalog representing archaeologic artifacts. This dataset does not
come with rules. Therefore, the authors added some manually generated rules. Another dataset
is DBpedia, representing structured data extracted from the Wikipedia info boxes. Same as
for Claros, the authors extended the dataset with manually generated rules. Both previous
datasets are real-world datasets. In the following we briefly introduce two synthetical datasets
called LUBM [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] and an extension of LUBM called UOBM [12]. LUBM is widely used for RDF
systems with an ontology describing universities, departments, professors, and students. The
data is generated by specifying the number of universities and comes with 14 well-designed
test queries. Those queries cover the features of traditional reasoning systems. However, rules
are not included, and the authors of Vlog and RDFox added them manually. UOBM extends
LUBM by addressing the sparsity of the data. Therefore, external links between members in
diferent university instances are added, resulting in exponential growth of the complexity for
scalability testing.
        </p>
        <p>14https://www.cs.utexas.edu/ftp/ops5-benchmark-suite/
15http://www.cs.ox.ac.uk/isg/tools/RDFox/2014/AAAI/</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion and Future Work</title>
      <p>
        In this paper, we presented a rule engine benchmarking framework called RUBEN, implemented
in Java. RUBEN evaluates the performance and scalability of rule engines; essential to select
a rule engine best fitting a given use case. In contrast to existing benchmarking frameworks
like OpenRuleBench [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], RUBEN provides a simple interface a rule engine needs to implement
to be included in the benchmark. Unfortunately, the first version of RUBEN requires adapting
the test cases to the newly introduced rule engine. So far, RUBEN provides most of the tests
initially introduced in the OpenRuleBench suite. The rule engines supported by RUBEN are a
subset of the engines introduced by OpenRuleBench. To this end, we welcome the community
to extend the rule engines and test cases list.
      </p>
      <p>
        In the future, we plan to add a general description of test cases by introducing a common rule
format. Then, each rule engine can include a translation method for converting the common
rule format into the required format. Still, keeping the possibility to provide optimized rule sets
for the cases. Besides, the list of rule engines will be extended by including RDFox [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Finally,
(maybe) with the help of the community, we will extend the test cases to fully support and
extend the list of test cases provided by OpenRuleBench. Further, it is planned to repeat the
benchmark every year.
[10] F. Bry, N. Eisinger, T. Eiter, T. Furche, G. Gottlob, C. Ley, B. Linse, R. Pichler, F. Wei,
Foundations of rule-based query answering, Reasoning Web International Summer School
(2007) 1–153.
[11] S. Bobek, P. Misiak, Framework for benchmarking rule-based inference engines, in:
International Conference on Artificial Intelligence and Soft Computing, Springer, 2017, pp.
399–410.
[12] L. Ma, Y. Yang, Z. Qiu, G. Xie, Y. Pan, S. Liu, Towards a complete owl ontology benchmark,
in: European Semantic Web Conference, Springer, 2006, pp. 125–139.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J.-F.</given-names>
            <surname>Baget</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Leclère</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.-L. Mugnier</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Rocher</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Sipieter</surname>
          </string-name>
          ,
          <article-title>Graal: A toolkit for query answering with existential rules</article-title>
          ,
          <source>in: International Symposium on Rules and Rule Markup Languages for the Semantic Web</source>
          , Springer,
          <year>2015</year>
          , pp.
          <fpage>328</fpage>
          -
          <lpage>344</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Nenov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Piro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Motik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Horrocks</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Banerjee</surname>
          </string-name>
          ,
          <article-title>Rdfox: A highly-scalable rdf store</article-title>
          , in: International Semantic Web Conference, Springer,
          <year>2015</year>
          , pp.
          <fpage>3</fpage>
          -
          <lpage>20</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D.</given-names>
            <surname>Carral</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Dragoste</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>González</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Jacobs</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Krötzsch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Urbani</surname>
          </string-name>
          ,
          <article-title>Vlog: A rule engine for knowledge graphs</article-title>
          , in: International Semantic Web Conference, Springer,
          <year>2019</year>
          , pp.
          <fpage>19</fpage>
          -
          <lpage>35</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Fodor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kifer</surname>
          </string-name>
          ,
          <string-name>
            <surname>Openrulebench:</surname>
          </string-name>
          <article-title>An analysis of the performance of rule engines</article-title>
          ,
          <source>in: Proceedings of the 18th international conference on World wide web</source>
          ,
          <year>2009</year>
          , pp.
          <fpage>601</fpage>
          -
          <lpage>610</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>B.</given-names>
            <surname>Bishop</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Fischer</surname>
          </string-name>
          ,
          <article-title>Iris-integrated rule inference system</article-title>
          ,
          <source>in: International Workshop on Advancing Reasoning on the Web: Scalability and Commonsense (ARea</source>
          <year>2008</year>
          ), sn,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Heflin</surname>
          </string-name>
          ,
          <article-title>Lubm: A benchmark for owl knowledge base systems</article-title>
          ,
          <source>Journal of Web Semantics</source>
          <volume>3</volume>
          (
          <year>2005</year>
          )
          <fpage>158</fpage>
          -
          <lpage>182</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>K. R.</given-names>
            <surname>Apt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. A.</given-names>
            <surname>Blair</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Walker</surname>
          </string-name>
          ,
          <article-title>Towards a theory of declarative knowledge, in: Foundations of deductive databases and logic programming</article-title>
          ,
          <source>Elsevier</source>
          ,
          <year>1988</year>
          , pp.
          <fpage>89</fpage>
          -
          <lpage>148</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>C. L.</given-names>
            <surname>Forgy</surname>
          </string-name>
          ,
          <article-title>Rete: A fast algorithm for the many pattern/many object pattern match problem</article-title>
          ,
          <source>in: Readings in Artificial Intelligence and Databases</source>
          , Elsevier,
          <year>1989</year>
          , pp.
          <fpage>547</fpage>
          -
          <lpage>559</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J. D.</given-names>
            <surname>Ullman</surname>
          </string-name>
          , Principles of Database and
          <string-name>
            <surname>Knowledge-Base Systems - Volume</surname>
            <given-names>I</given-names>
          </string-name>
          :
          <article-title>Classical Database Systems</article-title>
          , Computer Science Press, New York, Oxford,
          <year>1988</year>
          . URL: http://dl.acm. org/citation.cfm?id=
          <fpage>42790</fpage>
          ,
          <year>1995</year>
          <article-title>wurde der 8. Nachdruck des Buches publiziert</article-title>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>