<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Challenges for industrial-strength Information Retrieval on Databases</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Roberto Cornacchia Spinque</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Frank Dorssers Spinque</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Arjen P. de Vries Radboud University</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Michiel Hildebrand Spinque</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Implementing keyword search and other IR tasks on top of relational engines has become viable in practice, especially thanks to high-performance column-store technology. Supporting complex combinations of structured and unstructured search in real-world heterogeneous data spaces however requires more than \just" IR-on-DB. In this work, we walk the reader through our industrial-strength solution to this challenge and its application to a real-world scenario. By treating structured and unstructured search as rst-class citizens of the same computational platform, much of the integration e ort is pushed from the application level down to the data-management level. Combined with a visual design environment, this allows to model complex search engines without a need for programming.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        There is a growing demand for solving complex search
tasks in heterogeneous data spaces, such as enterprise
search [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], expert nding [
        <xref ref-type="bibr" rid="ref2 ref7">7, 2</xref>
        ], recommendation [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        These types of tasks require unstructured as well as
structured search. We argue that by implementing information
retrieval on a database it becomes easier to support complex
search tasks. Already in 1981, Crawford suggested in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]
that using standard query languages and proven relational
calculus: eases engineering; ensures repeatability of results
across systems; enables data-independence in text search
algorithms; allows search applications to bene t \for free"
from any advances in the database engine. In more recent
years, [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] emphasized these bene ts and showed that
relational technology can compete, performance-wise, with
specialized data structures, especially when implemented in
modern column-store engines optimized for online analytical
processing (OLAP) work-loads.
      </p>
      <p>In this paper we describe the challenges identi ed by
Spinque, a spin-o company from CWI Amsterdam, as
2017, Copyright is with the authors. Published in the Workshop
Proceedings of the EDBT/ICDT 2017 Joint Conference (March
21, 2017, Venice, Italy) on CEUR-WS.org (ISSN 1613-0073).
Distribution of this paper is permitted under the terms of the
Creative Commons license CC-by-nc-nd 4.0
it converted exciting results from research into
industrialstrength complex search solutions. We illustrate our
approach on a simpli ed as well as a real-world use case.
2.</p>
    </sec>
    <sec id="sec-2">
      <title>REAL-WORLD IR-on-DB</title>
      <p>This section describes the challenges for a uni ed
approach to structured and unstructured search and Spinque's
solution to these challenges:
1. e cient database implementation of IR tasks to search
unstructured data;
2. a exible data model to accommodate queries over any
type of structured data;
3. a mechanism to compute and propagate partial scores
consistently over unstructured and structured data;
4. an abstraction layer to model complex tasks easily.
The “toy” scenario. In the reminder of this section, let
us use the following running example: to perform keyword
search on a product database, but only consider the
description section of products in the category \toy".
2.1</p>
    </sec>
    <sec id="sec-3">
      <title>Keyword search in MonetDB/SQL</title>
      <p>
        The core of keyword search implementations is fast lookup
of query term occurrences within text documents. Inverted
index structures [
        <xref ref-type="bibr" rid="ref14 ref4">14, 4</xref>
        ] are used to map each term to its
\posting list" { the positions at which it appears in a
collection of text \documents". Terms occurrences are then
used to build the statistical data employed by the ranking
algorithm of choice.
      </p>
      <p>An inverted index can be easily implemented with any
relational DBMS. As shown in Figure 1, term lookup requires
an inner join on terms between a table containing query
terms and a table containing term occurrences.</p>
      <p>The ability to create such index structures on-demand is
crucial to support scenarios where keyword search is part of
more complex tasks, because their parameters (e.g.
stemming language) are often hard to decide upfront. Data fed
to our system undergoes almost no pre-processing, so that
the original text can be ranked at any time by e.g.
custom distance functions, tokenization strategies, stemming
choices. The only additions needed to MonetDB to support
on-demand indexing were two user-de ned functions to
implement a text tokenizer and Snowball stemmers for several
languages.
term
book
book
cake
history
term
book
cake
history
doc</p>
      <p>1term
term-doc
(b) Inverted index as a relational join on term</p>
      <p>Let us assume that the toy scenario introduced in
Section 2 has already been partially solved, so that a table
(productID int, description string) provides us with
pairs of \toy" products and their description, and products
have to be ranked according to the relevance of their
description to the query keywords. We show how Okapi BM25
ranking function can be implemented in MonetDB/SQL. The
following query turns a generic (docID int, data string)
table into the equivalent of a term-doc matrix:
CREATE VIEW term_doc AS
SELECT stem(lcase(token),'sb-english') as term, docID
FROM tokenize( (SELECT docID, data FROM docs) );</p>
      <p>From this doc-term matrix we can produce some simple
counts and a term dictionary:
CREATE VIEW doc_len AS
SELECT docID, count(*) as len
FROM term_doc GROUP BY docID;
CREATE VIEW termdict AS
SELECT row_number() over() as termID, terms.term
FROM (SELECT DISTINCT term FROM term_doc) AS terms;</p>
      <p>From a string-based doc-term matrix of boolean values, we
generate an integer-based doc-term matrix of frequencies:
CREATE VIEW tf AS
SELECT termdict.termID, term_doc.docID, count(*) as tf
FROM term_doc, termdict
WHERE term_doc.term = termdict.term
GROUP BY termdict.termID, term_doc.docID;</p>
      <p>Inverse document-frequency (IDF) of terms can be
formulated as follows:
CREATE VIEW idf AS
SELECT termID, log(
((SELECT count(*) FROM doc_len) - count(*) + 0.5)
/ ( count(*) + 0.5 ) ) as idf
FROM tf GROUP BY termID;</p>
      <p>BM25's term frequency is controlled by two free
parameters, k1 (saturation) and b (doc-length normalization):
CREATE VIEW tf_bm25 AS
SELECT tf.docID, tf.termID, tf.tf / (
tf.tf + (k1 * (1 - b + b * doc_len.len /
( SELECT avg(len) FROM doc_len ) ))) as tf
FROM tf, doc_len
WHERE tf.docID = doc_len.docID;</p>
      <p>We apply the normalization steps and dictionary mapping
seen above to the \query document" (a string singleton):
term
book
about
history
query</p>
      <p>CREATE VIEW qterms AS
SELECT termdict.termID
FROM tokenize((SELECT data from query)) AS qt, termdict
WHERE stem(lcase(qt.token),'sb-english') = termdict.term;</p>
      <p>Finally, the tf-idf contributions of all query terms are
summed up to de ne the relevance score of each document:
SELECT tf_bm25.docID, sum(tf_bm25.tf) as score
FROM tf_bm25, idf, qterms
WHERE tf_bm25.termID = qterms.termID
AND idf.termID = qterms.termID
GROUP BY tf_bm25.docID;</p>
      <p>
        Most alternative ranking functions would easily adapt or
reuse large parts of this implementation. Also, most of the
SQL queries above are independent of query-terms, which
allows to materialize intermediate results for reuse in di erent
search scenarios on the same data. While beating
specialized text retrieval systems on raw speed is not the focus of
this study, reaching reasonable performance is a requirement
for the development of real search solutions. In accordance
with [
        <xref ref-type="bibr" rid="ref10 ref5">5, 10</xref>
        ], we can report runtime performance in the range
of 20ms (hot data) for 3-term queries against a 2.3GB
collection of raw text (1.1M documents), on a standard Linux
desktop machine (8-core, Intel i7-3770S, 3.10GHz, 16GB
RAM, 256GB SSD), using MonetDB v11.23.14.
2.2
      </p>
    </sec>
    <sec id="sec-4">
      <title>Flexible data model</title>
      <p>Relational tables can store and query structured data
efciently, but they are not particularly application-friendly:
the schema of each table must be known by applications
using them, but the optimal schema depends on the data at
hand. Triple-stores can o er a valid alternative.</p>
      <p>Semantic triples are best known in the Semantic Web
community as the atomic data unit in the RDF data model.
Triples encode statements about resources, in the form of
(subject,property,object), and a collection of such
statements can be interpreted as a graph. Triple-stores are
(mostly relational) database systems speci cally designed to
store and manipulate this special kind of 3-column tables (or
4, when managing quadruples for named graphs), although
any relational DBMS would also serve the purpose. In fact,
we use the standard SQL interface of MonetDB to
implement and query a triple-store, with an important custom
addition that we describe in Section 2.3. One direct
advantage of using a standard SQL interface is that it allows to
exploit the strengths of both triples and standard tables.</p>
      <p>For the toy scenario of Section 2, the docs table to be
provided as input for keyword search (see Section 2.1) can
be generated at query time by the following SQL view:
CREATE VIEW docs as
SELECT t2.subject as docID, t2.object as data
FROM triples t1, triples t2
WHERE t1.property = 'category' AND t1.object = 'toy'
AND t2.property = 'description'
AND t1.subject = t2.subject;</p>
      <p>
        The triple model allows to write this and other
queries using simple and application-independent
patterns. However, this exibility comes at the price
of having to reconstruct the typical relational row
(product,category,description, ...) at query time,
which requires self-joins of a possibly large triples table.
Vertical partitioning of the set of triples can address
performance and scalability issues, as long as the right
partitioning approach is chosen for the application context at
hand. The only data-driven partitioning that we apply is
by the physical data type of objects (rather than
serializing every literal into strings). It is always applicable and
can improve e ciency, but does not solve scalability
issues. With the assumption that product categories and
descriptions are accessed often, storing (?,category,?) and
(?,description,?) triples into separate tables would
dramatically improve our scenario. This is the main reasoning
of [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], where the authors propose a vertical partitioning of
all properties into separate tables. However, [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] shows that
this solution is less scalable when the number of properties is
high. We use an on-demand approach to vertical
partitioning, which is applied not only to selections on the property
column of triples, but to any intermediate result generated in
our database. This creates an adaptive, query-driven set of
\cache" tables each corresponding to a speci c sub-query on
the original data. When the same computation is requested
several times, its full result is already materialized. An
interesting alternative to consider would be the detection of
emergent schemas [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], a data-driven technique to nd a
relational schema that is considered optimal for a given graph,
thus eliminating many join operations.
2.3
      </p>
    </sec>
    <sec id="sec-5">
      <title>Score propagation</title>
      <p>Sections 2.1 and 2.2 show how to combine ltering
(structured search) and ranking (unstructured search) of the text
collection de ned on-the- y by such ltering. What makes
these operations still disconnected is their inherently
different computational model: one produces certain answers
from facts, the other applies statistical methods to produce
likely relevant answers.</p>
      <p>We reduce this gap by implementing a probabilistic
relational database with tuple-level uncertainty: a probability
column p is appended to all tables, including triples, in our
RDBMS. Semantic triples no longer encode facts, but rather
uncertain events: (subject,property,object,p).
Probabilities smaller than 1 can originate from the data (e.g.
due to con dence-based data extraction techniques), or from
any intermediate computation that produces ranked results.
With probabilistic tuples, structured search need not be
restricted to boolean facts and can play alongside unstructured
search with the very same tools.</p>
      <p>
        Encoding probabilistic information is one part of the
solution. We still need to combine and propagate such
probabilities when tuples are processed by relational operators. For
this, we use a proprietary domain speci c language called
SpinQL, which implements the Probabilistic Relational
Algebra (PRA) developed in [
        <xref ref-type="bibr" rid="ref12 ref8">8, 12</xref>
        ], with particular focus on
e cient translation to SQL. SpinQL is used everywhere in
the system, including the implementation of BM25 and other
retrieval models. Extracting toy descriptions is expressed in
SpinQL as:
docs = PROJECT [$1,$6] (
      </p>
      <p>JOIN INDEPENDENT [$1=$1] (
SELECT [$2="category" and $3="toy"] (triples),
SELECT [$2="description"] (triples) ) );
and translates to:
CREATE VIEW docs as
SELECT t2.subject as docID, t2.object as data,
t1.p * t2.p as p
FROM triples t1, triples t2
WHERE t1.property = 'category' AND t1.object = 'toy'
AND t2.property = 'description'
AND t1.subject = t2.subject;
Each relational operator de nes how to compute
probability columns. For example, the query above joins
independent events from the two tables, which makes the resulting
probability of each join match be computed as the product
of two input tuples' probabilities. If applied correctly, this
algebra allows to keep the probabilistic computation sound.
Using SpinQL leads to more concise query plans and is less
error-prone than hand-specifying probability computations,
as these are only made explicit upon translation into SQL.
2.4</p>
    </sec>
    <sec id="sec-6">
      <title>Modeling complexity</title>
      <p>While SQL / SpinQL interfaces allow to express mixed
structured and unstructured search and can be evaluated
e ciently, they are not well suited for search engine
designers. A basic search engine would easily require tens of
queries with hundreds of lines of code. Therefore, we created
a graphical environment where a so-called search strategy is
modeled out of building blocks.</p>
      <p>Figure 2 shows the search strategy that de nes the toy
scenario used throughout Section 2. Block Rank by Text BM25
contains the BM25 implementation shown in Section 2.1,
though expressed in SpinQL rather than SQL. It takes a
probabilistic (docID, data) table on the left and a list of
query terms on the right. The sub-strategy on the left
corresponds to the sub-collection ltering of sections 2.2 and 2.3.</p>
      <p>Connecting blocks is a convenient way to express complex
search scenarios declaratively without programming e orts.
The SpinQL queries contained in each block are combined
automatically under the hood.
3.</p>
    </sec>
    <sec id="sec-7">
      <title>A REAL-WORLD SCENARIO</title>
      <p>Figure 3 depicts a simpli ed version (due to space and
con dentiality constraints) of a real strategy used by one of
our customers in the business of online auctions. Via the
website's search-bar, users activate this strategy to nd the
items they are interested in. The primary retrieval unit in
the database is a lot, which is an item or a set of items
for sale in an auction. Lots are connected to auctions via
triples like (lot23,hasAuction,auction12). Both lots and
auctions have their own identi er and a textual description,
as part of a rich semantic graph.</p>
      <p>Let us summarize the strategy in Figure 2 in a few steps:
1. The strategy rst selects nodes of type lot from the
graph, then it splits in two branches.
2. The branch on the left extracts the lot descriptions, on
which it ranks the lots with the given query keywords,
similarly to the toy scenario (Figure 2).
3. The branch on the right uses the same keywords
to rank lots by the description of the auctions in
which they are contained: it rst traverses
property hasAuction to nd the containing auction; then
ranks auctions by their description; nally, it traverses
hasAuction backward, to obtain lots again.
4. The two di erent ranked lists of lots are mixed via
linear combination, with the given weights.</p>
      <p>This strategy exempli es the concepts expressed in
Section 2:</p>
      <p>Keyword search. No speci c indexing con guration was
required. Two distinct inverted indices were created
ondemand, given the selected sub-collection.</p>
      <p>Flexible data model. Structured search, such as
ltering and graph traversal steps, rely on a data-agnostic
database schema. This allowed the design of strategy blocks
that work consistently on any collection.</p>
      <p>Score propagation. All the operations in this
strategy propagate probabilities through the graph, including the
rst ones, which carry unaltered probabilities (1.0) from
initial data. On the right branch however, the last traverse
operation nds lots with probabilities that depend on those
of their ranked auctions. This happens transparently, thanks
to the underlying probabilistic relational algebra layer.</p>
      <p>Strategy abstraction. Despite the mix of structured
and unstructured search that this strategy involves, it
remains understandable at a glance, with technical details
hidden, and can be engineered and modi ed easily.</p>
      <p>Industrial-strength implementation. The
production version of this strategy (which includes 5 parallel
keyword search branches and query expansion with synonyms
and compound terms), runs, together with several others,
on a single VM server (8-core, Intel Xeon E5-2620, 2.40GHz,
16GB RAM, 256GB SSD). It searches about 8 million lots
in 25 thousand auctions, 150,000 times per day (with peaks
of 450 per minute) with response times of about 150ms per
request (hot database). We consider this performance
adequate to the complexity of this task, but more importantly
it was achieved with no programming or optimization e ort.</p>
    </sec>
    <sec id="sec-8">
      <title>WRAP UP</title>
      <p>
        This work explored the long-standing IR and DB
integration issue with particular emphasis on the
implementation of industrial-strength search solutions. While [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]
already claimed that \databases form a exible rapid
prototyping tool ", we can add that \databases are also a solid and
viable solution for search in production environments ".
      </p>
      <p>We showed that by pulling information retrieval into a
database it becomes possible to realize a transparent
combination of structured and unstructured queries. This opens
up new ways to support complex search scenarios. With
the right abstractions on top of this infrastructure, realizing
e ective and e cient search solutions becomes a task for
domain and information specialists instead of programmers.
5.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>D. J.</given-names>
            <surname>Abadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Marcus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. R.</given-names>
            <surname>Madden</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Hollenbach</surname>
          </string-name>
          .
          <article-title>Scalable semantic web data management using vertical partitioning</article-title>
          .
          <source>VLDB '07</source>
          , pages
          <fpage>411</fpage>
          {
          <fpage>422</fpage>
          . VLDB Endowment,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>K.</given-names>
            <surname>Balog</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Fang</surname>
          </string-name>
          , M. de Rijke, P. Serdyukov, and
          <string-name>
            <given-names>L.</given-names>
            <surname>Si</surname>
          </string-name>
          .
          <article-title>Expertise retrieval</article-title>
          .
          <source>Foundations and Trends in Information Retrieval</source>
          ,
          <volume>6</volume>
          (
          <issue>2</issue>
          -3):
          <volume>127</volume>
          {
          <fpage>256</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Bellog n</surname>
          </string-name>
          , J.
          <string-name>
            <surname>Wang</surname>
            , and
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Castells</surname>
          </string-name>
          .
          <article-title>Bridging memory-based collaborative ltering and text retrieval</article-title>
          .
          <source>Inf</source>
          . Retr.,
          <volume>16</volume>
          (
          <issue>6</issue>
          ):
          <volume>697</volume>
          {
          <fpage>724</fpage>
          ,
          <string-name>
            <surname>Dec</surname>
          </string-name>
          .
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Bu</surname>
          </string-name>
          ttcher, C. Clarke, and
          <string-name>
            <given-names>G. V.</given-names>
            <surname>Cormack</surname>
          </string-name>
          . Information Retrieval:
          <article-title>Implementing and Evaluating Search Engines</article-title>
          . The MIT Press,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>R.</given-names>
            <surname>Cornacchia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Heman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zukowski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. P.</given-names>
            <surname>Vries</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Boncz</surname>
          </string-name>
          .
          <article-title>Flexible and e cient ir using array databases</article-title>
          .
          <source>The VLDB Journal</source>
          ,
          <volume>17</volume>
          (
          <issue>1</issue>
          ):
          <volume>151</volume>
          {
          <fpage>168</fpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>R. G. Crawford.</surname>
          </string-name>
          <article-title>The relational model in information retrieval</article-title>
          .
          <source>Journal of the American Society for Information Science</source>
          ,
          <volume>32</volume>
          (
          <issue>1</issue>
          ):
          <volume>51</volume>
          {
          <fpage>64</fpage>
          ,
          <year>1981</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Fazel-Zarandi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Huang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N.</given-names>
            <surname>Contractor</surname>
          </string-name>
          .
          <article-title>Expert recommendation based on social drivers, social network analysis, and semantic data representation</article-title>
          .
          <source>HetRec '11</source>
          , pages
          <fpage>41</fpage>
          {
          <fpage>48</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>N.</given-names>
            <surname>Fuhr</surname>
          </string-name>
          and
          <string-name>
            <given-names>T.</given-names>
            <surname>Ro</surname>
          </string-name>
          <article-title>lleke. A probabilistic relational algebra for the integration of information retrieval and database systems</article-title>
          .
          <source>ACM Trans. Inf</source>
          . Syst.,
          <volume>15</volume>
          (
          <issue>1</issue>
          ),
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>D.</given-names>
            <surname>Hawking</surname>
          </string-name>
          .
          <article-title>Challenges in enterprise search</article-title>
          .
          <source>ADC '04</source>
          , pages
          <fpage>15</fpage>
          {
          <fpage>24</fpage>
          ,
          <string-name>
            <surname>Darlinghurst</surname>
          </string-name>
          , Australia,
          <year>2004</year>
          . ACS, Inc.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>H.</given-names>
            <surname>Mu</surname>
          </string-name>
          hleisen, T. Samar,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lin</surname>
          </string-name>
          , and
          <string-name>
            <surname>A. de Vries</surname>
          </string-name>
          .
          <article-title>Old dogs are great at new tricks: Column stores for ir prototyping</article-title>
          .
          <source>SIGIR '14</source>
          , pages
          <fpage>863</fpage>
          {
          <fpage>866</fpage>
          , New York, NY, USA,
          <year>2014</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>M.-D.</surname>
            Pham and
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Boncz</surname>
          </string-name>
          .
          <article-title>Exploiting Emergent Schemas to Make RDF Systems More E cient</article-title>
          , pages
          <volume>463</volume>
          {
          <fpage>479</fpage>
          . Springer, Cham,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>T.</given-names>
            <surname>Roelleke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.</given-names>
            <surname>Azzam</surname>
          </string-name>
          .
          <article-title>Modelling retrieval models in a probabilistic relational algebra with a new operator: The relational bayes</article-title>
          .
          <source>The VLDB Journal</source>
          ,
          <volume>17</volume>
          (
          <issue>1</issue>
          ):5{
          <fpage>37</fpage>
          ,
          <string-name>
            <surname>Jan</surname>
          </string-name>
          .
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>L.</given-names>
            <surname>Sidirourgos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Goncalves</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kersten</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Nes</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Manegold</surname>
          </string-name>
          .
          <article-title>Column-store support for rdf data management: Not all swans are white</article-title>
          .
          <source>Proc. VLDB Endow</source>
          .,
          <volume>1</volume>
          (
          <issue>2</issue>
          ):
          <volume>1553</volume>
          {
          <fpage>1563</fpage>
          ,
          <string-name>
            <surname>Aug</surname>
          </string-name>
          .
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>J.</given-names>
            <surname>Zobel</surname>
          </string-name>
          and
          <string-name>
            <surname>A.</surname>
          </string-name>
          <article-title>Mo at. Inverted les for text search engines</article-title>
          .
          <source>ACM Comput. Surv.</source>
          ,
          <volume>38</volume>
          (
          <issue>2</issue>
          ),
          <year>July 2006</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>