<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Creating and Querying Personalized Versions of Wikidata on a Laptop</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Hans Chalupsky</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pedro Szekely</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Filip Ilievski</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Daniel Garijo</string-name>
          <email>daniel.garijo@upm.es</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kartik Shenoy</string-name>
          <email>kshenoyg@isi.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Information Sciences Institute, University of Southern California</institution>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Ontology Engineering Group, Universidad Politecnica de Madrid</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Application developers today have three choices for exploiting the knowledge present in Wikidata: they can download the Wikidata dumps in JSON or RDF format, they can use the Wikidata API to get data about individual entities, or they can use the Wikidata SPARQL endpoint. None of these methods can support complex, yet common, query use cases, such as retrieval of large amounts of data or aggregations over large fractions of Wikidata. This paper introduces KGTK Kypher, a query language and processor that allows users to create personalized variants of Wikidata on a laptop. We present several use cases that illustrate the types of analyses that Kypher enables users to run on the full Wikidata KG on a laptop, combining data from external resources such as DBpedia. The Kypher queries for these use cases run much faster on a laptop than the equivalent SPARQL queries on a Wikidata clone running on a powerful server with 24h time-out limits. Type: Research paper Code repository: https://github.com/usc-isi-i2/kgtk/</p>
      </abstract>
      <kwd-group>
        <kwd>Wikidata</kwd>
        <kwd>Knowledge Graphs</kwd>
        <kwd>KGTK</kwd>
        <kwd>Kypher</kwd>
        <kwd>Cypher</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Modern Knowledge Graphs (KGs) are increasingly focused on improving their
coverage of instances and statements and enhancing their expressivity in order
to support application needs such as question answering and entity linking. As a
result, Wikidata [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], a popular and representative KG, contains nearly 95 million
entities described with over 1.3 billion statements.1 Wikidata is also highly
expressive, using a rei cation model where each statement includes quali ers (e.g.,
to indicate temporal validity) and references (which provide the source(s) from
which the statement comes from).
      </p>
      <p>Copyright © 2021 for this paper by its authors. Use permitted under Creative
Commons License Attribution 4.0 International (CC BY 4.0).
1 https://grafana.wikimedia.org/d/000000175/wikidata-datamodel-statements</p>
      <p>
        Application developers today have three choices for exploiting the knowledge
present in Wikidata. They can download the Wikidata dumps in JSON or RDF
[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] format, they can use the Wikidata API to get data about individual entities,
or they can use the Wikidata SPARQL [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] endpoint for more elaborate and
complex queries.2 The public Wikidata SPARQL endpoint restricts queries to 5
minutes, returning an error when a query exceeds (or plans to exceed) that time.
To mitigate the time-limit restriction, developers can load the massive Wikidata
RDF dump on their own servers, a relatively complex process that requires a
large server and several days.
      </p>
      <p>
        This paper introduces Kypher, the query language and processor of the
KGTK Knowledge Graph Toolkit [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], which allows creating personalized
variants of Wikidata on a laptop, and enables running analytic queries faster than a
Wikidata SPARQL endpoint. Because Kypher uses the KGTK representation,
it is not restricted to Wikidata, and can be used to query RDF KGs such as
DBpedia [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The key advantages of Kypher over existing tooling are:
1. Ability to extract large amounts of data from Wikidata.
2. Ability to execute queries that retrieve large portions of the full Wikidata.
3. Ability to build personalized versions of Wikidata, and extending it with
other datasets for speci c use cases.
4. Easy installation as there are no databases to set up or administer.
5. Minimal hardware requirements, as Kypher can be used to query Wikidata
and DBpedia on a laptop.
      </p>
      <p>The rest of the paper is structured as follows. Section 2 describes the KGTK
toolkit, and section 3 presents an introduction to the Kypher query language.
Section 4 introduces ve representative use cases to illustrate the bene ts of
Kypher, and section 5 reports the times needed to address the use cases using
Kypher queries on a laptop; SPARQL queries on a clone of the Wikidata
endpoint; and SPARQL queries on the public Wikidata endpoint. Section 6 presents
conclusions, discussion of the results and directions for future work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Background</title>
      <p>
        The Knowledge Graph Toolkit (KGTK) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] is a comprehensive framework for
the creation and exploitation of large hyper-relational KGs, designed for ease
of use, scalability, and speed. KGTK represents KGs in tab-separated (TSV)
les with four columns: edge-identi er, head, edge-label, and tail. All KGTK
commands consume and produce KGs represented in this format, so they can
be composed into pipelines to perform complex transformations on KGs. KGTK
provides a suite of import commands to import Wikidata, RDF and popular
graph representations into the KGTK format. A rich collection of transformation
commands make it easy to clean, union, lter, and sort KGs; graph combination
commands support e cient intersection, subtraction, and joining of large KGs;
      </p>
      <sec id="sec-2-1">
        <title>2 https://query.wikidata.org/</title>
        <p>
          graph analytics commands support scalable computation of centrality metrics
such as PageRank, degrees, connected components and shortest paths; advanced
commands support lexicalization of graph nodes, and computation of multiple
variants of text and graph embeddings over the whole graph. In addition, a
suite of export commands supports the transformation of KGTK KGs into
commonly used formats, including the Wikidata JSON format, RDF triples, JSON
documents for ElasticSearch indexing and graph-tool.3 Finally, KGTK allows
browsing KGs in a UI using a variant of SQID;4 and includes a development
environment using Jupyter notebooks that provides seamless integration with
Pandas [
          <xref ref-type="bibr" rid="ref4 ref8">8, 4</xref>
          ]. KGTK can process Wikidata-sized KGs, with billions of edges, on
a laptop. Kypher (kgtk query) is one of 55 commands available in KGTK.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Kypher query language and processor</title>
      <p>
        Kypher stands for KGTK Cypher. Cypher [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] is a declarative graph query
language originally developed at Neo4j. OpenCypher5 is a corresponding
opensource development e ort for Cypher which forms the basis of the new Graph
Query Language (GCL).6 We chose Cypher since its ASCII-art pattern language
makes it easy even for novices to express complex queries over graph data.
      </p>
      <p>Kypher adopts many aspects of Cypher's query language, but has some
important di erences. Most notably, KGTK and therefore Kypher do not use the
property graph data model assumed by Cypher. Kypher only implements a
subset of the Cypher commands (for example, no update commands) and has some
minor di erences in syntax, for example, to support naming and querying over
multiple graphs. Kypher also does not yet support certain features such as
pathrange patterns, subqueries or unions which are planned as future extensions.</p>
      <p>To implement Kypher queries, we translate them into SQL and execute them
on SQLite, a lightweight le-based SQL database. Kypher queries are designed to
look and feel very similar to other le-based KGTK commands. They take
tabular le data as input and produce tabular data as output. There are no servers or
accounts to set up, and users do not need to know that there is in fact a database
used underneath to implement the queries. A cache mechanism makes multiple
queries over the same KGTK les very e cient. Kypher has been successfully
tested on Wikidata-scale graphs with 1.5B edges where queries executing on
a standard laptop run in milliseconds to minutes depending on selectivity and
result sizes. Additional information about Kypher and its capabilities can be
found online. 7
3 https://graph-tool.skewed.de
4 https://sqid.toolforge.org/
5 https://opencypher.org/
6 https://www.gqlstandards.org/home
7 https://kgtk.readthedocs.io/en/latest/transform/query/#overview</p>
      <p>Chalupsky et. al.</p>
    </sec>
    <sec id="sec-4">
      <title>Use Cases</title>
      <p>
        This section presents ve use cases that illustrate di erent ways to exploit the
data in Wikidata. We show how the use cases can be implemented using Kypher
queries and executed on a laptop. The equivalent implementations of the queries
in SPARQL have been tested against the public Wikidata SPARQL endpoint,
and against a large server with a 24 hour time out limit. The use cases have been
implemented in Python for Kypher (using a Jupyer Notebook) and SPARQL
(using a script) and are available online.8 All input datasets are available in
Zenodo under a public DOI [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. In the paper we illustrate each use case using
one Kypher query from the notebook.
4.1
      </p>
      <sec id="sec-4-1">
        <title>Retrieval of large amounts of data from Wikidata</title>
        <p>John is doing research on the popularity of rst names to improve his entity
resolution algorithm for people. He sees that Wikidata contains about 9 million
people, so he wants to get the distribution of counts of rst names from Wikidata.
He writes a SPARQL query, but it times out, so he downloads and loads the
Wikidata KGTK les on his laptop (which take approx. 98 minutes with the
latest dump) and writes a Kypher query. The query, shown in Figure 1, retrieves
all instances of human (Q5), gets their rst names using the P735 property ( rst
name) and returns the counts.
!$kypher -i items -i p31 -i labels
--match '
p31: (person)-[:P31]-&gt;(:Q5), # Q5 is person
items: (person)-[:P735]-&gt;(given_name), # P735 is first name
labels: (given_name)-[:label]-&gt;(given_name_label)'
--return 'distinct given_name as node1, count(given_name) as node2,
given_name_label as `node1;label`, "count_names" as label'
--order-by 'node2 desc'
-o "$OUT"/given-names.tsv</p>
        <p>John chooses standard KGTK names for the headers to generate the data
as a KGTK graph so that the output of the query is in the same format as all</p>
        <sec id="sec-4-1-1">
          <title>8 https://github.com/usc-isi-i2/kgtk-at-2021-wikidata-workshop/</title>
          <p>other KGTK les and can be used as input in future queries. In KGTK, triples
are represented using node1, label (i.e., property), and node2 headings. It is
also possible to include triples about node1 and node2 as additional columns by
using the semicolon notation. For example, node1;label represents the label
property of node1 (in KGTK we represent the Wikidata label attribute using
the label property).</p>
          <p>This example computes results over a large number of entities producing a
potentially large set of results. This use case cannot be handled using the public
Wikidata endpoint due to a timeout, while the Kypher query runs on a laptop
(8 minutes, 16 seconds). Our companion Jupyter notebook9 illustrates how this
query can be extended to measure the popularity of rst names over time.
4.2</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>Analytics on the full Wikidata</title>
        <p>Jessica is working with John on the entity resolution algorithm and her job is
to use the number of instances of each class in Wikidata as a feature. Jessica
just needs to count the number of instances of each class, summing over the
instances of all subclasses. She knows that there are over 1 million classes in
Wikidata (entities with a P279 property), so she knows it will not run on the
public SPARQL endpoint. Jessica copies the SQLite database from John so that
she does not have to wait the 98 minutes John had to wait to load the needed
Wikidata les on her laptop.
!$kypher -i p31 -i p279star
--match '
p31: (entity)-[:P31]-&gt;(class),
p279star: (class)-[:P279star]-&gt;(super_class)'
--return 'distinct super_class as node1, count(distinct entity) as
node2, "entity_count" as label'
--order-by 'node2 desc, node1'
-o "$OUT"/class.count.tsv.gz
node1
Q35120
Q99527517
Q488383</p>
        <p>The query (Figure 2) uses two les from the KGTK distribution of Wikidata.
The p31 le records the class of every instance, and the p279star le records
all the super-classes of every class using a new property called P279star. These
9
https://github.com/usc-isi-i2/kgtk-at-2021-wikidata-workshop/blob/main/wikidataworkshop-2021.ipynb
properties are commonly used so they are provided as separate les for the
convenience of users. The query retrieves the class from every entity from the
P31 le, retrieves all the super-classes of every class, and returns the entity count
for every super-class.
!$kypher -i p279star -i labels -i "$OUT"/class.count.tsv.gz --as count
--match '
p279star: (class)-[]-&gt;(:Q11424), # Q11424 is film
count: (class)-[:entity_count]-&gt;(count),
labels: (class)-[:label]-&gt;(class_label)'
--return 'class as node1, class_label as `node1;label`, count as node2'
--order-by 'cast(count, integer) desc'
--limit 10
node1
Q11424
Q24862
Q506240
node2
314889
33733
17310</p>
        <p>After coming back from lunch, the le is ready. The Kypher query ran in
88.97 minutes on Jessica's laptop, and contains data for 75K classes, as there
are many classes that do not have instances. The equivalent SPARQL query
timed out on the public SPARQL endpoint and did not complete after 24 hours
on our private SPARQL endpoint (see below for the reason why).</p>
        <p>Jessica is curious about the data, and wants to know the instance count for
all subclasses of lm (Q11424). Jessica could modify the query above to include
only instances of lm with p31: (entity)-[:P31]-&gt;(:Q11424), and the query
would return the counts. However, the output le she already computed has all
the data she needs, and it is a valid KGTK graph, so Jessica writes a query to
pick out the subset that she is interested in (Figure 3). She writes a query that
uses the output of the previous query (class.count.tsv.gz) as input, uses the
p279star le to get all the subclasses of lm including indirect subclasses, and
fetches the count from the counts le.</p>
        <p>The query returns results in 2.4 seconds, illustrating how Kypher makes it
easy to chain the results of queries, avoiding recomputation of expensive queries
to get results. The lm version of the query times out on the public SPARQL
endpoint and takes 114 seconds on the private SPARQL endpoint as there is
no easy way to reuse the results of the previous computation in a new query.
Jessica's Kypher query is e cient because she built a personalized version of
Wikidata on her laptop, choosing to add the entity count property to her KG
to make other queries run quickly. Using the le as input to her query was all
that Jessica had to do to add the data to her personalized version of Wikidata.
Kypher automatically loaded and indexed the data; in addition, Kypher will
check whether the les on disk have changed every time it runs a query that
uses the le, and will automatically reload and re-index the data as necessary.</p>
        <p>This use case is a further example of the previous use case where users want
to derive new data from Wikidata using queries that involves a large proportion
of the entities in Wikidata. KGTK supports these use cases by making it possible
to decompose complex use cases into independent queries, and allowing users to
build personalized versions of Wikidata where they extend Wikidata with the
results of previous queries. The p27star le that KGTK provides is also an
example of this capability, and is the reason why the instances query can run
in 80 minutes on a laptop but cannot produce results after 24 hours on a large
256GB server running SPARQL.
4.3</p>
      </sec>
      <sec id="sec-4-3">
        <title>Extraction of new graphs from Wikidata</title>
        <p>Bill is working on a project to nd networks of researchers working on speci c
topics. He wants to use publication data to nd relationships among authors.
Bill knows that he can get lots of publication data from Pubmed or Microsoft
Academic graph, but wants to give Wikidata a try as he heard that Wikidata
has close to 40 million publications, and that in Wikidata publications have links
to other entities such as main subjects.
!$kypher -i p31 -i p279star -i items -i time -i labels
--match '
p31: (pub)-[:P31]-&gt;(class),
p279star: (class)-[:P279star]-&gt;(:Q591041), # node for scientific publication
items: (pub)-[:P50]-&gt;(author1), # P50 is author
items: (pub)-[:P50]-&gt;(author2)'
--where 'author1 &gt; author2'
--return 'distinct author1 as node1, "Pcoauthor" as label,
author2 as node2, count(distinct pub) as count_publications'
--order-by 'count_publications desc'
-o "$TEMP"/coauthors.2019.tsv.gz</p>
        <p>Bill decides that the simplest experiment to try rst is to build a network
of authors of publications in Wikidata: he wants to create a graph of people in
Wikidata who authored papers, to put a link between two people if they
coauthored a paper, and to add a quali er with the count of papers they coauthored.
He knows the computation is expensive as there are around 40 million papers in
Wikidata, so the network will be large. He doesn't even try to write a SPARQL
query because he knows it will time out. Bill downloads the KGTK les and
writes the query shown in Figure 4. The query reuses the p31 and p279star
!$kypher -i p31 -i p279star -i items -i labels
--match '
p31: (pub)-[:P31]-&gt;(class),
p279star: (class)-[:P279star]-&gt;(:Q591041), # scientific publication
items: (pub)-[:P50]-&gt;(author1), # P50 is author
items: (pub)-[:P50]-&gt;(author2),
items: (pub)-[:P921]-&gt;(cancer_type), # P921 is main subject
p279star: (cancer_type)-[:P279star]-&gt;(:Q12078), # Q12078 is cancer
labels: (author1)-[:label]-&gt;(author1_label),
labels: (author2)-[:label]-&gt;(author2_label)'
--where 'author1 &gt; author2'
--return '
distinct author1 as node1, "Pcoauthor" as label, author2 as node1,
count(distinct pub) as count_publications,
author1_label as `node1;label`, author2_label as `node2;label`'
--order-by 'count_publications desc'
-o "$TEMP"/coauthors.cancer.tsv.gz</p>
        <p>les to retrieve all publications that are instances of any subclass of Q591041
(scienti c publication). He uses the P50 property to retrieve the authors and
uses two variables (author1 and author2) to retrieve multiple authors if they
are present. A where clause ensures that the variables are bound to di erent
authors, and the return clause constructs the output edges and quali er. Bill
uses the standard node1 and node2 headings to construct triples, using a new
property Pcoauthor. He also invents a new quali er count publications and
includes the labels of the authors so that he can read the output.</p>
        <p>Bill continues his investigation. He is interested in cancer research, so he
wants to build the same network but using only the papers about cancer. He
extends the query using the P921 property to retrieve the main subjects of a
paper, and again uses the p279star le to select subjects that are below cancer
(Q12078). He expects the query to be much faster because now it has strong
restrictions, so he gives it a try. The query, shown in Figure 5, takes 2.62 minutes
and produces a network with close to half a million edges.
4.4</p>
      </sec>
      <sec id="sec-4-4">
        <title>Queries combining multiple resources</title>
        <p>Abigail is working on a cultural heritage project, collaborating with the Getty
Research Institute who gave her a le with 27 thousand identi ers of artists
that she is interested in; the le has one identi er per line. The Getty uses
!$kypher -i items -i external_ids -i labels -i "$OUT"/ulan.tsv
--match '
ulan: (ulan_id)-[]-&gt;(),
# P214 is VIAF ID, P245 is Union List of Artist Names ID
external_ids: (viaf_id)&lt;-[:P214]-(artist)-[:P245]-&gt;(ulan_id),
labels: (artist)-[]-&gt;(artist_label)'
--return '
artist as node1, viaf_id as node1;P214, ulan_id as node1;P245,
artist_label as node1;label'
-o "$OUT"/ulan-to-viaf.tsv
node1 node1;P214 node1;P245 node1;label
Q1000596 "20822441" "500072302" 'Noemi Ferenczy'@en
Q1001063 "96418002" "500099612" 'Olga Fialka'@en
Q100156272 "309815799" "500335625" 'Gloria Lopez Cordova'@en</p>
        <p>ULAN identi ers10, and Abigail has a database indexed using VIAF identi ers.11
Abigail needs to map the ULAN identi ers to VIAF identi ers so that she can
use her database. She puts one of the ULAN identi ers in the Wikidata search
box and discovers that Wikidata has both ULAN and VIAF identi ers for many
artists, so she needs to write a query that retrieves artists using the ULAN
identi ers and returns the VIAF identi er when it is available in Wikidata.</p>
        <p>Abigail considers using SPARQL as it is easy to write a query to retrieve the
VIAF identi er given a ULAN identi er. This solution would require sending
27,000 queries to Wikidata (or 27 queries binding 1000 identi ers), and would
involve writing a script. She writes a script to produce a SPARQL query that
binds all 27,000 identi ers, but the query is too large and it is rejected in the
public SPARQL endpoint.</p>
        <p>Abigail solves the problem using Kypher. She rst renames the heading of
the ulan.tsv le to node1, so the le is now a valid KGTK graph le because
in KGTK any value can be used as node1 (subject). Her ULAN KGTK graph
contains 27,000 nodes and no edges for any node. She writes the Kypher query
shown in Figure 6. The query uses the ulan.tsv le as input, thereby
personalizing the Wikidata graph to include the ULAN identi er nodes. The query binds
the ulan id variable to the nodes in the ulan.tsv graph. The next clause uses
the external ids graph from the KGTK distribution of Wikidata to map the
ULAN ids to VIAF ids. She returns the data by using standard KGTK headers
so that she can use the resulting le in other queries. The query runs in 11.8
seconds and retrieves 8,116 VIAF ids.</p>
        <p>Wikidata has become a hub for identi ers as it contains a large number
of identi ers for entities (over 160 million identi ers). Abigail's use case is an
10 https://www.getty.edu/research/tools/vocabularies/ulan/
11 http://viaf.org
example of a common use case to exploit the Wikidata identi ers: a researcher
has an external source that contains identi ers present in Wikidata and wants
to retrieve the entities or map one type of identi er to another. Kypher queries
address this use case as external resources can be easily converted to KGTK
graphs and used in queries.
4.5</p>
      </sec>
      <sec id="sec-4-5">
        <title>Combination of Wikidata and DBpedia</title>
        <p>
          After mining her VIAF database, Abigail realizes that she needs more data
and wants to exploit the Wikipedia infoboxes. Abigail considers using SPARQL
federated queries [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] to combine Wikidata and DBpedia, but she faces the same
problem as before in that she has 27,000 identi ers and may need to issue a large
number of queries. Abigail downloads the DBpedia infobox data in RDF format
from the DBpedia Databus12 and uses KGTK commands to convert the data
to KGTK format. The resulting KGTK le contains close to 100 million edges
but the data is noisy as illustrated in the following excerpt. Abigail expects the
node2 column for these properties to contain Wikidata q-nodes, but sees that
often, literals are present.
        </p>
        <p>node1 label
Q466241 property:almaMater
Q466241 property:occupation
Q466241 property:spouse
!$kypher -i infobox -i p31 -i labels
--match '
infobox: (artist)-[:`property:spouse`]-&gt;(spouse),
p31: (spouse)-[]-&gt;(:Q5)'
--opt 'labels: (spouse)-[:label]-&gt;(spouse_label)'
--return 'artist as node1, "P26" as label, spouse as node2,
spouse_label as `node2;label`'
-o "$OUT"/spouses.dbpedia.qnodes.tsv</p>
        <p>Abigail is interested in the spouse data for her artists and writes the query
shown in Figure 7. The query uses as input an external le (infobox) and a
Wikidata le (p31), retrieves the spouse from the DBpedia le and veri es that
the value of spouse is and instance of human (Q5).
12 https://databus.dbpedia.org/dbpedia/</p>
        <p>Kypher loads the 100 million edge DBpedia le in 10.6 minutes and runs the
query in 3:4 minutes, a fast time considering that there are 322; 599 spouse edges
in the DBpedia graph, and all must be checked to be instances of human (Q5).
The query identi es 7; 325 high quality spouse statements in DBpedia infoboxes
that are not present in Wikidata.</p>
        <p>This use case illustrates building a personalized Wikidata extension that
augments Wikidata with noisy data and then uses Kypher to extract clean data.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Experiments</title>
      <p>We performed experiments to compare the execution times of the Kypher and
SPARQL implementations of the queries for the use cases presented in this
paper. We omitted the Wikibase API from our evaluation, as it does not support
the types of queries required by the use cases. We used two con gurations for
Kypher, MacBook Pro laptops with 16GB memory/256GB SSD disk and 32 GB
memory/1TB SSD. We used two con gurations for SPARQL, the public
Wikidata SPARQL endpoint and a local clone of Wikidata (June 2019) running on
a server with 24 Intel Xeon cores and 256 GB of memory and SSD. For the
Kypher queries we used the Wikidata February 15, 2021 distribution converted
to KGTK format. For the Wikidata clone we used the RDF dump from the June
15, 2019 distribution (we did not load the February 15, 2021 distribution as it
takes several days to load, and for the purpose of our experiments the earlier
dump is adequate as it is smaller).</p>
      <p>The Jupyter notebook for the Kypher queries was run twice. The rst run
of the notebook with an empty SQLite database took 349 minutes in the 32GB
laptop, and the second run of the notebook, after the data was loaded and indices
were built, took 164 minutes. The di erence, 185 minutes includes 98 minutes
to load the Wikidata data, 10 minutes to load the DBpedia infobox data, and
the rest, 77 minutes is time that Kypher used to build database indices.</p>
      <p>Table 1 shows the runtimes of the queries presented in the use cases. The
times are from the second run of the notebook after the data was loaded in the
Kypher SQLite database and indices had been created.</p>
    </sec>
    <sec id="sec-6">
      <title>Discussion and Conclusions</title>
      <p>The main objective of KGTK and Kypher is to democratize the exploitation of
Wikidata so that anyone with modest computing resources can take advantage of
the vast amounts of knowledge present in Wikidata. Our tools focus on use cases
that use large portions of Wikidata to distill new knowledge. The experiments
show how expensive queries (e.g., class instances use case, Section 4.2) that
cannot run in one day on a powerful server, complete in about one hour on a
laptop. Analytic queries ( rst names use case, Section 4.1) become possible on a
laptop in a few minutes, and distillation of knowledge for analysis (author and
cancer network use case, Section 4.3) become practical. Kypher enables users
to easily combine Wikidata with external sources to extract relevant Wikidata
knowledge (ULAN use case, Section 4.4), or to enhance Wikidata with knowledge
from external sources (DBpedia use case, Section 4.5).</p>
      <p>Kypher is not meant to address use cases that require the most up-to-date
data in Wikidata. KGTK uses the Wikidata JSON dumps published every few
days, and the KGTK work ow to process the JSON dump takes one day.</p>
      <p>The comparison with the Wikidata SPARQL endpoints is preliminary as we
have not controlled for caching in the triple store and in the operating system, or
performed systematic variations of the complexity of the queries. A more detailed
and controlled comparison is reserved for a future paper. Here we speculate on
the reasons why Kypher seems to perform signi cantly better than the Wikidata
SPARQL endpoints on the presented use cases:
1. Compact data model: the KGTK data model allows us to translate 1.2B
Wikidata statements very directly into 1.2B edges, while the RDF
translation requires rei cation and generates O (10B) triples. KGTK also does not
require the use of namespaces which makes data values more compact.
2. Smaller database size: more compact data translates directly into smaller
database sizes, for example, 142GB for the Kypher graph cache vs. 718GB
for the local Wikidata endpoint. This gives generally better locality for table
and index lookups and better caching of data pages.
3. Specialized tables: representing specialized data slices such as P279star in
their own graph tables makes their reuse very e cient and their indexes
more focused, compact, and cache-friendly.
4. Read-only processing: Kypher does not need to support ne-grained updates
of tables and indexes, which need to be supported by the public Wikidata
endpoint. This requires additional machinery that slows down performance.
5. Use case selection: triple stores and databases are optimized to support a
large number of use cases. Our set of use cases samples a small slice of that
space, and performance might be very di erent for other types of queries.</p>
      <p>The contribution of this paper is to show that KGTK and Kypher are e
ective tools for complex analytic use cases. The paper demonstrates that Kypher
supports a variety of use cases cases that are impractical with existing tooling.
Kypher allows researchers and developers to investigate use cases on their own
laptop, exploring extensions of Wikidata that would not be possible on shared
resources, with minimal setup, using a simple query language.</p>
      <p>Our experience with Kypher so far shows that new users, e.g., computer
science students, are able to quickly grasp the language, especially when they
have prior experience with SPARQL or SQL. In the future, we will perform a
formal user study in order to quantify the e ort needed to get familiar with
Kypher. We will also benchmark the performance of KGTK and Kypher against
multi-processing parsing and processing of the Wikidata dumps.</p>
      <p>Acknowledgements: This material is based on research sponsored by Air
Force Research Laboratory under agreement number FA8750-20-2-10002. The
U.S. Government is authorized to reproduce and distribute reprints for
Governmental purposes notwithstanding any copyright notation thereon. The views
and conclusions contained herein are those of the authors and should not be
interpreted as necessarily representing the o cial policies or endorsements, either
expressed or implied, of Air Force Research Laboratory or the U.S. Government.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Auer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kobilarov</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lehmann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cyganiak</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ives</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          :
          <article-title>Dbpedia: A nucleus for a web of open data</article-title>
          .
          <source>In: The semantic web</source>
          , pp.
          <volume>722</volume>
          {
          <fpage>735</fpage>
          . Springer (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Francis</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Green</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guagliardo</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Libkin</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lindaaker</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marsault</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Plantikow</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rydberg</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Selmer</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Taylor</surname>
          </string-name>
          , A.:
          <article-title>Cypher: An evolving query language for property graphs</article-title>
          .
          <source>In: Proceedings of the 2018 International Conference on Management of Data</source>
          . pp.
          <volume>1433</volume>
          {
          <issue>1445</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Ilievski</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garijo</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chalupsky</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Divvala</surname>
            ,
            <given-names>N.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yao</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rogers</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , Liu,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Schwabe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Szekely</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.</surname>
          </string-name>
          :
          <article-title>Kgtk: a toolkit for large knowledge graph manipulation and analysis</article-title>
          .
          <source>In: International Semantic Web Conference</source>
          . pp.
          <volume>278</volume>
          {
          <fpage>293</fpage>
          . Springer (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Wes</surname>
            <given-names>McKinney</given-names>
          </string-name>
          :
          <article-title>Data Structures for Statistical Computing in Python</article-title>
          . In: Stefan van der Walt, Jarrod Millman (eds.)
          <source>Proceedings of the 9th Python in Science Conference</source>
          . pp.
          <volume>56</volume>
          {
          <issue>61</issue>
          (
          <year>2010</year>
          ). https://doi.org/10.25080/Majora-92bf1922-00a
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Quilitz</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leser</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          :
          <article-title>Querying distributed rdf data sources with sparql</article-title>
          .
          <source>In: European semantic web conference</source>
          . pp.
          <volume>524</volume>
          {
          <fpage>538</fpage>
          . Springer (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Seaborne</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harris</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>SPARQL 1.1 query language</article-title>
          .
          <source>W3C recommendation</source>
          ,
          <source>W3C (Mar</source>
          <year>2013</year>
          ), https://www.w3.org/TR/2013/REC-sparql11
          <string-name>
            <surname>-</surname>
          </string-name>
          query-20130321/
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Szekely</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Datasets for Creating and Querying Personalized Versions of Wikidata in a Laptop (Jul</article-title>
          <year>2021</year>
          ). https://doi.org/10.5281/zenodo.5139550
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <article-title>8. the pandas development team: pandas-dev/pandas: Pandas (Feb</article-title>
          <year>2020</year>
          ). https://doi.org/10.5281/zenodo.3509134
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Vrandecic</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , Krotzsch, M.:
          <article-title>Wikidata: a free collaborative knowledgebase</article-title>
          .
          <source>Communications of the ACM</source>
          <volume>57</volume>
          (
          <issue>10</issue>
          ),
          <volume>78</volume>
          {
          <fpage>85</fpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Wood</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cyganiak</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lanthaler</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>RDF 1.1 concepts and abstract syntax</article-title>
          .
          <source>W3C recommendation</source>
          ,
          <source>W3C (Feb</source>
          <year>2014</year>
          ), https://www.w3.org/TR/2014/RECrdf11-concepts-20140225/
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>