<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>⋆ Generating RDF for Application Testing</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Daniel Blum</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sara Cohen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>daniel.blum@mail</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>sara@cs}.huji.ac.il</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>School of Computer Science and Engineering The Hebrew University of Jerusalem</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Application testing is a critical component of application development. Testing of Semantic Web applications requires large RDF datasets, conforming to an expected form or schema, and preferably, to an expected data distribution. Finding such datasets often proves impossible, while generating input datasets is often cumbersome. The GRR (Generating Random RDF) system is a convenient, yet powerful, tool for generating random RDF, based on a SPARQLlike syntax. In this poster and demo, we show how large datasets can be easily generated using intuitive commands.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        RDF (SIMILE [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], RBench [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]) and none of these programs can produce data that
conforms to a specific given structure, and thus, again, will not have the specified properties.
      </p>
      <p>In this demo, we present the GRR (Generating Random RDF) system for generating
RDF that satisfies both desirable properties given above. Thus, GRR is not a benchmark
system, but rather, a system to use for Semantic Web application testing. Using intuitive
data generation commands with a SPARQL-like syntax, GRR can produce data with a
complex graph structure, as well as draw the data values from desirable domains. Data
generation commands are translated into a series of SPARQL queries and update
commands which are applied directly to an RDF storage system.1 A video demonstration of
GRR is available online,2 and the system is available upon request.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Motivating Example</title>
      <p>
        As a motivating example, we discuss the problem of generating the data described in
the LUBM Benchmark. Note that GRR is not limited to creating benchmark data. In our
demo, we will demonstrate using GRR to generate other types of data, such as FOAF [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
(Friend of a Friend) datasets, which are used in social network applications.
      </p>
      <p>
        LUBM [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] is a collection of data describing university classes (i.e., entities), such
as departments, faculty members, students, courses, etc. These classes have a plethora
of properties (i.e., relations) between them, e.g., faculty members work for departments
and head departments, students take courses and are advised by faculty members, etc.
      </p>
      <p>In order to capture a real-world scenario, LUBM defines interdependencies between
the entities. For example, the number of students in a department is a function of the
number of faculty members. Specifically, LUBM requires there to be a 1:8-14 ratio of
faculty members to undergraduate students. As another example, the cardinality of a
property may be specified, such as each department must have a single head of
department (who must be a full professor). Properties may also be required to satisfy
additional constraints, e.g., courses, taught by faculty members, must be pairwise disjoint.</p>
      <p>
        In the next section, we describe the GRR data generation language, and demonstrate
commands for producing LUBM benchmark data. Due to space limitations, we do not
provide all commands used to reproduce LUBM. However, we note that the number of
words needed in all data generation commands (in order to reproduce LUBM), is only
about twice as many as used in the intuitive description of LUBM, provided by [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]!
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Data Generation Commands</title>
      <p>Data is generated by a sequence of data generation commands (dg-commands, for
short) c1, . . . , cn, when given as input a (possibly empty) RDF dataset R. The first
command c1 is evaluated over R, while each consecutive command ci is evaluated over
the output of the previous command ci−1.</p>
      <p>
        The general syntax of a single dg-command appears below. Note that square
brackets are used to denote optional portions, and the “*” indicates a component that can
appear any number of times.
1 The Jena Semantic Web Framework for Java [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] is used in our implementation.
2 http://www.cs.huji.ac.il/˜danieb12/
(FOR (EACH | sampling-method)
[WITH (GLOBAL DISTINCT | LOCAL DISTINCT | REPEATABLE)]
{list of classes}
[WHERE {list of conditions}] )*
[CREATE i-j {list of classes}]
[CONNECT {list of connections}]
      </p>
      <p>
        A dg-command contains any number of FOR clauses, and then optionally a CREATE
and/or CONNECT clause. Intuitively, the FOR clauses choose portions of the RDF input,
the CREATE clause creates new nodes in the RDF graph, and the CONNECT clause
connects nodes in the RDF graph. We require that at least one among the CREATE
and CONNECT clauses be present in every dg-command. We now describe each clause,
briefly. (Full language semantics appears in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]).
      </p>
      <p>– The FOR Clause: Each FOR clause defines (1) a query which will applied against
the RDF input, as well as (2) a method to choose a subset of the query results. For
(1), the user provides a list of classes whose instances should be chosen (similar to
a SPARQL SELECT clause), as well as any conditions (similar to a SPARQL WHERE
clause). The correspondence to SPARQL is not precise as we allow for certain
syntactic shortcuts, which avoid explicit variable use, and make dg-commands more
readable. For (2), the user defines both the method with which answers should be
sampled, as well as whether the sampling process is with/without repetition.
– The CREATE Clause: The CREATE clause defines nodes that should be created.</p>
      <p>The user provides both a list of RDF classes, and a range determining how many
instances of these classes should be created.3
– The CONNECT Clause: The CONNECT clause determines the edges that should be
generated in the RDF graph, by providing a list of triples.</p>
      <p>Several examples of dg-commands appear below. Explanations follow.
(c1) CREATE 1-5 {ub:Univ}
(c2) FOR EACH {ub:Univ}</p>
      <p>CREATE 15-25 {ub:Dept}</p>
      <p>CONNECT {ub:Dept ub:subOrg ub:Univ}
(c3) FOR EACH {ub:Faculty, ub:Dept}</p>
      <p>WHERE {ub:Faculty ub:worksFor ub:Dept}</p>
      <p>CREATE 8-14 {ub:Undergrad}</p>
      <p>CONNECT {ub:Undergrad ub:memberOf ub:Dept}
(c4) FOR EACH {ub:Dept}</p>
      <p>FOR 1 {ub:FullProf}
WHERE {ub:FullProf ub:worksFor ub:Dept}</p>
      <p>CONNECT {ub:FullProf ub:headOf ub:Dept}
3 Dg-commands do not directly define how textual (or other atomic) properties are created and
associated with class instances. This information is provided in a simple auxilliary file, e.g.,
which associates each textual property with a sampling method or dictionary.
(c5) FOR 20%-20% {ub:Undergrad, ub:Dept}</p>
      <p>WHERE {ub:Undergrad ub:memberOf ub:Dept}</p>
      <p>FOR 1 {ub:Prof}
WHERE {ub:Prof ub:memberOf ub:Dept}</p>
      <p>CONNECT {ub:Undergrad ub:advisor ub:Prof}
(c6) FOR EACH {ub:Undergrad}</p>
      <p>FOR 2-4 WITH LOCAL DISTINCT {ub:UndergradCourse}</p>
      <p>CONNECT {ub:Undergrad ub:takeCourse ub:UndergradCourse}
(c7) FOR EACH {foaf:Person ?p1}</p>
      <p>FOR 15-25 {foaf:Person ?p2} WHERE {FILTER( ?p1 != ?p2 )}
CONNECT {?p1 foaf:knows ?p2}</p>
      <p>Command c1 creates between 1 and 5 universities, and command c2 adds 15–25
departments as suborganizations for each university. Command c3 iterates over all pairs
of faculty members4 and departments, and adds 8-14 students, per pair to the
department (therby achieving the required 1:8-14 ratio of faculty members to undergraduates).
Command c4 chooses one full professor as the head of each department. Command c5
adds an advisor for 20% of all undergraduates. Command c6 assigns 2-4 courses for
each undergraduate. Note the use of WITH LOCAL DISTINCT which ensures that
the set of courses chosen per student does not contain repetition, while allowing
different students to be assigned the same courses. Finally, c7 demonstrates advanced features
including variables and a filter command, to connect people (in an FOAF RDF dataset)
to one another.</p>
      <p>In our poster and demo, we will show how to recreate the LUBM benchmark using
24 dg-commands, of the style seen above. In addition, we will show how to create
interesting datasets for the FOAF schema. We will also allow those interested to write
their own dg-commands, which we will evaluate in GRR to create an RDF dataset.
4 The faculty members were created with an additional dg-command, which was omitted due to
lack of space.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schultz</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>The Berlin SPARQL benchmark</article-title>
          .
          <source>International Journal of Semantic Web Information Systems</source>
          <volume>5</volume>
          (
          <issue>2</issue>
          ),
          <fpage>1</fpage>
          -
          <lpage>24</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Blum</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cohen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          : Grr:
          <article-title>Generating random RDF</article-title>
          .
          <source>Tech. rep.</source>
          , The Hebrew University of Jerusalem (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <article-title>3. The friend of a friend (FOAF) project</article-title>
          . http://www.foaf-project.org
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Guo</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pan</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heflin</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>LUBM: a benchmark for OWL knowledge base systems</article-title>
          .
          <source>Journal of Web Semantics</source>
          <volume>3</volume>
          (
          <issue>2-3</issue>
          ),
          <fpage>158</fpage>
          -
          <lpage>182</lpage>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <article-title>Jena-a Semantic Web framework for Java</article-title>
          . http://jena.sourceforge.net
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>6. RBench website. http://139.91.183.30:9090/RDF/RBench/index.html</mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Schmidt</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hornung</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lausen</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pinkel</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>SP2Bench: a SPARQL performance benchmark</article-title>
          .
          <source>In: ICDE</source>
          . pp.
          <fpage>222</fpage>
          -
          <lpage>233</lpage>
          . Shanghai, China (Mar
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>8. Simile website. http://simile.mit.edu/</mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>