<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Scalable graph analytics with GRADOOP</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Erhard Rahm</string-name>
          <email>rahm@informatik.uni-leipzig.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Leipzig Augustusplatz 10 Leipzig</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2015</year>
      </pub-date>
      <abstract>
        <p />
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Many Big Data applications in business and science require
the management and analysis of huge amounts of graph
data. Previous approaches for graph analytics such as graph
databases and parallel graph processing systems (e.g., Pregel)
either lack sufficient scalability or flexibility and
expressiveness. We are therefore developing a new end-to-end
approach for graph data management and analysis at the Big
Data center of excellence ScaDS Dresden/Leipzig. The
system is called Gradoop (Graph analytics on Hadoop). Gradoop
is designed around the so-called Extended Property Graph
Data Model (EPGM) which supports semantically rich,
schemafree graph data within many distinct graphs. A set of
highlevel operators is provided for analyzing both single graphs
and sets of graphs. The operators are usable within a
domainspecific language to define and run data integration
workflows (for integrating heterogeneous source data into the
Gradoop graph store) as well as analysis workflows. The
Gradoop data store is currently utilizing HBase for distributed
storage of graph data in Hadoop clusters. An initial version
of Gradoop is operational and has been used for analyzing
graph data for business intelligence and social network
analysis.</p>
      <p>About the Author
Erhard Rahm is full professor for databases at the
computer science institute of the University of Leipzig. His
current research focusses on Big Data and data
integration. He has authored several books and more than 200
peer-reviewed journal and conference publications. His
research on data integration and schema matching has been
awarded several times, in particular with the renowned
10year best-paper award of the conference series VLDB (Very
Large Databases) and the Influential Paper Award of the
conference series ICDE (Int. conf. on Data Engineering).
Prof. Rahm is one of the two scientific coordinators of the
new German center of excellence on Big Data ScaDS
(competence center for SCAlable Data services and Solutions)
Dresden/Leipzig that started its operation in Oct. 2014.</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>