<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Next Generation Data Integration (for the Life Sciences)</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ulf Leser</string-name>
          <email>leser@informatik.hu-berlin.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Humboldt-Universität zu Berlin Institute for Computer Science</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2014</year>
      </pub-date>
      <abstract>
        <p />
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Ever since the advent of high-throughput biology (e.g., the
Human Genome Project), integrating the large number of
diverse biological data sets has been considered as one of
the most important tasks for advancement in the
biological sciences. The life sciences also served as a blueprint
for complex integration tasks in the CS community, due to
the availability of a large number of highly heterogeneous
sources and the urgent integration needs. Whereas the early
days of research in this area were dominated by virtual
integration, the currently most successful architecture uses
materialization. Systems are built using ad-hoc techniques and
a large amount of scripting. However, recent years have
seen a shift in the understanding of what a "data
integration system" actually should do, revitalizing research in this
direction. In this tutorial, we review the past and current
state of data integration (exempli ed by the Life Sciences)
and discuss recent trends in detail, which all pose challenges
for the database community.</p>
      <p>About the Author
Ulf Leser obtained a Diploma in Computer Science at the
Technische Universitat Munchen in 1995. He then worked as
database developer at the Max-Planck-Institute for
Molecular Genetics before starting his PhD with the Graduate
School for "Distributed Information Systems" in Berlin. Since
2002 he is a professor for Knowledge Management in
Bioinformatics at Humboldt-Universitat zu Berlin.</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>