<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Data intensive analysis approaches in genomics and proteomics: ELIXIR initiatives</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Proceedings of the XVII International Conference «Data Analytics and Management in Data Intensive Domains» (DAMDID/RCDL'2015)</institution>
          ,
          <addr-line>Obninsk</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Alexander A. Kanapin Department of Oncology, University of Oxford</institution>
          ,
          <addr-line>Oxford</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <fpage>257</fpage>
      <lpage>259</lpage>
      <abstract>
        <p>Breakthrough in genome sequencing technologies resulted in the unprecedented growth of data volumes in genomics and proteomics. New paradigm of precision medicine signifies wide practical usage of these types of data. ELIXIR, a pan-European bioinformatics consortium meets the challenges arising from production, storage and analysis of massive data collections in genomics and proteomics and proposes several pilot programs, which aim to develop standards and algorithms for the data analysis. The interdisciplinary initiatives of the consortium, such as "BILS-ProteomeXchange integration using EUDAT resources” and "Interoperability of protein resources for drug discovery: Improving Links Between the Human Protein Atlas (HPA) and EMBL-EBI Protein Resources” are of great interest and their successful implementation requires collaboration of researchers and IT engineers. The article also describes general principles of the consortium organization and potential ways of participation in its collaboration projects and programs.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>Biology traditionally was a science based on
quantitative observations, and in contrast to physics, it
produced relatively small amounts of qualitative data.
The situation dramatically changed in a last quarter of
XX century. A rapid progress in new technologies of
analysis of living systems (cells and organisms) on
molecular level resulted in a burst of data, primarily
describing features of biological molecules, such as
nucleic acids and proteins. A matching appearance of
personal computers and global networks facilitated the
storage and processing of such information in both
small and large scale.</p>
      <p>
        As a result, a new discipline emerged in 1989, when
the term “bioinformatics” was mentioned in a title of a
scientific paper for the first time [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. The first databases
of primary structures of nucleic acids [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and proteins
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] were published in 1985 and 1991 respectively. From
the very beginning and up to present time, the majority
of the data deposited in the biological databanks
consists of sequences of biopolymers, namely nucleic
acids and proteins.
      </p>
      <p>
        A successful sequencing of human genome draft in
2001 [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] presented a next big step in the development of
bioinformatics and gave a tremendous momentum to
creation of new computational engineering solutions
and design of novel algorithms for genomic and
proteomic data analysis [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <p>A progress in biological data acquisition
technologies still remains one of major driving forces in
bioinformatics. Next Generation Sequencing (NGS)
techniques allow to obtain complete genome sequences
in a cheap and fast way. This development may change
paradigm of traditional medicine towards personal and
precise approaches to each of individual patients [11].
However, at the same time it creates new challenges in
data intensive analytics for both data storage and
manipulation technologies and algorithmic approaches.</p>
      <p>The practical solution of such tasks is only possible
in a framework of international consortia and
collaboration. ELIXIR, a pan-European consortium in
bioinformatics opens new opportunities for successful
establishment of collaboration in the pilot initiatives of
the consortium.</p>
    </sec>
    <sec id="sec-2">
      <title>2 Bioinformatics resources</title>
      <p>Data management in bioinformatics gradually
evolves with the increasing volumes of the biological
data. Historically, the protein and nucleic acids
databanks delivered the information via CD and other
similar media. Later, when networking bandwidth
allowed downloading large amounts of data, the
databases became available as downloadable flat files.
At present time the bioinformatics resources may be
classified using the following rough categories:
Data repositories. The public or commercial data
banks containing primary sequences and structures
of biopolymers. The repositories also contain tools
to analyse data provided by user in a context of the
resource. Examples: UniProt, GenBank, RSCB
PDB.</p>
      <p>Analytical toolboxes. The complex portals
providing exclusive algorithms for user data
analysis.</p>
      <p>Bioinformatics cloud resources.</p>
    </sec>
    <sec id="sec-3">
      <title>3 ELIXIR: pan-European collaboration in bioinformatics</title>
      <p>The ELIXIR consortium was founded in 2006 by
European Laboratory for Molecular Biology (EMBL).
The consortium officially started as a fully functional
body in December 2013 when the consortium
agreement was signed by the first member states. At
present it includes 12 full members and 6 observers.</p>
      <p>
        The major goal of ELIXIR is coordination of efforts
in quality control and archiving of life sciences data in
pan-European scale. The complexity of the data and its
heterogeneity calls for creation of infrastructure and
system of standards as well as development of proper
training programs. ELIXIR will act as a sustainable
repository for life science data that has been funded by
the public [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>The consortium is organized as a network of
interactions between central hub (Hinxton, UK) and
national nodes in each of the member states. The
participation in research pilot initiatives is opened to all
scientific organizations of the member states.</p>
      <p>
        Currently, ELIXIR is unfolding its activities through
series of pilot programs and initiatives. The scientific
program of the consortium proposes several research
and development avenues along the main directions of
future development of data intensive analytics in
biological sciences [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
    </sec>
    <sec id="sec-4">
      <title>4 Data intensive analytical programs in proteomics and genomics</title>
      <sec id="sec-4-1">
        <title>4.1 Integrative genomics initiatives in ELIXIR</title>
        <p>Comprehensive resources of various data modalities
in genomics is essential prerequisite for modern
research in biological sciences and translational
medicine. EMBL-EBI pioneers the initiative since the
creation of one of the first nucleotide sequences
database, EMBL-base. Now, as a part of ELIXIR
services it provides a diverse spectrum of genomics
data, the most outstanding of them are:</p>
        <p>ENA – European nucleotide archive, centred
around nucleotide sequencing. The resource
contains raw sequencing data, sequence assembly
and functional annotation of the data
EnsEMBL – unique genome annotation resource
containing high-quality integrated annotation on
vertebrate genomes. The resource comprises data
mining interface, BioMart for data retrieval.</p>
        <p>European Variation Archive – a recent
development of the novel approach to genomic
data, the database contains all types of genetic
variation data
Expression Atlas – RNA-related portal, collecting
information about gene expression patterns in</p>
        <p>different species and various biological conditions.
High quality manual curation and verification of the
information in the databases ensures the reliability of
the data available. Internal connectivity and integration
between the different resources in the Institute allows
high level of data integrity and consistency.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2 Proteomics in ELIXIR</title>
        <p>Proteomics research makes a significant part of the
consortium scientific programme. Several protein and
protein expression resources have been established in
Europe, containing valuable information for biomedical
research. Seamless navigation between these resources
is an important prerequisite for scientists to make
informed decisions about their research into new drug
targets and are exploring links between different
proteins in healthy and diseased tissues. Swedish
national node of ELIXIR plays an important role in this
action, working with EMBL-EBI. The consolidated
efforts make the Human Protein Atlas interoperable
with such proteomic resources as PRIDE, InterPro, and
the Gene Expression Atlas.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3 BILS - ProteomeXchange</title>
        <p>An arrival of tremendous volumes of biological data
calls for a need for distributed data storage and
replication and reliable and scalable data access
interface. One of the ELIXIR pilot initiatives aims to
integrate the raw data repositories for mass
spectrometry proteomics data run by Bioinformatics
Infrastructure for Life Sciences (BILS, Sweden) and
ProteomeXchange consortium via the PRIDE database,
hosted in EMBL-EBI, UK. The key point in the
infrastructure is provided by the European infrastructure
EUDAT (http://www.eudat.eu/). The ProteomeXchange
consortium facilitates submission and standardization of
dissemination practices for proteomics data resources.
The main goal of the consortium is to develop a
framework to allow standard data submission and
dissimentaion pipelines between main proteomic
repositories, such as PeptideAtlas, PRIDE and
MassIVE. The consortium encompasses 1963
proteomics datasets as of may 2015. PRIDE, one of key
participants, stores MS-based proteomics data, such as
protein expression data, post-translational
modifications, raw MS data and technical metadata.</p>
        <p>BILS is a distributed national research
infrastructure, supported by the Swedish Research
Council, its bioinformatics networks includes 6 nodes in
major Swedish universities. Proteios, a multi-user
platform for analysis and management of proteomics
data was developed as an essential part of the
integrative initiatives of BILS.</p>
        <p>EUDAT is a pan-European project aiming at
building and operating of global collaborative data
infrastructure for preserving and exchange of scientific
data in various disciplines. Essential components of its
software ecosystem, such as B2SAFE and iRODS
ensure robust, safe and highly available data access.
B2SAFE software is a key component of the
ProteomeXchange data infrastructure.</p>
        <p>The initiative could serve as an example of
engagement of various types of data storage services in
ELIXIR and demonstrate the potential of collaboration
among research infrastructures and e-infrastructures to
better manage the data deluge.</p>
      </sec>
      <sec id="sec-4-4">
        <title>4.4 Protein resources in drug discovery</title>
        <p>Important aspects of many genetic diseases are
reflected in potentially different roles of proteins and
pathways in diverse cell lineages. Interoperability
between databases providing tissue-specificity
information and describing expression of genes and
proteins in multiple tissues at different stages of
development in different diseased conditions becomes
critically important for the modern approaches in drug
discovery. The heterogeneity of the data representation
in these expression resources poses a challenge as they
often complement each other and different providers
follow different rules to annotate and provide the
information. The major goal of the ELIXIR pilot is to
define and implement standards and tools to facilitate
access and integration of the data for the scientific
community. The proteomics and expression resources in
the framework include:</p>
        <p>
          The Human Protein Atlas (HPA) [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], a database of
protein expression profiles based on
immunohistochemistry.
        </p>
        <p>
          The PRoteomics IDEntifications database (PRIDE)
[
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], a public data repository for protein and peptide
identifications.
        </p>
        <p>
          The Gene Expression Atlas (GXA) [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], an enriched
database of gene expression patterns.
        </p>
        <p>The project proposes the following integration
strategies. First, summaries of information from
different databases based on a single entry point and on
a common format will be created. The approach was
successfully introduced before by the EMBL-EBI
search portal and includes an amalgamation of service
layers on top of a database providing summary data in a
standard manner, while the original resources do not
change their data or schemas. The non-intrusive
approach ensures the independence of the original
sources and provides on demand integration. The
second approach was adopted by Biosapiens consortium
and defines a common terminology and format to
describe minimum information for specific data entries.
It provides a common language and standard format of
the data to integrate and compare protein annotations
for 39 databases. The strategy requires an agreement on
control vocabularies and changes that might affect data
content and annotation process and is therefore more
challenging task for the data providers.</p>
        <p>Distributed Annotation System (DAS) was used as
a communication fabrics to disseminate protein
expression summary data and protein sequence
annotations, as GXA and PRIDE use DAS to provide
expression data. In collaboration with HPA a new DAS
service was created to provide expression summaries.
Collaboration with other resources, such as UniProt,
PDB, pFam, InterPro, PRIDE and IntAct continues,
aiming to to create a BioJS component to standardize
the visualization of protein features which will be used
to represent related expression data such as antibody
binding and protein identifications.</p>
        <p>One of major challenges for expression
information integration among the listed sources is the
metadata annotation. The metadata harmonisation
implementation is planned as a next step, based on
Experimental Factors Ontology (EFO) as a reference
system. HPA also proposes XML solution, which is
more standardized, and flexible than DAS and might
suit better as means of data exchange.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Bairoch</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boeckmann</surname>
            <given-names>B</given-names>
          </string-name>
          .
          <article-title>The SWISS-PROT protein sequence data bank</article-title>
          .
          <source>Nucl. Ac. Res</source>
          ., v.
          <volume>19</volume>
          , p.
          <fpage>2247</fpage>
          -
          <lpage>2249</lpage>
          ,
          <year>1991</year>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Blomberg</surname>
            <given-names>N.</given-names>
          </string-name>
          <article-title>ELIXIR: Data for life</article-title>
          .
          <year>2014</year>
          , https://www.elixireurope.org/system/files/ELIXIR _2014_
          <article-title>brochure_full</article-title>
          .pdf
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Burks</surname>
            <given-names>C</given-names>
          </string-name>
          , et al.
          <article-title>The GenBank nucleic acid sequence database</article-title>
          .
          <source>Comput. Appl</source>
          . Biosci., v.
          <volume>4</volume>
          , p.
          <fpage>225</fpage>
          -
          <lpage>233</lpage>
          ,
          <year>1985</year>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Colwill</surname>
            <given-names>K</given-names>
          </string-name>
          ; Renewable Protein Binder Working Group, Gräslund S.
          <article-title>A roadmap to generate renewable protein binders to the human proteome</article-title>
          .
          <source>Nat Methods</source>
          ., v.
          <volume>15</volume>
          , p.
          <fpage>551</fpage>
          -
          <lpage>558</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <article-title>[5] ELIXIR consortium</article-title>
          .
          <source>Scientific programme 2014- 2018</source>
          . Executive summary.
          <year>2015</year>
          , https://www.elixireurope.org/system/files/ELIXIR-ExecutiveSummary-2015_Digital.pdf
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Lander</surname>
            <given-names>E.</given-names>
          </string-name>
          et al.
          <article-title>Initial sequencing and analysis of the human genome</article-title>
          .
          <source>Nature</source>
          , v.
          <volume>409</volume>
          , p.
          <fpage>860</fpage>
          -
          <lpage>921</lpage>
          ,
          <year>2001</year>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Masys</surname>
            <given-names>D.</given-names>
          </string-name>
          <article-title>New directions in bioinformatics</article-title>
          .
          <source>J. of Res. Nat. Inst. Stand. and Techn</source>
          . v.
          <volume>94</volume>
          , p.
          <fpage>59</fpage>
          -
          <lpage>63</lpage>
          ,
          <year>1989</year>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Petryszak</surname>
            <given-names>R</given-names>
          </string-name>
          , et al. Expression
          <string-name>
            <surname>Atlas</surname>
          </string-name>
          update-
          <article-title>-a database of gene and transcript expression from microarray- and sequencing-based functional genomics experiments</article-title>
          .
          <source>Nucleic Acids Res</source>
          ., v.
          <volume>42</volume>
          , p.
          <fpage>926</fpage>
          -
          <lpage>932</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Reisinger</surname>
            <given-names>F.</given-names>
          </string-name>
          et al.
          <article-title>Introducing the PRIDE Archive RESTful web services</article-title>
          .
          <source>Nucleic Acids Res</source>
          .,
          <source>pii: gkv382</source>
          ,
          <fpage>2015</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Thornton</surname>
            <given-names>J.</given-names>
          </string-name>
          <article-title>The future of bioinformatics</article-title>
          . Trends in Biotechn., v.
          <volume>17</volume>
          , p.
          <fpage>30</fpage>
          -
          <lpage>31</lpage>
          ,
          <year>1998</year>
          . Topol E.
          <article-title>Individualized medicine from prewomb to tomb</article-title>
          . Cell, v.
          <volume>157</volume>
          , p.
          <fpage>241</fpage>
          -
          <lpage>253</lpage>
          ,
          <year>2014</year>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>