<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Using semantic web technology to accelerate plant breeding.</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Pierre-Yves Chibon</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Benoˆıt Carr`eres</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Heleena de Weerd</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Richard G. F. Visser</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Richard Finkers</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Centre for BioSystems Genomics</institution>
          ,
          <addr-line>Wageningen, 6708 PB, Wageningen</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Experimental Plant Sciences, Wageningen University and Research Centre</institution>
          ,
          <addr-line>6708 PB Wageningen</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Wageningen UR Plant Breeding, Wageningen University and Research Centre</institution>
          ,
          <addr-line>6708 PB Wageningen</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>One goal within plant breeding is to find the causal gene(s) explaining a given phenotype. Semantic web technology brings opportunities to integration data and information accross spread data sources. Chebi2gene and Marker2sequence are two applications relying on this semantic web technology to integration genes, proteins, metabolites, pathways, literature. Their web-based interface allows biologists to use and explore this network of information.</p>
      </abstract>
      <kwd-group>
        <kwd>Semantic web applications</kwd>
        <kwd>Data integration</kwd>
        <kwd>Plant Breeding</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>”Producing 70 percent more food for an additional 2.3 billion people by 2050
while at the same time combating poverty and hunger, using scarce natural
resources more efficiently and adapting to climate change are the main challenges
world agriculture will face in the coming decades” (http://www.fao.org/news/story/en/item/35571/).
Plant breeding is part of the answer to this challenge. The FAO itself recognize
it: ”Plant breeding techniques can lead to improved crop varieties that increase
yields, decrease losses” (http://www.fao.org/news/story/en/item/35686/). In
order to improve crop varieties, breeders introgress genes of interest from one
accession to another. The challenge is to pinpoint the gene(s) responsible for the
improved traits. To find these genes, plant breeders use all the new types of
information, which become available using high-throughput technologies, such as
nextgeneration sequencing technology, RNASeq, proteomics and/or metabolomics.</p>
      <p>
        As a daily practice, plant breeders associate these large datasets to one or
several regions of the genome using advanced statistical methodology. These
regions are called Quantitative Trait Loci (QTLs) and can be introgressed from
one variety to another with the goal to developed an improved variety. However,
a typical QTL region may contain over hundreds of genes, including genes
negatively influencing the breeding goals. Complete genome sequences of many crop
plants are becoming available, including the genome of important food crops
such as tomato [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and potato [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The availability of structural and functional
genome annotations makes it possible to investigate the QTL region for genes
positively or negatively influencing the trait of interest.
      </p>
      <p>Plant breeding, as most research areas nowadays, faces the problem of
spreading data resources. Most plant species have their own website, ideally
crossreferenced to major cross-species database such as UniProt or GO, but the
number of resources available keeps increasing every year. When a researcher
starts to investigate the genes in a specific QTL region on a genome, he will
have to browse through an increasing number of websites and databases to
collect and integrate information about each of these genes. One solution to this
problem is to use semantic web technologies to aggregate and integrate the data
from different resources in a way that would be and automated and expandable
to new resources as they become available.</p>
      <p>Within Wageningen UR Plant Breeding, we have developed two new tools
relying on semantic web technologies to help breeders face this challenge, namely:
Chebi2gene and Marker2sequence.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Materials and Methods</title>
      <p>Our tools have been primary developed for usage with the Tomato genome
sequence, however, they should are implemented in a universal manner. The
limiting factor is that the annotation of a genome sequence should be available
in RDF. To convert the Tomato genome annotation into RDF we developed
a simple tool: gff2RDF, which retrieves and parses the gff file and outputs a
RDF document of the annotation. This tool is used for Tomato and is being
extended to work with Potato and Arabidopsis thaliana. In the conversion the
gene annotation linked to external database have been converted to use the URI
of these database, i.e.: the GO terms associated with the genes are identified
using the URI provided by the OWL file of the Gene Ontology consortium, and
similarly for the protein identifier against UniProt. The resulting RDF file has
been uploaded into a Virtuoso OSE server (v 6.1.3) together with the RDF files
provided by EBI for UniProt-core, UniProt-pathways, UniProt-go and
UniProtcitations (all version 2011 10), CHEBI (version 2011 09), Rhea (release 33), the
Gene Ontology OWL file as provided by the Gene Ontology Consortium (version
2011 11 03). Each resource is stored in its own graphs (Fig 1), allowing easier
upgrades, and SPARQL is used to do the integration.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Chebi2gene, linking metabolites to genes</title>
      <p>Metabolites are chemical compounds produced by an organism, in plants their
actions can be related to traits of major interest such as flavor or disease
resistance. The link between a metabolite and the genes involved in its expression is
not always straightforward. Chebi2gene is a proof of concept allowing breeders
to go from one metabolite to the associated gene(s). This allows biologists to
Fig. 1. This figure represents the integration of the different resources in our triple
store. The green box defines the different RDF graphs and the red ellipse the type of
information we extract from these graphs. The dashed gray line between around Rhea
is for the fact that Rhea do not reuse the URI of Chebi or UniProt when refering to
Chebi compound or a UniProt protein. The mapping is then indirect, as opposed to
the plain gray line where the URI are consistent and shared.
find all the genes in the genome related to a metabolite. To find these
associations, it uses CHEBI, Rhea, and UniProt databases and our RDF version of the
tomato genome annotation. The input is either a CHEBI identifier or the name
of the metabolite. If the input is a name, Chebi2gene will search the CHEBI
database for all compounds having this name in their name and optionally in
their synonyms. Once the metabolite has been uniquely identified with a CHEBI
identifier, Chebi2gene search in Rhea all the chemical reactions in which it is
involved, then all the proteins which are involved in these reactions and finally all
the genes from the genome annotation which are related to these proteins. These
searches are performed using SPARQL queries on our Virtuoso server across the
different graphs of the different resources.</p>
      <p>Chebi2gene is available at: http://www.plantbreeding.wur.nl/chebi2gene
For example, when searching for ”beta-carotene” in chebi2gene, three molecules
containing ”beta-carotene” are returned: ”beta-carotene”, ”beta-carotene
5,6epoxide”, and ”(5S,6R)-beta-carotene 5,6-epoxide”. From these three molecules,
the first one is the molecule of interest. It has the CHEBI identifier 17579.
Searching with this identifier in chebi2gene, we can find that this compound is
involved in 4 reactions which are associated with 10 proteins. Amongs these
proteins is ”Lycopene beta cyclase” (UniProt ID: Q38933) which is
associated with two pathways: ”Carotenoid biosynthesis; beta-carotene biosynthesis”
and ”Carotenoid biosynthesis; beta-zeacarotene biosynthesis” and four genes:
Solyc04g040190.1.1 (chromosome 4) and Solyc10g079480.1.1 (chromosome 10)
which are both ”Beta-lycopene cyclase” and Solyc06g074240.1.1 (chromosome
6) and Solyc12g008980.1.1 (chromosome 12) which are both ”Lycopene beta
cyclase”.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Marker2sequence explore a genome region for a candidate gene</title>
      <p>
        Marker2sequence [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] aims at mining quantitative trait loci (QTLs) for candidate
genes. For each gene, within the QTL region, marker2sequence uses semantic
web integration technology to integrate putative gene function with associated
Gene Ontology terms, proteins, pathways and literature. This integration is
performed using SPARQL queries against our triple-store. As mention earlier, a
typical QTL region easily contains several hundreds of genes, this gene list can
then be further filtered using a keyword based query on the aggregated
annotations. This single query search for the given keyword in the gene annotation
and GO terms, proteins and literature associated with this gene. More precisely,
it searches the keyword in he name and definition and synonym of the Gene
Ontology term associated with the gene. It searches the keyword in the protein
name and description of each protein associated with the gene. It searches the
keyword in the pathway name of each pathways associated to these proteins and
finally it searches in all title and abstract of the literature associated with these
proteins. If any of these elements contains the searched keyword the gene is
selected as a potentially interesting gene and returned to the user Marker2sequence
will help breeders to identify potential candidate genes for their traits of interest.
      </p>
      <p>
        Marker2sequence is available at: http://www.plantbreeding.wur.nl/BreeDB/marker2seq
For example, β-carotene content is a trait influencing the color of tomatoes
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Based on our QTL analysis, using data from the Solanum lycopersicum x
Solanum galapagense LA0483 RIL population [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], this compound has one QTL
on chromosome 6 (between TG253 and TG314). Marker2sequence identifies 988
genes in this region of the chromosome 6. A query with the keyword:
betacarotene, returns the gene Solyc06g074240.1.1. This gene, Solyc06g074240.1.1, is
associated with the GO term for carotenoid biosynthetic process, the pathway for
Carotenoid biosynthesis and more specifically the part of beta-carotene
biosynthesis. Information for each gene can be quickly mined using Marker2sequence
and this gene is the candidate for our trait of interest.
5
      </p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>Both Chebi2gene and Marker2sequences are tools presenting the potential of
semantic web integration for the plant breeding domain.</p>
      <p>All the tools presented in this abstract have been licensed under free license
and are available at: http://github.com/PBR/</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>The</given-names>
            <surname>Tomato Genome Consortium</surname>
          </string-name>
          .
          <article-title>The tomato genome sequence provides insights into fleshy fruit evolution</article-title>
          .
          <source>Nature</source>
          <volume>485</volume>
          (
          <year>2012</year>
          )
          <fpage>635641</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>The</given-names>
            <surname>Potato Genome Sequencing Consortium</surname>
          </string-name>
          .
          <article-title>Genome sequence and analysis of the tuber crop potato</article-title>
          .
          <source>Nature</source>
          <volume>475</volume>
          (
          <year>2011</year>
          )
          <fpage>189195</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Chibon</surname>
            ,
            <given-names>P-Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schoof</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Visser</surname>
            ,
            <given-names>R.G.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Finkers</surname>
          </string-name>
          , R.:
          <article-title>Marker2sequence, mine your QTL regions for candidate genes</article-title>
          .
          <source>Bioinformatics</source>
          <volume>28</volume>
          (
          <year>2012</year>
          )
          <fpage>1921</fpage>
          -
          <lpage>1922</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Lincoln</surname>
            ,
            <given-names>R.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Porter</surname>
            ,
            <given-names>J.W.</given-names>
          </string-name>
          :
          <article-title>Inheritance of beta-carotene in tomatoes</article-title>
          .
          <source>Genetics</source>
          <volume>35</volume>
          (
          <year>1949</year>
          )
          <fpage>206</fpage>
          -
          <lpage>211</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Paran</surname>
            ,
            <given-names>I. Goldman</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            ,
            <surname>Tanksley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.D.</given-names>
            ,
            <surname>Zamir</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          :
          <article-title>Recombinant inbred lines for genetic mapping in tomato</article-title>
          .
          <source>Theor. Appl. Genet</source>
          .
          <volume>90</volume>
          (
          <year>1995</year>
          )
          <fpage>542</fpage>
          -
          <lpage>548</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>