<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Automatic Synthesis of Bioconductor Pipelines: A Domain Modeling Challenge</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Anna-Lena Lamprecht</string-name>
          <email>anna-lena.lamprecht@lero.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tiziana Margaria</string-name>
          <email>tiziana.margaria@lero.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Lero - The Irish Software Research Centre, University of Limerick</institution>
          ,
          <country country="IE">Ireland</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>GNU R is a widely used programming language and software environment for statistical data analysis and visualization. Bioconductor [1] is a collection of bioinformatics packages that extends R's standard range of functionality by comprehensive libraries of functions and meta-data predominantly for the analysis of data from high-throughput genomics and molecular biology experiments, and additionally provides several example data sets that are useful for testing, benchmarking and demonstration purposes. Reference manuals and additional manuscripts provided with the packages at the Bioconductor web site describe a variety of data analysis procedures based on the available functionality. The described bioinformatics procedures or work ows are typically referred to as data analysis pipelines, as they typically have a simple, in fact mostly linear, structure. (This is largely due to the fact that the often considerable complexity of the individual analysis steps is encapsulated by a services with a simple interface, and hence hidden from the user at the provided level of abstraction.) Thus, automatic work ow composition functionality like the method described in [6], which is based on a linear-time logic synthesis algorithm, can be easily applied here to generate the complete analysis pipelines automatically. Making use of the constraint-driven work ow composition functionality of the PROPHETS framework [5], which is based on such a linear-time logic synthesis algorithm, we created a prototype of a corresponding synthesis framework [3]. Focusing on DNA microarray analysis, this prototype comprises around 30 services and provides a framework for user-level construction of analysis pipelines based on appropriately wrapped and integrated Bioconductor functionality. It helps handling the variability that is inherent in microarray data analysis at an useraccessible level and emphasizes the agility of a model-driven and service-oriented approach to work ow design. PROPHETS is the current reference implementation of the loose programming paradigm [4], aiming to simplify work ow development in order to reach application experts without programming background. Working with PROPHETS consists of two major phases: 1. Domain Modeling, where PROPHETS is prepared for the application domain by describing services and their interfaces formally, de ning service and type taxonomies (simple ontologies with is-a) relations only) and specifying domain-speci c constraints, and</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>2. Work ow Design, where the actual work ow synthesis takes place in the
fashion of an iterative \playing" with the speci cation, synthesis, and
constraint and parameter re nement, until a satisfying solution is found.</p>
      <p>Hence, for making new application domains accessible to work ow designers,
adequate domain modeling is crucial. For the prototype application we set up
the domain model manually. On the component level (i.e. when dealing with
the individual services), we could largely reuse available information from the
packages' documentation les. Fortunately, consistent terminology is used there
for referring to data types and functions, although there is no formal ontology or
taxonomy model behind that would explicitly de ne a controlled vocabulary. On
the ontological level (i.e. when de ning the service and type taxonomies of the
domain model), however, the process was not so straightforward. The package
descriptions did not contain any explicit relation annotations, except from the
hierarchical information that can be derived from the package structure itself.
Hence, we had to de ne abstract semantic categories for the hierarchical grouping
of services and types ourselves.</p>
      <p>
        In contrast to other case studies described in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], we could furthermore not
directly apply terminology from the EMBRACE Data and Methods Ontology
(EDAM) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] here. Although it contains several terms for the eld of
microarray data processing, they were mostly too generic for the annotation and
typesafe composition of services that perform very speci c operations like in our
case study, so that they had to be re ned to an adequate level of detail that
matched the individual functionalities. We are now investigating if and how the
gap between EDAM and the Bioconductor package documentation regarding
domain modeling for synthesis applications can be bridged automatically based
on information that is already available. Solving this challenge would enable the
automatic synthesis of Bioconductor pipelines at a larger scale, with ultimately
thousands of services available for automatic composition into analysis pipelines.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Gentleman</surname>
            ,
            <given-names>R.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carey</surname>
            ,
            <given-names>V.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bates</surname>
            ,
            <given-names>D.M.</given-names>
          </string-name>
          , et al.:
          <article-title>Bioconductor: open software development for computational biology and bioinformatics</article-title>
          .
          <source>Genome Biology</source>
          <volume>5</volume>
          (
          <issue>10</issue>
          ),
          <source>R80</source>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Ison</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kalas</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jonassen</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          , et al.:
          <article-title>EDAM: an ontology of bioinformatics operations, types of data and identi ers, topics and formats</article-title>
          .
          <source>Bioinformatics</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Lamprecht</surname>
            ,
            <given-names>A.L.</given-names>
          </string-name>
          :
          <article-title>User-Level Work ow Design - A Bioinformatics Perspective</article-title>
          , Lecture Notes in Computer Science, vol.
          <volume>8311</volume>
          . Springer (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Lamprecht</surname>
            ,
            <given-names>A.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Naujokat</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Margaria</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ste</surname>
            <given-names>en</given-names>
          </string-name>
          , B.:
          <article-title>Synthesis-Based Loose Programming</article-title>
          .
          <source>In: 7th Int. Conference on the Quality of Information and Communications Technology (QUATIC</source>
          <year>2010</year>
          ). pp.
          <volume>262</volume>
          {
          <issue>267</issue>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Naujokat</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lamprecht</surname>
            ,
            <given-names>A.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ste</surname>
            <given-names>en</given-names>
          </string-name>
          , B.:
          <article-title>Loose Programming with PROPHETS</article-title>
          .
          <source>In: 15th Int. Conference on Fundamental Approaches to Software Engineering (FASE</source>
          <year>2012</year>
          ). LNCS, vol.
          <volume>7212</volume>
          , pp.
          <volume>94</volume>
          {
          <fpage>98</fpage>
          . Springer (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6. Ste en,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Margaria</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Freitag</surname>
          </string-name>
          ,
          <string-name>
            <surname>B.</surname>
          </string-name>
          :
          <article-title>Module Con guration by Minimal Model Construction</article-title>
          .
          <source>Tech. rep., Fakultat fu</source>
          r Mathematik und Informatik, U. Passau (
          <year>1993</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>