<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Ontology-based Design of Experiments on Big Data Solutions?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Maximilian Zocholl</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Elena Camossi</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anne-Laure Jousselme</string-name>
          <email>anne-laure.jousselmeg@cmre.nato.int</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Cyril Ray</string-name>
          <email>fcyril.ray@ecole-navale.frg</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institut de Recherche de l'Ecole Navale</institution>
          ,
          <addr-line>Brest</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>NATO STO Centre for Maritime Research and Experimentation</institution>
          ,
          <addr-line>La Spezia</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper the ontology-based approach is proposed to support the evaluation of big data systems. Firstly, the approach formalises a decomposition and recombination of the big data solution, allowing for the aggregation of component evaluation results at intercomponent level. Secondly, existing work on Design of Experiments (DoE) is translated into an ontology for supporting the selection of experiments. It exploits domain and inter-domain speci c restrictions on the factor combinations in order to select from the very large number of possible experiments a representative subset. Contrary to existing approaches, the proposed use of ontologies is not limited to the assertional description and exploitation of past experiments but o ers richer terminological descriptions for the development of a DoE from scratch. As an application example, a DoE is developed for a maritime big data solution.</p>
      </abstract>
      <kwd-group>
        <kwd>Ontology</kwd>
        <kwd>Big Data Solutions</kwd>
        <kwd>Big Data Variations</kwd>
        <kwd>Eval- uation</kwd>
        <kwd>Design of Experiments (DoE)</kwd>
        <kwd>Situational Awareness</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The assessment of a big data system poses signi cant challenges in terms of the
large number of data variations that must be considered by the experiments. To
improve the evaluation e ciency, experiments must focus on a representative
subset demonstrating the system ability to scale along the considered big data
dimensions.</p>
      <p>
        In many disciplines, Design of Experiments (DoE) is used to organise the
evaluation of processes, systems, and products. DoE is a collective of principles,
statistical approaches and models for planning and performing experiments as
well as analysing their results [
        <xref ref-type="bibr" rid="ref2 ref6">2, 6</xref>
        ]. Typically, the experimental unit is modelled
as a system with input and output variables. Some or all controllable input
variables, the so called factors, are varied according to an experimental plan
that speci es the values of the variables, or factor levels, of each experiment.
After feeding the experimental unit with a set of input variables, the output is
observed. As the output depends on the system behavior and both controllable
and possibly uncontrollable input variables, the goal of DoE is the quanti cation
of the functional relation between the input and output of the system, e.g. by
the analysis of variance or covariance. A large body of methods and knowledge
exists, but barely formalized in a machine interpretable way.
      </p>
      <p>Two main challenges arise during the assessment of a big data system with
a DoE. Firstly, big data variations translate seamlessly into a large number of
factors with a multiplicity of possible factor levels. Secondly, the di erent
components of the big data system implement either deterministic or non-deterministic
processes and yield di erent output types, e.g. continuous or multinomial. Thus,
for a thorough assessment the system needs to be unfolded into its components.
Again, this increases the number of necessary experiments but additionally and
more importantly introduces the necessity of using completely di erent types of
DoE. Choosing the wrong DoE results in a reduction of statistical e ciency or
the lack of consistency of the results.</p>
      <p>
        In this work, we apply ontology based DoE to the experimental evaluation
of big data systems. The proposed formalisation encompasses the
decomposition of the big data system, supporting the roll-up of the experiment results
at component level to obtain the inter-component level evaluation. In addition,
complementarily to the related work discussed in Section 2, in this paper we
propose to expand the existing formalisations on DoE and leverage the domain
knowledge to drive the selection of experiments. Similarly to [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], the knowledge
is supposed to be captured in the T-Box of the Ontology, thus supporting DoE
which starts without prior knowledge of the domain or instances in the A-Box.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        Do et al. provide an overview of empirical techniques for software testing
identifying two approaches [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]: rstly, controlled experiments rely on the precise
variation of given variables; complementarily, case studies follow possible scenarios
of usage. Both approaches aim for the replicability of the performed
experiments, the possibility to aggregate their results beyond the anticipated scope
and by these means to validate the signi cance of their results in form of
models. Precursors for enabling these bene ts can be seen in the interpretability of
the experimental results, e.g. by the documentation or standardisation, and in
infrastructures allowing to share and connect artifacts [
        <xref ref-type="bibr" rid="ref1 ref5">1, 5</xref>
        ]. In the data
mining and machine learning domain recent e orts lead to the W3C ML Schema
Community Group with the goal to support the development of a data exchange
standard for experimental data by unifying existing, more speci c schemata [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
More speci cally, the group pools former e orts on Data Mining OPtimization
(DMOP) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], Expose, the OpenML related Ontology of Vanschoren [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], as well
as the contributions of Soldatova et al. in the form of EXPO and OntoDM [
        <xref ref-type="bibr" rid="ref7 ref9">9, 7</xref>
        ],
focussing on the process of scienti c experiments. A very mature contribution is
Ontology-based DoE on Big Data Solutions
WINGS, a semantic work ow system that eases the development of data mining
work ows by its user friendly interface [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Ontology-based DoE and Evaluation</title>
      <p>The goal of DoE is to choose an experimental plan that is statistically e cient
whilst allowing for an aggregation of the experimental results that are consistent
with the context of the experiments, as well as to accept or reject the research
hypothesis. The choice of the DoE depends on multiple criteria, including the
features of the experimental unit, the mathematical function to model the
behavior of the experimental unit, the restrictions on the experimental space. The
relation between these criteria and di erent DoE are modelled as ontology
axioms, and the resulting ontology enables an ontology-based DoE. The following
are a subset of the axioms modelled in OWL2:</p>
      <sec id="sec-3-1">
        <title>DoEW ithoutReplication v DoE u 9hasExpU nit:Deterministic:</title>
        <p>DoEW ithReplication
N onDeterministic</p>
      </sec>
      <sec id="sec-3-2">
        <title>DoE u 9hasExpU nit:Deterministic:</title>
      </sec>
      <sec id="sec-3-3">
        <title>9hasComponent:N onDeterministic:</title>
        <p>DoEW ithBlocking</p>
      </sec>
      <sec id="sec-3-4">
        <title>DoE u 9hasN uisanceF actor:Controllable:</title>
        <p>DoEW ithRandomization</p>
      </sec>
      <sec id="sec-3-5">
        <title>DoE u 9hasN uisanceF actor:U ncontrollable:</title>
        <p>(1)
(2)
(3)
(4)
(5)</p>
        <p>
          For the evaluation of the ontology the maritime rule sets from [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] are used.
The rule set is part of a big data solution currently developed in the execution of
the datAcron project3. Spatial and critical rule sets are modelled as components
of the composite rule sets.
        </p>
        <p>
          Axiom (1) ensures only experimental units with deterministic behavior to
be assigned to DoEs without replications. If an experimental unit with
nondeterministic behavior is assigned to a DoE without replication, the ontology
becomes inconsistent. All ve rule sets in [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] are modeled as experimental units,
namely \Vessel within Area", \Vessel under Way", \Aground", \Trawling" and
\Rendez-vous". As all rule sets are deterministic, the DoE instances related
to these di erent experimental units can be assigned correctly to the class
DoEWithoutReplication. Assuming a non-deterministic behavior of an arbitrary
component of a rule set, such as the user-driven detection of the same events in
[
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], axiom (2) and (3) induces an automatic classi cation of the respective DoE
as instance of the class DoEWithReplication. In case of a rule set is assigned
to the class Deterministic and a component of this rule is asserted as instance
of the class Non-Deterministic, an inconsistency is created. As the Axioms (4)
and (5) follow the same design pattern as (2), the available reasoning techniques
presented in Table 1 allow for a similar support during the process of selecting
a suitable DoE. All rule sets are deterministic and have no nuisance factors,
neither controllable nor uncontrollable.
3 www.datacron-project.eu
DoEWithReplication
DoEWithoutReplication
DoEWithBlocking
DoEWithoutBlocking
DoEWithRandomization
        </p>
        <p>DoEWithoutRandomization
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusion</title>
      <p>As big data systems ingest a large number of variables with a large range of
possible values, their evaluation requires a methodological choice of experiments.
The presented approach uses the descriptions of components of an existing big
data solution in order to reduce the design space of possible experiments
according to well understood concepts of DoE. By excluding infeasible or unnecessary
experiments, the number of experiments is reduced.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Do</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elbaum</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Rothermel</surname>
          </string-name>
          , G.:
          <article-title>Supporting Controlled Experimentation with Testing Techniques: An Infrastructure and its Potential Impact</article-title>
          .
          <source>Empirical Software Engineering</source>
          <volume>10</volume>
          (
          <issue>4</issue>
          ),
          <volume>405</volume>
          {
          <fpage>435</fpage>
          (
          <year>2005</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2. Fisher, R. A.:
          <article-title>The design of experiments</article-title>
          . Oliver And Boyd, Edinburgh, London, (
          <year>1937</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Gil</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ratnakar</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gonzlez-Calero</surname>
            ,
            <given-names>P. A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Groth</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , Moody, J.,
          <string-name>
            <surname>Deelman</surname>
          </string-name>
          , E.:
          <article-title>WINGS: Intelligent Work ow-Based Design of Computational Experiments</article-title>
          .
          <source>Intelligent Systems</source>
          , IEEE, (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Keet</surname>
            ,
            <given-names>C.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lawrynowicz</surname>
          </string-name>
          , A.,
          <string-name>
            <surname>d'Amato</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kalousis</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nguyen</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Palma</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stevens</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Hilario</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>The Data Mining OPtimization Ontology</article-title>
          .
          <source>Web Semantics: Science, Services and Agents on the World Wide Web</source>
          <volume>32</volume>
          ,
          <issue>43</issue>
          {
          <fpage>53</fpage>
          , (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>ML</given-names>
            <surname>Schema</surname>
          </string-name>
          <article-title>Core Speci cation</article-title>
          , http://www.w3.org/
          <year>2016</year>
          /10/mls/.
          <source>Last accessed: 8 Feb</source>
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Montgomery</surname>
            ,
            <given-names>D.C.</given-names>
          </string-name>
          :
          <article-title>Design and Analysis of Experiments</article-title>
          . John Wiley &amp; Sons, (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Panov</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soldatova</surname>
          </string-name>
          , L., Dzeroski, S.:
          <article-title>Ontology of core data mining entities</article-title>
          .
          <source>Data Mining and Knowledge Discovery</source>
          <volume>28</volume>
          (
          <issue>5</issue>
          ),
          <volume>1222</volume>
          {
          <fpage>1265</fpage>
          , (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Pitsikalis</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kontopoulos</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Artikis</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alevizos</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Delaunay</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pouessel</surname>
            ,
            <given-names>J.- E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dreo</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ray</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Camossi</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jousselme</surname>
            ,
            <given-names>A.-L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hadzagic</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Composite Event Patterns for Maritime Monitoring</article-title>
          .
          <source>In Proceedings of 10th Hellenic Conference on Arti cial Intelligence SETN</source>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Soldatova</surname>
            ,
            <given-names>L. N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>King</surname>
            ,
            <given-names>R. D.:</given-names>
          </string-name>
          <article-title>An ontology of scienti c experiments</article-title>
          .
          <source>Journal of the Royal Society Interface</source>
          <volume>3</volume>
          (
          <issue>11</issue>
          ),
          <volume>795</volume>
          {
          <fpage>803</fpage>
          , (
          <year>2006</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Vanschoren</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blockeel</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pfahringer</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Holmes</surname>
          </string-name>
          , G.:
          <article-title>Experiment databases</article-title>
          .
          <source>Machine Learning</source>
          <volume>87</volume>
          (
          <issue>2</issue>
          ),
          <volume>127</volume>
          {
          <fpage>158</fpage>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>