<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A conceptual modelling-based approach to generate data value through the end-user interactions: A case study in the genomics domain</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Carlos Iñiguez-Jarrín</string-name>
          <email>carlos.iniguez@epn.edu.ec</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Escuela Politécnica Nacional, Departamento de Informática y Ciencias de la Computación, Facultad de Ingeniería de Sistemas Ladrón de Guevara</institution>
          ,
          <addr-line>E11-253, Quito</addr-line>
          ,
          <country country="EC">Ecuador</country>
          <addr-line>P.O. Box 17-01-2759</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Universitat Politècnica de València</institution>
          ,
          <addr-line>Valencia</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In the current Big data ecosystem, identifying the data with the real value to an organization, or in other words the “data value", is a key issue for the decision making process. Understanding data implies a challenging cognitive process, which involves the know-how of domain experts. We propose an approach based on conceptual modelling to discover the data value through the interactions made by users when exploring the data. Our main ideas are: 1) To create a base of domain knowledge represented by interactions; 2) To formalize the interactions of the users with the data ecosystem. Our goal is to express high-level interactions between end-users and the data scenario, which represents the cognitive process followed to enact value from data. Such interactions together a subjacent conceptual model will be the mechanisms to recommend the next data exploring steps. In this paper we provide a solution design to generate value from the huge amount of data.</p>
      </abstract>
      <kwd-group>
        <kwd />
        <kwd>Data value</kwd>
        <kwd>interaction design</kwd>
        <kwd>visual data exploration</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        "Big data" is a umbrella term for expressing the huge ecosystem of structured and
unstructured data, which has become a issue of great interest for industry due to it
represents a potential resource to get an in-depth understanding of data. Because the
heterogeneous characteristic of data, "Big data" requires data analysis mechanisms
more powerful than the traditional ones. [1]. “Decision makers of all kinds, (…),
would like to base their decisions and actions on this data” [2]. Turning data into
meaningful information becomes a challenge [3][
        <xref ref-type="bibr" rid="ref9">4</xref>
        ]. Understanding the data is a
complex task that involves to locate relevant information from the large amount of
existent data and the cognitive process involved to obtain the “data value”.
      </p>
      <p>But, what is the meaning of "data value"? We define the "data value" like the
meaningful information obtained from a sense-making process that adds value to the
accomplishment of user goal. Data become information when they are ascribed value
[3]. In a sense-making process, domain experts create, modify, and evaluate schemas
of relations between items [5] [6] . Such schemas are representations of conceptual
models consisting of concepts and their relationships that allow people to abstract a
problem. In the Software Engineering and Data Base areas, conceptual models are
usually used to organize and shape the body of knowledge, as the skeleton fulfils its
function in the human body. However, understanding such conceptual models in order
to explore and generate value from data is not an easy task even by skilled technical
people. Therefore, the need for an intuitive mechanism for manipulating the
conceptual model as support for sense-making process by non-skilled users becomes a
challenge.</p>
      <p>This is where the prediction based on user-interactions and the interactive
visualization come into play. In the data sense-making process scenario, predictive
interactions could be the way to facilitate the conceptual model understanding. Capturing
and formalizing the resulting domain expert interactions become source of tracks
about their preferences, behaviour and decisions. Thus, inferring over such
interactions in order to predict the next steps to take by users together a suitable visual guide
can help them to decrease the complexity on understanding the conceptual model.</p>
      <p>The aim of this research is obtaining the data value through a sense-making
process to provide meaning from data. To achieve such a goal, in this paper we describe
the design of an approach based on underlying conceptual model and end-user
interactions in order to generate data value from a huge data set. The domain expert will
work on visualizing and exploring data and the set of registered interactions together
an underlying conceptual model will be the source of recommendation to generate
meaningful information.</p>
      <p>In order to apply our research, we selected the genetic diagnosis of diseases as a
case of study, where from a huge amount of genetic data, domain experts (clinicians,
physicians, scientists, etc.) carry out a sense-making process to detect whether a
person has an illness or estimate the risk of developing some disease. From a genetic
sample taken from one o many patients, the practitioners explore and analyse
manually the genetic data to find relationships between sample genetic mutations and
relevant information of genetic diseases, which is available on public genetic databases.
This is like finding a needle in a haystack. Finally, the relevant findings that
contribute to the diagnosis are consolidated y reported.</p>
      <p>Our proposal is related to the analysis phase, where the coalescence between
appropriate interactive mechanisms and the genome conceptual model [7] may be a
suitable mechanism to generate value of the large amount of genetic data.</p>
      <p>The paper is structured as follows: the next section summarizes the related work
then, in Section 3, we present the Design Science research methodology [8] applied to
this research. In Section 4, we describe the problem statement and identify the
research questions we plan to answer in the proposed work. Finally, section 5 presents
an overview of our solution design.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Works</title>
      <p>In this section, we briefly review relevant works from Human-computer Interaction,
Information Retrieval, and Data Visualization and Machine Learning that are related
to improving the way in which users get meaning from the data.</p>
      <p>Le et al. [9] discuss the use of analytic trails technology as part of Smarter Decisions
to support the users when conducting visual data analysis. Although the use of the
analytic trails is focused on the aspects of analytic provenance, asynchronous
collaboration, and reuse of analyses, it is important to mention that the captured interactions
become the source of “trails” of analysis tasks representing the analytic steps taken by
the user during visual data exploration.</p>
      <p>
        Athukorala [
        <xref ref-type="bibr" rid="ref9">4</xref>
        ] analyses how user interaction modelling can be applied to provide
better support in exploratory information–seeking. He proposes to model the user
behavior to allow information retrieval systems to infer the state of exploration from
observable aspects of user interactions.
      </p>
      <p>JIT interactive analytics [10] is a proposal to join computational and visual
techniques. JIT analytics is performed in real-time on data that users are interacting with
to guide visual-analytic exploration where enriching visualizations with annotations
suggests to users possible insights to examine further.</p>
      <p>In [11], RockQuery is proposed as a Visual Query System (VQS)/Data Query Tool
that combines ontology views with interaction community techniques to reduce the
overload of information by presenting visually to the user only the information that is
required by the data exploration task at hand.</p>
      <p>In [12], a knowledge navigation infrastructure is presented to explore literature related
to dengue disease. A content acquisition engine drives the delivery of dengue-specific
literature from public repositories by means of previously identified keywords.
In [13], Optique is presented as a tool to enable the access to Big data from a ontology
based approach to formulate visual queries.</p>
      <p>In [14], focused on exploratory search, is presented an approach where users can
directly manipulate document features (keywords) through a graphic user interface to
indicate their interests and reinforcement learning is used to model the user by
allowing the system to trade off between exploration and exploitation.</p>
      <p>All the works cited describe mechanisms aimed to allow users to get insight of large
dataset. Although they use annotations and keywords as source of automatic learning
approaches, they do not focus on the interactions set as a source of valuable
information. We propose an approach where the process to get value from data is
supported by both the inference over the user-interactions as guide to discover related data
and the conceptual model as guide to allow users to visually navigate on obtaining
value from data.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Research Methodology</title>
      <p>This research will follow the guidelines of the Design Science Methodology [8].
Design Science is oriented to information systems and software engineering research,
an appropriate approach to the nature of our research. The main object of study of this
methodology is an artifact in a problem context. In our case, the artifact is:
A conceptual modelling approach to generate data value</p>
      <p>based on user interactions
And the problem context consist of:</p>
      <p>Data analysis.</p>
      <p>Our research is considered as a utility-driven exploratory research project.
Exploratory research since our stakeholders (SENESCYT and Escuela Politécnica Nacional)
are willing to sponsor an exploratory research and utility-driven since the research
results must accomplish with its budgets and goals. Such goals can be expressed in a
hierarchical structure where the achievement of high-level goals is the result of the
achievement of every low-level goal. Our goals are defined and ordered from highest
(G1) to lowest (G4) goals level, as follow:
•
•
•
•</p>
      <p>G1: Develop a conceptual modelling-based approach to generate data value
through the end-user interactions.</p>
      <p>G2: Determine the conceptual modelling-based approach to generate data
value through the interactions.</p>
      <p>G3: Predict the effects caused by the approach in the context of use.</p>
      <p>G4: Make a literature review to determine the existent approaches to address
the generation of data value from both the conceptual model and interaction
perspectives.</p>
      <p>The defined goals derive problems classified on design problems and knowledge
questions. Solving design problems imply changing the real world to suit human
purposes, in contrast, solving knowledge questions imply acquiring knowledge about the
world without necessarily changing it [15]. In order to deal with the mentioned
problems, the methodology provides two rational nested problem-solving cycles, which
consist of several tasks. Both cycles work in a nested way ensuring the design and
evaluation of artifacts. The Fig. 1 shows the two nested cycles applied to our research
project. The Design Cycle contains three tasks (from T1 to T3) to address design
problems, whereas the Empirical Cycle contains five tasks (from T4 to T8) to address
the knowledge questions. It is important to mention that, the design cycle is a
subcycle from Engineering Cycle, for this reason not all tasks from engineering cycle are
regarded into the design cycle. Design science projects are always restricted to the
first three tasks of the engineering cycle [8]. The treatment implementation (T9) is out
of scope of this research project.</p>
      <p>T 1
T 3</p>
      <p>T 1.- PROBLEM INVESTIGATION
T 1.1 Investigate the need
for generating data value from a huge
amount of data
T 1.2 Investigate the conceptual
framework attached to the problem.</p>
      <p>T 1.3 Investigate existent approaches
related to the problem
T 1.4 Define criteria to evaluate the
approach</p>
      <p>T 7</p>
      <p>T 6</p>
      <p>T 8
EMPIRIC AL</p>
      <p>CYCLE
Validate the
proposed
approach</p>
      <p>T 5
value throug user-interactions. To achieve this goal, We have defined the research
questions showed in the Table 1.</p>
      <sec id="sec-3-1">
        <title>Code</title>
        <p>RQ 1
RQ 2
RQ 2.1
RQ 2.2
RQ3</p>
      </sec>
      <sec id="sec-3-2">
        <title>Research Question</title>
        <p>What approaches exist for data value generation
supported by conceptual models and end-user
interactions?
interactions?</p>
        <p>How to develop a conceptual modelling
apG1
proach to generate data value through the end-user</p>
        <sec id="sec-3-2-1">
          <title>Identify the user-interactions that can be suitable to support generating data value low user to generate value from data. Define the user interface characteristics that al</title>
        </sec>
        <sec id="sec-3-2-2">
          <title>What effects does the approach applied to the context of use cause?</title>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>Goal</title>
        <p>G2, G4
G1, G4
G1, G4
G3</p>
      </sec>
      <sec id="sec-3-4">
        <title>Problem Type</title>
        <p>KQ
DP
DP
DP
KQ</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Solution Design Proposal</title>
      <p>Our proposed approach is oriented to provide an environment where a domain user
can be guided by predictive interactions when exploring a large dataset. The process
is depicted in the Figure 2. In the first step, a graphical user interface allows user
interact with conceptual model elements expressed in meta-information in order to
order to explore and make sense of data showed. In the second step, the interactions
performed are stored in a knowledge database. Every action performed when the user
explores the data, represents an interaction consisting of a) the data attributes to
specialize the exploration and b) the track of conceptual model nodes visited until while.
In the step 3, an inference engine consisting of predictive algorithms takes the stored
interactions in order to suggest possible steps in the search process. The outcomes are
transformed to a suitable input for a DSL depicted in the step 4. The DSL use the set
of data attributes to request from database, the related data that accomplish the search
constraints, whereas the track of conceptual model selected is used to reasoning over
the conceptual model in order to find out the connected concepts related to the last
visited element. Those two sorts of information: related data and connected concepts
are communicated to the graphical user interface as depicted in the step 5. The
showed data represent the actual state of search whereas the connected concepts
become on the possible next steps that systems suggest to the user in order to guide the
process of get meaning of data.</p>
      <p>Unlike data-driven traditional approaches, where the users requires knowing the
conceptual model in order to create tailored queries to their needs, we aim to provide
an interactive and predictive mechanism that combines both the data-driven and
model-driven perspectives, that allows the user to obtain answers to their questions
while discovering the knowledge guided by the conceptual model of the problem.
Hence, end-users do not need to know the whole conceptual model which represents a
complex task.</p>
      <p>The main contributions to be expected from this work are:
1. Exploratory study of interactions of genetists when diagnosing genetic diseases.
2. Design a prototype to generate data value through user interactions, applied to the
genetic diseases diagnosis
3. Formalize the user-interactions involved on diagnosing genetic diseases.
4. Evaluate the user-interactions based approach on a real scenario.
6</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgements</title>
      <p>The author gratefully acknowledge the financial support provided by the Escuela
Politécnica Nacional, Secretaría Nacional de Educación, Ciencia y Tecnología
(SENESCYT) and IDEO project for the development of this research project. Thanks
to supervisor Óscar Pastor for his invaluable support and advice.
7
[1]</p>
      <p>M. Chen, S. Mao, and Y. Liu, “Big data: A survey,” in Mobile Networks and</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Applications</surname>
          </string-name>
          ,
          <year>2014</year>
          , vol.
          <volume>19</volume>
          , no.
          <issue>2</issue>
          , pp.
          <fpage>171</fpage>
          -
          <lpage>209</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>M. R. Lissack</surname>
          </string-name>
          , “
          <article-title>Of chaos and complexity: managerial insights from a new science,” Manag</article-title>
          . Decis., vol.
          <volume>35</volume>
          , no.
          <issue>3</issue>
          , pp.
          <fpage>205</fpage>
          -
          <lpage>218</lpage>
          , Apr.
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>K. M. Athukorala</surname>
          </string-name>
          , “
          <article-title>Enhancing Exploratory Information-Seeking through Interaction Modeling,” in USER MODELING, ADAPTATION, AND PERSONALIZATION</article-title>
          , UMAP
          <year>2014</year>
          ,
          <year>2014</year>
          , vol.
          <volume>8538</volume>
          , pp.
          <fpage>478</fpage>
          -
          <lpage>483</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>D. H. Chau</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Kittur</surname>
            ,
            <given-names>J. I.</given-names>
          </string-name>
          <string-name>
            <surname>Hong</surname>
            , and
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Faloutsos</surname>
          </string-name>
          , “Apolo,”
          <source>in Proceedings of the 2011 annual conference on Human factors in computing systems - CHI '11</source>
          ,
          <year>2011</year>
          , vol.
          <volume>46</volume>
          , p.
          <fpage>167</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Levin</surname>
          </string-name>
          , “
          <article-title>Conceptual modeling of human genome: Integration challenges</article-title>
          ,
          <source>” Lecture Notes in Computer Science including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics</source>
          , vol.
          <volume>7260</volume>
          LNCS.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          pp.
          <fpage>231</fpage>
          -
          <lpage>250</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>R.</given-names>
            <surname>Wieringa</surname>
          </string-name>
          , Design science methodology.
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>J.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pan</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Lai</surname>
          </string-name>
          , “
          <article-title>Analytic trails: Supporting provenance, collaboration, and reuse for visual data analysis by business users</article-title>
          ,
          <source>” in Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)</source>
          ,
          <year>2011</year>
          , vol.
          <volume>6949</volume>
          LNCS, no.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <source>PART 4</source>
          , pp.
          <fpage>256</fpage>
          -
          <lpage>273</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>E.</given-names>
            <surname>Kandogan</surname>
          </string-name>
          , “
          <article-title>Just-in-time interactive analytics: Guiding visual exploration of data,”</article-title>
          <source>IBM J. Res. Dev.</source>
          , vol.
          <volume>59</volume>
          , no.
          <issue>2</issue>
          /3, p.
          <volume>12</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>12</lpage>
          :
          <fpage>10</fpage>
          ,
          <string-name>
            <surname>Mar</surname>
          </string-name>
          .
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>J.</given-names>
            <surname>Lozano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Carbonera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pimenta</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Abel</surname>
          </string-name>
          , “
          <article-title>RockQuery An Ontology-based Data Querying Tool</article-title>
          ,” vol.
          <volume>3</volume>
          , pp.
          <fpage>25</fpage>
          -
          <lpage>33</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Rajapakse</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kanagasabai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. T.</given-names>
            <surname>Ang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Veeramani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Schreiber</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C. J. O.</given-names>
            <surname>Baker</surname>
          </string-name>
          , “
          <article-title>Ontology-centric integration and navigation of the dengue literature,”</article-title>
          <string-name>
            <given-names>J.</given-names>
            <surname>Biomed</surname>
          </string-name>
          . Inform., vol.
          <volume>41</volume>
          , no.
          <issue>5</issue>
          , pp.
          <fpage>806</fpage>
          -
          <lpage>815</lpage>
          , Oct.
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Horrocks</surname>
          </string-name>
          , “
          <article-title>OptiqueVQS: towards an ontology-based visual query system for big data</article-title>
          ,
          <source>” Proc. Fifth Int. Conf. Manag. Emergent Digit. Ecosyst. - MEDES '13</source>
          , pp.
          <fpage>119</fpage>
          -
          <lpage>126</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Jacucci</surname>
          </string-name>
          , “
          <article-title>Directing exploratory search: Reinforcement Learning from User Interactions with Keywords,”</article-title>
          <source>Proc. 2013 Int. Conf. Intell. user interfaces - IUI '13</source>
          , pp.
          <fpage>117</fpage>
          -
          <lpage>128</lpage>
          , Mar.
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>R.</given-names>
            <surname>Wieringa</surname>
          </string-name>
          , “
          <article-title>Design Science as nested problem solving</article-title>
          ,
          <source>” 4th Int. Conf. Des.</source>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>Sci. Res. Inf. Syst. Technol.</surname>
          </string-name>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>12</lpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>