<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>D. Martínez Minguet);</journal-title>
      </journal-title-group>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>CoMoDID: Combining explainable artificial intelligence and conceptual modeling for data intensive-domains management</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Oscar Pastor</string-name>
          <email>opastor@dsic.upv.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Diana Martínez Minguet</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jose Fabián Reyes Román</string-name>
          <email>jreyes@pros.upv.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alberto García S.</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ana Leon</string-name>
          <email>aleon@vrain.upv.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mireia Costa</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ferran Pla</string-name>
          <email>fpla@dsic.upv.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Data-Intensive Domains, Conceptual Modeling, Explainable Artificial Intelligence, Precision Medicine</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>ER2023: Companion Proceedings of the 42nd International Conference on Conceptual Modeling: ER Forum</institution>
          ,
          <addr-line>7th SCME</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Project Exhibitions</institution>
          ,
          <addr-line>Posters and Demos, and Doctoral Consortium</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Valencia</institution>
          ,
          <addr-line>46022</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Valencian Research Institute for Artificial Intelligence (VRAIN). Universitat Politècnica de València</institution>
          ,
          <addr-line>Camí de Vera S/N</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>1969</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>The large and heterogeneous data sets that characterize Data-Intensive Domains (DID) pose a challenge to developing data analysis and management approaches. A successful and eficient data knowledge extraction from DID-based systems is determined by assembling and analyzing such data sets, but integrating their diferent sources is arduous work. Finding sound solutions for this problem has become a relevant research goal, that existing DID-based systems are not solving in a final, convincing way. To solve this problem, a conceptual characterization of the data sets that constitute DID-based systems is essential. The use of foundational ontologies and conceptual modeling provides an adequate strategy to face the complexity of this problem by clarifying the data structure that is to be analyzed and managed. In this project we tackle this principle, by defining a method grounded on a conceptual model to develop eficient DID-based systems, and by making use of a well-grounded combination of Explainable Artificial Intelligence (XAI) and Machine Learning (ML) techniques to perform data analytics. In addition, the characterization of a platform for the implementation of the method is going to be designed and developed. The project's chosen domain of application is genomics, specifically in predicting critical diseases before symptoms manifest. Leveraging XAI and ML with genomic information can contribute to the advancement of precision medicine, allowing for the prediction of future diseases based on the available genomic data. The ML dimension will cover the predictive knowledge (is a disease present in a patient?), while the XAI dimension will deal with the explainable part (why the patient has the disease).</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org
CEUR
Workshop
Proceedings</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        Data has become an invaluable asset in today’s society, and its production is unparalleled,
continually increasing. This presents significant challenges for modern software platforms,
which must store, analyze, and quickly provide access to data for numerous users. Consequently,
various research fields related to data management and processing have undergone profound
transformations[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. One of the most current, relevant challenges in the software development
context is dealing with DID-based systems, which require extensive and heterogeneous datasets
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] to create knowledge from data. To develop efective and eficient methods and facilities for
data analysis and management, software developers must integrate complex, distributed, and
heterogeneous datasets from increasingly diverse data-generating technologies (e.g., sensors, the
internet, genome sequencing machines, and other sophisticated devices). Therefore, managing
this massive amount of data to find the most critical and actionable pieces of knowledge has
become a significant challenge.
      </p>
      <p>
        A fascinating example of DID-based systems are those that analyze the human genome [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
Understanding the human genome is a significant scientific challenge, requiring the
application of sound conceptual modeling techniques to manage such complex systems adequately.
The continuous generation of genomic data from improved sequencing technologies [
        <xref ref-type="bibr" rid="ref4 ref5 ref6">4, 5, 6</xref>
        ]
necessitates selecting the right data management strategy for software platforms. Developing
software systems to deal with these DID are key for a proper genome analysis that would lead
to anticipating future illness in the human population [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>
        To address these issues, this proposal will be grounded on an interdisciplinary scientific
policy especially interested in combining two strong lines of research: conceptual modeling
(CM) [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and explainable artificial intelligence (XAI) [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. To this aim, two main components
need to be explored, designed, and developed: i) A method to deal with DIDs problem’s
management (the methodological perspective) correctly and eficiently, and ii) a “materialization” of
the method in the form of a platform intended to assess the solution’s value in a challenging
and specially selected DID as the one related to the understanding of the human genome (the
practical perspective).
      </p>
      <p>
        In this scenario, applying a methodological framework based on XAI and CM to address
DIDs concerns efectively becomes a relevant, promising strategy that forms the basis of the
scientific approach used to achieve the project’s major goal. On the one hand, CM is recognized
as crucial for developing data-oriented computer systems, ensuring an accurate representation
of the application domain independently of the system that will be developed to address a
real-world problem. This is especially relevant when we want to “understand data” in a DID
context, which in our case applies to genomics. On the other hand, there is the application
of XAI principles [10, 11], which describe a system in which humans can easily understand
the results that an AIsystem provides, focusing primarily on understanding exactly “how” and
“why” decisions are taken to reach results [
        <xref ref-type="bibr" rid="ref9">9, 12</xref>
        ]. For DID-based systems, where the right
representation of concepts becomes a crucial step, CM becomes the perfect partner for a useful
XAI application [10] since by visualizing the relevant concepts, the structure of meaning people
use to understand the domain is clearly represented.
      </p>
      <p>Our approach -both methodological (a method) and practical (a platform for the genomics
domain)- is based on the group’s expertise [13, 14], focusing on understanding data’s true nature,
employing CM techniques, and addressing challenges such as data volume and processing.</p>
      <sec id="sec-2-1">
        <title>1.1. Details of the project</title>
        <p>The project combines XAI and CM for Data Intensive-Domains Management (CoMoDID). It is a
four-year project (Sept. 2022 – Dec. 2025). Currently, the Research team is constituted by Óscar
Pastor López, Juan Carlos Casamayor Ródenas, Tanja E. Vos, Lluís-F. Hurtado, Encarna Segarra,
Ferran Pla, Fernando García Granada, José F. Reyes Román (Postdoctoral Researcher), Alberto
García Simón (Postdoctoral Researcher) and Diana Martínez Minguet (Predoctoral Researcher),
in collaboration with the Genomics Team of the PROS research group. The project is supported
by the Generalitat Valenciana through the CIPROM/2021/023 project.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>2. Project goals, tangible outputs &amp; expected outcomes</title>
      <p>The research proposed in this project focuses on the design of solutions for DIDs problems since
existing frameworks to build DID-based systems lack a sound conceptual modeling grounding,
and too frequently, ad-hoc implementations are built. Both the method and the platform to
materialize the solution in order to show how it works for a selected DID conform to the two
major objectives of this project:
• Definition of a general method (the so-called DELFOS method (Figure 1), as the method
to be used for the CoMoDID project) applicable to any DID for facing its analysis and
design, which is based on a sound combination of conceptual modeling techniques and
XAI technologies.
• Development of a technological platform, that will instantiate and support the method in
a particularly challenging and complex DID context: the genomic domain.</p>
      <p>To address the methodological and practical components of the approach, we break these
objectives into specific goals (G) with associated work packages (WPs) which the tangible
outputs obtained so far result from.</p>
      <sec id="sec-3-1">
        <title>2.1. Specific goals for a general method</title>
        <p>G1. Ontological characterization of DIDs. (WP1: Ontological characterization of DIDs.):
The study and analysis of existing foundational ontologies related to DID characterization
in conjunction with the state of the art about existing solutions for the development of
DID platforms are reflected in the doctoral thesis:
– García Simón, A. (2022). Understanding the Code of Life: Holistic Conceptual
Modeling of the Genome. Universitat Politècnica de València. https:// doi.org/ 10.
4995/ Thesis/ 10251/ 191432
A preliminary definition of a foundational ontology for the development of DID platforms
is published in:
G2. Integration of XAI techniques for data management and exploitation. (WP2:
Integration of XAI techniques for data management): Focusing on the case study of the
genomic domain as one of the major DIDs that currently exist, a study and statistical
comparison of diferent data sources with information associated with two groups of
diseases: cancer and heart disease has been carried out. The results of these studies have
been reported in:
– Costa, M., García S, A., &amp; Pastor, O. (2022). A Comparative Analysis of the
Completeness and Concordance of Data Sources with Cancer-Associated
In order to replicate some of the criteria used by clinicians for genomic data to streamline
the process of selecting relevant data by reducing the efort of manual activities
performed by experts, diferent XAI techniques have been established according to the data
collections to be analyzed, and diferent data storage and integration techniques have
been evaluated.
(WP4: Analysis and design of a tool to help with the writing of medical case reports in
a genomic domain): As a first draft, a Deep Learning (Transformers) model has been
trained for a multi-class, multi-label classification task of radiology medical reports using
Transfer Learning techniques:</p>
      </sec>
      <sec id="sec-3-2">
        <title>2.2. Specific goals for a technological platform</title>
        <p>The instantiation of the method in the Genomics domain aims to validate that the method is
a very complex DID, to provide a technological platform to collect, manage and analyze the
generated data in practical settings in order to improve the understanding of the human genome
challenge, and ultimately to obtain relevant value by the extraction of knowledge from the data.
To achieve these objectives, the following goals are defined:</p>
        <p>G3. Definition of the interaction mechanisms for DID-based systems. (WP3: Interaction
Definition for DID-based Systems ): The analysis of the interaction requirements for
DIDbased systems and the elicitation of such requirements to design sustainable interfaces
are developed in:</p>
        <p>Remaining goals to be tackled are G4. Development of a platform to support the DELFOS
method and G5. DELFOS method and platform validation, with the associated
homonymous word packages (WP5 and WP6). This WPs leverage on all the previous ones which are in
the process of improvement and development.</p>
        <p>Finally, WP7: Communication, dissemination and exploitation of the results, is ubiquitous and
ensures consistent dissemination, visibility and outreach to all relevant stakeholders.</p>
        <p>The expected outcomes of the project are the development of a solid method that can be
applied to any complex DID, and the implementation of a platform that enables the instantiation
of the method for a particular DID-based system (genomics), providing the technological support.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>3. Relevance for ER</title>
      <p>The proposed project is aligned with several research topics relevant to the conceptual
modeling community. It is highly relevant to the topics of Ontological and cognitive foundations
and Semantics in conceptual modeling since the incorporation of foundational ontologies and
conceptual modeling in the project contributes to a solid theoretical foundation concerning
DID-based systems, in combination with the project’s focus on developing standardized
approaches for data integration and analysis which involves addressing semantic aspects. In the
same line, the project is relevant to the topic of Complex management of large conceptual
models, given that the project addresses the challenge of managing large and heterogeneous
data sets in complex DID-based systems.</p>
      <p>In another direction, the project aims to develop a method and platform to automate the
development of DID-based systems, including data modeling. In this context, using Artificial
Intelligence is useful for optimizing and automatizing data analysis. However, in the Precision
Medicine field, where the practical instantiation of the project is embedded, the necessity of
transparency and minimization of uncertainties is essential for the resulting decisions to be
explainable. XAI satisfies these requirements, thus being suitable for data analysis counceling.
The use of XAI and ML techniques involves knowledge representation and reasoning for
accurate data analysis, being directly related to the topic of Logic-based knowledge representation
and reasoning.</p>
      <p>Overall, the proposed project’s alignment with various research topics highlights its relevance
and potential contributions to conceptual modeling, as well as knowledge representation and
reasoning in the context of DIDs. It aims to address existing challenges and improve the
eficiency and accuracy of DID-based systems, ofering valuable insights for data analysts in
diverse research fields based on conceptual modeling techniques and foundational ontologies.</p>
    </sec>
    <sec id="sec-5">
      <title>4. Current Project Status</title>
      <p>The project is in the first quarter of its development and is well on schedule. So far, the project
tasks have involved the exhaustive characterization of the framework elements, as well as the
development of precursory solutions. The ongoing tasks concern further investigation and
the extension and improvement of the proposed solutions, for instance, the generation of a
preliminary platform prototype and precursory validation tests for both the method and the
platform. On the other hand, fruitful discussions with external companies have revealed future
lines of research which are being addressed in a new Ph.D. thesis conducted within the scope of
this project, regarding domain-specific concerns related to the genomic field.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>This work was supported by the Generalitat Valenciana through the CoMoDiD project (CIPROM/2021/023),
through a GVA-Predoctoral Research Grant (ACIF/2021/117), a Margarita Salas Grant, and the
Spanish State Research Agency through the DELFOS (PDC2021-121243-I00,MICIN/AEI/10.13039/501
100011033) and SREC (PID2021-123824OB-I00) projects, and co-financed with ERDF and the
European Union Next Generation EU/PRTR.
[10] O. Pastor, A. Palacio, J. Reyes Román, J. Casamayor, Modeling Life: A Conceptual
Schema-centric Approach to Understand the Genome, 2017, pp. 25–40. doi:10.1007/
978-3-319-67271-7_3.
[11] A. Barredo Arrieta, N. Díaz-Rodríguez, J. Del Ser, A. Bennetot, S. Tabik, A.
Barbado González, S. Garcia, S. Gil-Lopez, D. Molina, V. R. Benjamins, R. Chatila, F. Herrera,
Explainable artificial intelligence (xai): Concepts, taxonomies, opportunities and
challenges toward responsible ai, Information Fusion 58 (2019). doi:10.1016/j.inffus.2019.
12.012.
[12] J. Qin, M. Žumer, X. Wang, W. Fan, Conceptual models and ontological schemas for
semantically sustainable digital libraries, 2020, pp. 441–442. doi:10.1145/3383583.3398545.
[13] L. Kalinichenko, A. Volnova, E. Gordov, K. Nadezhda, D. Kovaleva, O. Malkov, I. Okladnikov,
N. Podkolodnyy, A. Pozanenko, N. Ponomareva, S. Stupnikov, A. Fazliev, Data access
challenges for data intensive research in russia 10 (2016) 2–22. doi:10.14357/19922264160101.
[14] S. Spreeuwenberg, Choose for AI and for Explainability, 2020, pp. 3–8. doi:10.1007/
978-3-030-40907-4_1.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Margara</surname>
          </string-name>
          , G. Cugola,
          <string-name>
            <given-names>N.</given-names>
            <surname>Felicioni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Cilloni</surname>
          </string-name>
          ,
          <article-title>A model and survey of distributed dataintensive systems</article-title>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Elizarov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Novikov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Stupnikov</surname>
          </string-name>
          ,
          <source>Data Analytics and Management in Data Intensive Domains</source>
          , Springer International Publishing,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Felderer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Russo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Auer</surname>
          </string-name>
          ,
          <source>On Testing Data-Intensive Software Systems</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>129</fpage>
          -
          <lpage>148</lpage>
          . doi:
          <volume>10</volume>
          .1007/978- 3-
          <fpage>030</fpage>
          - 25312-
          <issue>7</issue>
          _
          <fpage>6</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Cowley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Davis</surname>
          </string-name>
          ,
          <article-title>Next-generation sequencing and emerging technologies</article-title>
          ,
          <source>Seminars in Thrombosis and Hemostasis</source>
          <volume>45</volume>
          (
          <year>2019</year>
          ). doi:
          <volume>10</volume>
          .1055/s- 0039- 1688446.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>D.</given-names>
            <surname>Rigden</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Fernandez</surname>
          </string-name>
          ,
          <article-title>The 27th annual nucleic acids research database issue and molecular biology database collection</article-title>
          ,
          <source>Nucleic Acids Research</source>
          <volume>48</volume>
          (
          <year>2020</year>
          )
          <fpage>D1</fpage>
          -
          <lpage>D8</lpage>
          . doi:
          <volume>10</volume>
          . 1093/nar/gkz1161.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>W.</given-names>
            <surname>Mccombie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>McPherson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Mardis</surname>
          </string-name>
          ,
          <article-title>Next-generation sequencing technologies</article-title>
          ,
          <source>Cold Spring Harbor Perspectives in Medicine 9</source>
          (
          <year>2018</year>
          )
          <article-title>a036798</article-title>
          . doi:
          <volume>10</volume>
          .1101/cshperspect. a036798.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>B.</given-names>
            <surname>Louie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mork</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Martin-Sanchez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Halevy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Tarczy-Hornoch</surname>
          </string-name>
          ,
          <article-title>Data integration and genomic medicine</article-title>
          ,
          <source>Journal of biomedical informatics 40</source>
          (
          <year>2007</year>
          )
          <fpage>5</fpage>
          -
          <lpage>16</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.jbi.
          <year>2006</year>
          .
          <volume>02</volume>
          .007.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Olivé</surname>
          </string-name>
          ,
          <source>Conceptual Modeling of Information Systems</source>
          ,
          <year>2007</year>
          . doi:
          <volume>10</volume>
          .1007/ 978- 3-
          <fpage>540</fpage>
          - 39390- 0.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Spreeuwenberg</surname>
          </string-name>
          , AIX:
          <article-title>Artificial Intelligence needs explanation: Why and how transparency increases the success of AI solutions</article-title>
          ., Amsterdam: LibRT:
          <article-title>the Lab for Intelligent Business Rules Technology</article-title>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>