<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>smart data holistic approach for context-aware data analytics (AETHER-UA)</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ana Lavalle</string-name>
          <email>alavalle@dlsi.ua.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alejandro Maté</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Juan Trujillo</string-name>
          <email>jtrujillo@dlsi.ua.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Miguel A. Teruel</string-name>
          <email>materuel@dlsi.ua.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alexander Sánchez Díaz</string-name>
          <email>alexander.sanchez@ua.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Lucentia Research Group - Department of Software and Computing Systems, University of Alicante</institution>
          ,
          <addr-line>Carretera de San</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Project Exhibitions</institution>
          ,
          <addr-line>Posters and Demos, and Doctoral Consortium</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Smart Data</institution>
          ,
          <addr-line>Data Integration, Data Bias, Conceptual Modeling, User's requirements, Modeling of Machine</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Vicent del Raspeig</institution>
          ,
          <addr-line>s/n, 03690, San Vicent del Raspeig</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>A smart data holistic approach for context-aware data analytics (AETHER-UA) is one of the four subprojects, developed in the University of Alicante, being part of the whole project AETHER. This project is being developed by four partners: (i) University of Malaga - Coordinator; (ii) University of Alicante, (iii) University of Castilla La-Mancha, and (iv) University of Seville. The project is funded by the Ministry of Science and Innovation. The main goal of this project is to advance towards a knowledge-based framework integrating novel solutions for data, process and business analytics. The research activities for designing and developing Aether will mainly focus on three main challenges: the characterization of the datasets, the improvement and automation of the algorithms, and the generation of mechanisms to enhance model explainability and interpretation of the results. The project is highly related to data processing, integration, analysis and modeling. More concretely, within the AETHER-UA project, several proposals are being developed for the modeling of user's requirements for Machine Learning applications, the developing of a framework based on Model Driven Development (MDD) for eXplanable Artificial Intelligence and several approaches for the data bias analysis.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR</p>
      <p>ceur-ws.org
CEUR
Workshop
Proceedings
the datasets such as semantics, topology, or statistical features, is critical for data optimization,
process and business analytics. These properties can be used to automatically select the best
algorithms depending on the problem and the input data, and fine-tune their configuration to
get high-quality results.</p>
      <p>However, the characterization of the data at the beginning of the analysis processes will
discover properties that may not be valid as the data are transformed or on the inferred knowledge.
So, the second challenge will be to research how these properties are transmitted throughout
the analysis workflows or how new properties can be derived, and how they can be used to:
facilitate or automate the selection and configuration of the algorithms needed in the analysis of
the process; measure the actionability of the results based on values of authority and provenance
or reliability and confidence; derive an explanation of why the result is the one we obtain and
be able to make it intelligible. Thus, we aim to move from the use of analysis algorithms such
as black boxes to facilitate the understanding and interpretability of the results by analysts.</p>
      <p>All of these steps forward cannot be isolated from some fundamental aspects: (1) data tend
to be managed through workflows that can be represented by using business processes, which
implies new challenges; (2) the quality of the data is crucial for later analysis; (3) not only the
data are relevant for analysis, cybersecurity aspects must also be included, and (4) the advance
after the application of the various techniques must be measured by using KPIs.</p>
      <p>All these research activities will be applied in real contexts by developing pilots in industrial
contexts in diferent application domains taking advantage of existing collaborations with other
research organizations or companies in a multidisciplinary approach: precision agriculture,
personalized medicine, bioinformatics, Smart Cities, logistics, industry 4.0 and remote sensing.</p>
      <p>This project is being developed by four partners: (i) University of Malaga - Coordinator; (ii)
University of Alicante, (iii) University of Castilla La-Mancha, and (iv) University of Seville. The
project is funded by the Ministry of Science and Innovation and has a duration of four years.
2. Project information
• Project name: A smart data holistic approach for context-aware data analytics (Aether)
• Duration: From Jan, 1st, 2021 until Dec, 31st, 2024
• Partners: In the following, we classify the participants grouped by the corresponding
partner participating in the project
– Aether-UMA: KHAOS Research group - University of Málaga (UMA). Khaos
main background is related to semantics technologies.
– Aether-UCLM: Alarcos and GSyA research groups - University of Castilla-La
Mancha (UCLM). The Alarcos Research Group main background is data quality,
the quality of databases, data warehouses, data models, and data itself. The GSyA
research group’s main background is related to software development security and
security management and governance.
– Aether-UA: Lucentia Research group - University of Alicante (UA). The
Lucentia Group has extensive experience in the field of Business Intelligence (BI),
Data Analytics, Big Data, and Conceptual modeling of BI applications.
– Aether-US: IDEA Research group - University of Sevilla (US). The IDEA Group
is devoted to research on issues related to business process models and diagnosis of
systems based on data analysis.
• Funding agency: Ministry of Science and Innovation
• Website: https://aether.es/</p>
    </sec>
    <sec id="sec-2">
      <title>3. Project goals</title>
      <p>• General goal 1 (responsible UMA): To design the semantic core for the Aether
framework to deal with relevant aspects of data, process and business analytics.
• General goal 2 (Responsible UCLM): To investigate data characterization mechanisms
for the analysis of datasets (not only Big Data) to obtain semantic (domain), topological,
statistical (data biases, kind of distribution, data type, descriptive statistics, categories)
and other properties (e.g., quality, security, provenance).
• General goal 3: (Responsible US) To design and develop new solutions for improved
data analytics, process analytics and business analytics, taking advantage of the analysis
of the datasets and processes.
• General goal 4 (Responsible UA): To facilitate the interpretability and the explainability
of analytic results, with particular focus on aligning Machine Learning techniques with
business objectives. This will lead to new insights and the identification of novel, relevant
analysis.</p>
      <p>– Specific goal 4.1 (I. Navas): To study the use of knowledge graphs and ontologies
to enhance the explicability of the models (of the algorithms) used in the analysis
of data, to improve the actionability of results maintaining the privacy rules about
data determined.
– Specific goal 4.2 (J.C. Trujillo): To define a methodology to derive indicators
and visualizations from the metadata generated by the various investigations and
consolidated in BIGOWL, taking into account user goals, semantics and machine
learning algorithms outputs.
– Specific goal 4.3 (A.J. Varela): Reasoning about the security vulnerabilities and
vulnerable configurations to improve the interpretability and usability based on the
results over the datasets and the analysis of their security requirements, features,
configurations, and the information of the system.
– Specific goal 4.4 (J.A. Cruz): To define quantitative (synthetic) indicators to
measure the interpretability of the results of data analysis and machine learning
algorithms. These indicators will allow us to warn users about underlying aspects
that may afect their interpretation.
– Specific goal 4.5 (M.T. Gómez): To study the alignment between the applied
techniques in the data analytics of business processes, and the obtained information,
using the knowledge graphs and ontologies to facilitate the interpretability and
usability of the results (of the algorithms) of the data analysis even before the
algorithms are applied to ascertain the complexity and the usability of the used
techniques.
– Specific goal 4.6 (A. Maté): To develop reasoning techniques that use the
semantics of the analytic workflow and KPI definitions to facilitate the identification of
potential causes for underperforming objectives as well as their implications for the
organization.
• General goal 5 (Responsible UMA): To develop ”Piloting” actions, through the
development of pilot experiences with companies and organizations in the application of the
results of the project.</p>
    </sec>
    <sec id="sec-3">
      <title>4. Methodology and Work Plan</title>
      <sec id="sec-3-1">
        <title>4.1. Methodology</title>
        <p>Our work methodology builds on the reference framework provided by the Enterprise Unified
Process (Enterprise Unified Process. Several authors. Enterprise Unified Process Publishing,
2020, ISBN: 978-1867420781), which proposes an iterative, incremental life cycle that is very
appropriate for research projects. It has been adapting it for the last decade, so it is very
solid nowadays. As a summary, the main phases into which the project was divided and
then the disciplines that must be performed are as follows: (i) Inception, (ii) Preparation, (iii)
Construction, (iv) Transition, and (v) Recapping.</p>
        <p>• Inception: the goal is to identify the working hypotheses and the objectives, verify that
they are realistic and aligned with current trends, ensure that they are well aligned with
the existing research programs, and check that they will have positive scientific, technical,
social and economic impacts. It is a good idea to use brainstorming meetings in this phase,
so that researchers may contribute with their ideas. This phase has been developed at the
beginning of the project proposal preparation.
• Preparation: the goal is to work on the general goal and the motivation of the project
plus a summary of the state of the art that supports them. It is a good idea to use a
brainstorming meeting so that every researcher can contribute to his or her ideas. After
that, the focus must shift towards devising the work packages, their specific goals, the
dissemination, transfer, and contingency plans. It is a good idea to work in small groups
in brainstorming meetings so that every researcher can contribute ideas. This phase
has been developed during the project proposal preparation. Still, it will be revisited in
the first month of the project development to update it according to any new insight
discovered.
• Construction: this phase focuses on defining the goals within each work package and
dividing these work packages into several tasks. It is a good idea to work in small groups
that organise a kick-of meeting per goal in which they agree on the tasks to be done,
the workflow, the intermediate milestones, and the key performance indicators. Then,
the researchers work individually or very small groups and meet periodically in plenary
workshops in which they present their results, discuss new ideas, and identify deviations
and corrective actions.
• Transition: the goal is to transfer the research results to the end-users and the industry. We
will pay special attention to transferring them to the companies that support the proposal
using seminars and workshops. We will also pay special attention to complementary
projects in which we can apply our results in the context of their industrial projects.
• Recapping: the goal is to recap on how the project has gone on to identify strong and
weak points that may help work better in future projects.</p>
        <p>Our software development methodology combines Kanban and iterative prototyping. Kanban
is an agile methodology based on keeping a board to visualise the tasks to carry out (backlog), in
progress, and finish. Thus, the team members can see all the tasks in their context. When some
tasks are completed, the next one is taken from the backlog. The idea behind using prototypes
is to develop incomplete versions of the software project to get a working system quickly. So, it
can be easily tested, and the users can provide feedback early. Once a prototype is ready, a new
one is planned to cover new features, and so on. Prototyping can be combined with continuous
integration, so that software is produced in short cycles to ensure that it can be ready to be
released at any time by frequently applying the steps of testing, building, and releasing. The
project will be managed with the help of OpenProject (https://www.openproject.org/), which
was specifically designed to help with project planning and scheduling, to define user stories
and tasks, to assign them to developers, and to report on timings and costs.</p>
        <p>The source code of the software projects will be monitored by the Travis ( https://www.
travis-ci.org) continuous integration system. The applications will be packed into containers
using Docker (https://www.docker.com), which has proven to help deploy the results very
easily.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5. Work plan</title>
      <p>The table shown in Figure 2 summarises the work plan and how it is related to interdisciplinary
dimensions of the project (* means that the task involves several dimensions). In the following
Section 5.1, we describe the main work Package developed by the University of Alicante pointing
out the main role of conceptual modeling in diferent areas.</p>
      <sec id="sec-4-1">
        <title>5.1. WP4: Model interpretability and result explainability (WP Leader: A. Maté)</title>
        <p>The goal of this WP is to facilitate more confident and accurate decisions by improving the
interpretability and explainability of analytic workflows and outputs. The aim is to provide
decision-makers with additional pieces of evidence and insights on the rationale behind
analytical workflows such as data subsets that lead to the activation of Machine Learning &amp; Artificial
Intelligence rules, the semantic relationship between the data being analyzed and other data
sets stored by the organization, or the connection between analytic outputs and business goals
and their implications.</p>
        <p>• Task 4.1. (UMA, UA, US, UCLM): Knowledge graph based explainability (9 Months).</p>
        <p>This task will analyse the use of knowledge graphs to enhance the understanding of the
datasets and the models used to transform them through data analytics. The semantics
defined at the diferent levels of knowledge will enable the exploration of these knowledge
graphs to extract relevant knowledge. Data privacy will be preserved when exploring
external knowledge graphs.</p>
        <p>– Deliverable 4.1 [REPORT] Analysis of knowledge graphs to improve model
explainability (M39).
• Task 4.2. (UA, UMA): Semantically enhanced dashboards and scorecards (9 months).</p>
        <p>This task aims at developing a methodology for deriving dashboards and scorecards
exploiting the semantic information in the Aether framework. To this aim, previous
works of the UA will be combined with the expertise on semantics from UMA to create a
methodology that takes into account not only user goals and the data at hand, but also
data semantics to create richer visualizations and dashboards.</p>
        <p>– Deliverable 4.2 [REPORT] Methodology for the derivation of semantically
enhanced dashboards and scorecards (M39).
• Task 4.3. (US, UCLM): Testing of Vulnerabilities (9 months). Automatic creation
of security tests from exploiting databases based on the extracted vulnerabilities from
relevant databases to create models that describe the possible vulnerabilities and evolution
of the systems. These models will be used to verify and diagnose the running configuration
further to define tests according to their potential vulnerabilities.</p>
        <p>– Deliverable 4.3 [REPORT, SOFTWARE] Report the automatic extraction of the
models from databases, and software to verify and diagnose configurations, and
security tests (M41).
• Task 4.4. (UCLM, UMA, UA): Data Analytics Results’ Interpretability (12 months). The
interpretation of the results of the Machine Learning models is vital because (1) it helps to
correct unwanted biases, (2) it facilitates the detection of noise that alters the prediction
and (3) it allows to confirm that the variables that the model has selected are significant
in the framework of the problem. However, when faced with diferent possible models,
there is currently no mechanism for measuring the level of interpretability of the model.
Therefore, in this task, a series of experiments will be carried out to identify possible
metrics that will allow the user to define the level of interpretability of the algorithms.
– Deliverable 4.4 [REPORT] Experimental results in measuring interpretability of
data analytics results in the project (M44).
• Task 4.5. (US, UMA, UA): Alignment among processes and discovery techniques (10
months). How the best configuration of the process mining techniques can be ascertained
before the techniques are applied. To fulfil this task, it is necessary to characterise the
processes to discover the relation between the parameters of the techniques and the
processes where they are applied.</p>
        <p>– Deliverable 4.5 [REPORT] Methodology to facilitate the alignment of the most
suitable algorithms and setting according to data sets in process discovery (M44)
• Task 4.6. (UA, UMA): Business reasoning framework (10 months). This work package
aims at developing a novel methodology for reasoning on the current performance of the
organization by making the system aware of existing KPIs and analytic workflows (ML
outputs and data transformations) to provide as much insights as possible on the root
cause of underperforming business objectives.</p>
        <p>– Deliverable 4.6 [REPORT] Business reasoning framework description (M46).</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>6. Project relevance</title>
      <p>Aether project aims to enhance the whole data analytics process, seamlessly augmenting
datasets and algorithms with semantic and domain knowledge, while also improving their
security, quality and actionability. The expected result is a new way to approach data analytics,
enabling advanced reasoning and applications with less efort, augmenting the data catalogue
incorporated in the business decisions. In turn, this change has profound implications for
academia, society, and industry. Numerous organizations, such as hospitals, universities or
industries, can harness this new approach to easily integrate new datasets into their analysis
while having a clearer view of their implications for their objectives.</p>
    </sec>
    <sec id="sec-6">
      <title>7. Current state</title>
      <p>
        As it can be summarized, the project is highly focused on “Data”. In a more particular way,
we should point out that conceptual modeling has a relevant role in many Work Packages and
transversal tasks. One of the main novel research lines under development in this project is the
modeling of user’s requirements for Machine Learning Applications. Currently, we are working
on the extension of the iStar [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] modeling approach in order to be able to gather the main ML
requirements.
      </p>
      <p>
        On the other hand, data bias is one of the most relevant issues to be tackled when working
with external data and/or developing ML models. In this way, we have developed several works
where conceptual modeling and visualization play relevant roles in finding the main data bias
in a set of data [
        <xref ref-type="bibr" rid="ref3 ref4 ref5 ref6">3, 4, 5, 6</xref>
        ]. These works can be applied to small and also Big Data sources [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ].
On the other hand, one of the current research lines will be focused on the modeling of ML
models in order to model the most relevant data to train the ML models [
        <xref ref-type="bibr" rid="ref10 ref11 ref12 ref13 ref9">9, 10, 11, 12, 13</xref>
        ].
      </p>
      <p>
        We also apply diferent Process Mining techniques for the discovery of defacto process models,
conformance analysis and extensions by combining AI models. We use algorithms such as
Fuzzy Miner to manage levels of abstraction in the models by tackling both correlation and
relevance variables, and the Trace Alignment algorithm, for clustering based on Agglomerative
Hierarchical Clustering (AHC) to detect patterns in process traces [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>C.</given-names>
            <surname>Barba-González</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>García-Nieto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. del Mar</given-names>
            <surname>Roldán-García</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Navas-Delgado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. J.</given-names>
            <surname>Nebro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. F.</given-names>
            <surname>Aldana-Montes</surname>
          </string-name>
          ,
          <article-title>Bigowl: Knowledge centered big data analytics</article-title>
          ,
          <source>Expert Systems with Applications</source>
          <volume>115</volume>
          (
          <year>2019</year>
          )
          <fpage>543</fpage>
          -
          <lpage>556</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Barrera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Reina-Reina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lavalle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Maté</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Trujillo</surname>
          </string-name>
          ,
          <article-title>An extension of istar for machine learning requirements by following the prise methodology</article-title>
          ,
          <source>Available at SSRN</source>
          <volume>4358075</volume>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Lavalle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Maté</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Trujillo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>García-Carrasco</surname>
          </string-name>
          ,
          <article-title>Law modeling for fairness requirements elicitation in artificial intelligence systems</article-title>
          ,
          <source>in: Proceedings of the 41st International Conference on Conceptual Modeling, ER 2022</source>
          , Springer,
          <year>2022</year>
          , pp.
          <fpage>423</fpage>
          -
          <lpage>432</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M. del Mar</given-names>
            <surname>Roldán-García</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>García-Nieto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Maté</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Trujillo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. F.</given-names>
            <surname>Aldana-Montes</surname>
          </string-name>
          ,
          <article-title>Ontology-driven approach for kpi meta-modelling, selection and reasoning</article-title>
          ,
          <source>International Journal of Information Management</source>
          <volume>58</volume>
          (
          <year>2021</year>
          )
          <fpage>102018</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Lavalle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Maté</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Trujillo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>García</surname>
          </string-name>
          ,
          <article-title>A methodology based on rebalancing techniques to measure and improve fairness in artificial intelligence algorithms</article-title>
          ,
          <source>in: Proceedings of the 24nd International Workshop on Design, Optimization, Languages and Analytical Processing of Big Data DOLAP@EDBT/ICDT</source>
          <year>2022</year>
          , volume
          <volume>3130</volume>
          <source>of CEUR Workshop Proceedings, CEUR-WS.org</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>81</fpage>
          -
          <lpage>85</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Lavalle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mate</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Trujillo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Teruel</surname>
          </string-name>
          ,
          <article-title>A data analytics methodology to visually analyze the impact of bias and rebalancing</article-title>
          ,
          <source>IEEE Access 11</source>
          (
          <year>2023</year>
          )
          <fpage>56691</fpage>
          -
          <lpage>56702</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Maté</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Peral</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Trujillo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Blanco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>García-Saiz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Fernández-Medina</surname>
          </string-name>
          ,
          <article-title>Improving security in nosql document databases through model-driven modernization</article-title>
          ,
          <source>Knowledge and Information Systems</source>
          <volume>63</volume>
          (
          <year>2021</year>
          )
          <fpage>2209</fpage>
          -
          <lpage>2230</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>C.</given-names>
            <surname>Blanco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>García-Saiz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. G.</given-names>
            <surname>Rosado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Santos-Olmo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Peral</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Maté</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Trujillo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Fernández-Medina</surname>
          </string-name>
          ,
          <article-title>Security policies by design in nosql document databases</article-title>
          ,
          <source>Journal of Information Security and Applications</source>
          <volume>65</volume>
          (
          <year>2022</year>
          )
          <fpage>103120</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>García-Ponsoda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>García-Carrasco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Teruel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Maté</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Trujillo</surname>
          </string-name>
          ,
          <article-title>Feature engineering of eeg applied to mental disorders: a systematic mapping study</article-title>
          ,
          <source>Applied Intelligence</source>
          (
          <year>2023</year>
          )
          <fpage>1</fpage>
          -
          <lpage>41</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J.</given-names>
            <surname>García-Carrasco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Maté</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Trujillo</surname>
          </string-name>
          ,
          <article-title>A data-driven methodology for guiding the selection of preprocessing techniques in a machine learning pipeline</article-title>
          ,
          <source>in: International Conference on Advanced Information Systems Engineering</source>
          , Springer,
          <year>2023</year>
          , pp.
          <fpage>34</fpage>
          -
          <lpage>42</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>R.</given-names>
            <surname>Tardío</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Maté</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Trujillo</surname>
          </string-name>
          ,
          <article-title>Beyond tpc-ds, a benchmark for big data olap systems (bdolap-bench)</article-title>
          ,
          <source>Future Generation Computer Systems</source>
          <volume>132</volume>
          (
          <year>2022</year>
          )
          <fpage>136</fpage>
          -
          <lpage>151</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>J. M. Barrera</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Reina</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Mate</surname>
            ,
            <given-names>J. C.</given-names>
          </string-name>
          <string-name>
            <surname>Trujillo</surname>
          </string-name>
          ,
          <article-title>Fault detection and diagnosis for industrial processes based on clustering and autoencoders: a case of gas turbines</article-title>
          ,
          <source>International Journal of Machine Learning and Cybernetics</source>
          <volume>13</volume>
          (
          <year>2022</year>
          )
          <fpage>3113</fpage>
          -
          <lpage>3129</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>A.</given-names>
            <surname>Reina Reina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Barrera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Valdivieso</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.-E. Gas</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Maté</surname>
            ,
            <given-names>J. C.</given-names>
          </string-name>
          <string-name>
            <surname>Trujillo</surname>
          </string-name>
          ,
          <article-title>Machine learning model from a spanish cohort for prediction of sars-cov-2 mortality risk and critical patients</article-title>
          ,
          <source>Scientific Reports</source>
          <volume>12</volume>
          (
          <year>2022</year>
          )
          <fpage>5723</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>J.-F. Rodríguez-Quintero</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Sánchez</surname>
            ,
            <given-names>L. Iriarte</given-names>
          </string-name>
          <string-name>
            <surname>Navarro</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Maté</surname>
            ,
            <given-names>M. Marco</given-names>
          </string-name>
          <string-name>
            <surname>Such</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Trujillo</surname>
          </string-name>
          ,
          <article-title>Fraud audit based on visual analysis: A process mining approach,</article-title>
          <year>2021</year>
          -
          <volume>05</volume>
          -21.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>