<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Artificial Intelligence in the New Scenario of Data Spaces</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Juan Trujillo</string-name>
          <email>jtrujillo@dlsi.ua.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alejandro Reina-Reina</string-name>
          <email>alejandro.reina@ua.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gustavo Candela</string-name>
          <email>gcandela@dlsi.ua.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Data Spaces, Collaborative Data Management, Data Governance, Federated Machine Learning</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Lucentia Research Group (DLSI), University of Alicante</institution>
          ,
          <addr-line>San Vicent del Raspeig</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Data spaces present significant opportunities for organizations to collaborate and leverage decentralized, interoperable, and secure data for advanced analytics and Federated Machine Learning. However, challenges such as ensuring data quality, managing privacy, and integrating heterogeneous data formats and semantics remain critical. Addressing these challenges requires robust data governance, real-time quality assurance, and the adoption of AI-data integration tools to improving decision-making and generating value in dynamic business ecosystems.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Data Spaces and Collaborative</title>
    </sec>
    <sec id="sec-2">
      <title>Models: Challenges and</title>
    </sec>
    <sec id="sec-3">
      <title>Opportunities in Data</title>
    </sec>
    <sec id="sec-4">
      <title>Management</title>
      <p>
        Today, the eficient use of data is a key factor in the success of
an organization. The advancement of organizations toward
more collaborative and interconnected models can foster
the emergence of new ways to manage and leverage data
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        In this regard, data spaces present opportunities to
enhance advanced analytics and foster collaboration between
companies, representing a crucial evolution that could shift
the big data paradigm, currently based on data warehouses
and data lakes. However, this evolution requires a high
level of maturity, with big data management and AI
development being key. Organizations must adapt to a dynamic
and complex data infrastructure, considering aspects such
as decentralization, interoperability, and data quality, which
pose essential technical and organizational challenges to
optimize the use of data in strategic decision-making and
value creation [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
    </sec>
    <sec id="sec-5">
      <title>2. Data spaces in business ecosystems</title>
      <p>
        To date, organizations rely primarily on traditional big data
architectures such as data warehouses and data lakes. These
are designed to consolidate structured data in a centralized
environment. However, with the advent of data spaces,
there is the challenge of managing data in a distributed
and decentralized manner while simultaneously adhering
to the principles of interoperability, security, privacy, and
data governance [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Furthermore, data spaces must address
the challenge of fostering trust-based collaboration among
participating organizations [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
nEvelop-O
Despite the challenges, data spaces enable organizations
to collaborate with each other and share information
without the need to move data from their original sources. The
advantages of data spaces are invaluable, as they facilitate
new opportunities for the metrics necessary for KPI
monitoring through distributed analytics models [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Moreover,
data spaces can support Federated Machine Learning [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ],
allowing companies to learn collaboratively while ensuring
that personal data remain confidential and private, without
leaving the organization’s private environment. This also
helps to comply with privacy regulations.
      </p>
    </sec>
    <sec id="sec-6">
      <title>3. Evaluating and improving the quality of data-driven services</title>
      <p>
        An inherent challenge that users will face in data spaces
is the evaluation of the quality and reliability of the data
they consume [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. It is true that data spaces ofer a rich
ecosystem of information sources, but their heterogeneity
raises critical questions: How can we ensure that the shared
data is accurate, up-to-date, and free from biases? What
mechanisms can guarantee that data-driven services are
consistent and reliable enough to be integrated into my
business process?
      </p>
      <p>
        Furthermore, in a context where companies can consume
third-party information to feed their business models, it is
crucial to have metrics that allow the evaluation of data
quality aspects [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], such as completeness, accuracy, or
consistency. This is because if the data are of low quality, the
results obtained, as well as AI-based decisions, may be
erroneous [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>To address this issue, organizations must implement
robust data governance processes, combined with
technological tools that enable the auditing and certification of data
quality, even in real time. Furthermore, the use of quality
labels or certifications could become a standard in these data
spaces, similar to how products or services are evaluated in
other sectors.</p>
    </sec>
    <sec id="sec-7">
      <title>4. Use of Heterogeneous Data with</title>
    </sec>
    <sec id="sec-8">
      <title>Diferent Formats and Semantics</title>
      <p>Data heterogeneity has been a recurring obstacle in
traditional information management systems, but it becomes
even more pronounced in distributed and decentralized
CEUR</p>
      <p>
        ceur-ws.org
ecosystems [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. In this regard, organizations face another
key challenge, which is the ability to work with data that
does not necessarily share the same format, structure, or
semantics as their own data [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. For example, a transport
service provider may receive trafic data from one source
and weather data from another, but analytical models must
be able to integrate these datasets to extract valuable
insights that can support decision-making.
      </p>
      <p>
        The solution to this challenge is not trivial; however, it
necessarily involves the use of standards such as
ontologies, as well as advanced technologies, including
integration and automatic transformation tools. In this context,
approaches based on embeddings [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] or AI [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] may
facilitate the alignment of heterogeneous data, enabling its
efective integration and utilization.
      </p>
    </sec>
    <sec id="sec-9">
      <title>Acknowledgments</title>
      <p>This work has been co-funded by the AETHER-UA project
(PID2020-112540RB-C43) funded by the Spanish Ministry of
Science and Innovation; the ENIA Chair of Artificial
Intelligence from the University of Alicante (TSI-100927-2023-6)
funded by the Recovery, Transformation and Resilience Plan
from the European Union Next Generation through the
Ministry for Digital Transformation and the Civil Service; and
the BALLADEER (PROMETEO/2021/088) project funded
by the Conselleria de Innovación, Universidades, Ciencia
y Sociedad Digital (Generalitat Valenciana), and Spanish
Ministry of Digital Transformation (TSI-100121-2024-10).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>L. M.</given-names>
            <surname>Camarinha-Matos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Afsarmanesh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Galeano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Molina</surname>
          </string-name>
          ,
          <article-title>Collaborative networked organizations - concepts and practice in manufacturing enterprises</article-title>
          ,
          <source>Computers &amp; Industrial Engineering</source>
          <volume>57</volume>
          (
          <year>2009</year>
          )
          <fpage>46</fpage>
          -
          <lpage>60</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.cie.
          <year>2008</year>
          .
          <volume>11</volume>
          .024.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>F.</given-names>
            <surname>Möller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Jussen</surname>
          </string-name>
          , V. Springer,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gieß</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. C.</given-names>
            <surname>Schweihof</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gelhaar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Guggenberger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Otto</surname>
          </string-name>
          ,
          <article-title>Industrial data ecosystems and data spaces</article-title>
          ,
          <source>Electronic Markets</source>
          <volume>34</volume>
          (
          <year>2024</year>
          )
          <article-title>41</article-title>
          . doi:
          <volume>10</volume>
          .1007/s12525-024-00724-0.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Huber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Wessel</surname>
          </string-name>
          , G. Brost,
          <string-name>
            <given-names>N.</given-names>
            <surname>Menz</surname>
          </string-name>
          , Building Trust in Data Spaces, Springer International Publishing,
          <year>2022</year>
          , pp.
          <fpage>147</fpage>
          -
          <lpage>164</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -93975-
          <issue>5</issue>
          _
          <fpage>9</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>R.</given-names>
            <surname>Kalmar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Rauch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dörr</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Liggesmeyer</surname>
          </string-name>
          ,
          <source>Agricultural Data Space</source>
          , Springer International Publishing,
          <year>2022</year>
          , pp.
          <fpage>279</fpage>
          -
          <lpage>290</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -93975-5_
          <fpage>17</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>B.</given-names>
            <surname>Farahani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. K.</given-names>
            <surname>Monsefi</surname>
          </string-name>
          ,
          <article-title>Smart and collaborative industrial iot: A federated learning and data space approach</article-title>
          ,
          <source>Digital Communications and Networks</source>
          <volume>9</volume>
          (
          <year>2023</year>
          )
          <fpage>436</fpage>
          -
          <lpage>447</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.dcan.
          <year>2023</year>
          .
          <volume>01</volume>
          .022.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S. R.</given-names>
            <surname>Carroll</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Garba</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O. L.</given-names>
            <surname>Figueroa-Rodríguez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Holbrook</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Lovett</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Materechera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Parsons</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Raseroka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Rodriguez-Lonebear</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Rowe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Sara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. D.</given-names>
            <surname>Walker</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Anderson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>Hudson, The CARE principles for indigenous data governance</article-title>
          ,
          <source>Data Sci. J</source>
          .
          <volume>19</volume>
          (
          <year>2020</year>
          )
          <article-title>43</article-title>
          . URL: https://doi.org/10.5334/dsj-2020-043. doi:
          <volume>10</volume>
          .5334/DSJ-2020-043.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>I. ISO</surname>
          </string-name>
          , Iec 25012.
          <article-title>software engineering-software product quality requirements and evaluation (square)-data quality model</article-title>
          ,
          <source>International Organization for Standardization</source>
          (
          <year>2000</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Alkatheeri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ameen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Isaac</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Al-Shibami</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Nusari</surname>
          </string-name>
          ,
          <article-title>The mediation efect of management information systems on the relationship between big data quality and decision making quality</article-title>
          ,
          <source>Test Engineering and Management</source>
          <volume>82</volume>
          (
          <year>2020</year>
          )
          <fpage>12065</fpage>
          -
          <lpage>74</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Naeem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Jamal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Diaz-Martinez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Butt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Montesano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. I.</given-names>
            <surname>Tariq</surname>
          </string-name>
          , E. D. la
          <string-name>
            <surname>Hoz-Franco</surname>
          </string-name>
          , E.
          <string-name>
            <surname>De-La-Hoz-Valdiris</surname>
          </string-name>
          ,
          <article-title>Trends and future perspective challenges in big data</article-title>
          ,
          <source>in: Advances in Intelligent Data Analysis and Applications</source>
          , volume
          <volume>253</volume>
          , Springer, Singapore,
          <year>2022</year>
          , pp.
          <fpage>309</fpage>
          -
          <lpage>325</lpage>
          . doi:
          <volume>10</volume>
          .1007/
          <fpage>978</fpage>
          -981-16-5036-9_
          <fpage>30</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>W.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. E.</given-names>
            <surname>Madnick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Cheung</surname>
          </string-name>
          ,
          <article-title>Discovering and reconciling value conflicts for numerical data integration</article-title>
          ,
          <source>Information Systems</source>
          <volume>26</volume>
          (
          <year>2001</year>
          )
          <fpage>635</fpage>
          -
          <lpage>656</lpage>
          . doi:
          <volume>10</volume>
          .1016/S0306-
          <volume>4379</volume>
          (
          <issue>01</issue>
          )
          <fpage>00043</fpage>
          -
          <lpage>6</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>C.</given-names>
            <surname>Koutras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Siachamis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ionescu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Psarakis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Brons</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fragkoulis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Lofi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bonifati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Katsifodimos</surname>
          </string-name>
          , Valentine:
          <article-title>Evaluating matching techniques for dataset discovery</article-title>
          ,
          <year>2021</year>
          . URL: https://arxiv.org/abs/
          <year>2010</year>
          .07386. arXiv:
          <year>2010</year>
          .07386.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>J.</given-names>
            <surname>García-Carrasco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Reina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lavalle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Maté</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Trujillo</surname>
          </string-name>
          ,
          <article-title>A conceptual model-based approach for exploiting large language model embeddings in automatic data integration</article-title>
          ,
          <year>2024</year>
          . doi:
          <volume>10</volume>
          .2139/ssrn.5024901.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>