<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>March</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Exploring Content-Based Catalogs for Enhanced Discovery Services in Data Spaces</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Adriana Morejón</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alberto Berenguer</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lucía de Espona</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>David Tomás</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jose-Norberto Mazón</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University Institute for Computing Research, Department of Software and Computing Systems, University of Alicante</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2025</year>
      </pub-date>
      <volume>25</volume>
      <issue>2025</issue>
      <fpage>0000</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>In the realm of data sharing, efective data discovery is critical for fostering collaboration and innovation across organizations. Data spaces merge the interoperability and secure sharing features of data ecosystems with the transactional and economic aspects of data markets, enabling seamless data exchange among data providers and consumers while preserving data sovereignty. Central to this paradigm is the data catalog, an entity that handles metadata and provides it for data discovery services. However, traditional data catalogs rely heavily on metadata related to high-level representation of data, as well as the overall features of datasets (e.g., dataset name, license or keywords, following standards such as DCAT). This limitation hampers data discovery, as these metadata alone may not efectively describe datasets to fully convey the relevance of datasets. To address this challenge, this paper proposes content-based catalogs to enhance data discovery within data spaces. Our approach for content-based catalogs incorporates three key assets for data consumers to discover relevant datasets and evaluate their utility before accessing them: (i) high-quality structural metadata for datasets; (ii) representative data samples from the dataset; and (iii) a discovery service for searching relevant datasets. Our novel content-based catalog reduces the risk of unintended exposure while still showcasing the value of data, thus preserving the interests of both data providers and consumers within data spaces.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;data space</kwd>
        <kwd>data discovery</kwd>
        <kwd>data catalog</kwd>
        <kwd>data samples</kwd>
        <kwd>metadata</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        In today’s data-driven landscape, the efective management
and exchange of data are central to both organization
strategy and innovation. Data ecosystems and data markets,
in particular, represent critical frameworks for facilitating
data exchange between participants. While data ecosystems
focus on allowing organizations to share data in secure
environments with an emphasis on interoperability and
privacy, data markets are transactional platforms where
data is monetized. Data spaces emerge from the confluence
of data ecosystems and data markets [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. They leverage
the collaborative and interoperable nature of ecosystems to
ensure seamless data sharing among participating
organizations, while also enabling structured transactions and value
exchanges, as in data markets. This combined approach
responds to both the strategic and operational data needs and
ofers of participants (data consumers and data providers),
fostering a dynamic scenario where data must be
discoverable before being reused in a variety of applications [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ],
while data sovereignty is preserved [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. This brings
attention to the critical role of data discovery in data spaces, as
data consumers must identify the appropriate datasets from
data providers to develop products and services. This
necessity has driven a growing focus on creating innovative
mechanisms that enable the eficient and efective
discovery of relevant data [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Traditionally, data discovery and
exploration rely on metadata provided from data catalogs to
help data consumers find relevant data for their needs [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ],
even within enterprises [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ].
      </p>
      <p>
        Data catalogs in data spaces are typically based on DCAT
standard [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], which have proven efective in the context of
open data. DCAT metadata includes high-level descriptors
of datasets, such as title, description, keywords, identifier,
classification, publisher, license, and technical metadata (e.g.,
access URLs, format, and size). Although DCAT metadata
facilitates the retrieval of basic data sets on open data portals,
they are inherently limited to address more complex data
discovery needs [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. This limitation can hinder the
identification of datasets that precisely meet specific requirements
or use cases, particularly in complex or domain-specific
scenarios such as data spaces [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Data catalogs for data spaces
must semantically describe data content and structure of
data, thus allowing data consumers to eficiently discover,
understand, and assess the relevance of available data for
specific use cases [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. In order to preserve organizational
data sovereignty, current approaches for data space catalogs
ofer metadata as semantic descriptions of available data
sources without ofering the content itself [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>
        Interestingly, leveraging the content of datasets
themselves can enhance discovery processes by enabling data
consumers to evaluate data and verify alignment with their
needs. However, these content-based catalogs pose a
significant challenge to data spaces due to the Arrow paradox [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
This economic theory highlights the inherent dificulty of
valuing data before it is shared because, once the data is
disclosed, its value is essentially transferred to the recipient. If
content-based catalogs are used in data spaces for improving
discoverability, this creates a dilemma for data providers:
to attract potential consumers, they must reveal enough
about their datasets to demonstrate value, but doing so risks
exposing data in a way that diminishes its sovereignty.
Similarly, data consumers face uncertainty about the usefulness
of data before obtaining it, making it challenging to justify
investment or commitment in the data space.
      </p>
      <p>Therefore, novel content-based catalog solutions must
balance the need for richer discovery mechanisms with the
constraints of data sovereignty. This requires a new
generation of data catalogs capable of managing two levels of
content-based metadata: (i) descriptions of dataset content
derived from its internal structure (e.g., field names and
descriptions), and (ii) controlled, privacy-preserving data
samples that facilitate dataset evaluation without disclosing
all the data.</p>
      <p>
        To face this challenge, this paper proposes content-based
catalogs that include: (i) structural metadata by using
Datapackage of Frictionless Data project [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]; (ii) data samples as
the subset of data that behaves in the same way that the
original dataset for data discovery [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], as well as (iii) a discovery
service [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] built upon a metadata repository harvesting
the above information across the multiple data providers of
the data space. Our novel catalog for data sapces
streamlines the data discovery process for consumers, assessing
the relevance of data without fully exposing its content
until agreements for sharing data between consumers and
providers are in place.
      </p>
      <p>The remainder of this paper is structured as follows:
Section 2 presents related work to the field of data spaces
architectures; Section 3 describes our architecture for data
spaces that considers a novel content-based catalog for data
discovery services; content-based catalog for data spaces
is described in Section 4; finally, Section 5 sketches out
conclusions and future work.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        A data space is a federated data infrastructure that supports
trustworthy data sharing among data providers and data
consumers [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. To enable this, data spaces rely on
connectors [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], which act as interfaces between participants
and the ecosystem. Connectors ensure technical
interoperability and enforce access policies, allowing providers to
keep data sovereignty when sharing data while granting
consumers seamless access without compromising control.
Hence, data spaces enable new options for value creation
where providers and consumers easily interact and work
together to find, access, publish, consume, and reuse data,
as well as to stimulate innovation [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ].
      </p>
      <p>
        Numerous research initiatives have resulted in conceptual
frameworks and reference architecture models to accelerate
the development of data spaces [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ].
      </p>
      <p>
        The i4Trust initiative emerges as a collaboration program,
targeting the creation of data spaces by proposing an
architecture based on commonly agreed building blocks [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ].
They are a combination of components from FIWARE and
iSHARE, two data space foundations. The iSHARE
Foundation [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] maintains a trust framework for data spaces. With
the iSHARE Trust Framework, base components are
available for data spaces, aligned with the European Strategy
for Data1 and reference architectures like the International
Data Spaces Association (IDSA) and Gaia-X described
below. The convergence of these architectures enable a unified
data platform, contributing to a digital maturity model for
building a data ecosystem that strives for standardization.
The FIWARE Foundation [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] introduces an open source
framework that promotes the development of interoperable,
smart solutions.
      </p>
      <p>
        Another innovative data space technical framework has
been developed by Gaia-X, aiming to establish a trustworthy
ecosystem where data is shared, maintaining the user’s
digital sovereignty of the data. This standard-based framework
allows the implementation of distributed data systems in
all European countries in a legally secure manner, enabling
compliance with GDPR and other data regulations [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ].
      </p>
      <p>
        The International Data Spaces Association (IDSA) is a
non-profit organization focusing on establishing and
promoting standards for data spaces as trusted environments
1https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX%
3A52020DC0066
where organizations can share data while retaining full
control over its use [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]. The IDS Reference Architecture Model
(IDS-RAM) [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ], materializes the standard developed from
collecting requirements from various industries and the
results gained from the model’s implementation. Additionally,
the IDS-RAM connectors fit the GAIA-X principles and
architecture model [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ]. IDSA-RAM is suitable for industrial
applications, as shown by the architecture developed by the
Fraunhofer-Gesellschaft [
        <xref ref-type="bibr" rid="ref1 ref17 ref28">28, 1, 17</xref>
        ].
      </p>
      <p>One of the core building blocks of those reference
architectures is the data catalog, which serves as a structured
inventory of available datasets and associated metadata and access
policies. These catalogs enable data discovery while
preserving the sovereignty and privacy of the underlying datasets.
Unfortunately, current reference architectures incorporate
catalogs that are insuficient for efective data discovery,
posing two key limitations that hinder data consumers from
accurately evaluating relevance of shared datasets:
• Limited metadata: existing data catalogs primarily
rely on high-level metadata to describe datasets but
lack detailed structural information of dataset
content, such as field names or descriptions, which are
essential for precise dataset assessment.
• Content blindness: since federated environments
must preserve data sovereignty, current data
catalogs exclude actual dataset content, relying solely
on metadata. This prevents data consumers from
gaining deeper insights into dataset relevance before
access is granted.</p>
      <p>To overcome these challenges in data discovery, reference
architectures for data spaces require complementing data
catalogs with content-based metadata while protecting data
sovereignty. In response to these requirements, the
following section proposes an evolution of a reference architecture
that integrates content-based metadata and data discovery
mechanisms that respect data sovereignty.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Architecture</title>
      <p>
        From the reference architectures described in the previous
section, we have chosen the one proposed by IDSA as it has
shown compatibility with others. Additionally, it provides
an open source ready-to-use Docker-based implementation
of a Minimally Viable Data Space (IDS-MVDS) [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ]. The
IDS-MVDS consists of the following elements:
• Identity Provider: the IDS ofers identity
management across participants of the data space, according
to modern standards with low organizational
hurdles. Two elements compose the Identity provider
in the IDS-MVDS:
– Certificate Authority (CA) : responsible for
issue and manage technical identity claims.
– Dynamic Attribute Provisioning Service
(DAPS): provides short-lived tokens with
upto-date information about connectors.
• Metadata Broker: The Metadata Broker contains
an endpoint for the registration, publication,
maintenance, and query of Self-Descriptions from the data
providers.
• Data consumers and data providers: IDS
connectors that request or ofer data within the data space.
A connector can be simultaneously providing its
own data and consuming from another connector.
      </p>
      <p>
        The IDS Connectors metadata and data in the IDS
ecosystem are structured hierarchically according to the data
model based on the structure of the IDS information model.
The main entity of this data model is the resource as it
contains the core metadata of a data object, including the title,
description and license information. Related groups of
resources are organized by catalogs, the top entity in the data
model. A resource also has a list of representations that
describes the format of the ofered data. Below the
representation, there is the artifact that has a 1:1 relation to the
raw data and includes low level information such as
checksum and byte size. Each artifact has a reference to contract
agreements, which describe the agreed usage between data
provider and data consumer. Contract ofers can contain
multiple rules representing IDS Usage Control Patterns.
Further details about the data model entities and their attributes
are available at [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ].
      </p>
      <p>Among the previously listed IDS-MVDS components, the
Metadata Broker plays a crucial role in aggregating metadata
that describes the variety of connectors and catalogs in the
entire data space ecosystem. Serving as a central repository,
the broker primarily focuses on receiving metadata from
data providers and delivering it to data consumers within
the data space. The broker ofers search functionalities
but only contains metadata up to the catalog level which
is insuficient for a potential data consumer to discover
relevant data in the data space. Furthermore, it will require
the consumer to access each suitable data provider catalog
to request detailed metadata to make an informed decision
about starting the negotiation for a particular data ofering.</p>
      <p>The efectiveness of the metadata broker’s search service
largely depends on the accuracy and completeness of the
metadata provided by data providers at the resource level.
However, because these metadata entries are often
manually curated, they may lack suficient detail to fully describe
the dataset, limiting the precision of search and retrieval
processes. Additionally, data space catalogs do not provide
metadata at the artifact level or dataset samples,
preventing users from evaluating dataset relevance before access
is granted. Unlike open data portals, where datasets can
be freely downloaded for assessment, data spaces impose
stricter access controls, making discovery more challenging.</p>
      <p>Our proposed solution overcomes the limitations of the
IDS-MVS architecture by introducing a discovery service
that leverages an enhanced data provider catalog with
content-based metadata at artifact level, i.e. structural
metadata describing data schema (field names, data types, and
descriptions of the fields), as well as sample data instances.
These content-based metadata is extracted from the datasets
and combines it with the eventually existing metadata at
resource level (title, keywords, etc.) to generate a
comprehensive description of the data ofering. This information
is then integrated into the data provider catalog to make
it available for the data consumers. Specifically, for each
ofered dataset, an additional sample resource is generated
automatically and added to the catalog. This sample
resource contains both structural metadata, and a
representative subset of the full dataset. The original resource contains
a link to this sample resource, following the IDS metadata
standard which contains a specific element with this aim.
This sample resource is stored locally on the data provider
for a simplified retrieval as shown in Figure 1.</p>
      <p>
        These enhancements to the data space reference
architecture empower potential data consumers with detailed
insights into a dataset’s structure and representative samples
before initiating full access negotiations. This is achieved
through advanced data discovery services. In this context,
we introduce InferIA [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], an search engine, based on
artificial intelligence, that utilizes our content-based catalog to
enhance search accuracy and relevance, outperforming
traditional metadata brokers. To connect this service with the
IDS data space ecosystem, we have developed the following
suite of components depicted on Figure 1:
• Data consumer: essential connector for the
discovery service that allows the browser to act as a data
consumer and collect metadata from the data space.
• Browser: the browser component queries
periodically the Metadata Broker through the data
consumer API to obtain the list of data providers and
their catalogs. For each of those catalogs, it requests
a list of ofered resources and all the available
metadata at every entity level including the providers’
data catalogs, ofered resources and artifacts. The
linked sample resources are also requested to obtain
their associated content-based metadata at artifact
level. The browser dumps the collected data on a
temporary database available for the search engine.
• Metadata storage: our approach utilizes an
ElasticSearch database to structure and optimize the data
collected by the browser, as well as to store the word
embeddings generated by the embedding
microservice (described below). This setup facilitates eficient
retrieval of relevant data, making it easier for data
consumers to find what they need.
• Embedding microservice: embeddings for
datasets are generated using a large language model
(LLM). For example, in Spanish, we use OpenAI’s
text-embedding-3-small model.
• Search microservice: this component provides the
algorithm to retrieve the data from the metadata
storage, re-rank, and sort the results from a given
query according to our previous work [
        <xref ref-type="bibr" rid="ref10 ref31 ref32">10, 31, 32</xref>
        ].
• API: used as a middleware to consume the search
microservice.
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Content-based Catalogs</title>
      <p>
        Our approach of content-based catalog for data spaces adds
metadata at both the resource and data artifact levels. At
the resource level, the metadata follows the IDS standard,
composed of multiple fields describing the entity that has
been added by the data provider. This resource-level
metadata is integrated with the content-based metadata obtained
automatically from the data and formatted using the Data
Package standard from Frictionless Data project [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] and
then combined together with the data sample to be stored
locally in the data provider at the artifact level (as shown
in Figure 1). This way, the content-based metadata will be
obtained together with the data sample when requesting
the sample resource.
      </p>
      <p>
        The Data Package standard is a simple container format
for describing a coherent collection of data in a single
package. It provides the basis for the convenient delivery,
installation, and management of datasets [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ]. The metadata
within the Data Package that applies to resource level is the
same as the one included in the sample resource at a higher
level of IDS entity hierarchy adapted to the Data Package
IDS CONNECTOR
Data Provider (detail)
      </p>
      <p>CATALOG
SAMPLE RESOURCE
(sample offer)
metadata</p>
      <p>RESOURCE</p>
      <p>(offer)
sample
link
metadata
REPRESENTATION</p>
      <p>REPRESENTATION</p>
      <p>ARTIFACT
Dpackage
metadata</p>
      <p>SAMPLE</p>
      <p>DATA
(local)</p>
      <p>ARTIFACT</p>
      <p>CSV DATA
(local or remote)
IDS CONNECTOR
Data Consumer
IDS CONNECTOR</p>
      <p>Data Provider
IDS CONNECTOR</p>
      <p>Data Provider
LLM</p>
      <p>Embedding
Microservice</p>
      <p>
        Search
Microservice
standard. Both of them are compatible with DCAT W3C
standard [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], ensuring interoperability.
      </p>
      <p>
        Furthermore, additional metadata on schema of the data
is included: the name of each field, the description of each
ifeld, the data type, and a sample value. This information
can be added for each field by using Frictionless [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
      </p>
      <p>
        Finally, to eficiently generate representative data
samples while maintaining computational eficiency, we employ
an approach that uses word embeddings to assess data
similarity according to our previous work [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Traditional
methods for generating semantic representations can be
resource-intensive, especially for large datasets. To mitigate
this, we hypothesize that a substantial portion of a dataset’s
content can be removed without significantly afecting its
semantic representation. To validate this hypothesis, we
evaluated diferent reduction techniques, demonstrating
that retaining only a small percentage of the original data
preserves its representational quality while reducing
computational costs. Three distinct methods were applied: (i)
random reduction, which randomly discards a fixed percentage
of data; (ii) duplicate removal, which eliminates redundant
entries to enhance uniqueness; and (iii) TF-IDF selection,
which prioritizes the most informative elements based on
their statistical significance. Specifically, among the three
proposed reduction strategies, TF-IDF demonstrated
superior performance overall compared to random reduction and
duplicate removal [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. These techniques not only reduce
data volume but also improve the performance of search
and retrieval systems, enabling more eficient and scalable
data discovery within the data space.
      </p>
      <p>In this paper, we propose a novel architecture for data spaces
designed to integrate content-based metadata generation
and discovery mechanisms, ensuring both compliance with
data sovereignty principles and enhanced efectiveness in
identifying relevant datasets within the data space
ecosystem. We addressed the challenges of improving data
discovery in data spaces by proposing an innovative approach
that combines descriptive metadata from data schema with
data sampling. While traditional metadata catalogs serve
as an essential foundation for organizing and describing
datasets, their limitations in conveying dataset-specific
details often impede efective data discovery. Our approach
mitigates these limitations by leveraging both, structural
metadata and data sampling techniques that balance data
value demonstration with the protection of data sovereignty.
By empowering data consumers to assess dataset relevance
while safeguarding data provider sovereignty, our approach
establishes a significant step toward operationalizing
efective data discovery in federated environments, such as data
spaces. On the basis on this research for improving data
discovery in data spaces, there are several promising
directions for future work. Adaptive sampling methodologies
that can dynamically adjust to the characteristics of diverse
datasets across data providers is an important next step. This
could involve leveraging advanced clustering algorithms,
domain-specific heuristics, or adaptive machine learning
models to optimize the sampling process. Also, ensuring
that the sampling techniques align with privacy
requirements is crucial. Future work could also explore integrating
privacy-preserving mechanisms to enhance trust and
compliance with legal and ethical standards. Finally, evaluation
of content-based catalogs in various real-world sectoral data
spaces, such as healthcare or environment, would provide
insights into its adaptability to specific domains.</p>
      <p>This work is part of the project TED2021-130890B-C21,
funded by MCIN/AEI/10.1 3039501100011033 and by the
European Union NextGenerationEU/PRTR.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions and Future Work</title>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments References</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>F.</given-names>
            <surname>Möller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Jussen</surname>
          </string-name>
          , V. Springer,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gieß</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. C.</given-names>
            <surname>Schweihof</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gelhaar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Guggenberger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Otto</surname>
          </string-name>
          ,
          <article-title>Industrial data ecosystems and data spaces</article-title>
          ,
          <source>Electronic Markets</source>
          <volume>34</volume>
          (
          <year>2024</year>
          )
          <fpage>1</fpage>
          -
          <lpage>17</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>R. C.</given-names>
            <surname>Fernandez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Subramaniam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Franklin</surname>
          </string-name>
          ,
          <article-title>Data market platforms: trading data assets to solve data problems</article-title>
          ,
          <source>Proceedings of the VLDB Endowment</source>
          <volume>13</volume>
          (
          <year>2020</year>
          )
          <fpage>1933</fpage>
          -
          <lpage>1947</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>P.</given-names>
            <surname>Hummel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Braun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Tretter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dabrock</surname>
          </string-name>
          ,
          <article-title>Data sovereignty: A review</article-title>
          ,
          <source>Big Data &amp; Society</source>
          <volume>8</volume>
          (
          <year>2021</year>
          )
          <fpage>2053951720982012</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Bhandari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Fariha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Price</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Vanterpool</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bowne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>McEvoy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Gadepally</surname>
          </string-name>
          , et al.,
          <article-title>Examples are all you need: Iterative data discovery by example in data lakes</article-title>
          ,
          <source>in: CIDR</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>C.</given-names>
            <surname>Gröger</surname>
          </string-name>
          ,
          <article-title>There is no AI without data</article-title>
          ,
          <source>Communications of the ACM</source>
          <volume>64</volume>
          (
          <year>2021</year>
          )
          <fpage>98</fpage>
          -
          <lpage>108</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>R.</given-names>
            <surname>Eichler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Gröger</surname>
          </string-name>
          , E. Hoos,
          <string-name>
            <given-names>H.</given-names>
            <surname>Schwarz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Mitschang</surname>
          </string-name>
          ,
          <article-title>Data shopping-how an enterprise data marketplace supports data democratization in companies</article-title>
          ,
          <source>in: International Conference on Advanced Information Systems Engineering</source>
          , Springer,
          <year>2022</year>
          , pp.
          <fpage>19</fpage>
          -
          <lpage>26</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>C.</given-names>
            <surname>Labadie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Legner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Eurich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fadler</surname>
          </string-name>
          ,
          <article-title>Fair enough? enhancing the usage of enterprise data with data catalogs</article-title>
          ,
          <source>in: 2020 IEEE 22nd Conference on Business Informatics (CBI)</source>
          , volume
          <volume>1</volume>
          , IEEE,
          <year>2020</year>
          , pp.
          <fpage>201</fpage>
          -
          <lpage>210</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>R.</given-names>
            <surname>Albertoni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Browning</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Cox</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. N.</given-names>
            <surname>GonzalezBeltran</surname>
          </string-name>
          , A. Perego,
          <string-name>
            <given-names>P.</given-names>
            <surname>Winstanley</surname>
          </string-name>
          ,
          <article-title>The w3c data catalog vocabulary, version 2: Rationale, design principles, and uptake</article-title>
          ,
          <source>Data Intelligence</source>
          <volume>6</volume>
          (
          <year>2024</year>
          )
          <fpage>457</fpage>
          -
          <lpage>487</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>P.</given-names>
            <surname>Křemen</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Nečasky`, Improving discoverability of open government data with rich metadata descriptions using semantic government vocabulary</article-title>
          ,
          <source>Journal of Web Semantics</source>
          <volume>55</volume>
          (
          <year>2019</year>
          )
          <fpage>1</fpage>
          -
          <lpage>20</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A.</given-names>
            <surname>Berenguer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Tomás</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mazón</surname>
          </string-name>
          ,
          <article-title>Tabular open government data search for data spaces based on word embeddings</article-title>
          , in: E.
          <string-name>
            <surname>Gallinucci</surname>
          </string-name>
          , L. Golab (Eds.),
          <source>Proceedings of the 25th International Workshop on Design, Optimization, Languages and Analytical Processing of Big Data (DOLAP) co-located with the 26th International Conference on Extending Database Technology and the 26th International Conference on Database Theory (EDBT/ICDT</source>
          <year>2023</year>
          ), Ioannina, Greece, March
          <volume>28</volume>
          ,
          <year>2023</year>
          , volume
          <volume>3369</volume>
          <source>of CEUR Workshop Proceedings, CEUR-WS.org</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>61</fpage>
          -
          <lpage>70</lpage>
          . URL: https://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>3369</volume>
          /paper6.pdf .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hauf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. M.</given-names>
            <surname>Comet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Moosmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Lange</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Chrysakis</surname>
          </string-name>
          , J. Theissen-Lipp,
          <article-title>FAIRness in dataspaces: The role of semantics for data management</article-title>
          , in: The Second International Workshop on Semantics in Dataspaces, co-located
          <source>with the Extended Semantic Web Conference</source>
          ,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>N.</given-names>
            <surname>Jahnke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Otto</surname>
          </string-name>
          ,
          <article-title>Data catalogs in the enterprise: applications and integration</article-title>
          ,
          <source>Datenbank-Spektrum</source>
          <volume>23</volume>
          (
          <year>2023</year>
          )
          <fpage>89</fpage>
          -
          <lpage>96</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Azcoitia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Laoutaris</surname>
          </string-name>
          ,
          <article-title>A survey of data marketplaces and their business models</article-title>
          ,
          <source>ACM SIGMOD Record</source>
          <volume>51</volume>
          (
          <year>2022</year>
          )
          <fpage>18</fpage>
          -
          <lpage>29</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>D.</given-names>
            <surname>Fowler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Barratt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Walsh</surname>
          </string-name>
          ,
          <article-title>Frictionless data: making research data quality visible</article-title>
          ,
          <source>International Journal of Digital Curation</source>
          <volume>12</volume>
          (
          <year>2017</year>
          )
          <fpage>274</fpage>
          -
          <lpage>285</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>A.</given-names>
            <surname>Berenguer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Tomás</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mazón</surname>
          </string-name>
          ,
          <article-title>Evaluating the impact of content deletion on tabular data similarity and retrieval using contextual word embeddings</article-title>
          , in: N.
          <string-name>
            <surname>Goharian</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Tonellotto</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>He</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Lipani</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>McDonald</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Macdonald</surname>
          </string-name>
          , I. Ounis (Eds.),
          <source>Advances in Information Retrieval - 46th European Conference on Information Retrieval</source>
          ,
          <string-name>
            <surname>ECIR</surname>
          </string-name>
          <year>2024</year>
          , Glasgow, UK, March
          <volume>24</volume>
          -28,
          <year>2024</year>
          , Proceedings,
          <string-name>
            <surname>Part</surname>
            <given-names>II</given-names>
          </string-name>
          , volume
          <volume>14609</volume>
          of Lecture Notes in Computer Science, Springer,
          <year>2024</year>
          , pp.
          <fpage>433</fpage>
          -
          <lpage>447</lpage>
          . URL: https://doi.org/10.1007/978-3-
          <fpage>031</fpage>
          -56060-6_
          <fpage>28</fpage>
          . doi:
          <volume>10</volume>
          . 1007/978-3-
          <fpage>031</fpage>
          -56060-6\_
          <fpage>28</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>A.</given-names>
            <surname>Berenguer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Alcaraz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Tomás</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mazón</surname>
          </string-name>
          ,
          <article-title>From research on data-intensive software to innovation in data spaces: A search service for tabular data</article-title>
          ,
          <source>IEEE Softw</source>
          .
          <volume>41</volume>
          (
          <year>2024</year>
          )
          <fpage>59</fpage>
          -
          <lpage>66</lpage>
          . URL: https://doi.org/10.1109/ MS.
          <year>2024</year>
          .
          <volume>3359333</volume>
          . doi:
          <volume>10</volume>
          .1109/MS.
          <year>2024</year>
          .
          <volume>3359333</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>B.</given-names>
            <surname>Otto</surname>
          </string-name>
          ,
          <article-title>A federated infrastructure for european data spaces</article-title>
          ,
          <source>Communications of the ACM</source>
          <volume>65</volume>
          (
          <year>2022</year>
          )
          <fpage>44</fpage>
          -
          <lpage>45</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>A.</given-names>
            <surname>Gieß</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Hupperz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Schoormann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Möller</surname>
          </string-name>
          ,
          <article-title>What does it take to connect? unveiling characteristics of data space connectors</article-title>
          , in: T. X.
          <string-name>
            <surname>Bui</surname>
          </string-name>
          (Ed.),
          <source>57th Hawaii International Conference on System Sciences, HICSS</source>
          <year>2024</year>
          ,
          <article-title>Hilton Hawaiian Village Waikiki Beach Resort</article-title>
          , Hawaii, USA, January 3-
          <issue>6</issue>
          ,
          <year>2024</year>
          , ScholarSpace,
          <year>2024</year>
          , pp.
          <fpage>4238</fpage>
          -
          <lpage>4247</lpage>
          . URL: https://hdl.handle. net/10125/106895.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>A.</given-names>
            <surname>Gieß</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Möller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Schoormann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Otto</surname>
          </string-name>
          ,
          <article-title>Design options for data spaces</article-title>
          ,
          <source>in: Thirty-first European Conference on Information Systems (ECIS</source>
          <year>2023</year>
          ),
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>M.</given-names>
            <surname>Bacco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kocian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Chessa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Crivello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Barsocchi</surname>
          </string-name>
          ,
          <article-title>What are data spaces? systematic survey and future outlook, Data in Brief 57 (</article-title>
          <year>2024</year>
          )
          <fpage>110969</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <fpage>i4Trust</fpage>
          ,
          <article-title>B2B Data Sharing Playbook - the i4Trust approach to Data Sharing,</article-title>
          <year>2021</year>
          . URL: https://i4trust.org/wp-content/uploads/i4Trust_ DataSharingPlaybook.pdf .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22] iSHARE Foundation, iSHARE - Trust
          <source>Framework for Data Spaces</source>
          ,
          <year>2021</year>
          . URL: https://ishare.eu/.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>FIWARE</given-names>
            <surname>Foundation</surname>
          </string-name>
          , FIWARE - Open APIs for Open Minds,
          <year>2024</year>
          . URL: https://www.fiware.org.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Gaia-X European</surname>
          </string-name>
          <article-title>Association for Data and Cloud AISBL</article-title>
          ,
          <string-name>
            <surname>Gaia-X Architecture Document</surname>
          </string-name>
          ,
          <year>2024</year>
          . URL: https://docs.gaia-x.eu/technical-committee/ architecture-document/latest/.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>International</given-names>
            <surname>Data Spaces Association</surname>
          </string-name>
          ,
          <source>International Data Spaces Association</source>
          ,
          <year>2022</year>
          . URL: https:// internationaldataspaces.org/.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>International</given-names>
            <surname>Data Spaces Association</surname>
          </string-name>
          ,
          <source>IDS Reference Architecture Model, Version 4.0</source>
          ,
          <year>2022</year>
          . URL: https://docs.internationaldataspaces.org/ ids-knowledgebase/ids-ram-4/introduction/1_1_ goals_
          <article-title>of_the_international_data_spaces.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>International</given-names>
            <surname>Data Spaces Association</surname>
          </string-name>
          , Position Paper - GAIA
          <string-name>
            <surname>-X and</surname>
            <given-names>IDS</given-names>
          </string-name>
          ,
          <year>2021</year>
          . URL: https://internationaldataspaces. org/wp-content/uploads/dlm_uploads/ IDSA-Position-
          <article-title>Paper-</article-title>
          <string-name>
            <surname>GAIA-X-</surname>
          </string-name>
          and
          <string-name>
            <surname>-IDS</surname>
          </string-name>
          .pdf .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <surname>Fraunhofer-Institut für</surname>
          </string-name>
          Software- und
          <string-name>
            <surname>Systemtechnik</surname>
            <given-names>ISST</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fraunhofer - Software</surname>
          </string-name>
          (-
          <source>architecture)</source>
          ,
          <year>2024</year>
          . URL: https://www.dataspaces.fraunhofer.de/en/software. html.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>International</given-names>
            <surname>Data Spaces Association</surname>
          </string-name>
          , Idsa github repository,
          <year>2024</year>
          . URL: https://github.com/ International-Data-Spaces-Association.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>International</given-names>
            <surname>Data Spaces Association</surname>
          </string-name>
          ,
          <source>Idsa data model documentation</source>
          ,
          <year>2024</year>
          . URL: https: //international-data
          <article-title>-spaces-association</article-title>
          .github.io/ DataspaceConnector/Documentation/v5/DataModel.
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>J.</given-names>
            <surname>Pilaluisa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Tomás</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Navarro-Colorado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mazón</surname>
          </string-name>
          ,
          <article-title>Contextual word embeddings for tabular data search and integration</article-title>
          ,
          <source>Neural Comput. Appl</source>
          .
          <volume>35</volume>
          (
          <year>2023</year>
          )
          <fpage>9319</fpage>
          -
          <lpage>9333</lpage>
          . URL: https://doi.org/10.1007/s00521-022-08066-8. doi:
          <volume>10</volume>
          .1007/S00521-022-08066-8.
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>A.</given-names>
            <surname>Berenguer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mazón</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Tomás</surname>
          </string-name>
          ,
          <article-title>Word embeddings for retrieving tabular data from research publications, Mach</article-title>
          . Learn.
          <volume>113</volume>
          (
          <year>2024</year>
          )
          <fpage>2227</fpage>
          -
          <lpage>2248</lpage>
          . URL: https: //doi.org/10.1007/s10994-023-06472-0. doi:
          <volume>10</volume>
          .1007/ S10994-023-06472-0.
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>Data</given-names>
            <surname>Package</surname>
          </string-name>
          Working Group,
          <source>Data Package Standard</source>
          ,
          <year>2024</year>
          . URL: https://datapackage.org/standard/ data-package/.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>