<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>X. Liu);</journal-title>
      </journal-title-group>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>MDMapper: A Framework for Aligning Master Data Models using Ontology Matching Techniques</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Xianhao Liu</string-name>
          <email>xianliu@dtu.dk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jesper Grode</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michael R. Hansen</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Ontology Matching, Master Data Management, Data Exchange</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Stibo Systems A/S</institution>
          ,
          <addr-line>Axel Kiers Vej 11, 8270 Højbjerg</addr-line>
          ,
          <country country="DK">Denmark</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Technical University of Denmark</institution>
          ,
          <addr-line>2800 Kgs-Lyngby</addr-line>
          ,
          <country country="DK">Denmark</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>This paper introduces a matching framework tailored for master data model matching, incorporating techniques from the field of ontology matching. We present a new quantitative approach for heuristic similarity estimation between hierarchical data structures, which involves heterogeneous data. We also introduce a relation-based navigation technique and an availability management method based on restrictions that support eficient and progressive matching processes. This integration of ontology matching techniques into master data model matching not only improves alignment consistency and quality, but also facilitates more automatic data exchange solutions. The experiments on OAEI Anatomy and Conference tracks indicate that our approach may be competitive, while an experiment on industrial classification standards shows that our approach performs significantly better than the considered baseline approaches.</p>
      </abstract>
      <kwd-group>
        <kwd>Matching Techniques</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>The digital information supply chain refers to the comprehensive process through which digital
data is generated, processed, stored, transmitted, and ultimately utilized. Unlike the physical
supply chain, which deals with tangible goods, the digital information supply chain deals with
intangible data flows, requiring robust infrastructure and sophisticated protocols to ensure
eficiency and security. The key actors in the digital information supply chain are data producers
such as manufacturers, distributors and retail chains, and data consumers such as end users,
consumers, and online shoppers. In addition, regulatory bodies often oversee the flow of
information to ensure compliance with legal, ethical, and security standards, or simply require
businesses to adhere to regulatory compliance standards.</p>
      <p>Actors within the digital information supply chain often operate a so-called Master Data
Management (MDM) platform as a central hub for data exchange. However, they predominantly
use diverse data models to categorize their products. To ensure accurate, consistent, and efective
data exchange, it is essential to align data between the actors.</p>
      <p>
        Master data models are, typically, hierarchical concept classifications [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] where attributes may
be attached to the concepts. A product in an MDM system is an instance of a particular concept,
having data fields corresponding to the attribute descriptions of that concept. Master data models
can be considered special ontologies, having a simple hierarchical ontological structure, with
rich descriptions of attributes comprising types, units etc. Consequently, ontology matching
becomes a centerpiece for these actors to exchange data.
      </p>
      <p>
        A scenario: Ontology Matching in MDM: Consider three businesses,  ,  and  , that
exchange data described by diferent industrial classification standards: ETIM [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and ECLASS
[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].  is a manufacturer, sending data to a distributor  , who in turn sends data to a retailer  .
 ,  and  each operate their own business systems and use their own ontologies to manage
the data representing the physical goods they exchange. We call this “product data” in the
following:
•  classifies and manages all product data according to the ETIM classification standard.
•  uses ECLASS to classify and manage their product data.
      </p>
      <p>•  uses a custom made classification to classify and manage the product data
For this example, let us assume that  manufactures coaxial cables,  distributes electrical and
electronic products, including cables, and  is an online store that sells audio gear to musicians.
Table 1 is the partial data fields (attributes) of a coaxial cable in diferent ontologies.</p>
      <p>What can be noticed is that the data itself, for instance the color of the cable is the same (as it
is one and the same physical cable), but the name of the data fields are diferent across the three
representations.</p>
      <p>In this simple information supply chain,  →  →  , we already face two problems:
• Either  needs to have information about how  expects to receive data and map the
data from its own representation to that of  , or  will have to receive  ’s representation
of data and then map to its own ontology.
•  is faced with the same problem as  - either map from ECLASS to that of the recipient
 , which implies that  needs to know about  ’s representation of data, or  needs to
map from  ’s representation to its own.</p>
      <p>In addition to this, even before being able to map the data fields, the correct
category/classification must be determined for each product sent from  and received by  (and from  ,
received by  ). A specific data record of the coax cable product may be classified as Table 2.
Notice that the custom hierarchy for  focuses more on the selling aspects of a product.</p>
      <p>In a typical information supply chain, these two problems multiply;  will receive data from
many data providers, and  needs to send data to many data consumers, and all may operate
diferent data standards.</p>
      <p>We present a framework for ontology matching tailored to MDM systems, which typically are
simple classification hierarchies with rich attributes of concepts with data types and units. In
Section 2, we model typical MDM matching tasks, including key properties that narrow the
matching scope and enhance eficiency and quality. We also introduce a heuristic similarity
measure that incorporates descendants and attributes. The framework is detailed in Section
3. Section 4 covers experiments using OAEI tracks, Anatomy and Conference, as well as the
industrial classification standards, ETIM and ECLASS.</p>
      <p>
        Related Work: Eine et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] presents the feasibility of ontology-based big data management,
with applications in data integration using ontology alignment. Ramzy et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] present a
methodology for master data management based on knowledge graphs, which relies on the
establishment of a knowledge graph (KG) layer to build a common understanding of key business
entities and semantic mappings from and to the original data sources. These works demonstrate
the potential of applying ontology technologies in master data management.
      </p>
      <p>
        Euzenat et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] ofer comprehensive coverage of ontology matching within a uniform
framework, and OAEI (Ontology Alignment Evaluation Initiative) [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] provides a global benchmark
for evaluating schema or ontology matching methods. These works formalized the uniform
knowledge and evaluation benchmarks utilized in this paper.
      </p>
      <p>
        LogMap [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] is one of the leading systems in the Ontology Alignment Evaluation Initiative
(OAEI). It performs matching based on logic consistency and output coherence alignment.
LogMap uses a Horn propositional logic representation of the extended hierarchy of each
ontology with all existing mappings, and applys Dowling-Gallier algorithm [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] for unsatisfiability
checking. Although its approach is eficient in checking the consistency of mappings, it does
not directly identify conflicting mappings. Additionally, since the expansion of mappings is
based on previously discovered mappings, some true mappings may remain undiscovered if
they were not covered within the expansion range of existing mappings.
      </p>
      <p>We use a matching space maintaining a pool of correspondences that are not in conflict
with the current alignments. That pool shrinks when new correspondences are added. The
manner in which this pool is maintained gives fewer false negatives compared to LogMap for
the considered data sets.</p>
      <p>
        Hansen et al. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] present an approach using formal methods to ensure consistent alignment
for simple ontologies in the digital information supply chain. Their work provides models,
methods, and tools for ontology matching that can guarantee the consistency of alignments,
particularly in the context of master data management systems. Furthermore, they proposed
the potential of guiding the search for consistent correspondences while eliminating irrelevant
ones, which inspired the conflict-based restriction approach proposed in this paper.
      </p>
      <p>
        COMA/COMA++ [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] presents approaches for the flexible combination of similarity
measurements, demonstrating that strategically combining diferent similarity measures can lead to
improved performance. Transformer-based models [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] like BERT [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] provide the ability to
capture semantic similarity more accurately from text context and have been used more widely
in the field of ontology matching. Furthermore, OLaLa [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] utilizes large language models
(LLMs) as a similarity measure and achieves competitive results on OAEI benchmarks. These
works inspired the aggregate similarity measurement in this paper.
      </p>
      <p>
        Background knowledge is another important factor to improve the accuracy of similarity
estimation and has been widely used [
        <xref ref-type="bibr" rid="ref14 ref15 ref8">8, 14, 15</xref>
        ]. Portisch et al. [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] comprehensively reviewed
background knowledge in ontology matching from the perspective of methods and applications.
Appropriate background knowledge can enhance the detectability of matches between
domainspecific terms in MDM, which would be a key point to further enhance our approach.
      </p>
    </sec>
    <sec id="sec-3">
      <title>2. Modelling for Aligning MDM</title>
      <p>We now present the basic relations we shall use, the definition of an MDM ontology  and the
basic properties of  , and some fundamental properties of consistent alignment.
2.1. Relating concepts
A correspondence is defined as a triple ( 1,  2,  ) , where  1 and  2 are concepts, and  represents
one of the following relations:
• isEqual (=): Indicates that two concepts are equivalent.
• lessEqual (≤): Indicates that the first concept is a specialization of the second one.
• largerEqual (≥): Indicates that the first concept is a generalization of the second one.
• disjoint (∅): Indicates that two concepts are disjoint (or exclusive).
• partialOverlap (≠∪∩): Indicates that two concepts overlap but are not equal, and neither is
a subset of the other.
2.2. MDM ontology
An MDM ontology is a hierarchical classification of concepts, that is, a tree  , where the nodes
are concepts. Furthermore, the concepts in  satisfy certain properties (making  a classification).</p>
      <p>Let  1 and  2 be concepts (nodes) of  :
• If  1 is a child of  2, then ( 1,  2, ≤).</p>
      <p>• If  1 and  2 are siblings in  , then ( 1,  2, ∅).</p>
      <p>That is, a child is a specialization of its parent and siblings are mutually exclusive concepts.</p>
      <p>The following properties are consequences of above properties:
• If  1 is a descendant of  2, then ( 1,  2, ≤). (This property can also be expressed in terms of
ancestor and ≥.)
• If  1 and  2 are diferent concepts in  and neither is a descendant of the other,
then ( 1,  2, ∅).</p>
      <p>Furthermore, we have</p>
      <p>• If  1 and  2 are diferent concepts in  , then ¬( 1,  2, =).</p>
      <p>That is,  cannot contain two diferent equivalent concepts.
2.3. Conflict-based Restriction
When a set of correspondences  is established between two MDM ontologies   and   , it is
easy to reach an inconsistent situation. We shall now formulate some consistency constraints
on  . These constraints will later be exploited in order to prune the space that is explored when
searching for new correspondences.</p>
      <p>A correspondence between   and   is a relation (  ,   ,  ) , where   is a concept in   and   is
a concept in   .</p>
      <p>Let  be the alignment that contains only correspondences with  relations.  is conflict
free if the following restrictions are satisfied for every correspondence (  ,   , =) ∈  :
1. for every   ′ in   , where   ′ is not   : (  ′,   , =) ∉  ,
2. for every   ′ in   , where   ′ is not   : (  ,   ′, =) ∉  ,
3. for every ancestor   ′ of   in   and for every descendant   ′ of   in   : (  ′,   ′, =) ∉  ,
4. for every descendant   ′ of   in   and for every ancestor   ′ of   in   : (  ′,   ′, =) ∉  ,
5. for every linear relative   ′ of   in   and for every non-linear relative   ′ of   in   :
(  ′,   ′, =) ∉  , and
6. for every non-linear relative   ′ of   in   and for every linear relative   ′ of   in   :
(  ′,   ′, =) ∉  ,
where linear relatives of a concept  in a hierarchy  are the ancestors and descendants of  ,
while non-linear relatives of a concept  in a hierarchy  are all other concepts in  that are
neither ancestors nor descendants of  .
2.4. Heuristic Similarity Between Hierarchies
We propose a heuristic approach to estimate the overall similarity between two hierarchies by
evaluating potential correspondences based on their similarity matrix between entities, where
an entity can be either a concept or an attribute.</p>
      <p>- Let   denote the set of source entities.
- Let   denote the set of target entities.</p>
      <p>- (,  ′) presents the similarity score between a source entity  ∈   and a target entity  ′ ∈   ,
the value falls within the range of 0 to 1.</p>
      <p>Using a specified threshold  , we filter out entities whose highest similarity score to their
counterparts exceeds this threshold, calling them prominent entities. The set of prominent
entities of source and target are denoted as   and   :
  = { ∈   ∣ m′∈a x (,  ′) &gt; },
  = { ′ ∈   ∣ max (,  ′) &gt; }
∈ 
(1)
  is the set of source entities that are considered compatible with at least one target entity.
Similarly,   is the set of target entities that are considered to be compatible with at least one
source entity. Both identify key entities in their respective hierarchies as potential candidates
for alignment.</p>
      <p>
        The highest similarity scores associated with these prominent entities are referred to as
prominent scores and are collectively denoted by  , with the type  ∶ (  ∪   ) → [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ].
      </p>
      <p>For each entity  in   or   , the prominent score  [] is given by:</p>
      <p>max {(,  ′) ∣ (,  ′) &gt; } if  ∈   ,
 [] = { ′∈  (2)
max {( ′, ) ∣ ( ′, ) &gt; } if  ∈   .</p>
      <p>′∈</p>
      <p>The sequence  thus consists of these maximum similarity scores corresponding to the entities
in   and   . Each entity in   and   has a corresponding value in  , therefore, | | = |  | + |  |.</p>
      <p>The ratio  of the size of prominent entities to the total number of entities is used to represent
the scale of potential mappings:
  = |  | + |  | (3)</p>
      <p>|  | + |  |
The mean value of the prominent scores represents the quality of these potential mappings:
 =̄
1</p>
      <p>∑  []
| | ∈(  ∪  )
(4)</p>
      <p>The final heuristic value ℎ is determined by combining the ratio  and the mean prominent
score  :̄
ℎ =   ×  =̄ 1 ∑  [] (5)</p>
      <p>|  | + |  | ∈(  ∪  )</p>
      <p>The heuristic score ℎ, defined as the product of the   of prominent entities and the mean
prominent score  ,̄ ranges from 0 to 1 due to their individual constraints.   represents the
proportion of prominent entities (0 to 1), while  ̄ is the average of normalized similarity scores
above a threshold (0 to 1). Thus, ℎ provides a normalized score indicating the extent and quality
of potential correspondence, from no correspondence (0) to perfect correspondence (1). Table 5
demonstrates the impact of this method on enhancing the classification-level similarity matrix
with attribute-level similarity.</p>
    </sec>
    <sec id="sec-4">
      <title>3. Framework</title>
      <p>We propose MDMapper, which is a framework specifically designed for MDM matching tasks.
The architecture of MDMapper is shown in Figure 1.</p>
      <p>In the Pre-processing phase, we initiate with Data Extraction, which parses both source and
target ontologies into internal representations (   ℎ   and     ). The next
step involves the Similarity Measure module, which employs a variety of methods to estimate
the pairwise similarity of source-target concepts.
3.1. Similarity Measure
We employ similarity measurement approaches to estimate the similarity scores of pairwise
source-target concepts to capture diferent characteristics. The final similarity matrix is a linear
combination of multiple similarity matrices.</p>
      <p>
        We use the iSUB [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] algorithm, which emphasizes the importance of shared contiguous
substrings, and the pre-trained language model, Sentence Transformer [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], which demonstrates
superior performance in capturing semantic meaning.
      </p>
      <p>Additionally, we enhance our similarity matrix with our heuristic approach (see Section
2.4) via bottom-up similarity propagation, efectively augmenting the direct comparisons with
inherited similarities across the data hierarchy.
(1) This part comprises two hierarchies (  and   ) and a correspondence (
(2) For this part, the table on the left represents the updated matching space according to the
correspondence (

,   , =); some cells are blocked with status -1 according to the restrictions in Section 2.3; blue
cells are blocked according to Restriction 1 and 2, yellow cells are blocked according to Restriction 3
and 4, and green cells are blocked according to Restriction 5 and 6. The tables on the right display the
narrowed matching spaces for further matching process
3.2. Matching
Throughout the matching process, the framework maintains a Matching Space to manage
discovered correspondences and prevent potential conflicting matches, thereby preserving the
consistency of the alignment and enhancing the eficiency of matching.</p>
      <p>The Matching Space is initialized as a matrix of size |  | × |  |, with all cells set to open (0)
status. For each newly discovered correspondence, the corresponding cell in the Matching Space
is marked as aligned (1). Restrictions are added by blocking cells, which are set to blocked (-1)
status. These cells correspond to potential correspondences that conflict with the accepted ones,
as described in Section 2.3. It is illustrated with a simple case in Figure 2 how the matching
space maintains alignment and restrictions.</p>
      <p>Initial Matching:</p>
      <p>During an initial matching phase, high-confidence correspondences, known
as anchors, are identified. They are added to the matching space and will never be removed.</p>
      <p>
        An entity matching framework CollaborEM [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] utilizes Iterative KG Completion [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] to
generate positive labels for their self-supervised Entity Matching task. We refined this approach
for anchor identification, which gives stricter constraints that ensure higher-quality anchors
compared to merely applying a threshold on similarity scores.
      </p>
      <p>Our approach employs a significantly mutually most similar rule to identify anchors based on
the similarity matrix. Each identified anchor (, , =) ,  ∈   ,  ∈   , must satisfy:
1. Their similarity score must exceed the given threshold.
2. They must be mutually the most similar to each other.</p>
      <p>3. There should be a margin between their similarity and the second most similar pair.
Matching phase: The matching phase performs a local matching for each candidate in
the candidate pool, which is initially populated with anchors. Candidates at lower levels are
prioritized for extraction from the pool. For each candidate (, ,  ) , the first step is to decide the
pair of concepts ( ′,  ′) to explore in the following local matching.</p>
      <p>The navigation function determines ( ′,  ′) based on the hierarchical positions of (, ) and
their relation  , as follows:</p>
      <p>⎧(parent(), parent())
( ′,  ′) = (parent(), )
⎨
⎩(, parent())
if  ∈ {" = ", " ≠∩∪ ", "∅"} and parent(), parent() exist,
if  = " ≤ " and parent() exists,
if  = " ≥ " and parent() exists.</p>
      <p>The basic ideology of the function design: On the one hand, valid correspondences increase the
likelihood of relations existing between their hierarchical neighbors; ‘=’, ‘≤’, and ‘≥’ suggest
possible matches among parents or siblings. On the other hand, ‘≠∪∩’ and ‘∅’ indicate the need
to expand the matching scope.</p>
      <p>Since only correspondences with an ‘=’ relation can be identified via similarity score, for a
given (, ) with unknown relation, we determine its relation by evaluating the overlap of the
descendants of them, using the following criteria:
(6)
(7)
 =</p>
      <p>=∶
⎧
⎪ ≤∶</p>
      <p>≥∶
⎨⎪ ∅ ∶</p>
      <p>∪
⎩ ≠∩∶
(∀ ∈ (), ∃ ∈ () ∶  = 
(∀ ∈ (), ∃ ∈ () ∶  = 
(∀ ∈ (), ∃ ∈ () ∶  = 
(∀ ∈ (), ∀ ∈ () ∶  ≠ 
Otherwise
) ∧ (∀ ∈ (), ∃ ∈ () ∶  = 
) ∧ (∃ ∈ (), ∀ ∈ () ∶  ≠ 
) ∧ (∃ ∈ (), ∀ ∈ () ∶  ≠ 
)
)
)
)
 denotes all descendant concepts of the given concept.</p>
      <p>Local matching: This phase identifies new correspondences within a narrower matching
scope tailored to the suggested pair of concepts ( ′,  ′). For the given pair of concepts ( ′,  ′),
we identify the correspondences between their descendants and the concepts themselves. A
local similarity matrix is constructed, related to the sub-hierarchies of  ′ and  ′, which is then
ifltered through the matching space to delineate the matching scope.</p>
      <p>Furthermore, depending on the relations between their linear relatives, we apply varying
thresholds to improve the matching quality. For example, for a pair of concepts ( ″,  ″) in the
local matching scope, if ( ( ″),  ( ″), =) exists, a lower threshold should be applied
for matching  ″ and  ″, since the equivalence in their parents indicates a higher likelihood of
an ‘=’ relation between them.</p>
      <p>Local matching may find new correspondences that are added to the candidate pool.
Furthermore, the matching space is updated accordingly. This iterative process continues until no
further correspondences can be identified.</p>
    </sec>
    <sec id="sec-5">
      <title>4. Experiments</title>
      <p>Although MDMapper is designed primarily to address MDM matching tasks, it is also capable of
handling simple ontology matching tasks. To compare it with other ontology matching systems,
we apply our framework to the OAEI Anatomy and Conference tracks. Furthermore, to evaluate
our performance in solving real-world MDM matching problem, we applied our framework
to match ETIM and ECLASS in various versions: with and without attributes. Both versions
showed outperforming results compared to the baseline approaches. All experiments were
conducted on a MacBook with an Apple M2 Max chip and 64GB of RAM.
4.1. OAEI Tracks
We applied our approach to the OAEI Anatomy and Conference tracks. Table 3 shows that our
approach achieves an F1 measure of 0.905, which ranks 4th based on the anatomy track (2023).
Table 4 shows that our F1 measure is 0.65 on the conference track (2023), which ranks 3rd. Our
approach is eficient, with runtimes of 58 seconds on the Anatomy track and 97 seconds on the
Conference track. In addition, our approach ensures coherent alignment.
4.2. Industry Classification Standards Matching
Data Description: ETIM-7 has 4,878 classifications across 3 layers, while ECLASS-11 includes
86,468 classifications in 6 layers. ETIM Germany is working on aligning ETIM to ECLASS, but
the current reference alignment is incomplete. It includes 2,762 ETIM categories mapped to</p>
      <p>The attribute features we use include name, datatype, and unit.
2,435 ECLASS categories, totaling 2,875 mappings. All correspondences are between the ETIM
leaf nodes and the parent nodes of the ECLASS leaf nodes.</p>
      <p>We extracted subtrees from both ETIM and ECLASS on the basis of these mappings, preserving
root-to-leaf paths. The resulting Sub-ETIM subtree consists of 2,873 categories in 3 layers, with
10,745 attributes aligned with leaf nodes. The Sub-ECLASS subtree contains 5,561 categories in
6 layers, with 8,928 attributes aligned with leaf nodes.</p>
      <p>Given the incomplete alignment, we evaluated the matching results using the filtered
correspondences within specific layers of the classification hierarchies. While valid correspondences
may exist beyond this scope, they cannot be assessed without complete reference mappings.
Analysis: To benchmark our approach, we used StringEquiv and Optimal Threshold (the best
achievable results based on the threshold applied to the similarity matrix) for matching tasks
between Sub-ETIM and Sub-ECLASS. These techniques are among the most commonly used in
current MDM matching solutions. Since the ETIM and ECLASS datasets are not available in
OWL or RDF format, we were unable to directly apply other Ontology Matching Systems.</p>
      <p>We conducted two experimental variants with our approach (see Table 5): one excludes
attributes, while the other incorporates them into the similarity matrix using the heuristic
method described in Section 2.4. Our approach, both with and without attributes, significantly
outperforms the baseline methods. The improvement achieved by incorporating attributes
demonstrates the efectiveness of using attribute features in category similarity measurements.
Excluding pre-processing, the matching process takes 5.73 seconds.</p>
    </sec>
    <sec id="sec-6">
      <title>5. Discussion</title>
      <p>We have introduced a framework for ontology matching geared towards the Master Data
Management domain. In addition to known ontology matching techniques, we developed a
relation-based navigation approach that narrows the matching scope using existing
correspondences. Furthermore, a heuristic approach is proposed to estimate the overall similarity between
subtrees, integrating heterogeneous entities into a unified measure.</p>
      <p>Experiments on data from OAEI tracks indicate that our approach may be competitive, while
the experiment on industrial classification standards shows outperforming results in solving
real-world MDM matching problems compared to selected baseline techniques.</p>
      <p>The framework is in a prototype stage and under development for further enhancements.
Future work includes integrating domain-specific external knowledge resources to refine
matching quality, incorporating global optimization and callback mechanisms to resolve conflicts, and
developing a user interface to select diferent conflict-based alignment versions.
Acknowledgement
This study was funded by Innovation Fund Denmark (grant number 2050-00004B).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M. R.</given-names>
            <surname>Hansen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Grode</surname>
          </string-name>
          ,
          <article-title>Consistent alignments for simple ontologies in the digital information supply chain</article-title>
          ,
          <source>in: The Practice of Formal Methods</source>
          , Springer,
          <year>2024</year>
          , pp.
          <fpage>175</fpage>
          -
          <lpage>194</lpage>
          . Chapter 9.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>ETIM</given-names>
            <surname>International</surname>
          </string-name>
          ,
          <source>Etim classification model version 7</source>
          .0,
          <year>2019</year>
          . URL: http://oaei. ontologymatching.org/, accessed:
          <fpage>2024</fpage>
          -07-17.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3] eCl@ss e.V.,
          <source>ecl@ss standard version 11.0</source>
          ,
          <year>2021</year>
          . URL: https://www.eclass.eu/, accessed:
          <fpage>2024</fpage>
          -07-17.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>B.</given-names>
            <surname>Eine</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jurisch</surname>
          </string-name>
          , W. Quint,
          <article-title>Ontology-based big data management</article-title>
          ,
          <source>Systems</source>
          <volume>5</volume>
          (
          <year>2017</year>
          )
          <fpage>45</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>N.</given-names>
            <surname>Ramzy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Durst</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Schreiber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Auer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chamanara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Ehm</surname>
          </string-name>
          ,
          <article-title>Knowgraph-mdm: A methodology for knowledge-graph-based master data management</article-title>
          ,
          <source>in: 2022 IEEE 24th Conference on Business Informatics (CBI)</source>
          , volume
          <volume>2</volume>
          , IEEE,
          <year>2022</year>
          , pp.
          <fpage>9</fpage>
          -
          <lpage>16</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Euzenat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Shvaiko</surname>
          </string-name>
          , et al.,
          <article-title>Ontology matching</article-title>
          , volume
          <volume>18</volume>
          , Springer,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>OAEI</surname>
          </string-name>
          ,
          <article-title>Ontology alignment evaluation initiative, 2023</article-title>
          . URL: http://oaei.ontologymatching. org/, accessed:
          <fpage>2024</fpage>
          -07-17.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>E.</given-names>
            <surname>Jiménez-Ruiz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. Cuenca</given-names>
            <surname>Grau</surname>
          </string-name>
          ,
          <article-title>Logmap: Logic-based and scalable ontology matching</article-title>
          ,
          <source>in: The Semantic Web-ISWC</source>
          <year>2011</year>
          : 10th International Semantic Web Conference, Bonn, Germany,
          <source>October 23-27</source>
          ,
          <year>2011</year>
          , Proceedings,
          <source>Part I 10</source>
          , Springer,
          <year>2011</year>
          , pp.
          <fpage>273</fpage>
          -
          <lpage>288</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>W. F.</given-names>
            <surname>Dowling</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. H.</given-names>
            <surname>Gallier</surname>
          </string-name>
          ,
          <article-title>Linear-time algorithms for testing the satisfiability of propositional horn formulae</article-title>
          ,
          <source>The Journal of Logic Programming</source>
          <volume>1</volume>
          (
          <year>1984</year>
          )
          <fpage>267</fpage>
          -
          <lpage>284</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>H.-H. Do</surname>
          </string-name>
          , E. Rahm,
          <article-title>Coma-a system for flexible combination of schema matching approaches</article-title>
          ,
          <source>in: VLDB'02: Proceedings of the 28th International Conference on Very Large Databases, Elsevier</source>
          ,
          <year>2002</year>
          , pp.
          <fpage>610</fpage>
          -
          <lpage>621</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaswani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Parmar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Gomez</surname>
          </string-name>
          , Ł. Kaiser,
          <string-name>
            <surname>I. Polosukhin</surname>
          </string-name>
          ,
          <article-title>Attention is all you need</article-title>
          ,
          <source>Advances in neural information processing systems</source>
          <volume>30</volume>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , Bert:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          , arXiv preprint arXiv:
          <year>1810</year>
          .
          <volume>04805</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>S.</given-names>
            <surname>Hertling</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Paulheim</surname>
          </string-name>
          , Olala:
          <article-title>Ontology matching with large language models</article-title>
          ,
          <source>in: Proceedings of the 12th Knowledge Capture Conference</source>
          <year>2023</year>
          ,
          <year>2023</year>
          , pp.
          <fpage>131</fpage>
          -
          <lpage>139</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>D.</given-names>
            <surname>Faria</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Santos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. S.</given-names>
            <surname>Balasubramani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. C.</given-names>
            <surname>Silva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. M.</given-names>
            <surname>Couto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Pesquita</surname>
          </string-name>
          , Agreementmakerlight, Semantic
          <string-name>
            <surname>Web</surname>
          </string-name>
          (
          <year>2013</year>
          )
          <fpage>1</fpage>
          -
          <lpage>13</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>D.</given-names>
            <surname>Ngo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Bellahsene</surname>
          </string-name>
          , Overview of yam++
          <article-title>-(not) yet another matcher for ontology alignment task</article-title>
          ,
          <source>Journal of Web Semantics</source>
          <volume>41</volume>
          (
          <year>2016</year>
          )
          <fpage>30</fpage>
          -
          <lpage>49</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>J.</given-names>
            <surname>Portisch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hladik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Paulheim</surname>
          </string-name>
          ,
          <article-title>Background knowledge in ontology matching: A survey, Semantic Web (</article-title>
          <year>2022</year>
          )
          <fpage>1</fpage>
          -
          <lpage>55</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>G.</given-names>
            <surname>Stoilos</surname>
          </string-name>
          , G. Stamou,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kollias</surname>
          </string-name>
          ,
          <article-title>A string metric for ontology alignment</article-title>
          ,
          <source>in: The Semantic Web-ISWC</source>
          <year>2005</year>
          : 4th International Semantic Web Conference,
          <string-name>
            <surname>ISWC</surname>
          </string-name>
          <year>2005</year>
          , Galway, Ireland, November 6-
          <issue>10</issue>
          ,
          <year>2005</year>
          . Proceedings 4, Springer,
          <year>2005</year>
          , pp.
          <fpage>624</fpage>
          -
          <lpage>637</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>N.</given-names>
            <surname>Reimers</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Gurevych</surname>
          </string-name>
          ,
          <article-title>Sentence-bert: Sentence embeddings using siamese bertnetworks</article-title>
          ,
          <source>in: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, Association for Computational Linguistics</source>
          ,
          <year>2019</year>
          . URL: https: //arxiv.org/abs/
          <year>1908</year>
          .10084.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>C.</given-names>
            <surname>Ge</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <article-title>Collaborem: A self-supervised entity matching framework using multi-features collaboration</article-title>
          ,
          <source>IEEE Transactions on Knowledge and Data Engineering</source>
          <volume>35</volume>
          (
          <year>2021</year>
          )
          <fpage>12139</fpage>
          -
          <lpage>12152</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>W.</given-names>
            <surname>Zeng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <article-title>Degree-aware alignment for entities in tail</article-title>
          ,
          <source>in: Proceedings of the 43rd international ACM SIGIR conference on research and development in information retrieval</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>811</fpage>
          -
          <lpage>820</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>