<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>GenOM: Ontology Matching with Description Generation and Large Language Model</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yiping Song</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jiaoyan Chen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Renate A. Schmidt</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, The University of Manchester</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2025</year>
      </pub-date>
      <fpage>0000</fpage>
      <lpage>0003</lpage>
      <abstract>
        <p>Ontology matching (OM) plays an essential role in enabling semantic interoperability and integration across heterogeneous knowledge sources, particularly in the biomedical domain which contains numerous complex concepts related to diseases and pharmaceuticals. This paper introduces GenOM, a large language model (LLM)-based ontology alignment framework, which enriches the semantic representations of ontology concepts via generating textual definitions, retrieves alignment candidates with an embedding model, and incorporates exact matching-based tools to improve precision. Extensive experiments conducted on the OAEI Bio-ML track demonstrate that GenOM can often achieve competitive performance, surpassing many baselines including traditional OM systems and recent LLM-based methods. Further ablation studies confirm the efectiveness of semantic enrichment and few-shot prompting, highlighting the framework's robustness and adaptability.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Ontology Matching</kwd>
        <kwd>Large Language Model</kwd>
        <kwd>Semantic Embedding</kwd>
        <kwd>Definition Generation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        In recent years, the rapid growth of domain-specific ontologies has led to a growing need for semantic
interoperability and knowledge integration across diverse knowledge systems [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Ontologies are
developed to formally represent concepts and relationships in a domain, yet they are often constructed in
isolation, following distinct modelling choices, terminological conventions, and structural assumptions.
This independence has resulted in significant heterogeneity across ontologies, which poses considerable
challenges to the integration and reuse of knowledge.
      </p>
      <p>
        Ontology Matching (OM) as known as ontology alignment, the task of identifying semantic
correspondences between entities in diferent ontologies, has therefore become a crucial area of research
[
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ]. By establishing links such as equivalence (indicating that two concepts represent the same or
highly similar meaning) or subsumption (where one concept is a more general or specific variant of
the other), ontology alignment facilitates accurate knowledge translation and consistent information
exchange between systems. However, matching concepts across ontologies is far from straightforward.
Three major sources of heterogeneity commonly hinder this process: (1) Terminological diferences,
where the same concept may be described using diferent labels or synonyms; (2) Structural diferences,
reflecting the varying levels of complexity in ontology design—from deeply nested hierarchies to flat
enumerative lists; (3) Granularity diferences, where the same domain knowledge may be captured with
difering levels of detail or abstraction.
      </p>
      <p>
        These variations significantly increase the cognitive and computational burden of OM. The challenge
is further magnified by the rapidly growing scale of modern ontologies. For example, SNOMED-CT
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], a widely adopted clinical terminology, contains several hundred thousand medical concepts. As
ontologies continue to expand in size and complexity, manual alignment methods become increasingly
infeasible, underscoring the necessity for automated or semi-automated alignment techniques capable
of operating at scale.
      </p>
      <p>
        Traditional OM systems, such as LogMap [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and AML [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], primarily rely on string matching, indexing
and structure matching technologies. They often fall short in capturing or fully utilising the underlying
semantic information of concepts. With the advent of large language models (LLMs), their remarkable
capabilities in text understanding and generalisation have attracted significant attention. Recently,
several ontology alignment systems have begun to incorporate LLMs to better capture the semantics
of complex concepts across heterogeneous ontologies. These approaches typically leverage LLMs for
identifying semantic similarities between concept pairs via embedding, and/or making direct alignment
decisions via generation. For example, LLM4OM [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] employs ChatGPT and OpenAI embeddings for
pairwise matching (see Section 2 for more related work analysis). However, current LLM-based OM
approaches still exhibit notable limitations. Some of them, despite leveraging powerful language models,
struggles to deliver satisfactory performance on more complex OM tasks. Some others can achieve
promising results on certain benchmarks, but they may rely on LLMs with very large-scale parameters
(e.g., 70B LLaMA-2 used in Olala [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]) that imposes substantial computational demands, raising concerns
about scalability.
      </p>
      <p>To address these limitations, we propose a novel ontology matching framework named GenOM
which utilises LLMs and extended textual definitions of concepts. The framework begins by extracting
both lexical and structural information from the source and target ontologies. This information is
then semantically enhanced using an LLM, resulting in more informative and context-aware concept
descriptions. Subsequently, these enriched representations are embedded into vector space, enabling
the retrieval of candidate alignments based on semantic similarity. A lightweight 7B-parameter LLM is
employed to assess the equivalence of candidate pairs through a classification-based approach, while
traditional exact matching techniques are incorporated to supplement and refine the alignment results.
The efectiveness of GenOM is demonstrated through comprehensive experiments evaluated using
standard metrics such as precision, recall, F1-score, mean reciprocal rank (MRR), and Hit@K. The model
achieves competitive performance compared to several state-of-the-art ontology alignment systems,
highlighting its robustness and practical applicability. Extensive ablation studies were also conducted
to demonstrate the efectiveness of the proposed framework across multiple dimensions.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        At present, OM approaches can be broadly categorised into four main types: traditional
knowledgebased systems, machine learning-based systems, pre-trained language model-based systems and more
recently, large language models (LLMs)-based systems. Traditional systems, such as LogMap and AML,
rely primarily on lexical similarity, structural heuristics, and external resources like UMLS or WordNet
[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. LogMap [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] extracts class names and searches for matches via external lexicons while addressing
logical inconsistencies by selecting alignments with higher confidence scores. AML [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] employs multiple
matching strategies, including exact and character-based matchers, and has shown strong performance
on medical datasets. However, these systems depend heavily on curated lexicons and often fail to fully
capture and utilise complex semantics of diferent kinds.
      </p>
      <p>With the rise of machine learning, several approaches have emerged that model alignment as
a classification or (embedding-based) similarity learning task [ 10, 11, 12, 13, 14, 15]. For instance,
DeepAlignment [11] vectorises class names and computes Euclidean distances to assess similarity, while
the CNN-based system [15] use character-level embeddings and hierarchical context to train binary
classifiers for equivalence detection. Although these methods ofer improvements over traditional
techniques, they often require large annotated datasets and extensive parameter tuning. Moreover,
their domain-specific nature limits transferability across ontologies in diferent fields.</p>
      <p>Using encoder-based pre-trained language models like BERT [16] have leveraged the representational
power of contextual embeddings and a memorization based on large-scale parameters learned from
corpora to address some of these limitations. BioSTransformers [17] adopt a Siamese architecture
based on domain-specific BERT models to compute semantic similarity between biomedical concepts.
Built on the BERT architecture [18], BERTMap fine-tunes a domain-specific BERT model on ontology
alignment corpora, enabling it to capture subtle semantic diferences even when lexical overlap is low.
This approach has demonstrated strong results, particularly in biomedical applications. BERTSubs [19]
takes a similar architecture as BERTMap but focuses on the subsumption relationship.</p>
      <p>
        More recently, LLMs such as GPT-3.5, GPT-4, and T5-XXXL have been applied to ontology alignment,
ofering stronger generalisation and semantic reasoning capabilities. Several studies have explored
prompt-based querying and retrieval-augmented generation to support alignment tasks. For example,
Norouz et al. [20] used GPT-4 to align ontology via prompts, observing high recall but reduced precision
due to the model incorrectly classifying subclass relationships as equivalence relations. Yuan et al.
[21] tested both open- and closed-source models on medical alignment tasks, using structural context
to enhance predictions. Other work has introduced hybrid frameworks combining vector similarity
retrieval (e.g., using SBERT) with LLM verification stages [
        <xref ref-type="bibr" rid="ref7">7, 22</xref>
        ], aiming to reduce hallucination
and improve alignment quality. The Olala system [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] further integrates embedding-based candidate
ifltering and post-processing with LLaMA-2 for final alignment decisions. While LLM-based approaches
have shown considerable potential, they still face challenges including scalability to large ontologies
such as SNOMED CT, struggles to deliver satisfactory performance on more complex OM tasks and
computational cost, particularly when relying on closed-source or very large models.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology</title>
      <sec id="sec-3-1">
        <title>3.1. Task Formulation</title>
        <p>
          The OM task can be formally defined as follows. Given two ontologies, referred to as the source ontology
s and the target ontology t, let  denote the set of named concepts in s and  denote the set of
named concepts in t. The objective is to identify a set of mappings, where each mapping consists of a
concept pair (, ), with  ∈  and  ∈ , that are considered semantically related. Formally, the
alignment output is represented as:
 = {(, ,  ) |  ∈ ,  ∈ ,  ∈ [
          <xref ref-type="bibr" rid="ref1">0, 1</xref>
          ]}
(1)
where  denotes the confidence score, quantifying the degree of semantic equivalence between  and
.
        </p>
        <p>This score often serves as a basis for selecting high-confidence mappings.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. System Architecture</title>
        <p>As shown in Figure 1, GenOM comprises five main components:
1. Ontology Data Extraction: Structural and lexical information is extracted for each named
concept from both the source and target ontologies. This includes labels, synonyms, parent
concepts, and axioms of concept equivalence.
2. Definition Generation : An LLM is prompted to generate natural language definitions or
paraphrased descriptions of each concept, based on the extracted information. This step enhances
the semantic representation of concepts, especially for those lacking explicit textual definitions.
3. Candidate Mapping Generation: Using the embedding model, pairwise cosine similarity scores
are computed between the source and target concepts to generate top- candidate mappings.
4. LLM-Based Equivalence Judgement: For each candidate mapping, an LLM is queried to
determine whether the mapping’s two concepts are semantically equivalent, with prompts that
incorporate enriched definitions and structural context.
5. Post-processing and Result Fusion: Filtering is performed based on both the probability
distribution of the LLM outputs and the cosine similarity scores, retaining only high-confidence
alignment results. In addition, exact matching modules are applied to recall highly confident
matches based on identical labels.</p>
        <p>By concept definition generation, candidate retrieval, and alignment judgement, GenOM integrates
the strengths of embedding-based similarity, exact lexical matching, and LLM reasoning.</p>
        <sec id="sec-3-2-1">
          <title>3.2.1. Ontology Data Extraction</title>
          <p>Both lexical and structural characteristics of each concept from the source and target ontologies are
extracted and utilized to support the following alignment process. Specifically, the extracted information
includes the concept’s label (defined by the annotation property rdfs:label), a set of synonyms (retrieved
using the annotation properties listed in Table 1), and its parent concepts. For concepts defined using
EquivalentClass axioms, the built-in verbalisation module in DeepOnto [23] is used to convert
logical expressions into natural language descriptions (Table 2).</p>
        </sec>
        <sec id="sec-3-2-2">
          <title>Property IRI</title>
          <p>Label
Synonym
http://www.w3.org/2000/01/rdf-schema#label
http://www.geneontology.org/formats/oboInOwl#hasSynonym
http://www.geneontology.org/formats/oboInOwl#hasExactSynonym
http://www.ebi.ac.uk/efo/alternative_term
http://www.orpha.net/ORDO/Orphanet_#symbol
http://purl.org/sig/ont/fma/synonym
http://www.w3.org/2004/02/skos/core#altLabel
http://www.w3.org/2004/02/skos/core#prefLabel
http://ncicb.nci.nih.gov/xml/owl/EVS/Thesaurus.owl#P108
http://ncicb.nci.nih.gov/xml/owl/EVS/Thesaurus.owl#P90</p>
        </sec>
        <sec id="sec-3-2-3">
          <title>Description Logic Axiom</title>
          <p>Product containing only betamethasone and calcipotriol (medicinal product) ≡
MedicinalProduct ⊓ ∃ RoleGroup.(∃ hasActiveIngredient.Betamethasone) ⊓
∃ RoleGroup.(∃ hasActiveIngredient.Calcipotriol)</p>
        </sec>
        <sec id="sec-3-2-4">
          <title>Verbalized Description</title>
          <p>Medicinal product (product) that Role group (attribute) something that Has active ingredient
(attribute) Betamethasone (substance) and something that Has active ingredient (attribute)
Calcipotriol (substance).</p>
        </sec>
        <sec id="sec-3-2-5">
          <title>3.2.2. Definition Generation</title>
          <p>LLMs encode extensive general purpose and domain knowledge. They are leveraged and integrated with
concept information explicitly represented in the ontology to enhance the semantic representation of the
concept. In this framework, definitions for medical concepts are generated by prompting an LLM with
background information extracted from the ontology, including the concept’s label, synonyms, parent
concepts, and, where applicable, the verbalised descriptions of logical expressions from EquivalentClass
axioms.</p>
          <p>Table 3 presents the prompt template employed for definition generation. In this template, the LLM
is provided with both the available concept-specific information and the name of the source ontology.
This additional context helps the LLM recall relevant domain knowledge encoded in its parameters,
thereby generating definitions that are more accurate and context-aware. The inclusion of the ontology
name in the prompt acts as supplementary guidance, particularly for sparsely annotated concepts.</p>
          <p>The amount and richness of information associated with a concept can vary considerably across
ontologies. Some concepts are well-described, including multiple synonyms, hierarchical structure, and
even formal axioms. In contrast, other concepts may be sparsely defined, often limited to a label with
little or no supporting context. In such cases, the LLM’s internalised knowledge becomes essential in
compensating for missing semantics.</p>
        </sec>
        <sec id="sec-3-2-6">
          <title>Prompt Templates for Concept Definition Generation</title>
        </sec>
        <sec id="sec-3-2-7">
          <title>Role: System</title>
          <p>You are generating a definition for a concept from the {s name} ontology. The definition will be used to align
it with candidate concepts in the {t name} ontology.</p>
          <p>You are a biomedical ontology expert. Your task is to generate a concise, alignment-friendly definition for a
given biomedical concept. The definition should be semantically precise, distinguishable from related terms, and
suitable for matching across ontologies.</p>
          <p>Only return the definition.</p>
        </sec>
        <sec id="sec-3-2-8">
          <title>Role: User</title>
          <p>Concept: Product containing only betamethasone and calcipotriol (medicinal product)
Synonyms: Betamethasone and calcipotriol only product
Parents: Product containing betamethasone and calcipotriol (medicinal product)</p>
          <p>Description: Medicinal product (product) that Role group (attribute) something that Has active ingredient
(attribute) Betamethasone (substance) and something that Has active ingredient (attribute) Calcipotriol (substance)</p>
        </sec>
        <sec id="sec-3-2-9">
          <title>Output:</title>
          <p>A medicinal product specifically formulated to contain solely betamethasone and calcipotriol as its active
ingredients, designed for the treatment or management of specific dermatological conditions.</p>
        </sec>
        <sec id="sec-3-2-10">
          <title>3.2.3. Candidate Mapping Generation</title>
          <p>To generate candidate concept pairs for alignment, this framework adopts an embedding-based retrieval
strategy. Concepts from both the source ontology s and the target ontology t are first encoded
into fixed-size vector representations. Each concept is represented using a combination of its label,
synonyms, and its enriched definition produced in the previous step. Figure 2 shows an example of
the input text fed into the embedding model. Structural information such as hierarchical relations
is deliberately excluded at this stage to reduce complexity in embedding. The framework adopts the
text-embedding-3-small1 model for embedding generation. Compared to traditional
encoderbased models such as Sentence-BERT [24], this LLM-based embedding model demonstrates superior
capability in distinguishing complex biomedical concepts.</p>
          <p>Once the embeddings are obtained, a cosine similarity-based retrieval process is applied to identify, for
each source concept, the top- most semantically similar concepts from the target ontology. Candidate
selection is based on vector similarity, allowing the system to retrieve a shortlist of potentially equivalent
concept pairs. These candidates are subsequently passed to the next stage for semantic equivalence
assessment.</p>
          <p>Embedding Input Example
Label: Product containing only betamethasone and calcipotriol (medicinal product);
Synonyms: Betamethasone and calcipotriol only product; Definition: A medicinal product
specifically formulated to contain solely betamethasone and calcipotriol as its active ... ;</p>
        </sec>
        <sec id="sec-3-2-11">
          <title>3.2.4. LLM-Based Equivalence Judgement</title>
          <p>In this stage, an LLM is employed to determine whether each candidate concept pair represents a
semantic equivalence. Rather than prompting the model to generate full descriptive justifications, which
would be time-consuming and potentially verbose, a lightweight classification strategy is adopted.
Specifically, each concept pair is presented via a prompt designed to elicit a binary response — YES if
the concepts are equivalent, and NO otherwise.</p>
          <p>To support this, a prompt (as shown in Table 4) is constructed with strong instructional guidance,
encouraging the model to respond using only a single classification token.</p>
          <p>The predicted equivalence score is then computed based on the probability of the YES token, extracted
directly from the model’s output logits. Specifically, given the model’s output logits z ∈ R at the final
decoding position (where  is the vocabulary size), the softmax function is applied to convert the logits
into a probability distribution:
 (YES) = softmax(z)YES =</p>
          <p>exp(YES)
∑︀=1 exp()</p>
          <p>Here, YES denotes the logit corresponding to the token YES. The resulting probability serves as
the model’s confidence in semantic equivalence for a given concept pair. Concept pairs with  (YES)
exceeding a predefined threshold are retained for alignment.</p>
          <p>This probability-based scoring approach significantly reduces inference time and simplifies
decisionmaking, as it avoids generating full-length text responses and instead relies on a single-token
classification strategy, while maintaining high alignment precision.
1https://platform.openai.com/docs/models/text-embedding-3-small</p>
        </sec>
        <sec id="sec-3-2-12">
          <title>Prompt Template for Equivalence Judgement</title>
        </sec>
        <sec id="sec-3-2-13">
          <title>System Message:</title>
          <p>You are an expert in biomedical concept classification. You will be given two biomedical concepts. Based
on the information provided, determine whether the two concepts refer to the same real-world entity
(ontology matching). Only respond with YES or NO.</p>
        </sec>
        <sec id="sec-3-2-14">
          <title>User Message:</title>
        </sec>
        <sec id="sec-3-2-15">
          <title>Concept A</title>
          <p>Name: {lateral rectus nerve}
Synonyms: {abducens nerve ...}
Superclass: {peripheral nerve of head and neck (body structure) ...}
Definition: {the lateral rectus nerve, also known as ...}</p>
        </sec>
        <sec id="sec-3-2-16">
          <title>Concept B</title>
          <p>Name: {abducent nerve [vi]}
Synonyms: {nervus abducens ...}
Superclass: {right posterior crico-arytenoid ligament ...}</p>
          <p>Definition: {the abducent nerve [vi] is a branch of the cranial nerve vi that innervates ...’]}</p>
        </sec>
        <sec id="sec-3-2-17">
          <title>3.2.5. Post-processing and Result Fusion</title>
          <p>To ensure the quality and reliability of the final alignment output, a post-processing stage is applied
to filter and refine the results generated by the LLM. First, a threshold  prob is imposed on the
tokenlevel probability associated with the YES response. Candidate pairs with confidence scores below this
threshold are discarded. In parallel, the cosine similarity scores obtained during candidate generation are
also considered, and pairs with lower than  cs embedding similarity are removed to prevent semantically
distant matches from being retained.</p>
          <p>After this dual-filtering step, To further enhance precision, outputs from the LLM module were merged
with the results of two exact matching systems, LogMapLt2 and BERTMapLt3, which are lightweight
versions of their original models, both simplified to include only the string matching component. This
fusion combines semantic reasoning with surface-level matching, improving overall coverage while
preserving precision. The resulting set constitutes the final alignment output.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Experiment and Evaluation</title>
      <sec id="sec-4-1">
        <title>4.1. Datasets and Evaluation Metrics</title>
        <p>The experiments were conducted on the OAEI 2024 Bio-ML4 track [25] whose benchmarks are
designed for biomedical ontology alignment tasks. The dataset comprises five sub-tasks involving six
widely used biomedical ontologies: Systematized Nomenclature of Medicine - Clinical
Terms (SNOMED-CT), National Cancer Institute Thesaurus (NCIT), Foundational
Model of Anatomy (FMA), Human Disease Ontology (DOID), Orphanet Rare Disease
Ontology (ORDO), and Online Mendelian Inheritance in Man (OMIM). The two oficial
evaluation protocols of the Bio-ML track were adopted: global matching which focuses on ranking the
correct target concept among a list of candidates, and local ranking which is to evaluate the system’s
ability to identify correct mappings among all the possible concept pairs across two ontologies. Precision
(P), Recall (R), and F1-score are measured for global matching, and Mean Reciprocal Rank (MRR) and
Hit@1 are calculated for local ranking. These metrics provide a comprehensive view of the OM systems.
2https://github.com/ernestojimenezruiz/logmap-matcher
3https://github.com/KRR-Oxford/DeepOnto/tree/main/src/deeponto/align/bertmap
4https://krr-oxford.github.io/OAEI-Bio-ML/2024/index.html</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Experiment Setup</title>
        <p>To systematically evaluate the proposed framework, we designed distinct experimental settings
tailored to each module of the pipeline. Definition Generation: For semantic enrichment, the
Qwen2.5-7B-Instruct-1M5 LLM was utilised to generate concise definitions based on concept-level
contextual information. The generation process was controlled with a temperature of 0.7 and a top_p
value of 0.9, ensuring that the definitions remained focused and semantically aligned with the
underlying concepts. Candidate Mapping Generation: Cosine similarity computations and HNSW-based
indexing were implemented via the faiss library to eficiently retrieve the top- 10 most similar concepts
for each source entity. LLM-based Judgement: The same Qwen2.5-7B-Instruct-1M model was
applied to perform binary equivalence classification over the candidate pairs. To accelerate inference,
we adopted the float16 data type. Post-processing and Result Fusion. In the final stage, results
were filtered using a token probability threshold  prob of 0.99 and a cosine similarity threshold  cs of
0.97. These thresholds were initially optimised on the SNOMED-NCIT (neoplas) task, then fixed and
consistently applied across all the other tasks. Moreover, since BERTMapLt demonstrated stronger
performance on the neoplas task among the exact matching models, it was selected for application to
the remaining tasks as well.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Overall Results</title>
        <p>Table 5 presents the comparison between our proposed method and several state-of-the-art ontology
alignment systems on the Bio-ML track. The results demonstrate that our model achieves consistently
strong performance across all tasks, ranking among the top three in every case. Moreover, in tasks
where our method ranks second, the performance gap with the best-performing system is marginal—for
instance, only 0.006 in the SNOMED-NCIT (pharm) task and as small as 0.001 in the SNOMED-FMA
(body) task.</p>
        <p>To ensure a fair assessment and avoid overestimating the performance of our model, we excluded the
neoplas task from the overall evaluation, as it was used during threshold tuning. When evaluated
solely on the remaining unseen tasks—where no threshold optimisation was performed—our framework
still achieved the highest average F1 score of 0.769. This surpasses the second-best method, BERTMap
(0.762), and the third-best, LogMapBio (0.760). These results indicate that the proposed approach not
only performs competitively on tuned datasets but also maintains strong and consistent performance
across previously unseen tasks, highlighting its robustness and generalisability in biomedical ontology
alignment.</p>
        <p>In addition, when compared to the other LLM-based OM system LLM4OM which leverages
ChatGPT3.5 and OpenAI’s embedding model as reported in its original paper, our approach GenOM delivers
consistently better performance across all evaluated tasks, including the last four tasks where GenOM
generalises the hyper parameter settings optimised from the first task.</p>
      </sec>
      <sec id="sec-4-4">
        <title>4.4. Ablation Study</title>
        <sec id="sec-4-4-1">
          <title>4.4.1. Impact of Definition Enrichment on Local Ranking and Candidate Retrieval</title>
          <p>We study the impact of the generated concept definition on each of the two stages: LLM-based
equivalence judgement and candidate generation. For LLM-based equivalence judgement, we compare using
only the concept label, and using both the label and the generated definition. The local ranking result
is shown in Table 6. In particular, both MRR and Hit@1 show noticeable gains in all tasks except for
SNOMED-FMA (body), where performance remains comparable.</p>
          <p>For the candidate generation stage, we additionally report Hit@5 and Hit@10, as these metrics are
essential for determining the appropriate top- value—that is, how many candidate concepts should be
passed to the LLM for equivalence judgement. As shown in Table 7, incorporating definition information
led to improvements across all three metrics (Hit@1, Hit@5, Hit@10) in all tasks, except for a slight
5https://huggingface.co/Qwen/Qwen2.5-7B-Instruct-1M
SNOMED-NCIT (Neoplas)
SNOMED-NCIT (Pharm)
SNOMED-FMA (Body)
OMIM-ORDO
NCIT-DOID</p>
          <p>System
LogMap
LogMapBio
LogMapLt
Matcha
BERTMap
BERTMapLt
BioSTransMatch
LLM4OM
GenOM
LogMap
LogMapBio
LogMapLt
Matcha
BERTMap
BERTMapLt
BioSTransMatch
LLM4OM
GenOM
LogMap
LogMapBio
LogMapLt
Matcha
BERTMap
BERTMapLt
BioSTransMatch
LLM4OM
GenOM
LogMap
LogMapBio
LogMapLt
Matcha
BERTMap
BERTMapLt
BioSTransMatch
LLM4OM
GenOM
LogMap
LogMapBio
LogMapLt
Matcha
BERTMap
BERTMapLt
BioSTransMatch
LLM4OM
GenOM
Overall performance on the five OAEI 2024 Bio-ML tasks. NA indicates no results as the systems do not
supporting the calculation of the metrics. Bold indicates the best performance, while underline indicates the
second-best. All the baseline results are from the track’s website, where the LogMapLt and the BERTMap series
are reimplemented with system submissions, while the other baselines submitted the final result files without
system reimplementation, probably using settings optimised for each task.
drop in Hit@10 on the SNOMED-NCIT (pharm) task. These results suggest that enriched definitions
help retrieve a greater number of correct candidates, thereby increasing the likelihood of including the
true match within the top- shortlist.</p>
        </sec>
        <sec id="sec-4-4-2">
          <title>4.4.2. Efectiveness Compared to Original Exact Matching Methods</title>
          <p>We compare the results of GenOM with the results of the stand-alone exact matching system that
GenOM adopts in the final stage (either BERTMapLt or LogMapLt). The BERTMapLt and LogMapLt
Task
OMIM_ORDO (with)
OMIM_ORDO (without)
NCIT_DOID (with)
NCIT_DOID (without)
SNOMED_NCIT_pharm (with)
SNOMED_NCIT_pharm (without)
SNOMED_NCIT_neoplas (with)
SNOMED_NCIT_neoplas (without)
SNOMED_FMA_body (with)
SNOMED_FMA_body (without)
results reported in this section are reproduced within the scope of this work, and may therefore difer
slightly from the results presented earlier. As shown in Table 8, GenOM consistently outperforms the
original exact matchers across all evaluated tasks.</p>
          <p>GenOM (BERTMapLt) achieved an average F1 score of 0.771 across the five benchmark tasks,
outperforming the original BERTMapLt model, which obtained an average of 0.744. A similar improvement
was observed with LogMapLt: while the standalone LogMapLt achieved an average F1 score of only
0.631, the GenOM-enhanced version reached 0.716. These results suggest that while exact matching
provides a solid foundation for identifying high-confidence correspondences, it remains limited in
capturing more nuanced semantic equivalence. By integrating LLM-based reasoning and enriched
conceptual representations, GenOM is able to significantly enhance both coverage and accuracy over
the base exact matching techniques. This is reflected in a notable increase in recall: GenOM achieves,
on average, an 8% improvement in recall over BERTMapLt, and an even more substantial 24% increase
when compared to LogMapLt.</p>
        </sec>
        <sec id="sec-4-4-3">
          <title>4.4.3. Efectiveness of Few-Shot Prompting</title>
          <p>This experiment also investigates the efect of few-shot prompting on the LLM-based equivalence
judgement stage. The results are shown inTable 9, where few-shot prompting is set to 2 examples, all
results are reported prior to the integration of exact matching, and the threshold for cosine similarity
was kept consistent with the earlier setting at 0.97. Aside from the inclusion of few-shot examples, all
other settings are identical. The evaluation was conducted without incorporating results from the exact
matching module, as the impact of the few-shot strategy tends to be diminished once exact matching is
applied. The results indicate that incorporating two-shot examples consistently improves performance
across most tasks. Except for the NCIT-DOID task, where performance remains unchanged, all other
tasks exhibit notable gains in F1 score. This demonstrates that few-shot prompting can efectively guide
the LLM towards more accurate classification, especially in borderline cases where single-instance
reasoning may be insuficient.
SNOMED-NCIT-neoplas
SNOMED-NCIT-pharm
SNOMED-FMA-body
NCIT-DOID
OMIM-ORDO
GenOM(BERTMapLt) 0.795
BERTMapLt 0.831
GenOM(LogMapLt ) 0.869
LogMapLt 0.952
GenOM(BERTMapLt) 0.989
BERTMapLt 0.981
GenOM(LogMapLt ) 0.988
LogMapLt 0.996
GenOM(BERTMapLt) 0.944
BERTMapLt 0.979
GenOM(LogMapLt ) 0.876
LogMapLt 0.971
GenOM(BERTMapLt) 0.912
BERTMapLt 0.919
GenOM(LogMapLt ) 0.939
LogMapLt 0.955
GenOM(BERTMapLt) 0.803
BERTMapLt 0.834
GenOM(LogMapLt ) 0.839
LogMapLt 0.937</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion, Discussion and Future Work</title>
      <p>This paper presents GenOM, a general-purpose framework for ontology alignment that integrates
concept semantic enrichment with LLM-based textual definition generation, embedding-based candidate
retrieval, LLM prompting-based equivalence judgement, and exact matching in a modular design. The
approach demonstrates strong performance across five biomedical ontology alignment tasks of OAEI
Bio-ML, outperforming many baselines, and the efectiveness of its important components has been
verified via extensive ablation studies. In particular, the framework shows its ability to generalise across
datasets while maintaining alignment accuracy, without relying on handcrafted features or extensive
task-specific engineering.</p>
      <p>Although GemOM has achieved promising performance for equivalence mappings in OM, several
challenges remain:
1. It is dificult to consistently assess the degree of equivalence between concept pairs. This challenge
afects both the LLM-based judgement stage and the choice of cosine similarity threshold for
candidate retrieval.
2. The definition of equivalence which can vary subtly across tasks: concept pairs deemed equivalent
in one alignment task may not be considered so in another, leading to inconsistencies in judgement.
Accordingly, the optimal similarity threshold becomes task-dependent. For example, a cosine
similarity above 0.80 indicates equivalence in some tasks, but some other tasks may require a
threshold of 0.95 to ensure equivalence.
3. The alignment performance of LLMs is highly sensitive to the prompt. In evaluation, we observed
that vague prompts such as simply asking the model to “determine whether two concepts are
equivalent” often fails to elicit correct predictions; in many cases, the LLM almost never produces
a “YES” output. This highlights the importance of prompt specificity in steering LLM behaviour
and underscores a practical challenge in applying LLMs to alignment in a generalisable way.</p>
      <p>For the future work, one key direction is to expand the scope of GenOM to include additional
alignment types beyond equivalence, such as subsumption. Another key direction involves addressing
the variability in how equivalence is defined across diferent ontologies and tasks. In many alignment
scenarios, the threshold for considering two concepts equivalent may depend on contextual or
domainspecific nuances, which are dificult to capture using a fixed similarity score or binary decision. To
tackle this, future research will explore task-adaptive alignment criteria, including dynamic threshold
selection and prompt-based calibration techniques that allow the LLM to assess the strength or type of
correspondence more flexibly. Additionally, incorporating finer-grained semantic similarity measures
and confidence estimation strategies could help better reflect the spectrum of equivalence relations
observed in practice.</p>
    </sec>
    <sec id="sec-6">
      <title>Declaration on Generative AI</title>
      <p>During the preparation of this work, the authors used Grammarly in order to grammar and spell check,
and improve the text readability. After using the tool, the authors reviewed and edited the content as
needed to take full responsibility for the publication’s content.
[10] A. Doan, J. Madhavan, P. Domingos, A. Halevy, Ontology matching: A machine learning approach,
in: Handbook on ontologies, Springer, 2004, pp. 385–403.
[11] P. Kolyvakis, A. Kalousis, D. Kiritsis, Deepalignment: Unsupervised ontology matching with
refined word vectors, in: Proceedings of the 16th Annual Conference of the North American
Chapter of the Association for Computational Linguistics: Human Language Technologies, 1-6
June 2018, 2018.
[12] I. Nkisi-Orji, N. Wiratunga, S. Massie, K.-Y. Hui, R. Heaven, Ontology alignment based on word
embedding and random forest classification, in: Joint European Conference on Machine Learning
and Knowledge Discovery in Databases, Springer, 2018, pp. 557–572.
[13] L. L. Wang, C. Bhagavatula, M. Neumann, K. Lo, C. Wilhelm, W. Ammar, Ontology alignment in
the biomedical domain using entity definitions and context, in: Proceedings of the BioNLP 2018
workshop, 2018, pp. 47–55.
[14] J. Chen, E. Jiménez-Ruiz, I. Horrocks, D. Antonyrajah, A. Hadian, J. Lee, Augmenting ontology
alignment by semantic embedding and distant supervision, in: European Semantic Web Conference,
Springer, 2021, pp. 392–408.
[15] A. Bento, A. Zouaq, M. Gagnon, Ontology matching using convolutional neural networks, in:</p>
      <p>Proceedings of the Twelfth Language Resources and Evaluation Conference, 2020, pp. 5648–5653.
[16] Y. He, J. Chen, D. Antonyrajah, I. Horrocks, Bertmap: a bert-based ontology alignment system, in:</p>
      <p>Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, 2022, pp. 5684–5691.
[17] S. Menad, W. Laddada, S. Abdeddaïm, L. F. Soualmia, Biostransformers for biomedical ontologies
alignment., in: KEOD, 2023, pp. 73–84.
[18] J. Devlin, Bert: Pre-training of deep bidirectional transformers for language understanding, arXiv
preprint arXiv:1810.04805 (2018).
[19] J. Chen, Y. He, Y. Geng, E. Jiménez-Ruiz, H. Dong, I. Horrocks, Contextual semantic embeddings
for ontology subsumption prediction, World Wide Web (WWW) 26 (2023) 2569–2591. URL:
https://doi.org/10.1007/s11280-023-01169-9. doi:10.1007/S11280-023-01169-9.
[20] S. S. Norouzi, M. S. Mahdavinejad, P. Hitzler, Conversational ontology alignment with chatgpt,</p>
      <p>ArXiv abs/2308.09217 (2023). URL: https://api.semanticscholar.org/CorpusID:261031024.
[21] Y. He, J. Chen, H. Dong, I. Horrocks, Exploring large language models for ontology alignment,
arXiv preprint arXiv:2309.07172 (2023).
[22] Z. Qiang, W. Wang, K. Taylor, Agent-om: Leveraging llm agents for ontology matching, arXiv
preprint arXiv:2312.00326 (2024).
[23] Y. He, J. Chen, H. Dong, I. Horrocks, C. Allocca, T. Kim, B. Sapkota, DeepOnto: A python
package for ontology engineering with deep learning, Semantic Web 15 (2024) 1991–2004. URL:
download/2024/HeCDHAKS24.pdf. doi:10.3233/SW-243568.
[24] N. Reimers, I. Gurevych, Sentence-bert: Sentence embeddings using siamese bert-networks, arXiv
preprint arXiv:1908.10084 (2019).
[25] Y. He, J. Chen, H. Dong, E. Jiménez-Ruiz, A. Hadian, I. Horrocks, Machine learning-friendly
biomedical datasets for equivalence and subsumption ontology matching, in: International
Semantic Web Conference, Springer, 2022, pp. 575–591.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.</given-names>
            <surname>Staab</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Studer</surname>
          </string-name>
          , Handbook on ontologies, Springer Science &amp; Business
          <string-name>
            <surname>Media</surname>
          </string-name>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>P.</given-names>
            <surname>Shvaiko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Euzenat</surname>
          </string-name>
          ,
          <article-title>Ontology matching: State of the art and future challenges</article-title>
          ,
          <source>IEEE Transactions on Knowledge and Data Engineering</source>
          <volume>25</volume>
          (
          <year>2013</year>
          )
          <fpage>158</fpage>
          -
          <lpage>176</lpage>
          . doi:
          <volume>10</volume>
          .1109/TKDE.
          <year>2011</year>
          .
          <volume>253</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Euzenat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Shvaiko</surname>
          </string-name>
          , Ontology matching, 2nd ed., Springer-Verlag, Heidelberg (DE),
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>SNOMED</given-names>
            <surname>International</surname>
          </string-name>
          ,
          <source>SNOMED CT: The Global Clinical Terminology</source>
          , https://www.snomed.org,
          <year>2024</year>
          . Accessed:
          <fpage>2025</fpage>
          -06-23.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>E.</given-names>
            <surname>Jiménez-Ruiz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. Cuenca</given-names>
            <surname>Grau</surname>
          </string-name>
          ,
          <article-title>Logmap: Logic-based and scalable ontology matching</article-title>
          ,
          <source>in: The Semantic Web-ISWC</source>
          <year>2011</year>
          : 10th International Semantic Web Conference, Bonn, Germany,
          <source>October 23-27</source>
          ,
          <year>2011</year>
          , Proceedings,
          <source>Part I 10</source>
          , Springer,
          <year>2011</year>
          , pp.
          <fpage>273</fpage>
          -
          <lpage>288</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>D.</given-names>
            <surname>Faria</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Pesquita</surname>
          </string-name>
          , E. Santos,
          <string-name>
            <given-names>M.</given-names>
            <surname>Palmonari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. F.</given-names>
            <surname>Cruz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. M.</given-names>
            <surname>Couto</surname>
          </string-name>
          ,
          <article-title>The agreementmakerlight ontology matching system, in: On the Move to Meaningful Internet Systems: OTM 2013 Conferences: Confederated International Conferences: CoopIS</article-title>
          ,
          <string-name>
            <surname>DOA-Trusted</surname>
            <given-names>Cloud</given-names>
          </string-name>
          ,
          <source>and ODBASE</source>
          <year>2013</year>
          , Graz, Austria, September 9-
          <issue>13</issue>
          ,
          <year>2013</year>
          . Proceedings, Springer,
          <year>2013</year>
          , pp.
          <fpage>527</fpage>
          -
          <lpage>541</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>H. B.</given-names>
            <surname>Giglou</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. D'Souza</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Engel</surname>
          </string-name>
          , S. Auer,
          <article-title>Llms4om: Matching ontologies with large language models</article-title>
          ,
          <source>arXiv preprint arXiv:2404.10317</source>
          (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S.</given-names>
            <surname>Hertling</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Paulheim</surname>
          </string-name>
          , Olala:
          <article-title>Ontology matching with large language models</article-title>
          ,
          <source>in: Proceedings of the 12th Knowledge Capture Conference</source>
          <year>2023</year>
          ,
          <year>2023</year>
          , pp.
          <fpage>131</fpage>
          -
          <lpage>139</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Anam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y. S.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. H.</given-names>
            <surname>Kang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <article-title>Review of ontology matching approaches and challenges</article-title>
          ,
          <source>International Journal of Computer Science and Network Solutions</source>
          <volume>3</volume>
          (
          <year>2015</year>
          )
          <fpage>1</fpage>
          -
          <lpage>27</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>