<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Initial achievements in relation extraction from RNA-focused scientific papers</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Emanuele Cavalleri</string-name>
          <email>emanuele.cavalleri@unimi.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mauricio Soto-Gomez</string-name>
          <email>mauricio.soto@unimi.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ali Pashaeibarough</string-name>
          <email>ali.pashaeibarough@unimi.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dario Malchiodi</string-name>
          <email>dario.malchiodi@unimi.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Harry Caufield</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Justin Reese</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christopher J. Mungall</string-name>
          <email>CJMungall@lbl.gov</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Peter N. Robinson</string-name>
          <email>peter.robinson@bih-charite.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Elena Casiraghi</string-name>
          <email>elena.casiraghi@unimi.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giorgio Valentini</string-name>
          <email>giorgio.valentini@unimi.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marco Mesiti</string-name>
          <email>marco.mesiti@unimi.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Berlin Institute of Health - Charité</institution>
          ,
          <addr-line>Universitätsmedizin, Berlin, 13353</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Computer Science, Università di Milano</institution>
          ,
          <addr-line>Via Celoria 18, 20133 Milano</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>ELLIS, European Laboratory for Learning and Intelligent Systems, Milan Unit</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Environmental Genomics and Systems Biology Division, Lawrence Berkeley National Laboratory</institution>
          ,
          <addr-line>Berkeley, CA</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Relation extraction from the scientific literature to comply with a domain ontology is a well-known problem in natural language processing and is particularly critical in precision medicine. The advent of large language models (LLMs) has paved the way for the development of new efective approaches to this problem, but the extracted relations can be afected by issues such as hallucination, which must be minimized. In this paper, we present the initial design and preliminary experimental validation of SPIREX, an extension of the SPIRES-based system for the extraction of RDF triples from scientific literature involving RNA molecules. Our system exploits schema constraints in the formulations of LLM prompts along with our RNA-based KG, RNA-KG, for evaluating the plausibility of the extracted triples. RNA-KG contains more than 9M edges representing diferent kinds of relationships in which RNA molecules can be involved. Initial experimental results on a controlled data set are quite encouraging.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;RNA-based Knowledge Graphs</kwd>
        <kwd>relation discovery</kwd>
        <kwd>LLM</kwd>
        <kwd>Prompt Engineering</kwd>
        <kwd>Link Prediction</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Ribonucleic acid (RNA) plays a critical role in the central dogma of molecular biology, serving as
the intermediary between DNA and proteins, the building blocks of life. Beyond its traditional
role in protein synthesis, RNA is involved in a variety of cellular processes, including gene
regulation and catalysis, highlighting its importance in understanding the complexities of
biological systems. RNA-KG [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] is the first ontology-based knowledge graph for representing
coding and non-coding RNA molecules and their interactions with other biomolecular data
as well as with pathways, abnormal phenotypes and diseases to support the study and the
discovery of the biological role of RNA. RNA-KG contains around 9M edges extracted from
more than 50 public data sources and can be exploited to study RNA molecules and develop
innovative graph algorithms to support knowledge discovery in data science.
      </p>
      <p>
        The manual ingestion of triples in a knowledge graph by expert curators is a time-consuming
and costly operation and tools supporting them in the extraction of biological entities and their
relationships from plain texts are highly demanding. The advent of LLMs [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] has paved the
way for the development of new efective tools for this problem [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. However, these techniques
have shown diferent limitations, such as generating incorrect statements due to hallucinations
(inaccurate, nonsensical, or irrelevant output in the given context) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and insensitivity to
negations [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], that cannot be tolerated in sensitive domains like precision medicine. SPIRES
(Structured Prompt Interrogation and Recursive Extraction of Semantics) [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] is a recently
proposed knowledge extraction approach that exploits LLMs to identify instances of a knowledge
schema expressed in terms of LinkML [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Since the schema contains a conceptualization of a
given domain in terms of concepts, relationships, and properties we are interested in, it can
be used for defining more efective LLM prompts. Additionally, SPIRES allows grounding of
atomic textual elements as concepts taken from a variety of OBO Foundry ontologies [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>Even if SPIRES has proven its eficiency in the extraction of triples from plain text according
to bio-ontologies, there is the need to evaluate the reliability of the extracted triples both in
terms of the generated identifiers (i.e. they correctly represent the identified entities) and the
accuracy of their data source. In this paper, we address this problem by exploiting RNA-KG
as a gold standard in the RNA world because it contains many interactions involving RNA
molecules and can be used to evaluate the plausibility of the extracted triples.</p>
      <p>
        To leverage SPIRES for its ability to extracting triples from texts and supporting experts in
their validation, we present SPIREX, a system for the extraction of reliable triples from scientific
papers. There are two main backbones of the system. On one side, SPIRES and the LinkML
representation of the RNA-KG schema [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] allow the extraction of RDF triples compliant with the
domain Ontology. On the other side, we use RNA-KG as a gold standard providing knowledge
about interactions involving RNA molecules and use link prediction techniques to validate the
’plausibility’ of the extracted triples; i.e., the likelihood of the triple to be part of RNA-KG. The
initial experimental results on a manually curated testbed of 60 scientific texts are encouraging.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. RNA-KG and SPIRES</title>
      <p>
        RNA-KG [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] is the first knowledge graph encompassing biological knowledge about RNAs
gathered from more than 50 public databases, integrating functional relationships with genes,
proteins, chemicals, and ontologically grounded biomedical concepts. The current release
of RNA-KG has a single component containing around 600K nodes and 9M edges and can be
queried via SPARQL endpoint at https://RNA-KG.anacleto.di.unimi.it. Nodes are usually mapped
to reference biomedical vocabularies and ontologies such as NCBI Gene Entrez identifiers for
uniquely identifying genes and many kinds of non-coding RNAs (ncRNAs), Human Phenotype
Ontology (HPO [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]) for phenotypes, Monarch merged disease ontology (Mondo [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]) for
species 0.4%(2,148)
sequence 0.4%(2,363)
other terms 0.4%(2,427)
pathway 0.5%(2,606)
vaccine 1.1%(6,246)
anatomy 2.5%(14,232)
phenotype 2.9%(16,865)
      </p>
      <p>
        disease 4.0%(23,270)
diseases, and Gene Ontology (GO [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]) for annotating ncRNAs. Moreover, all the possible
interactions are represented by means of the Relation Ontology (RO [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]). This ensures common
semantics for the diferent relationships that can be extracted from the sources.
      </p>
      <p>
        Figure 1a shows the distribution of nodes contained in RNA-KG (details in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]). Nodes can be
classified into nodes representing ontological terms and bio-entities lacking a direct mapping to
ontological terms. Bio-entities have been further subdivided into RNA nodes, and non-RNA
nodes (named other bio-entities) that contain, for instance, gene and nodes describing
genomics features (e.g., nucleotide substitution). Figure 1b shows the distribution of edges in
RNA-KG. Edges have been subdivided into three categories: ) edges representing RO properties
that characterize interactions among RNA molecules in the considered sources; ) other edges
not belonging to RO properties; ) edges representing the subClassOf relationships. The
edges of the last two categories are introduced from the integration of bio-ontologies into
RNA-KG and the lack of a dedicated ontology for RNA molecules. When RNA molecules cannot
be precisely mapped to a reference ontology, they are included as subClassOf an appropriate
class within Sequence Ontology (SO [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]).
      </p>
      <p>
        SPIRES [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] is a recently proposed approach to information extraction that creates and refines
prompts to maximize the efectiveness of LLMs by exploiting domain knowledge encapsulated
through a schema expressed in LinkML [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. By identifying and extracting relevant information
from an input text, it adopts zero-shot or few-shot learning to identify and extract relevant
entities and relationships among them, which are then normalized and grounded through ontologies
and vocabularies. SPIRES is a general-purpose approach that can be used across a variety of
domains and does not require specific training/tuning on the considered domain. SPIRES adopts
an engineering approach for creating prompts for interacting with an LLM (like GPT3, GPT4)
to improve the quality of the generated responses through the use of domain-specific schema.
In this way, technical challenges for generative AI (e.g., constructing comprehensive real-world
regulates activity of
n miRNA
+ id: String
+ description: String
+ sequence: String
n + family name: String
n
      </p>
      <p>gene
n regulates activity of n + id: Integer</p>
      <p>+ type: String
n interacts with n + symbol: String</p>
      <p>protein
1 gene product of 1n ++ idde:sSctrriipntgion: String
nn rheagsulgaetenseapcrtoivdituycotf n + synonym: String list
+ ortholog: String list
+ sequence: String
causally related to
n</p>
      <p>n
disease
+ id: String</p>
      <p>n
n ++ sdyensocrnipymtio:nS:tSritnrginlgist n
causes or contributes to condition
causes or contributes to condition
knowledge and improving the accuracy of automated responses) can be addressed.</p>
      <p>The specification of this schema in LinkML contains the classes of entities and relationships
among them within the specified domain. Classes can also include attributes (e.g., name,
type, and list of synonyms) to enrich entity description. The LinkML schema is automatically
processed to generate a list of prompts through which SPIRES interacts with a LLM. Each prompt
of the list is submitted to the LLM for collecting information that is exploited for completing
the following prompt by eventually considering the bio-ontologies (e.g., for changing a protein
symbol with the corresponding identifier in an ontology). This recursive refinement process
improves the quality of the information gathered through the LLM.</p>
    </sec>
    <sec id="sec-3">
      <title>3. The SPIREX system</title>
      <p>
        As shown in the architecture in Figure 3, SPIREX is composed of two modules: the SPIRES
module is used for extracting the RDF triples from scientific abstracts. Then, an embedding of
RNA-KG is used to validate the generated triples and score their level of plausibility.
SPIRES module for RNA-KG. Through the study of the scientific literature about RNA
interactions, and the analysis of more than 50 data sources [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] all over the world, we have
identified the kinds of relationships that can involve RNA molecules and reported them in a
meta-graph [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Figure 2 shows an excerpt of the UML schema describing the entities that are
connected to miRNA molecules through diferent kinds of relationships in the meta-graph.
      </p>
      <p>
        Starting from its LinkML representation, a list of prompts specific for the RNA domain are
generated according to which entities and the relationships contained in a text are extracted
by considering the schema constraints. Moreover, SPIRES adopts bio-ontology of our domain
(details in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]) for producing source and target identifiers according to the RNA-KG identification
scheme and RO predicates.
      </p>
      <p>RNA-KG module for link prediction. The validation of new potential relations derived
from the SPIRES module can be modeled as a link prediction task on RNA-KG, performed
via either Graph Neural Networks (GNNs) or Random-Walk (RW) based methods for Graph
Representation Learning. GNN approaches usually present scalability issues, while RW-based
graph embedding overcomes this problem by the use of random-walk approaches that sample
the graph to construct a representation of the nodes (and edges) in a lower dimensional vector
space that feeds traditional ML models.</p>
      <p>
        In SPIREX we have used Node2vec [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] for the embedding of RNA-KG. Node2vec is a
wellknown random walk-based approach that aims to capture the graph topology from the node
neighborhoods. The model generates a set of second order random walks across the graph, that
are used to train a shallow neural network to compute a vector representation of the graph
components. One of the key features of Node2vec is the possibility to generate paths that
focus either on the local or global structure of the graph, providing a great flexibility in the
graph representation. Our system uses the implementation available in GRAPE [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], a software
resource specifically designed for the manipulation and embedding of large graphs.
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Preliminary experimental results</title>
      <p>
        Experiments have been realized for both modules of SPIREX. For the first module, we evaluated
the prediction accuracy of SPIRES in extracting triples in a set of manually annotated documents.
We also compared SPIRES with base LLMs to verify the advantage of using LinkML in the
specification of the domain schema. For the second module, we checked if the simple predictive
model can generate reasonable scores on RNA-KG. Finally, we assessed the ability of the
predictor to evaluate the plausibility of triples extracted through SPIRES according to RNA-KG.
SPIRES prediction accuracy and comparison with base LLMs. As described in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], a
corpus of 60 scientific articles related to RNA molecules and their interactions has been gathered
from PubMed, ResearchGate, and Google Scholar. Starting from them, we have identified
abstracts, discussions, or specific subsections within the domain of interest. They have been
manually annotated with the entities and the six kinds of interactions that can be extracted
from them (reported in the y-axis of the diagram in Figure 4a).
      </p>
      <p>For evaluating the predictions, we have used standard metrics (precision, recall, and F-score)
by considering the True Positive (TP), False Positive (FP), and False Negative (FN) according to
the manually tagged paragraphs. As shown in Figure 4a, the obtained results, using
GPT3.5turbo in SPIRES for each category of interaction, indicate a consistent trend where TP rate
tends to be higher with respect to both FP and FN rates. The only exception is for
proteindisease relations, where FN rate is higher than TP rate. We noticed that many protein-disease
relations are undetected, often because they are expressed in complex ways and this can lead to
inaccurate entity recognition. Despite this, the overall precision remains remarkably high and,
in biomedicine, this is preferable because it prioritizes certainty over ambiguity.</p>
      <p>
        We have assessed the performance of SPIRES by considering as baseline approaches OpenAI
GPT (ver. GPT3.5-turbo) and Llama 2 [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] (ver. llama-2-70b-chat). As back-end LLM of SPIRES,
we have considered both GPT3.5-turbo and GPT4-turbo. We have manually grounded instances
and relationships that can be extracted from 20 documents among those considered in the
previous experiment. Regarding the prompt to be used with the base LLM system, we have
considered a simple one requesting to extract triples from the considered text with an explicit
request for mapping the extracted concepts to appropriate terminologies. Given that both
OpenAI GPT and Llama 2 caution that the ontology identifiers provided are hypothetical and
might not align with actual identifiers in the ontologies, and considering the general community
advice against relying on IDs from an LLM [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], we decided to substitute the grounding process
with our manually curated look-up tables [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>As shown in Figure 4b, SPIRES outperforms baseline LLMs used alone both in terms of
precision and recall. The histogram points out a high increment in TP rate and a decrease in FP
and FN rates when adopting SPIRES for extracting relations that adhere to a specified schema
within texts. Furthermore, when adopting GPT4-turbo in SPIRES the recall metric improves
due to the lower FN rate with a positive efect on the F-score.</p>
      <p>Evaluation of the plausibility of SPIREX predictions. For evaluating the plausibility, a
restricted RNA-KG view has been considered that roughly corresponds to the schema in Figure 2
focusing on the predictions of miRNA-disease relationships. More precisely, we have considered
two diferent settings of the hold-out procedure to evaluate prediction performance.
0</p>
      <p>0</p>
      <p>0.8
0.8
1
1
0
0.2</p>
      <p>0.4 0.6
Prediction Probabilities</p>
      <p>0.8
(c)</p>
      <p>
        In the first one, named RNA-KG Δ , the test set corresponds to triples involving miRNAs
and diseases from the source RNAdisease [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], while training set corresponds to the remaining
miRNA-disease triples of RNA-KG; in the second one, named RNA-KGΔ10% , we randomly
included in the test set 10% of miRNA-disease triples, and in the training set the remaining
90%, independent of their original source, guaranteeing to maintain the graph connectivity,
according to a connected Monte-Carlo hold-out strategy [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
      </p>
      <p>We directly applied node2vec to the prediction of miRNA-disease edges according to the
RNAKGΔ and RNA-KGΔ10% experimental settings, using a Multi-Layer-Perceptron trained
on the node2vec edge embeddings. The default parameters adopted in GRAPE have been
chosen. Figure 5a and 5b show that the triples in the test set exhibit high probabilities, in both
experimental settings. As reported in Figure 5a, with RNA-KGΔ , ∼ 63% of these triples are
associated with a score higher than 0.6. In the case of RNA-KGΔ10% , we notice that ∼ 88% of the
test set was correctly classified with a probability higher than 0.6 and ∼ 74% with a probability
higher than 0.8 (see Figure 5b). Node2vec is thus a reasonable predictor of miRNA-disease edges
to be included in the RNA-KG and can be used to assess the plausibility of SPIRES predictions.</p>
      <p>To assess the ability of node2vec in evaluating the plausibility of the triples extracted by
SPIRES, we have considered true positive triples extracted from our manually curated dataset
involving miRNAs and diseases. Figures 5c and 5d show the distribution of the probabilities
predicted by node2vec on the miRNA-disease edges extracted by SPIRES. Specifically, blue
columns represent the number of miRNA-disease triples that are already included in RNA-KG,
whereas orange columns represent the number of triples that are missing in RNA-KG. In both
cases, node2vec is able to correctly classify almost all the tuples already present in the partial
KG but can also discriminate between plausible and implausible new triples, ofering a potential
validation tool. Indeed in both experimental settings, we can identify a set of edges included in
RNA-KG that are predicted with a high probability by both SPIRES and node2vec (blue bars), but
also a set of edges extracted by SPIRES and predicted with a high probability by node2vec, even
if these edges are not present in RNA-KG (orange bars). These last edges can be considered as
possible new candidates for miRNA-disease relationships. In Figures 5c and 5d, the orange bars
denote relationships identified by SPIRES, yet assigned a low probability by node2vec. These
edges can be considered “uncertain” in the sense that they are not confirmed by an independent
edge prediction method that exploits the topological characteristics of the RNA-KG. We believe
these results can be improved by considering expanded views of RNA-KG and more complex
ML methods capable of accommodating its inherent heterogeneity.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Concluding remarks</title>
      <p>
        In this paper we have described the initial steps in the design and development of the SPIREX
system for the extraction of meaningful triples from scientific papers that exploit RNA-KG as
a gold standard for checking the plausibility of the extracted triples. The initial experimental
results are encouraging of the efectiveness of the proposed tool. At the current stage, we
have used a basic link prediction measure for assessing the relationship’s plausibility according
to the knowledge graph’s current state. However, a much more accurate measure should be
developed that takes into account other factors (like the number of times the relationship
has been identified in diferent sources, the presence of the relationship in other sources of
information, or the coherence of the relationship with respect to the other triples extracted
from the same scientific paper). We are also considering the adoption of other link prediction
methodologies, especially those for heterogeneous graphs that can easily scale with big KGs.
Finally, even if the approach has been tested in the context of RNA-KG, we would like to
generalize it to other application domains that exploit biomedical KGs (e.g. [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]) for extracting
new facts from texts.
      </p>
      <p>Datasets. Experiments have been realized by using the following datasets: (schema and docs)
https://doi.org/10.5281/zenodo.10671796; (RNA-KG): https://doi.org/10.5281/zenodo.10078876.
Acknowledgements. This research was in part supported by the “National Center for Gene Therapy
and Drugs based on RNA Technology”, PNRR-NextGeneration EU program [G43C22001320007] and
in part by the MUSA - Multilayered Urban Sustainability Action - Project, funded by the
PNRR-NextGeneration EU program ([G43C22001370007], Code ECS00000037).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>E.</given-names>
            <surname>Cavalleri</surname>
          </string-name>
          , et al.,
          <string-name>
            <surname>RNA-KG</surname>
          </string-name>
          :
          <article-title>An ontology-based knowledge graph for representing interactions involving RNA molecules</article-title>
          ,
          <year>2023</year>
          . arXiv:
          <volume>2312</volume>
          .
          <fpage>00183</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>R.</given-names>
            <surname>Bommasani</surname>
          </string-name>
          , et al.,
          <source>On the opportunities and risks of foundation models</source>
          ,
          <year>2021</year>
          . arXiv:
          <volume>2108</volume>
          .
          <fpage>07258</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A. J.</given-names>
            <surname>Thirunavukarasu</surname>
          </string-name>
          , et al.,
          <article-title>Large language models in medicine</article-title>
          ,
          <source>Nature Medicine</source>
          <volume>29</volume>
          (
          <year>2023</year>
          )
          <fpage>1930</fpage>
          -
          <lpage>1940</lpage>
          . doi:
          <volume>10</volume>
          .1038/s41591-023-02448-8.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Ji</surname>
          </string-name>
          , et al.,
          <article-title>Survey of hallucination in natural language generation</article-title>
          ,
          <source>ACM Computing Surveys</source>
          <volume>55</volume>
          (
          <year>2023</year>
          )
          <fpage>1</fpage>
          -
          <lpage>38</lpage>
          . doi:
          <volume>10</volume>
          .1145/3571730.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ettinger</surname>
          </string-name>
          ,
          <article-title>What BERT is not: Lessons from a new suite of psycholinguistic diagnostics for language models</article-title>
          ,
          <source>Transactions of the Association for Computational Linguistics</source>
          <volume>8</volume>
          (
          <year>2020</year>
          )
          <fpage>34</fpage>
          -
          <lpage>48</lpage>
          . doi:
          <volume>10</volume>
          .1162/tacl_a_
          <fpage>00298</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J. H.</given-names>
            <surname>Caufield</surname>
          </string-name>
          , et al.,
          <article-title>Structured Prompt Interrogation and Recursive Extraction of Semantics (SPIRES): a method for populating knowledge bases using zero-shot learning</article-title>
          ,
          <source>Bioinformatics</source>
          <volume>40</volume>
          (
          <year>2024</year>
          )
          <article-title>btae104</article-title>
          . doi:
          <volume>10</volume>
          .1093/bioinformatics/btae104.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>S.</given-names>
            <surname>Moxon</surname>
          </string-name>
          , et al.,
          <article-title>The Linked Data Modeling Language (LinkML): A General-Purpose Data Modeling Framework Grounded in Machine-Readable Semantics</article-title>
          , in: International Conference on Biomedical Ontologies,
          <year>2021</year>
          , pp.
          <fpage>148</fpage>
          -
          <lpage>151</lpage>
          . URL: https://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>3073</volume>
          /paper24.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>R.</given-names>
            <surname>Jackson</surname>
          </string-name>
          , et al.,
          <string-name>
            <surname>OBO</surname>
          </string-name>
          <article-title>Foundry in 2021: operationalizing open data principles to evaluate ontologies</article-title>
          ,
          <year>Database 2021</year>
          (
          <year>2021</year>
          ). doi:
          <volume>10</volume>
          .1093/database/baab069.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>E.</given-names>
            <surname>Cavalleri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mesiti</surname>
          </string-name>
          ,
          <article-title>On the extraction of meaningful RNA interactions from scientific publications through LLMs and SPIRES</article-title>
          ., In: 8th Int'
          <article-title>l workshop on Data Analytics solutions for Real-LIfe APplications</article-title>
          .,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>P. N.</given-names>
            <surname>Robinson</surname>
          </string-name>
          , et al.,
          <article-title>The human phenotype ontology: A tool for annotating and analyzing human hereditary disease</article-title>
          ,
          <source>The American Journal of Human Genetics</source>
          <volume>83</volume>
          (
          <year>2008</year>
          )
          <fpage>610</fpage>
          -
          <lpage>615</lpage>
          . doi:
          <volume>10</volume>
          .1016/j. ajhg.
          <year>2008</year>
          .
          <volume>09</volume>
          .017.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>N. A.</given-names>
            <surname>Vasilevsky</surname>
          </string-name>
          , et al.,
          <article-title>Mondo: Unifying diseases for the world</article-title>
          , by the world,
          <year>2022</year>
          . doi:
          <volume>10</volume>
          .1101/
          <year>2022</year>
          .04.13.22273750.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M.</given-names>
            <surname>Ashburner</surname>
          </string-name>
          , et al.,
          <article-title>Gene ontology: tool for the unification of biology</article-title>
          ,
          <source>Nature Genetics</source>
          <volume>25</volume>
          (
          <year>2000</year>
          )
          <fpage>25</fpage>
          -
          <lpage>29</lpage>
          . doi:
          <volume>10</volume>
          .1038/75556.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>C.</given-names>
            <surname>Mungall</surname>
          </string-name>
          , et al., oborel/obo-relations:
          <fpage>2023</fpage>
          -08-18 release,
          <year>2023</year>
          . doi:
          <volume>10</volume>
          .5281/zenodo.8263469.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>K.</given-names>
            <surname>Eilbeck</surname>
          </string-name>
          , et al.,
          <article-title>The sequence ontology: a tool for the unification of genome annotations</article-title>
          ,
          <source>Genome Biology</source>
          <volume>6</volume>
          (
          <year>2005</year>
          ). doi:
          <volume>10</volume>
          .1186/gb-2005-6-5-r44.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>E.</given-names>
            <surname>Cavalleri</surname>
          </string-name>
          , et al.,
          <article-title>A meta-graph for the construction of an rna-centered knowledge graph</article-title>
          , in: Bioinformatics and Biomedical Engineering, Springer Nature Switzerland, Cham,
          <year>2023</year>
          , pp.
          <fpage>165</fpage>
          -
          <lpage>180</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>031</fpage>
          -34953-9_
          <fpage>13</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>A.</given-names>
            <surname>Grover</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Leskovec</surname>
          </string-name>
          , node2vec:
          <article-title>Scalable feature learning for networks</article-title>
          ,
          <year>2016</year>
          . arXiv:
          <volume>1607</volume>
          .
          <fpage>00653</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>L.</given-names>
            <surname>Cappelletti</surname>
          </string-name>
          , et al.,
          <article-title>GRAPE for fast and scalable graph processing and random-walk-based embedding</article-title>
          ,
          <source>Nature Computational Science</source>
          <volume>3</volume>
          (
          <year>2023</year>
          )
          <fpage>552</fpage>
          -
          <lpage>568</lpage>
          . doi:
          <volume>10</volume>
          .1038/s43588-023-00465-8.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>H.</given-names>
            <surname>Touvron</surname>
          </string-name>
          , et al.,
          <source>Llama</source>
          <volume>2</volume>
          :
          <article-title>Open foundation and fine-tuned chat models</article-title>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>T.</given-names>
            <surname>Groza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Caufield</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Gration</surname>
          </string-name>
          , G. Baynam,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Haendel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. N.</given-names>
            <surname>Robinson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. J.</given-names>
            <surname>Mungall</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. T.</given-names>
            <surname>Reese</surname>
          </string-name>
          ,
          <article-title>An evaluation of GPT models for phenotype concept recognition</article-title>
          ,
          <year>2023</year>
          . arXiv:
          <volume>2309</volume>
          .
          <fpage>17169</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>J.</given-names>
            <surname>Chen</surname>
          </string-name>
          , et al.,
          <source>RNADisease v4</source>
          .
          <article-title>0: an updated resource of RNA-associated diseases, providing RNA-disease analysis, enrichment and prediction</article-title>
          ,
          <source>Nucleic Acids Research</source>
          <volume>51</volume>
          (
          <year>2022</year>
          )
          <fpage>D1397</fpage>
          -
          <lpage>D1404</lpage>
          . doi:
          <volume>10</volume>
          .1093/nar/gkac814.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>T.J.</given-names>
            <surname>Callahan</surname>
          </string-name>
          , et al.,
          <article-title>An Open-Source Knowledge Graph Ecosystem for the Life Sciences</article-title>
          ,
          <source>Scientific Data</source>
          <volume>11</volume>
          (
          <issue>1</issue>
          ) (
          <year>2024</year>
          )
          <article-title>363</article-title>
          . doi:s41597-
          <fpage>024</fpage>
          -03171-w.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>