<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Enhancing Scholarly Understanding: A Comparison of Knowledge Injection Strategies in Large Language Models</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andrea Cadeddu</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alessandro Chessa</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vincenzo De Leo</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gianni Fenu</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Enrico Motta</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francesco Osborne</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Diego Reforgiato Recupero</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Angelo Salatino</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luca Secchi</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Business and Law, University of Milano Bicocca</institution>
          ,
          <addr-line>Milan</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Mathematics and Computer Science, University of Cagliari</institution>
          ,
          <addr-line>Cagliari</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Knowledge Media Institute, The Open University</institution>
          ,
          <addr-line>Milton Keynes</addr-line>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Linkalab s.r.l.</institution>
          ,
          <addr-line>Cagliari</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The use of transformer-based models like BERT for natural language processing has achieved remarkable performance across multiple domains. However, these models face challenges when dealing with very specialized domains, such as scientific literature. In this paper, we conduct a comprehensive analysis of knowledge injection strategies for transformers in the scientific domain, evaluating four distinct methods for injecting external knowledge into transformers. We assess these strategies in a single-label multi-class classification task involving scientific papers. For this, we develop a public benchmark based on 12k scientific papers from the AIDA knowledge graph, categorized into three fields. We utilize the Computer Science Ontology as our external knowledge source. Our findings indicate that most proposed knowledge injection techniques outperform the BERT baseline.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Knowledge Graphs</kwd>
        <kwd>Natural Language Processing</kwd>
        <kwd>BERT</kwd>
        <kwd>Classification Tasks</kwd>
        <kwd>Feature Engineering</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Transformer models, such as BERT [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and GPT-41, have achieved state-of-the-art performance
in a range of natural language processing tasks. However, despite these advancements, they
still encounter considerable limitations, particularly in dealing with intricate concepts specific
to specialized domains. This issue becomes particularly pronounced in the field of scientific
research, which demands a nuanced grasp of highly specific concepts and their relations [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. A
key challenge in this context is the accurate classification of scientific articles [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. This task is
crucial for structuring and retrieving scientific knowledge, thereby supporting researchers in
staying up-to-date with the latest advancements [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        A prevalent method to enhance the competence of transformers within specific domains
involves continuous pretraining on domain-specific documents [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. However, this approach
poses significant hurdles, particularly due to the need to process a significant volume of
unlabeled, domain-specific text to efectively fine-tune the model parameters [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. To address these
limitations and increase the accuracy of transformers within specialized domains, researchers
have begun to explore the paradigm of knowledge injection [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. These techniques incorporate
external knowledge into transformer models, with the aim of enhancing their comprehension
and, subsequently, their performance in pertinent tasks. They can handle a variety of structured
data, but knowledge graphs (KGs) are now emerging as the prevalent choice [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        In this paper, we analyse four primary strategies for integrating external knowledge into
transformers and evaluate their efectiveness in the task of scientific article classification. For
this purpose, we introduce a new benchmark for scientific article classification using 12k articles
from AIDA KG [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] linked to the Computer Science Ontology (CSO) [
        <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
        ]. Our findings reveal
interesting insights about the eficacy of diferent strategies and their impact on scientific text
classification, with a hybrid approach exploiting BERT and an MLP architecture yielding the
best performance.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        The limitations of transformers prompted the research community to develop several strategies
for knowledge injection [
        <xref ref-type="bibr" rid="ref12 ref13">12, 13</xref>
        ]. Yang et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] proposed a classification for
KnowledgeEnhanced Pre-trained Transformers (KEPTs), based on knowledge granularity, injection method,
and degree of symbolic knowledge parameterization.
      </p>
      <p>
        Having referred exclusively to the classification based on the knowledge injection method
(KIM), this study focuses on two types of KEPT: i) data-structure-unified and ii)
embeddingcombined KEPTs, which are most pertinent in low-resource settings. Data-structure-unified
KEPTs transform knowledge graph (KG) triplets into token sequences, providing a unified
learning algorithm [
        <xref ref-type="bibr" rid="ref14 ref15">14, 15</xref>
        ]. The implementation of this general strategy can vary significantly,
depending on the injection heuristics. Embedding-combined KEPTs use representation learning
algorithms to translate symbolic knowledge into an embedding space, improving resolution
capabilities by combining resultant vectors with additional knowledge vectors [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. In this
paper, we will both consider a simple method based on direct text injection and K-BERT [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], a
more complex technique that controls the visibility of the injected triples to reduce noise.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Background</title>
      <p>This paper investigates various strategies for injecting knowledge into transformer models
and utilizes the task of classifying scientific papers as a use case to evaluate these approaches.
Specifically, we designed diferent versions of a BERT-based text classifier using diferent
strategies for knowledge injection. This process involves training the proposed models to
perform a single-label multi-class classification task. The objective is to correctly identify and
assign the research field of a paper based on its title and abstract.</p>
      <p>To evaluate the classifiers, we created a benchmark of 12K papers split equally between the
ifelds of Artificial Intelligence (AI), Software Engineering (SE), and Human-Computer Interaction
(HCI). As a source of additional knowledge to be integrated, we utilized a KG of research topics,
which was extracted from CSO. The upcoming sections discuss BERT, AIDA KG, and CSO.</p>
      <p>
        BERT [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] is a transformer-based pre-trained model. As a “masked language model”, BERT is
pre-trained by predicting missing words in a sentence, using both the left and right context. The
bidirectional training allows it to capture comprehensive language representations. BERT can
be fine-tuned for a wide range of NLP tasks. In the case of text classification, BERT is extended
by adding a classification layer on top of the pre-trained model. Once the text is tokenized,
special tokens [CLS] and [SEP] are added, segment and position embeddings are assigned, and
the task-specific classification layer is set, the model can be fine-tuned using labelled data.
      </p>
      <p>The Computer Science Ontology (CSO - https://w3id.org/cso) is a large-scale ontology
which provides a comprehensive taxonomy of Computer Science topics, including over 14K
topics. It was generated by applying the Klink-2 algorithm over a corpus of 16M scientific
articles. CSO incorporates three primary semantic relationships: superTopicOf, relatedEquivalent,
and preferentialEquivalent. It is used by academic institutions and commercial organizations for
supporting a range of relevant tasks, including scholarly data exploration, community detection,
document retrieval, and article recommendation.</p>
      <p>
        The Academia/Industry DynAmics KG (AIDA KG - https://w3id.org/aida) [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] is a KG
that describes 21M articles and 8M patents, categorized according to the research topics from
CSO [
        <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
        ] and relevant industrial sectors (e.g., automotive, financial, energy). It was generated
by integrating data from various data sources in this space, such as Microsoft Academic Search,
Dimensions, OpenAlex, DBLP, the Research Organization Registry (ROR), and DBpedia.
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. The Proposed Benchmark</title>
      <p>In this study, we propose a new benchmark for comparing knowledge injection strategies in
scientific paper classification that comprises three components: i) 12K labelled paper abstracts,
ii) a KG of related research topics, and iii) supplementary data to support various KIMs. We used
a venue-based labelling approach to categorize papers into Artificial Intelligence (AI), Software
Engineering (SE), and Human-Computer Interaction (HCI) fields, selecting papers published
after 2010 and with at least three citations. The data set is balanced with 4K papers per category.
The KG for supporting the KIM was generated by extracting the portion of CSO describing the
ifelds of AI, SE, and HCI. This includes 4,629 topics and 9,258 triples.</p>
      <p>The supplementary data enables the KIMs to select the KG portions pertinent to a specific
article. They include a paper-topic map and a specificity score for each topic. The
papertopic map links papers to relevant topics within the KG. The specificity score reflects how
discriminative is a topic for the classification task and is equal to the highest frequency among
the frequencies with which a topic is distributed among the three categories of texts (AI, SE,
and HCI) of the training dataset.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Knowledge Injection Methodologies (KIMs)</title>
      <sec id="sec-5-1">
        <title>This section outlines the four KIMs that have been compared.</title>
        <p>
          Direct Text Injection. Relevant information can be directly integrated into the text, a
process akin to prompt extension [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. This solution appends relevant triples at the end of
the abstracts. Specifically, for each entity linked to the paper, we append two relevant triples.
The triples are modified by converting semantic relationships (e.g., subTopicOf ) into English
expressions (“is a narrower concept than”). The resulting sentences are added to the text.
        </p>
        <p>
          K-BERT. K-BERT [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] is a popular technique for knowledge injection that augments the text
with triples from the KG. It incorporates a Knowledge Layer that identifies the KG entities in
the input text2 and appends to them relevant triples, creating a “sentence tree”. The sentence
tree is processed by the Embedding Layer, assigning positional embeddings, and the Seeing
Layer, filtering noise via the visible matrix. This matrix ensures that injected predicates and
objects only influence the embeddings of the entities they were attached to. The output from
the previous layers is then processed by the Mask-Transformer, which adapts the self-attention
mechanism to accommodate the visible matrix. K-BERT has shown improved performance
compared to BERT in specific domains like finance, medicine, and law [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ].
        </p>
        <p>To use K-BERT for our case, we adapted its implementation3 to process English texts and
incorporate the KG detailed in Section 4. We adjusted the Knowledge Layer to recognize CSO
topics’ surface form in each sentence and append relevant ontology triples.</p>
        <p>
          Integration of Additional Features Using a Multilayer Perceptron. We adapted the
embedding-combined KEPTs method from [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ], which extends BERT with non-textual data
using a multilayer perceptron (MLP). The method uses the BERT model to process the text,
concatenates the resulting embeddings with additional features derived from relevant metadata,
and feeds them into the MLP. We modified the original implementation in [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] to fit our purpose.
We used a standard English BERT model and introduced a component for generating a vector
of features from the KG of research topics. For each article, we selected three topics with
the highest specificity and concatenated them. We then used Sentence BERT (SBERT) [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] to
transform the resulting string into an embedding. Finally, we concatenated this vector with
original text’s embedding and fed it to the MLP, adjusting the final SoftMax layer to give an
output probability for each category.
        </p>
        <p>
          Domain-specific Pre-training. Additional pre-training of BERT on a specific domain can
be seen as a knowledge injection [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. It involves masking selected input tokens and tasking
BERT with predicting them using the surrounding context. The best results are often achieved
by masking about 15% of input tokens. We started with the standard bert-base-uncased model4
and extended its pre-training using text representations of CSO ontology triples as input.
        </p>
      </sec>
      <sec id="sec-5-2">
        <title>2K-BERT uses string match to identify entity labels in the text. 3Available at https://github.com/autoliuweijie/K-BERT 4https://huggingface.co/bert-base-uncased</title>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Evaluation and Conclusions</title>
      <p>
        In this section, we present the performance of both the vanilla BERT and the BERT models
enhanced with the four KIMs, evaluated on the benchmark described in Section 4. In summary,
we compare the following methods: i) BERT, the uncased BERT model trained on text features,
used as baseline; ii) Direct Text Injection (BERT-DTI), which appends additional knowledge at
the end of the input text; iii) K-BERT, which appends additional triples to entities in the text; iv)
Integration of Additional Features Using a MLP (BERT-MLP), which combines the BERT outputs
with additional features; v) Domain-specific Pre-training ( BERT-PT), which extends BERT
pretraining on all triples in the KG. In all experiments, we utilized a balanced test set consisting of
1,500 documents. The size of the training datasets was varied, with trials conducted using 3,000,
6,000, and 9,000 articles, to assess the impact of difering training sizes. In line with the findings
reported in [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], we run each configuration with 10 diferent random seeds. The standard
deviations of the F1-scores were typically below 1%, as detailed in Table 1, demonstrating
consistency. BERT-MLP outperformed the other methods for the largest training size (9K),
yielding an F1-score of 0.880. K-BERT exhibit the best performance for 3K and 6K training
sizes, yielding respectively 0.869 and 0.871. BERT-DTI exhibited marginal improvements over
the vanilla BERT, indicating that even a basic solution that appends knowledge to the text can
yield some benefits. Nevertheless, more sophisticated methods seem to produce much better
results. Finally, BERT-PT performed worse than the standard BERT, likely attributable to the
relatively small knowledge base utilized for pretraining. This outcome may suggest that in
similar cases, KIMs that enhance input could be more efective than pretraining the model on
domain-specific data.
      </p>
      <p>In summary, BERT-MLP and K-BERT seem the best options for this task, with BERT-MLP
showing an advantage when larger training data is available. Future research will broaden the
analysis of KIMs in complex domains. This will involve exploring other fields and use cases
further to understand the potential and limitations of these methods. As future directions, we
will extend this work with more complete evaluation metrics and statistical significance tests;
we will also share a public repository with a full code and examples.</p>
      <p>Train size</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , BERT:
          <article-title>Pre-training of Deep Bidirectional Transformers for Language Understanding</article-title>
          ,
          <source>Naacl-Hlt</source>
          <year>2019</year>
          (
          <year>2018</year>
          )
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          . doi:
          <volume>10</volume>
          . 18653/v1/
          <fpage>N19</fpage>
          -1423.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Alawad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. T.</given-names>
            <surname>Young</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gounley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Schaeferkoetter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. J.</given-names>
            <surname>Yoon</surname>
          </string-name>
          , X.
          <string-name>
            <surname>-C. Wu</surname>
            ,
            <given-names>E. B.</given-names>
          </string-name>
          <string-name>
            <surname>Durbin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Doherty</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Stroup</surname>
          </string-name>
          , et al.,
          <article-title>Limitations of transformers on clinical text classification</article-title>
          ,
          <source>IEEE journal of biomedical and health informatics 25</source>
          (
          <year>2021</year>
          )
          <fpage>3596</fpage>
          -
          <lpage>3607</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.-W.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.-M. Gil</surname>
          </string-name>
          ,
          <article-title>Research paper classification systems based on tf-idf and lda schemes</article-title>
          ,
          <source>Human-centric Computing and Information Sciences</source>
          <volume>9</volume>
          (
          <year>2019</year>
          )
          <fpage>1</fpage>
          -
          <lpage>21</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Salatino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Osborne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Birukou</surname>
          </string-name>
          , E. Motta,
          <article-title>Improving editorial workflow and metadata quality at springer nature, in: The Semantic Web-ISWC 2019: Auckland</article-title>
          , New Zealand,
          <source>October 26-30</source>
          , Springer,
          <year>2019</year>
          , pp.
          <fpage>507</fpage>
          -
          <lpage>525</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>F.</given-names>
            <surname>Barbieri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Camacho-Collados</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. Espinosa</given-names>
            <surname>Anke</surname>
          </string-name>
          , L. Neves,
          <article-title>TweetEval: Unified benchmark and comparative evaluation for tweet classification, in: Findings of the Association for Computational Linguistics</article-title>
          : EMNLP,
          <year>2020</year>
          , pp.
          <fpage>1644</fpage>
          -
          <lpage>1650</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>K. S.</given-names>
            <surname>Kalyan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rajasekharan</surname>
          </string-name>
          , S. Sangeetha,
          <article-title>AMMUS : A survey of transformer-based pretrained models in natural language processing</article-title>
          ,
          <source>CoRR abs/2108</source>
          .05542 (
          <year>2021</year>
          ). URL: https://arxiv.org/abs/2108.05542. arXiv:
          <volume>2108</volume>
          .
          <fpage>05542</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Xiao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Peng</surname>
          </string-name>
          ,
          <article-title>A survey of knowledge enhanced pre-trained models</article-title>
          ,
          <source>CoRR abs/2110</source>
          .00269 (
          <year>2021</year>
          ). URL: https://arxiv.org/abs/ 2110.00269. arXiv:
          <volume>2110</volume>
          .
          <fpage>00269</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>C.</given-names>
            <surname>Peng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Xia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Naseriparsa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Osborne</surname>
          </string-name>
          ,
          <article-title>Knowledge graphs: opportunities and challenges</article-title>
          ,
          <source>Artificial Intelligence Review</source>
          (
          <year>2023</year>
          )
          <fpage>1</fpage>
          -
          <lpage>32</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Angioni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Salatino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Osborne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. R.</given-names>
            <surname>Recupero</surname>
          </string-name>
          , E. Motta,
          <article-title>Aida: A knowledge graph about research dynamics in academia and industry</article-title>
          ,
          <source>QSS</source>
          <volume>2</volume>
          (
          <year>2021</year>
          )
          <fpage>1356</fpage>
          -
          <lpage>1398</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Salatino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Osborne</surname>
          </string-name>
          , E. Motta,
          <article-title>Augur: Forecasting the emergence of new research topics</article-title>
          ,
          <source>in: Proc. of the 18th ACM/IEEE on JCDL, JCDL '18</source>
          , ACM, NY, USA,
          <year>2018</year>
          , p.
          <fpage>303</fpage>
          -
          <lpage>312</lpage>
          . URL: https://doi.org/10.1145/3197026.3197052. doi:
          <volume>10</volume>
          .1145/3197026.3197052.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Salatino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Thanapalasingam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mannocci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Osborne</surname>
          </string-name>
          ,
          <string-name>
            <surname>E. Motta,</surname>
          </string-name>
          <article-title>The computer science ontology: A large-scale taxonomy of research areas</article-title>
          ,
          <source>in: The Semantic Web - ISWC 2018</source>
          , Springer,
          <year>2018</year>
          , pp.
          <fpage>187</fpage>
          -
          <lpage>205</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -00668-6_
          <fpage>12</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Namazifar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Hazarika</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Padmakumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Hakkani-Tür</surname>
          </string-name>
          ,
          <article-title>Kilm: Knowledge injection into encoder-decoder language models</article-title>
          ,
          <year>2023</year>
          . arXiv:
          <volume>2302</volume>
          .
          <fpage>09170</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>D.</given-names>
            <surname>Emelin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Bonadiman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Alqahtani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Mansour,
          <article-title>Injecting domain knowledge in language models for task-oriented dialogue systems</article-title>
          ,
          <year>2022</year>
          . arXiv:
          <volume>2212</volume>
          .
          <fpage>08120</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>P.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Neubig</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Yuan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Fu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Hayashi</surname>
          </string-name>
          , G. Neubig,
          <article-title>Pre-train , Prompt , and Predict : A Systematic Survey of Prompting Methods in Natural Language Processing</article-title>
          ,
          <source>ACM Comput. Surv</source>
          .
          <volume>55</volume>
          (
          <year>2023</year>
          )
          <fpage>1</fpage>
          -
          <lpage>46</lpage>
          . doi:
          <volume>10</volume>
          .1145/3560815. arXiv:
          <volume>2107</volume>
          .
          <year>13586v1</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>W.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Ju</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <surname>K-</surname>
          </string-name>
          <article-title>BERT: enabling language representation with knowledge graph</article-title>
          , CoRR abs/
          <year>1909</year>
          .07606 (
          <year>2019</year>
          ). URL: http://arxiv.org/ abs/
          <year>1909</year>
          .07606. arXiv:
          <year>1909</year>
          .07606.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>M.</given-names>
            <surname>Ostendorf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bourgonje</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Berger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Schneider</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Rehm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Gipp</surname>
          </string-name>
          ,
          <article-title>Enriching BERT with knowledge graph embeddings for document classification</article-title>
          , CoRR abs/
          <year>1909</year>
          .08402 (
          <year>2019</year>
          ). URL: http://arxiv.org/abs/
          <year>1909</year>
          .08402. arXiv:
          <year>1909</year>
          .08402.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>N.</given-names>
            <surname>Reimers</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Gurevych</surname>
          </string-name>
          ,
          <article-title>Sentence-bert: Sentence embeddings using siamese bert-networks</article-title>
          ,
          <year>2019</year>
          , pp.
          <fpage>3973</fpage>
          -
          <lpage>3983</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>D19</fpage>
          -1410.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>J.</given-names>
            <surname>Dodge</surname>
          </string-name>
          , G. Ilharco,
          <string-name>
            <given-names>R.</given-names>
            <surname>Schwartz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Farhadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Hajishirzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. A.</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <article-title>Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping</article-title>
          , CoRR abs/
          <year>2002</year>
          .06305 (
          <year>2020</year>
          ). URL: https://arxiv.org/abs/
          <year>2002</year>
          .06305. arXiv:
          <year>2002</year>
          .06305.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>