<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>LLM-Resilient Bibliometrics: Factual Consistency Through Entity Triplet Extraction⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alexander Sternfeld</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrei Kucharavy</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dimitri Percia David</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alain Mermoud</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Julian Jang-Jaccard</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Cyber-Defence Campus, armasuisse, Science and Technology</institution>
          ,
          <addr-line>Thun</addr-line>
          ,
          <country country="CH">Switzerland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Institute of Entrepreneurship Management, HES-SO Valais-Wallis</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>The increase in power and availability of Large Language Models (LLMs) since late 2022 led to increased concerns with their usage to automate academic paper mills. In turn, this poses a threat to bibliometrics-based technology monitoring and forecasting in rapidly moving fields. We propose to address this issue by leveraging semantic entity triplets. Specifically, we extract factual statements from scientific papers and represent them as (subject, predicate, object) triplets before validating the factual consistency of statements within and between scientific papers. This approach heavily penalizes blind usage of stochastic text generators such as LLMs while not penalizing authors who used LLMs solely to improve the readability of their paper. Here, we present a pipeline to extract such triplets and compare them. While our pipeline is promising and sensitive enough to detect inconsistencies between papers from diferent domains, the intra-paper entity reference resolution needs to be improved to ensure that triplets are more specific. We believe that our pipeline will be useful to the general research community working on the factual consistency of scientific texts.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Bibliometrics</kwd>
        <kwd>Entity Extraction</kwd>
        <kwd>Machine Learning</kwd>
        <kwd>Technological Forecasting</kwd>
        <kwd>Quantum Computing</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        For firms to make informed investment decisions, sound
forecasts on the development of technologies are necessary.
One prominent method for technology forecasting is
bibliometrics, which uses the information in scholarly books
and journals [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Modern bibliometric methods leverage
the increase in available data by applying machine-learning
methods. For example, Percia David et al. (2023) analyse
arXiv pre-prints to evaluate the security development of
information technologies [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        While scientific publications are thus increasingly more
important for technology forecasting, the quality of the
papers must be evaluated critically. The publish-or-perish
pressure led to a record growth in the number of
scientific publications per author, often with minimal peer
review [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ]. In such a setting, if LLMs can generate text that
suficiently resembles a scientific article to pass for one on a
cursory reading, they are likely to be used to generate scores
of articles. Unfortunately, this eventuality is already likely
to be a reality, given that Majovsky et al. (2023) showed
that ChatGPT can create an authentic-looking neurosurgery
scientific article [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        Recently, there has been a growing interest in identifying
text generated by LLMs. As early as 2019, Zellers et al.
showed that a GPT2-like LLM Grover could detect its own
output. However, recent research suggests that, in general,
LLM detectors either do not work or are easy to evade [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ].
Overall, for a minimally competent attacker who wants to
evade detection, LLM detectors cannot be relied upon.
      </p>
      <p>
        Unfortunately, the situation is serious enough for some
of the most reputable providers of proxies of the impact
of scientific articles to have modified their algorithms to
only consider publications adhering to stringent criteria
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Due to the velocity of innovation and the reliance on
preprint repositories, such an approach is not adapted to
technology monitoring in the domains adjacent to
cybersecurity and machine learning. Because of these factors, we
investigate if factual consistency could be used for
LLMresilient bibliometrics instead.
      </p>
      <p>Specifically, we represent facts as entity triplets of the
form (subject, predicate, object) that are extracted from the
claims of the paper. The entity triplet plays a crucial role as
it serves as a proxy to understand the primary claims of the
paper and subsequently validates factual consistency
compared to other works in the domain. Our paper describes the
workflow involved in entity triplet extraction and provides
an overview of our initial findings regarding the
efectiveness of the entity triplets and their relation to the number
of clusters generated around the subject. The code of this
project is available at https://github.com/technometrics-lab/
0-Factual_Consistency_Through_Entity_Triplets, at
commit c7b01e4.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related work</title>
      <p>
        Previous approaches for claim extraction can be categorized
into heuristics and machine learning methods. The
advantage of an approach based on heuristics is that no training
data is required and the computational cost tends to be low.
However, machine learning approaches can capture more
complex patterns, leading to the extraction of triplets of
higher quality. Such methods have been developed most
prominently in the biomedical domain. For example, Li et
al. (2021) use BiLSTMs to extract the factual statements
presented in papers [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Although less labeled data is available,
there has been work focusing on claim extraction from
papers in diferent domains. For instance, Binder et al. (2022)
use BiLSTMs for argumentative discourse unit recognition
and argumentative relation extraction [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <p>
        The majority of existing triplet extraction models use
supervised training. Two notable examples are RECON and
sPERT, which require labeled training data [
        <xref ref-type="bibr" rid="ref11 ref12">11, 12</xref>
        ]. The
disadvantage of supervised methods is the need for training
data and the dependency on the relations that are present in
the dataset. In contrast, unsupervised methods do not need
training data and use either heuristics or machine learning
methods to extract triplets. One example of such a model is
      </p>
      <p>
        Stanford OpenIE, which extractsrelational tuples without
the need to specify a schema in advance. However, it has
been shown that OpenIE tends to extract too aggressively,
resulting in the presence of non-useful relations [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. We
contribute by providing a method that can provide triplets
from scientific papers with a high precision, while the user
only needs to specify the desired research categories.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology</title>
      <p>The process of extracting informative triplets from raw PDFs
consists of four main stages. First, PDFs are converted to
text files, after which they are preprocessed to remove word
breaks and citations, expand abbreviations and lemmatize
words. Then, we extract the sentences from the paper that
convey its core ideas, which we refer to as claims. From
these claims, we can extract subject-predicate-object triplets.
The last step is to process these triplets further so that they
can be used in a comparative analysis. The entire pipeline
is displayed in Figure 1. In the following subsections, we
elaborate on each of the steps. While we focus on arXiv, the
approach is generalizable to all scientific PDFs.</p>
      <sec id="sec-3-1">
        <title>3.1. Preprocessing</title>
        <p>
          To convert the pdf to text we use the PyMuPDF library [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ].
We then further clean the text by removing bracketed
citations and merging words that were split due to line breaks.
We then expand abbreviations by using a rule-based
algorithm introduced by Schwartz and Hearst (2003) [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]. In
appendix 6.1 we show that the Schwartz-Hearst algorithm
outperforms the scispaCy [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] and NLPRe [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] abbreviation
detection methods, which are built on spaCy. Moreover, due
to the rule-based nature of the algorithm, it is relatively
fast. The example below shows the transition from a raw
sentence to a preprocessed result.
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Claim and triplet extraction</title>
        <p>
          After preprocessing the text, we identify the sentences
that convey the authors’ claims. Specifically, we use the
ClaimDistiller framework developed by Wei et al. (2023)
[
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]. In their work, both CNN’s and BiLSTM’s are trained for
claim extraction on the PubMED-RCT and SciARK datasets
[
          <xref ref-type="bibr" rid="ref19 ref20">19, 20</xref>
          ]. Although the usage of supervised contrast training
improves the performance of the model, it causes a
computational overhead. We use the BiLSTM without supervised
contrast learning to strike a balance between performance
and computational eficiency. We choose to extract the
claims from the papers such that the subsequent triplet
extraction will have to be performed on fewer sentences.
        </p>
        <p>
          Next, we want to reduce the claims to (subject,
predicate, object) triplets, analogous to the Resource Description
Framework (RDF) format commonly used in the
representation of OWL ontologies. We choose this representation, as it
will facilitate the comparison of claims across papers. We use
the Python library textacy, which is built on spaCy [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ]
and has a built-in extraction method that does not require
the specification of relations in advance.
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Post-processing of triplets</title>
        <p>Our goal is to compare triplets across papers. Therefore,
we further process the triplets so that we can pair triplets
from diferent papers that refer to the same subject. The
following steps are followed:
1. Lowercase all words in the triplet
2. Remove triplets where either the subject or object
contains more than 6 words
3. Remove stopwords from the triplets based on the
list included in NLTK
4. Remove any character that is not text
5. Lemmatize verbs and nouns in the triplet
6. Remove words containing less than 3 characters
7. Filter the triplets non-specific to scientific work by
comparison with a general book corpus
8. Filter the triplets characteristic of scientific works
in general by comparison with arXiv articles from
diferent categories</p>
        <p>In the second step, we choose this cutof, as we expect
that phrases of over 6 words may contain nuances that
cannot be captured in a simple subject-predicate-object
relation. Words with less than 3 characters are removed, as we
observed that such words were often noise. Moreover, as
abbreviations are expanded we expect all informative terms
to be at least of length 3.</p>
        <p>We use the general-purpose Gutenberg book corpus to
iflter the triplets that carry little information. We define
the number of times term  appears at least 5 times in a
document in the book corpus and in the paper corpus as
, and ,, respectively. We then assign a score  to each
term :
 =
⎪
⎩∞
⎧
⎪−∞
⎨( , ) − ( , )
, &lt; 10
if , ≥ 10 and , &gt; 0
if , = 0 and , ≥ 10</p>
        <p>Terms that are not present in at least 10 papers, thus get
a score of −∞ . If the term is present in at least 10 papers,
the score increases when the frequency of the term in the
book corpus is lower. We keep the triplets with subjects in
the top 10% of the term scores.</p>
        <p>In the last step, we aim to keep only the triplets that
carry domain-specific information. Therefore, we sample a
random subset of 1000 arXiv papers from December 2023
from diferent categories than our target papers. We then
only keep the triplets with subjects present in a maximum
of 15 papers.</p>
      </sec>
      <sec id="sec-3-4">
        <title>3.4. Clustering</title>
        <p>
          As we use the extracted triplets to compare papers, it is
necessary to cluster them based on the subject and object.
Both SciBERT [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] encodings and spaCy [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] embeddings
were considered. Based on visual inspection, the
resulting clusters are most coherent when using SciBERT, which
is a language model based on BERT [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ], pretrained on a
large multi-domain corpus of scientific publications. After
encoding both the subjects and objects, we utilize an
agglomeration hierarchical clustering algorithm from Scipy [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ],
which compares the average distance between clusters. By
visually inspecting the dendograms, we set a cutof to
obtain the final clusters. Figure 5 in the appendix shows an
example of the dendograms for a subset of the subjects and
objects. The threshold is chosen at the height where the
distance between clusters begins to noticeably increase.
        </p>
      </sec>
      <sec id="sec-3-5">
        <title>3.5. Triplet comparison</title>
        <p>
          After clustering, we compare the triplets within the same
cluster based on the predicates. We take a first step in this
direction by analysing embedding inversions, as simple vector
arithmetic can provide valuable insights into word
relationships, such as negation or gender variants [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ]. Specifically,
we subtract the spaCy embeddings of the predicates and
study the tokens closest to the resulting vector. Although
SciBERT encodings likely contain more semantic
information, the input sequence is embedded along with its context,
hence it cannot be easily inverted to a token.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <p>4.1. Data
We consider two diferent datasets to evaluate our method.
First, to leverage in-house expertise in the domain of
computer science and natural language processing (NLP), we
focus on publications relevant to the domain and retrieve
data from the arXiv categories cs.AI, cs.CL and cs.LG.
Specifically, we retrieve all papers from December 2023,
which amounts to a total of 4225 research articles.</p>
      <p>
        Second, for a quantitative analysis of the usefulness of
triplets for factual consistency evaluation, we consider two
surveys. To validate our approach in an independent
domain, we considered both a survey on LLMs and a survey
on Quantum Computing. Specifically, we analyse a survey
on LLMs by Zhao et al. (2019) [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] and a survey on quantum
computing technologies by Gyongyosi and Imre (2019) [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ].
We construct a dataset comprising these two surveys and
the arXiv preprints cited by these surveys. We limit
ourselves to the papers for which an arXiv ID was provided in
the references of the survey, leading to a total of 188 papers.
In the subsequent section, we refer to the papers related to
the LLM survey as the LLM data and to the papers related
to the quantum computing survey as the quantum data.
      </p>
      <sec id="sec-4-1">
        <title>4.2. Cluster analysis</title>
        <sec id="sec-4-1-1">
          <title>4.2.1. CS articles December 2023</title>
          <p>After preprocessing the articles, we extract the triplets. For
the triplet extraction, we used the hyperparameters
displayed in Table 5 in the appendix. In total, 79,986 triplets
were extracted from the research articles. We cluster these
triplets based on both the subject and object embeddings,
which resulted in 37,076 clusters. Figure 3 in the appendix
shows the distribution of the number of triplets per cluster.
This shows that most clusters contain less than 25 triplets,
but that there are outliers that contain over 200 triplets.</p>
        </sec>
        <sec id="sec-4-1-2">
          <title>4.2.2. LLM and quantum computing surveys</title>
          <p>For a more in-depth analysis, we consider the triplets
extracted from the LLM and quantum computing surveys and
their cited papers. In total, 1895 triplets are extracted from
the 188 papers. We cluster these triplets based on the
subjects so we can evaluate the diferences between the objects
in a cluster. Figure 4 in the appendix shows that larger
clusters often have triplets from multiple categories, whereas
small clusters tend to have triplets from only one category.</p>
          <p>Next, we make pairwise comparisons between the object
embeddings within a cluster. Figure 2 shows the pairwise
distances (L2 norm) between objects for each cluster size.
We find that for clusters below size 8, the distance between
objects from the LLM data and quantum data is larger than
the distances of objects within a category. For larger cluster
sizes, this efect disappears. This indicates that for smaller
clusters with triplets from both categories, the objects are
more diverse. Furthermore, we see that for clusters with
triplets from one category, the distance between objects
increases for larger sizes. This confirms that larger clusters
are more domain-agnostic and contain more varied objects.</p>
          <p>Table 2 shows manually selected clusters with sizes 2, 4
and 8. The column with mixed data clusters shows
clusters that contain both triplets from the LLM data and the
quantum data, where the triplets from the quantum data
are displayed in italics. For smaller clusters, the triplets
within a dataset tend to be similar and vary only slightly. In
contrast, the objects difer more for the mixed data clusters
as they are domain-specific. On the other hand, the results
suggest that larger clusters contain more domain-agnostic
subjects and objects, such as lab. Consequently, the distance
between the objects from diferent datasets difers less than
between objects from the same data. Further manual
inspection supports this hypothesis with large clusters containing
subjects such as appendix and conclusion.</p>
          <p>Overall, we argue that it means that triplets as extracted
by our pipelines can be used as proxies for factual
consistency, but that additional refinement is needed to avoid
extracting overly generic statements.</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>4.3. Factual consistency</title>
        <sec id="sec-4-2-1">
          <title>4.3.1. Predicate comparisons</title>
          <p>
            As a first step in evaluating the factual consistency between
papers, triplets in the same cluster are compared based
on the predicates. Specifically, two triplets are considered
consistent when the predicates are synonyms, hypernyms
or hyponyms. If the predicates are antonyms, the triplets
are considered inconsistent. The VerbOcean and WordNet
databases are used to label pairs of predicates [
            <xref ref-type="bibr" rid="ref28 ref29">28, 29</xref>
            ].
          </p>
        </sec>
        <sec id="sec-4-2-2">
          <title>4.3.2. Embedding inversion</title>
          <p>To do a more qualitative assessment, we invert the
diferences of the embeddings of predicates from the same cluster.
Table 6 in the appendix shows a manual selection of 9 of
these embedding inversions. The results show that an
embedding inversion does not provide informative results in
this context. In general, we do not find that there is a
noticeable diference between embedding inversions of predicates
that are consistent or inconsistent.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>This paper presents an unsupervised method for the
extraction of triplets from scientific work. Whereas
previous methods either require labeled training data or the
prespecification of entity relations, we allow for entity triplet
extraction through domain specification.</p>
      <p>The results show that the extracted triplets accurately
reflect the domain from the corresponding scientific work.
When we cluster triplets based on the subject, we find that
smaller clusters tend to be domain-specific. In contrast,
larger clusters are more generic and often contain triplets
from diferent domains. We interpret it as extracted triplets
being suitable for evaluating factual consistency, but
requiring further refinement for a more specific extraction. We
believe this is due to insuficient resolution of excessively
general nouns (e.g. lab, conclusion). To compare the triplets,
an embedding inversion was implemented on the diference
of the verb embeddings for similar triplets. Our findings
show that an embedding inversion does not allow us to
discriminate between consistent and inconsistent triplets.</p>
      <p>Our results suggest that the next steps for the usage of
the extracted triplets for the development of LLM-resilient
proxies should focus on better filtering of domain-agnostic
subjects, for them to be informative about factual
consistency. Then, a semantic network can be built based on the
similarities between the triplets for the entirety of the
scientific publications in a domain of interest. By leveraging
this network, we can identify papers that are factually
inconsistent or excessively consistent and use the remainder
of the corpus for a bibliometric analysis.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Appendix</title>
      <sec id="sec-6-1">
        <title>6.1. Abbreviation detection algorithms</title>
        <p>
          During the preprocessing of papers, we expand abbreviations and map them to their long form. To compare the performance
of diferent abbreviation detection algorithms, we evaluate them on the paper Fundamentals of Generative Large Language
Models and Perspectives in Cyber-Defense, of which we have a thorough understanding [
          <xref ref-type="bibr" rid="ref30">30</xref>
          ].
        </p>
        <p>
          Table 4 shows the performance of the Schwartz-Hearst [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ], scispaCy [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] and NLPRe [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] abbreviation detection methods.
The results show that the Schwartz-Hearst algorithm performs the best, though the scispaCy implementation has a similar
performance. However, the Schwartz-Hearst algorithm is much faster, so we chose this approach.
        </p>
      </sec>
      <sec id="sec-6-2">
        <title>6.2. Triplet extraction</title>
      </sec>
      <sec id="sec-6-3">
        <title>6.3. Triplet clustering</title>
      </sec>
      <sec id="sec-6-4">
        <title>6.4. Embedding inversion</title>
        <p>Original triplets
(example, illustrates, behavior),
(example, mimic, behavior)
(architecture, accomplishes, score),
(architecture, achieves, score)
(type rnns, perform, baseline model),
(type rnns, outperform, baseline model)
(image representation, extract, concept),
(image representation, capture, concept)
(language model, incurs, cost),
(language model, slash, cost)
(text, represents, knowledge),
(text, requires, knowledge)
(subgradient method, not ensure, convergence),
(gradient algorithm, enjoy, convergence)
not ensure</p>
        <p>enjoy
Verb 1
illustrates</p>
        <p>Verb 2
mimic
accomplishes</p>
        <p>achieves
performs</p>
        <p>outperform
extract
incurs
capture
slash
represents
requires
(knowledge transfer, demonstrates, improvement),
(knowledge transfer, not maintain, improvement)
demonstrates</p>
        <p>not maintain</p>
        <p>Top 10 embedding inversions
ILLUSTRATES, ILLUSTRATING,
schematically, EXEMPLARY,
SUMMARIZES, DEPICTS,
EMBODIMENT, ILLUSTRATED,
ILLUSTRATIVE, DESCRIBES
rigamarole, ERRAND, busywork,
AFTERWORDS, canvasing, thigns,
harrasing, forementioned,
explaning, Busy-Work
PERFORMS, PERFORMING,
PERFORMED, PERFORM, SINGS,
CONCERT, ACTs, SONG,
RENDITION, PLAYS
COMPLIANCE, COMPLY, INSUFFICIENT,
IDENTIFIED, ENSURE, AUDIT,
INDICATED, NON-COMPLIANCE,
DETERMINES, IMPROPERLY
EXTRACT, EXTRACTS, DECOCTION,
TINCTURE, GINSENG, TURMERIC,
Comfrey, KOLA, ALOE, Stevia
INCURS,Accrues, ASCERTAINS, INCUR,
INCURRING, incure, howsoever,
INCURRED, internalizes, Indemnified
REPRESENTS, REPRESENTED,
REPRESENTING, RepresENT,
ABSCISSA, symbolises, PERSONIFIES,
DEPICTS, symbolised, Respresents
DEMONSTRATES, demonstates, demostrates,
EXEMPLIFIES, Dissects, Elucidates,
DECONSTRUCTS, ILLUSTRATES,
explicates, EXPLORES
Distance</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. L.</given-names>
            <surname>Porter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Cunningham</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Chiavetta</surname>
          </string-name>
          ,
          <string-name>
            <surname>N.</surname>
          </string-name>
          <article-title>Newman, Parallel or intersecting lines? intelligent bibliometrics for investigating the involvement of data science in policy analysis</article-title>
          ,
          <source>IEEE Transactions on Engineering Management</source>
          <volume>68</volume>
          (
          <year>2020</year>
          )
          <fpage>1259</fpage>
          -
          <lpage>1271</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D. Percia</given-names>
            <surname>David</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Maréchal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Lacube</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gillard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Tsesmelis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Maillart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mermoud</surname>
          </string-name>
          ,
          <article-title>Measuring security development in information technologies: A scientometric framework using arxiv e-prints</article-title>
          ,
          <source>Technological Forecasting and Social Change</source>
          <volume>188</volume>
          (
          <year>2023</year>
          )
          <article-title>122316</article-title>
          . URL: https://www.sciencedirect.com/ science/article/pii/S004016252300001X. doi:https:// doi.org/10.1016/j.techfore.
          <year>2023</year>
          .
          <volume>122316</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Hanson</surname>
          </string-name>
          , P. G. Barreiro,
          <string-name>
            <given-names>P.</given-names>
            <surname>Crosetto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Brockington</surname>
          </string-name>
          ,
          <article-title>The strain on scientific publishing</article-title>
          ,
          <source>CoRR abs/2309</source>
          .15884 (
          <year>2023</year>
          ). URL: https: //doi.org/10.48550/arXiv.2309.15884. doi:
          <volume>10</volume>
          .48550/ ARXIV.2309.15884. arXiv:
          <volume>2309</volume>
          .
          <fpage>15884</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J. P. A.</given-names>
            <surname>Ioannidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Klavans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. W.</given-names>
            <surname>Boyack</surname>
          </string-name>
          ,
          <article-title>Thousands of scientists publish a paper every five days</article-title>
          ,
          <source>Nature</source>
          <volume>561</volume>
          (
          <year>2018</year>
          )
          <fpage>167</fpage>
          -
          <lpage>169</lpage>
          . URL: https://api.semanticscholar.org/ CorpusID:52198631.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Majovsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Černý</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kasal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Komarc</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Netuka</surname>
          </string-name>
          ,
          <article-title>Artificial intelligence can generate fraudulent but authentic-looking scientific medical articles: Pandora's box has been opened</article-title>
          ,
          <source>Journal of Medical Internet Research</source>
          <volume>25</volume>
          (
          <year>2023</year>
          ). doi:
          <volume>10</volume>
          .2196/46924.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>C.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Shu</surname>
          </string-name>
          ,
          <string-name>
            <surname>Can</surname>
          </string-name>
          llm-generated
          <source>misinformation be detected?</source>
          ,
          <year>2023</year>
          . arXiv:
          <volume>2309</volume>
          .
          <fpage>13788</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D. S. G.</given-names>
            <surname>Henrique</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kucharavy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Guerraoui</surname>
          </string-name>
          ,
          <article-title>Stochastic parrots looking for stochastic parrots: Llms are easy to fine-tune and hard to detect with other llms</article-title>
          ,
          <year>2023</year>
          . arXiv:
          <volume>2304</volume>
          .
          <fpage>08968</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Clarivate</surname>
          </string-name>
          ,
          <source>2024 journal citation reports</source>
          , https://clarivate.com/blog/2024-journal
          <article-title>-citationreports-changes-in-journal-impact-factor-categoryrankings-to-enhance-</article-title>
          <string-name>
            <surname>transparency-</surname>
          </string-name>
          and-inclusivity/,
          <year>2024</year>
          . Accessed:
          <fpage>2024</fpage>
          -02-29.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. A.</given-names>
            <surname>Burns</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Peng</surname>
          </string-name>
          ,
          <article-title>Scientific discourse tagging for evidence extraction</article-title>
          , in: P. Merlo,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tiedemann</surname>
          </string-name>
          , R. Tsarfaty (Eds.),
          <source>Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics:</source>
          Main Volume,
          <source>EACL</source>
          <year>2021</year>
          , Online,
          <source>April 19 - 23</source>
          ,
          <year>2021</year>
          , Association for Computational Linguistics,
          <year>2021</year>
          , pp.
          <fpage>2550</fpage>
          -
          <lpage>2562</lpage>
          . URL: https://doi.org/10.18653/v1/
          <year>2021</year>
          .eacl-main.
          <volume>218</volume>
          . doi:
          <volume>10</volume>
          .18653/V1/
          <year>2021</year>
          .EACL-MAIN.
          <year>218</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A.</given-names>
            <surname>Binder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Verma</surname>
          </string-name>
          , L. Hennig,
          <article-title>Full-text argumentation mining on scientific publications</article-title>
          ,
          <source>CoRR abs/2210</source>
          .13084 (
          <year>2022</year>
          ). URL: https: //doi.org/10.48550/arXiv.2210.13084. doi:
          <volume>10</volume>
          .48550/ ARXIV.2210.13084. arXiv:
          <volume>2210</volume>
          .
          <fpage>13084</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>A.</given-names>
            <surname>Bastos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Nadgeri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. O.</given-names>
            <surname>Mulang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Shekarpour</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hofart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kaul</surname>
          </string-name>
          , Recon:
          <article-title>Relation extraction using knowledge graph context in a graph neural network</article-title>
          ,
          <source>in: Proceedings of the Web Conference</source>
          <year>2021</year>
          , WWW '21,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2021</year>
          , p.
          <fpage>1673</fpage>
          -
          <lpage>1685</lpage>
          . URL: https://doi.org/10.1145/ 3442381.3449917. doi:
          <volume>10</volume>
          .1145/3442381.3449917.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M.</given-names>
            <surname>Eberts</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ulges</surname>
          </string-name>
          ,
          <article-title>Span-based joint entity and relation extraction with transformer pre-training</article-title>
          , in: G. D.
          <string-name>
            <surname>Giacomo</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Catalá</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Dilkina</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Milano</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Barro</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Bugarín</surname>
          </string-name>
          , J. Lang (Eds.),
          <source>ECAI 2020 - 24th European Conference on Artificial Intelligence</source>
          ,
          <volume>29</volume>
          <fpage>August</fpage>
          -8
          <source>September</source>
          <year>2020</year>
          , Santiago de Compostela, Spain,
          <source>August 29 - September 8, 2020 - Including 10th Conference on Prestigious Applications of Artificial Intelligence (PAIS</source>
          <year>2020</year>
          ), volume
          <volume>325</volume>
          <source>of Frontiers in Artificial Intelligence and Applications</source>
          , IOS Press,
          <year>2020</year>
          , pp.
          <fpage>2006</fpage>
          -
          <lpage>2013</lpage>
          . URL: https://doi.org/10.3233/FAIA200321. doi:
          <volume>10</volume>
          .3233/FAIA200321.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>L.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Omidvar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Ma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Agrawal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>An</surname>
          </string-name>
          ,
          <article-title>Unsupervised knowledge graph generation using semantic similarity matching</article-title>
          , in: C.
          <string-name>
            <surname>Cherry</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Fan</surname>
            , G. Foster,
            <given-names>G. R.</given-names>
          </string-name>
          <string-name>
            <surname>Hafari</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Khadivi</surname>
            ,
            <given-names>N. V.</given-names>
          </string-name>
          <string-name>
            <surname>Peng</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Ren</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Shareghi</surname>
          </string-name>
          , S. Swayamdipta (Eds.),
          <source>Proceedings of the Third Workshop on Deep Learning for LowResource Natural Language Processing</source>
          , Association for Computational Linguistics, Hybrid,
          <year>2022</year>
          , pp.
          <fpage>169</fpage>
          -
          <lpage>179</lpage>
          . URL: https://aclanthology.org/
          <year>2022</year>
          .deeplo-
          <volume>1</volume>
          .18. doi:
          <volume>10</volume>
          .18653/v1/
          <year>2022</year>
          .deeplo-
          <volume>1</volume>
          .
          <fpage>18</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Artifex</surname>
          </string-name>
          , Pymupdf, https://pypi.org/project/PyMuPDF/,
          <year>2024</year>
          . Accessed:
          <fpage>2024</fpage>
          -02-29.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>A. S.</given-names>
            <surname>Schwartz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Hearst</surname>
          </string-name>
          ,
          <article-title>A simple algorithm for identifying abbreviation definitions in biomedical text</article-title>
          , in: R. B.
          <string-name>
            <surname>Altman</surname>
            ,
            <given-names>A. K.</given-names>
          </string-name>
          <string-name>
            <surname>Dunker</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Hunter</surname>
          </string-name>
          , T. E. Klein (Eds.),
          <source>Proceedings of the 8th Pacific Symposium on Biocomputing, PSB</source>
          <year>2003</year>
          , Lihue, Hawaii, USA, January 3-
          <issue>7</issue>
          ,
          <year>2003</year>
          ,
          <year>2003</year>
          , pp.
          <fpage>451</fpage>
          -
          <lpage>462</lpage>
          . URL: http://psb.stanford.edu/psb-online/proceedings/ psb03/schwartz.pdf .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>M.</given-names>
            <surname>Neumann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>King</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Beltagy</surname>
          </string-name>
          , W. Ammar,
          <article-title>ScispaCy: Fast and robust models for biomedical natural language processing</article-title>
          , in: D.
          <string-name>
            <surname>Demner-Fushman</surname>
            ,
            <given-names>K. B.</given-names>
          </string-name>
          <string-name>
            <surname>Cohen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Ananiadou</surname>
          </string-name>
          , J. Tsujii (Eds.),
          <source>Proceedings of the 18th BioNLP Workshop</source>
          and Shared Task, Association for Computational Linguistics, Florence, Italy,
          <year>2019</year>
          , pp.
          <fpage>319</fpage>
          -
          <lpage>327</lpage>
          . URL: https://aclanthology.org/W19-5034. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>W19</fpage>
          -5034.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>H. B. Travis Hoppe</surname>
          </string-name>
          , Nlpre, https://github.com/ NIHOPA/NLPre,
          <year>2024</year>
          . Accessed:
          <fpage>2024</fpage>
          -04-15.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>X.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. R. U.</given-names>
            <surname>Hoque</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Claimdistiller: Scientific claim extraction with supervised contrastive learning</article-title>
          , in: C. Zhang,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , P. Mayr,
          <string-name>
            <given-names>W.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Suominen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          , Y. Ding (Eds.),
          <source>Proceedings of Joint Workshop of the 4th Extraction</source>
          and
          <article-title>Evaluation of Knowledge Entities from Scientific Documents (EEKE2023) and the 3rd AI + Informetrics (AII2023) co-located with the JCDL 2023</article-title>
          ,
          <string-name>
            <given-names>Santa</given-names>
            <surname>Fe</surname>
          </string-name>
          , New Mexico, USA and Online,
          <volume>26</volume>
          June,
          <year>2023</year>
          , volume
          <volume>3451</volume>
          <source>of CEUR Workshop Proceedings, CEUR-WS.org</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>65</fpage>
          -
          <lpage>77</lpage>
          . URL: https://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>3451</volume>
          /paper11.pdf .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>A.</given-names>
            <surname>Fergadis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Pappas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Karamolegkou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Papageorgiou</surname>
          </string-name>
          ,
          <article-title>Argumentation mining in scientific literature for sustainable development</article-title>
          , in: K.
          <string-name>
            <surname>Al-Khatib</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Hou</surname>
          </string-name>
          , M. Stede (Eds.),
          <source>Proceedings of the 8th Workshop on Argument Mining</source>
          , Association for Computational Linguistics, Punta Cana, Dominican Republic,
          <year>2021</year>
          , pp.
          <fpage>100</fpage>
          -
          <lpage>111</lpage>
          . URL: https://aclanthology.org/
          <year>2021</year>
          .argmining-
          <volume>1</volume>
          .10. doi:
          <volume>10</volume>
          .18653/v1/
          <year>2021</year>
          .argmining-
          <volume>1</volume>
          .
          <fpage>10</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>F.</given-names>
            <surname>Dernoncourt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. Y.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <article-title>PubMed 200k RCT: a dataset for sequential sentence classification in medical abstracts</article-title>
          , in: G. Kondrak, T. Watanabe (Eds.),
          <source>Proceedings of the Eighth International Joint Conference on Natural Language Processing</source>
          (Volume
          <volume>2</volume>
          :
          <string-name>
            <surname>Short</surname>
            <given-names>Papers)</given-names>
          </string-name>
          ,
          <source>Asian Federation of Natural Language Processing</source>
          , Taipei, Taiwan,
          <year>2017</year>
          , pp.
          <fpage>308</fpage>
          -
          <lpage>313</lpage>
          . URL: https://aclanthology.org/I17-2052.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>M.</given-names>
            <surname>Honnibal</surname>
          </string-name>
          , I. Montani,
          <string-name>
            <given-names>S. Van</given-names>
            <surname>Landeghem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Boyd</surname>
          </string-name>
          , spaCy: Industrial-strength
          <source>Natural Language Processing in Python (</source>
          <year>2020</year>
          ). doi:
          <volume>10</volume>
          .5281/zenodo.1212303.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>I.</given-names>
            <surname>Beltagy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lo</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Cohan,
          <article-title>SciBERT: A pretrained language model for scientific text</article-title>
          , in: K. Inui,
          <string-name>
            <given-names>J.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Ng</surname>
          </string-name>
          ,
          <string-name>
            <surname>X.</surname>
          </string-name>
          Wan (Eds.),
          <source>Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLPIJCNLP)</source>
          ,
          <article-title>Association for Computational Linguistics</article-title>
          , Hong Kong, China,
          <year>2019</year>
          , pp.
          <fpage>3615</fpage>
          -
          <lpage>3620</lpage>
          . URL: https:// aclanthology.org/D19-1371. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>D19</fpage>
          - 1371.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          ,
          <article-title>BERT: pre-training of deep bidirectional transformers for language understanding</article-title>
          , in: J.
          <string-name>
            <surname>Burstein</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Doran</surname>
          </string-name>
          , T. Solorio (Eds.),
          <source>Proceedings of the</source>
          <year>2019</year>
          <article-title>Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis</article-title>
          , MN, USA, June 2-7,
          <year>2019</year>
          , Volume
          <volume>1</volume>
          (Long and Short Papers),
          <source>Association for Computational Linguistics</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          . URL: https://doi.org/10.18653/v1/n19-
          <fpage>1423</fpage>
          . doi:
          <volume>10</volume>
          .18653/V1/N19-1423.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>P.</given-names>
            <surname>Virtanen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Gommers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. E.</given-names>
            <surname>Oliphant</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Haberland</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Reddy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Cournapeau</surname>
          </string-name>
          , E. Burovski,
          <string-name>
            <given-names>P.</given-names>
            <surname>Peterson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Weckesser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bright</surname>
          </string-name>
          ,
          <string-name>
            <surname>S. J. van der Walt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Brett</surname>
          </string-name>
          , J. Wilson,
          <string-name>
            <given-names>K. J.</given-names>
            <surname>Millman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Mayorov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. R. J.</given-names>
            <surname>Nelson</surname>
          </string-name>
          , E. Jones,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kern</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Larson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. J.</given-names>
            <surname>Carey</surname>
          </string-name>
          , İ. Polat,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Feng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. W.</given-names>
            <surname>Moore</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. VanderPlas</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Laxalde</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Perktold</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Cimrman</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <string-name>
            <surname>Henriksen</surname>
            ,
            <given-names>E. A.</given-names>
          </string-name>
          <string-name>
            <surname>Quintero</surname>
            ,
            <given-names>C. R.</given-names>
          </string-name>
          <string-name>
            <surname>Harris</surname>
            ,
            <given-names>A. M.</given-names>
          </string-name>
          <string-name>
            <surname>Archibald</surname>
            ,
            <given-names>A. H.</given-names>
          </string-name>
          <string-name>
            <surname>Ribeiro</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Pedregosa</surname>
          </string-name>
          , P. van Mulbregt,
          <source>SciPy 1</source>
          .0 Contributors, SciPy
          <volume>1</volume>
          .
          <article-title>0: Fundamental Algorithms for Scientific Computing in Python</article-title>
          ,
          <source>Nature Methods</source>
          <volume>17</volume>
          (
          <year>2020</year>
          )
          <fpage>261</fpage>
          -
          <lpage>272</lpage>
          . URL: https://doi.org/10.1038/s41592-019-0686-2. doi:
          <volume>10</volume>
          .1038/s41592-019-0686-2.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>K.</given-names>
            <surname>Ethayarajh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Duvenaud</surname>
          </string-name>
          , G. Hirst,
          <article-title>Towards understanding linear word analogies</article-title>
          , in: A.
          <string-name>
            <surname>Korhonen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Traum</surname>
          </string-name>
          , L. Màrquez (Eds.),
          <article-title>Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics</article-title>
          , Florence, Italy,
          <year>2019</year>
          , pp.
          <fpage>3253</fpage>
          -
          <lpage>3262</lpage>
          . URL: https: //aclanthology.org/P19-1315. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>P19</fpage>
          - 1315.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>L.</given-names>
            <surname>Gyongyosi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Imre</surname>
          </string-name>
          ,
          <article-title>A survey on quantum computing technology</article-title>
          ,
          <source>Comput. Sci. Rev</source>
          .
          <volume>31</volume>
          (
          <year>2019</year>
          )
          <fpage>51</fpage>
          -
          <lpage>71</lpage>
          . URL: https://doi.org/10.1016/j.cosrev.
          <year>2018</year>
          .
          <volume>11</volume>
          .002. doi:
          <volume>10</volume>
          .1016/J.COSREV.
          <year>2018</year>
          .
          <volume>11</volume>
          .002.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>W. X.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Hou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Min</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          , P. Liu,
          <string-name>
            <given-names>J.</given-names>
            <surname>Nie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wen</surname>
          </string-name>
          ,
          <article-title>A survey of large language models</article-title>
          ,
          <source>CoRR abs/2303</source>
          .18223 (
          <year>2023</year>
          ). URL: https: //doi.org/10.48550/arXiv.2303.18223. doi:
          <volume>10</volume>
          .48550/ ARXIV.2303.18223. arXiv:
          <volume>2303</volume>
          .
          <fpage>18223</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>G. A.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <article-title>WordNet: A lexical database for english</article-title>
          ,
          <source>Communications of the ACM</source>
          <volume>38</volume>
          (
          <year>1995</year>
          )
          <fpage>39</fpage>
          -
          <lpage>41</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>T.</given-names>
            <surname>Chklovski</surname>
          </string-name>
          , P. Pantel,
          <article-title>VerbOcean: Mining the web for fine-grained semantic verb relations</article-title>
          , in: D.
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          Wu (Eds.),
          <source>Proceedings of the 2004 Conference on Empirical Methods in Natural Language Processing</source>
          , Association for Computational Linguistics, Barcelona, Spain,
          <year>2004</year>
          , pp.
          <fpage>33</fpage>
          -
          <lpage>40</lpage>
          . URL: https://aclanthology.org/ W04-3205.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>A.</given-names>
            <surname>Kucharavy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Schillaci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Maréchal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Würsch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Dolamic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Sabonnadiere</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. P.</given-names>
            <surname>David</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mermoud</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Lenders</surname>
          </string-name>
          ,
          <article-title>Fundamentals of generative large language models and perspectives in cyber-defense</article-title>
          ,
          <year>2023</year>
          . arXiv:
          <volume>2303</volume>
          .
          <fpage>12132</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>