<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Detection from GitHub Leveraging Large Language Models</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Lu Gan</string-name>
          <email>lu.gan@gesis.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Martin Blum</string-name>
          <email>blumma@uni-trier.de</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Danilo Dessí</string-name>
          <email>danilo.dessi@gesis.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Brigitte Mathiak</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ralf Schenkel</string-name>
          <email>schenkel@uni-trier.de</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stefan Dietze</string-name>
          <email>stefan.dietze@gesis.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>. Introduction</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Background</string-name>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Named Entity Recognition, Large Language Model, Knowledge Graph</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>GESIS - Leibniz Institute for the Social Sciences</institution>
          ,
          <addr-line>Köln</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Heinrich Heine University Düsseldorf</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Trier</institution>
          ,
          <addr-line>Trier</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Named entity recognition is an important task when constructing knowledge bases from unstructured data sources. Whereas entity detection methods mostly rely on extensive training data, Large Language Models (LLMs) have paved the way towards approaches that rely on zero-shot learning (ZSL) or few-shot learning (FSL) by taking advantage of the capabilities LLMs acquired during pretraining. Specifically, in very specialized scenarios where large-scale training data is not available, ZSL / FSL opens new opportunities. This paper follows this recent trend and investigates the potential of leveraging Large Language Models (LLMs) in such scenarios to automatically detect datasets and software within textual content from GitHub repositories. While existing methods focused solely on named entities, this study aims to broaden the scope by incorporating resources such as repositories and online hubs where entities are also represented by URLs. The study explores diferent FSL prompt learning approaches to enhance the LLMs' ability to identify datasets and software mentions within repository texts. Through analyses of LLM efectiveness and learning strategies, this paper ofers insights into the potential of advanced language models for automated entity detection.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR</p>
      <p>
        ceur-ws.org
has observed unprecedented advancements in the Natural Language Processing (NLP) field.
However, the exploration for the KG construction task has not been fully explored due to
the huge diversity of domains and usable inputs as sources [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. In fact, despite the recent
advancements, significant limitations persist, especially when it comes to handling complex
concepts unique to specialized domains [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Furthermore, whereas entity detection methods
mostly rely on extensive training data, LLMs have paved the way for approaches that utilize
zero-shot learning (ZSL) or few-shot learning (FSL) by taking advantage of the capabilities LLMs
acquired during pretraining. Thus, in very specialized scenarios where large-scale training data
is not available, ZSL / FSL opens new opportunities. This is especially prominent in scientific
research where a deep understanding of particular entities and their interconnections is required,
while no suficient training data exists.
      </p>
      <p>
        Existing works build KGs either manually or by adopting supervised pipelines on scientific
literature [
        <xref ref-type="bibr" rid="ref9">9, 10</xref>
        ], ignoring resources that are released along with or cited by research papers,
such as source code repositories, datasets, or Machine Learning (ML) models. For example,
ML models are available on hubs such as Huggingface [11] and PyTorch Hub [12], software
source code is found on hosting services such as GitHub [13] and BitBucket [14], and datasets
are provided in repositories such as Zenodo [15]. However, other places exist where code and
datasets can be stored, e.g., they can be released as downloads from servers of institutions or
organizations and, specifically to datasets, they can be linked to repositories where an associated
source code is hosted. This lack of standardization in releasing software and datasets is an
issue in making research reproducible and reusable, thus contributing to the state of the art’s
crisis [16]. Focusing on where datasets are shared raises two main issues: i) datasets shared
in repositories are not easily findable and their reuse is limited, and ii) it is not easy to track
state-of-the-art results on these datasets since they are not referenced uniformly (i.e., they
do not have an associated Digital Object Identifier (DOI) and citations are often embedded in
continuous text or footnotes instead of bibliographies).
      </p>
      <p>To address these issues, in this paper, we present our eforts to automatically discover datasets
and software hidden on README pages in GitHub repositories using Large Language Models,
and describe our experience in extracting them for a later knowledge graph population. This
task diverges from traditional NER approaches for three main reasons: first, comprehending
the content linked to a URL solely from its context poses significant dificulty; second, URLs are
not frequent tokens encountered in vocabulary, leading language models to potentially lack
suficient information for accurate interpretation; third, it is not trivial to retrieve suficient
contexts to infer the intention of URLs from the structured README pages. Specifically, our
study examines the efectiveness of LLMs in two aspects: first, their capability to identify
software and datasets represented by URLs within GitHub repository READMEs, and second,
their performance in classifying URLs found in these repositories. To do so, we present our
analysis in FSL settings due to the lack of available training data. In summary, the contributions
of this paper are: i) an analysis of two LLMs and their quantized models to detect datasets
and software in text from GitHub repositories, ii) an analysis of few-shot prompts to teach an
LLM to detect these mentions, and iii) a manually annotated dataset of 811 GitHub repositories
containing 1,439 URLs and their context. All the resources of this paper are available here.</p>
    </sec>
    <sec id="sec-2">
      <title>2. LLMs Exploration for Software and Dataset Mentions</title>
    </sec>
    <sec id="sec-3">
      <title>Extraction</title>
      <p>This section describes the environment for our analyses, the selected LLMs, the used data, and
the parsing of the LLMs’ output.</p>
      <sec id="sec-3-1">
        <title>2.1. Dataset and Software Mentions Extraction Tasks</title>
        <sec id="sec-3-1-1">
          <title>We investigate the performance of LLMs on two tasks:</title>
          <p>Extraction and Classification (E+CL) Task. The LLM is prompted to extract URLs and
classify them into one of the classes described in Section 2.3. For this task, two prompts are
used. Prompt-1 describes the task and provides four static examples (i.e., provides the same
examples for each request to the LLM), and Prompt-2 describes the task and provides four
dynamic examples (i.e., the provided examples are selected based on their textual similarity to
the context passed to the LLM).</p>
          <p>Classification (CL) Task The LLM is provided with the URL and its context and must classify
them into one of the predefined classes. For this task, two prompts are used. Prompt-3 describes
the task and provides four static examples, and Prompt-4 describes the task and provides four
dynamic examples.</p>
          <p>For both tasks, the LLM is not asked to annotate the provided examples (static and dynamic
ones), i.e., the examples are not used in the evaluation. All four prompt templates can be found
here and we demonstrate one prompt template for E+CL task in Figure 1.</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>2.2. Large Language Models</title>
        <sec id="sec-3-2-1">
          <title>In this section, we describe the LLMs used in our exploration:</title>
          <p>LLaMA 2. LLaMA [17], which stands for “Large Language Model for AI”, represents a series
of foundational models developed by Meta AI. Launched in two iterations, LLaMA-1 (2022)
and LLaMA-2 (2023), it ofers various model sizes (7B to 70B parameters) trained on publicly
available datasets. Notably, LLaMA-2 surpassed GPT-3 on certain benchmarks while promoting
open access with model weights and free use for research and commercial applications.
Mistral 7B [18] is an LLM released in 2023. It exploits grouped-query attention (GQA) [19] and
sliding window attention (SWA) [20, 21] to speed up the inference and reduce memory
requirements during decoding, facilitating the accommodation of larger batch sizes. Furthermore, SWA
enables the management of longer sequences, a common problem of several LLMs. These LLMs
are suitable to handle long prompts needed to instruct the model to perform the required tasks.
LLaMA 2 with quantization &amp; Mistral 7B with quantization. Running the original LLaMA
2 and Mistral models on a GPU demands a relevant amount of vRAM. Since these requirements
currently exceed the hardware available on many research computers, quantization has been
applied to reduce the models’ 16-bit floating-point weights to values between 2 and 8 bits. This
modification enables low-setting machines to run these models while accepting small quality
losses in the expected outcomes.
# Example 2:
Input: {{{ EXAMPLE TEXT }}}
Output: {{{ EXAMPLE OUTPUT JSON }}}
# Example 3:
...
# Example 4:
...
[/INST]
[INST]
# to annotate
Input: {{{INSERT INPUT HERE}}}[/INST]</p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>2.3. Gold Standard Data</title>
        <p>For our study, we use GitHub URLs extracted from the unarXiv [22, 23] dataset as initial seed.
The data and the source code utilized for the extraction process are available here. For each
repository, its README file is parsed to retrieve all outgoing URLs and their surrounding
contexts. The selected sample used to investigate the LLMs’ performance on the mentioned
tasks covers 811 GitHub repositories containing a total of 1,439 URLs in their README files.
Each entry is then manually labeled as one of the following classes:
• Dataset Direct Link. This class includes URLs that point to a dataset file (e.g., JSON, CSV
or TXT files with tabular data containing labels and/or floating-point numbers), archives
(e.g., “.zip” or “.tar.gz”) containing dataset files, or machine learning models (e.g., “.gguf”).
• Dataset Landing Page. This class contains URLs pointing to an index page or directory
that allows the download of one or more dataset files, a software repository that contains
source code generating the dataset or downloading it from an external source (e.g. Google
Drive), or the URL points to a file in a GitHub repository containing a dataset in some
other file in the same repository.
• Software . This class includes URLs that point to software snippets, notebooks, source
code repositories, etc.
• Other. This class includes all the URLs that could not be mapped as Dataset Direct Link,</p>
        <p>Dataset Landing Page, or Software .</p>
        <p>All URLs are manually annotated by researchers from the computer science domain. During the
annotation, it is asked to open the URL in a web browser and to assign only one class. The final
annotation distribution for the classes is: 120 Dataset Direct Link, 678 Dataset Landing Page, 355
Software , 286 Other.</p>
      </sec>
      <sec id="sec-3-4">
        <title>2.4. Output Parsing</title>
        <p>LLMs take our prompts containing instructions as textual input and generate the most probable
response according to the models’ parameters. These create outputs that are not structured and,
consequently, cannot be automatically elaborated on. This might make it unfeasible to apply
LLMs on a large scale. Thus, to enable and facilitate automatic evaluation of LLMs’ generated
outputs, the following post-processing steps are applied:
Remove additional conversational opening phrases. Although the LLMs are instructed to
reply using a JSON object without extra words as the output format, they often include opening
phrases that do not add any value to the required tasks (e.g., “I hope this helps! ”). Therefore,
we apply a set of text replacement rules to all generated outputs, filtering out these opening
phrases.</p>
        <p>Consolidate and convert the structured format string to JSON. Due to the generative
nature, the prompt output string may not present in the requested structured format or may be
incomplete (e.g., instead of returning one JSON array containing JSON objects, it might generate
a newline-separated enumeration of JSON objects). Thus, we consolidate the LLM’s response as
much as possible and then convert the JSON fragments into our designed structured format.
URL Matching. Multiple URLs can appear in the input context presented to the LLMs. To map
the detected URLs with the ground-truth URLs, a 1-to-1 bipartite graph matching solution is
used. More precisely, for each ground-truth URL the most similar unmatched detected URL is
selected. The similarity is based on the longest common substring ratio against the ground-truth
URL. Spurious URLs are matched with the empty set. Despite our efort to extract as much
information from the LLMs’ outputs as possible, it is not always possible to parse them into an
appropriate format for evaluation. We report these parsing statistics for diferent settings in the
evaluation section.</p>
      </sec>
      <sec id="sec-3-5">
        <title>2.5. Evaluation Setting</title>
        <p>We perform the evaluation of the LLMs for the tasks described in Section 2.1 using precision
and recall as defined in [ 24]. More precisely, we compute these scores following this schema:
• Strict. The LLM’s prediction precisely matches the gold standard annotation in both
boundary surface string and entity type.
• Exact. The LLM’s prediction exactly matches the boundary surface URL of the gold
standard annotation, regardless of the entity type.
• Partial. The LLM’s prediction partially overlaps with the boundary surface URL of the
gold standard annotation, regardless of the entity type.
• Type. The LLM’s predicted entity types are matched with the gold standard annotation,
even if the boundary surface string does not fully match.</p>
        <p>The diferent evaluation categories capture various levels of correctness in LLMs’ predictions.
Furthermore, we also analyze a binary setting where the LLMs’ output is only distinguished
between those URLs which refer to datasets (labeled as Dataset Direct Link or Dataset Landing
Page) and those which do not. This simpler task can serve as a baseline and help to better
analyze the potential trade-of between complexity and accuracy.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>3. Results and Discussion</title>
      <p>This section describes the evaluation we conducted to understand how LLMs can be leveraged
to identify dataset and software mentions from GitHub.</p>
      <p>Evaluation on proper output generation. Table 1 reports statistics regarding the number
of generated outputs from the models that correctly adhered to the required input prompt for
parsing. The percentage of generated outputs that were not analyzable ranges between 3.1%
to 14.2%. This issue was particularly prevalent for the Llama models across both static and
dynamic settings, whereas Mistral demonstrated greater stability in producing the expected
outcomes. This discrepancy significantly impacted subsequent analyses, as outputs deviating
from the expected format were considered invalid, negatively afecting the overall models’
performance.</p>
      <p>Evaluation on the E+CL Task. The performance of the LLMs for the E+CL Tasks can be
observed in Table 2. Due to a low ratio (e.g., 18/733 for Llama 2 7b) of parsable outputs in ZSL
setting, which leads to near-zero precision and recall scores, we only report the FSL results.
The LLMs can detect URLs based on the exact and partial results. However, their success rate is
lower than that of known non-LLM-based alternative methods (e.g., regular expressions). This
can be explained by our observation that sometimes one or more of the URLs contained in the
input context are missed by the LLMs, or nonexisting URLs are hallucinated and appended to the
generated output. Among the four LLMs in Table 2, we observe that Mistral 7b outperforms the
other models in strict and type settings, which corresponds to the best classification capability
regardless of taking the URL extraction boundary into account. A generally reduced performance
of the quantized models is visible when compared with the respective original models, except
for the dynamic setting of Llama 2 7b. Also, the quantized Mistral model surpasses Llama 2
7b original in the classification task. Furthermore, in contrast with recent literature [ 25], our
experiments do not show an improvement in the models’ performance when prompted with
dynamic samples. This could indicate that examples need to be carefully selected when feeding
the models to achieve optimal model performance.</p>
      <p>Evaluation on the CL Task. The results of the CL Tasks are reported in Table 3. This task
should be relatively straightforward compared to the E+CL tasks, given that the URL to be
classified was provided as part of the input. Here, in terms of the E+CL task, one of our findings
is that the LLMs are unable to perform the task with reasonable precision. Furthermore, our
investigation reveals a notable challenge: the models frequently struggled to accurately match
the input URL with the URL provided in the context. This results in mismatches between the
two URLs or even the generation of entirely new URLs, thereby reducing the efectiveness of
the richer input context.</p>
      <sec id="sec-4-1">
        <title>3.1. Findings and Limitations</title>
        <p>The experience in testing LLMs in ZSL and FSL for the E+CL and CL tasks can be summarized
as follows:
Limited parsing capabilities. LLMs, due to their generative nature, do not appear suitable to
perform tasks requiring parsing and extraction of complex entities such as the ones subject of
this study. Although LLMs generally can detect URLs, they lack the precision ofered by other
non-LLM-based methods (e.g., regular expression based heuristics) and sometimes hallucinate
non-existing URLs. Additionally, they struggle to classify URLs into similar but nonidentical
classes (in the case of distinguishing dataset direct link and dataset landing page), thus making
their use for certain KG population tasks unfeasible.</p>
        <p>Understanding issue. Sometimes, the models do not understand the request and provide
replies which are not useful or irrelevant for the entity detection task. Examples are “Sure! I’m
ready to annotate the URLs in the input. Please provide the input text.” and “Of course, I can’t
predict the future, and I don’t know what will happen to me or to the world.”.</p>
        <p>More input same results limitation. The fact that a richer input in the CL task did not help
the model to perform better highlights a significant limitation in the models’ ability to properly
integrate and leverage contextual information for niche challenges, emphasizing the need for
further research to enhance their contextual understanding and performance in such tasks.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. Conclusion and Outlook</title>
      <p>In this paper, we investigated two LLMs and their quantized models, that are open to the
scientific community free of charge, for the niche task of aiming at organizing datasets and
software into a KG. Despite the considerable enthusiasm surrounding LLMs, our investigation
reveals a sobering reality: of-the-shelf models are inadequate for addressing intricate tasks
demanding high precision and recall, particularly in the identification and classification of URLs.
Efective organization following semantic web best practices necessitates a level of precision and
recall that these models fail to achieve. Our findings emphasize the importance of tempering
expectations regarding the applicability of LLMs to complex tasks and highlight the need for
further research and development to enhance their suitability for such endeavors.
[10] D. Dessí, F. Osborne, D. Reforgiato Recupero, D. Buscaldi, E. Motta, Cs-kg: A large-scale
knowledge graph of research entities and claims in computer science, in: International
Semantic Web Conference, Springer, 2022, pp. 678–696.
[11] HuggingfaceURL, Hugging Face – The AI community building the future., https://
huggingface.co/, 2024. Accessed: 2024-02-28.
[12] PyTorchHubURL, PyTorch Hub, https://pytorch.org/hub/, 2024. Accessed: 2024-02-28.
[13] GitHubURL, GitHub: Let’s build from here, https://github.com/, 2024. Accessed:
2024-0228.
[14] BitBucketURL, Bitbucket | Git solution for teams using Jira, https://bitbucket.org/, 2024.</p>
      <p>Accessed: 2024-02-28.
[15] ZenodoURL, Zenodo, https://zenodo.org/, 2024. Accessed: 2024-02-28.
[16] M. Ferrari Dacrema, P. Cremonesi, D. Jannach, Are we really making much progress? a
worrying analysis of recent neural recommendation approaches, in: Proceedings of the
13th ACM conference on recommender systems, 2019, pp. 101–109.
[17] H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière,
N. Goyal, E. Hambro, F. Azhar, et al., Llama: Open and eficient foundation language
models, arXiv preprint arXiv:2302.13971 (2023).
[18] A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. d. l. Casas, F. Bressand,
G. Lengyel, G. Lample, L. Saulnier, et al., Mistral 7b, arXiv preprint arXiv:2310.06825
(2023).
[19] J. Ainslie, J. Lee-Thorp, M. de Jong, Y. Zemlyanskiy, F. Lebrón, S. Sanghai, Gqa: Training
generalized multi-query transformer models from multi-head checkpoints, arXiv preprint
arXiv:2305.13245 (2023).
[20] R. Child, S. Gray, A. Radford, I. Sutskever, Generating long sequences with sparse
transformers, arXiv preprint arXiv:1904.10509 (2019).
[21] I. Beltagy, M. E. Peters, A. Cohan, Longformer: The long-document transformer, arXiv
preprint arXiv:2004.05150 (2020).
[22] T. Saier, M. Färber, unarXive: A Large Scholarly Data Set with Publications’ Full-Text,
Annotated In-Text Citations, and Links to Metadata, Scientometrics 125 (2020) 3085–3108.</p>
      <p>URL: https://doi.org/10.1007/s11192-020-03382-z.
[23] T. Saier, M. Färber, unarXive: A Large Scholarly Data Set with Publications’ Full-Text,
Annotated In-Text Citations, and Links to Metadata, 2020. URL: https://doi.org/10.5281/
zenodo.4313164. doi:10.5281/ZENODO.4313164, version 4.
[24] N. Chinchor, B. Sundheim, MUC-5 evaluation metrics, in: Fifth Message Understanding
Conference (MUC-5): Proceedings of a Conference Held in Baltimore, Maryland, August
25-27, 1993, 1993. URL: https://aclanthology.org/M93-1007.
[25] B. Ding, C. Qin, L. Liu, Y. K. Chia, B. Li, S. Joty, L. Bing, Is GPT-3 a good data annotator?,
in: A. Rogers, J. Boyd-Graber, N. Okazaki (Eds.), Proceedings of the 61st Annual Meeting
of the Association for Computational Linguistics (Volume 1: Long Papers), Association
for Computational Linguistics, Toronto, Canada, 2023, pp. 11173–11195. URL: https://
aclanthology.org/2023.acl-long.626. doi:10.18653/v1/2023.acl- long.626.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.</given-names>
            <surname>Verma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Bhatia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Harit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Batish</surname>
          </string-name>
          ,
          <article-title>Scholarly knowledge graphs through structuring scholarly communication: a review</article-title>
          ,
          <source>Complex &amp; Intelligent Systems</source>
          <volume>9</volume>
          (
          <year>2023</year>
          )
          <fpage>1059</fpage>
          -
          <lpage>1095</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>H.</given-names>
            <surname>Ye</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , H. Chen,
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <article-title>Generative knowledge graph construction: A review</article-title>
          ,
          <source>arXiv preprint arXiv:2210.12714</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D.</given-names>
            <surname>Dessí</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Osborne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. R.</given-names>
            <surname>Recupero</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Buscaldi</surname>
          </string-name>
          , E. Motta,
          <article-title>Scicero: A deep learning and nlp approach for generating scientific knowledge graphs in the computer science domain</article-title>
          ,
          <source>Knowledge-Based Systems</source>
          <volume>258</volume>
          (
          <year>2022</year>
          )
          <fpage>109945</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>T.</given-names>
            <surname>Al-Moslmi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. G.</given-names>
            <surname>Ocaña</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. L.</given-names>
            <surname>Opdahl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Veres</surname>
          </string-name>
          ,
          <article-title>Named entity extraction for knowledge graphs: A literature overview</article-title>
          ,
          <source>IEEE Access 8</source>
          (
          <year>2020</year>
          )
          <fpage>32862</fpage>
          -
          <lpage>32881</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>N.</given-names>
            <surname>Milošević</surname>
          </string-name>
          , W. Thielemann,
          <article-title>Comparison of biomedical relationship extraction methods and models for knowledge graph creation</article-title>
          ,
          <source>Journal of Web Semantics</source>
          <volume>75</volume>
          (
          <year>2023</year>
          )
          <fpage>100756</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , J. Kang,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , W. Zhang, H. Chen,
          <article-title>Meta-learning with dynamicmemory-based prototypical network for few-shot event detection</article-title>
          ,
          <source>in: Proceedings of the 13th International Conference on Web Search and Data Mining</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>151</fpage>
          -
          <lpage>159</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Rossi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Barbosa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Firmani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Matinata</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Merialdo</surname>
          </string-name>
          ,
          <article-title>Knowledge graph embedding for link prediction: A comparative analysis</article-title>
          ,
          <source>ACM Transactions on Knowledge Discovery from Data (TKDD) 15</source>
          (
          <year>2021</year>
          )
          <fpage>1</fpage>
          -
          <lpage>49</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S.</given-names>
            <surname>Carta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Giuliani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Piano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. S.</given-names>
            <surname>Podda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Pompianu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. G.</given-names>
            <surname>Tiddia</surname>
          </string-name>
          ,
          <article-title>Iterative zero-shot llm prompting for knowledge graph construction</article-title>
          ,
          <source>arXiv preprint arXiv:2307.01128</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>D.</given-names>
            <surname>Schindler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Zapilko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Krüger</surname>
          </string-name>
          ,
          <article-title>Investigating software usage in the social sciences: A knowledge graph approach</article-title>
          , in: European Semantic Web Conference, Springer,
          <year>2020</year>
          , pp.
          <fpage>271</fpage>
          -
          <lpage>286</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>