<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Qualitative Coding in the Age of AI: An Ontology-Driven Approach</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Daniil Dobriy</string-name>
          <email>daniil.dobriy@wu.ac.at</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Axel Polleres</string-name>
          <email>axel.polleres@wu.ac.at</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Workshop</string-name>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Qualitative Coding, Ontology Engineering, Knowledge Graph Construction, Large Language Models</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Vienna University of Economics and Business</institution>
          ,
          <country country="AT">Austria</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Qualitative coding is an essential methodological tool in qualitative research. Although various tools exist to support manual qualitative coding, the process remains highly resource-intensive, requiring significant time and expertise. It is further complicated by inconsistent reporting of coding protocols, difering interpretations of inter-coder reliability metrics, and dificulties in achieving conceptual agreement among coders - all of which contribute to the broader replication crisis in science. Meanwhile, large language models have demonstrated remarkable abilities to understand context across diverse tasks and are increasingly applied to information and, in combination with semantic web technologies, knowledge extraction. In this work, we define qualitative coding process as a knowledge base construction task and propose an ontology-mediated approach to automating qualitative coding, formalising key aspects of the methodology and reliability assessment using semantic web standards. We evaluate the reliability of this approach through a concrete implementation and case study, finding that qualitative coding can benefit substantially from AI-driven automation - especially when the underlying coding ontology is well-defined and domain-relevant constraints are explicitly codified.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Qualitative coding is an essential methodological tool of qualitative research in which codes
systematically assign descriptors to portions of natural language text or visual data [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. More concisely, it
is “the simple operation of identifying segments of meaning in your data and labelling them with
a code,” [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] where a code is “a word or short phrase that symbolically assigns a summative, salient,
essence-capturing, and/or evocative attribute for a portion of language-based or visual data” [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The
methodology is applied to research data sources, such as interviews and other text excerpts, visual data,
audio recordings, and other unstructured data, to both create various degrees of structured data for
further qualitative and quantitative content analysis, and as a stand-alone method for theory-building,
notably in classical [
        <xref ref-type="bibr" rid="ref1 ref4">1, 4</xref>
        ] and constructivist [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] Grounded Theory.
      </p>
      <p>
        Whenever qualitative coding is intended beyond simple feature extraction, the coding process is
applied iteratively. With each iteration, codes are refined in various ways: redefined, modified, combined,
split, reassigned etc. Through multiple iterations, a better understanding of the underlying data is
sought after, which includes uncovering potential connections and insights. Each coding step can
be deductive (predefined codes or categories are applied to text, in a top-down approach) as well as
inductive (codes are “learned” from text, in a bottom-up approach), and often researchers employ both
approaches interchangeably in the coding process [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>Despite advances in supporting software tools (see Section</p>
      <sec id="sec-1-1">
        <title>2.2), including qualitative data analysis</title>
        <p>
          systems (QDAS), qualitative coding remains heavily dependent on manual labour. The manual
subtasks of qualitative coding demand substantial resources in terms of time and skilled personnel. The
time required for data analysis is a primary issue in qualitative coding, as conceptual tasks associated
with qualitative analysis cannot be easily expedited [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. Furthermore, besides the methodology being
particularly daunting for novice researchers, with available training resources less pervasive than
https://dobriy.org (D. Dobriy); http://polleres.net (A. Polleres)
        </p>
        <p>CEUR</p>
        <p>
          ceur-ws.org
for quantitative methods [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], the requirement for multiple coders to independently process the same
material – a practice necessary for ensuring reliability – significantly multiplies the resource intensity
[
          <xref ref-type="bibr" rid="ref9">9</xref>
          ].
        </p>
        <p>
          Another challenge associated with qualitative research and qualitative coding studies is their
reproducibility. Though qualitative methods often do not easily subject themselves to replication, aiming at
the objectivity of research and striving towards reproducibility shall be nonetheless attempted [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ].
In this context, inter-coder reliability (ICR) is considered a good practice in qualitative research [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ].
For example, in health education and behaviour research, double coding of all transcripts was the
most common coding method, employed in 47.9% of qualitative coding studies, highlighting the broad
acceptance of the collaborative methodology [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. However, achieving a high ICR score becomes ever
more challenging as the conceptual complexity of codes increases [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ].
        </p>
        <p>
          Besides the practical dificulties of achieving conceptual agreement, coding protocols or essential
elements of the process are occasionally simply not reported [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], as for example, 31.3% of qualitative
coding studies examined in health education and behaviour research failed to clearly describe the coding
approach altogether [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. Even when reported, authors often have diferent interpretations of similar
ICR metrics, and the percentage of data used in ICR tests varies across studies [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. Thus, the issues
associated with reproducibility in qualitative coding research contribute to the overall replication crisis
[
          <xref ref-type="bibr" rid="ref13 ref15">15, 13</xref>
          ], most prominently in disciplines relying on qualitative methodologies, such as psychology [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ],
social sciences [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ], healthcare [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] etc. These issues have motivated the standardisation of qualitative
coding practices [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ], including the development of coding reliability frameworks, consensus-based
approaches and collaborative protocols.
        </p>
        <p>
          Recently, pre-trained large language models (LLMs) have demonstrated impressive emergent
capabilities, especially in understanding context across diferent tasks and scenarios. In natural language
processing (NLP), LLMs have achieved state-of-the-art performance across many tasks, becoming the de
facto baseline methods [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]. Accordingly, LLMs are increasingly used for information and knowledge
extraction tasks in various domains [
          <xref ref-type="bibr" rid="ref20 ref21 ref22 ref23">20, 21, 22, 23</xref>
          ]. They have also been proposed for document
information extraction from visually-rich documents [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ], scientific texts [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ] and proceedings [
          <xref ref-type="bibr" rid="ref26 ref27">26, 27</xref>
          ]
and medical descriptions [
          <xref ref-type="bibr" rid="ref28">28</xref>
          ]. In this context, the extraction process commonly employs carefully
crafted prompts to specify both the desired information to be extracted and its structural representation
[
          <xref ref-type="bibr" rid="ref29 ref30">29, 30</xref>
          ].
        </p>
        <p>
          Complementing information extraction, semantic web [
          <xref ref-type="bibr" rid="ref31">31</xref>
          ] provides a universal framework for
knowledge representation, allowing the publication of interconnected data on the web and enabling
machines to process the explicit semantics of data through the use (most notably, reuse) of ontologies,
which, in turn, enable reasoning [
          <xref ref-type="bibr" rid="ref32">32</xref>
          ]. Also notable is the synergetic application of semantic web in
machine learning [
          <xref ref-type="bibr" rid="ref33">33</xref>
          ], where ontologies encode semantics and enable grounding of machine learning
methods and their outputs. Semantic web technologies are especially efective in reducing hallucinations
in LLMs [
          <xref ref-type="bibr" rid="ref34 ref35 ref36 ref37">34, 35, 36, 37</xref>
          ], for example, though knowledge injection in prompts [
          <xref ref-type="bibr" rid="ref38 ref39 ref40 ref41 ref42 ref43 ref44">38, 39, 40, 41, 42, 43, 44</xref>
          ]
and outputs [
          <xref ref-type="bibr" rid="ref45">45</xref>
          ].
        </p>
        <p>This research paper aims to build upon the recent advances in combining LLMs with semantic web
technologies, and defines qualitative coding from the ontology engineering perspective, proposing an
ontology-mediated approach for qualitative coding automation. The approach formalises parts of the
qualitative coding methodology and reliability assessment approaches using semantic web standards.</p>
        <p>This work is structured as follows: Section 2 surveys prior work on AI-assisted qualitative coding and
existing tools; Section 3 introduces the case study used for evaluation; Section 4 details our
ontologymediated approach, including the formalisation of qualitative coding with semantic web standards,
the design of domain-specific and linking ontologies, and the automated workflow; Section 5 presents
the AutoCod implementation; Section 6 outlines the evaluation design, and Section 7 reports results
for both deductive and inductive experiments; Section 8 interprets the findings within qualitative
methodology and ontology engineering, noting limitations and implications; and Section 9 summarises
the contributions and sketches directions for future work.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>This section reviews prior work relevant to the automation of qualitative coding and the software
tools that support it. We first discuss approaches to automating qualitative coding, including both
traditional rule-based methods and recent advances leveraging large language models (Section 2.1). We
then provide an overview of widely used qualitative coding tools together with an analysis of their
features (Section 2.2).</p>
      <sec id="sec-2-1">
        <title>2.1. Qualitative Coding Automation</title>
        <p>
          A number of studies have investigated automated approaches to qualitative coding, particularly for the
initial deductive phase in which a predefined codebook is applied to the data. The majority of existing
tools implement some form of keyword or regular expression matching, typically supplemented by
human validation or human-in-the-loop designs to ensure accuracy. In general, most studies highlight
the importance of the combination of automated and manual coding, as well as the need for transparency
and interpretability of the coding process. Tietz et al. [
          <xref ref-type="bibr" rid="ref46">46</xref>
          ] presented refer, a WordPress-based tool
for semantic annotation, which combines automatic named entity linking and manual annotation,
concluding that the combination of both automated and manual annotation achieve best results.
        </p>
        <p>
          Rietz and Maedche [
          <xref ref-type="bibr" rid="ref47">47</xref>
          ] present Cody, an interactive system that combines editable code rules
(keyword matching with boolean operators) with supervised ML to extend manual codes to seen and
unseen data. The authors report improvements in ICR to simple keyword matching, and note that
coding rules provide structure and transparency. Cai et al. [
          <xref ref-type="bibr" rid="ref48">48</xref>
          ] propose nCoder+, a widely-used tool
for automatic coding of large datasets. The tool relies on regular expressions and manual validation to
establish reliability of codes. Marathe and Toyama [49] conducted a user-centred study of qualitative
coding practices, noting potential for automation after the codebook has been created and highlighting
transparency requirements for automatic coding tools. They then built a prototype for first-pass
semi-automatic coding using simple NLP methods (keyword matching with boolean operators), and
demonstrated high agreement with human coders, noting the potential of advanced NLP techniques
to improve on the result. However, a notable limitation of such tools is their reliance on predefined
keywords or rules, which are infeasible for the coding of complex, under-represented or
counterintuitive concepts. Furthermore, such an approach limits the set of codes to predefined or learned tags,
whereby more complex semi-structured descriptions (e.g., features, relationships) could be required.
        </p>
        <p>Recently, studies have employed LLMs for content analysis in specific domains and identified their
potential to assist in creating coding schemes [50] or in-vivo categories [51]. Other coding automation
approaches have focused on thematic analysis, as e.g., Feinerer and Wild [52], Lennon et al. [53] and
Bryda and Sadowski [54].</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Qualitative Coding Tools</title>
        <p>
          A variety of dedicated software tools have been developed to facilitate qualitative coding and qualitative
data analysis (QDAS). Review studies have analysed the use of such tools in qualitative coding in the
healthcare (incl. medical education) domain [
          <xref ref-type="bibr" rid="ref9">55, 9</xref>
          ]. Table 1 gives an overview of those tools, their use
in studies, and their basic features, including the coding functionality (i.e., assigning codes) and code
aggregation, search and visualisation functions, automatic transcription, collaborative features, support
for multilingual annotation, statistical analysis as well as keyword-based autocoding (automatic code
assignment) capabilities.
        </p>
        <p>The majority of these tools provide core functionality for supporting manual qualitative coding
processes, including document search, aggregation, and coding capabilities. Most tools also ofer data
visualisation features and collaborative coding support. Additional features vary across platforms: some
tools include transcription functionality and multilingual support, while statistical analysis capabilities
are limited to MAXQDA (and underutilised). Keyword-based autocoding functionality is available only
in NVivo, Atlas.ti and MAXQDA. Thus, existing tools primarily streamline manual coding workflows
and facilitate collaboration, with the conceptual and interpretive aspects of qualitative coding remaining
the responsibility of human researchers and coders.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Use Case Study</title>
      <p>To illustrate and evaluate our methodology, we apply it to a published qualitative coding study by van
Gend and Zuiderwijk [56], which employs both deductive and inductive coding phases. In this regard,
we note that any mischaracterisation, omission, or misinterpretation of the original study methodology
or findings remains the sole responsibility of the present authors, and maintain that the chosen study
is an exemplary application of the qualitative coding methodology motivated by open research data
sharing and reuse.</p>
      <p>Van Gend and Zuiderwijk conduct a qualitative single-case study at TU Delft to examine how
combinations of institutional and infrastructural arrangements afect open research data sharing and
reuse in context. Their contributions are twofold [56]: (1) providing a contextualised overview of
institutional and infrastructural arrangements that stimulate open research data sharing and reuse;
and (2) discussing the potential impact of implementing certain arrangements in the case of a Dutch
university pursuing open science.</p>
      <p>As part of the study, the authors collect seven semi-structured interviews. The interviews are
fully transcribed, anonymised, member-checked, and qualitatively coded in ATLAS.ti, with Figure 1
illustrating an interview transcript coded using the tool.</p>
      <p>The case study implements the qualitative coding process in the following way: first, it applies a
theory-driven codebook derived from prior literature in a theory-driven approach [57], followed by
open coding [58] (creating initial categories), axial coding [59, 60] (relating categories into themes,
specifying properties and dimensions), and focused coding [59] (refining salient codes), distinguishing
the codes by actor type (researchers vs. policymakers/support staf). Figure 2 illustrates the qualitative
coding workflow from the case study.</p>
      <p>The analysis then synthesises these codes to identify efective arrangements. For the reference on
the infrastructural and institutional arrangements investigated, we refer the reader to the original paper
and the open data repository containing the interviews as well as the codebook1 (licensed under CC
BY 4.02). While the approach developed in Section 4 is evaluated in comparison to the results of this</p>
      <sec id="sec-3-1">
        <title>1See https://data.4tu.nl/datasets/aa72daa3-6e9a-47a1-b35b-05ade49e16a5/1 2See https://creativecommons.org/licenses/by/4.0/deed</title>
        <p>workflow, it difers from the methodology used by the authors of the case study by conflating the open,
axial and focused coding into a composite inductive coding phase.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Methodology</title>
      <p>This section presents our ontology-mediated approach to automating qualitative coding. We illustrate
aspects of the approach using a case study introduced in Section 3, reproducing a published qualitative
coding analysis to enable comparison with manual coding results. Our methodology encompasses three
key components: (1) the formalisation of qualitative coding concepts using semantic web standards,
(2) the implementation of an automated coding pipeline in the AutoCod tool, and (3) a preliminary
evaluation comparing automated results against human-coded baselines of the case study.</p>
      <sec id="sec-4-1">
        <title>4.1. Ontology-Based Formalisation of Qualitative Coding</title>
        <p>Our ontology-driven approach transforms traditional qualitative coding practices by formalising
codebook elements as semantic web constructs in a domain-specific ontology. Where traditional codebooks
provide linear lists of codes with textual descriptions, ontologies encode this knowledge in a
machinereadable format by minting specific URIs for categories and then including explicit semantics using
semantic web standards. The mapping between codebook elements and ontological constructs follows
established patterns: code names become rdfs:label properties providing human-readable identifiers;
code descriptions and coding guidelines are captured in rdfs:comment annotations that preserve the
interpretive guidance, textual examples, and other contextual information including textual inclusion
and exclusion criteria for the code; hierarchical code structures translate to class hierarchies using
rdfs:subClassOf relationships; besides being included verbatim, inclusion/exclusion criteria are also
formalised as SHACL constraints. This formalisation enables automated validation while maintaining
the natural language contextual explanations underlying the coding schemes, as illustrated in Figure 3
(a) and (b).</p>
        <p>Code Name:
Data Storage Solutions
Description:
Infrastructure for storing
research data including
repositories and archives
Include:
- Institutional repositories
- Cloud storage
- Data archives
semantic web constructs
rdfs:label
“Data Storage Solutions”
rdfs:comment
“Infrastructure for storing
research data including
repositories and archives”
sh:NodeShape
Constraint: must have
:storageType property
Parent Code:
Infrastructural Instruments</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Domain Ontology Creation</title>
        <sec id="sec-4-2-1">
          <title>In our use case, we define a domain-specific ontology</title>
          <p>3 for infrastructural and institutional arrangements.</p>
          <p>This reflects the common theory-based beginning of the qualitative coding approach. The resulting
ontology defines a core class</p>
          <p>Instrument, which is specialised into InfrastructuralInstrument and
InstitutionalInstrument.</p>
          <p>Improving
Usability Of</p>
          <p>Infrastructures
Support Metadata</p>
          <p>Deposit And
Browsing
sharing. The ontology defines a hierarchical class structure using rdfs:subClassOf relationships (blue arrows)
throughout the class hierarchy. In practice, instances of instruments are linked to category classes through the
hasCategory object property.</p>
          <p>Orienting on the literature-based taxonomy introduced in the case study, each branch is organised into
InstrumentCategory classes (e.g., improving usability, legal assistance, selection support, sustainability
for infrastructures; credit/recognition, researcher skills, legal assistance, cost coverage, and delegation
3The domain-specific ontology together with additional materials can be found in the code repository:
https://github.com/semantisch/autocod/tree/main/paper/resources .
of research data management tasks for institutions), under which specific instrument classes (e.g.,
metadata browsing, data quality indicators; data literacy programmes, administrative/financial support)
are defined. Relationships include hasCategory linking instruments to their categories, as well as
textual definitions ( rdfs:comment) specifying scope and intent of each class in line with the theoretical
framework described in the paper. Figure 4 illustrates the resulting class hierarchy.</p>
          <p>In the next step, we refine the ontology with SHACL constraints for downstream validation. The
constraints enforce branch consistency (ensuring instruments are categorised within their appropriate
branch), prevent cross-branch categorisation, and restrict single typing to avoid dual classification
as both infrastructural and institutional. Unlike OWL restrictions such as rdfs:domain/rdfs:range
statements which serve inferential purposes and operate under the open-world assumption, SHACL
constraints provide closed-world validation against data graphs, enabling detection of constraint
violations, that could remain undetected in OWL’s monotonic reasoning framework.</p>
        </sec>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Linking Ontology</title>
        <p>Beyond the domain-specific ontology placing codes in a taxonomy, qualitative coding requires a
linking mechanism to connect identified concepts to their textual evidence in the source material. This
necessitates a separate linking ontology that bridges the semantic representations with the original
data sources.</p>
        <p>“I think what the university library tries to do is to give researchers integrated research support, and they do that with something we call the research cycle.”
:TextSource
:Statement
rdf:Statement</p>
        <p>The linking ontology must satisfy three core requirements: (1) provenance tracking to maintain
traceability between extracted concepts and their source locations within transcripts, (2) evidence
attachment enabling direct quotation and contextual information to support coding decisions, and (3)
entity resolution to handle references to the same real-world entities across diferent textual mentions.
The resulting linking ontology reuses the PROV ontology4 (notably, by specifying :evidencedBy
as the sub-property of the prov:wasDerivedFrom) and employs a reification pattern where coding
statements are modelled as first-class objects with properties linking to both the semantic assertion
(e.g., an instrument instance having a particular category) or resource, and their textual justification
(specific quotes, paragraph references, and contextual metadata from the source interviews). Figure 5
illustrates the resulting linking ontology.</p>
        <sec id="sec-4-3-1">
          <title>4See https://www.w3.org/TR/prov-o/</title>
        </sec>
      </sec>
      <sec id="sec-4-4">
        <title>4.4. Automated Coding Workflow</title>
        <p>In summary, the coding pipeline relies on the two above-mentioned ontologies: a domain-specific
ontology defining coding concepts and significant constraints for validation, and a (generic) linking
ontology that connects textual evidence to semantic statements. Our pipeline operates through an
LLMbased approach where a large language model is prompted to extract domain concepts from transcripts
using only the vocabulary defined in the provided ontologies, and tasked to retrieve well-formed RDF
data. At each LLM-extraction step, RDF syntax as well as constraints are validated, and a retry loop
(including the previous unsuccessful validation report) is implemented in case of errors. The pipeline
implements the deductive and inductive coding phases in the following way (see Figure 6):
Linking
Ontology</p>
        <p>Domain
Ontology</p>
        <p>SHACL
Constraints</p>
        <p>Textual
Source</p>
        <p>Linking
Ontology</p>
        <p>Knowledge Graph
(with coded text)</p>
        <p>Domain
Ontology</p>
        <p>SHACL
Constraints
LLM Processing
(Theory-driven) No</p>
        <p>RDF Output
Knowledge</p>
        <p>Graph</p>
        <p>Yes</p>
        <p>Syntax &amp;
SHACL
Valid?</p>
        <p>LLM Processing
(Pattern Discovery) No</p>
        <p>Ontology
Extensions
Extended
Ontology</p>
        <p>Yes</p>
        <p>Syntax &amp;
SHACL
Valid?
(a) Deductive approach
(b) Inductive approach</p>
        <p>Deductive Coding Phase (Figure 6a): The LLM receives three inputs: (1) the domain ontology in
Turtle format, defining instrument classes, categories, and their relationships; (2) the linking ontology;
and (3) the input data in plain text. The model is constrained (including, through automatic validation) to
instantiate only classes and properties explicitly defined in the domain ontology, ensuring theory-driven
coding aligned with the predefined conceptual framework. Each instantiated concept must be supported
by direct quotes from the transcript, linked using the linking ontology. The output is a well-formed
RDF graph.</p>
        <p>
          Inductive Coding Phase (Figure 6b): In the inductive phase, text excerpts are re-analysed to
discover emergent concepts not captured in the initial ontology using predefined strategies common in
qualitative coding [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] (see Table 2). This mirrors the open coding approach in traditional qualitative
analysis, where new categories emerge from the data. The LLM follows encoded strategies to propose
extensions to the domain ontology based on patterns observed in the coded excerpts, which are then
validated against SHACL constraints before incorporation.
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Pipeline Implementation</title>
      <p>We then implement the ontology-driven qualitative coding workflows described in Section 4.4 as an
automated qualitative coding tool which accepts domain ontologies and textual transcripts as inputs
and utilises large language models to extract concepts constrained by the ontological vocabulary. The
implementation validates extracted RDF against SHACL constraints and maintains links between coded
concepts and their supporting text excerpts through the linking ontology. It processes input sources
in both deductive mode (applying predefined ontological concepts) and inductive mode (discovering
new concepts to extend the ontology), with all outputs forming validated RDF knowledge graphs. More
detailed architectural description of the tool and its documentation are available in the repository.5</p>
      <p>The tool implements a number of supporting features extending the general (deductive and inductive)
approaches described in Section 4.4. It can be used as a web-based application, via API or as a
commandline tool. For larger single datasets, the tool allows definition of separators for chunking as well as
their parallel processing, which is followed by the integration of the resulting chunk-based knowledge
graphs. In the file explorer, the tool highlights the coded excerpts for better overview. All input sources
in a workspace are integrated into the combined coding knowledge graph that can be exported. All
intermediate results, as well as resources including the domain ontology, linking ontology as well as
prompts can be edited at any stage to adjust the coding process. Figure 7 shows the AutoCod web
interface.6</p>
    </sec>
    <sec id="sec-6">
      <title>6. Evaluation</title>
      <p>Our evaluation methodology relies on a comparative analysis between automated and manual coding
results across both deductive and inductive phases:</p>
      <p>For the deductive coding phase, we assess the alignment between excerpts identified by the
automated approach and those coded manually in the original study. We examine the intersection of
coded segments, noting that while the automated approach tends to identify precise text fragments at
the sentence or phrase level, manual coding in the original study typically operated at the paragraph
level, resulting in coarser-grained text segmentation.</p>
      <p>For the inductive coding phase, we evaluate the semantic correspondence between ontology
extensions generated by our automated approach and the emergent categories identified through
manual analysis in the original study. This comparison focuses on conceptual overlap rather than exact
terminological matching, as the automated approach produces formal ontology classes while manual
coding yields informal category labels. Finally, we assess the degree to which automatically generated</p>
      <sec id="sec-6-1">
        <title>5See https://github.com/semantisch/autocod</title>
        <p>6The AutoCod tool and its documentation are available at https://github.com/semantisch/autocod
superclasses and relationships capture the theoretical insights derived from traditional qualitative
analysis in the case study.</p>
        <p>For both phases, we conduct the evaluation using the GPT-5 model (08-08-2025, temperature
hyperparameter not supported), processing transcripts that have been chunked on double newlines, with a
maximum context window of 20,000 tokens per chunk. This ensures that the model can handle long
input documents while maintaining coherence and context.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>7. Results</title>
      <p>As the result of the deductive pass on the seven interview transcripts using the domain-specific ontology
defined in Section 4.2, a combined knowledge graph of 865 triples is created. Table 3 summarises the
number of triples and text excerpts/codes created for each transcript. The resulting generated knowledge
graphs are available in the online repository.7</p>
      <p>While the automated approach selects less codes than the case study, the manual codes are often
overlapping (cf. Figure 1). In terms of the overlap in codes, we calculate the precision and recall of
the automated approach compared to the non-overlapping manual codes. Precision is the number of
correctly identified codes divided by the total number of codes identified by the automated approach.
Recall is the number of correctly identified codes divided by the total number of the non-overlapping
codes in the case study.</p>
      <sec id="sec-7-1">
        <title>7See https://github.com/semantisch/autocod/tree/main/paper/resources</title>
        <p>Interview</p>
        <p>In the next stage, open coding on the selected excerpts is used to extend the ontology itself. Here,
we (1) compare the results from cumulative coding strategies (after one step of inductive coding) to
the codes from the study and (2) analyse the additional identified concepts/diferences. The extended
ontology is available in the online repository.</p>
        <p>The automated inductive coding generated 67 new ontology classes and properties, which we compare
to the 55 codes from the manual codebook. Table 4 shows the mapping between automated concepts
and manual codes:</p>
        <p>Automated Ontology Class
:DataSteward
:Library
:FrontOficeUnit, :BackOficeUnit
:DMPAuthoring
:DatasetCuration, :DataDocumentation
:ProvideEthicsAndPrivacyReviewSupport
:DeliverCarpentriesWorkshops
:ProvideDataRefinementSmallGrants
:RepositorySelection
:PerformSubmissionQualityChecks</p>
        <p>Manual Code(s)
Organisational Structures</p>
        <p>Data Steward [PM/SP],
Data Steward [R]
Library’s role</p>
        <p>Support implementation tips
Activities and Support</p>
        <p>DMP implementation tips
Infrastructure usability
Policy implementation tips
Education
Financial Support
Funds</p>
        <p>Infrastructure
Choosing infrastructures,
Discovering infrastructures</p>
        <p>Infrastructure requirements</p>
        <p>The automated approach created formal class hierarchies implicit in the manual coding
(e.g., :OrganisationalUnit with subclasses :Library, :Faculty, :CentralAdministration).
Relationships between concepts were formalised (e.g., :administeredBy, :supportsActivity,
:hasTargetAudience) whereas manual codes captured these as implicit connections. Explicit classes
for :PhDStudents and :Researchers emerged, while manual codes distinguished perspectives with
[PM/SP] and [R] sufixes.</p>
        <p>Conversely, the manual codebook captured a number of aspects not identified by the automated
approach: 13 codes focusing on experiences with various RDM aspects (e.g., “Experience with DMPs”,
“Experience with infrastructures”), 6 codes addressing motivations for (not) sharing/reusing data,
personal/community-related codes like “Colleagues in department”, “Peer-to-peer”, and “Trust in peers
and the community” as well as direct mentions of 4TU repository and specific implementation details
that were naturally not as prominent in the automatic approach since they primarily represent implicit
knowledge/focus of researchers.</p>
        <p>Thus, the automated approach identified formal structural relationships and created reusable
ontological patterns, achieving approximately 35% direct and partial overlap with manual codes. However,
it noticeably missed subjective, experiential, and motivational aspects which human coders intuitively
capture.</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>8. Discussion</title>
      <p>We have shown that, in a single automated, inductive coding pass, it is possible to identify and formalise
key structural codes within an ontology, replicating the results of inductive coding phases. While
the resulting codes are comparable to the use case (high precision of 85.5%), the automated approach
noticeably lacks in coverage (recall of 68%).</p>
      <p>The relatively high precision achieved in deductive coding is dependent on exact and extensive
rdfs:comment annotations that capture definitions, coding guidelines, inclusion/exclusion criteria.
Although not included in this study, few-shot examples could also have a positive impact on precision.
This suggests that ontology engineers should invest considerable efort in creating detailed, stand-alone
natural language descriptions that go beyond minimal definitions. Also, such a perspective establishes
a new “consumer” of ontologies and of the ontology engineering process – language models that utilise
them as an interface for knowledge extraction, and suggests that future ontology engineering should
explicitly consider human-AI collaboration patterns.</p>
      <p>Another advantage of the semantic web stack is the validation layer which prevents logically
inconsistent coding. In this regard, constraint definition (especially, SHACL) should be considered an essential
component of ontologies that could potentially be utilised for knowledge extraction or knowledge base
construction.</p>
      <p>The recall results could be also attributed to the singular pass of generation. In this regard, the
few-shot, top-k, majority rule, cumulative, LLM-as-a-judge and agentic components are expected to
considerably improve it. Our results also show that subjective and experiential codes are less readily
captured automatically. Addressing these aspects may require additional strategies to guide the coding
process, including human-in-the-loop (HITL) approaches or retrieval-augmented generation (RAG).</p>
      <p>Notable limitations of the performed case study include reliance on only one, closed-source LLM for
generation, potential data leakage (dataset published before the LLM training date), limited nature of the
golden standard used in the case study, basic nature of the precision/recall evaluation without a dedicated
ICR study and lack of an ablation study on the pipeline components or output normalisation via, e.g.,
aggregation, ensemble decoding and canonicalisation. Furthermore, while the iterative constraint
validation is implemented via a SHACL validation engine, the repair process is currently relying solely
on prompting. While such an approach is efective, more advanced repair techniques [ 61] are expected
to improve repair performance.</p>
      <p>In general, the proposed approach reveals synergies between traditionally distinct fields. Qualitative
coding’s emphasis on iterative refinement aligns naturally with ontology evolution and refinement as
part of the ontology engineering methodologies [62]. In this regard, semantic web standards provide
the formalisation that qualitative methods often lack, and the definition of codebooks as ontologies is
well suited for their FAIR publication [63] and reuse in line with Linked Data Principles [64].</p>
      <p>Finally, with regard to the replicability of qualitative coding, the contribution of coding automation
is twofold. On the one hand, the use of automatic coding tools enable automatic documentation and
protocol generation for the coding process, ensuring protocol adherence but also producing detailed
reproducible protocols. On the other hand, the precise definition of the initial configuration (initial
ontology, language model and hyper-parameters, coding pipeline) and the availability of reproducible
protocols enables direct and, potentially, automatic replication and sensitivity analyses with regard to
pipeline configurations, language model hyper-parameters, prompt engineering and ontology variations.</p>
    </sec>
    <sec id="sec-9">
      <title>9. Conclusions</title>
      <p>Our first contribution is to demonstrate that codebooks can be formally defined as and mapped to
ontologies, which enables LLM-supported knowledge extraction. In our use case, we created a domain
ontology directly from the manual codebook and demonstrated its practical utility in automated coding
of the case study interviews.</p>
      <p>Then, we have created pipelines for deductive and inductive ontology-driven and LLM-supported
coding which incorporated validation steps, and implemented the pipelines as an open source automatic
coding tool. As a supporting resource to linking instances and statements in the generated coding
knowledge graph, we have created the minimal linking ontology defining relationships between text
excerpts/codes, extracted instances and statements.</p>
      <p>Following the framework of the case study, we have evaluated the performance of the pipelines on
case study interviews, testing both deductive and inductive automated coding phases and comparing
them to the manually coded gold standard of the case study. We have demonstrated the general
feasibility of ontology-mediated qualitative coding, with potential implications for both qualitative
research methodology and ontology engineering practices. By design, the quality of the resulting
knowledge base depends on the quality of the domain-specific ontology itself, as well as the associated
constraints used for validation and repair of violations.</p>
      <sec id="sec-9-1">
        <title>Future work</title>
        <p>In addition to the need for a more robust pipeline design and a more robust, extensive evaluation
(addressing limitations discussed in Section 8), one clear avenue for future work aimed at
evaluating the reliability of the technique, consequently, facilitating adoption is defining a benchmark for
evaluating automated qualitative coding pipelines, which would include source data, codebooks,
domainspecific ontologies, coding tasks (deductive, inductive, mixed) and coding gold standards, as well as
considering aspects that underpin trust in the coding process, such as transparency, explainability and
interpretability.</p>
        <p>A number of improvements on the coding pipeline mentioned in Section 8 could be implemented
before a comprehensive evaluation on a larger benchmark dataset is attempted. In this context, a
framework for the ICR evaluation of the structured coding results (coding knowledge graph) has to be
established, which will include the definition of the level of reliability testing (triple-level, entity-level,
code-level evaluations), the semantic matching criteria and adaptations to the conventional ICR metrics
(especially, in the context of collaborative/agentic coding with multiple LLMs). Finally, while general
synthesis of the coding results, even in inductive case, remains outside the scope of current work it is
an ambitious and logical future direction for qualitative coding automation.</p>
      </sec>
    </sec>
    <sec id="sec-10">
      <title>Acknowledgments</title>
      <p>This research was funded in whole or in part by the Austrian Science Fund (FWF) 10.55776/COE12.</p>
    </sec>
    <sec id="sec-11">
      <title>Declaration on Generative AI</title>
      <p>During the preparation of this work, the author(s) used Claude 4 Opus in order to: Grammar and
spelling check. After using these tool(s)/service(s), the author(s) reviewed and edited the content as
needed and take(s) full responsibility for the publication’s content.
for improving recall of ncoder coding, in: International Conference on Quantitative Ethnography,
Springer, 2019, pp. 41–54.
[49] M. Marathe, K. Toyama, Semi-automated coding for qualitative research: A user-centered inquiry
and initial prototypes, in: Proceedings of the 2018 CHI conference on human factors in computing
systems, 2018, pp. 1–12.
[50] R. Bijker, S. S. Merkouris, N. A. Dowling, S. N. Rodda, Chatgpt for automated qualitative research:</p>
      <p>Content analysis, Journal of medical Internet research 26 (2024) e59050.
[51] D. Dobriy, M. Beno, A. Polleres, Smw cloud: A corpus of domain-specific knowledge graphs
from semantic mediawikis, in: A. Meroño Peñuela, et al. (Eds.), The Semantic Web. ESWC
2024, volume 14665 of Lecture Notes in Computer Science, Springer, Cham, 2024. URL: https:
//doi.org/10.1007/978-3-031-60635-9_9. doi:10.1007/978- 3- 031- 60635- 9_9.
[52] I. Feinerer, F. Wild, Automated coding of qualitative interviews with latent semantic analysis, in:
Information systems technology and its applications–6th international conference–ISTA 2007,
Gesellschaft für Informatik e. V., 2007, pp. 66–77.
[53] R. P. Lennon, R. Fraleigh, L. J. Van Scoy, A. Keshaviah, X. C. Hu, B. L. Snyder, E. L. Miller, W. A.</p>
      <p>Calo, A. E. Zgierska, C. Grifin, Developing and testing an automated qualitative assistant (aqua)
to support qualitative analysis, Family medicine and community health 9 (2021) e001287.
[54] G. Bryda, D. Sadowski, From words to themes: Ai-powered qualitative data coding and analysis,
in: World conference on qualitative research, Springer, 2024, pp. 309–345.
[55] I. G. Raskind, R. C. Shelton, D. L. Comeau, H. L. F. Cooper, D. M. Grifith, M. C. Kegler, A Review
of Qualitative Data Analysis Practices in Health Education and Health Behavior Research, Health
Education &amp; Behavior 46 (2019) 32–39. URL: https://journals.sagepub.com/doi/10.1177/109019811
8795019. doi:10.1177/1090198118795019.
[56] T. van Gend, A. Zuiderwijk, Open research data: A case study into institutional and infrastructural
arrangements to stimulate open research data sharing and reuse, Journal of Librarianship and
Information Science 55 (2023) 782–797.
[57] J. T. DeCuir-Gunby, P. L. Marshall, A. W. McCulloch, Developing and using a codebook for the
analysis of interview data: An example from a professional development research project, Field
methods 23 (2011) 136–155.
[58] T. R. Lindlof, B. C. Taylor, Qualitative communication research methods, Sage publications, 2017.
[59] K. Charmaz, Constructing grounded theory: A practical guide through qualitative analysis, sage,
2006.
[60] J. Corbin, A. Strauss, Unending work and care: managing chronic illness at home, Jossey-Bass
health series (1988).
[61] T. Pellissier Tanon, C. Bourgaux, F. Suchanek, Learning how to correct a knowledge base from the
edit history, in: The World Wide Web Conference, 2019, pp. 1465–1475.
[62] K. I. Kotis, G. A. Vouros, D. Spiliotopoulos, Ontology engineering methodologies for the evolution
of living and reused ontologies: status, trends, findings and recommendations, The Knowledge
Engineering Review 35 (2020) e4.
[63] M. D. Wilkinson, M. Dumontier, I. J. Aalbersberg, G. Appleton, M. Axton, A. Baak, N. Blomberg,
J.-W. Boiten, L. B. da Silva Santos, P. E. Bourne, et al., The fair guiding principles for scientific data
management and stewardship, Scientific data 3 (2016) 1–9.
[64] C. Bizer, T. Heath, T. Berners-Lee, Linked data: Principles and state of the art, in: World wide web
conference, volume 1, Citeseer, 2008, p. 40.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>B. G.</given-names>
            <surname>Glaser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. L.</given-names>
            <surname>Strauss</surname>
          </string-name>
          ,
          <string-name>
            <surname>E. Strutzel,</surname>
          </string-name>
          <article-title>The discovery of grounded theory; strategies for qualitative research</article-title>
          ,
          <source>Nursing research 17</source>
          (
          <year>1968</year>
          )
          <fpage>364</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Skjott Linneberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Korsgaard</surname>
          </string-name>
          ,
          <article-title>Coding qualitative data: A synthesis guiding the novice</article-title>
          ,
          <source>Qualitative research journal 19</source>
          (
          <year>2019</year>
          )
          <fpage>259</fpage>
          -
          <lpage>270</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Saldaña</surname>
          </string-name>
          ,
          <article-title>The Coding Manual for Qualitative Researchers</article-title>
          , 4th ed., SAGE Publishing Inc.,
          <string-name>
            <surname>Thousand</surname>
            <given-names>Oaks</given-names>
          </string-name>
          , California,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Williams</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Moser</surname>
          </string-name>
          ,
          <article-title>The art of coding and thematic exploration in qualitative research</article-title>
          ,
          <source>International management review 15</source>
          (
          <year>2019</year>
          )
          <fpage>45</fpage>
          -
          <lpage>55</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>K.</given-names>
            <surname>Charmaz</surname>
          </string-name>
          , Constructivist grounded theory,
          <source>The journal of positive psychology 12</source>
          (
          <year>2017</year>
          )
          <fpage>299</fpage>
          -
          <lpage>300</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Ritchie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ormston</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            McNaughton
            <surname>Nicholls</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <article-title>Qualitative research practice: A guide for social science students and researchers (</article-title>
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>C. K.</given-names>
            <surname>Russell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. M.</given-names>
            <surname>Gregory</surname>
          </string-name>
          ,
          <article-title>Issues for consideration when choosing a qualitative data management system</article-title>
          ,
          <source>Journal of Advanced Nursing</source>
          <volume>18</volume>
          (
          <year>1993</year>
          )
          <fpage>1806</fpage>
          -
          <lpage>1816</lpage>
          . URL: https://onlinelibrary.wiley.com/ doi/10.1046/j.1365-
          <fpage>2648</fpage>
          .
          <year>1993</year>
          .
          <volume>18111806</volume>
          .x. doi:
          <volume>10</volume>
          .1046/j.1365-
          <fpage>2648</fpage>
          .
          <year>1993</year>
          .
          <volume>18111806</volume>
          .x.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>E.</given-names>
            <surname>Childs</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. B.</given-names>
            <surname>Demers</surname>
          </string-name>
          ,
          <article-title>Qualitative coding boot camp: an intensive training and overview for clinicians, educators, and administrators</article-title>
          ,
          <source>MedEdPORTAL</source>
          <volume>14</volume>
          (
          <year>2018</year>
          )
          <fpage>10769</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S. O.</given-names>
            <surname>Clarke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. C.</given-names>
            <surname>Coates</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Jordan</surname>
          </string-name>
          ,
          <article-title>A practical guide for conducting qualitative research in medical education: Part 3-Using software for qualitative analysis</article-title>
          ,
          <source>AEM Education and Training</source>
          <volume>5</volume>
          (
          <year>2021</year>
          )
          <article-title>e10644</article-title>
          . URL: https://onlinelibrary.wiley.com/doi/10.1002/aet2.10644. doi:
          <volume>10</volume>
          .1002/aet2.
          <fpage>10644</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>P.</given-names>
            <surname>TalkadSukumar</surname>
          </string-name>
          , R. Metoyer,
          <article-title>Replication and transparency of qualitative research from a constructivist perspective</article-title>
          ,
          <source>OSF Preprints</source>
          (
          <year>2019</year>
          ). URL: https://osf.io/6efvp/.
          <source>doi:10.31219/osf.i o/6efvp, accessed March 6</source>
          ,
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>C. O'Connor</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Jofe</surname>
            , Intercoder Reliability in Qualitative Research: Debates and
            <given-names>Practical</given-names>
          </string-name>
          <string-name>
            <surname>Guidelines</surname>
          </string-name>
          ,
          <source>International Journal of Qualitative Methods</source>
          <volume>19</volume>
          (
          <year>2020</year>
          )
          <article-title>1609406919899220</article-title>
          . URL: https://journals.sagepub.com/doi/10.1177/1609406919899220. doi:
          <volume>10</volume>
          .1177/1609406919899220.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>I. G.</given-names>
            <surname>Raskind</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. C.</given-names>
            <surname>Shelton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. L.</given-names>
            <surname>Comeau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. L.</given-names>
            <surname>Cooper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. M.</given-names>
            <surname>Grifith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. C.</given-names>
            <surname>Kegler</surname>
          </string-name>
          ,
          <article-title>A review of qualitative data analysis practices in health education and health behavior research</article-title>
          ,
          <source>Health Education &amp; Behavior</source>
          <volume>46</volume>
          (
          <year>2019</year>
          )
          <fpage>32</fpage>
          -
          <lpage>39</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>N. C.</given-names>
            <surname>Nelson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ichikawa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chung</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>M. Malik, Mapping the discursive dimensions of the reproducibility crisis: A mixed methods analysis</article-title>
          ,
          <source>PLOS ONE 16</source>
          (
          <year>2021</year>
          )
          <article-title>e0254090</article-title>
          . URL: https: //dx.plos.
          <source>org/10</source>
          .1371/journal.pone.0254090. doi:
          <volume>10</volume>
          .1371/journal.pone.
          <volume>0254090</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>K. K. C. Cheung</surname>
            ,
            <given-names>K. W. H.</given-names>
          </string-name>
          <string-name>
            <surname>Tai</surname>
          </string-name>
          ,
          <article-title>The use of intercoder reliability in qualitative interview data analysis in science education</article-title>
          ,
          <source>Research in Science &amp; Technological Education</source>
          <volume>41</volume>
          (
          <year>2023</year>
          )
          <fpage>1155</fpage>
          -
          <lpage>1175</lpage>
          . URL: https://www.tandfonline.com/doi/full/10.1080/02635143.
          <year>2021</year>
          .
          <volume>1993179</volume>
          . doi:
          <volume>10</volume>
          .1080/02635143.2
          <fpage>021</fpage>
          .1993179.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M.</given-names>
            <surname>Baker</surname>
          </string-name>
          ,
          <volume>1</volume>
          ,
          <string-name>
            <surname>500</surname>
          </string-name>
          <article-title>scientists lift the lid on reproducibility</article-title>
          ,
          <source>Nature</source>
          <volume>533</volume>
          (
          <year>2016</year>
          )
          <fpage>452</fpage>
          -
          <lpage>454</lpage>
          . URL: https: //www.nature.com/articles/533452a. doi:
          <volume>10</volume>
          .1038/533452a.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>Open</given-names>
            <surname>Science</surname>
          </string-name>
          <string-name>
            <surname>Collaboration</surname>
          </string-name>
          ,
          <article-title>Estimating the reproducibility of psychological science</article-title>
          ,
          <source>Science</source>
          <volume>349</volume>
          (
          <year>2015</year>
          )
          <article-title>aac4716</article-title>
          . URL: https://www.science.org/doi/10.1126/science.aac4716. doi:
          <volume>10</volume>
          .1126/science. aac4716.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>C. F.</given-names>
            <surname>Camerer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dreber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Holzmeister</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.-H.</given-names>
            <surname>Ho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Huber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Johannesson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kirchler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Nave</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. A.</given-names>
            <surname>Nosek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Pfeifer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Altmejd</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Buttrick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Chan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Forsell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gampa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Heikensten</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Hummer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Imai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Isaksson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Manfredi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Rose</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.-J.</given-names>
            <surname>Wagenmakers</surname>
          </string-name>
          , H. Wu,
          <article-title>Evaluating the replicability of social science experiments in Nature and Science between 2010 and 2015</article-title>
          ,
          <source>Nature Human Behaviour</source>
          <volume>2</volume>
          (
          <year>2018</year>
          )
          <fpage>637</fpage>
          -
          <lpage>644</lpage>
          . URL: https://www.nature.com/articles/s41562-018-0399-z. doi:
          <volume>10</volume>
          .1038/s41562- 018- 0399- z.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>A.</given-names>
            <surname>Chandrasekar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. E.</given-names>
            <surname>Clark</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Martin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Vanderslott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. C.</given-names>
            <surname>Flores</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Aceituno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Barnett</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Vindrola-Padros</surname>
          </string-name>
          ,
          <string-name>
            <surname>N.</surname>
          </string-name>
          <article-title>Vera San Juan, Making the most of big qualitative datasets: a living systematic review of analysis methods, Frontiers in big Data 7 (</article-title>
          <year>2024</year>
          )
          <fpage>1455399</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>A.</given-names>
            <surname>Zubiaga</surname>
          </string-name>
          ,
          <article-title>Natural language processing in the era of large language models</article-title>
          ,
          <source>Frontiers in Artificial Intelligence</source>
          <volume>6</volume>
          (
          <year>2024</year>
          )
          <article-title>1350306</article-title>
          . URL: https://www.frontiersin.org/articles/10.3389/frai.
          <year>2023</year>
          .1350306 /full. doi:
          <volume>10</volume>
          .3389/frai.
          <year>2023</year>
          .
          <volume>1350306</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>R.</given-names>
            <surname>Han</surname>
          </string-name>
          ,
          <string-name>
            <surname>C</surname>
          </string-name>
          . Yang,
          <string-name>
            <given-names>T.</given-names>
            <surname>Peng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Tiwari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wan</surname>
          </string-name>
          , L. Liu,
          <string-name>
            <given-names>B.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>An empirical study on information extraction using large language models</article-title>
          ,
          <source>arXiv preprint arXiv:2409.00369</source>
          (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>D.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Peng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          , E. Chen,
          <source>Large Language Models for Generative Information Extraction: A Survey</source>
          ,
          <year>2024</year>
          . URL: http://arxiv.org/ab s/2312.17617. doi:
          <volume>10</volume>
          .48550/arXiv.2312.17617, arXiv:
          <fpage>2312</fpage>
          .17617 [cs].
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>X.-Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S.
          <string-name>
            <surname>-M. Cai</surname>
            ,
            <given-names>X.-R.</given-names>
          </string-name>
          <string-name>
            <surname>Shen</surname>
            , Y. Han,
            <given-names>W.-H.</given-names>
          </string-name>
          <string-name>
            <surname>Hu</surname>
            ,
            <given-names>Y.-R.</given-names>
          </string-name>
          <string-name>
            <surname>Zhang</surname>
          </string-name>
          ,
          <source>Eficient Unified Information Extraction Model Based on Large Language Models</source>
          ,
          <year>2024</year>
          . URL: https://www.ssrn.com/abstract= 5053609. doi:
          <volume>10</volume>
          .2139/ssrn.5053609.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>F.</given-names>
            <surname>Polat</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Tiddi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Groth</surname>
          </string-name>
          ,
          <article-title>Testing prompt engineering methods for knowledge extraction from text</article-title>
          , Semantic
          <string-name>
            <surname>Web</surname>
          </string-name>
          (
          <year>2024</year>
          )
          <fpage>1</fpage>
          -
          <lpage>34</lpage>
          . URL: https://www.medra.org/servlet/aliasResolver?alias=iospress &amp;doi=10.3233/SW-243719. doi:
          <volume>10</volume>
          .3233/SW-243719.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>V.</given-names>
            <surname>Perot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Luisier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Su</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. S.</given-names>
            <surname>Boppana</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , C.-
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Hua</surname>
          </string-name>
          , LMDX:
          <article-title>Language model-based document information extraction and localization</article-title>
          , in: L.
          <string-name>
            <surname>-W. Ku</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Martins</surname>
          </string-name>
          , V. Srikumar (Eds.),
          <source>Findings of the Association for Computational Linguistics: ACL</source>
          <year>2024</year>
          ,
          <article-title>Association for Computational Linguistics</article-title>
          , Bangkok, Thailand,
          <year>2024</year>
          , pp.
          <fpage>15140</fpage>
          -
          <lpage>15168</lpage>
          . URL: https://aclanthology.org/
          <year>2024</year>
          .findings-acl.
          <volume>899</volume>
          /. doi:
          <volume>10</volume>
          .18653/v1/
          <year>2024</year>
          .findings-acl.
          <volume>8</volume>
          <fpage>99</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>J.</given-names>
            <surname>Dagdelen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dunn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Walker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. S.</given-names>
            <surname>Rosen</surname>
          </string-name>
          , G. Ceder,
          <string-name>
            <given-names>K. A.</given-names>
            <surname>Persson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jain</surname>
          </string-name>
          ,
          <article-title>Structured information extraction from scientific text with large language models</article-title>
          ,
          <source>Nature Communications</source>
          <volume>15</volume>
          (
          <year>2024</year>
          )
          <fpage>1418</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>N.</given-names>
            <surname>Mihindukulasooriya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Tiwari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Dobriy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Å. Nielsen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. R.</given-names>
            <surname>Chhetri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Polleres</surname>
          </string-name>
          ,
          <article-title>Scholarly wikidata: Population and exploration of conference data in wikidata using llms</article-title>
          ,
          <source>in: International Conference on Knowledge Engineering and Knowledge Management</source>
          , Springer,
          <year>2024</year>
          , pp.
          <fpage>243</fpage>
          -
          <lpage>259</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>D.</given-names>
            <surname>Dobriy</surname>
          </string-name>
          ,
          <article-title>Employing rag to create a conference knowledge graph from text</article-title>
          , in: ESWC'
          <year>2024</year>
          :
          <article-title>The 21st Extended Semantic Web Conference</article-title>
          , Hersonissos, Greece,
          <year>2024</year>
          . URL: https://ceur-ws.org/Vo l-
          <volume>3747</volume>
          /text2kg_paper4.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>R.</given-names>
            <surname>Alharbi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.</given-names>
            <surname>Ahmed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Dobriy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Łajewska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Menotti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Saeedizade</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Dumontier, Exploring the role of generative ai in constructing knowledge graphs for drug indications with medical context, 15th International Semantic Web Applications and Tools for Healthcare and Life Sciences (SWAT4HCLS</article-title>
          <year>2024</year>
          )
          <article-title>(</article-title>
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vijayan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A Prompt</given-names>
            <surname>Engineering</surname>
          </string-name>
          <article-title>Approach for Structured Data Extraction from Unstructured Text Using Conversational LLMs</article-title>
          ,
          <source>in: 2023 6th International Conference on Algorithms, Computing and Artificial Intelligence</source>
          ,
          <string-name>
            <given-names>ACM</given-names>
            ,
            <surname>Sanya</surname>
          </string-name>
          <string-name>
            <surname>China</surname>
          </string-name>
          ,
          <year>2023</year>
          , pp.
          <fpage>183</fpage>
          -
          <lpage>189</lpage>
          . URL: https://dl.acm.org/doi/10. 1145/3639631.3639663. doi:
          <volume>10</volume>
          .1145/3639631.3639663.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>M.</given-names>
            <surname>Moundas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>White</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. C.</given-names>
            <surname>Schmidt</surname>
          </string-name>
          ,
          <article-title>Prompt patterns for structured data extraction from unstructured text</article-title>
          ,
          <source>in: Proceedings of the 31st Conference on Pattern Languages of Programs</source>
          , People, and
          <string-name>
            <surname>Practices</surname>
          </string-name>
          (PLoP
          <year>2024</year>
          ), ACM,
          <year>2024</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>15</lpage>
          . URL: https://www.cs.wm.edu/~dcschmidt /PDF/Prompt_Patterns_for_Structured_Data_Extraction_from_Unstructured_Text___Final.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>T.</given-names>
            <surname>Berners-Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hendler</surname>
          </string-name>
          ,
          <string-name>
            <surname>O. Lassila,</surname>
          </string-name>
          <article-title>The semantic web: A new form of web content that is meaningful to computers will unleash a revolution of new possibilities</article-title>
          ,
          <source>in: Linking the World's Information: Essays on Tim Berners-Lee's Invention of the World Wide Web</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>91</fpage>
          -
          <lpage>103</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>A.</given-names>
            <surname>Scherp</surname>
          </string-name>
          , G. Groener,
          <string-name>
            <given-names>P.</given-names>
            <surname>Škoda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Hose</surname>
          </string-name>
          , M.-E. Vidal, Semantic Web: Past, Present, and
          <article-title>Future, Transactions on Graph Data and Knowledge (TGDK) 2 (</article-title>
          <year>2024</year>
          ) 3:
          <fpage>1</fpage>
          -
          <lpage>3</lpage>
          :
          <fpage>37</fpage>
          . URL: https://drops.da gstuhl.de/entities/document/10.4230/TGDK.2.
          <issue>1</issue>
          .3. doi:
          <volume>10</volume>
          .4230/TGDK.2.
          <issue>1</issue>
          .3, artwork Size: 37 pages,
          <volume>1598870</volume>
          bytes Medium: application/pdf Publisher:
          <article-title>Schloss Dagstuhl - Leibniz-Zentrum für Informatik</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>A.</given-names>
            <surname>Breit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Waltersdorfer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. J.</given-names>
            <surname>Ekaputra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sabou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ekelhart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Iana</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Paulheim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Portisch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Revenko</surname>
          </string-name>
          , A. t. Teije, et al.,
          <article-title>Combining machine learning and semantic web: A systematic mapping study</article-title>
          ,
          <source>ACM Computing Surveys</source>
          <volume>55</volume>
          (
          <year>2023</year>
          )
          <fpage>1</fpage>
          -
          <lpage>41</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Min</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          , G. Qi,
          <string-name>
            <surname>Can ChatGPT Replace Traditional KBQA</surname>
          </string-name>
          <article-title>Models? An In-Depth Analysis of the Question Answering Performance of the GPT LLM Family</article-title>
          , in: T. R.
          <string-name>
            <surname>Payne</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Presutti</surname>
            , G. Qi,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Poveda-Villalón</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Stoilos</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Hollink</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Kaoudi</surname>
          </string-name>
          , G. Cheng, J.
          <source>Li (Eds.)</source>
          ,
          <source>The Semantic Web - ISWC</source>
          <year>2023</year>
          , volume
          <volume>14265</volume>
          , Springer Nature Switzerland, Cham,
          <year>2023</year>
          , pp.
          <fpage>348</fpage>
          -
          <lpage>367</lpage>
          . URL: https://link.springer.com/10.1007/978-3-
          <fpage>031</fpage>
          -47240-4_
          <fpage>19</fpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>031</fpage>
          -47240-4_19, series Title: Lecture Notes in Computer Science.
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>G.</given-names>
            <surname>Agrawal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Kumarage</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Alghami</surname>
          </string-name>
          , H. Liu,
          <article-title>Can Knowledge Graphs Reduce Hallucinations in LLMs? : A Survey (</article-title>
          <year>2023</year>
          ). URL: https://arxiv.org/abs/2311.07914. doi:
          <volume>10</volume>
          .48550/ARXIV.2311.07 914, publisher: arXiv Version Number:
          <fpage>1</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>K.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y. E.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X. L.</given-names>
            <surname>Dong</surname>
          </string-name>
          , Head-to-Tail:
          <article-title>How Knowledgeable are Large Language Models (LLM)? A.K.A</article-title>
          .
          <string-name>
            <surname>Will LLMs Replace Knowledge</surname>
          </string-name>
          <article-title>Graphs? (</article-title>
          <year>2023</year>
          ). URL: https: //arxiv.org/abs/2308.10168. doi:
          <volume>10</volume>
          .48550/ARXIV.2308.10168, publisher: arXiv Version Number:
          <fpage>1</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>A.</given-names>
            <surname>Martino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Iannelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Truong</surname>
          </string-name>
          ,
          <article-title>Knowledge Injection to Counter Large Language Model (LLM) Hallucination</article-title>
          , in: C. Pesquita,
          <string-name>
            <given-names>H.</given-names>
            <surname>Skaf-Molli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Efthymiou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kirrane</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ngonga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Collarana</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Cerqueira</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Alam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Trojahn</surname>
          </string-name>
          , S. Hertling (Eds.),
          <source>The Semantic Web: ESWC 2023 Satellite Events</source>
          , volume
          <volume>13998</volume>
          , Springer Nature Switzerland, Cham,
          <year>2023</year>
          , pp.
          <fpage>182</fpage>
          -
          <lpage>185</lpage>
          . URL: https://link.s pringer.
          <source>com/10</source>
          .1007/978-3-
          <fpage>031</fpage>
          -43458-7_
          <fpage>34</fpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>031</fpage>
          -43458-7_34, series Title: Lecture Notes in Computer Science.
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , J. Liu, G. Liu,
          <article-title>An Enhanced Prompt-Based LLM Reasoning Scheme via Knowledge Graph-Integrated Collaboration (</article-title>
          <year>2024</year>
          ). URL: https://arxiv.org/abs/2402.04978. doi:
          <volume>10</volume>
          .48550/ARX IV.
          <year>2402</year>
          .04978, publisher: arXiv Version Number:
          <fpage>1</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [39]
          <string-name>
            <given-names>J.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Mei</surname>
          </string-name>
          , J. Ma, Can LLMs Efectively Leverage Graph Structural Information: When and
          <string-name>
            <surname>Why</surname>
          </string-name>
          (
          <year>2023</year>
          ). URL: https://arxiv.org/abs/2309.16595. doi:
          <volume>10</volume>
          .48550/ARXIV.2309.16595, publisher: arXiv Version Number:
          <fpage>2</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          [40]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <article-title>Evaluating and Enhancing Large Language Models for Conversational Reasoning on Knowledge Graphs (</article-title>
          <year>2023</year>
          ). URL: https://arxiv.org/abs/2312.11282. doi:
          <volume>10</volume>
          .48550/ARXIV.2312.11282, publisher: arXiv Version Number:
          <fpage>2</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          [41]
          <string-name>
            <given-names>A.</given-names>
            <surname>Martino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Iannelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Truong</surname>
          </string-name>
          ,
          <article-title>Knowledge Injection to Counter Large Language Model (LLM) Hallucination</article-title>
          , in: C. Pesquita,
          <string-name>
            <given-names>H.</given-names>
            <surname>Skaf-Molli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Efthymiou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kirrane</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ngonga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Collarana</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Cerqueira</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Alam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Trojahn</surname>
          </string-name>
          , S. Hertling (Eds.),
          <source>The Semantic Web: ESWC 2023 Satellite Events</source>
          , volume
          <volume>13998</volume>
          , Springer Nature Switzerland, Cham,
          <year>2023</year>
          , pp.
          <fpage>182</fpage>
          -
          <lpage>185</lpage>
          . URL: https://link.s pringer.
          <source>com/10</source>
          .1007/978-3-
          <fpage>031</fpage>
          -43458-7_
          <fpage>34</fpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>031</fpage>
          -43458-7_34, series Title: Lecture Notes in Computer Science.
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          [42]
          <string-name>
            <given-names>J.</given-names>
            <surname>Baek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. F.</given-names>
            <surname>Aji</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Safari</surname>
          </string-name>
          ,
          <article-title>Knowledge-Augmented Language Model Prompting for Zero-Shot Knowledge Graph Question Answering</article-title>
          ,
          <year>2023</year>
          . URL: http://arxiv.org/abs/2306.04136. doi:
          <volume>10</volume>
          .48550 /arXiv.2306.04136, arXiv:
          <fpage>2306</fpage>
          .04136 [cs].
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          [43]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Sun,</surname>
          </string-name>
          <article-title>MindMap: Knowledge Graph Prompting Sparks Graph of Thoughts in Large Language Models (</article-title>
          <year>2023</year>
          ). URL: https://arxiv.org/abs/2308.09729. doi:
          <volume>10</volume>
          .48550/ARXIV.2308. 09729, publisher: arXiv Version Number:
          <fpage>4</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          [44]
          <string-name>
            <given-names>H.</given-names>
            <surname>Abu-Rasheed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. H.</given-names>
            <surname>Abdulsalam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Weber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fathi</surname>
          </string-name>
          ,
          <article-title>Supporting Student Decisions on Learning Recommendations: An LLM-Based Chatbot with Knowledge Graph Contextualization for Conversational Explainability and Mentoring (</article-title>
          <year>2024</year>
          ). URL: https://arxiv.org/abs/2401.08517. doi:
          <volume>10</volume>
          .48550/ARXIV.2401.08517, publisher: arXiv Version Number:
          <fpage>3</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref45">
        <mixed-citation>
          [45]
          <string-name>
            <given-names>X.</given-names>
            <surname>Guan</surname>
          </string-name>
          , Y. Liu,
          <string-name>
            <given-names>H.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Han</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L</given-names>
            . Sun,
            <surname>Mitigating Large Language Model Hallucinations via Autonomous Knowledge</surname>
          </string-name>
          Graph-based
          <string-name>
            <surname>Retrofitting</surname>
          </string-name>
          (
          <year>2023</year>
          ). URL: https://arxiv.org/abs/2311.13314. doi:
          <volume>10</volume>
          .48550/ARXIV.2311.13314, publisher: arXiv Version Number:
          <fpage>1</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref46">
        <mixed-citation>
          [46]
          <string-name>
            <given-names>T.</given-names>
            <surname>Tietz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Jäger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Waitelonis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Sack</surname>
          </string-name>
          ,
          <article-title>Semantic annotation and information visualization for blogposts with refer</article-title>
          .,
          <source>in: VOILA@ ISWC</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>28</fpage>
          -
          <lpage>40</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref47">
        <mixed-citation>
          [47]
          <string-name>
            <given-names>T.</given-names>
            <surname>Rietz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Maedche</surname>
          </string-name>
          ,
          <string-name>
            <surname>Cody:</surname>
          </string-name>
          <article-title>An ai-based system to semi-automate coding for qualitative research</article-title>
          ,
          <source>in: Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>14</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref48">
        <mixed-citation>
          [48]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Cai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Siebert-Evenstone</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Eagan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. W.</given-names>
            <surname>Shafer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. C.</given-names>
            <surname>Graesser</surname>
          </string-name>
          , ncoder+:
          <article-title>a semantic tool</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>