<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Stanford University, Palo Alto, California, USA, March</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>A Hybrid Intelligent Approach for the Support of Higher Education Students in Literature Discovery</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ryan Prater</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Emanuele Laurenzi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>School of Business FHNW University of Applied Sciences and Arts Northwestern Switzerland</institution>
          ,
          <addr-line>CH-4600 Olten</addr-line>
          ,
          <country country="CH">Switzerland</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2022</year>
      </pub-date>
      <volume>2</volume>
      <fpage>1</fpage>
      <lpage>23</lpage>
      <abstract>
        <p>In this paper, we present a hybrid intelligent approach that combines knowledge engineering, machine learning, and human intervention to automatically recommend literature resources relevant for a high quality of literature discovery. The primary target group that we aim to support is higher education students in their first experiences with research works. The approach builds a knowledge graph by leveraging a logistic regression algorithm which is first parameterized and then influenced by the interventions of a supervisor and a student, respectively. Both interventions allow continuous learning based on both the supervisor's preferences (e.g. high score of H-index) and the student's feedback to the resulting literature resources. The creation of the hybrid intelligent approach followed the Design-Science Research methodology and is instantiated in a working prototype named PaperZen. The evaluation was conducted in two complementary ways: (1) by showing how the design requirements manifest in the prototype, and (2) with an illustrative scenario in which a corpus of a research project was taken as a source of truth. A small subset from the corpus was entered into the PaperZen and Google Scholar, independently. The resulting literature resources were compared with the corpus of a research project and show that PaperZen outperforms Google Scholar.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;literature discovery</kwd>
        <kwd>hybrid intelligence</kwd>
        <kwd>knowledge graph</kwd>
        <kwd>logistic regression</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The review of literature in research works like master’s theses serve as the foundation upon
which that research is built [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The presence of a research-worthy problem, for example, is
established through the literature [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The discovery of content covering existing research,
theories and evidence (i.e., literature discovery) is prerequisite for a literature review and creates
the conditions for critically evaluating and discussing this content [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        Conducting literature discovery is an intense learning process for higher education students,
especially for those novel to research. The initial guidance of a supervisor is gradually weaned
over time by the increasingly acquired self-regulated learning ability of the student, until the
discovery can be conducted autonomously. Self-regulated learning refers to students’ cognitive
and motivational strategies for learning [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], and shall be supported for the benefit of the student’s
learning process, e.g., see [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        Poor support on student’s literature discovery might bring frustration such as sense of lost
time or inability to make progress. This is problematic as it can lead to serious consequences
such as the student abandoning the thesis topic or dropping the module [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        In response to this problem, we present a hybrid intelligent approach that aims to foster
self-regulated learning in literature discovery for research works like, but not limited to, master’s
theses. To accomplish this, the approach implements the Zone of Proximal Development (ZPD)
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. That is, a student in the ZPD of a particular task can accomplish it if appropriate assistance
is provided (see the middle layer in Figure 1).
      </p>
      <p>In our approach, the assistance is provided by incorporating:
1. the supervisor’s preferences about paper discovery, e.g., a high H-index of authors or
Journal paper over conference papers. This refers to the "what I can’t do” layer for a
student (see Fig. 1);
2. the student’s feedback over the discovered research resources. This refers to the "what I
can do” layer for a student (see Fig. 1);
3. the automatic suggestion of new resources, based on the combination of both (1) and (2).</p>
      <p>This refers to the "what I can do with help” layer for a student (see Fig. 1).</p>
      <p>While the supervisors’ preferences can be engineered, discovery of research resources is
tacit knowledge which we aim to learn with the help of the student’s feedback. Therefore, a
hybrid intelligent approach is implemented that combines Knowledge Engineering and Machine
Learning approaches with human interventions.</p>
      <p>This paper is structured as follows. First, in Section 2 background concepts are explored,
including a brief overview of learning theory, imparting of knowledge, as well as related work
and design patterns for hybrid intelligent approaches. The methodology of the approach is then
explained in Section 3. The approach is then described in Section 4.2 in the form of a boxology.
The approach is instantiated into a working prototype as illustrated in Section 4.3. Section
5 discusses how the value delivered to students and supervisors was measured by
existencematching through comparison of the corpus of a research project. Finally, a conclusion and
suggestions for future research are made in Section 6.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Background and Related Work</title>
      <p>This section introduces the theoretical background for learning and imparting of knowledge
within the learning context. Next, approaches for literature discovery are discussed. Finally,
design patterns for hybrid intelligence are introduced.</p>
      <sec id="sec-2-1">
        <title>2.1. Learning Theory and Imparting of Knowledge</title>
        <p>
          Learning theory strives to explain how students receive, process and retain knowledge during
learning [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. Prior knowledge or experiences play a fundamental part in how understanding is
acquired or changed and knowledge and skills retained.
        </p>
        <p>
          Vygotsky [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] introduced the notion of a Zone of Proximal Development (ZPD) as "the distance
between the actual developmental level as determined by independent problem-solving and the
level of potential development as determined through problem-solving under adult guidance or
in collaboration with more capable peers". According to Vygotsky, when a learner is in the ZPD
for a particular task, she/he can achieve it if appropriate assistance is provided.
        </p>
        <p>
          Similarly, connectivism [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] postulates that learning occurs when connections are made
between nodes in a learner’s network - where a node can be anything from knowledge in the
learner’s mind, to a digital artifact, or another person. This implies that new knowledge must
be connected to existing knowledge or experiences which can be understood as a concretization
of the ZPD and that such connection can be mediated by human interaction in the digital
environment.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Approaches for Literature Discovery</title>
        <p>
          This section elaborates on existing approaches for literature discovery which follow a subset of
the system outlines and characteristics of existing literature discovery systems by Beel et al. [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ].
The relevant systems are categorized in approaches that first rely on knowledge graphs, then
machine learning, and finally their hybrid combination.
        </p>
        <sec id="sec-2-2-1">
          <title>2.2.1. Knowledge Graphs in Literature Discovery</title>
          <p>
            Knowledge graphs (KG) use labeled schemas to enable semantical reasoning of their contained
data [
            <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
            ], e.g., searching papers which are cited from a specific author. Knowledge Graphs
are often stored in graphs databases, which are designed to facilitate knowledge retrieval and
logical inferencing. Elsevier1 and SemanticScholar2 are examples of scholarly knowledge graphs
which aim to structure properties and information of resources related to literature such as
publication date, journal ISSN, and publishing authors.
          </p>
          <p>1https://www.elsevier.com/connect/how-ai-and-knowledge-graphs-can-make-your-research-easier
2https://www.semanticscholar.org/about</p>
          <p>
            Knowledge graphs are also used in scholarly literature for approaches to literature discovery.
Ammar et al. [
            <xref ref-type="bibr" rid="ref12">12</xref>
            ] described the construction of SemanticScholar through the use of labeled
nodes and edges, and how they can be used for discovery. Liang et al. [
            <xref ref-type="bibr" rid="ref13">13</xref>
            ] constructed a citation
graph with a citation link schema, e.g., Comparable, Based-On, and General, and developed
literature recommendations based on schema weights and data distance. Auer et al. [
            <xref ref-type="bibr" rid="ref14">14</xref>
            ]
demonstrated the human-machine readability of knowledge graph literature discovery via three
use-cases: computer science, COVID-19, and materials science, and highlight the importance of
semantic structuring.
          </p>
          <p>Pure knowledge graph-based approaches however, require contextual experience for
semantical understanding which novice researchers may lack. An approach which learns student
behaviors via machine learning may provide a better medium for self-regulated learning.</p>
        </sec>
        <sec id="sec-2-2-2">
          <title>2.2.2. Machine Learning in Literature Discovery</title>
          <p>
            Some literature discovery systems leverage machine learning to make predictions based on
scoring of literature properties without the use of knowledge graphs. The LitSuggest system
[
            <xref ref-type="bibr" rid="ref15">15</xref>
            ] used a logistic regression classifier to analyze a flat-list of PubMed articles based on various
properties, e.g., text analysis of abstracts, and achieved competitive results compared to similar
systems. CiteSeer [
            <xref ref-type="bibr" rid="ref16">16</xref>
            ], one of the earliest literature recommendation systems, parsed related
citations into tabular format, evaluated various methods of intelligent string comparison, and
noted dificulties with "identification and disambiguation of authors".
          </p>
          <p>A drawback of pure machine learning approaches is the lack of contextual knowledge and
generalizability, which hinder the identification of resources that are relevant to a research
problem, especially those targeting diferent application domains. A hybrid combination of
knowledge engineering and machine learning can be used to overcome the drawback of both
approaches.</p>
        </sec>
        <sec id="sec-2-2-3">
          <title>2.2.3. Hybrid Approaches in Literature Discovery</title>
          <p>
            Literature recommendation systems exist which use both knowledge graphs and machine
learning. For example, Middleton et al. [
            <xref ref-type="bibr" rid="ref17">17</xref>
            ] used a multi-class nearest-neighbor algorithm and
compared a flat-list data representation against an instantiated is-a ontology of hierarchical
literature topics and found the ontology approach resulted in improved accuracy and "rounder
profiles".
          </p>
          <p>
            Such multi-faceted systems may also include a human component which can satisfy or
supplement the recommendation task. Beel et al. [
            <xref ref-type="bibr" rid="ref9">9</xref>
            ] explored studies using collaborative filtering ,
which leverage social systems for improved performance. Pennock et al. [
            <xref ref-type="bibr" rid="ref18">18</xref>
            ] categorized users
according to personality type and applied it to the CiteSeer [
            <xref ref-type="bibr" rid="ref16">16</xref>
            ] system. Tang and McCalla
[
            <xref ref-type="bibr" rid="ref19">19</xref>
            ] applied pedagogical features, e.g., learner background knowledge and interest, to literature
discovery and emphasized the their importance for expanding suggestions outside of localized
results.
          </p>
          <p>Although such hybrid approaches provide greater support in literature discovery, they lack
human intervention, which could be helpful for the students to achieve an adequate Zone of
Proximal Development.</p>
          <p>
            By integrating student-supervisor interaction into a hybrid approach, we aim to support
students who are novel to research to achieve an adequate Zone of Proximal Development.
That is, literature resources are ofered to students by taking into account both their feedback
and supervisors’ preferences. The proposed hybrid intelligent approach also contribute to
address the research call made by Auer et al. [
            <xref ref-type="bibr" rid="ref14">14</xref>
            ], who asked "How can we increasingly involve
specialist scientists in the curation process?".
2.3. Patterns for Hybrid Approaches
van Harmelen and ten Teije [
            <xref ref-type="bibr" rid="ref20">20</xref>
            ] proposed an abstraction mechanism for all hybrid approaches
that combine learning and knowledge engineering which led to a set of reusable and
compositional design patterns. Witschel et al. [
            <xref ref-type="bibr" rid="ref21">21</xref>
            ] builds on their work to abstract systems with a human
in the loop. For example, the output of a machine learning may be in a human-interpretable
form which could be used by a human expert to create or improve a knowledge representation,
which could then be used for deductive reasoning. We leverage these research findings on
design patterns for hybrid (intelligent) approaches to create our approach (see Section 4.2).
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology</title>
      <p>
        The hybrid intelligent approach was developed by following Design Science Research [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]
which is a problem-solving methodology via the process of awareness of problem, approach,
development, evaluation, and conclusion. This methodology was used by other studies on
literature discovery [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] and was selected because instantiation of an innovative artifact would
enable evaluation of our approach to self-regulated learning of literature discovery.
      </p>
      <p>In the awareness of problem phase (Section 4.1), semi-structured interviews were conducted
to supplement literature findings on approaches for literature discovery. A set of requirements
were then derived and addressed with the hybrid intelligent boxology approach by the designed
artifact (Section 4.2 - approach phase). These requirements were incorporated in the approach
instantiation, which comes in the form of a running prototype (Section 4.3 - development phase).
In the evaluation phase (Section 5), the hybrid intelligent approach was evaluated through the
prototype, first with respect to the requirements and then by performing test runs to evaluate
on the breadth of the approach. The conclusion phase then takes places in Section 6.</p>
    </sec>
    <sec id="sec-4">
      <title>4. The Hybrid Intelligent Approach</title>
      <p>This section describes the proposed hybrid intelligent approach to support higher education
students in literature discovery. An awareness of the problem is derived from stakeholder
interviews which are translated into a set of requirements. These requirements are manifested in
the components of the approach. The latter is then instantiated in the technical implementation
called PaperZen.</p>
      <sec id="sec-4-1">
        <title>4.1. Awareness of Problem</title>
        <p>For an in-depth understanding of the problem addressed in this work, literature findings were
supplemented with primary data from semi-structured interviews. The latter were conducted
face-to-face physically or via calls, with duration of one hour each. The selection criteria (C1-4)
for interviewees were as follows:</p>
        <p>C1 a university research supervisor and reviewer with at least 5 years of experience in
both supervising research works and reviewing literature discovery guided by other
supervisors.</p>
        <p>C2 a university research supervisor with at least 5 years of experience in supervising research
works.</p>
        <p>C3 a master student engaged in writing a master’s thesis, thus being supervised.
C4 a master student in the last year of studies having at least one experience with a supervised
research work but not yet started their theses.</p>
        <p>Respectively, interviews were conducted with: (C1) one research supervisor and reviewer
a university professor with 10 years experience in both supervising and reviewing. (C2) one
research supervisor - a university lecturer and researcher with 5 years experience in supervising,
(C3) two master’s students writing theses, (C4) and 17 masters students at their last year of
studies with 1 to 3 research works under supervision. The higher number in category (C4) is
justified by less experience in research works.</p>
        <p>To elicit as much insights as possible, open-ended questions were asked, e.g.,
• As a supervisor to thesis writers, what pain points do you experience when guiding
literature discovery? Please underpin the answers by referring to concrete cases from
your past supervision experience (C1 and C2).
• As a master’s student writing your thesis, how do you discern literature quality? Please
underpin the answers by referring to concrete cases from your past experience (C3).
• As a master’s student with previous experience in research works, what constraints or
oversight does your supervisor have on your literature discovery? Please underpin the
answers by referring to concrete cases from your past experience (C4).</p>
        <p>Interviews results were analysed and the main findings ( F1-4) were consolidated.
F1 : No clear guidance makes students insecure about their direction (C3 and C4).
F2 : There is no scientific method to convert research questions into literature searches
(C1 and C2). This prompts supervisors to create varying and sometimes inconsistent
constraints for their students (C4).</p>
        <p>F3 : Students must review the results of their literature discovery, many of which might not
be relevant to the topic (C3 and C4).</p>
        <p>F4 : If a student loses direction, supervisors only have the opportunity for correction at the
next review of the students work (C1 and C2).</p>
        <p>F5 : From the perspective of (C1), neglect of proper supervision from (C2), e.g. of lack of
time, impacts the quality of literature discovery from (C3).</p>
        <p>Finally, findings were mapped into three main requirements ( R1-3), which set the basis for
designing the novel hybrid intelligent approach.</p>
        <p>R1 : the approach shall enable the students to be guided when selecting literature as well as
being empowered to decide on which resource to review (F1, F3).</p>
        <p>R2 : the approach shall enable the supervisors to influence student literature discovery
outside of the classroom (F4, F5).</p>
        <p>R3 : the approach shall enable to adequately consider the knowledge of both students and
supervisors in the literature discovery result (F2, F4).</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. The Suggested Boxology for the Hybrid Intelligent Approach</title>
        <p>
          Our approach is illustrated in Figure 2 using the boxology approach proposed by van Harmelen
and ten Teije [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] and later extended by Witschel et al. [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ]. Human represents a human
component, sym a model-based input or output, data a non-model-based input or output, ML an
inductive machine learning inference, and KR a deductive knowledge representation inference.
Numbers are for reference purposes.
        </p>
        <p>
          Three boxology design patterns influenced this approach. The Feedback-Based Learning
Pattern of Witschel et al. [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] was used to address R1 and manifested in boxes 1-6 to reuse a
student’s input (box 2) to incorporate human judgement into the the machine learning weights
of box 3 which feeds back to box 1. These weights are then used in box 5 in combination with the
supervisor scores (box 12) to influence the suggestions (box 6), which addresses R2 and follows
the Explainable learning systems with background knowledge pattern of van Harmelen and ten
Teije [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ]. The combination of the student and supervisor scores (boxes 4, 12) to produce a
combined output (box 6) follows the Learning an intermediate abstraction for reasoning pattern,
which addresses R3.
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Instantiation and Steps of the Hybrid Intelligent Approach</title>
        <p>The hybrid intelligent approach was instantiated in a working prototype called PaperZen.
Neo4j3 was chosen as the knowledge graph database for storage. The literature resources were
obtained via API of SemanticScholar. The LogisticRegression model provided by the Python
library scikit-learn with default parameters was used to implement machine learning with
logistic regression (LR).</p>
        <p>The instantiation consists of several steps and components which were derived from the
boxology pattern of Figure 2. A student (box 1) begins their discovery journey by adding
literature papers to their library (student corpus, box 2) which are saved to a graph database.
During or before this time, a supervisor (box 9) sets their search preferences for the student
by adjusting multiple feature weights (box 10). The user then initiates a search. The approach
gathers (box 7) a number of related papers or authors (Potentials, box 8) which are trained on a
logistic regression model (box 3) of previously saved or rejected search suggestions from the
student (box 8). The Potentials are then evaluated against the trained model and a vector of
probabilities for the saved class are outputted (vector-y-student, box 4). The normalized
feature columns of the test set are multiplied by the supervisor preferences (box 11) and each
row is then mean-averaged into a vector of probabilities (vector-y-supervisor, box 12).
vector-y-student and vector-y-supervisor are then mean-averaged together (box 5)
to produce the final per-Potential prediction vector-y-prediction (box 6) which is then
ordered by prediction value and presented in groups of three to the student (back to box 1).
The student can choose to save or reject each suggestion individually (box 2). Based on these
decisions, the LR model (box 3) is further trained and the process begins again.</p>
        <p>Boxes 7, 11, 3, 5, and 1 are further detailed below.</p>
        <sec id="sec-4-3-1">
          <title>4.3.1. Knowledge Graph Component</title>
          <p>To seed (start) their corpus, a student may search for and add a paper by DOI, paper title, or
BibTeX upload. To represent the various literature resource types while supporting eficient
storage and efective human and machine readability, the following labelling schema was used:
• Corpus - Saved: Nodes which have been saved to the student’s corpus.
• Corpus - Rejected: Nodes which have been rejected by the student during search.
• Resource - Paper: Nodes which are papers, e.g., journal article, conference paper,
book, etc. This is predetermined by SemanticScholar.
• Resource - Author: Nodes which are authors of papers.
• Resource - Citation: Nodes (implicitly papers) listed as a citations to papers.
• Resource - Reference: Nodes (implicitly papers) listed as a references from papers.</p>
          <p>The labels are added at diferent stages of student interaction, and multiple labels may be
applied to a single node, e.g., a student may be suggested a Resource - Paper which is both a
Resource - Reference and a Resource - Citation, and upon saving the suggestion the
node receives the Corpus - Saved label. Currently, rejected suggestions are never presented
again unless the student manually adds them to their corpus.</p>
          <p>Upon addition of resources to the student’s library, PaperZen preemptively discovers and
hydrates (represents) Potential resources (box 7). To ensure a distributed approach, each node
saved in the student’s corpus is iterated following the process of Listing 1. A configuration
number, NUM_LOOKUP_PER_NODE, controls how many Potentials of each resource can be
extracted per node, and another, NUM_TARGET_PER_RESOURCE, defines the target amount of
Potentials per resource type. To prevent redundancy, no lookup is done if a node has already
been hydrated by the API.</p>
          <p>Listing 1: Pseudocode of aggregating Potentials
def a g g r e g a t e _ p o t e n t i a l s ( s t u d e n t s _ s a v e d _ c o r p u s ) :</p>
          <p>NUM_TARGET_PER_RESOURCE = 100
NUM_LOOKUP_PER_NODE = 3
f o r node in s t u d e n t s _ s a v e d _ c o r p u s :
f o r l a b e l in [ ’ a u t h o r ’ , ’ c i t a t i o n ’ , ’ r e f e r e n c e ’ ] :
e x i s t i n g _ p o t e n t i a l s = g e t _ h y d r a t e d _ n o d e s ( c o r p u s , l a b e l</p>
          <p>)
p o t e n t i a l s = g e t _ c o n n e c t i o n s _ t o _ n o d e ( node )
n e w _ p o t e n t i a l s = e x i s t i n g _ p o t e n t i a l s − p o t e n t i a l s
new_nodes = [ ]
f o r p o t e n t i a l in n e w _ p o t e n t i a l s :
i f len ( e x i s t i n g _ p o t e n t i a l s ) + len ( new_nodes ) &gt;</p>
          <p>NUM_LOOKUP_PER_NODE :
break
new_node = h y d r a t e _ a n d _ s t o r e _ n o d e ( p o t e n t i a l )
new_nodes += new_node
i f ( len ( e x i t i n g _ p o t e n t i a l s ) + len ( new_nodes ) ) &gt;</p>
          <p>NUM_TARGET_PER_RESOURCE :
break</p>
        </sec>
        <sec id="sec-4-3-2">
          <title>4.3.2. Human Preference Component</title>
          <p>Box 11 symbolizes the encoding of Potentials’ various resource characteristics as features to be
preferenced by the supervisor. To do this, the following properties are used as feature weights:
• L: Labels indicating resource type. This is one-hot-encoded into L-A (Resource</p>
          <p>Author), L-C (Resource - Citation), and L-R (Resource - Reference).
• HI: H-Index of the author
• CC: Citation count of the paper
• RC: Reference count of the paper
• PY: Publication year of the paper
• RI: Number of relationships inbound to a node
• RO: Number of relationships outbound from a node</p>
          <p>These weights are available to supervisors via a user interface which allows customization of
preferences (supervisor weights) for the prevalence of such features in search suggestions and
addresses R2 (See Figure 3 (2)). For example, by setting the weight for L-A to 100 and the weight
of L-C and L-R both to zero, the supervisor efectively influences the search to only return
authors. These weights are also used as features to the LR model (box 3). This also addresses R1
by quantifying various features of the literature discovery to learn the search preferences of
students.</p>
        </sec>
        <sec id="sec-4-3-3">
          <title>4.3.3. Machine Learning Component</title>
          <p>
            Upon initialization of PaperZen, the LR (box 3) training dataset is empty, so results are
compensated by the supervisor weights which are defaulted to 0.5. If not empty, existing decisions
(saved or rejected) of a student are used for training the LR model. HI, CC, RC, and PY are
commonly used in literature recommendation systems [
            <xref ref-type="bibr" rid="ref16">16</xref>
            ]. These values are set to zero when
not applicable to certain resource types, e.g., a paper will not have an H-Index, but publication
year is set to the mean of the publication years.
          </p>
          <p>
            To ensure balanced weighting across features, each feature is individually normalized using
a min-max scalar with values in the range of [
            <xref ref-type="bibr" rid="ref1">0,1</xref>
            ]. The resulting training set matrix is referred
to as matrix_X_train_student. The Potentials are then prepared and normalized using the
same process for training data to produce matrix_X_test_student. The output of evaluation
is a vector of prediction probabilities (vector_y_student) for Corpus - Saved with values
in the range of [
            <xref ref-type="bibr" rid="ref1">0,1</xref>
            ] which represents the vectorization of R1.
          </p>
        </sec>
        <sec id="sec-4-3-4">
          <title>4.3.4. Recommendation Component</title>
          <p>Box 5 represents the combination of LR and supervisor weights to form literature resource
suggestions. Following the LR of the Potentials and the output of the probabilities vector, the loaded
model coeficients for each feature are extracted. This is a vector which represent the LR weights
the trained model assigns to each feature in order to calculate the prediction for each sample.
The supervisor weights are then matrix-multiplied across matrix_X_test_student to
produce a matrix of human-influenced features matrix_X_test_supervisor and is then
meanaveraged row-wise to produce the prediction vector vector_y_supervisor which represents
the supervisor preference of each resource. vector_y_student and vector_y_supervisor
are then concatenated and mean-averaged row-wise to produce vector_y_prediction which
represents the hybrid student and supervisor preference, addressing R2.</p>
        </sec>
        <sec id="sec-4-3-5">
          <title>4.3.5. Feedback Loop Component</title>
          <p>The results of matrix_y_hybrid are index-matched and sliced from
matrix_X_train_supervisor, ordered by descending value, and returned to the
calling search function. The result is a list of Potentials which are ordered by hybrid
student-supervisor probability of the student saving the resource. This list is then displayed to
the student (box 1) in batches of three.</p>
          <p>When presented the suggestions, the student may choose to save or reject each suggestion
individually. This represents the guided choice of R1. If rejected, the resource (node) is labeled
with Corpus - Rejected and the resource is never again suggested to the student. If saved,
the label Corpus - Saved is applied to the resource, and the resource will be added to the
student’s corpus. These resources are later included in the next search as training data, which
means the LR model dataset grows proportional to use and R3 is iteratively addressed.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Evaluation</title>
      <p>The hybrid intelligent approach was evaluated in two manners: (1) by showing how the
requirements manifest in the technical prototype and (2) by proving the utility of the proposed
approach through tests. The former is described in Section 5.1 and the latter in Section 5.2.</p>
      <sec id="sec-5-1">
        <title>5.1. Requirements Traceability</title>
        <p>Figure 3 depicts the three main user interfaces of PaperZen with numbers for reference. The
student may upload resources which are added to their library (1). The supervisor may then
adjust their preference weights (2) which influence the suggestions in the student’s search (3).
The three main requirements for the approach manifest as follows:</p>
        <p>R1 : the approach shall enable the students to be guided when selecting literature as well as
being empowered to decide on which resource to review. Fig. 3 (3) shows how PaperZen
suggests resources for review to the student.</p>
        <p>R2 : the approach shall enable the supervisors to influence student literature discovery
outside of the classroom. Fig. 3 (2) shows how PaperZen allows the supervisor to adapt
their preference weights.</p>
        <p>R3 : the approach shall enable to adequately consider the knowledge of both students and
supervisors in the literature discovery result. This requirement manifests in terms of
displayed resources (see Fig. 3 (1)) after both the supervisor and the student provide the
preferences for and the feedback to the resources, respectively.</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Utility of the PaperZen Approach</title>
        <p>
          This section describes how the approach was evaluated with respect to its utility. This refers
to the ability of the PaperZen approach to return satisfying results. For this, the research
literature of the Swiss research project "Digital Self-Study Assistant" (DSSA) [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ] was used as
a ground truth dataset (truth corpus). Hence, five test runs were performed covering various
configurations to explore the breadth of PaperZen’s features. The results were then compared
to Google Scholar, to which the title of the paper(s) in each of the test runs were entered in the
search bar. Details are reported in the below sub-sections 5.2.1 to 5.2.3. Finally, a discussion
about the evaluation results is provided in sub-section 5.2.4.
        </p>
        <sec id="sec-5-2-1">
          <title>5.2.1. The Truth Corpus</title>
          <p>
            The truth corpus4 contained 43 items, one of which was a website, therefore the sample size n
was 42. The 42 articles were searched in SemanticScholar and four articles could not be found,
meaning the best accuracy possible for PaperZen would be 90.48%. Additional characteristics of
the truth corpus were: 22 of the literature were journal articles, 12 were conference proceedings,
and eight were books or book excerpts. The mean publication year of the literature was 2011
and the median was 2014. The total amount of authors (separated from their papers) was 107.
Figure 4 depicts an excerpt of the knowledge graph from the corpus of the DSSA research
project [
            <xref ref-type="bibr" rid="ref24">24</xref>
            ].
          </p>
          <p>Five evaluations were run on the truth corpus:
• Eval0: One random article from the truth corpus was selected, LR weights were set at
default, supervisor weights were set at default
• Eval1: The same article from Eval0 was selected, LR weights were reused from Eval0,
supervisor weights were set at default, except for Author=1.0
• Eval2: Three random articles from the truth corpus were selected, LR weights were set
at default, supervisor weights were set at default
• Eval3: The same articles from Eval2 were selected, LR weights were reused from Eval2,
supervisor weights were set at default, except for Publication Year=1.0
• Eval4: Nine random articles from the truth corpus were selected, LR weights were set at
default, supervisor weights were set at default, except for Author=0.0</p>
          <p>
            Given the truth corpus of size n and because PaperZen provides recommendations in groups
of 3, n/3 iterations of PaperZen search took place for each evaluation. Accuracy was measured
(true positives/n), which is a common evaluation method of recommendation systems [
            <xref ref-type="bibr" rid="ref9">9</xref>
            ].
          </p>
          <p>The seed article(s) for matching the truth corpus were chosen by random integer between
0 and n and used to index-select the truth corpus literature sorted by ascending publication
year. When indicated, the LR weights were reset to default, meaning the training began with
the seed article(s). When indicated, the supervisor weights were reset to default values of 0.5
per feature. After every evaluation, the database was reset, allowing suggestions to reappear in
subsequent tests.</p>
        </sec>
        <sec id="sec-5-2-2">
          <title>5.2.2. Evaluation Results</title>
          <p>The results of the five evaluations are show in Table 1 and the resulting LR weights in Table 2.</p>
          <p>In Eval0, three suggestions matched the truth corpus, with mean and median publication
years of 2011 and 2014 respectively. From 42 suggestions, two were duplicates, but were not
matches to the truth corpus. All three matches were literature articles with publication years of
2018, 1999, and 2019 and appeared in the 7th, 9th, and 14th search iterations respectively.</p>
          <p>In Eval1, the LR weights were kept from Eval0 and L-A was increased to 1 to test the efect
on resource types returned. There were six matches to the truth corpus - two literature articles
and four authors. The literature article publication years were both 2017 and appeared in the
4th and 5th search iterations respectively.</p>
          <p>Eval2 evaluated the efect of using multiple seed articles. The three articles did not include
the article used in Eval0 or Eval1. The LR weights and supervisor weights were reset to
default (0.5). Out of 42 suggestions, there were zero matches. Two suggestions were duplicates.
Both the mean and median publication years for Eval2 were 1993.
Evaluation
Eval0
Eval1
Eval2
Eval3
Eval4</p>
          <p>Matches
3
6
0
9
0</p>
          <p>Eval3 used the same seed articles as Eval2, but set PY to 1. The LR weights and supervisor
weights were reset to default. 33 literature articles were returned (9 authors) with mean and
median publication years of 2021 and 2019 respectively. There were nine matches, all authors
of the seed papers. The author matches were spread among the 5th-12th search iterations.</p>
          <p>Eval4 used nine seed articles to evaluate the efect of using 20% of the truth corpus for
seeding. The LR and supervisor weights were reset to default, except for L-A set to 0 to avoid
author-resource-centralization. There were zero matches to the truth corpus.</p>
        </sec>
        <sec id="sec-5-2-3">
          <title>5.2.3. Comparison to Google Scholar</title>
          <p>For a comparison to Google Scholar, the titles of the selected articles of the evaluations were
concatenated together into the search of Google Scholar, and the top n results were recorded,
with n being equal to the number of articles in the truth corpus. This comparison was done
for Eval0, Eval2, and Eval4 because Eval2 and Eval3 reused the articles from Eval0 and
Eval2 respectively.</p>
          <p>For Eval0, Google Scholar returned 29,300 results in 0.05 seconds, of which the first 42
articles were compared to the truth corpus. One article which was present in the truth corpus,
resulting in an accuracy of 0.0238. For Eval2, Google Scholar returned 25 results in 0.06 seconds,
of which all results were compared to the truth corpus. Zero articles were present in the truth
corpus, resulting in an accuracy of 0. For Eval4, Google Scholar returned 0 results. This was
most likely due to the concatenation of the nine article titles.</p>
        </sec>
        <sec id="sec-5-2-4">
          <title>5.2.4. Discussion on Evaluation Results</title>
          <p>The evaluation of matching accuracy was generally low, with a 7% literature match in Eval0,
5% literature match in Eval1, and 0% in other evaluations. However, PaperZen outperformed
Google Scholar which matched 2% of the literature in Eval0 and 0% in other evaluations.
Further, many of the returned articles showed characteristics applicable to the research topic
of DSSA (e.g., titles such as "Simulating Instructional Roles through Pedagogical Agents") but
were not in the truth corpus, perhaps indicating a small sample size. Additionally, the first
suggestions of Eval4 were unrelated in title to DSSA, but had a large number of citations,
possibly indicating min-max stratification of the CC weight and a local search maximum.</p>
          <p>The feature weights in Table 2 appeared to be susceptible to clustering, such as RI and RO.
Given the low matching accuracy, we are hesitant to conclude that RI and RO are actually as
important as indicated by the data. However, Eval1 confirmed the ability of L-A to influence
suggested resource types and Eval3 confirmed the ability of PY to influence the publication
year of the literature. This indicates feasibility of supervisor influence on student literature
suggestions during discovery and addresses.</p>
          <p>Four authors in Eval1 and nine in Eval3 were suggested, all authors of the seed literature.
The relation of these resources to the seed literature indicates the connectivity of related resource
types and verifies the integrity of the knowledge graph architecture.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>The presented hybrid intelligent approach aims to support higher education students who
are novel to research to achieve an adequate Zone of Proximal Development in literature
discovery. That is, literature resources are ofered to students by taking into account both their
feedback and supervisors’ preferences. Specifically, the suggestions derive from a combination
of literature discovery techniques of knowledge graphs, machine learning, and collaborative
ifltering and allow the research-supervisor to influence what is presented to the student while
the student’s self-regulated learning is supported by learning their personal capabilities. The
approach is instantiated in a working prototype called PaperZen, which incorporates techniques
of knowledge engineering, machine learning, and human interventions. By matching output
suggestions to a known research corpus, PaperZen showed matches could be made and displayed
comparative and better performance to Google Scholar.</p>
      <p>Limitations included the small sample size of the 42 resources in the single truth corpus,
the suggestion of author resources without inclusion in the seed corpus, and the diference in
search inputs for Google Scholar (keyword text input) compared to PaperZen (resource based).
In the future, additional features will be explored over a greater evaluation sample to ensure
reproducible results, text analysis will be included to reduce local search maximums, and both
students and supervisors will be involved in the evaluation for a deeper analysis of learning.
1–17</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Ellis</surname>
          </string-name>
          ,
          <article-title>A Systems Approach to Conduct an Efective Literature Review in Support of Information Systems Research</article-title>
          , Informing Science: The
          <source>International Journal of an Emerging Transdiscipline</source>
          <volume>9</volume>
          (
          <year>2006</year>
          )
          <fpage>181</fpage>
          -
          <lpage>212</lpage>
          . doi:
          <volume>10</volume>
          .28945/479.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>T.</given-names>
            <surname>Ellis</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y. Levy</surname>
          </string-name>
          ,
          <article-title>Framework of Problem-Based Research: A Guide for Novice Researchers on the Development of a Research-Worthy Problem</article-title>
          , Informing Science: The
          <source>International Journal of an Emerging Transdiscipline</source>
          <volume>11</volume>
          (
          <year>2008</year>
          )
          <fpage>17</fpage>
          -
          <lpage>33</lpage>
          . doi:
          <volume>10</volume>
          .28945/438.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Friedrich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Jonkmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Nagengast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Schmitz</surname>
          </string-name>
          , U. Trautwein,
          <article-title>Teachers' and students' perceptions of self-regulated learning and math competence: Diferentiation and agreement</article-title>
          ,
          <source>Learning and Individual Diferences</source>
          <volume>27</volume>
          (
          <year>2013</year>
          )
          <fpage>26</fpage>
          -
          <lpage>34</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.lindif.
          <year>2013</year>
          .
          <volume>06</volume>
          .005.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>English</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kitsantas</surname>
          </string-name>
          ,
          <article-title>Supporting student self-regulated learning in problem-and project-based learning</article-title>
          ,
          <source>Interdisciplinary Journal of Problem-based Learning 7</source>
          (
          <year>2013</year>
          )
          <fpage>128</fpage>
          -
          <lpage>150</lpage>
          . doi:
          <volume>10</volume>
          .7771/
          <fpage>1541</fpage>
          -
          <lpage>5015</lpage>
          .
          <fpage>1339</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>D.</given-names>
            <surname>Hallberg</surname>
          </string-name>
          , U. Olsson,
          <article-title>Self-regulated learning in students' thesis writing</article-title>
          ,
          <source>International Journal of Teaching and Education</source>
          <volume>1</volume>
          (
          <year>2017</year>
          )
          <fpage>13</fpage>
          -
          <lpage>24</lpage>
          . doi:
          <volume>10</volume>
          .20472/TE.
          <year>2017</year>
          .
          <volume>5</volume>
          .1.002.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>L. S.</given-names>
            <surname>Vygotsky</surname>
          </string-name>
          , Mind in Society: Development of Higher Psychological Processes, Harvard University Press, Cambridge, MA,
          <year>1978</year>
          . doi:
          <volume>10</volume>
          .2307/j.ctvjf9vz4.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>O. J.</given-names>
            <surname>Ellis</surname>
          </string-name>
          , Human Learning, 6th ed.,
          <source>Pearson</source>
          ,
          <year>2012</year>
          . URL: https://www.pearson.com/us/ higher-education/product/Ormrod-Human-Learning
          <string-name>
            <surname>-</surname>
          </string-name>
          6th-Edition/9780132595186.html.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>G.</given-names>
            <surname>Siemens</surname>
          </string-name>
          ,
          <article-title>Connectivism: A learning theory for the digital age</article-title>
          ,
          <source>International Journal of Instructional Technology and Distance Learning</source>
          <volume>2</volume>
          (
          <year>2004</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>Beel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Gipp</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Langer</surname>
          </string-name>
          , C. Breitinger,
          <article-title>Research-paper recommender systems: A literature survey</article-title>
          ,
          <source>International Journal on Digital Libraries</source>
          <volume>17</volume>
          (
          <year>2016</year>
          )
          <fpage>305</fpage>
          -
          <lpage>338</lpage>
          . doi:
          <volume>10</volume>
          .1007/ s00799-015-0156-0.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>P. A.</given-names>
            <surname>Bonatti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Decker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Polleres</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Presutti</surname>
          </string-name>
          , Knowledge Graphs:
          <article-title>New Directions for Knowledge Representation on the Semantic Web (Dagstuhl Seminar 18371), Dagstuhl Reports 8 (</article-title>
          <year>2019</year>
          )
          <fpage>29</fpage>
          -
          <lpage>111</lpage>
          . doi:
          <volume>10</volume>
          .4230/DagRep.8.9.29.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>A.</given-names>
            <surname>Hogan</surname>
          </string-name>
          , E. Blomqvist,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cochez</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>D'amato</article-title>
          , G. D.
          <string-name>
            <surname>Melo</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Gutierrez</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Kirrane</surname>
            ,
            <given-names>J. E. L.</given-names>
          </string-name>
          <string-name>
            <surname>Gayo</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Navigli</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Neumaier</surname>
            ,
            <given-names>A.-C. N.</given-names>
          </string-name>
          <string-name>
            <surname>Ngomo</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Polleres</surname>
            ,
            <given-names>S. M.</given-names>
          </string-name>
          <string-name>
            <surname>Rashid</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Rula</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Schmelzeisen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Sequeda</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Staab</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Zimmermann</surname>
          </string-name>
          ,
          <source>Knowledge Graphs, ACM Computing Surveys</source>
          <volume>54</volume>
          (
          <year>2021</year>
          )
          <fpage>1</fpage>
          -
          <lpage>37</lpage>
          . doi:
          <volume>10</volume>
          .1145/3447772.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>W.</given-names>
            <surname>Ammar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Groeneveld</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bhagavatula</surname>
          </string-name>
          , I. Beltagy,
          <string-name>
            <given-names>M.</given-names>
            <surname>Crawford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Downey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dunkelberger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Elgohary</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Feldman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Ha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kinney</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kohlmeier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Murray</surname>
          </string-name>
          , H.
          <string-name>
            <surname>-H. Ooi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Peters</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Power</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Skjonsberg</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Wilhelm</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Yuan</surname>
          </string-name>
          , M. van
          <string-name>
            <surname>Zuylen</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          <string-name>
            <surname>Etzioni</surname>
          </string-name>
          ,
          <article-title>Construction of the Literature Graph in Semantic Scholar, in: Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</article-title>
          , Volume
          <volume>3</volume>
          (
          <string-name>
            <surname>Industry</surname>
            <given-names>Papers)</given-names>
          </string-name>
          ,
          <source>Association for Computational Linguistics</source>
          , New Orleans - Louisiana,
          <year>2018</year>
          , pp.
          <fpage>84</fpage>
          -
          <lpage>91</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>N18</fpage>
          -3011.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Qian</surname>
          </string-name>
          ,
          <article-title>Finding relevant papers based on citation relations</article-title>
          , in: H.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Oyama</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Hu</surname>
          </string-name>
          , T. Qian (Eds.),
          <source>Web-Age Information Management</source>
          , volume
          <volume>6897</volume>
          LNCS, Springer Berlin Heidelberg, Berlin, Heidelberg,
          <year>2011</year>
          , pp.
          <fpage>403</fpage>
          -
          <lpage>414</lpage>
          . doi:
          <volume>10</volume>
          .1007/ 978-3-
          <fpage>642</fpage>
          -23535-1_
          <fpage>35</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>S.</given-names>
            <surname>Auer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Oelen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Haris</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Stocker</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. D'Souza</surname>
            ,
            <given-names>K. E.</given-names>
          </string-name>
          <string-name>
            <surname>Farfar</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Vogt</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Prinz</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Wiens</surname>
            ,
            <given-names>M. Y.</given-names>
          </string-name>
          <string-name>
            <surname>Jaradeh</surname>
          </string-name>
          ,
          <article-title>Improving Access to Scientific Literature with Knowledge Graphs</article-title>
          ,
          <source>Bibliothek Forschung und Praxis</source>
          <volume>44</volume>
          (
          <year>2020</year>
          )
          <fpage>516</fpage>
          -
          <lpage>529</lpage>
          . doi:
          <volume>10</volume>
          .1515/bfp-2020-
          <year>2042</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>A.</given-names>
            <surname>Allot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Luo</surname>
          </string-name>
          ,
          <string-name>
            <surname>Z.</surname>
          </string-name>
          <article-title>lu, LitSuggest: A web-based system for literature recommendation and curation using machine learning</article-title>
          ,
          <source>Nucleic Acids Research</source>
          <volume>49</volume>
          (
          <year>2021</year>
          )
          <fpage>W352</fpage>
          -
          <lpage>W358</lpage>
          . doi:
          <volume>10</volume>
          .1093/nar/gkab326.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>C. L.</given-names>
            <surname>Giles</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. D.</given-names>
            <surname>Bollacker</surname>
          </string-name>
          , S. Lawrence,
          <string-name>
            <surname>Citeseer:</surname>
          </string-name>
          <article-title>An automatic citation indexing system</article-title>
          ,
          <source>in: Proceedings of the Third ACM Conference on Digital Libraries</source>
          , DL '98,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>1998</year>
          , pp.
          <fpage>89</fpage>
          -
          <lpage>98</lpage>
          . doi:
          <volume>10</volume>
          .1145/276675. 276685.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>S. E.</given-names>
            <surname>Middleton</surname>
          </string-name>
          ,
          <string-name>
            <surname>D. C. De Roure</surname>
            ,
            <given-names>N. R.</given-names>
          </string-name>
          <string-name>
            <surname>Shadbolt</surname>
          </string-name>
          ,
          <article-title>Capturing knowledge of user preferences: Ontologies in recommender systems</article-title>
          ,
          <source>in: Proceedings of the 1st International Conference on Knowledge Capture</source>
          , number January in K-CAP '
          <fpage>01</fpage>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2001</year>
          , p.
          <fpage>100</fpage>
          -
          <lpage>107</lpage>
          . doi:
          <volume>10</volume>
          .1145/500737.500755.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>D. M. Pennock</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Horvitz</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Lawrence</surname>
            ,
            <given-names>C. L.</given-names>
          </string-name>
          <string-name>
            <surname>Giles</surname>
          </string-name>
          ,
          <article-title>Collaborative filtering by personality diagnosis: A hybrid memory and model-based approach</article-title>
          ,
          <source>in: Proceedings of the 16th Conference on Uncertainty in Artificial Intelligence</source>
          , UAI '
          <fpage>00</fpage>
          , Morgan Kaufmann Publishers Inc., San Francisco, CA, USA,
          <year>2000</year>
          , p.
          <fpage>473</fpage>
          -
          <lpage>480</lpage>
          . URL: https://dl.acm.org/doi/10.5555/647234. 720062.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>T. Y.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <surname>G. I. McCalla</surname>
          </string-name>
          ,
          <article-title>The pedagogical value of papers: a collaborative-filtering based paper recommender</article-title>
          ,
          <source>J. Digit. Inf</source>
          .
          <volume>10</volume>
          (
          <year>2009</year>
          ). URL: http://dblp.uni-trier.de/db/journals/jodi/ jodi10.html#TangM09.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>F. van Harmelen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ten
            <surname>Teije</surname>
          </string-name>
          ,
          <article-title>A boxology of design patterns for hybrid learning and reasoning systems</article-title>
          ,
          <source>Journal of Web Engineering</source>
          <volume>18</volume>
          (
          <year>2019</year>
          )
          <fpage>97</fpage>
          -
          <lpage>124</lpage>
          . doi:
          <volume>10</volume>
          .13052/ jwe1540-
          <fpage>9589</fpage>
          .18133, abstract in CEUR vol.
          <volume>2491</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>H. F.</given-names>
            <surname>Witschel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Pande</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Martin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Laurenzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Hinkelmann</surname>
          </string-name>
          ,
          <article-title>Visualization of patterns for hybrid learning and reasoning with human involvement</article-title>
          , in: R. Dornberger (Ed.),
          <source>New Trends in Business Information Systems and Technology: Studies in Systems, Decision and Control</source>
          , volume
          <volume>294</volume>
          , Springer International Publishing, Cham,
          <year>2021</year>
          , pp.
          <fpage>193</fpage>
          -
          <lpage>204</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -48332-6_
          <fpage>13</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>V. K.</given-names>
            <surname>Vaishnavi</surname>
          </string-name>
          ,
          <source>Design Science Research Methods and Patterns: Innovating Information and Communication Technology</source>
          , 1st ed.,
          <string-name>
            <surname>Auerbach</surname>
            <given-names>Publications</given-names>
          </string-name>
          , New York,
          <year>2007</year>
          . doi:
          <volume>10</volume>
          . 1201/9781420059335.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>B.</given-names>
            <surname>Sturm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sunyaev</surname>
          </string-name>
          ,
          <article-title>Design Principles for Systematic Search Systems: A Holistic Synthesis of a Rigorous Multi-cycle Design Science Research Journey</article-title>
          ,
          <source>Business Information Systems Engineering</source>
          <volume>61</volume>
          (
          <year>2019</year>
          )
          <fpage>91</fpage>
          -
          <lpage>111</lpage>
          . doi:
          <volume>10</volume>
          .1007/s12599-018-0569-6.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>C.</given-names>
            <surname>Pande</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. F.</given-names>
            <surname>Witschel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Martin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Montecchiari</surname>
          </string-name>
          ,
          <article-title>Hybrid conversational ai for intelligent tutoring systems</article-title>
          , in: A.
          <string-name>
            <surname>G. D. L. R. S</surname>
          </string-name>
          . F. v. H.
          <string-name>
            <surname>A. M. K. Hinkelmann</surname>
          </string-name>
          , H.-G. Fill (Ed.),
          <source>Proceedings of the AAAI 2021 Spring Symposium on Combining Machine Learning and Knowledge Engineering (AAAI-MAKE</source>
          <year>2021</year>
          ), Stanford University, Palo Alto, California, USA, March
          <volume>22</volume>
          -24,
          <year>2021</year>
          , volume
          <volume>2846</volume>
          <source>of CEUR Workshop Proceedings, CEUR-WS.org</source>
          ,
          <year>2021</year>
          . URL: http://ceur-ws.
          <source>org.</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>