<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Building a Knowledge Graph for Recommending Experts</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Pittsburgh</institution>
          ,
          <addr-line>Pittsburgh PA 15260</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Identifying experts is an important challenge in many contexts. In this paper, we present a method to build a knowledge graph by integrating data from Google Scholar and Wikipedia to help students nd a research advisor or thesis committee member. This knowledge graph is used to power the exploratory search interface to recommend similar</p>
      </abstract>
      <kwd-group>
        <kwd>Data Integration</kwd>
        <kwd>Knowledge Graph</kwd>
        <kwd>Recommender Systems</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>Identifying experts is an important challenge in many contexts. The nature of
this challenge is to nd a knowledgeable person with an advance expertise in
one or more target topics among a large number of potential candidates. A
wellexplored example of this task is nding an expert for a speci c project within a
large company or nding a doctor with advance knowledge of a speci c disease
in a large city. While in these two contexts, large companies and hospitals use
knowledge management techniques to catalogue key areas of expertise and use
it to represent information about each expert, nding experts in other contexts
could be more challenging.</p>
      <p>The context that we target in this paper involves students nding a research
advisor. Each year, undergraduates, master-level and doctorate students face
the di cult challenge of nding a research advisor. While large universities
have many highly knowledgeable faculty, nding one with the expertise that
matches the student's interests, requirements, and preparation is a challenging
task. Whether the task is nding an advisor for a summer research project, a
faculty sponsor for an independent study, or a committee member for a doctoral
thesis, online sources frequently fail the students, and they resort to 'word of
mouth' within a limited circle of instructors, classmates, and university sta .
One problem in using online sources is the wide variety of sources with relevant
information that can exist (e.g., department directories, publication sites,
funding agency pages, personal home pages, etc.) Each of these sources covers only
some aspects of the faculty member's expertise and frequently represent only
a subset of available advisors. Despite these di erent sources, there is typically
a lack of "expertise catalogs". A university usually o ers a catalog of courses
and majors, but not a ne-grained catalog of expertise areas covered by faculty.
As a result, students frequently cannot even properly name their target area of
interest or formulate a Web search query when looking for advisors.</p>
      <p>The focus of our project is to o er a single-access-point exploratory search
system, which allows students to discover their target areas of interest and nd
relevant advisors within these areas. In its core, the platform uses a knowledge
expertise graph, which represents multiple connections between research topics
and prospective research advisors within a large university or a large research
eld. We built this graph by processing several knowledge sources about faculty
and their research interests. This paper brie y reviews the type of knowledge
graph we built, the process of extracting information for its development, and
the information exploration system powered by this knowledge.
2</p>
    </sec>
    <sec id="sec-2">
      <title>BACKGROUND</title>
      <p>
        In the past, there have been attempts to build \a map of science" representing
most important areas of research expertise and their connections with experts;
however, the lack of proper information sources makes it hard to produce maps
that are suitable for nding advisors. Examples of this attempt to build a map
of science is presented in [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] and [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], where academic journals are used as
proxies of expertise areas, and a map of science is built by clustering journals by
co-publication links. While this map is useful as a \big picture" of science, its
use in the context of nding research advisors is problematic since it represents
expertise on a very coarse-grain level and does not capture many prospective
advisors who are not frequent journal authors. However, the emergence of modern
sites powered by a combination of advanced information processing and
collective wisdom makes the task of building a ne-grain knowledge network of experts
and expertise areas feasible. In our work we rely most extensively on two of these
sites - Google Scholar and Wikipedia.
      </p>
      <p>
        Google Scholar has been long recognized as one of the best freely accessible
academic information sources in terms of coverage and accessibility [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. It has
been compared positively with a number of similar citation services namely Web
of Science [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], PubMed, and Scopus [
        <xref ref-type="bibr" rid="ref5 ref7">7, 5</xref>
        ]. Yet, although Google Scholar contains
nearly 160 million documents [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] covering a large portion of published documents,
the lack of semantic connections between concepts and keywords within these
documents makes it di cult to use the system for nding advisors, especially by
less experienced students.
      </p>
      <p>
        Wikipedia is commonly used by researchers to compute the semantic
relatedness of concepts between and within documents [
        <xref ref-type="bibr" rid="ref1 ref8 ref9">9, 8, 1</xref>
        ], extract Open
Information [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], and mine meaning using relations, facts, and descriptions to extract
and use concepts [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>BUILDING THE KNOWLEDGE GRAPH</title>
      <p>To support students in nding advisors, we created a knowledge graph using
data from Google Scholar and enriched it semantically using Wikipedia. In turn,
this graph was used to power an interactive exploratory recommendation
interface, which makes the task of advisor- nding easier, especially for students with
a limited level of knowledge and familiarity with the subject of research. To
support several advisor- nding scenarios, we built several versions of the graph.
The graph presented in this paper is focused on the task of nding a top expert
in a speci c topic of interest within some broad eld of research (such as
Articial Intelligence) across many universities. This is a typical task for a student
selecting a doctoral program to join or for a senior doctoral student looking for
an external thesis committee member.
3.1</p>
      <sec id="sec-3-1">
        <title>Data Sources</title>
        <p>Google Scholar We utilized the information of 1000 active scholars in two
popular elds of computer science: Arti cial Intelligent and Computer
Architecture (focusing on the top 500 scholars in each eld). For each individual, we
extracted the following information (see Table 1):
{ Name: Full name of the scholar.
{ A liation: The university or research institution the scholar is a liated
with.
{ Veri ed Email Domain: Used to check the validity of the scholar pro le.
{ Self-De ned Keywords: A list of up to ve keywords de ned by scholars
to describe their research interests.
{ Citations: The total number of citations received by all of the scholar's
publications.
{ h-index: The h-index measures the citation impact and productivity of
a scholar's publications. We use this measure alongside other quantitative
scores to re-order the results of the recommendations.
{ i10-Index: i10-Index describes the total number of the scholar's publications
with 10 citations or more. This score, which is only used by Google Scholar,
was also used to re-rank the results of the recommendations.
{ Recent publications (20): We used the 20 most recent publications to
generate additional keywords representing the current interests of each scholar.
The keywords were extracted from the titles of recent publications as follows:
After removing stop-words, we generated all of the possible keyword
candidates as uni-grams, bi-grams, and tri-grams. Next, we only kept the keywords
that have an entry in Wikipedia (see keyword veri cation below).
{ Top Co-Authors (10): For each scholar, we extracted a list of the top 10
co-authors from their Google Scholar pro le.
Wikipedia We used Wikipedia to add a semantic layer to pro les extracted
from Google Scholar. Throughout this process, we also obtained useful
information that led to a stronger connection between keywords and enables us to add
weight to each scholar-keyword relation. The Wikipedia API has been used for
the following purposes:
{ Keyword veri cation: As mentioned before, we collected two sets of
keywords for each scholar: self-de ned keywords and keywords extracted from
recent publications. The Wikipedia API has been used to verify the validity
of these keywords by using fuzzy match techniques to nd the Wikipedia
entry describing the keyword. We removed all keywords that did not match
with any article in Wikipedia. While Wikipedia might miss articles for some
less popular research topics, we need to have all topic keywords explained for
the student audience and a match to a Wikipedia article was the best way to
assure it. For all remaining keywords, we calculated the association weight
between a keyword and a scholar as cosine similarity between the full-text
Wikipedia entry of each keyword and concatenated text from the scholar's
recent publications.
{ Entry Summary: To o er student users a short description of each topic
keyword, we collected page summaries for all keywords using the Wikipedia
API.
{ Top relevant keywords (10): Most (if not all) Wikipedia pages have
multiple links to similar or related articles. We collected the top 10 links
based on the number of their occurrences in each page. We employ these
links to create a highly connected network of keywords.
{ Entry Categories: Wikipedia uses categories to group similar articles. We
extracted all categories associated with a page and used a full Wikipedia
category hierarchy schema to nd relationships between categories in our
data set.</p>
      </sec>
      <sec id="sec-3-2">
        <title>Graph Representation</title>
        <p>We used the Neo4j graph database to represent information about all scholars
and keywords. The overall schema of the knowledge graph is represented in
Figure 1. As is shown, there are three distinct node types in the graph:</p>
        <p>{ Individual: This node type conveys demographic information about
scholars including full name, a liation, veri ed email domain, and URLs of
personal homepages and Wikipedia pages (if they exist). Individual nodes are
connected via "works with" links which represent the co-authorship relations
between scholars. An individual node also connects to several keyword nodes
that represent the scholar's research interests and expertise. The individual
nodes with a dashed border represent scholars added via co-authorship
extraction who are not among the top 500 extracted scholars. These nodes are
not considered in the nal recommendations and only used to indicate the
connections of the top scholars.
{ Keyword: There are three types of keyword nodes. Self-De ned keywords,
keywords extracted from recent publications, and relevant keywords that
represent the connection between two other types (shown by a dashed border)
and will not appear in the recommendations. The relationship between
keywords represented by "has link" arc is established if the target node has been
mentioned in the source node's entry page. Keyword nodes are connected to
individual nodes via "has key" and to category nodes via "belongs to" arcs.
{ Category: We employed a full hierarchical schema of Wikipedia categories
to represent the inter-connectivity between categories in our data set. These
relations are presented as the "has child" arc in Figure 1. The category nodes
are used to nd the semantic relationship between keywords.</p>
        <p>https://en.wikipedia.org/wiki/Neo4j</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>USING THE KNOWLEDGE GRAPH</title>
      <p>4.1</p>
      <sec id="sec-4-1">
        <title>Interface Design</title>
        <p>The knowledge graph is used to power the exploratory search interface for nding
advisors. The interface consists of four main sections.</p>
        <p>Instant search box Users can use the search box to search for topic keywords or
scholars of interest (Figure 2:B). When the user starts typing, a list of matching
keywords and scholars appears. When an item is selected from the list, it is
automatically added to the proper location on the left or the right side of the
interface. At the same time, an updated list of recommendations is presented to
the user.
Favorite keywords This section (Figure 2:A) shows users' \favorite" keywords.
Users add keywords to this list using the instant search box or by clicking on the
plus button next to each recommended keyword. Users can interact with three
buttons on the right side of each keywords to (1) see more information about that
keyword (including its Wikipedia summary, similar keywords, and other scholars
with this research interest), (2) remove the keyword from the favorite list, and
(3) enable/disable the e ect of this keyword on the list of recommendations.
Favorite scholars Similar to favorite keywords, users can assemble a list of
favorite scholars (Figure 2:C). A new favorite scholar can be added to the list
from the instant search results or a list or recommended scholars by clicking
on the plus button next to a recommended scholar. The three buttons on the
right side of each favorite scholar can be used to obtain more details about the
scholar (a liation, full list of research interests, and similar scholars), remove
the scholar from the list, and enable/disable the e ect of this scholar in the
nal recommendation. Together with the favorite keywords list explained above,
the list of favorite scholars form the users' pro les of interests, which the users
gradually assemble while exploring possible areas of interests and scholars. In
turn, the pro les of interests are used to generate further recommendations as
explained below.</p>
        <p>Recommendations This section (Figure 2:D) consists of two subsections.
Recommended keywords (Figure 2:D1) shows the list of the three most relevant
additional keywords, which are suggested given already selected (and enabled)
favorite keywords and scholars. Users can see more information about the
keyword (similar to the favorite keywords section) and also add these recommended
keywords to their favorite lists using two circular buttons on the right side of
each keyword. Recommended scholars (Figure 2:D2) shows a list of recommended
scholars, which are most relevant to the active (enabled) favorite topic keywords
and most similar to the active favorite scholars. For each recommended scholar,
the list shows basic personal and academic information. Users can also see the
similarity between the recommended scholars and their pro les of interests
represented by favorite keywords and scholars.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Recommendation method</title>
        <p>We generate the recommendations using Cypher Query Language in Neo4j. In
the following we explain how we generate recommendations for keywords and
scholars.</p>
        <p>Keyword Recommendations In order to recommend similar keywords, we
use the user's favorite keywords and scholars. Each keyword is connected to other
keywords in two ways: (1) via the similar research interest between scholars and
(2) via similar relevant keywords and categories. We consider both of these
relations to nd similar keywords. In the nal list, we sorted the keywords based on
the number of occurrences then we chose the top three keywords to be presented
to the user.</p>
        <p>Scholar Recommendations Similar to keyword recommendations, we use
both favorite keywords and scholars. There are three criteria for scholar
recommendations: the scholar's weighted research interests, co-authorship relationship
between scholars, and connection between the scholar's interests through
relevant keywords and categories. After generating the list of candidate scholars,
we sort it based on the similarity score (calculated based on weighted similarity
score for each of the three criteria) and present the top ten results to the user.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>DISCUSSION AND FUTURE</title>
    </sec>
    <sec id="sec-6">
      <title>WORK</title>
      <p>We presented a method to build a knowledge graph by integrating data from
Google Scholar and Wikipedia to help students with limited knowledge about a
subject nd a research advisor or thesis committee member. Although Google
scholar covers a variety of publications and patents, additional sources of
information (e.g., the scholar's active research projects, funding information, etc.)
could make the knowledge graph more connected and provide the users with
additional critical information when it comes to nding an advisor. We plan to
re ne our keyword extraction techniques. More sophisticated methods of
extraction using natural language processing and machine learning could potentially
improve the semantic relations between concepts and provide users with a more
realistic set of research interests for scholars. We have also designed a series of
controlled user studies and eld studies to evaluate the usability and value of
the exploratory search interface. We hope that these user studies will provide
valuable insights for improving the knowledge graph and the interface.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Behnam</given-names>
            <surname>Rahdari</surname>
          </string-name>
          and
          <string-name>
            <given-names>Peter</given-names>
            <surname>Brusilovsky</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>User-controlled hybrid recommendation for academic papers</article-title>
          .
          <source>In Proceedings of the 24th International Conference on Intelligent User Interfaces: Companion. ACM</source>
          ,
          <volume>99</volume>
          {
          <fpage>100</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Colin</given-names>
            <surname>Murray</surname>
          </string-name>
          , Weimao Ke, and Katy Borner.
          <year>2006</year>
          .
          <article-title>Mapping Scienti c Disciplines and Author Expertise Based on Personal Bibliography Files</article-title>
          .
          <source>In Tenth International Conference on Information Visualisation (IV'06)</source>
          .
          <volume>258</volume>
          {
          <fpage>263</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Enrique</given-names>
            <surname>Ordun</surname>
          </string-name>
          <article-title>~a-Malea, Juan Manuel Ayllon, Alberto Mart n-Mart n</article-title>
          ,
          <source>andEmilio Delgado Lopez-Cozar</source>
          .
          <year>2014</year>
          .
          <article-title>About the size of Google Scholar: playing the numbers</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Fei</given-names>
            <surname>Wu</surname>
          </string-name>
          and
          <string-name>
            <given-names>Daniel S.</given-names>
            <surname>Weld</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Open Information Extraction Using Wikipedia</article-title>
          .
          <source>In Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics (ACL '10)</source>
          .
          <article-title>Association for Computational Linguistics</article-title>
          , Stroudsburg,PA, USA,
          <volume>118</volume>
          {
          <fpage>127</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>John</given-names>
            <surname>Mingers</surname>
          </string-name>
          and Martin Meyer.
          <year>2017</year>
          .
          <article-title>Normalizing Google Scholar data for use in research evaluation</article-title>
          .
          <source>Scientometrics</source>
          <volume>112</volume>
          ,
          <issue>2</issue>
          (
          <year>2017</year>
          ),
          <volume>1111</volume>
          {
          <fpage>1121</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Joost CF De Winter</surname>
          </string-name>
          , Amir A.
          <string-name>
            <surname>Zadpoor</surname>
            , and
            <given-names>Dimitra</given-names>
          </string-name>
          <string-name>
            <surname>Dodou</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>The expansion of Google Scholar versus Web of Science: a longitudinal study</article-title>
          .
          <source>Scientometrics</source>
          <volume>98</volume>
          ,
          <issue>2</issue>
          (
          <year>2014</year>
          ),
          <volume>1547</volume>
          {
          <fpage>1565</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Matthew</surname>
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Falagas</surname>
            ,
            <given-names>Eleni I. Pitsouni</given-names>
          </string-name>
          , George A.
          <string-name>
            <surname>Malietzis</surname>
            , and
            <given-names>Georgios</given-names>
          </string-name>
          <string-name>
            <surname>Pappas</surname>
          </string-name>
          .
          <year>2008</year>
          . Comparison of PubMed, Scopus, Web of Science, and Google Scholar:
          <article-title>Strengths and weaknesses</article-title>
          .
          <source>The FASEB Journa l22</source>
          ,
          <volume>2</volume>
          (
          <year>2008</year>
          ),
          <volume>338</volume>
          {
          <fpage>342</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8. Max Volkel, Markus Krotzsch, Denny Vrandecic, Heiko Haller, and
          <string-name>
            <given-names>Rudi</given-names>
            <surname>Studer</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Semantic Wikipedia</article-title>
          .
          <source>In Proceedings of the 15th international conference on World Wide Web</source>
          .
          <volume>585</volume>
          {
          <fpage>594</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>Michael</given-names>
            <surname>Strube</surname>
          </string-name>
          and Simone Paolo Ponzetto.
          <year>2006</year>
          .
          <article-title>WikiRelate! Computing semantic relatedness using Wikipedia</article-title>
          .
          <source>InAAAI</source>
          , Vol.
          <volume>6</volume>
          . 1419{
          <fpage>1424</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Olena</surname>
            <given-names>Medelyan</given-names>
          </string-name>
          , David Milne,
          <string-name>
            <given-names>Catherine</given-names>
            <surname>Legg</surname>
          </string-name>
          , and
          <string-name>
            <surname>Ian</surname>
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Witten</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Mining meaning from Wikipedia</article-title>
          .
          <source>International Journal of Human-Computer Studies 67</source>
          ,
          <issue>9</issue>
          (
          <year>2009</year>
          ),
          <volume>716</volume>
          {
          <fpage>754</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>Philipp</given-names>
            <surname>Mayr</surname>
          </string-name>
          and
          <string-name>
            <surname>Anne-Kathrin Walter</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>An exploratory study of Google Scholar</article-title>
          .
          <source>Online Information Review</source>
          <volume>31</volume>
          ,
          <issue>6</issue>
          (
          <year>2007</year>
          ),
          <volume>814</volume>
          {
          <fpage>830</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12. Richard M.
          <article-title>Shi rin and Katy Borner</article-title>
          .
          <year>2004</year>
          .
          <article-title>Mapping knowledge domains</article-title>
          .
          <source>Proceedings of the National Academy of Sciences101, suppl 1</source>
          (
          <year>2004</year>
          ),
          <volume>5183</volume>
          {
          <fpage>5185</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>