<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Semantic-based Expert Search in Textbook Research Archives</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Marco Pavan</string-name>
          <email>P@10</email>
          <email>P@3</email>
          <email>P@5</email>
          <email>marco.pavan@uniud.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ernesto William De Luca</string-name>
          <email>deluca@gei.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Georg-Eckert-Institut - Leibniz-Institute for international Textbook Research</institution>
          ,
          <addr-line>Celler Stra e 3, Braunschweig</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Udine</institution>
          ,
          <addr-line>Via delle Scienze 206, Udine</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2015</year>
      </pub-date>
      <fpage>18</fpage>
      <lpage>29</lpage>
      <abstract>
        <p>Expert nding and the identi cation of similar professionals are important tasks for many services provided by companies and institutions. Nowadays, the rapid growth of web services and social and professional networks, allowed di erent kind of users to share personal data and increased the amount of information available. Most of research works focus on a limited set of users, characterized by the same kind of main activities, e.g., researchers, or exploit external knowledge, such as prede ned ontologies. An heterogeneous environment, with possible lack of information, and not well structured data, puts forward new challenges, to address the problem of adapting user pro ling and consequently expert search. In this paper, we rst provide a general perspective on studies on expert nding, similar people identi cation and social recommender systems, highlighting some critical issues related to information extraction and user pro le de nition. We then present a rst attempt to create an expert search system to support users in Library and Archiving Communities, such as researchers, students, authors, in the eld of Textbook Research, in nding other experts to get in contact with or to start a cooperation. This is organized in three phases: rst, we collect information about users in order to build structured pro les; then, we build a Community Knowledge Graph (CKG) which de nes relationships and weights among terms that occur in the pro le sections, emphasizing information shared by users in the entire analyzed community. As third step, we exploit the CKG structure to de ne similarity values among users based on weights got by their common terms in the CKG, and their distances in the graph. We conjecture that the CKG allows to model users emphasizing new semantic aspects of relationships among pro le elements, and helps to improve the similarity computation and the expert search. We present a preliminar experimental evaluation on real users.</p>
      </abstract>
      <kwd-group>
        <kwd>Semantic Enrichment</kwd>
        <kwd>Expert Search</kwd>
        <kwd>User Modeling</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>2
1</p>
      <p>Marco Pavan and Ernesto William De Luca</p>
    </sec>
    <sec id="sec-2">
      <title>Introduction</title>
      <p>In recent years, the increasing pervasiveness of social network platforms has
made users able to share personal data, also for professional purposes, in order
to nd positions and collaborations o ering competence and knowledge. Also
several companies and institutions have strong interest in exploiting that kind
of information to nd people with particular skills and expertise that t their
needs. With these premises it is clear how important is addressing the expert
search problem and exploiting new sources of potentially interesting
information, therefore, accurate expert search systems enable users and companies to
quickly nd the right desirable people without being overwhelmed by irrelevant
information or losing too much time seeking on several web platforms.</p>
      <p>
        There has been an increasing interest among researchers in improving the
search results of expert nding, and several solutions had been proposed, even
recently [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. However, there are still many problems when adopting
those solutions in real situation.
      </p>
      <p>
        Most of works in expert search focus on a single assumption about expert
similarity (i.e., researchers are similar if they have similar papers, or similar
topics; people on Q&amp;A platforms are similar if their questions and answers cover
similar topics, etc.). While this is working well on speci c platforms for
speci c tasks, a more general and adaptable framework could be needed to take in
care multiple possible sources for user similarity, i.e. research papers if available,
tweets if users use Twitter, etc., and to allow us to build an adaptive system
able to handle di erent search tasks. In real situations users could have di erent
needs and di erent approaches on expert search. Someone could be an expert
who is looking for partners with same skills, thus, she could try to search pro les
similar to herself; on the other hand, an inexperienced user, such as a student,
might need general knowledge about a topic, with no particular requirements on
user expertise, just with the su cient background or interest. Yet others, could
be interested in nding users that have collected a lot of information about a
topic, thus, not related to work or skills, but with great interest on a subject,
such as a particular sport, political party, etc. On this basis, it is clear how a
more general pro ling could help, and how information about users obtained
from heterogeneous information sources could enrich the pro le, providing
different insights on user similarity. Recent studies [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] demonstrate
that information from social networks can be exploited to improve accuracy of
recommendations and user similarities, therefore it seems reasonable to continue
on this research direction.
      </p>
      <p>Another critical issue is how to get the correct knowledge base, that helps
to connect texts to semantic entities, such as categories or topics, and allows us
to properly model users and build their pro les. Research works, focused on a
single domain, exploit speci c ontologies or databases to address this problem,
but obviously this approach is limited to a set of users of the same type, for
instance researchers in a speci c eld, employees in a particular company, or
customers interested in speci c products.</p>
      <p>One of the major challenges of digital archiving for Textbook Research is
how to deal with changing technologies and changing user communities, which
necessitate tools to formalize, detect and measure knowledge evolution.
Semantic representations of contextual knowledge about cultural heritage objects and
users, especially in Textbook Research, will enhance organization and access
of data and knowledge, because usually the relationships among them are not
emphasized or even identi ed.</p>
      <p>In this paper, we focus on a novel semantic-Blser approach for expert search.
We create a Community Knowledge Graph (CKG) used for semantic
enrichment. The graph is dynamically built during the users set analysis and allows
to de ne user similarities by comparing the relationships among terms and
entities. In particular, we collected structured information similarly to a CV, and
taking into account the terms occuring together in di erent elds to emphasize
their relationships to other users sharing the same content. Then we build the
knowledge behind their pro les. Indeed, we observed that some users who work
on similar projects might have similar biographical data, or share same
interests. On the other hand, students with interests on particular topic, could be
connected to experts through similar pro le content. With a speci c ontology
or database the community analysis is limited and it is not possible to properly
model a general user who could seek for experts with di erent pro le but with
some intersection.</p>
      <p>The novelty of our approach is the semantic enrichment of the pro les based
on a community graph in order to improve the results of similarities scores
among users. We exploit the CKG structure, with its weighted relationships,
and we compare users considering the graph distances of their common terms in
the CKG.</p>
      <p>The paper is structured in 6 sections. After the introduction, we discuss
related work in Section 2. We focus on the problem statement in Section 3
and describe our approach in Section 4. Section 5 describes our preliminary
evaluation and Section 6 concludes the paper and shows our future work.
2</p>
    </sec>
    <sec id="sec-3">
      <title>Related Work</title>
      <p>
        The expert search problem has arised great interest among researchers, and
several groups addressed the main problems behind this important topic, such as
lack of information about users, or de ning supporting knowledge needed to
build models and improve the results. Gollapalli et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] propose a graph-based
model for expertise retrieval with the objective of enabling search using either a
topic or a name. El-korany [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] proposes a novel cascaded model for expert
recommendation using aggregated knowledge. He exploits social networks in order
to extract useful contents for building a vector space model. With this approach
he computes the relevance of contents respect to a speci c query and applies
PageRank algorithm to rank candidate experts. Fang et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] investigate how
to merge and weight heterogeneous knowledge sources to improve the expert
nding process. Cetintas et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] propose a discriminative probabilistic model
4
      </p>
      <p>
        Marco Pavan and Ernesto William De Luca
that identi es latent content and graph classes for people with similar pro le.
Moreira et al. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] exploit unsupervised rank aggregation methods. They
combine multiple estimators of expertise, derived from the textual contents, from
a graph-structure of the citation patterns, and information extracted from user
pro les. Yang et al. [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] utilize Normalized Google Distance (NGD) to enhance
the relevance between initial query and extended query, and to improve the
accuracy of the search results of the expert nding system.
      </p>
      <p>
        Other researchers focus on speci c platform, such as Twitter, to address
speci c expert search problem. Stankovic at al. [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] propose a method for suggesting
potential collaborators for solving challenges online, based on their competence
and interests. Yet others focus on issues related to collaborative ltering. Spaeth
et al. [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] explore text-mining techniques to improve classical collaborative
ltering methods for matching people who are looking for expert advice on a speci c
topic. Gujral et al. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] try to go further and proposed a knowledge prototype
for expert knowledge synthesis. Plumbaum et al. [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] propose a personalized
recommendation application by combining semantic and structured information
available in research communities. In particular they exploit encyclopedic
knowledge sources, a large news article dataset, and collected implicit user feedback.
      </p>
      <p>
        All of these related works have signi cant value, given the importance of all
the issues addressed to improve expert search systems, and it is clear how one of
the most important issues underlying these systems is the user modeling and the
related eventual enrichment process. In this direction other authors recently
presented their work. Abel et al. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] propose a model based on Twitter posts linked
to related news articles to identify activities. Noureddine et al. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] address the
problem of researchers pro ling, exploiting structured and unstructured data
from di erent heterogenous web sources. They propose an ontology-based
architecture that correlates information coming from several sources. Also other
researchers exploited multiple external sources to address the pro le enrichment
problem. Orlandi, in two works [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] proposes a semantic approach for
interlinking social websites and provenance management on the Web of Data; and
a methodology for the automatic creation and aggregation of interoperable and
multi-domain user pro les of interests using semantic technologies. The author
build user pro les based on qualitative and quantitative measures about user
activities across social sites. Song [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] demonstrates how integrating multiple
social networks as external sources outperforms the use of only a single source,
by proposing a conceptual volunteering decision model. Mizzaro et al. [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]
exploit Twitter and Wikipedia as external sources for enriching short texts and
build a network-based user model to improve similarity scores. Al-Kouz et al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]
analyze the users' social graph and the users' interactions with attention on posts
and group memberships to model user interests and elds of expertise.
      </p>
      <p>
        Another important issue often present in research works related to expert
search is the use of a supporting knowledge base, usually de ned with an
ontology. Some researchers focused on ontology-based user interests modeling and
matching. Cena et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] propose an approach for propagating user interests on
in ontology-based user models, in order to solve the cold-start problem in
recommender systems by exploiting the ontological structure of the domain. Koh
et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] present an iconcept-matching approach to measure degrees of similarity
among users. They exploit Kullback-Leiber distance to measure similarity on
users represented by concept hierarchy.
3
      </p>
    </sec>
    <sec id="sec-4">
      <title>Problem statement</title>
      <p>To better understand what are the main problems and di culties emerging with
expert search tasks and related user pro ling, we list a set of conceptual problems
presented by the current state-of-the-art solutions. Many expert search systems,
and even recommender systems, as rst step, need to extract the starting
information about users in order to build the pro le. Most of existing systems
rely on a single data source, with the risk to have incomplete information about
users. Depending of the nature of that source, the obtained information could
be related on only one aspect of the users, i.e. only their work expertise or
skills, or purchases preferences; or it could be very schematic and represented by
very short texts, therefore likely with poor knowledge about them. So two rst
conceptual problems are:</p>
      <sec id="sec-4-1">
        <title>P1a: Poor pro le problem - lack of information. With a single data source,</title>
        <p>the user pro le could be incomplete and focused only on certain aspects, therefore
ignoring a more complete user overview.</p>
      </sec>
      <sec id="sec-4-2">
        <title>P1b: Poor pro le problem - short texts. With a single data source, the</title>
        <p>user pro le could be composed of short texts that make di cult the information
extraction, due to their brevity that does not provide su cient word occurrences.</p>
        <p>Moreover, the texts that compose the initial user data, most of time are
not structured, and users are represented as bag-of-words, with no information
about what kind of data could be related to personal data, expertise, or general
interests. We can de ne this problem as:</p>
      </sec>
      <sec id="sec-4-3">
        <title>P2: Structural problem - unstructured information. Extracted user data</title>
        <p>do not have structure that allow us to identify what kind of information we
have. This issue makes di cult the process of dividing information in sections
to properly model users under several aspects.</p>
        <p>
          Very recent works [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] have highlighted how an enrichment process can
overcome the lack of information and improve the results for several purposes.
Some other new works in the literature [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ], [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ], [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ], [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] introduced
multisource enrichment approaches in order overcome these problems. However some
techniques exploit ontologies or other prede ned knowledge bases as supporting
structure. This approach is useful if the set of user to analyzed has characterized
by the same interests, or they work in the same eld, or even they share similar
CV, but in heterogeneous environments such as Seek&amp;O er job platforms or
Q&amp;A services, a xed knowledge base could be not perfect suited for that
purpose. Moreover, the dynamic aspect of web platforms, where active communities
        </p>
        <p>Marco Pavan and Ernesto William De Luca
are not always the same, arise the need of having dynamic also the knowledge
support, evolving over time. On this basis we de ne other two problems:</p>
      </sec>
      <sec id="sec-4-4">
        <title>P3a: Supporting knowledge problem - prede ned and domain-dependent</title>
        <p>semantics. Ontologies and any supporting knowledge bases are usually crafted
for speci c domains, therefore they might not be well suited for any set of users.</p>
      </sec>
      <sec id="sec-4-5">
        <title>P3b: Supporting knowledge problem - xed semantic structure. Ontolo</title>
        <p>gies and any supporting knowledge bases are usually xed structures, therefore
not able to change over time in adaptable frameworks.</p>
        <p>Another important issue in expert search systems is related to user similarity
computation, used for de ning the scores between couples of users and obtain a
ranked list of suggestions that meet the needs expressed by the requesting user.
Most of the state-of-the-art proposed systems use a single function to compare
the user models they build, but in heterogeneous environments could be helpful
to emphasize only some aspects of users, in order to better t the expressed
need, taking in care what kind of user is the requester, or even analyzing for
what purpose is the request itself.</p>
      </sec>
      <sec id="sec-4-6">
        <title>P4: Similarity computation problem - Only one similarity score. The</title>
        <p>computation of only one similarity score between two users is a global value that
does not take in care the several aspects of a user pro le. Moreover, it does not
give di erent weights to part of pro le that need more or less emphasis, based
on the current request.</p>
        <p>
          Recent research works [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ], [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] addressed problems P1 and P2 exploiting
information from external sources to get useful additional data, therefore in this
paper we focus on problem P3, related to the supporting knowledge base in
heterogeneous environments.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Proposed approach</title>
      <p>In the following we describe our approach organized into three main phases.
During the rst phase we collect information about users in order to build structured
pro les; the second phase consists in building a graph composed of all texts from
all users, with attention on enriching the network with semantic relationships
among entities obtained by the pro les structure; in the last phase we compute
similarity scores among users exploiting the network structure that allows us to
consider distances among terms which represent users.</p>
      <p>Figure 1 shows an overall representation of our approach. Each phase is
described in full details in the following sections.
4.1</p>
      <sec id="sec-5-1">
        <title>Step 1: Structured user pro les</title>
        <p>We have chosen a real situation to analyze as case study, in order to then
investigate on the e ectiveness of our approach. We started a pilot study at the
Semantic-based Expert Search in Textbook Research Archives
7
Georg-Eckert-Institute (GEI), to collect data about people who work in the
international textbook research eld, but with di erent roles.</p>
        <p>A pilot study is a standard scienti c tool for testing a research question
in a \soft" way, allowing scientists to conduct a preliminary analysis before
starting a full-blown experiment. In our case, we decided to start a survey in
order to analyse how we can help researchers in nding other researchers that
are interested in a very speci c research area, namely Textbook Research. We
wanted to nd out how we can create semantic user pro les being di erent from
other research portals like \research gate" or \academia".</p>
        <p>We selected a sample of 32 users distributed as follows: 33% men, 67%
women (di erent nationalities); 18% with age under 30, 55% between 31 and
40, 27% more than 40; 70% graduates, 30% doctors; 67% with activities related
to research. We collected data related to biographical information, their current
position, projects where they are/were involved in, their interests, etc., all for
researchers, authors and students/trainees. We then structured information by
grouping terms which compose the collected textual descriptions in a tree
structure with a set of pre-selected entities. It is clear how an heterogeneous set of
users could led us to have di erent pro les with di erent kind of shared interests
or expertise have been collected into the community knowledge graph.</p>
        <p>Figure 2 shows an example of such a structured user pro le with the related
entities and terms.</p>
        <p>To test the feasibility, equipment and methods, we started the pilot study for
nding out what are the important relations that should be collected for
enriching the base pro le within semantic information that can be derived from the
needs of the researchers we asked to participate. For this study we focus on
digi</p>
        <p>Marco Pavan and Ernesto William De Luca
tal archives eld and in particular on the case study of Textbook Research, that
does not involve only researcher, but, as mentioned before, people with di erent
backgrounds and interests. Therefore, this preliminary test is also important for
training inexperienced users, explaining them how to interact with the system,
what is the knowledge we extract, in order to get useful information to nd out
what are the sub-group of experiments we have to take into account after this
pilot phase. Pilots are rapidly becoming an essential pre-cursor to many research
projects, in order to reduce costs. At the same time we could nd out how to
set the next experimental phase, because the users gave us feedback of unclear
statements or questions.
4.2</p>
      </sec>
      <sec id="sec-5-2">
        <title>Step 2: Community Knowledge Graph</title>
        <p>As second step we build a knowledge graph in order to have a structure that
represent the entire community. We de ne this structure as Community
Knowledge Graph (CKG). Based on the structure de ned for user pro les we keep the
same entities as nodes and we add directed edges in order to connect words used
by users to those entities which represent the important aspects that
characterize those users. We also add a new set of nodes in order to represent users and
connect them to the words they used. In this way it is possibile to analyze the
network by user. We scan all users' data in order to dynamically build the CKG
and semantically enrich it with higher weights on edges where the relationship
between a word and an entity is repeated. With this approach we emphasize the
important relationships inside a community facing the problems P3 described in
Section 3.</p>
        <p>More formally, let CKG = (V; E) the Community Knowledge Graph, where
V = fv1; v2; :::; vng is the set of all entities and words extracted from user pro les,
Semantic-based Expert Search in Textbook Research Archives
9
and E = fe1; e2; :::; emg is the set of undirected edges. We say that there exists
evi;vj 2 E , vj isRelatedTo vi. The property \isRelatedTo" could be intended
in several ways in order to de ne multiple semantic relationships. The
TermEntity relation is used for words belonging to the same concept, the Entity-Entity
relation represents concepts related to a more general category of information
which characterize users and the User-Term relation connects users with the
words they use.</p>
        <p>Figure 3 shows an example of CKG where it is possible to see how users
are connected to the words they use, words to Entities, and nally to the more
general entities that we call Categories. During the building process, if some
words or entities are already present in the model the edges weights are increased
in order to emphasize those relationships3. With this methodology we can build
the real structure that users' aspects have. For instance, if a lot of people with a
speci c skill work in a particular company, we can have that strong relationship
as high weight for the edge that connect the word with company name and the
entity which represent the skills.
4.3</p>
      </sec>
      <sec id="sec-5-3">
        <title>Step 3: Similarity computation</title>
        <p>The creation of the CKG lets us comparing di erent users extracting information
from their user pro les, de ning similarity scores between them. The network
structure allows us to have connections between words and entities based on
the actual relationships found in the original texts, therefore we extract only
the sharing nodes of two users, and we measure a similarity score based on the
weights of the relationships computed as in the following:</p>
        <p>Starting from two analyzed nodes vi and vj we follow the path in the graph
until they share a common node. If this is the case, then we compute the
similarity score based on the valid relationships, otherwise we discard it. In the
3 For an easier reading, in Figure 3 multiple edges are shown.</p>
        <p>Marco Pavan and Ernesto William De Luca
case there are multiple paths we consider the path with the highest total edges
weight, in order to emphasize the strong relationship between the two nodes that
CKG has extracted.</p>
        <p>In Figure 3, the purple dashed line highlights an example of selected path
for the weighted distance. Formally, Let Fi;j E be the set of edges into the
selected path from vi to vj de ned with our rule previously described, we de ne
w(vi; vj ), the weight of the relationships between vi and vj , as follows:
w(vi; vj ) =</p>
        <p>w(Fi;j )
plen(vi; vj )
where w(Fi;j ) is sum of edges weights into the respective set of edges, and
plen(vi; vj ) is the length of the path from vi to vj .</p>
        <p>By repeating this process for all couples of nodes we can get a global score
for user similarity.</p>
        <p>Let Vua V and Vub V be respectively the sets of nodes of hypothetical
user ua and user ub. We de ne the similarity score as follow:
sim(ua; ub) =</p>
        <p>X w(vi; vj )
with vi; vj 2 Vua _ vi; vj 2 Vub .
5</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Preliminary Evaluation</title>
      <p>In order to study the e ectiveness of our approach, we have set up a preliminary
experimental evaluation. We have run three di erent algorithms on the dataset
composed of the 32 users described in section 4.1. The rst one, called
bows, is a classic bag of words approach which exploit only the information users
provide about their biography, therefore with no additional information about
skills, current projects, interests. It simply counts the term frequency in order
to build the user pro le. The second one, called bow, is the same bag of words
approach but using all information available about users, not only the \short
bio". The third one, called pei, is our approach based on the semantic-based
enrichment process that uses our CKG. We have run all algorithms 32 times in
order to get the top ten suggested users for each one of the analyzed user. For
the evaluation we have chosen a sample of 6 testers: 3 internals, therefore people
who are working at the GEI and who know all users pro les and activities inside
the institute; and 3 external, people who do not know them, but with access to
the dataset with all text inserted.</p>
      <p>Table 1 shows rst results obtained by each approach, using the Precision
metric at several levels, with resulting ranked lists with 3, 5 and 10 elements.</p>
      <p>Table 2 displays the Discounted Cumulative Gain scores, using the Precision
metric at several levels, with the same granularity levels.</p>
      <p>Analyzing the results, we can notice that bow is less e ective than bow-s,
even if it exploit more text and information. The pei approach outperforms the
Semantic-based Expert Search in Textbook Research Archives
11
other approaches and con rms the assumption that semantic enrichment can
increase the performance of the retrieval system personalizing the results.</p>
      <p>By looking at the Precision scores on Table 1 it is possible to notice how
bow-s tends to keep the scores around a certain value, di erently than bow that
increases value as the resulting ranked list get longer. This issue highlights how
the use of additional information could help on systems who provide longer lists
of results but with no high precision on top ranked elements. Our proposed
approach pei overcomes this issue by providing good top ranked results using
more semantic information.</p>
      <p>Table 2, with DCG scores, shows how all algorithms provide more gain for
users, as the resulting ranked list get longer, and it is possibile to see how our
technique pei outperforms the others at all levels.
6</p>
    </sec>
    <sec id="sec-7">
      <title>Conclusions and future work</title>
      <p>In this paper, we provided a rst perspective on studies on expert nding,
similar people identi cation and social recommender systems, highlighting some
critical issues related to information extraction and user pro le de nition. We
presented our semantic-based enrichment approach for expert search that bases
on a Community Knowledge Graph, which de nes relationships and weights
among textbook researchers. The semantic relations help in nding the di erent
experts that could be of interest and could be recommended.</p>
      <p>For future work, we plan to expand our approach and investigate how the
network structure could be exploited to improve the results, and to run the
next evaluation with detailed user tests. Moreover, we want to explore the
possibility to extend our approach to resources that users interact with, such as
textbooks, manuscripts, or other cultural heritage objects, in order to improve
semantic search and information retrieval tasks in digital archives and libraries,
considering the relationships among them and with users.</p>
      <p>Marco Pavan and Ernesto William De Luca</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>F.</given-names>
            <surname>Abel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Gao</surname>
          </string-name>
          , G. j. Houben, and
          <string-name>
            <given-names>K.</given-names>
            <surname>Tao</surname>
          </string-name>
          .
          <article-title>Semantic enrichment of twitter posts for user pro le construction on the social web</article-title>
          .
          <source>ESWC</source>
          ,
          <year>June 2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>A.</given-names>
            <surname>Al-Kouz</surname>
          </string-name>
          , E. W. De Luca, and
          <string-name>
            <given-names>S.</given-names>
            <surname>Albayrak</surname>
          </string-name>
          .
          <article-title>Latent semantic social graph model for expert discovery in facebook</article-title>
          . IICS,
          <year>June 2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>F.</given-names>
            <surname>Cena</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Likavec</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F.</given-names>
            <surname>Osborne</surname>
          </string-name>
          .
          <article-title>Propagating user interests in ontology-based user model</article-title>
          .
          <source>In Arti cial Intelligence Around Man and Beyond</source>
          , volume
          <volume>6934</volume>
          of Lecture Notes in Computer Science, pages
          <volume>299</volume>
          {
          <fpage>311</fpage>
          .
          <year>September 2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>S.</given-names>
            <surname>Cetintas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rogati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Si</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Fang</surname>
          </string-name>
          .
          <article-title>Identifying similar people in professional social networks with discriminative probabilistic models</article-title>
          .
          <source>SIGIR</source>
          ,
          <year>July 2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>A.</surname>
          </string-name>
          <article-title>El-korany. Integrated expert recommendation model for online communities</article-title>
          .
          <source>IJWesT</source>
          , Vol.
          <volume>4</volume>
          ,
          <year>October 2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Y.</given-names>
            <surname>Fang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Si</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A. P.</given-names>
            <surname>Mathur</surname>
          </string-name>
          .
          <article-title>Discriminative probabilistic models for expert search in heterogeneous information sources</article-title>
          .
          <source>Journal of Information Retrieval</source>
          , Vol.
          <volume>14</volume>
          ,
          <year>April 2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>S. D.</given-names>
            <surname>Gollapalli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mitra</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C. L.</given-names>
            <surname>Giles</surname>
          </string-name>
          .
          <article-title>Ranking experts using author-documenttopic graphs</article-title>
          .
          <source>JCDL</source>
          ,
          <year>July 2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>M.</given-names>
            <surname>Gujral</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Chandra</surname>
          </string-name>
          .
          <article-title>Beyond recommenders and expert nders, processing the expert knowledge</article-title>
          .
          <source>International Journal of Computer Science Issues</source>
          , Vol.
          <volume>11</volume>
          ,
          <year>January 2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>W.</given-names>
            <surname>Koh</surname>
          </string-name>
          and
          <string-name>
            <given-names>L.</given-names>
            <surname>Mui</surname>
          </string-name>
          .
          <article-title>An information theoretic approach to ontology-based interest matching</article-title>
          .
          <source>IJCAI</source>
          ,
          <year>August 2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>S.</given-names>
            <surname>Mizzaro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pavan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and I.</given-names>
            <surname>Scagnetto</surname>
          </string-name>
          .
          <article-title>Content-based similarity of twitter users</article-title>
          . ECIR,
          <year>March 2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>S.</given-names>
            <surname>Mizzaro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pavan</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Scagnetto</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Valenti</surname>
          </string-name>
          .
          <article-title>Short text categorization exploiting contextual enrichment and external knowledge</article-title>
          .
          <source>SoMeRA</source>
          , SIGIR,
          <year>July 2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>C. Moreira</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Martins</surname>
            , and
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Calado</surname>
          </string-name>
          .
          <article-title>Using rank aggregation for expert search in academic digital libraries</article-title>
          .
          <source>CoRR</source>
          ,
          <year>January 2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13. H.
          <string-name>
            <surname>Noureddine</surname>
            , I. Jarkass,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Hazimeh</surname>
            ,
            <given-names>O. A.</given-names>
          </string-name>
          <string-name>
            <surname>Khaled</surname>
            , and
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Mugellini</surname>
          </string-name>
          . Carp:
          <article-title>Correlation-based approach for researcher pro ling</article-title>
          .
          <source>SEKE</source>
          ,
          <year>July 2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <given-names>F.</given-names>
            <surname>Orlandi</surname>
          </string-name>
          <article-title>. Multi-source provenance-aware user interest pro ling on the social semantic web</article-title>
          .
          <source>UMAP</source>
          ,
          <year>July 2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <given-names>F.</given-names>
            <surname>Orlandi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Breslin</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Passant</surname>
          </string-name>
          .
          <article-title>Aggregated, interoperable and multi-domain user pro les for the social web. I-SEMANTICS</article-title>
          ,
          <year>September 2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <given-names>T.</given-names>
            <surname>Plumbaum</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lommatzsch</surname>
          </string-name>
          , E. W. De Luca, and
          <string-name>
            <given-names>S.</given-names>
            <surname>Albayrak</surname>
          </string-name>
          . Serum:
          <article-title>Collecting semantic user behavior for improved news recommendations</article-title>
          .
          <source>UMAP</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <given-names>X.</given-names>
            <surname>Song</surname>
          </string-name>
          .
          <article-title>Enrichment of user pro les across multiple online social networks for volunteerism matching for social enterprise</article-title>
          .
          <source>SIGIR</source>
          ,
          <year>July 2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <given-names>A.</given-names>
            <surname>Spaeth and M. C.</surname>
          </string-name>
          <article-title>Desmarais. Combining collaborative ltering and text similarity for expert pro le recommendations in social websites</article-title>
          .
          <source>In UMAP</source>
          , volume
          <volume>7899</volume>
          of Lecture Notes in Computer Science, pages
          <volume>178</volume>
          {
          <fpage>189</fpage>
          .
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>M. Stankovic</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Rowe</surname>
            , and
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Laublet</surname>
          </string-name>
          .
          <article-title>Finding co-solvers on twitter with a little help from linked data</article-title>
          .
          <source>In The Semantic Web: Research and Applications</source>
          , volume
          <volume>7295</volume>
          of Lecture Notes in Computer Science, pages
          <volume>39</volume>
          {
          <fpage>55</fpage>
          .
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <given-names>K.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lin</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Chuang</surname>
          </string-name>
          .
          <article-title>Using google distance for query expansion in expert nding</article-title>
          . ICDIM, IEEE,
          <year>September 2014</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>