<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>BIR</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>New Datasets and a Benchmark of Document Network Embedding Methods for Scientific Expert Finding</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Robin Brochier</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Antoine Gourru</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Adrien Guille</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Julien Velcin</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Universite ́ de Lyon</institution>
          ,
          <addr-line>Lyon 2 ERIC EA3083</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <volume>14</volume>
      <fpage>16</fpage>
      <lpage>29</lpage>
      <abstract>
        <p>The scientific literature is growing faster than ever. Finding an expert in a particular scientific domain has never been as hard as today because of the increasing amount of publications and because of the ever growing diversity of expertise fields. To tackle this challenge, automatic expert finding algorithms rely on the vast scientific heterogeneous network to match textual queries with potential expert candidates. In this direction, document network embedding methods seem to be an ideal choice for building representations of the scientific literature. Citation and authorship links contain major complementary information to the textual content of the publications. In this paper, we propose a benchmark for expert finding in document networks by leveraging data extracted from a scientific citation network and three scientific question &amp; answer websites. We compare the performances of several algorithms on these different sources of data and further study the applicability of embedding methods on an expert finding task.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Many tools offer to search and filter the vast data sources available on the Web. In
particular, there is a multitude of platforms directed to the scientific community. From the
simple search engine for publications to the social network for researchers, all consume
and produce valuable data for searching scientific content of interest. Expert finding
is one the the most challenging problem that finds application in both academia and
the industry. To tackle this challenge, recent advances in document network
embedding (DNE) has the potential to inspire new unsupervised models that can deal with the
heterogeneous network of documents of the scientific literature. However, the design
of such efficient algorithms heavily depends on the development of strong evaluation
frameworks.</p>
      <p>In this paper, we propose a methodology and provide 4 datasets that extend the
limited scope of expertise retrieval evaluation frameworks. Furthermore, we provide
experiment results computed with unsupervised methods and we extend document network
embedding algorithms to this specific task.</p>
      <p>Our contributions are the following:
2
– we provide 4 datasets for expert finding extracted from a scientific publication
network and three question &amp; answer (Q&amp;A) websites and make them publicly
available 1;
– we describe an evaluation methodology based on the ranking of expert candidates
given a set of labeled document queries;
– we report experiment results that give some insights on this expert finding task;
– we explore and analyze the use of state-of-the-art document network embedding
algorithms for expert finding and we show that further research is needed to bridge
the gap between DNE methods and expert finding.</p>
      <p>The rest of the paper is organized as follows. In Section 2, we survey related works.
We detail in Section 3 our evaluation methodology, the datasets we extracted, the
evaluation measures and the algorithms we use. In Section 4 we show and analyze the results
of our experiments. Finally, in Section 5, we discuss our findings and provide future
directions.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Works</title>
      <p>In this section, we first present a formal definition for expert finding. Then we present
algorithms of the literature that address expert finding. Finally, we describe recent
methods for document network embedding that have the potential to deal with this particular
task.
2.1</p>
      <sec id="sec-2-1">
        <title>Formal definition of expert finding</title>
        <p>The concept of expert finding can cover a large range of tasks. The main principle
behind expertise retrieval is the search for candidates given a query. To match these two,
an algorithm will be provided with some data to link the output space, a ranking of
candidates, with the input space, which is often a textual content. However, many different
types of data can be considered to address this challenge. To fairly compare algorithms,
we choose a fixed structure for the data which reflects common use cases. Furthermore,
if supervised methods benefit from labeled fields of expertise associated with the
candidates, they are beyond the scope of this paper which focuses on unsupervised methods
only. Our goal is to compare methods that do not require sometimes costly annotations.</p>
        <p>
          Early works in expert search [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] usually consider a small set of topical queries. The
direct namings of these topics are used to retrieve a list of candidates by leveraging
a collection of documents they published (e.g., emails, scientific papers). This type of
evaluation is used across several public datasets [
          <xref ref-type="bibr" rid="ref15 ref16 ref26">15, 16, 26</xref>
          ].
        </p>
        <p>
          More recently, the concept of expert finding has been merged into the wider
concept of entity retrieval [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. As more and more complex data are produced on the Web,
expert finding becomes a particular application of entity search. At the same time, Q&amp;A
websites such as Stack Overflow 2 generate and make publicly accessible a big amount
of questions with expert answers, collaboratively curated by their users. Several works
1 https://github.com/brochier/expert_finding
2 https://stackoverflow.com/
address the search for experts in such websites [
          <xref ref-type="bibr" rid="ref20 ref28">20,28</xref>
          ]. Often, the task consists in either
finding the exact list of users who answered a specific question or ranking the answers
according to the user votes. In the first case, the task involves considering the evolution
of the users across time and, in the second case, the task involves understanding the
intrinsic quality of a written answer. Nevertheless, [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ] reviews several models for
expert finding in Q&amp;A websites. Their experiments show that matrix factorization-based
methods perform better than tree based and ranking based methods.
        </p>
        <p>
          In this paper, we adopt the document-query methodology recently proposed in [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
The expert search is performed given a set of queries that are particular textual instances
of some expert topics (or fields of expertise). Given a query, an algorithm should rank
first the candidates that are associated to the same fields of expertise. We provide 4
datasets for which we annotated experts and document queries. Each dataset consists
in candidates and documents linked by authorship relations (candidate-document e.g.
authorship) and by response relations (document-document e.g citation or answer). A
query is therefore one of the documents (e.g. a scientific paper or a question) for which
we aim to retrieve some experts of the topics depicted in it. This configuration reflects
many real case scenarios such as (1) the automatic search for scientific reviewers, (2) the
recommendation of expert users in Q&amp;A websites or even (3) the retrieval of interesting
profiles for job offers.
2.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Algorithms for expert finding</title>
        <p>
          Numerous works have addressed automatic expertise retrieval. We describe here the
main approaches and some interesting recent methods. P@noptic Expert [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] creates
meta-documents for a candidate by concatenating the contents of all documents she
produced. In this manner, ranking the candidates given a query becomes a similarity
search between the query representation and the meta-documents representations. A
voting model [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] computes the similarities between the query and the documents. The
algorithm then aggregates these scores at the candidate level by using a fusion
technique such as the reciprocal rank [
          <xref ref-type="bibr" rid="ref27">27</xref>
          ]. A propagation model [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] takes advantage of the
links between candidates and documents to propagate the similarities between the query
and the documents. Using random walks with restart [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ], the iterative propagation of
the scores converges in a few steps to a stationary distribution over the candidates.
WISER [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] models each candidate as a small, weighted, sub-graph of the Wikipedia
Knowledge Graph. Information derived from these graphs and traditional document
retrieval techniques are combined to identify experts w.r.t a query. Note that methods
leveraging external data are out of the scope of our benchmark. LT Expertfinder is an
evaluation framework for expert finding [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] based on an interactive tool. It integrates
various existing algorithms (such as [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]) in a user-friendly way. The underlying corpus
used by this tool is the ACL Anthology Network. However, it does not include a
wellestablished ground truth to assess who are the experts. Indeed, the evaluation is purely
done in an online manner since the user has to evaluate the degree of expertise based on
several features, such as author’s citations, h-index, keywords, etc. Recent works [
          <xref ref-type="bibr" rid="ref23 ref9">9,23</xref>
          ]
propose ad hoc embedding techniques, whereas, in this work, we’re interested in
measuring the performance of conventional network embedding techniques.
4
2.3
Network embedding [
          <xref ref-type="bibr" rid="ref12 ref19">12, 19</xref>
          ] provides an efficient approach to represent nodes in a low
dimensional vector space, suitable for solving various machine learning tasks. Recent
techniques extend NE for document networks. Text-Associated DeepWalk (TADW)
[
          <xref ref-type="bibr" rid="ref24">24</xref>
          ] extends DeepWalk to deal with textual attributes. Yang et al. prove, following the
work in [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ], that Skip-Gram with hierarchical softmax can be equivalently formulated
as a matrix factorization problem. TADW then consists in constraining the
factorization problem with a pre-computed representation of the documents by using Latent
Semantic Analysis (LSA) [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. Graph2Gauss (G2G) [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] is an approach that embeds
each node as a Gaussian distribution instead of a vector. The algorithm is trained by
passing node attributes through a non-linear transformation via a deep neural network
(encoder). GVNR-t [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] is a matrix factorization approach for document network
embedding, inspired by GloVe [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ], that simultaneously learns word, node and document
representations by optimizing a least-square objective over a co-occurrence matrix of
the nodes constructed by truncated random walks. IDNE [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] introduces a Topic-Word
Attention mechanism, trained from the connections of a document network, to represent
documents as mixtures of topics.
        </p>
        <p>DNE algorithms do not directly apply to expert finding data since they are not
designed to handle multiple types of nodes, in particular candidate nodes. In this paper,
we show (1) two methods to extend their applicability to the task of expert finding and
(2) the impact of their representations when they are used as document representations
for traditional expert finding algorithms.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Evaluation Methodology</title>
      <p>We present in this section the evaluation methodology that we follow to access the
performances of several algorithms for expert finding. We first describe the task we seek
to solve, then we describe the datasets that we extracted and explain how we annotated
them in order to access the quality of the algorithms’ outputs. Finally, we detail the
models used in our experiments.
3.1</p>
      <sec id="sec-3-1">
        <title>Ranking expert candidates from document queries</title>
        <p>
          Expert finding is a complex task that can be formalized in multiple ways. Early works
define this task as a ranking problem given several topic-queries where the naming of
these topics are directly used as queries to retrieve the expert candidates. However, in
many real world applications, a user is asked to provide a specific and detailed query.
In a Q&amp;A website for instance, a user usually exposes the problem she faces in full
detail and does not necessarily know the exact naming of the fields of expertise needed
to solve her problem. Furthermore, querying an algorithm with a small set of
topicqueries can lead to poor evaluation measures due to the usually small number of fields
of expertise associated with the dataset. For this reasons, we follow the document-query
evaluation methodology proposed in [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] by processing 4 datasets for which a set of
document-queries is manually annotated.
        </p>
        <p>The expert finding task in this paper is a ranking problem. Given a document labeled
with a ground truth set of fields of expertise, an algorithm is queried to rank a set
of candidates, among which a subset of experts are associated with the same set of
labels. The data provided to the algorithms consists in a corpus of nd documents D,
nc candidates C, a network of authorship with adjacency matrix Adc 2 Nnd nc and
a network of documents with adjacency matrix Add 2 Nnd nd . Figure 1 shows an
hypothetical dataset used in this paper. The ranking is performed in an unsupervised
setting, that is, no ground truth labels of expertise are given to the algorithms. The
set of labeled documents (the queries) can be smaller than nd and the set of labeled
candidates (experts) can be smaller than nc (i.e not all documents and candidates are
labeled).</p>
        <sec id="sec-3-1-1">
          <title>Expertise labels</title>
          <p>C1</p>
          <p>C2</p>
          <p>C3</p>
          <p>C4</p>
          <p>C5
D1</p>
          <p>D2</p>
          <p>D3</p>
          <p>D4</p>
          <p>D5</p>
          <p>D6
Adc
Add
the rankings D1 7! C3C4C5C1C2 and D6 7! C4C5C3C2C1.</p>
          <p>
            To evaluate the candidate scores provided by the algorithms, we compare the
resulting rankings with the ground truth fields of expertise. If a document is associated with
three different labels, we expect the algorithm to rank first all experts associated to at
least one of these labels. We report the area under the ROC curve (AUC), the precision
6
at 10 (P@10) and the average precision (AP) and we compute their standard
deviation along the queries. That is, we evaluate the robustness of the algorithms against the
variety of document-queries.
We consider 4 datasets. The first one is an extract of DBLP [
            <xref ref-type="bibr" rid="ref22">22</xref>
            ] in which a list 199
experts in 7 fields are annotated [
            <xref ref-type="bibr" rid="ref26">26</xref>
            ] by human judgments 3. Our dataset only considers
the annotated experts and the other candidates that are close in the co-authorship
network which explains the relatively small size of our network compared to the original
one. In addition to the expert annotations, our evaluation framework requires document
annotations since we adopt the document-query methodology for expertise retrieval. We
asked two PhD students in computer science to associate independently 20 randomly
drawn documents per field of expertise (140 in total). Then, only the labels on which
the two annotators agreed were kept, leaving 114 annotated papers. The mean Cohen’s
kappa coefficient across the labels is 0:718. An advantage of our methodology is that
we can evaluate the algorithms on more queries (114 documents) than the traditional
method (7 labels). This allows us to assess the robustness of the algorithms by
computing the standard deviations of the ranking metrics along all queries. However, one
might suggest that these 7 labels do not reflect a representative set of expertise as there
are too broad. For this reason, we seek for a wider granularity of expertise by the use of
well-know question &amp; answer website.
          </p>
          <p>If scientific publication networks are easy to find on the Web, scientific expertise
annotations are rarely available for both authors and publications. We use data
downloaded in June 2019 from Stack Exchange 4 to create datasets for expert finding
collected from three communities closely related to research. Academia 5 is dedicated to
academics and higher education. Mathoverflow 6 gathers professional mathematicians
and is widely used by researchers. Stats 7 (also known as Cross Validated) addresses
statistics, machine learning and data mining issues. For each dataset, we first keep
questions with at least 10 user votes that have at least one answer with 10 user votes or more.
We build the networks by linking questions with their answers and by linking answers
with the users who published them. The field of expertise are the tags associated with
the questions. Only the tags that occur at least 50 times are kept. We annotate an expert
with the tags of a question if her answer to that question received at least 10 votes. Note
that the tags are first provided by the users who ask the questions but they are thereafter
verified by experimented users.</p>
          <p>The general properties of our 4 datasets are presented in Table 1. The annotations
and the preprocessed datasets are made publicly available.
3 https://lfs.aminer.cn/lab-datasets/expertfinding/#expert-list
4 https://archive.org/details/stackexchange
5 https://academia.stackexchange.com/
6 https://mathoverflow.net/
7 https://stats.stackexchange.com/</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>Datasets and Benchmark of DNE Methods for Expert Finding 7</title>
        </sec>
        <sec id="sec-3-1-3">
          <title>DBLP Stats Academia Mathoverflow</title>
          <p>We run the experiments with 4 baseline algorithms and 4 document network embedding
algorithms. The laters are adapted with two aggregation schemes in order to deal with
the candidates since they are primarily designed for document network only. These
aggregations are arbitrary and are voluntarily the most straightforward way to run DNE
algorithms on bipartite networks of authors-documents. We further discuss these choices
in section 4.</p>
          <p>
            Baselines We run the experiments with the same models as in [
            <xref ref-type="bibr" rid="ref3">3</xref>
            ], using the tf-idf
representations and the cosine similarity measure. Also, we add a random model to
have reference metrics:
– Random model: we randomly draw scores between 0 and 1 for each candidate;
– P@noptic model [
            <xref ref-type="bibr" rid="ref7">7</xref>
            ]: we concatenate the textual content of each document
associated to the candidates, use their tf-idf representations and compute the cosine
similarity to produce the scores;
– Voting model [
            <xref ref-type="bibr" rid="ref14">14</xref>
            ]: we use the reciprocal rank to aggregate the scores at the
candidate level;
– Propagation model [
            <xref ref-type="bibr" rid="ref21">21</xref>
            ]: we concatenate the two adjacency matrices Adc and
Add to construct a transition matrix between candidates and documents such that
A = AAdd|dc A0dc . The initial scores are the cosine similarities between the tf-idf
representations of the query and the documents. The scores are propagated
iteratively until convergence with a restart probability of 0:5.
          </p>
          <p>We also run the voting and propagation models using document representations
produced by IDNE in place of the tf-idf vectors. The document network provided to
IDNE has adjacency matrix Ad = AdcAd|c + Add.</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>Extending DNE algorithms for expert finding DNE methods usually operate in net</title>
        <p>works of documents, with no candidate nodes. To apply them in the context of expert
finding, we propose two straightforward approaches:
– pre-aggregation: as in the P@noptic model, meta-documents are generated by
aggregating the documents produced by each candidates. Furthermore, an
adjacency matrix of a meta-network between candidates and documents is constructed.
We compute a candidate network as Ac = Ad|cAdc and a document network as
8
For all datasets, the propagation model performs generally better than the other
algorithms, particularly in terms of precision. Both aggregation schemes yield to poor
results and none of these two methods appear to be better than the other. GVNR-t is
the best algorithm among the document network embedding models. We believe that,
if DNE algorithms are well suited for document network representation learning, the
gap between simple tasks such as node classification and link prediction and the task of
expert finding is too big for our naive aggregation schemes to perform well. Especially,
the network structure changes significantly between an homogeneous network and an
Datasets and Benchmark of DNE Methods for Expert Finding
9
(a) Propagation model with tf-idf
representations: the curve has a nice shape which means
the ranking of candidates are good even for the
last ranked experts.
(b) Propagation model with IDNE
representations: the first ranked candidates are good but
the algorithm tends to wrongly rank last many
true experts.
heterogeneous network. Moreover, expert finding algorithms often benefit from
information about the centrality of the candidates and documents. DNE algorithms do not
particularly preserve this information neither do our aggregation schemes.
4.2</p>
      </sec>
      <sec id="sec-3-3">
        <title>Using DNE as document representations for the baselines</title>
        <p>Since the baseline algorithms perform well, we study the possibility to apply them using
a DNE algorithm for the representations of the documents. We only report the results
with the representations computed with IDNE but we observe the same behaviors with
other DNE models. First, these representations constantly improve the voting model,
which achieves best results in terms of AUC on Stats and Mathoverflow. Then, the most
surprising effect is the significant decrease of performance of the propagation model. If
the precision for the first ranked candidates is not affected, the AUC score significantly
drops for the three Q&amp;A datasets. We believe that document network embeddings
captures too long-range dependencies between the documents in the network, which are
then subsequently exaggerated by the propagation procedure. Figure 2 shows the effect
of the representations used with the propagation model on the ROC curve.
4.3</p>
      </sec>
      <sec id="sec-3-4">
        <title>Differences between the datasets</title>
        <p>The results achieved by the algorithms on all three Stack Exchange datasets are
consistent. However, they do not behave the same with DBLP. First, DNE methods get closer
scores to the baselines on DBLP. In the Q&amp;A datasets, the interactions are more isolated
i.e. there are more users having fewer interactions. This difference of network
properties might disadvantage DNE methods who are usually trained on scale-free networks
whose degree distribution follows a power law. Moreover, the propagation method does
not suffer with DBLP from the decrease of performance induced by the IDNE
representations. We hypothesize that the low number of expertise fields associated with this
dataset largely reduces the effect described in the previous section.
5</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Discussion and Future Work</title>
      <p>In this paper, we provide experiment materials for expert finding with the help of four
annotated datasets and further report results based on several baseline algorithms.
Moreover, we study the ability of document network embedding methods to tackle the expert
finding challenge. We show that DNE algorithms can not be trivially adapted to achieve
state-of-the-art scores. However, we reveal that document network embeddings can
improve the voting model but diminish the propagation model.</p>
      <p>In future work, we would like to find an efficient way to bridge the gap between
DNE algorithms and expert finding. To do so, taking the heterogeneity into account
should help better capturing the real similarity between a document and a candidate.
Furthermore, a deeper analysis of the interplay between the candidates and the text
content of the documents appears to be a necessary way to better understand the task of
expert finding.</p>
      <sec id="sec-4-1">
        <title>Datasets and Benchmark of DNE Methods for Expert Finding 11</title>
        <p>R. Brochier et al.</p>
      </sec>
      <sec id="sec-4-2">
        <title>Datasets and Benchmark of DNE Methods for Expert Finding 13</title>
        <p>R. Brochier et al.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Balog</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Serdyukov</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>De Vries</surname>
            ,
            <given-names>A.P.</given-names>
          </string-name>
          :
          <article-title>Overview of the trec 2010 entity track</article-title>
          .
          <source>Tech. rep.</source>
          ,
          <source>NORWEGIAN UNIV OF SCIENCE AND TECHNOLOGY TRONDHEIM</source>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bojchevski</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , Gu¨nnemann, S.:
          <article-title>Deep gaussian embedding of graphs: Unsupervised inductive learning via ranking</article-title>
          .
          <source>In: International Conference on Learning Representations</source>
          . pp.
          <fpage>1</fpage>
          -
          <lpage>13</lpage>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Brochier</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guille</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rothan</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Velcin</surname>
          </string-name>
          , J.:
          <article-title>Impact of the query set on the evaluation of expert finding systems</article-title>
          .
          <source>Proceedings of the 3rd Joint Workshop on Bibliometric-enhanced Information Retrieval and Natural Language Processing for Digital Libraries (BIRNDL</source>
          <year>2018</year>
          )
          <article-title>co-located with the 41st</article-title>
          <source>International ACM SIGIR Conference</source>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Brochier</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guille</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Velcin</surname>
          </string-name>
          , J.:
          <article-title>Global vectors for node representations</article-title>
          .
          <source>In: The World Wide Web Conference</source>
          . pp.
          <fpage>2587</fpage>
          -
          <lpage>2593</lpage>
          . ACM (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Brochier</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guille</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Velcin</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>Inductive document network embedding with topic-word attention</article-title>
          .
          <source>In: Proceedings of the 42nd European Conference on Information Retrieval Research</source>
          . Springer (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Cifariello</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferragina</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ponza</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Wiser: A semantic approach for expert finding in academia based on entity linking</article-title>
          .
          <source>Inf. Syst</source>
          .
          <volume>82</volume>
          ,
          <fpage>1</fpage>
          -
          <lpage>16</lpage>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Craswell</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hawking</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vercoustre</surname>
            ,
            <given-names>A.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wilkins</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          : P@
          <article-title>noptic expert: Searching for experts not just for documents</article-title>
          .
          <source>In: Ausweb Poster Proceedings, Queensland, Australia</source>
          . vol.
          <volume>15</volume>
          , p.
          <volume>17</volume>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Craswell</surname>
          </string-name>
          , N.,
          <string-name>
            <surname>de Vries</surname>
            ,
            <given-names>A.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soboroff</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Overview of the trec 2005 enterprise track</article-title>
          .
          <source>In: Trec</source>
          . vol.
          <volume>5</volume>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>7</lpage>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>Dargahi</given-names>
            <surname>Nobari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Sotudeh Gharebagh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Neshati</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          :
          <article-title>Skill translation models in expert finding</article-title>
          .
          <source>In: Proceedings of the 40th international ACM SIGIR conference on research and development in information retrieval</source>
          . pp.
          <fpage>1057</fpage>
          -
          <lpage>1060</lpage>
          . ACM (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Deerwester</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dumais</surname>
            ,
            <given-names>S.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Furnas</surname>
            ,
            <given-names>G.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Landauer</surname>
            ,
            <given-names>T.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harshman</surname>
          </string-name>
          , R.:
          <article-title>Indexing by latent semantic analysis</article-title>
          .
          <source>Journal of the American Society for Information Science</source>
          <volume>41</volume>
          (
          <issue>6</issue>
          ),
          <fpage>391</fpage>
          -
          <lpage>407</lpage>
          (
          <year>1990</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Fischer</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Remus</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Biemann</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Lt expertfinder: An evaluation framework for expert finding methods</article-title>
          .
          <source>In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (Demonstrations)</source>
          . pp.
          <fpage>98</fpage>
          -
          <lpage>104</lpage>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Grover</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leskovec</surname>
          </string-name>
          , J.: node2vec:
          <article-title>Scalable feature learning for networks</article-title>
          .
          <source>In: Proceedings of the 22nd ACM SIGKDD international conference on Knowledge discovery and data mining</source>
          . pp.
          <fpage>855</fpage>
          -
          <lpage>864</lpage>
          . ACM (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Levy</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goldberg</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Neural word embedding as implicit matrix factorization</article-title>
          .
          <source>In: Advances in neural information processing systems</source>
          . pp.
          <fpage>2177</fpage>
          -
          <lpage>2185</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Macdonald</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ounis</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Voting for candidates: adapting data fusion techniques for an expert search task</article-title>
          .
          <source>In: Proceedings of the 15th ACM international conference on Information and knowledge management</source>
          . pp.
          <fpage>387</fpage>
          -
          <lpage>396</lpage>
          . ACM (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Macdonald</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ounis</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soboroff</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Overview of the trec 2007 blog track</article-title>
          .
          <source>In: TREC</source>
          . vol.
          <volume>7</volume>
          , pp.
          <fpage>31</fpage>
          -
          <lpage>43</lpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Mislevy</surname>
            ,
            <given-names>R.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Riconscente</surname>
            ,
            <given-names>M.M.</given-names>
          </string-name>
          :
          <article-title>Evidence-centered assessment design</article-title>
          .
          <source>In: Handbook of test development</source>
          , pp.
          <fpage>75</fpage>
          -
          <lpage>104</lpage>
          .
          <string-name>
            <surname>Routledge</surname>
          </string-name>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Page</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Motwani</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Winograd</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>The pagerank citation ranking: Bringing order to the web</article-title>
          .
          <source>Tech. rep.</source>
          ,
          <string-name>
            <surname>Stanford InfoLab</surname>
          </string-name>
          (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Pennington</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Socher</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manning</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          : Glove:
          <article-title>Global vectors for word representation</article-title>
          .
          <source>In: Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP)</source>
          . pp.
          <fpage>1532</fpage>
          -
          <lpage>1543</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Perozzi</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Al-Rfou</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Skiena</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          : Deepwalk:
          <article-title>Online learning of social representations</article-title>
          .
          <source>In: Proceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and data mining</source>
          . pp.
          <fpage>701</fpage>
          -
          <lpage>710</lpage>
          . ACM (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Riahi</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zolaktaf</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shafiei</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Milios</surname>
          </string-name>
          , E.:
          <article-title>Finding expert users in community question answering</article-title>
          .
          <source>In: Proceedings of the 21st International Conference on World Wide Web</source>
          . pp.
          <fpage>791</fpage>
          -
          <lpage>798</lpage>
          . ACM (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Serdyukov</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rode</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hiemstra</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Modeling multi-step relevance propagation for expert finding</article-title>
          .
          <source>In: Proceedings of the 17th ACM conference on Information and knowledge management</source>
          . pp.
          <fpage>1133</fpage>
          -
          <lpage>1142</lpage>
          . ACM (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Tang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , Zhang, J.,
          <string-name>
            <surname>Yao</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , Zhang,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Su</surname>
          </string-name>
          ,
          <string-name>
            <surname>Z.</surname>
          </string-name>
          :
          <article-title>Arnetminer: extraction and mining of academic social networks</article-title>
          .
          <source>In: Proceedings of the 14th ACM SIGKDD international conference on Knowledge discovery and data mining</source>
          . pp.
          <fpage>990</fpage>
          -
          <lpage>998</lpage>
          . ACM (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Van Gysel</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>de Rijke</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Worring</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Unsupervised, efficient and semantic expertise retrieval</article-title>
          .
          <source>In: WWW</source>
          . vol.
          <year>2016</year>
          , pp.
          <fpage>1069</fpage>
          -
          <lpage>1079</lpage>
          . The International World Wide Web Conferences Steering Committee (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          :
          <article-title>Network representation learning with rich text information</article-title>
          .
          <source>In: Twenty-Fourth International Joint Conference on Artificial Intelligence</source>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Yuan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , Zhang,
          <string-name>
            <given-names>Y.</given-names>
            ,
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Hall</surname>
          </string-name>
          ,
          <string-name>
            <surname>W.</surname>
          </string-name>
          , Cabota`,
          <string-name>
            <surname>J.B.</surname>
          </string-name>
          :
          <article-title>Expert finding in community question answering: a review</article-title>
          .
          <source>Artificial Intelligence Review</source>
          <volume>53</volume>
          (
          <issue>2</issue>
          ),
          <fpage>843</fpage>
          -
          <lpage>874</lpage>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Zhang</surname>
          </string-name>
          , J.,
          <string-name>
            <surname>Tang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Expert finding in a social network</article-title>
          .
          <source>In: International Conference on Database Systems for Advanced Applications</source>
          . pp.
          <fpage>1066</fpage>
          -
          <lpage>1069</lpage>
          . Springer (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Song</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ma</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jiang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jin</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ma</surname>
          </string-name>
          , S.:
          <article-title>Expansionbased technologies in finding relevant and new information: Thu trec 2002: Novelty track experiments</article-title>
          .
          <source>NIST SPECIAL PUBLICATION SP 251</source>
          ,
          <fpage>586</fpage>
          -
          <lpage>590</lpage>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>He</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ng</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          :
          <article-title>Expert finding for question answering via graph regularized matrix completion</article-title>
          .
          <source>IEEE Transactions on Knowledge and Data Engineering</source>
          <volume>27</volume>
          (
          <issue>4</issue>
          ),
          <fpage>993</fpage>
          -
          <lpage>1004</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>