<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Multilingual Expert Search using Linked Open Data as Interlingual Representation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Daniel M. Herzig</string-name>
          <email>herzig@kit.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hristina Taneva</string-name>
          <email>hristina.taneva@student.kit.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute AIFB Karlsruhe Institute of Technology 76128 Karlsruhe</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Most Information Retrieval models take documents as Bagof-Words and are thereby bound to the language of the documents. In this paper, we present an approach using Linked Open Data resources, i.e. URIs, as interlingual document representations. Documents and queries are summarized by the resources they contain. We show the applicability of our approach for multilingual retrieval with a case study on expert search.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>When encountering a problem, there are often two ways to get to a solution.
Either acquire the knowledge, in order to solve the problem by oneself, or ask
for outside help, preferably somebody who has expertise and experience in the
needed domain. The rst case is often not feasible or would require too much
time. In the second case, the subsequent problem of nding the right expert
arises. We address this problem and present an approach for expert search in
this paper.</p>
      <p>Identifying who an expert is for a certain domain can be done in many
ways. One possible solution is to use documents and assume that the authors
have expertise on the topics they wrote about. We apply this assumption and
consider documents for the identi cation of experts.</p>
      <p>Obviously, the more speci c the problem is the harder is it to nd an expert.
Thus, extending the considered search space even across languages improves the
situation. The scenario of considering documents in di erent languages is not an
arti cial one, e.g. global companies have product documentations in many
languages or online developer forums have discussion threads in di erent languages.
Our approach addresses the problem of how to deal with di erent languages by
applying an interlingual representation for documents based on Linked Open
Data resources.</p>
      <p>
        Expert Search Track at CriES Our approach participated at the expert search
track of the Cross-lingual Expert Search Workshop (CriES) at CLEF 2010 [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ].
The setting and the evaluations presented in this paper are provided by the
workshop. The task of the expert search track was to nd experts for 60 topics
consisting of 15 topics in each of the four languages English, Spanish, French,
and German in the Yahoo! Answers data corpus [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]1. The data corpus consists
of 780193 threads, i.e. questions and answers, in the categories "Health",
"Computer &amp; Internet", and "Science &amp; Math.", in four languages written by 169819
users, i.e. experts. Table 1 gives an overview of the data set. More details and
an overview of the results of the workshop can be found in [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ].
      </p>
      <sec id="sec-1-1">
        <title>Threads</title>
      </sec>
      <sec id="sec-1-2">
        <title>Users</title>
        <p>This paper is organized as follows. After the introduction in this section, we
describe the usage of Linked Open Data as an interlingual representation in
Section 2. In Section 3, we present our model for expert search, how we create
pro les between resources and experts and how we estimate parameters.
Section 4 presents the evaluation and Section 5 discusses related work. Finally, we
conclude in Section 6.
2</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Multilingual IR based using Linked Open Data</title>
      <p>
        Most common models in Information Retrieval see documents as Bag-of-Words,
i.e. they resolve the order of the words and take the collection of words as the
representation of a document. As a consequence, this representation is directly
bound to the language of the document. When using keyword queries in one
language, relevant documents in another languages are probably not retrieved.
We propose an approach using Linked Open Data resources as document
representation. Linked Open Data (LOD) refers to interlinked, publicly available,
and structured datasets on the web using semantic web standards, in particular
the Resource Description Framework (RDF) [
        <xref ref-type="bibr" rid="ref4 ref7">4, 7</xref>
        ].
      </p>
      <p>The rst principle of LOD states that things should be identi ed by Uniform
Resource Identi ers (URIs), where things, i.e. resources, can be virtually
everything. The notion is not limited to physical things, but comprises also abstract or
intangible concepts, like happiness or re alarm. URIs are not necessarily human
readable, since they are meant to be processed by machines. Therefore, human
readable labels are often assigned to URIs. Since there can be multiple labels
in di erent languages for one URI, the URI itself can be seen as an interlingual
representation for the resource it identi es. Figure 1 illustrates an example about
the resource representing Germany and its labels in several languages.</p>
      <p>
        The resource in Figure 1 is taken from DBpedia2. DBpedia is a popular
LOD dataset extracted from Wikipedia, which exploits the interlanguage links
1 This dataset is provided by the Yahoo! Research Webscope program (see http://
research.yahoo.com/) under the following ID:L6. Yahoo! Answers Comprehensive
Questions and Answers (version 1.0)
2 http://dbpedia.org, Aug 4 2010
3 http://www.fao.org/agrovoc/, Aug 4 2010
4 http://www.illc.uva.nl/EuroWordNet/, Aug 4 2010
exploit additional information from other Linked Open Data sources, e.g. the
resource representing Germany from Figure 1 is linked through the typed link
owl : sameAs [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] to resources representing the same thing, e.g. to the resource
from The New York Times5 or from Geonames6. Beside information about the
same thing, also the typed links between di erent resources can be exploited. It
allows to enrich the representation with additional information and adapt it to
speci c use cases or domains.
      </p>
      <p>document  1:  
Bulgaria's   best   World   Cup   performance   was   in   the  
1994  World  Cup  in  the  United  States,  where  they  beat  
defending   champions   Germany   to   reach   the   semi-­‐
finals.  
document  2:  
Deutschland   ist   ein   föderalis'scher   Staat   in  
MiOeleuropa.   Deutschland   ist   Gründungsmitglied   der  
Europäischen   Union   und   mit   knapp   82   Millionen  
Einwohnern  deren  bevölkerungsreichstes  Land.  
bag  of  resouces  1:  
db:  Bulgaria  
db:  FIFA_World_Cup  
db:  United_States  
db:  Germany  
bag  of  resouces  2:  
db:  Germany  
db:  Sovereign_state  
db:  Central_Europe  
db:  Ci'zen  
The Yahoo! Answers data corpus contains discussion threads consisting of an
initial question and subsequent answers. The problem of expert search in this
context is to nd users, who are likely able to answer a given question q, i.e. a
topic, based on the threads in the data corpus.</p>
      <p>We apply mixture language models. Potential experts are ranked according
to the probability that the expert ex 2 E can answer the given question q 2 Q,
i.e. P (exjq). A question q is modeled as a Bag-of-Resources: q = fr1; :::; rng.</p>
      <p>P (exjq) / P (ex) P (qjex) = P (ex)
(1)
n
Y P (rijex)
i=1</p>
      <p>We apply Bayes's theorem and assume P (q) and the prior P (ex) to be equal
to 1. The probability P (rijex) is approximated as a weighted sum of several
features f and smoothed by information over the entire corpus C.</p>
      <p>P (rijex) = X( f Pf (rijex)) + C PC (ri)
f
s:t: X
f
f + C = 1
5 http://data.nytimes.com/55761554936313344161, Aug 4 2010
6 http://sws.geonames.org/2921044/about.rdf, Aug 4 2010</p>
      <sec id="sec-2-1">
        <title>Expert - Resource Pro les</title>
        <p>One answer per thread is marked by the questioner or by votes of other users
as the best answer to the question. The user, who gave the best answer, is
identi able by its ID. All other answers do not have a user ID. We exploit this
setting by building two di erent models. These models are illustrated in Figure 4
and explained below.</p>
        <p>aBest 
expert  ID 
q  
a  
 a* 
a  
a  
aAl  </p>
        <p>Best-Answer Model
This model takes the question q and the best answer a together as abest and
relates abest to the expert who gave the best answer, as illustrated in Figuren 4.
The idea behind this model is that the user obviously understood the question,
because he was able to give the best answer. Therefore he holds expertise about
the covered resources. Formally, the model is de ned as follows, where f req(r; a)
is the frequency of resource r in a.</p>
        <p>Pbest(rjex) = X P (rjabest)
abest</p>
        <p>P (exjabest) P (abest)</p>
        <p>P (ex)
with</p>
        <p>P (rjabest) = Pr2abest f req(r; abest)</p>
        <p>f req(r; abest)
P (abest) =</p>
        <p>; P (ex) =
P (exjabest) = 1; i ex author of abest, 0 otherwise
1 1
jQj
jEj
All-other-Answers Model
This model relates all answers aall, except the best answer, to the expert, who
gave the best answer. The assumption behind this model is that an expert,
who gave the best answer, might also say that other answers are not correct.
Therefore, we assume that the expert has expertise about the resources covered
by these answers as well, at least to some extent. Formally, the model is de ned
analogously to the previous one:</p>
        <p>Pall(rjex) =</p>
        <p>X P (rjaall)
aall</p>
        <p>P (exjaall) P (aall)</p>
        <p>P (ex)
3.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Parameter Estimation</title>
        <p>The mixture model presented in the previous section allows to balance the
inuence of each model through the corresponding parameter f . In our case, we
need to determine best, the weight for the Best-Answer-Model and all, the
weight for the All-other-Answers-Model. The smoothing parameter C remains
xed at C = 0:1. Hence, best + all =! 0:9 must hold.</p>
        <p>In order to examine the e ect of di erent parameter con gurations on the
performance of the retrieval, we used the 60 topics provided by workshop along with
the given a priori relevance information. The a priori relevance information is
directly taken from the data set, i.e. each topic has exactly one relevant expert,
namely the one, who wrote the best answer for this topic. This setting is not
optimal, since these questions are part of the data corpus itself and not distinct
from it. Furthermore, it can be assumed that there are more than one relevant
expert per question and that judging the performance by the occurrence of just
one expert in the result set will deviate from the actual result. However, it
allows at least to roughly estimate a parameter con guration. We used the Mean
Average Precision (MAP) to measure the performance.</p>
        <p>Since the 60 topics are part of the Best-Answer-Model a correlation between
best and the M AP can be assumed in this setting. The M AP was measured
in steps of 0:05 from best = 0, i.e. the performance without the
Best-AnswerModel, to best = 0:9, the performance of the Best-Answer-Model alone. The
observed MAP for the di erent parameters is shown in the left plot of Figure 5.
As assumed, a correlation between best and the M AP can be observed.
Remarkably, the MAP decreases for best = 0:9, despite the assumed correlation.
This suggests that information is lost, if the All-Answers-Model is not involved
and as a consequence, the optimal parameter con guration can not be the
maximal observed MAP. We used least-square curve tting to approximate the
observed values, i.e. the red line in Figure 5. The maximum of the tted curve is
best = 0:66. Comparing the estimated values with the actual MAP computed
ex post with the entire assessments shows that the maximum is even lower at
about best = 0:53, see right plot of Figure 5.
4</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Evaluation</title>
      <p>We submitted three runs with di erent con gurations, see Table 2 for an overview
of the results. Some of the 60 topics are very short and many are written
in rather colloquial language and grammar or use abbreviations, e.g. "Why
0
.
00.0 0.2 0.4
λbest
0.6 0.8
do women get PMS?" or "hab es runtergeladen wie kann ich bei msn
chatten?", which caused problems for the Wikipedia Miner to identify resources.
For 11 topics the Wikipedia Miner did not identify any resources. In these
cases, we extracted the resources manually, e.g. the resources db:Woman and
db:Premenstrual syndrome for the rst question mention before and the
resources db:MSN and db:Online chat for the latter. We did this for run1 and
run3 and left the topics untouched for run2, in order to see how the approach
performs without any manual intervention.</p>
      <p>Strict Lenient Parameters
Run Id P@10 MRR P@10 MRR best all C
run3 0.49 (+157%) 0.76 (+90%) 0.87 (+123%) 0.93 (+48%) 0.7 0.2 0.1
run1 0.48 (+153%) 0.77 (+93%) 0.86 (+121%) 0.94 (+49%) 0.6 0.3 0.1
run2 0.35 (+84%) 0.65 (+63%) 0.61 (+56%) 0.74 (+17%) 0.6 0.3 0.1
BM25 + Z-Score 0.19 0.40 0.39 0.63</p>
      <p>
        Table 2 shows the results for the top 10 retrieved experts. Precision at
cuto level 10 (P@10) and Mean Reciprocal Rank (MRR) are used as evaluation
measures. Precision/Recall curves for each run are presented in Figure 6 using
strict and lenient assessments [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. All three runs exceed the standard IR
baseline, BM25 + Z-Score [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. The baseline uses machine translation to translate
the topics in the four languages and matches them against monolingual indexes.
The results retrieved from the four monolingual indexes are combined for each
expert using the Z-Score [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
      </p>
      <p>Beside retrieving relevant experts for a topic, one main aim of our approach
was to cross the language barrier and nd experts regardless of their language.
Figure 7 visualizes the language distribution of the retrieved experts for each
topic language for run1. In order to facilitate the comparison, the distribution
of threads and experts in the data set is displayed on the right. One can see
that indeed experts in all four languages were retrieved in most cases. Further,
the domination of english speaking experts is due to the proportions in the data
set and in addition due to the larger, underlying resource space, as illustrated
in Figure 2. However, a bias towards the language of the topic is also
observable, because not all resource have labels in all other languages as discussed in
Section 2.</p>
      <p>10 ●
8 ●
6
4 ●
2 ●
0 EN ES FR DE</p>
      <p>English Topics
0
1
8
6
4
2 ●
0 EN ES FR DE</p>
      <p>Spanish Topics
01 ●
8
6
4 ●
2 ●
0 E●N ES FR DE</p>
      <p>French Topics
0
1
8
64 ● ●●
2 ● ●
0 EN ES FR DE</p>
      <p>
        German Topics
Our approach uses Language Models, which have been studied by [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and applied
by [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] for expert search. The latter compares two models with di erent search
strategies. The rst model collects all documents for every candidate and then
identi es the relevant topics in these documents. The second model nds rst
the signi cant documents for a given topic and then discovers the associated
experts. Our approach is based on model similar to the second model. We nd
rst the documents comprising the resources of the query and then relate the
resources to the expert who gave the best answer.
      </p>
      <p>
        Using concepts instead of terms was studied by [
        <xref ref-type="bibr" rid="ref14 ref16 ref17">14, 16, 17</xref>
        ]. These approaches
use Explicit Semantic Analysis and match topics to documents in a concept space
consisting of Wikipedia articles. Our approach uses also Wikipedia as
background knowledge and a representation similar to the concept representation, as
long as only the resources are considered. However, as discussed in Section 2,
using URIs instead of concepts allows to draw information from other sources
and facilitates the usage of connections between the resources. In order to
extract the resources from the documents, we screened several tools that deal with
Named Entity Recognition and Extraction. The Enrycher web service [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] tries
to extract not just resources, but also triples, which connect these resources.
The OpenCalais Web Service analyzes text and returns semantic metadata[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
However, these approaches do not work with all four languages. The advantage
of the Wikipedia Miner [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], beside fast and precise results, is that the resource
space is clearly de ned, i.e. all articles of Wikipedia, and that it supports the
language of the loaded Wikipedia le. Hence, we choose [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] for our approach.
      </p>
      <p>
        Not just indexing the terms of a document, but the idea of indexing what a
document is about, i.e. topic indexing, was introduced by [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Another approach
to topic indexing by embedding background knowledge derived from Wikipedia
was introduced by [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. All relevant topics being mentioned in a document are
linked to Wikipedia articles. The titles of the articles are used as index terms.
A similar approach to ours, however our approach is not limited to Wikipedia
and the usage of URIs instead of terms allows to exploit the links between the
URIs in a later stage.
      </p>
      <p>
        A di erent approach to multilingual IR was introduced by [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], who uses a
multilingual ontology to map a term with the appropriate concept. However, it
does not consider disambiguation of terms. An aspect covered by our approach,
since URIs are not ambiguous and the URIs are determined using the
disambiguation of [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] examines the impact of the use of semantic annotations
on the performance in monolingual and cross-language retrieval.
6
      </p>
    </sec>
    <sec id="sec-4">
      <title>Conclusions and Future Work</title>
      <p>
        We presented an approach for the Expert Search challenge of the CriES
Workshop at CLEF 2010 using Linked Open Data resources, i.e. URIs, as interlingual
document representations. We used Wikipedia as the corpus of resources, but the
approach is not limit to the usage of Wikipedia. Resources are extracted from the
documents using the Wikipedia Miner Toolkit [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] and used to create
ExpertResource pro les. A mixture model is applied for the retrieval and ranking of
experts for a given topic. Also topics are represented as a Bag-of-Resources.
      </p>
      <p>Our approach yielded solid results by exceeding the standard BM25 +
ZScore baseline from 17% to 157% regarding Mean Reciprocal Rank and Precision
at 10. Another advantage of our approach is that not the entire documents need
to be indexed, but just a summary consisting of several URIs, which decreases
the index size.</p>
      <p>In future, we plan to use more features of Linked Open Data for IR. In
particular, exploiting the links between resources to leverage the interconnection.
7</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <p>We thank Philipp Sorg for the helpful discussions and his valuable feedback.
We also thank David Milne for developing the Wikipedia Miner and making it
available as open source. Research reported in this paper was supported by the
German Federal Ministry of Education and Research (BMBF) under the iGreen
project (grant 01IA08005K).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1. Opencalais. http://www.opencalais.
          <source>com, Aug</source>
          <volume>4</volume>
          2010.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <article-title>Owl web ontology language overview</article-title>
          .
          <source>W3c recommendation, World Wide Web Consortium</source>
          ,
          <year>February 2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>K.</given-names>
            <surname>Balog</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Azzopardi</surname>
          </string-name>
          , and
          <string-name>
            <surname>M. de Rijke</surname>
          </string-name>
          .
          <article-title>A language modeling framework for expert nding</article-title>
          .
          <source>Information Processing &amp; Management</source>
          ,
          <volume>45</volume>
          (
          <issue>1</issue>
          ):
          <fpage>119</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>C.</given-names>
            <surname>Bizer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Heath</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Berners-Lee</surname>
          </string-name>
          .
          <article-title>Linked data - the story so far</article-title>
          .
          <source>Int. J. Semantic Web Inf. Syst.</source>
          ,
          <volume>5</volume>
          (
          <issue>3</issue>
          ):1{
          <fpage>22</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>P.</given-names>
            <surname>Cimiano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Schultz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sizov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Sorg</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Staab</surname>
          </string-name>
          .
          <article-title>Explicit vs. latent concept models for cross-language information retrieval</article-title>
          .
          <source>In Proceedings of the Int. Joint Conf. on Arti cial Intelligence (IJCAI)</source>
          , pages
          <fpage>1513</fpage>
          {
          <fpage>1518</fpage>
          . AAAI Press,
          <year>July 2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6. G. de Melo and
          <string-name>
            <given-names>G.</given-names>
            <surname>Weikum</surname>
          </string-name>
          .
          <article-title>Towards a universal wordnet by learning from combined evidence</article-title>
          .
          <source>In Proc. of the 18th ACM Conf. on Information and Knowledge Management (CIKM</source>
          <year>2009</year>
          ), pages
          <fpage>513</fpage>
          {
          <fpage>522</fpage>
          , New York, NY, USA,
          <year>2009</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>E. M. Frank</given-names>
            <surname>Manola</surname>
          </string-name>
          .
          <article-title>RDF primer</article-title>
          . http://www.w3.org/TR/rdf-primer/, Feb.
          <year>2004</year>
          . W3C Recommendation.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>J.</given-names>
            <surname>Guyot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Radhouani</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Falquet</surname>
          </string-name>
          .
          <article-title>Ontology-based multilingual information retrieval</article-title>
          .
          <source>In CLEF Workhop, Working Notes Multilingual Track</source>
          , pages
          <volume>21</volume>
          {
          <fpage>23</fpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>D.</given-names>
            <surname>Hiemstra</surname>
          </string-name>
          .
          <article-title>Using language models for information retrieval</article-title>
          . University of twente,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>M.</given-names>
            <surname>Maron</surname>
          </string-name>
          .
          <article-title>On indexing, retrieval and the meaning of about</article-title>
          .
          <source>Journal of the American Society for Information Science</source>
          , pages
          <volume>38</volume>
          {
          <fpage>43</fpage>
          ,
          <year>1977</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>O.</given-names>
            <surname>Medelyan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. H.</given-names>
            <surname>Witten</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Milne</surname>
          </string-name>
          .
          <article-title>Topic indexing with wikipedia</article-title>
          .
          <source>In Proc. of the AAAI WikiAI workshop</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <given-names>D.</given-names>
            <surname>Milne</surname>
          </string-name>
          and
          <string-name>
            <given-names>I. H.</given-names>
            <surname>Witten</surname>
          </string-name>
          .
          <article-title>Learning to link with wikipedia</article-title>
          .
          <source>In Proc. of the 17th ACM Conf. on Information and knowledge management (CIKM)</source>
          , pages
          <fpage>509</fpage>
          {
          <fpage>518</fpage>
          ,
          <string-name>
            <surname>Napa</surname>
            <given-names>Valley</given-names>
          </string-name>
          , California, USA,
          <year>2008</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <given-names>D.</given-names>
            <surname>Milne</surname>
          </string-name>
          and
          <string-name>
            <given-names>I. H.</given-names>
            <surname>Witten</surname>
          </string-name>
          .
          <article-title>An open-source toolkit for mining wikipedia</article-title>
          .
          <source>In Proc. New Zealand Computer Science Research Student Conf.</source>
          , volume
          <volume>9</volume>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>M. Potthast</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Stein</surname>
            , and
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Anderka</surname>
          </string-name>
          .
          <article-title>A wikipedia-based multilingual retrieval model</article-title>
          .
          <source>In ECIR</source>
          , pages
          <volume>522</volume>
          {
          <fpage>530</fpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <given-names>J.</given-names>
            <surname>Savoy</surname>
          </string-name>
          .
          <article-title>Data fusion for e ective european monolingual information retrieval</article-title>
          .
          <source>In Multilingual Information Access for Text, Speech and Images</source>
          , pages
          <volume>233</volume>
          |
          <fpage>244</fpage>
          .
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <given-names>P.</given-names>
            <surname>Sorg</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Cimiano</surname>
          </string-name>
          .
          <article-title>Cross-lingual information retrieval with explicit semantic analysis</article-title>
          .
          <source>In Working Notes for the CLEF 2008 Workshop</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <given-names>P.</given-names>
            <surname>Sorg</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Cimiano</surname>
          </string-name>
          .
          <article-title>An experimental comparison of explicit semantic analysis implementations for cross-language retrieval</article-title>
          .
          <source>In Proc. of the Int. Conf. on Applications of Natural Language to Information Systems (NLDB)</source>
          , pages
          <fpage>36</fpage>
          {
          <fpage>48</fpage>
          . Springer,
          <year>June 2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18. P. Sorg,
          <string-name>
            <given-names>P.</given-names>
            <surname>Cimiano</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Sizov</surname>
          </string-name>
          .
          <article-title>Overview of the cross-lingual expert search (CriES) pilot challenge</article-title>
          .
          <source>In Working Notes of the CLEF 2010 Lab Sessions</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>M. Surdeanu</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Ciaramita</surname>
            , and
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Zaragoza</surname>
          </string-name>
          .
          <article-title>Learning to rank answers on large online QA collections</article-title>
          .
          <source>In Proc. of the 46th Annual Meeting for the Association for Computational Linguistics: Human Language Technologies (ACL-08: HLT)</source>
          ,
          <source>page 719727</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <given-names>M.</given-names>
            <surname>Volk</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Buitelaar</surname>
          </string-name>
          .
          <article-title>Ontologies in cross-language information retrieval</article-title>
          .
          <source>In Proceedings of WOW2003</source>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <given-names>T.</given-names>
            <surname>Stajner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Rusu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Dali</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Fortuna</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Mladenic</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Grobelnik</surname>
          </string-name>
          .
          <article-title>EnrycherService oriented Text Enrichment</article-title>
          .
          <source>Proc. of SiKDD</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>