<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview of the Cross-lingual Expert Search (CriES) Pilot Challenge</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Philipp Sorg</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Philipp Cimiano</string-name>
          <email>cimiano@cit-ec.uni-bielefeld.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Antje Schultz</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sergej Sizov</string-name>
          <email>sizovg@uni-koblenz.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Cognitive Interaction Technology, Center of Excellence (CITEC), Bielefeld University</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Information Systems &amp; Semantic Web, University of Koblenz</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Institute AIFB, Karlsruhe Institute of Technology</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper provides an overview of the cross-lingual expert search pilot challenge as part of the cross-lingual expert search (CriES) workshop collocated with the CLEF 2010 conference. We present a detailed description of the dataset used in the challenge. This dataset is a subset of an o cial crawl of Yahoo! Answers published in the context of the Yahoo! Webscope program. Further we describe the selection process of the 60 multilingual topics used in the challenge. The Gold Standard for these topics was created by human assessors who evaluated pooled results of submitted runs. We present data showing that the experts relevant for our chosen topics indeed speak di erent languages. This corroborates the fact that we need to design retrieval systems that build on a cross-lingual notion of relevance for the expert retrieval task. Finally we summarize the results of the four groups that participated in this challenge using standard evaluation measures. Additionally we also analyze the overlap of retrieved experts in the submitted runs.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The CriES workshop | Cross-lingual Expert Search: Bridging CLIR and Social
Media | addresses the problem of multilingual expert search in social media
environments. The main topics are multilingual expert retrieval methods, social
media analysis with respect to expert search, selection of datasets and evaluation
of expert search results.</p>
      <p>In this paper we describe the pilot challenge as part of the CriES workshop.
This includes a detailed description of the dataset, the selection process for the
topics used in the challenge and the evaluation methodology including relevance
assessment. We also present an overview of the results submitted by the
participating groups.</p>
      <p>Motivation. Online communities generate major economic value and form
pivotal parts of corporate expertise management, marketing, product support, CRM,
product innovation and advertising. In many cases, large-scale online
communities are multilingual by nature (e.g. developer networks, corporate knowledge
bases, blogospheres, Web 2.0 portals). Nowadays, novel solutions are required to
deal with both the complexity of large-scale social networks and the complexity
of multilingual user behavior.</p>
      <p>At the same time, it becomes more and more important to e ciently identify
and connect the right experts for a given task across locations, organizational
units and languages. The key objective of the lab is to consider the problem of
multilingual retrieval in the novel setting of modern social media leveraging the
expertise of individual users.</p>
      <p>Pilot Challenge Topic and Goals. We instantiate the problem setting by an
expert nding task, i.e. our goal is to identify the expertise of online community
members and to provide expert suggestions for solving new problems, questions,
or help requests in multilingual social media. In many cases, expert users in
online communities are multilingual, i.e. they participate in discussions in several
languages. Frequently, the actual expertise of the user is language-independent,
so he/she could provide meaningful assistance and support to questions and
requests stated in any of the known languages. The combined analysis of
multilingual user contributions (e.g. answers or postings from the past) together with
mining of his social environment (e.g. interaction with other community
members in the past, contact/favorite lists, etc.) may provide better indications that
the user has the necessary expertise for addressing the request irrespective of
the language. The key research challenges addressed by the expert nding task
can be summarized as follows:
{ User characterization: the use of multilingual evidence of social media for
building expert pro les;
{ Community analysis: mining of social relationships in collaborative
environments for multilingual retrieval scenarios;
{ User-centric recommender algorithms: development of retrieval and
recommendation algorithms that allow for similarity search and ranked retrieval of
expert users in online communities (in contrast to more common document
retrieval tasks).
2</p>
    </sec>
    <sec id="sec-2">
      <title>Dataset</title>
      <p>
        We used the dataset from the Yahoo! Answers portal introduced by Surdeanu
et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].4 Yahoo! Answers is currently the biggest community QA portal.
According to Google ad planner statistics5 the portal attracts 97M unique visitors
and 1.1B page views per month.6
4 This dataset is provided by the Yahoo! Research Webscope program (see http://
research.yahoo.com/) under the following ID: L6. Yahoo! Answers Comprehensive
Questions and Answers (version 1.0)
5 http://www.google.com/adplanner/
6 Statistics from 2010/05/17
      </p>
      <p>The dataset published by Yahoo! contains 4.5M questions with 35.9M
answers. For each question one answer is marked as best answer. In the portal,
the best answer is determined by either the user who submitted the question
or via other users ratings. The dataset contains IDs of authors of questions and
best answers, whereas authors of non-best answers are anonymous. Questions
are organized into categories which form a category taxonomy.</p>
      <p>The dataset used in the CriES pilot challenge is a subset of the Yahoo!
Answers Webscope dataset, considering only questions and answers in the topic
elds de ned by the following three top level categories including their sub
categories: \Computer &amp; Internet", \Health" and \Science &amp; Mathematics". As our
approach is targeted at the expert retrieval problem, by choosing very
\technical categories" our goal was to yield a dataset with a high number of technical
questions requiring domain expertise to be answered. As many questions in the
dataset serve the only purpose of diversion, it was important to nd categories
were the share of such questions is low.</p>
      <p>In the challenge, 4 languages are considered: English, German, French and
Spanish. As category names are language-speci c, questions and answers from
categories corresponding to the selected categories in other languages are also
included, e.g. \Gesundheit" (German), \Sante" (French) and \Salud" (Spanish)
that correspond to the \Health" category.</p>
      <p>The selected dataset consists of 780,193 questions, each question having
exactly one best answer. The answers were posted by 169,819 di erent users, i.e.
potential experts in our task. The answer count per expert follows a power log
distribution, i.e. 54% of the experts posted one answer, 93% 10 or less, 96% 20 or
less. 410 experts published answers in more than one language. These
multilingual experts posted 8,976 answers, which shows that they are active users in the
portal. The language of questions and answers are distributed over languages as
shown in the following table:</p>
      <sec id="sec-2-1">
        <title>Category</title>
        <p>Comp. &amp; Internet
Health
Science &amp; Math.</p>
      </sec>
      <sec id="sec-2-2">
        <title>Questions EN</title>
        <p>317,074 89%
294,944 95%
185,994 91%</p>
        <p>Language share</p>
        <p>DE FR
1% 3%
1% 2%
1% 2%</p>
        <p>ES
6%
2%
6%
2.1</p>
        <p>Topic Selection
The topics we use in the challenge consist of 15 questions in each language (60
topics in total). As these topics are questions posted by users of the portal, they
express a real information need. Considering the multilingual dimension of our
expert retrieval scenario, we de ned the following criteria for topic selection to
ensure that the selected topics are indeed applicable for our task:
{ International domain. People from other countries should be able to answer
the question. In particular, answering the question should not require
knowledge that is speci c to a geographic region, country or culture. Examples:
Pro: Why doesn't an optical mouse work on</p>
        <p>a glass table?</p>
        <p>Contra: Why is it so foggy in San Francisco?
{ Expertise questions. As the goal of our system is to nd experts in the domain
of the question, all questions should require domain expertise to answer
them. This excludes for example questions that ask for opinions or do not
expect an answer at all. Examples:</p>
        <p>Pro: What is a blog?
Contra: What is the best podcast to subscribe</p>
        <p>to?
We performed the following steps to select the topics:
1. Selection of 100 random questions per language from the dataset (total of
400 candidate topics).
2. Manual assessment of each candidate topic by three human assessors. They
were instructed to check the ful llment of the criteria de ned above.
3. For each question the language coverage was computed. The language
coverage tries to quantify how much potentially relevant experts are contained
in the dataset for each topic and for each of the di erent languages. The
language coverage was calculated by translating a topic into the di erent
languages (using Google Translate) and then using a standard IR system
to retrieve all the expert pro les that contain at least one of the terms in
the translated query. Topics were assigned high language coverage if they
matched an average number of experts in all of the languages. In this way
we ensure that the topics are well covered in the di erent languages under
consideration but do not match too many experts pro les in each language.
This is important for our multilingual task as we intend to nd experts in
di erent languages.
4. Candidate questions are sorted rst by the manual assessment and then by
language coverage. The top 15 questions in each language were selected as
topics.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Relvance Assessment</title>
      <p>We used result pooling for the evaluation of the retrieval results of the
participating groups. For each pooled run, the top 10 experts were pooled and evaluated.
The assessment of experts was based on expert pro les. Assessors received
topics and the complete pro le of experts, consisting of all answers posted by the
expert in question. Based on this information they assigned topic-expert tuples
to the following relevance classes:
2 Expert is likely able to answer.
1 Expert may be able to answer.
0 Expert is probably not able to answer.
0
2
0
1
5
3
0
3
5
2
0
2
5
1
0
1
0</p>
      <p>0
de
en
es
fr
de
en
es
fr
(a) English Topics
(b) German Topics
5</p>
      <p>5
0</p>
      <p>0
de
en
es
fr
de
en
es
fr
(c) French Topics
(d) Spanish Topics</p>
      <p>The assessors were instructed to only use evidence in the dataset for their
judgments. It is assumed that experts expressed all their knowledge in the answer
history and will not have expertise about other topics, unless it can be inferred
from existing answers.</p>
      <p>Overall, 6 assessors evaluated 7,515 pairs of topics and expert pro les. The
distribution of relevant users for the topics in the four di erent languages is
presented in Figure 1. In order to visualize the multilingual nature of the task
we also classi ed relevant users to languages using their answers in the dataset.
The distribution of relevant users for the topics in the four languages is shown
separately for each user group. The analysis of the relevant user distribution
shows that for all topics the main share of relevant users publish answers either
in the topic language or in English. This motivates the cross-language expert
retrieval task we consider, as mono-lingual retrieval in the topic language or
cross-lingual retrieval from the topic language to English do clearly not su ce.
The number of relevant experts posting in a di erent language than the topic
language or English constitute a small share. However the percentage is large
enough | for example Spanish experts for German topics | in order not to
consider these experts.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <p>
        Baseline. In addition to the submitted runs, we de ned a standard IR
baseline: BM25+Z-Score. This baseline uses language speci c indexes of expert text
pro les. These pro les consist of all former answers of each expert in a speci c
language. Topics are translated to each language using Google Translate and
the BM25 model [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] is used to get language speci c results. Using the Z-Score
normalization [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], the nal scores for each expert for a speci c topic are obtained
by aggregation.
      </p>
      <p>Evaluation of Submitted Runs. Four di erent groups participated in the pilot
challenge. Results based on the relevance assessment of the top 10 retrieved
experts are presented in Figure 1. In addition to the submitted runs we also
present results for the baseline de ne above. We use two di erent evaluation
measures: Precision at cuto level 10 (P@10) and Mean Reciprocal Rank (MRR).
The best runs achieved promising retrieval results with P@10 of :62 (strict
assessment, iftene run2) and :87 (lenient assessment, herzig 3-boe-07-02-01-q01m).
Both runs signi cantly improve the baseline that achieves P@10 of :19 (strict
assessment) and :39 (lenient assessments). Precision / Recall curves for each run
are presented in Figure 2 using strict assessment and in Figure 3 using lenient
assessment.</p>
      <p>Overlap of Retrieved Experts. The overlap of retrieved experts between runs is
presented in Table 2. Comparing any combination of two runs, the presented
numbers correspond to the count of retrieved experts for each topic that are not
retrieved by both runs.
0,1
0,2
0,3
0,4
0,5
0,6
0,7
Fig. 2. Precision/Recall Curves based on interpolated Recall (strict assessment).
1
0,9
0,8
0,7
0,6
0,5
0,4
0,3
0,2
0,1
0</p>
      <p>The presented statistics show that the overlap of retrieved experts across
the four groups is very low. Even the best performing runs of two di erent
groups (iftene run2, herzig 3-boe-07-02-01-q01m) have a small overlap of 14%
while having similar values for P@10 and MRR. This shows that the combination
of di erent approaches is an important topic for future work.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <p>This work was funded by the Multipla project7 sponsored by the German
Research Foundation (DFG) under grant number 38457858 as well as by the Monnet
project8 funded by the European Commission under FP7.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Robertson</surname>
            ,
            <given-names>S.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Walker</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>Some simple e ective approximations to the 2-Poisson model for probabilistic weighted retrieval</article-title>
          .
          <source>In: Proceedings of the 17th International Conference on Research and Development in Information Retrieval (SIGIR)</source>
          . pp.
          <volume>232</volume>
          |
          <fpage>241</fpage>
          . Springer, Dublin (
          <year>1994</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Savoy</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Data fusion for e ective european monolingual information retrieval</article-title>
          .
          <source>In: Multilingual Information Access for Text, Speech and Images</source>
          , pp.
          <volume>233</volume>
          |
          <issue>244</issue>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Surdeanu</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ciaramita</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zaragoza</surname>
          </string-name>
          , H.:
          <article-title>Learning to rank answers on large online QA collections</article-title>
          .
          <source>In: Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics (ACL)</source>
          . pp.
          <volume>719</volume>
          {
          <fpage>727</fpage>
          .
          <string-name>
            <surname>Columbus</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ohio</surname>
          </string-name>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>