<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>CLEF 2000 { 2014: Lessons Learnt from Ad Hoc Retrieval (Extended Abstract)?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nicola Ferro</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gianmaria Silvello</string-name>
          <email>silvellog@dei.unipd.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Information Engineering</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Padua</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper reports the outcomes of a longitudinal study on the CLEF Ad Hoc track in order to assess its impact in the last fteen years on the e ectiveness of monolingual, bilingual and multilingual information access and retrieval systems.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Experimental evaluation has been a key driver for research and innovation in
the Information Retrieval (IR) eld since its inception. Large-scale evaluation
campaigns such as Conference and Labs of Evaluation Forum (CLEF)1, are
known to act as catalysts for research by o ering carefully designed evaluation
tasks for di erent domains and use cases and, over the years, to have provided
both qualitative and quantitative evidence about which algorithms, techniques
and approaches are most e ective.</p>
      <p>
        As a consequence, some attempts have been made to determine their
impact [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], however, in the literature there have been few systematic longitudinal
studies about the impact of evaluation campaigns on the overall e ectiveness
of IR systems. One of the most relevant works compared the performances of
eight versions of the SMART system on eight di erent Text REtrieval
Conference (TREC) ad-hoc tasks (i.e. TREC-1 to TREC-8) and showed that the
performances of the SMART system has doubled in eight years [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. On the other
hand, these results \are only conclusive for the SMART system itself" [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and
this experiment is not easy to reproduce in the CLEF context because we would
need to use di erent versions of one or more systems { e.g. a monolingual, a
bilingual and a multilingual system { and to test them on many collections for
a great number of tasks. Furthermore, today's systems increasingly rely on
online linguistic resources (e.g. MT systems, Wikipedia, on-line dictionaries) which
continuously change over time, thus preventing comparable longitudinal studies
even when using the same systems.
      </p>
      <p>
        Therefore, we carry out a longitudinal study on the Ad-Hoc track of CLEF
in order to understand its impact on monolingual, bilingual, and multilingual
? The extended version of this abstract has been published in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
1 http://www.clef-initiative.eu/
retrieval by adopting the score standardization methodology proposed in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. This
methodology allows us to carry out inter-collection comparison between systems
by limiting the e ect of collections (i.e. corpora of documents, topics and
relevance judgments) and by making system scores interpretable in themselves.
      </p>
      <p>For this study we apply standardization to Average Precision (AP)
calculated for all the runs submitted to the ad-hoc tracks of CLEF (i.e. monolingual,
bilingual and multilingual tasks from 2000 to 2007) and to The European
Library (TEL) tracks (i.e. monolingual and bilingual tasks from 2008 to 2009).</p>
      <p>
        All the CLEF results that we analysed in this paper are available through
the Distributed Information Retrieval Evaluation Campaign Tool (DIRECT)
system2 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]; the software library (i.e. MATTERS) used for calculating measure
standardization as well as for analysing the performances of the systems is
publicly available at the URL: http://matters.dei.unipd.it/.
      </p>
      <p>
        In the following we report the main research questions we tackled and we
provide a short summary of the main ndings for each of them. More detailed
experimental results concerning those research questions could be found in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
Finally, we conclude by outlining the work we envision for the future.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Research Questions</title>
      <p>The longitudinal study we carried out was aimed at tackling four research
questions for which we report a brief insight of our ndings.</p>
      <p>RQ1. Do performances of monolingual systems increase over the
years? Are more recent systems better than older ones?</p>
      <p>From the analysis of mean standardized AP (sMAP) across monolingual tasks
we can see an improvement of performances, even though it is not always steady
from year to year, see Table 1. The best systems are rarely the most recent ones;
this may be due to a tendency towards tuning well performing systems relying
on established techniques in the early years of a task while focusing on
understanding and experimenting new techniques and methodologies in later years. In
general, the assumption for which the life of a task is summarized by increase in
system performances, plateau and termination oversimpli es reality: researchers
and developers do not just incrementally adding new pieces on existing
algorithms, rather they often explore completely new ways or add new components
to the systems, causing a temporary drop in performances. Thus, we do not have
a steady increase but rather a general positive trend.</p>
      <p>RQ2. Do performances of bilingual systems increase over the years
and what is the impact of source languages?</p>
      <p>System performances in bilingual tasks show a growing trend across the years
although it is not always steady and it depends on the number of submitted runs
as well as on the number of newcomers. The best systems for bilingual tasks
are often the more recent ones showing the importance of advanced linguistic
resources that become available and improved over the years. Source languages
2 http://direct.dei.unipd.it/</p>
      <p>Task
AH Mono ES
AH Mono DE
TEL Mono DE
AH Mono FR
TEL Mono FR
AH Mono IT
AH Mono NL
have a high impact on the performances of a given target language, showing
that some combinations are better performing than others { e.g. Spanish to
Portuguese has a higher median sMAP than German to Portuguese.</p>
      <p>RQ3. Do performances of multilingual systems increase over the
years?</p>
      <p>Multilingual systems show a steady growing trend of performances over the
years despite the variations in target and source languages from task to task.
We can identify a growing trend of performances especially for top systems. For
instance the multilingual task with four languages reports a major improvement
of median sMAP from 2002 to 2003 even though the top system of 2003 has lower
sMAP than the one of 2002; the multilingual task with 8 languages reports the
lowest median sMAP and, at the same time, the best performing system of all
multilingual tasks.</p>
      <p>RQ4. Do monolingual systems have better performances than
bilingual and multilingual systems?</p>
      <p>Systems which operate on monolingual tasks prove to be more performing
than bilingual ones in most cases, even though the di erence between top
monolingual and top bilingual systems reduces year after year and sometimes the ratio
is even inverted. In some cases, multilingual systems turn out to have higher
performances than bilingual ones and the top multilingual system, as shown in Table
2, has the highest sMAP of all the systems which participated in CLEF tasks
from 2000 to 2009: the work done for dealing with the complexity of multilingual
tasks pays o in terms of overall performances of the multilingual systems.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Future Works</title>
      <p>This study opens up diverse analysis possibilities and as future works we plan
to investigate several further aspects regarding the cross-lingual evaluation
activities carried out by CLEF; we will: (i) apply standardization to other
largelyadopted IR measures { e.g. Precision at 10, RPrec, Rank-Biased Precision, bpref
{ with the aim of analysing system performances from di erent perspectives; (ii)
aggregate and analyse the systems on the basis of adopted retrieval techniques
to better understand their impact on overall performances across the years; and
(iii) extend the analysis of bilingual and multilingual systems grouping them on
a source and target language basis thus getting more insights into the role of
language morphology and linguistic resources in cross-lingual IR.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>M.</given-names>
            <surname>Agosti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. Di</given-names>
            <surname>Buccio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          , I. Masiero,
          <string-name>
            <given-names>S.</given-names>
            <surname>Peruzzo</surname>
          </string-name>
          , and
          <string-name>
            <surname>G. Silvello.</surname>
          </string-name>
          <article-title>DIRECTions: Design and Speci cation of an IR Evaluation Infrastructure</article-title>
          . In T. Catarci,
          <string-name>
            <given-names>P.</given-names>
            <surname>Forner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Hiemstra</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Pen~as, and G. Santucci, editors,
          <source>Proc. of the 3rd Int. Conf. of the CLEF Initiative (CLEF</source>
          <year>2012</year>
          ), pages
          <fpage>88</fpage>
          {
          <fpage>99</fpage>
          . Lecture Notes in Computer Science (LNCS) 7488, Springer, Heidelberg, Germany,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>C.</given-names>
            <surname>Buckley</surname>
          </string-name>
          .
          <article-title>The SMART project at TREC</article-title>
          .
          <source>In TREC | Experiment and Evaluation in Information Retrieval</source>
          , pages
          <volume>301</volume>
          |
          <fpage>320</fpage>
          . MIT Press,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          and
          <string-name>
            <given-names>G.</given-names>
            <surname>Silvello. CLEF 15th</surname>
          </string-name>
          <article-title>Birthday: What Can We Learn From Ad Hoc Retrieval</article-title>
          ? In E. Kanoulas,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lupu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Clough</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sanderson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hall</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hanbury</surname>
          </string-name>
          , and E. Toms, eds,
          <source>Proc. of the 5th Int. Conf. of the CLEF Initiative (CLEF</source>
          <year>2014</year>
          ), pages
          <fpage>31</fpage>
          {
          <fpage>43</fpage>
          . Lecture Notes in Computer Science (LNCS) 8685, Springer, Heidelberg, Germany,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Rowe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. W.</given-names>
            <surname>Wood</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. L.</given-names>
            <surname>Link</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D. A.</given-names>
            <surname>Simoni</surname>
          </string-name>
          .
          <article-title>Economic Impact Assessment of NIST's Text REtrieval Conference (TREC) Program</article-title>
          . RTI International, USA,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>W.</given-names>
            <surname>Webber</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <article-title>Mo at, and</article-title>
          <string-name>
            <given-names>J.</given-names>
            <surname>Zobel</surname>
          </string-name>
          .
          <article-title>Score standardization for inter-collection comparison of retrieval systems</article-title>
          .
          <source>In SIGIR 2008</source>
          , pages
          <fpage>51</fpage>
          {
          <fpage>58</fpage>
          . ACM Press,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>