<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Cultural Heritage in CLEF (CHiC) 2013 - Multilingual Task Overview1</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Vivien Petras</string-name>
          <email>vivien.petras@ibi.hu-berlin.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Toine Bogers</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nicola Ferro</string-name>
          <email>ferro@dei.unipd.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ivano Masiero</string-name>
          <email>masieroi@dei.unipd.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Berlin School of Library and Information Science</institution>
          ,
          <addr-line>Humboldt-Universitätzu Berlin, Dorotheenstr. 26, 10117 Berlin</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Information Engineering, University of Padova</institution>
          ,
          <addr-line>Via Gradenigo 6/B, 35131Padova</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Royal School of Library and Information Science, Copenhagen University</institution>
          ,
          <addr-line>Birketinget 6, 2300 Copenhagen S</addr-line>
          ,
          <country country="DK">Denmark</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The Cultural Heritage in CLEF 2013 multilingual task comprised two sub-tasks: multilingual ad-hoc retrieval and semantic enrichment. The multilingual ad-hoc retrieval sub-task evaluated retrieval experiments in 13 languages (Dutch, English, German, Greek, Finnish, French, Hungarian, Italian; Norwegian, Polish, Slovenian, Spanish, Swedish). More than 140,000 documents were assessed for relevance on a tertiary scale. The ad-hoc task had 7 participants submitting 30 multilingual and 41 monolingual runs. The semantic enrichment task evaluated monolingual and multilingual semantic enrichments (suggestions based on a query) in the same 13 languages. Two participants submitted 10 runs. Results indicated that different languages contribute differently to the overall retrieval effectiveness, probably dependent on collection size. Experiments showed that using more or all of the provided languages usually increases retrieval effectiveness, but not always. For a multilingual task of this scale (13 languages), more participants are necessary in order to provide enough variations in runs to allow for comparative analyses.</p>
      </abstract>
      <kwd-group>
        <kwd>cultural heritage</kwd>
        <kwd>Europeana</kwd>
        <kwd>ad-hoc retrieval</kwd>
        <kwd>semantic enrichment</kwd>
        <kwd>multilingual retrieval</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Cultural heritage collections – preserved by archives, libraries, museums and other
institutions – consist of “sites and monuments relating to natural history, ethnography,
archaeology, historic monuments, as well as collections of fine and applied arts" [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
Cultural heritage content is often multilingual and multimedia (e.g. text, photographs,
images, audio recordings, and videos), usually described with metadata in multiple
formats and of different levels of complexity. Cultural heritage institutions have
dif1 Parts of this paper were already published in the CHIC 2013 LNCS Overview paper [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
ferent approaches to managing information and serve diverse user communities, often
with specialized needs. The targeted audience of the CHiC lab and its tasks are
developers of cultural heritage information systems, information retrieval researchers
specializing in domain-specific (cultural heritage) and / or structured information
retrieval on sparse text (metadata) and semantic web researchers specializing on semantic
enrichment with LOD data. Evaluation approaches (particularly system-oriented
evaluation) in this domain have been fragmentary and often non-standardized. CHiC aims
at moving towards a systematic and large-scale evaluation of cultural heritage digital
libraries and information access systems.
      </p>
      <p>After a pilot lab in 2012, where a standard ad-hoc information retrieval scenario
was tested together with two use-case-based scenarios (diversity task and semantic
enrichment task), the 2013 lab diversifies and becomes more realistic in its tasks
organization. The pilot lab has shown that cultural heritage is a truly multilingual area,
where information systems contain objects in many different languages. Cultural
heritage information systems also differ from some more specified information systems
in that ad-hoc searching might not be the prevalent form of access to this type of
content. The 2013 CHiC lab therefore focuses on multilinguality in the retrieval tasks and
adds an interactive task, where different usage scenarios for cultural heritage
information systems were tested. The multilingual tasks described in this paper required
multilingual retrieval in up to 13 languages, making CHiC the most multilingual
CLEF lab ever.</p>
      <p>CHiC has teamed up with Europeana2, Europe’s largest digital library, museum
and archive for cultural heritage objects to provide a realistic environment for
experiments. Europeana provided the document collection (digital representations of
cultural heritage objects) and queries from their query logs. The interactive task also
provided a topic clustering algorithm and a customized browsable portal based on
Europeana data.</p>
      <p>The paper is structured as follows: Chapter 2 introduces the Europeana document
collection. Chapters 3 and 4 describe the sub-tasks multilingual ad-hoc and
multilingual semantic enrichment in detail, their requirements, participants and results. The
conclusion provides an outlook on the future of CHiC and the potential synergies of
combining ad-hoc and interactive information retrieval evaluation.
2</p>
    </sec>
    <sec id="sec-2">
      <title>The Europeana Collection</title>
      <p>The Europeana information retrieval document collection was prepared for the CHiC
pilot lab in 2012 (Petras et al., 2012). It consists of the complete Europeana metadata
index as downloaded from the production system in March 2012. It contains
23,300,932 documents with a size of 132 GB. With the move of Europeana to an open
data license in the summer of 2012 and the subsequent changes in content, this test
document collection represents a snapshot of Europeana data from a particular time.
However, the overlap to the current content is about 80%.</p>
      <p>The collection consists of metadata records describing cultural heritage objects,
e.g. the scanned version of a manuscript, an image of a painting of sculpture or an
audio or video recording. Roughly, 62% of the metadata records describe images,
35% describe text, 2% describe audio and 1% video recordings.</p>
      <p>The collection was divided into 14 sub-collections according to the language of the
content provider of the record (which usually indicates the language of the metadata
record). A threshold was set: all languages with less than 100,000 documents were
grouped together under the name “Others”. The 13 language collections included
Dutch, English, German, Greek, Finnish, French, Hungarian, Italian; Norwegian,
Polish, Slovenian, Spanish, Swedish. For the CHiC 2013 experiments, all
subcollections except the “Others” were used, totaling roughly 20 million documents.
The 14 sub-collections are listed in table 1.
The XML metadata contains title and description data, media type and chronological
data as well as provider information. For ca. 30% of the records, content-related
enrichment keywords were added automatically by Europeana based on a mapping
between metadata terms and terms from controlled lists like DBpedia names. In the
Europeana portal, object records commonly also contain thumbnails of the object if it
is an image and links to related records. These were not included with the test
collection, but relevance assessors were able to look at them at the original source. Figure 1
shows an extract example record from the Europeana CHiC collection.
The sub- tasks are a continuation of the 2012 CHiC lab, using a similar task scenarios,
but requiring multilingual retrieval and results. Two sub-tasks were defined:
multilingual ad-hoc retrieval and multilingual semantic enrichment.</p>
      <p>The traditional multilingual ad-hoc retrieval task measures information retrieval
effectiveness with respect to user input in the form of queries. The 13 language
subcollections form the multilingual collection (ca. 20 million documents) against which
experiments were run. Participants were asked to submit ad-hoc information retrieval
runs based on 50 topics (provided in all 13 languages) and including at least 2 and at
most all 13 collection languages. For pooling purposes, participants were also asked
to submit monolingual runs choosing any of the collection languages. Because the
topics were provided in all collection languages, the focus of the task was not on topic
translation, but on multilingual retrieval across different collection languages.</p>
      <sec id="sec-2-1">
        <title>3.1 Topic Creation</title>
        <p>A new set of 50 topics was created for the 2013 edition of CHiC, where topic
selection was determined partially by the potential for retrieving a sufficient number of
relevant documents in each of the collection languages. CHiC 2012 used topics from
the Europeana query logs alone, which resulted in zero results for some of the 3
languages [13]. The problem of having zero relevant results is aggravated when
collection languages are varied, especially in the cultural heritage area. Many topics are
relevant for only a few languages or cultures. For 2013, more focus was put on testing
all topics in all languages for retrieving relevant documents, which resulted in fewer
zero relevant result topics. The topic creation process started with creating a pool of
candidate topics, which derived from four different sources:
 15 topics that showed promising retrieval performance were re-used from the
2012 topic set (only in 3 languages) to test their performance in 13 languages.
 Another 19 topics that were not specific to only a handful of languages were
taken from an annotated snapshot of the Europeana query log (the same
procedure was used for the 2012 topics).
 The Polish task also suggested topics, 17 were not considered to be relevant only
in Polish and input in the candidate pool.
 Finally, two of the track organizers generated another 21 test queries covering a
wide range of topics contained in Europeana’s collections that would span all
collection languages.</p>
        <p>
          These 73 candidate topics were then translated into all 13 languages by volunteers.
The translated candidate topics were run against the 13 language collections using
Indri 5.2 with default settings3. We retained the 50 topics that returned the highest
number of relevant documents for all thirteen languages. Another factor that affected
the final selection of the 2013 topics was the abundance of named-entity queries
(around 60%) in the 2012 topic set. While named-entity queries are a common type of
query for Europeana [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], they are less challenging than non-entity queries that
describe a more complex information need. For this we wished to down-sample the
proportion of named-entity queries to around 20%.
        </p>
        <p>The final topics set covers a wide range of topics and consisted of 12 topics from
the 2012 topic set, 13 log-based topics, 13 topics from the Polish subtask, and 12
intellectually derived queries. In form and type, the different query types are
indistinguishable and usually include 1-3 query terms (e.g. “silent film”, “ship wrecks”, and
“last supper”). The underlying information need for a query can be ambiguous if the
intention of the query is not clear. In this case, the track organizers discussed the
query and agreed on the most likely information need. These were not admissible for
information retrieval. Figure 2 shows an example of an English query.
&lt;topic lang="en"&gt;
&lt;identifier&gt;CHIC-004&lt;/identifier&gt;
&lt;title&gt;silent film&lt;/title&gt;
&lt;description&gt;documents on the history of silent film, silent film videos, biographies of
actors and directors, characteristics of silent film and decline of this genre&lt;/description&gt;
&lt;/topic&gt;
3Jelinek-Mercer smoothing with λ set to 0.4 and no stemming or stopword filtering.</p>
      </sec>
      <sec id="sec-2-2">
        <title>3.2 Pooling and Relevance Assessments</title>
        <p>This year, we produced 13 pools, one for each target language using different depths
depending on the language and the available number of documents. The pools were
created using all the submitted runs. A 14th pool, for the multilingual task, is the
union of the 13 pools described above. Table 2 provides details about the created pools,
their size, the number of relevant and not relevant documents, and the pooled runs.</p>
        <p>
          Size
Size
Size
Size
Size
150
11,640
941
342
10,357
43 out of 50
1
We used graded relevance, i.e. highly relevant, partially relevant, and not relevant. To
compute the standard performance measures reported in Section 3.3, we used binary
relevance and conflated highly relevant and partially relevant to just relevant. The
DIRECT system [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] was used to collect runs, perform relevance assessment, and
compute performances. The system’s interfaces and processes were also described in
last year’s CHiC Paper [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]
        </p>
        <p>For all languages except English, native language speakers performed the
relevance assessments. Fifteen assessors took 2 weeks to assess the ca. 140,000
documents. The assessors received detailed instructions on how to use the assessor
interface and guidelines, how the relevance assessments were to be approached. Constant
communication via a common mailing list ensured that assessors across languages
treated topics from the same perspective.</p>
        <p>Despite our efforts in topic creation, some topics in some languages did not have
any relevant documents in the pool. Besides not all queries having relevant documents
in the Europeana collection, the problem was exacerbated by receiving very few
monolingual runs that could be used for pooling, sometimes resulting in very small
pools. While 11 languages have at least 40 topics with relevant documents (5 with 48
or more topics with relevant documents), Finnish (only 16 topics with relevant
documents) and Slovenian (only 37 topics with relevant documents) give raise for concern
in comparative analyses.</p>
      </sec>
      <sec id="sec-2-3">
        <title>3.3 Participants and Runs</title>
        <p>Seven different teams participated in the 2013 edition of the ad-hoc track (table 3).</p>
      </sec>
      <sec id="sec-2-4">
        <title>Group</title>
        <p>CEA LIST
Department of Computer Science, University of Neuchâtel
MRIM/LIG, University of Grenoble
RSLIS, University of Copenhagen &amp; Aalborg University
School of Information, UC Berkeley
Technical University of Chemnitz
University of Westminster</p>
      </sec>
      <sec id="sec-2-5">
        <title>Country</title>
        <p>France
Switzerland
France
Denmark
USA
Germany
Great Britain
Out of the 71 runs submitted, 30 were multilingual runs using at least 2 collection
languages; 10 runs used all available languages for both topics and collections. All
languages were also represented in the monolingual or bilingual runs (41 total).
English, German, French and Italian were the popular languages for the monolingual runs,
all other languages had only 1 or 2 runs. Toine Bogers (RSLIS) provided 2 more
baseline runs for each language collection using the Indri information retrieval system
using language modelling with either the Dirichlet (no stopword list, no stemming) or
the Jelinek-Mercer smoothing algorithm (with stopword list, no stemming), which are
used in the comparison. Table 4 shows the submitted runs and their language
combinations including the baseline runs.
4
4
3
4
1
1
1
1
1
1</p>
        <p>Topic
Language(s)</p>
        <p>DE
EN
ES
FI
FR
IT
NL
EN
IT</p>
        <p>Collection
Language(s)
DE,EN,FR
DE,EN,FR
DE,EN,FR
DE,EN,FR
DE,EN,FR
DE, EN, FR
DE,EN,FR</p>
        <p>EN, IT
EN, IT
1
1
1
1
1
1
1
1
1
3.4</p>
      </sec>
      <sec id="sec-2-6">
        <title>Results &amp; Participant Approaches</title>
        <p>Because of the many variations in topic and collection language configurations,
comparisons between runs is difficult. Since language combinations are then varied by
different system configurations, the matrix of possible impact factors becomes very
big. However, several comparisons can give indications into further research
questions that should be analyzed.
3.4.1 Multilingual Runs: All Languages vs. Fewer languages
Table 5 shows the best multilingual run per participating group ordered by MAP
showing the topic and collection languages that were used for retrieval. Note that only
the best run is selected for each group, even if the group may have more than one top
run.
MULTILINGUALNOEXPANSION
UNINEMULTIRUN5
RSLIS_MULTI_FUSION_COMBS
UM
R005
BERKMLENFRDE19</p>
        <p>Topic
Languages</p>
        <p>All</p>
        <p>All NOT
EL, HU, SL</p>
        <p>All
All</p>
        <p>EN
EN,FR,DE</p>
        <p>Collection
Languages</p>
        <p>All</p>
        <p>All NOT
EL, HU, SL</p>
        <p>All</p>
        <p>All</p>
        <p>EN,IT
EN,FR,DE</p>
        <p>MAP
23.38%
18.78%
Figure 3 shows the best 5 multilingual runs in an interpolated recall vs. average
precision graph.
0,8
0,6
0,4
0,2
0
THOMAS_WILHELM.TUC_ALL_LA
THOMAS_WILHELM.TUC_ALL_HS</p>
        <p>MITRA_AKASEREH.UNINEMULTIRUN5
0
0,1
0,2
0,3
0,4
0,5
0,6
0,7
0,8
0,9
1
It is difficult to interpret these figures in terms of which languages have the most
input for retrieval success as the applied IR systems play a much bigger role in this
cross-system comparison.</p>
        <p>UC Berkeley compared experiments with different topic languages against a
multilingual collection of English, French and German combined. Results show that using
the exact same languages for topics achieves a slightly higher result than using just
one of the topic languages or even more languages (table 6). In this experiment,
differences between runs are probably not all statistically significant. However it is
interesting to note that English and French seem not to contribute to the retrieval
effectiveness as much as German, for example, and that a topic language, which is not
represented in the collection languages (ES) can still achieve almost as high a MAP as
the topic language English.
RSLIS used a similar approach with equivalent results: using one topic language
against the whole multilingual index did result in lower retrieval effectiveness than
the fusion runs using 3 topic languages (table 7).
Both groups found that the German topics seem to have the highest retrieval impact.
The Westminster group [11] showed in a similar experiment that English seemed to
have a higher impact than Italian. More runs would be necessary to be able to perform
a complete analysis.</p>
        <p>Unine experimented with removing topic and collection languages equally and
different fusion algorithms (merging results from separate language indexes) and
showed that leaving out the smaller collection languages can result in an increase in
performance, however, the impact of an individual language is unclear (table 8).
Finally, TU Chemnitz experimented with different stemming algorithms for all
languages and found that using a less aggressive stemmer worked best compared to the
standard rule-based stemmers used in Solr or a no-stemming approach (table 9).</p>
        <sec id="sec-2-6-1">
          <title>3.4.2 Monolingual Runs</title>
          <p>For pooling purposes, participants submitted monolingual runs as well. We can
compare them using the whole multilingual pool (results are also available in the
DIRECT4 system) or using the monolingual pools. While a multilingual pool is what
the real use case prescribes (all languages are potentially relevant), we can also look
at monolingual pools to achieve an improved system comparison (less variation
because of language). We will concentrate on the 4 languages with the most submitted
experiments: English (10), Italian (8), German and French (6). Table 10 shows the
best monolingual run for each participant in those languages.</p>
        </sec>
        <sec id="sec-2-6-2">
          <title>4 http://direct.dei.unipd.it</title>
          <p>Berkeley
RSLIS</p>
          <p>BERKMONOFR02
Unfortunately, only 2 groups (RSLIS &amp; CEA List) submitted runs to all 4 languages
so that a comparison among even those 4 languages becomes difficult.</p>
        </sec>
        <sec id="sec-2-6-3">
          <title>3.4.3 Participant Approaches Table 11 briefly summarizes the participants’ approaches to the ad-hoc track.</title>
        </sec>
      </sec>
      <sec id="sec-2-7">
        <title>Description of approach</title>
        <p>
          Apache Solr with special focus on comparing different types of
stemmers (generic, rule-based, dictionary-based) [12].
Query expansion of a Vector Space model with tf-idf weighting by
using related concepts extracted from Wikipedia using Explicit
Semantic Analysis [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
        </p>
        <p>BASELINE.GER2
CEALISTGERMA
NNOEXPANSION
BERKBIENDE09</p>
        <sec id="sec-2-7-1">
          <title>Neuchâtel</title>
        </sec>
        <sec id="sec-2-7-2">
          <title>RSLIS</title>
        </sec>
        <sec id="sec-2-7-3">
          <title>UC Berkeley</title>
        </sec>
        <sec id="sec-2-7-4">
          <title>Westminster</title>
          <p>
            Language modeling approach using Dirichlet smoothing and
Wikipedia as external document collection to estimate the word
probabilities in case of sparsity of the original term-document matrix [10].
Probabilistic IR using Okapi model with stopword filtering and light
stemming. Collection fusion on the results lists from 13 different
monolingual indexes using z-score normalization merging [
            <xref ref-type="bibr" rid="ref2">2</xref>
            ].
Language modeling with Jelinek-Mercer smoothing and no
stopword filtering or stemming. One run each for English, French, and
German where these topic languages are run against a multilingual
index. Two fusion runs using the CombSUM and CombMNZ
methods combining these three monolingual runs against the multilingual
index [
            <xref ref-type="bibr" rid="ref8">8</xref>
            ].
          </p>
          <p>
            Probabilistic text retrieval model based on logistic regression
together with pseudo-relevance feedback for all of the runs. Runs with
English, French, and German topic sets and sub-collections, as well
translations generated by Google Translate [
            <xref ref-type="bibr" rid="ref4">4</xref>
            ].
          </p>
          <p>Divergence from randomness algorithm using Terrier on the English
and Italian collections [11].
4</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>The CHiC Multilingual Semantic Enrichment Task</title>
      <p>The multilingual semantic enrichment task requires systems to present a ranked list of
related concepts for query expansion. Related concepts can be extracted from
Europeana data or from other resources in the Linked Open Data cloud or other external
resources (e.g. Wikipedia). Participants were asked to submit up to 10 query
expansion terms or phrases per topic. This task included 25 topics in all 13 languages.
Participants could choose to experiment on monolingual or multilingual semantic
enrichments. The suggested concepts were assessed with respect to their relatedness to
the original query terms or query category.</p>
      <p>Only 2 groups participated in the semantic enrichment task, making a comparison
more difficult. Almost all experiments contained either only English concepts or
concepts from several languages (multilingual). In total, 10 experiments were submitted.</p>
      <p>MRIM/LIG (Univ. of Grenoble) used Wikipedia as a knowledge base and the
query terms in order to identify related Wikipedia articles for enrichment candidates.
Both in-links and out-links to and from these related articles (in particular their titles)
were then used to extract terms for enrichment [10].</p>
      <p>
        CEA List used Explicit Semantic Analysis (documents are mapped to a semantic
structure) also with Wikipedia as a knowledge base. Whereas MRIM/LIG used the
title of Wikipedia articles and their in- and out-links for concept expansion, CEA List
concentrated on the categories and the first 150 characters within a Wikipedia article.
When Wikipedia category terms overlapped with query terms, these concepts were
boosted for expansion. In ad-hoc retrieval, the topic and expanded concepts were
matched against the collection and the results were then matched again to a
consolidated version of the topics (favoring more frequent concept phrases) before outputting
the result. For multilingual query expansion, the interlingua links to parallel language
versions of a Wikipedia article were used in a fusion model. For most expansion
experiments, only concepts were considered that appear in at least 3 Wikipedia language
versions, allowing for multilingual expansions [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>The semantic enrichments were evaluated using a tertiary relevance assessment
(definitely relevant, maybe relevant, not relevant) and P@1, P@3 and P@10
measurements. Table 12 shows the results for the best 2 runs for each participants using
either the strict relevance measurement (just definitely relevant) or the relaxed
relevance measurement (definitely relevant and maybe relevant).
Only CEA List experimented with multilingual enrichments. Interestingly, a
multilingual enrichment run was the best with a relaxed relevance measurement, while the
monolingual run was the best with a strict relevance measurement.
5</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusion and Outlook</title>
      <p>The results of this year’s multilingual CHiC task show that multilingual information
retrieval experiments are challenging not only because of the number of languages
that need to be processed but also because of the number of participants necessary in
order to produce comparable results. As the number of possible language variations
increases (CHiC had 13 source languages and 13 target languages), very few
experiments across participants can be compared. While this year’s results have shown that
searching in several languages increases the overall performance (an obvious result),
we could not show which languages contributed more to retrieval results. Future
research in the multilingual task needs to focus on narrower defined tasks (e.g.
particular source languages against the whole collection) or define a GRID experiment
where a particular information retrieval system performs all possible run variation to
arrive at better answers.</p>
      <p>The interactive study collected a rich data set of questionnaire and log data for
further use. Because the task was designed for easy entrance (predetermined system and
research protocol, this is somewhat different that the traditional lab and is planned to
follow a 2-year cycle (assuming the lab’s continuation). In year two, the data gathered
this year should be released to the community in aggregate form having been assessed
by the user interaction community with the goal of identifying a set of objects that
need to be developed. The ad-hoc retrieval tasks can benefit from the interactive task
by re-using the real queries in ad-hoc retrieval test scenarios – effectively merging
both evaluation methods.</p>
      <sec id="sec-4-1">
        <title>Acknowledgements.</title>
        <p>This work was supported by PROMISE (Participative Research Laboratory for
Multimedia and Multilingual Information Systems Evaluation, Network of Excellence
cofunded by the 7th Framework Program of the European Commission, grant agreement
no. 258191. We would like to thank Europeana for providing the data for collection
and topic preparation and providing valuable feedback on task refinement. We would
like to thank Maria Gäde, Preben Hansen, Anni Järvelin, Birger Larsen, Simone
Peruzzo, Juliane Stiller, Theodora Tsikrika and Ariane Zambiras for their invaluable
help in translating the topics. We would also like to thank our relevance assessors
Tom Bekers, Veronica Estrada Galinanes, Vanessa Girnth, Ingvild Johansen,
Georgios Katsimpras, Michael Kleineberg, Kristoffer Liljedahl, Giuliano Migliori,
Christophe Onambélé, Timea Peter, Oliver Pohl, Siri Soberg, Tanja Špec, Emma Ylitalo.
10. Tan, K., Almasri, M., Chevallet, J., Mulhem, P., Berrut, C. Multimedia Information
Modeling and Retrieval(MRIM)/Laboratoire d'Informatique de Grenoble (LIG) at CHiC2013. In
Proceedings CLEF 2013, Working Notes (2013).
11. Tanase, D. Using the Divergence Framework for Randomness: CHiC 2013 Lab Report. In</p>
        <p>Proceedings CLEF 2013, Working Notes (2013).
12. Wilhelm-Stein, T., Schürer, B., Eibl, M. Identifying the most suitable stemmer for the CHiC
multilingual ad-hoc task. In Proceedings CLEF 2013, Working Notes (2013).</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Agosti</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferro</surname>
          </string-name>
          , N.:
          <article-title>Towards an Evaluation Infrastructure for DL Performance Evaluation</article-title>
          . In Tsakonas, G. and
          <string-name>
            <surname>Papatheodorou</surname>
          </string-name>
          , C. (eds.),
          <article-title>Evaluation of Digital Libraries: An Insight to Useful Applications</article-title>
          and Methods, pp
          <fpage>93</fpage>
          -
          <lpage>120</lpage>
          . Chandos Publishing, Oxford, UK (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Akasereh</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Naji</surname>
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Savoy</surname>
            <given-names>J</given-names>
          </string-name>
          . UniNE at CLEF - CHIC
          <year>2013</year>
          .
          <source>In Proceedings CLEF</source>
          <year>2013</year>
          , Working Notes (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3. International Council of Museums (
          <year>2003</year>
          ).
          <article-title>Scope Definition of the CIDOC Conceptual Reference Model</article-title>
          . http://www.cidoc-crm.org/scope.html
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Larson</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <article-title>Pseudo-Relevance Feedback for CLEF-CHiC Adhoc</article-title>
          .
          <source>In Proceedings CLEF</source>
          <year>2013</year>
          , Working Notes (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Petras</surname>
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferro</surname>
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gäde</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Isaac</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kleineberg</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Masiero</surname>
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nicchio</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stiller</surname>
            <given-names>J</given-names>
          </string-name>
          .
          <article-title>Cultural Heritage in CLEF (CHiC) Overview 2012</article-title>
          .
          <source>In Proceedings CLEF-2012</source>
          , Working Paper (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Petras</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bogers</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toms</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hall</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Savoy</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Malak</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pawłowski</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferro</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Masiero</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <article-title>Cultural Heritage in CLEF (CHiC) 2013</article-title>
          .
          <source>In Proceedings of CLEF</source>
          <year>2013</year>
          , LNCS, Springer (forthcoming).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Popescu</surname>
            ,
            <given-names>A. CEA</given-names>
          </string-name>
          <article-title>LIST's participation at the CLEF CHiC 2013</article-title>
          .
          <source>In Proceedings CLEF</source>
          <year>2013</year>
          , Working Notes (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Skov</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bogers</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lund</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jensen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wistrup</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Larsen</surname>
            ,
            <given-names>B. RSLIS</given-names>
          </string-name>
          /AAU at CHiC
          <year>2013</year>
          .
          <source>In Proceedings CLEF</source>
          <year>2013</year>
          , Working Notes (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Stiller</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gäde</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Petras</surname>
          </string-name>
          ,
          <string-name>
            <surname>Vivien</surname>
          </string-name>
          (
          <year>2010</year>
          ).
          <article-title>Ambiguity of Queries and the Challenges for Query Language Detection</article-title>
          .
          <source>In CLEF 2010 LABs and Workshops</source>
          . Retrieved from http://clef2010.org/resources/proceedings/clef2010labs_submission_41.pdf
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>