<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview of the ImageCLEF 2013 Personal Photo Retrieval Subtask</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>David Zellhofer</string-name>
          <email>david.zellhoefer@tu-cottbus.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Brandenburg Technical University, Database and Information Systems Group</institution>
          ,
          <addr-line>Walther-Pauer-Str. 1, 03046 Cottbus</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>The subtask assesses the retrieval e ectiveness in di erent retrieval usage scenarios in a personal photo collection and with di erent user groups. That is, the subtask reveals whether a tested algorithm is stable in terms of e ectiveness for 7 di erent user groups. This perspective on retrieval performance evaluation separates the 2013 version of the subtask from its pilot phase although it relies on the same data set. The data set has been sampled from 19 layperson photographers and consists of 5,555 unprocessed digital photographs. To solve the subtask, the participants are asked to retrieve the 100 best matching documents for 74 sample information needs that consist of visual concepts and events. Each sample information need is modeled by at most one query-by-example document and up to three to browsed documents. The best performing groups, ISI and DBIS, used visual low-level features and metadata to solve the task. The current best-placed run achieves a nDCG at 20 of 0.7427 for the average user group using relevance feedback and all available modalities, i.e., visual data and metadata such as Exif or GPS information. Regarding the stability, roughly 50% of the submitted runs perform equally well over all user groups.</p>
      </abstract>
      <kwd-group>
        <kwd>Content-Based Image Retrieval</kwd>
        <kwd>Benchmark</kwd>
        <kwd>Experiments</kwd>
        <kwd>Personal Photograph Collection</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Following a pilot task phase in 2012, the personal photo retrieval task has become
an o cial subtask of the ImageCLEF photo annotation and retrieval challenge in
2013. The subtask focuses on di erent retrieval usage scenarios and user groups.
That is, the subtask reveals whether the tested algorithms are stable in terms
of retrieval quality for di erent user groups or not. This perspective on retrieval
performance evaluation separates the 2013 version of the subtask from its pilot
phase [
        <xref ref-type="bibr" rid="ref9">7</xref>
        ] although it relies on the same data set. The data set has been
sampled from 19 layperson photographers and consists of 5,555 unprocessed digital
photographs. A detailed description of the data set is available in [
        <xref ref-type="bibr" rid="ref8">6</xref>
        ].
      </p>
      <p>In contrast to system-centric (Cran eld-based) benchmarks, the subtask tries
to establish a more user-centered perspective on multimodal information
retrieval (MIR) and content-based image retrieval (CBIR). This objective is
reected by three design choices of the subtask.</p>
      <p>
        First, the subtask is not only providing sample information needs (IN) with
one or more relevant query documents that have to be used in a query by
example (QBE) fashion. In order to simulate the user's interaction with the MIR
system, browsed documents are provided in addition to a number of query
documents. Unlike the query documents, browsed documents are not necessarily
fully relevant for a given topic. Instead, they vary in their level of relevance and
can also be totally irrelevant, e.g., to model erroneous user input caused by a
click on an image document that has nothing to do with the current IN but that
grabbed the user's attention. From a wider perspective, this form of IN speci
cation re ects the transition between di erent search strategies that have been
described, e.g., by [
        <xref ref-type="bibr" rid="ref7">5</xref>
        ] or [
        <xref ref-type="bibr" rid="ref1 ref5 ref6">1</xref>
        ].
      </p>
      <p>
        Second, the subtask respects the gradual relevance of documents with respect
to an IN. That is, the subtask's ground truth is based on graded relevance
judgments. Consequently, an appropriate metric, nDCG [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], is used for the evaluation
of the participants' submission (see Section 4.1).
      </p>
      <p>
        Third, the subtask acknowledges the subjectivity of relevance assessments.
Because user groups that interact with an MIR system and their subjective
notion of relevance vary, multiple ground truths are provided for di erent user
groups. Hence, it becomes possible to assess the stability of an examined retrieval
algorithm in terms of retrieval e ectiveness. This experimental idea was
motivated by a preliminary study [
        <xref ref-type="bibr" rid="ref9">7</xref>
        ] with the data obtained from the participants of
the 2012 pilot task indicating that the algorithms' retrieval performances vary
amongst di erent user groups.
      </p>
      <p>The paper is structured as follows. The next section brie y describes the
resources of the subtask, i.e., the data set, the ground truth, and the
accompanying baseline system. Section 3 describes the sample information needs (or topics)
that are used for the assessment of the retrieval e ectiveness of an investigated
retrieval algorithm. Section 4 discusses the results of all participants of the
subtask, while the last Section 5 concludes the paper.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Resources of the Subtask</title>
      <p>
        The subtask relies on the Pythia dataset [
        <xref ref-type="bibr" rid="ref8">6</xref>
        ] which equates the data set of the
2012 pilot task on personal photo retrieval that will be described in this section.
Hence, the description resembles the publication of 2012 in large parts [7, cf.
pp. 1-12]. To complete the description of the provided resources, Section 2.2
will comment on the acquisition of the ground truth. The following section will
then discuss the elicitation of the browsing data o ered to the participants as
an additional resource.
      </p>
      <sec id="sec-2-1">
        <title>The Pythia Dataset</title>
        <p>
          To overcome limitations by binary relevance judgments often found in common
test collections, the Pythia collection [
          <xref ref-type="bibr" rid="ref8">6</xref>
          ] has been proposed. The collection is
aiming at providing a benchmark for user-centered or relevance feedback-related
experiments which are a ected by subjective relevance levels in particular. The
collection di ers from collections consisting of Flickr downloads or the like as
it has been sampled from 19 layperson photographers. In addition to the
image data, the contributors to the collection completed a survey (see Section 6)
asking for their photograph taking behavior, their demographics etc. To ensure
a variance in photographic motifs and style, the contributors have been chosen
from di erent demographic groups. Thus, one can interpret the content of the
collection as a mirror of a photographer's lifespan with typical changing usage
behaviors, cameras, topics, and places. The total size of the collection is 5,555
documents.
        </p>
        <p>The documents within the collection have neither been processed extensively
nor have duplicates been removed. Hence, the data can be considered a
realistic sample from a typical user's hard-disk. The collection is rich on metadata
including GPS, IPTC, EXIF, and information about events depicted on each
photography. All this information is available to the participants of the subtask.
For an overview, see Table 1.
To obtain the ground truth, 42 assessors were asked to participate. With the
help a web-based evaluation tool (see [7, Fig. 3]), the assessors could judge the
relevance of an image with respect to a sample IN (topic) on a graded scale
ranging from 0 (irrelevant) to 3 (fully relevant). All assessors had to judge all
documents with respect to a topic. The topics were associated with the assessors
by random. To keep them motivated, the assessors were allowed to work with
the collection from a place of their choice. Additionally, they could pause an
assessment run and continue from later on. A time constraint has not been
de ned. In average 2.69 topics were evaluated per assessor (standard deviation:
1.60). The individual assessments were saved separately in order to maintain
them for later usage.
not (see Section 3).</p>
        <p>In order to associate the relevance assessments with di erent user groups,
the assessors had to answer a questionnaire (see Section 6). The questionnaire's
outcome is listed in Table 5. The core characteristics of the assessor group can
be subsumed as follows. The majority of the assessors (28 out of 42) are male
and born between 1979 and 1991 (median: 1987). Most of the assessors are
students with a background in economics (26), the second largest group (13)
has a background in computer science and information technology. Regarding
their level of expertise in the eld of MIR or IR, 9 assessors took classes in MIR
while 11 heard IR. When asked directly about their knowledge of the
eld the
median lies at \little knowledge" with an average of 1.40, i.e., a trend towards
considering themselves as an `informed outsiders".</p>
      </sec>
      <sec id="sec-2-2">
        <title>Calculation of the Ground Truths for each Topic Based on the individual</title>
        <p>assessments, a ground truth for the average user group has been calculated.
First, the frequency of each graded relevance judgement (out of an interval from
0 (irrelevant) to 3 (fully relevant)) was counted per image and topic. Based on
these relevance judgment frequencies, an estimation value was calculated and
rounded. The rounded estimation value of the relevance of an image regarding a
topic was then used as the averaged graded relevance assessment for this image.
In consequence, each image could be associated with a graded relevance judgment
for each topic.</p>
        <p>In addition to the average user group ground truth, 6 representative user
groups could be de ned on the basis of the demographics of the assessors that
are listed in Table 5. For each of the user groups listed below, a distinct ground
truth was derived. In principle, the acquisition of the user group-speci c ground
truths follows the aforementioned process with the di erence that it relies only
on relevance assessments that are associated with the speci c user group (e.g.,
expert MIR users). In the event of a missing relevance assessment for the
topicuser group combination, the assessment is taken from the average ground truth.
The resulting user groups are as follows:
Experts A group of users that stated that they have an expertise with IR.
Non-Experts The complement of the experts group.</p>
        <p>Male/Female The assessors divided by gender.</p>
        <p>IT This groups consists of assessors with an IT background.</p>
        <p>Non-IT The complement of the IT group.</p>
        <p>Generation of the Browsing Information As we could not obtain real
browsing information, it had to be generated arti cially. Using the graded
relevance assessments, multiple images were chosen as browsing images. The
provided browsed images have a relevance grade ranging from 0 to 3, i.e., they range
from irrelevant to fully relevant for a given topic. In other words, the browsing
data consists of interesting images which were not satisfying the information
need of the modeled user and motivated him or her to proceed with the search.
In contrast to 2012, browsed images could also be irrelevant in order to include
erroneous user input.
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Baseline System</title>
        <p>In addition to the resources of the 2012 pilot task, the participants were given
access to a baseline system that can be used for feature extraction and similarity
calculation. The baseline system1 is available for Linux, Mac OS X, and Windows
as C++ source code and is licensed under the Apache License version 2.0. All
participants were free to use the system that o ers 17 global and local visual
features (and some variants).
3</p>
        <p>
          Description of the Sample Information Needs
Unlike in the subtask's pilot phase, the sample information needs (topics) are
no longer subdivided into events and visual concepts. An event in the sense
of this subtask can be a rock concert or a holiday trip to a region or city. In
contrast, a visual concept is a depiction of an object, e.g., a house or a street
scene. Table 2 lists all topics including their title and their associated event class
(see [
          <xref ref-type="bibr" rid="ref8">6</xref>
          ] on the WordNet-based event classes)2. The topics 50 to 81 are taken
1 See http://imageclef.org/2013/photo/retrieval.
2 Please note that the focus on events representing a holiday or a city trip is not
a freely chosen bias. Instead, it re ects the state of randomly picked images from
real-world personal photo collections [
          <xref ref-type="bibr" rid="ref8">6</xref>
          ].
without modi cations from the pilot phase's topic set [
          <xref ref-type="bibr" rid="ref9">7</xref>
          ]. The titles were not
made available to the participants of the subtask contrasting to 2012 in avoid
a manual optimization towards events or visual concepts based on the titles.
Additional training data was not released.
        </p>
        <p>For each topic, the sample IN is modeled by at most one fully relevant QBE
document and/or a sequence of up to three browsed documents of varying
relevance with respect to the IN. 10.81% of the topics (i.e., topics 15, 17, 21, 22, 24,
28, 36, and 42; see Table 2) contain irrelevant browsed documents. The number
of topics has been increased to 74 in comparison to 39 during the pilot phase.
To summarize, the subtask can be considered more complex in comparison to
the pilot, because the IN speci cation o ers less reliable information that can
be exploited for query construction.</p>
        <p>In consequence, the best matching documents for each topic are expected
to be retrieved ad hoc without additional knowledge about the user's context.
That is, all participants have to rely on at most one QBE document and/or
browsing data and are asked to nd the best matching documents illustrating
an event or depicting a visual concept. Thus, an additional objective of this task
is to nd out whether the participating retrieval systems can exploit data from
di erent search strategies, i.e., query-by-example and browsing data, in order
to nd both visual concepts and photos depicting events. To solve the task, the
participants have access to pre-extracted visual low-level features, metadata (e.g.
GPS information), but are also free to use their own techniques.
4</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <p>In comparison to the pilot phase of the subtask, the participation rate could be
increased by ca. 233 %. In 2012, only 3 groups submitted results. This year, 7
groups participated in the subtask, i.e., ca. 38 % of the groups that took part in
the ImageCLEF 2013 photo annotation and retrieval challenge. Unfortunately,
none of the last year's participants could be motivated to submit runs to the
2013 subtask.
4.1</p>
      <sec id="sec-3-1">
        <title>Evaluation Metrics</title>
        <p>
          As said in the introduction, the relevance of a document with respect to an IN
is both highly subjective and relative. That is, a document can be very relevant
for an IN while another can be of little value in comparison. To re ect this
fact, the presented ground truths are based on a gradual scale of relevance.
Unfortunately, traditional measurements such as the mean average precision
(MAP) or precision at n cannot deal with this kind of judgements. Hence, the
subtask's retrieval e ectiveness evaluation relies on the normalized discounted
cumulative gain (nDCG) measurement [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. As stated in [
          <xref ref-type="bibr" rid="ref8">6</xref>
          ] \DCG also provides
more appropriate means to evaluate relevance feedback (RF) or adaptive systems
as it can be used to measure slight changes or re-orderings of relevant documents
with varying degrees of relevance within the result list". The core idea of DCG
is to apply \a discount factor to the relevance scores in order to devaluate
lateretrieved documents" [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. In other words, the metric rewards highly relevant
documents at the rst positions in the result ranking and punishes systems
retrieving less relevant documents at the rst places. A full discussion of the
metric is available by Jarvelin and Kekalainen [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. For the scope of this task, the
DCG implementation of trec eval version 9.0 with standard discount settings
is used. For the sake of completeness, MAP at a cut-o level of 100 is also used.
4.2
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Results of the Participants</title>
        <p>{ visual features alone (IMG)
{ visual features and metadata (IMGMET)
{ visual features and browsing data (IMGBRO)
{ metadata alone (MET)
{ metadata and browsing data (METBRO)
{ browsing data alone (BRO)
{ a combination of all modalities (IMGMETBRO)</p>
        <p>The best performing groups, ISI and DBIS, used visual low-level features
and metadata to solve the task. While ISI used relevance feedback for 4 of
their 5 runs, DBIS used this technique only for run #3. Table 4 shows all the
results of the di erent runs ordered by nDCG at cut-o level 20 for the average
user group. The complete results for all other user groups are available on the
subtask's website3. In accordance with the ndings of the last years' ImageCLEF
tasks, there is evidence that the utilization of multiple modalities increases the
retrieval e ectiveness.
3 http://imageclef.org/2013/photo/retrieval#results</p>
        <p>
          The current best-placed run achieves a nDCG at 20 of 0.7427 (average user
group) using relevance feedback and all available modalities (IMGMETBRO). In
the last year, the best group achieved a nDCG at 20 of 0.5459 for visual concepts
(IMGMET, no RF) and a NDCG at 20 of 0.9697 for event retrieval (MET, no
RF). Please note that the values values are not meant to be compared directly
because of the adjustments made to the subtask in 2013. Unfortunately, none
of the former participants could be motivated to submit runs this year. Thus,
a statement about an in- or decrease of retrieval e ectiveness cannot be made
on the basis of the submitted runs. To complicate the matter, only the
rst two
groups have published their algorithms and approaches towards the task. Hence,
we cannot provide a complete methodology or retrieval type listing in Table 3.
For a description of the methods used by the two
rst groups, see [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] and [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ].
the e ectiveness variance over the di erent user groups. The y-axis shows the
obtained rank and the x-axis the run ID that is listed in Table 3. Figure 1
shows clearly that roughly 50% of the submitted runs have a low rank variance.
That is, they perform equally well for all examined user groups. The other half
{ predominantly the weak performing runs { is not very stable. Whether this
e ect correlates with the used features, matching algorithms, or other variables
remains an area for further research and cannot be investigated in this paper
due to the missing publications.
Although the participation rate in the ImageCLEF 2013 subtask on personal
photo retrieval is high, the low publication rate of the participants complicate
an interpretation of the results. Anyhow, the results of this subtask strengthen
the central nding of the last years of ImageCLEF: the combination of multiple
modalities does improve the retrieval e ectiveness.
        </p>
        <p>The interpretation of the stability of the submitted runs indicates that there
might be a correlation between the e ectivity and stability of an algorithm. In
other words, the better one's algorithm performs the more likely it is that it
will do so for di erent user groups. Whether this e ect is due to other (hidden)
variables remains an open question. Maybe this question motivates the missing
participants to publish their algorithms and approaches towards the solution of
the subtask. Because of the low publication rate, a general interpretation of the
results is hardly possible.</p>
        <p>Another interesting result of the conducted experiment is that both leading
groups { ISI and DBIS { perform almost equally well although ISI is relying on
sophisticated techniques such as Fisher vectors and local features while DBIS
uses global low-end features embedded in a logical query language. Given the
fact, that local features are computationally more intensive than global features,
one might further investigate the logical combination of global features in order
to achieve comparable results at less computational costs.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Content of the Usage Questionnaire</title>
      <p>{ Year of Birth
{ Gender
{ Job Type 1) Pupil, 2) In job training, 3) Student, 4) Fully employed, 5)</p>
      <p>Part-time employed, 6) Not employed, 7) Retired, 8) Other
{ Field of Study / Job Training
{ Course Level</p>
      <sec id="sec-4-1">
        <title>Q0: Have you visited one or more oft he following lectures?</title>
        <p>IR) Information Retrieval, MR) Multimedia Retrieval</p>
      </sec>
      <sec id="sec-4-2">
        <title>Q1: Are you familiar with the principles of content-based information retrieval?</title>
        <p>0) No, 1) A little, 2) I am an informed outsider, 3) Very much, 4) I am an
expert
Q2: Are you colorblind? 0) I don't know, 1) No, 2) Yes</p>
      </sec>
      <sec id="sec-4-3">
        <title>Q3: How many minutes do you use the internet per day?</title>
        <p>0) Not at all, 2) 1 - 30 minutes, 3) 31 - 60 minutes. 4) 61 - 90 minutes, 5)
91 - 120 minutes, 6) More than 120 minutes, 7) More than 240 minutes</p>
      </sec>
      <sec id="sec-4-4">
        <title>Q4: Do you know Web 2.0 services such as Flickr or Fotocommunity.de</title>
        <p>for sharing holiday, family or other photographs with friends?
0) Never heard of it, 2) Know it by name, 3) I have visited such websites, 4)
I do have an account</p>
      </sec>
      <sec id="sec-4-5">
        <title>Q5: How often do you use such Web 2.0 services to share photographs</title>
        <p>with friends? 0) Never, 1) Less than once a month, 2) More than once a
month, 3) Weekly, 4) Daily</p>
      </sec>
      <sec id="sec-4-6">
        <title>Q6: Which of the following services do you use to upload and administrate holiday, family or other photographs? (Choose one or more.)</title>
        <p>None, Facebook, Flickr, Fotocommunity.de, Picasa, Other</p>
      </sec>
      <sec id="sec-4-7">
        <title>Q7: How often do you take photographs?</title>
        <p>0) Seldom, 1) Only at special events, 2) Often, 3) Virtually always
h
t 5 1 1 2 1 0 0 1 0 0 1 1 0 0 0 2 1 0 1 0 0 0 0 0 0 0 1 0 1 0 0 0 0 1 2 1 1 0 0 1 1 1 1 0 2 0 5
5
r Q
.
0
o 6
f 4 3 1 3 1 2 2 0 0 0 3 0 1 2 3 2 3 0 1 1 2 2 0 2 1 1 3 2 2 2 1 1 1 2 2 3 2 1 1 3 3 3 2 0 3 2 7
, Q
.
1
s
r 3 3 6 5 6 6 5 3 3 3 4 6 5 5 5 3 5 6 5 2 6 5 6 3 6 5 6 6 6 4 6 5 3 6 6 6 4 5 5 3 5 6 6 2 6 5 8
8
o Q
.
4
s
s 2 0 1 1 1 0 1 1 1 1 0 0 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 0 1 1 0 1 1 1 0 1 1 1 1 1 1 1 1 0 1 1 3
8
e Q 0
.
s
f
o
9 9 9 9 9 9 9 9 9 9 9 9 9 9 9 9 9 9 9 9 9 9 9 9 9 9 9 9 9 9 9 9 9 9 9 9 9 9 9 9 9 9 9 9 9 9
Y B 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 M
9
2
.
in xa ian an
e
l
a
m
e
F
IrsssseoD rsssseo11 rsssseo18 rsssseo24 rsssseo27 rsssseo48 rsssseo13 rsssseo28 rsssseo26 rsssseo25 rsssseo10 rsssseo12 rsssseo16 rsssseo17 rsssseo30 rsssseo21 rsssseo31 rsssseo33 rsssseo22 rsssseo23 rsssseo20 rsssseo51 rsssseo37 rsssseo36 rsssseo41 rsssseo42 rsssseo2 rsssseo4 rsssseo5 rsssseo3 rsssseo3 rsssseo2 rsssseo3 rsssseo4 rsssseo5 rsssseo1 rsssseo1 rsssseo3 rsssseo4 rsssseo4 rsssseo1 rsssseo4 rsssseo4 eM M
4 0 2 9 9 8 6 3 5 9 5 5 3 4 9 7 M M d e
A a a a a a a a a a a a a a a a a a a a a a a a a a a a a a a a a a a a a a a a a a a
in xa ed ea
f
1 1
o IR 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0 1 0 0 0 0 1 1 1 0 1 1 0 1 0 0 0 0 1 0 0 1 1 0 0 0 0 M M M M 3 1
ic M 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 1 1 1 1 0 0 1 0 0 0 0 1 0 0 0 1 1 0 0 0
raprseou lvee ..cS ..cS ..cS ..cS hD .S .S .S .S .S .S .S .S .S .S .S .S .S .S .S .S .S .S .S hD
.c .c .c .c .c .c .c .c .c .c .c .c .c .c .c . . . .</p>
        <p>c c c c
.c .c .c .c c c c c</p>
        <p>. . . .
.S .S .S S hD .S .S .S .S
.</p>
        <p>.
c
S
.
g C L M B B B P M M M B M B M B M M M M M M M B M M M P
M M M B P M M M M M
n
ia n</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Belkin</surname>
          </string-name>
          , N.:
          <article-title>Intelligent information retrieval: Whose intelligence?</article-title>
          <source>In: ISI '96: Proceedings of the Fifth International Symposium for Information Science</source>
          . pp.
          <volume>25</volume>
          {
          <issue>31</issue>
          (
          <year>1996</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2. Bottcher, T., Zellhofer,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Schmitt</surname>
          </string-name>
          ,
          <string-name>
            <surname>I.</surname>
          </string-name>
          :
          <article-title>BTU DBIS' Personal Photo Retrieval Runs at ImageCLEF 2013</article-title>
          . In: CLEF 2013 Labs and Workshop, Notebook Papers,
          <fpage>23</fpage>
          -
          <lpage>26</lpage>
          September 2013, Valencia, Spain (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3. Jarvelin,
          <string-name>
            <surname>K.</surname>
          </string-name>
          , Kekalainen, J.:
          <article-title>Cumulated gain-based evaluation of IR techniques</article-title>
          .
          <source>ACM Trans. Inf. Syst</source>
          .
          <volume>20</volume>
          (
          <issue>4</issue>
          ),
          <volume>422</volume>
          {
          <fpage>446</fpage>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Mizuochi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Higuchi</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kamada</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harada</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          : MIL at ImageCLEF 2013:
          <article-title>Personal Photo Retrieval</article-title>
          .
          <source>In: CLEF 2013 Labs and Workshop</source>
          , Notebook Papers,
          <fpage>23</fpage>
          -
          <lpage>26</lpage>
          September 2013, Valencia, Spain (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <source>s 1 0</source>
          <volume>0 0 0 0 1 1 2 3 0 0 0 0 0 1 1 1 1 2 0 0 2 3 3 3 4 4 0 1 1 1 1 2 3 1 1 2 3 4 1 3 3 0 4 1 0 4</volume>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <source>a Q 1 .</source>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          5.
          <string-name>
            <surname>Reiterer</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          , Mu ler, G.,
          <string-name>
            <surname>Mann</surname>
            ,
            <given-names>M.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Handschuh</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>INSYDER - an information assistant for business intelligence</article-title>
          .
          <source>In: Proceedings of the 23rd annual international ACM SIGIR conference on Research and development in information retrieval</source>
          . pp.
          <volume>112</volume>
          {
          <fpage>119</fpage>
          . SIGIR '00,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2000</year>
          ), http://doi.acm.
          <source>org/10</source>
          .1145/345508.345559
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          6. Zellhofer, D.:
          <article-title>An Extensible Personal Photograph Collection for Graded Relevance Assessments and User Simulation</article-title>
          .
          <source>In: Proceedings of the ACM International Conference on Multimedia Retrieval. ICMR '12</source>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          7. Zellhofer, D.:
          <article-title>Overview of the Personal Photo Retrieval Pilot Task at ImageCLEF 2012</article-title>
          . In: Forner,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Karlgren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Womser-Hacker</surname>
          </string-name>
          ,
          <string-name>
            <surname>C</surname>
          </string-name>
          . (eds.)
          <article-title>CLEF 2012 Evaluation Labs</article-title>
          and Workshop (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>