<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>DCU-TCD@LogCLEF 2010: Re-ranking Document Collections and Query Performance Estimation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Johannes Leveling</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>M. Rami Ghorab</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Walid Magdy</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gareth J. F. Jones</string-name>
          <email>gjonesg@computing.dcu.ie</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vincent Wade</string-name>
          <email>vincent.wadeg@scss.tcd.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Centre for Next Generation Localisation (CNGL), Knowledge and Data Engineering Group, Trinity College Dublin</institution>
          ,
          <addr-line>Dublin 2</addr-line>
          ,
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Centre for Next Generation Localisation (CNGL), School of Computing, Dublin City University</institution>
          ,
          <addr-line>Dublin 9</addr-line>
          ,
          <country country="IE">Ireland</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes the collaborative participation of Dublin City University and Trinity College Dublin in LogCLEF 2010. Two sets of experiments were conducted. First, di erent aspects of the TEL query logs were analysed after extracting user sessions of consecutive queries on a topic. The relation between the queries and their length (number of terms) and position ( rst query or further reformulations) was examined in a session with respect to query performance estimators such as query scope, IDF-based measures, simpli ed query clarity score, and average inverse document collection frequency. Results of this analysis suggest that only some estimator values show a correlation with query length or position in the TEL logs (e.g. similarity score between collection and query). Second, the relation between three attributes was investigated: the user's country (detected from IP address), the query language, and the interface language. The investigation aimed to explore the in uence of the three attributes on the user's collection selection. Moreover, the investigation involved assigning di erent weights to the three attributes in a scoring function that was used to re-rank the collections displayed to the user according to the language and country. The results of the collection re-ranking show a signi cant improvement in Mean Average Precision (MAP) over the original collection ranking of TEL. The results also indicate that the query language and interface language have more in uence than the user's country on the collections selected by the users.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>LogCLEF at the Conference on Multilingual and Multimodal Information Access
Evaluation 2010 is an initiative to analyse search logs and discover patterns
of multilingual search behaviour. The logs come from user interactions with
the European Library (TEL)3, which is a portal forming a single interface for
searching across the content of many European national libraries.</p>
      <p>
        LogCLEF started as a task of the CLEF 2009 evaluation campaign [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ],
in which Trinity College Dublin and Dublin City University participated
collaboratively. Our previous participation focused on investigating users' query
reformulations and their mental model of the search engine as well as analysing
user behaviour for di erent linguistic backgrounds [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>In our LogCLEF 2010 participation, two sets of experiments were carried out
on action logs from TEL. The rst set of experiments investigated the assumption
that query performance predictors not only can be used to estimate performance,
but can also be utilised to analyse real user behaviour. For example, users' query
modi cations generally aim at improving the results and should re ect the fact
that the users' queries improve over successive reformulations. This should also
result in di erent values for performance estimates.</p>
      <p>
        The second set of experiments were concerned with studying three attributes
related to the users and the queries they submit, namely: the user's country (i.e.
the location from which the query was submitted), the language of the
submitted query, and the speci ed interface language by the user. The experiments
investigated the in uence of the three attributes on the collections (libraries or
online resources) selected by the user. In other words, the experiments aimed
at answering the question of whether or not the users' choice of collections is
in uenced by their location and language. The experiments involved re-ranking
the list of collections displayed to the user. Typically, result re-ranking involves
re-ranking a list of documents. However, in this study, the re-ranking process
was applied to the list of collections displayed to the user by TEL. A scoring
function that comprises the three attributes was used to determine the relevance
of each collection to the current search in terms of country and language.
Reranking alternatives were investigated by assigning di erent weights to the three
attributes in the scoring function. The results of the collection re-ranking
experiments showed a 27.4% improvement in Mean Average Precision (MAP) [
        <xref ref-type="bibr" rid="ref4 ref5">4,
5</xref>
        ] over the original list of collections displayed by TEL. The results also suggest
that the query language and interface language have more in uence than the
user's country on the collections selected by the users.
      </p>
      <p>The rest of this paper is organised as follows: Section 2 introduces related
work; Section 3 describes preprocessing of the log les for session reconstruction;
Section 4 introduces query performance predictors and details the analysis
carried out with these measures on the query logs; Section 5 describes the collection
re-ranking experiments and their results; and nally, conclusions and outlook on
future work are given in Section 6.</p>
    </sec>
    <sec id="sec-2">
      <title>3 http://www.theeuropeanlibrary.org/</title>
      <sec id="sec-2-1">
        <title>Related Work</title>
        <sec id="sec-2-1-1">
          <title>Query Performance Estimation</title>
          <p>
            Query performance estimators have been widely used to improve retrieval e
ectiveness by predicting query performance [6{9], or to reduce long queries [
            <xref ref-type="bibr" rid="ref10">10</xref>
            ].
Hau [
            <xref ref-type="bibr" rid="ref11">11</xref>
            ] gives a comprehensive overview over predicting performance of queries
and retrieval systems.
          </p>
          <p>Typically, these estimators are applied to user queries for automatic query
modi cation, either before retrieval (i.e. pre-retrieval), based on information
available at indexing time, or after an initial retrieval (i.e. post-retrieval), in
which case additional information such as relevance scores for the documents
can be employed as features for the performance estimators. The main
objective is to improve e ectiveness of information retrieval systems in general or
for speci c topics. However, these estimators have { to the best of the authors'
knowledge { not yet been applied to real user queries to investigate if real query
reformulations actually improve the search results.
2.2</p>
        </sec>
        <sec id="sec-2-1-2">
          <title>Result Re-ranking</title>
          <p>Result re-ranking is one of the well-known techniques used for search
personalisation. [12{14]. Result re-ranking adapts to the user's needs in that it brings
results that are more relevant to him/her to higher positions in the result list.
It takes place after an initial set of results have been retrieved by the system,
where an additional ranking phase is performed to re-order the results based on
various adaptation aspects (e.g. user's language, knowledge, or interests).</p>
          <p>Typically, the re-ranking process is applied to the list of retrieved documents.
However this study investigates applying the re-ranking process to the list of
collections displayed to the user, not the list of documents within a collection.
3</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>Data Preparation</title>
        <p>Two datasets (action log les) were provided at LogCLEF: the rst one (L1)
contained action logs (queries and user interactions) from January 2007 to June
2008 from the TEL web portal (approximately 1.8 million records), and the
second one (L2) contained action logs from January 2009 to December 2009
(approximately 76 thousand records). In addition, a third dataset (HTTP logs)
was provided, which contained more details of the search sessions and HTTP
connections, including the list of collections searched for each submitted query.</p>
        <p>
          For the rst set of experiments, queries from the two datasets of action logs
were used. The logs were pre-processed following the approach in our previous
experiments [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] where user sessions were reconstructed and consecutive queries
on the same topic were determined. Session reconstruction was done by
structuring all actions with the same session ID together and sorting them by timestamp.
If the time between consecutive actions exceeded 30 minutes, the start of a new
session was assumed. In addition, only English queries and sessions containing at
least three queries were considered. The resulting dataset contained 340 sessions
from L1 with an average session length of 3.61 queries per session. The queries
in this set consisted on average of 3.3 terms. Preprocessing of L2 resulted in 487
sessions with an average session length of 3.61 queries per session. Queries in
this set consisted of 4.16 terms on average.
        </p>
        <p>
          Preliminary experiments on the TEL data showed that only a few URLs
from the logs were functional. Retrieving result pages for queries via their URL
results in error codes of "not found" (HTTP response code 404). Furthermore,
referrer URLs seem to be truncated to a certain number of characters, yielding
malformed URLs. Thus, an external resource with reasonable coverage was used
to compute query statistics (e.g. term frequencies and collection frequencies).
The DBpedia4 collection of Wikipedia article abstracts was indexed and the
queries were evaluated against this index. The Lucene toolkit5 was employed to
index the document collection, using the BM25 model [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] for retrieval (although
no results were retrieved for the experiments, because only pre-retrieval query
performance estimators have been used).
        </p>
        <p>For the second set of experiments, data from both the action logs and the
HTTP logs of TEL 2007 was used. Due to time limitation and some technical
errors in the recorded HTTP logs, the experiments were only conducted on a
subset of the data. The selected subset comprised the submitted queries and
clicked collections during the month of February 2007. This included
approximately 1,800 queries from di erent languages.
4
4.1</p>
      </sec>
      <sec id="sec-2-3">
        <title>Experiments on Estimates of Query Performance</title>
        <sec id="sec-2-3-1">
          <title>Query Performance Estimators</title>
          <p>
            Several pre-retrieval performance predictors were employed. Pre-retrieval means
that the computation does not rely on calculating relevance scores or other
postretrieval information. Performance predictors are described in more detail in [
            <xref ref-type="bibr" rid="ref6">6</xref>
            ].
The following is the list of predictors used in our experiments.
          </p>
          <p>
            Query length. For our experiments, the query length is de ned as the number of
unique query terms after stopword removal (see Equation 1). Zhai and La erty
[
            <xref ref-type="bibr" rid="ref16">16</xref>
            ] have observed that the query length has a strong impact on the
smoothing method for language modeling. Similarly, He and Ounis [
            <xref ref-type="bibr" rid="ref17">17</xref>
            ] nd that query
length strongly a ects normalisation methods for probabilistic retrieval
models. Intuitively, longer queries should be less ambiguous because they contain
additional query terms which provide context for disambiguation.
          </p>
          <p>QL = number of unique non-stopword terms
(1)
where q is the query.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>4 http://dbpedia.org/</title>
    </sec>
    <sec id="sec-4">
      <title>5 http://lucene.apache.org/</title>
      <p>
        Query scope (QS). The query scope tries to measure the generality of a query
by its result set size [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], i.e. by the cardinality of the set of documents
containing at least one query term (see Equation 2). A query matching only very few
documents can be considered very speci c, while a query containing only high
frequency terms will return a larger portion of the document collection.
      </p>
      <p>QS =</p>
      <p>
        log(nq=N )
where nq is result set size and N is total number of documents in the collection.
IDF based score (IDFmm). IDF mm denotes the distribution of informative
amount of terms t in a query. This IDF -based feature is introduced as 2 by
He and Ounis [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. We use INQUERY's IDF formula (Equation 3)
idf (t) =
log2(N + 0:5)=Nt
      </p>
      <p>log2(N + 1)
IDF mm =
idfmax
idfmin
where t is a term and Nt is the number of documents containing t. In the IDF mm
formula (see Equation 4), idfmax and idfmin are the maximum and minimum
IDF of terms from the query q.</p>
      <p>SCQS =</p>
      <p>X(1 + ln(tfcoll)) ln(1 +
t2q</p>
      <p>N
Nt
)
where Nt is the number of documents containing term t.
(2)
(3)
(4)
(5)
(6)
Simpli ed query clarity score (SCS). The query clarity is inversely proportional
to the ambiguity of the query (i.e. clarity and ambiguity are contrary concepts).</p>
      <p>
        This score measures ambiguity based on the analysis of the coherence of
language usage in documents whose models are likely to generate the query
[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Because the computation of the original query clarity score relies on the
availability of relevance scores, its computation would be time-consuming and
result in post-retrieval estimates. We employ the de nition of a simpler version
proposed by He and Ounis [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], the simpli ed query clarity score (see Equation 5).
      </p>
      <p>SCS = Pt2q Pml(tjq) log2 PPmcoll(lt(jtq))</p>
      <p>
        ' Pt2q qqtlf log2 tfcollq=ttfo=kqelncoll
where Pml(tjq) is the maximum likelihood of the query model of term t in query
q and Pcoll(t) is the collection model. tfcoll is the collection frequency, tokencoll
is the number of tokens in the whole collection, qtf is the term frequency in the
query, and ql is the query length (the number of non-stopwords in the query).
Similarity score between collection and query (SCQS). Zhao, Scholer, et al. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]
compute a similarity score between the query and the document collection to
estimate pre-retrieval query performance (see Equation 6). For our experiments,
we use the sum of the contributions of all individual query terms. As pointed
out in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], this metric will be biased towards longer queries.
      </p>
      <p>
        Average inverse document collection term frequency (avICT F ). Kwok [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]
introduced the inverse collection term frequency (ICT F ) as an alternative to IDF .
As ICT F is highly correlated with search term quality, He and Ounis [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]
proposed using the average ICT F (avICT F ) to predict query performance.
avICT F =
log2(Qt2q totkfecnoclloll )
ql
(7)
      </p>
      <p>The denominator ql is the reciprocal of the maximum likelihood of the query
model of SCS in Equation 5. The use of avICT F is similar to measuring the
divergence of a collection model (i.e. ICT F ) from a query model. Thus, avICT F
and SCS should have similar query performance estimates.
4.2</p>
      <sec id="sec-4-1">
        <title>Experiments and Results</title>
        <p>For the rst part of our experiments on the TEL log data, the relation between
real user queries and the query performance estimators are investigated. All
searches from the reconstructed sessions from L1 and L2 were conducted on the
indexed DBpedia abstract collection and the average values for all query
performance estimators were calculated. Due to data sparsity, some results were
grouped by query length (number of terms in a query) and by position (whether
a query is the rst one submitted in a session or is a further reformulation) into
bins. Moreover, session estimates were analysed by distinguishing between
sessions beginning with a short query and those beginning with a long query. This
was performed with the aim of investigating if the rst query in a session
indicates the search behaviour during that session. For example, users may expand
short queries and reduce long queries if their initial search is not successful.</p>
        <p>
          For the experiments, short queries were de ned as queries containing one to
three terms and long queries were de ned as queries with more than three terms.
This value roughly corresponds to the well-known average number of terms in
web queries (2-3 terms) [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] and to the average number of terms in the TEL logs
(3-4 terms) [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. Figures 1-4 show results from the analyses (L1 data on left, L2
data on right in Figures 1-3). Note that there is a di erence between QL, which
counts the number of unique terms in a query, and the query length in tokens
(including duplicates), which is shown on the x-axis in Figure 1.
        </p>
        <p>Comparing all six query performance estimators investigated (i.e. QL, QS,
IDF mm, SCS, avICT F , and SCQS) to the query length, we nd that,
unsurprisingly, there is a high correlation between the query length (non-stopword
terms) and QL (unique non-stopword terms). Query scope (QS) increases slightly
for queries containing more than 3 terms, but remains almost constant for longer
queries. IDF mm slightly increases for longer queries. SCS has its maximum
for the shortest queries (length 1-3 terms) decreases for medium queries, and
increases again for most longer queries. Finally, avICT F also decreases for
medium-length queries, with higher values for short and long queries (see
Figure 1). All estimators show similar behaviour on the sessions of L1 and L2.</p>
        <p>Next, the estimator values were calculated and compared across di erent
positions in the sessions (see Figure 2). The positions correspond to users' query
reformulations. The sessions from L1 are longer than sessions from L2 (14 queries
maximum for L1 vs. 10 queries maximum for L2). With a few exceptions, most
estimators have almost constant values for the di erent reformulations. One of
the exceptions is avICT F , which drops after position 3 and has a sudden, but
constant maximum after position 12. However, SCS increases after position 7 in
both data sets. SCQS compared to the query position shows that with a higher
number of queries in a session (i.e. more query reformulations), SCQS drops
consistently (see Figure 4 left).</p>
        <p>Finally, dividing the data into sessions starting with a short query (less than
4 terms) and sessions starting with longer queries reveals the most interesting
results. QL, QS, and IDF mm are higher for longer queries in both data sets
investigated, while SCS and avICT F are higher for shorter queries (see
Figure 3). The highest di erence can be observed for SCQS (see Figure 4 right)
which shows much lower values for sessions with initial long queries.
4.3</p>
      </sec>
      <sec id="sec-4-2">
        <title>Discussion</title>
        <p>Most of the query performance predictors we experimented with the aim of
measuring clarity, speci city, or ambiguity of query terms. For example, query length
is presumed to be an indicator of result quality because the more terms a query
contains, the more speci c the query is, i.e. the more precise the results should
be. However, our experiments did not show that longer queries are consistently
less ambiguous, or more speci c (Figure 1).</p>
        <p>Our next analysis (Figure 2) was concerned with query reformulations in
a single session. We investigated whether queries become more successful after
reformulation, i.e. if a user's query reformulation will achieve better performance
(estimates). We also have not observed a simple or direct relationship between
the query position and the performance estimates. Instead, the extracted user
sessions might contain data from users with di erent search strategies or data
from users changing their search strategy in the middle of a session. This may be
an explanation for sudden drops and increases in some performance estimators
(see, for instance avICT F in Figure 2 on the L1 data on the left).</p>
        <p>Finally, SCQS showed the most interesting behaviour. For longer queries, its
value decreases consistently for both data sets. With higher query positions, its
value typically decreases, and the di erences for sessions beginning with short
and long queries are huge.</p>
        <p>A small sample of sessions was manually investigated in more detail. Several
users reformulate their queries by using the same search terms in di erent elds
(for the advanced eld-based search) or to search in di erent indexes (e.g. "lima
licinio" repeated for di erent elds). This type of modi cation might not actually
improve the retrieved results. However, the current TEL portal already supports
a catch-all eld ("any eld"), i.e. a eld covering a search in all indexes. Possibly
this eld was not available at the time the query logs were collected.</p>
        <p>Some users modify their original query, examine the results, and submit
the original query again (e.g. "environmental comunication" is modi ed into
"comunication" which is modi ed back to "environmental comunication" { note
the spelling error in all queries). A similar behaviour can be observed for longer
sessions. After several queries, the users seem to revert to previously submitted
queries, probably because the TEL portal does not provide query suggestions
and the users do not nd relevant or new information allowing them to change
their query formulation.</p>
        <p>Finally, some queries contain more than one change at once (e.g. \music
manuscripts anton bruckner" is changed to \bruckner anton sinfonie"). Thus,
changes in performance predictors might also incorporate several changes (i.e.
increases and decreases in quality of results) at once.
5
5.1</p>
        <sec id="sec-4-2-1">
          <title>Re-ranking the List of Collections</title>
        </sec>
      </sec>
      <sec id="sec-4-3">
        <title>Overview</title>
        <p>When users submit a search to the TEL portal, they are presented with a list
of collections on the left side of the screen and a list of results from the selected
collection on the right side of the screen. A collection is either a library or an
online resource that is associated with a certain country. The list of collections
is presented to the user in alphabetical order of the countries' acronyms (two
letters according to the ISO 3166-1 standard).</p>
        <p>The users' queries and interactions with the portal were recorded in two
datasets called: action logs and HTTP logs. This set of experiments involved
studying di erent aspects regarding three attributes exhibited in the logs: user's
country (location from which the query was submitted), query language, and
interface language. In our experiments, the user's country was determined from
the IP address recorded in the logs. The query language was detected using the
Google AJAX Language API6, which returns a con dence level associated with</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>6 http://code.google.com/apis/ajaxlanguage/</title>
      <p>the detected language of a query. Only queries that had a minimum value of
10% con dence level were used in the experiments, which reduced the number
of queries subject to experimentation to 566 queries (from the 1,800 queries of
February 2007).</p>
      <p>Furthermore, the experiments involved studying the list of collections that
were searched for each query, and the collections that the user clicked on. The
study investigated re-ranking the list of collections displayed to the user with the
aim of increasing the retrieval precision (across the collections, not the actual
result documents) by bringing collections that match the user's country or
language to higher ranks in the list. In other words, the re-ranking process focused
on studying if users have a higher tendency to click on collections that belong
to their country or language. For this part of the experiments, languages were
assigned to a collection based on the o cial languages that are spoken in the
country associated with that collection.
5.2</p>
      <sec id="sec-5-1">
        <title>Descriptive Statistics</title>
        <p>The following statistics are drawn from the selected subset of the logs. Figure 5
shows the distribution of users' countries on the left (i.e. countries that the
queries came from) and the language distribution of queries on the right. It is
noted that a large percentage (50%) of the queries were in English although
the total percentage of queries coming from English-speaking countries (United
States, United Kingdom, and Canada) was only 16%. This suggests that many
users from non-English countries do not submit queries in their native language.</p>
        <p>
          Figure 6 shows the relation between user's country, query language, and
interface language for all queries (left) and for non-English queries (right).
Surprisingly, it is noted that only 24% of the queries coming from a country are
in languages associated with that country. This is not necessarily an inclination
towards using English in search, as similar results were observed when
studying non-English queries, where less than 30% of the queries were submitted in
a language that matched the country. Therefore, this suggests that the
country attribute may not be very reliable to base the collection re-ranking decision
upon. This was con rmed by further experiments (discussed below).
In order to evaluate the retrieval precision over the list of collections presented
to the user, we used the collections that the user clicked on as implicit relevance
judgements (i.e. binary relevance judgements where the clicked collections are
assumed to be the relevant ones, and non-clicked collections are assumed to be
irrelevant). This follows on the method adopted in [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. However, the sense of
relevance in our experiments is in terms of matching country and language.
        </p>
        <p>
          Mean Average Precision (MAP) was used for evaluation as it is a common
evaluation metric in information retrieval that rewards relevant items being
ranked higher in the list [
          <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
          ]. MAP was calculated across the queries in the
selected subset of the data. The original ranked list of collections (i.e the one
presented to the user by TEL) was used as the baseline for evaluation of
retrieval precision (MAP score = 0.580). Several alternative re-ranked lists were
investigated and compared to the baseline.
        </p>
        <p>A result (collection) scoring function was used to re-rank the list of
collections. The function was based on matching the three attributes with the
collection's country and language as follows (where matching=1 and non-matching=0):
1. Mc: matching the user's country with the collection's country.
2. Mq: matching the query's language with the collection's language (i.e.
matching with any of the o cial languages spoken in the corresponding country).
3. Mi: matching the interface language with the collection's language.</p>
        <p>Each of the above attributes is multiplied by a scalar weight (Wc, Wq, Wi
respectively) so as to control (and test) the degree of contribution of each attribute
in the function. Thus the collection scoring function becomes Equation 8
N ewScore = (Wc</p>
        <p>Mc) + (Wq</p>
        <p>Mq) + (Wi Mi)
(8)
Finally, the collections are re-ranked based on descending order of the new score.</p>
        <p>Unlike other collections, collection #a0000 had no country associated with
it (a virtual collection of various European digitised books, images, maps, etc.).
This would have caused it to be permanently placed at the end of the re-ranked
list (i.e. get a zero score from the scoring function) because of not matching any
country or language. Therefore, an exception was made regarding this collection;
it was placed at the top of the re-ranked list (similar to TEL's ranking).
Table 1 shows the MAP value for some selected re-ranking runs with
alternative combinations of weights that ranged from 0.0 to 1.0. The results showed a
signi cant improvement of 27.4% in retrieval precision for the re-ranked
collection lists (with weights: Wc=0.1, Wq=0.3, Wi=0.6) over the baseline ranking.
This improvement is statistically signi cant as per the T-test (with p=0.01) and
Wilcoxon test (with con dence=99%). The results also suggest that there is a
relation between the query language and interface language on the one hand and
the clicked collections on the other hand. Moreover, these two attributes seem
to have more in uence than the user's country (location from which the query
was submitted) on the selected collections by the users.</p>
        <sec id="sec-5-1-1">
          <title>Conclusions and Future Work</title>
          <p>The experiments using query performance predictors show that SCQS has a
high correlation with query length and query position. Distinguishing between
sessions with an initial short or long query seems to indicate a di erent
subsequent user behaviour. For example, users starting with a short query are more
likely to expand their query to obtain more speci c results, while users starting
with a longer query probably reduce their query because the retrieved result set
is too small. To obtain more conclusive results, we want to perform more
experiments on larger query logs. The results might also di er from results obtained
from performing the same analysis on web search logs, because of the di erent
application domain (bibliographic search in the TEL portal), the slightly longer
queries and possibly longer sessions, which include more query reformulations.
Query performance estimators have been developed for web search and may thus
be optimised for shorter queries and fewer search iterations. In conclusion, more
consistent results may be found if the methods are applied on a larger data set.</p>
          <p>The experiments of collection re-ranking based on matching the query
language and the interface language with the collection's language showed a
signi cant improvement in precision. This suggests that there is opportunity for
improving the user's experience with multilingual search in the TEL portal if
the user's language is taken into consideration. Future work will involve
using machine learning techniques to learn a ranking function that optimises the
weights of the three attributes: country, query language, and interface language.
Moreover, we will attempt to work around the technical errors present in the
dataset of TEL HTTP logs, which will enable us to conduct the experiments on
the full dataset.</p>
        </sec>
      </sec>
      <sec id="sec-5-2">
        <title>Acknowledgements</title>
        <p>This research is supported by the Science Foundation of Ireland (grant 07/CE/I1142)
as part of the Centre for Next Generation Localisation (www.cngl.ie) at Dublin
City University and Trinity College Dublin.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Mandl</surname>
          </string-name>
          , T., di Nunzio, G.:
          <article-title>Overview of the LogCLEF track</article-title>
          . In Borri, F.,
          <string-name>
            <surname>Nardi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peters</surname>
          </string-name>
          , C., eds.:
          <article-title>Results of the CLEF 2009 Cross-Language System Evaluation Campaign</article-title>
          ,
          <source>Working Notes of the CLEF 2009 Workshop</source>
          . (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Mandl</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Agosti</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Di</given-names>
            <surname>Nunzio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            ,
            <surname>Yeh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Mani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            ,
            <surname>Doran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Schulz</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.M.:</surname>
          </string-name>
          <article-title>LogCLEF 2009: the CLEF 2009 Cross-Language Log le Analysis Track Overview</article-title>
          . In Peters,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Di Nunzio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            ,
            <surname>Kurimo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Mandl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Mostefa</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          , Pen~as,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Roda</surname>
          </string-name>
          , G., eds.
          <source>: Multilingual Information Access Evaluation Vol. I Text Retrieval Experiments: Proceedings 10th Workshop of the Cross-Language Evaluation Forum</source>
          ,
          <string-name>
            <surname>CLEF</surname>
          </string-name>
          <year>2009</year>
          , Corfu, Greece.
          <source>Revised Selected Papers. LNCS</source>
          , Springer (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Ghorab</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leveling</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wade</surname>
          </string-name>
          , V.:
          <article-title>TCD-DCU at LogCLEF 2009: An analysis of queries, actions, and interface languages</article-title>
          . In
          <string-name>
            <surname>Borri</surname>
          </string-name>
          , F.,
          <string-name>
            <surname>Nardi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peters</surname>
          </string-name>
          , C., eds.:
          <article-title>Results of the CLEF 2009 Cross-Language System Evaluation Campaign</article-title>
          ,
          <source>Working Notes of the CLEF 2009 Workshop</source>
          . (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Baeza-Yates</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ribeiro-Neto</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Modern Information Retrieval. 1st edn</article-title>
          . (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Manning</surname>
            ,
            <given-names>C.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Raghavan</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schutze</surname>
          </string-name>
          , H.:
          <article-title>Introduction to Information Retrieval. 1st edn</article-title>
          . Cambridge University Press (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>He</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ounis</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Inferring query performance using pre-retrieval predictors</article-title>
          . In Apostolico,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Melucci</surname>
          </string-name>
          , M., eds.
          <source>: String Processing and Information Retrieval</source>
          , 11th International Conference, SPIRE 2004, Padova, Italy, October 5-
          <issue>8</issue>
          ,
          <year>2004</year>
          , Proceedings. Volume
          <volume>3246</volume>
          of Lecture Notes in Computer Science., Springer (
          <year>2004</year>
          )
          <volume>43</volume>
          {
          <fpage>54</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Cronen-Townsend</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Croft</surname>
            ,
            <given-names>B.W.</given-names>
          </string-name>
          :
          <article-title>Predicting query performance</article-title>
          .
          <source>In: SIGIR 2002: Proceedings of the 25th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, August 11-15</source>
          ,
          <year>2002</year>
          , Tampere, Finland,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2002</year>
          )
          <volume>299</volume>
          {
          <fpage>306</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Scholer</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tsegay</surname>
          </string-name>
          , Y.:
          <article-title>E ective pre-retrieval query performance prediction using similarity and variability evidence</article-title>
          . In Macdonald,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Ounis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            ,
            <surname>Plachouras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            ,
            <surname>Ruthven</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            ,
            <surname>White</surname>
          </string-name>
          , R.W., eds.
          <source>: Advances in Information Retrieval, 30th European Conference on IR Research</source>
          , ECIR
          <year>2008</year>
          , Glasgow, UK, March 30- April 3,
          <year>2008</year>
          . Proceedings. Volume
          <volume>4956</volume>
          of Lecture Notes in Computer Science., Springer (
          <year>2008</year>
          )
          <volume>52</volume>
          {
          <fpage>64</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Plachouras</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ounis</surname>
            , I., van Rijsbergen,
            <given-names>C.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cacheda</surname>
          </string-name>
          , F.: University of Glasgow at the Web Track:
          <article-title>Dynamic application of hyperlink analysis using the query scope</article-title>
          .
          <source>In: Proceedings of TREC</source>
          <year>2003</year>
          ,
          <article-title>Gaithersburg</article-title>
          ,
          <string-name>
            <surname>MD</surname>
          </string-name>
          , NIST (
          <year>2003</year>
          )
          <volume>646</volume>
          {
          <fpage>652</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Kumaran</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carvalho</surname>
            ,
            <given-names>V.R.</given-names>
          </string-name>
          :
          <article-title>Reducing long queries using query quality predictors</article-title>
          . In Allan, J.,
          <string-name>
            <surname>Aslam</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sanderson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhai</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zobel</surname>
          </string-name>
          , J., eds.
          <source>: Proceedings of the 32nd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          ,
          <string-name>
            <surname>SIGIR</surname>
          </string-name>
          <year>2009</year>
          , Boston, MA, USA, July
          <volume>19</volume>
          -
          <issue>23</issue>
          ,
          <year>2009</year>
          , ACM (
          <year>2009</year>
          )
          <volume>564</volume>
          {
          <fpage>571</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Hau</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Predicting the E ectiveness of Queries and Retrieval Systems</article-title>
          .
          <source>PhD thesis</source>
          , University of Twente.
          <source>CTIT</source>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Speretta</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gauch</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Personalized search based on user search histories</article-title>
          . In: IEEE/WIC/ACM International Conference on Web Intelligence. (
          <year>2005</year>
          )
          <volume>622</volume>
          {
          <fpage>628</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Pretschner</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gauch</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Ontology based personalized search</article-title>
          .
          <source>In: 11th IEEE International Conference on Tools with Arti cial Intelligence</source>
          .
          <article-title>(</article-title>
          <year>1999</year>
          )
          <volume>391</volume>
          {
          <fpage>398</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Agichtein</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brill</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dumais</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Improving web search ranking by incorporating user behavior information</article-title>
          .
          <source>In: 29th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR</source>
          <year>2006</year>
          ).
          <article-title>(</article-title>
          <year>2006</year>
          )
          <volume>19</volume>
          {
          <fpage>26</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Robertson</surname>
            ,
            <given-names>S.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Walker</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hancock-Beaulieu</surname>
            ,
            <given-names>M.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gatford</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Okapi at TREC-3</article-title>
          . In Harman, D.K., ed.:
          <source>Overview of the Third Text Retrieval Conference (TREC-3)</source>
          , Gaithersburg,
          <string-name>
            <surname>MD</surname>
          </string-name>
          , USA,
          <source>National Institute of Standards and Technology (NIST)</source>
          (
          <year>1995</year>
          )
          <volume>109</volume>
          {
          <fpage>126</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Zhai</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>La</surname>
            <given-names>erty</given-names>
          </string-name>
          , J.:
          <article-title>A study of smoothing methods for language models applied to information retrieval</article-title>
          .
          <source>ACM Transactions on Information Systems</source>
          <volume>22</volume>
          (
          <issue>2</issue>
          ) (
          <year>2004</year>
          )
          <volume>179</volume>
          {
          <fpage>214</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>He</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ounis</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>A study of parameter tuning for term frequency normalization</article-title>
          .
          <source>In: Proceedings of the 2003 ACM CIKM International Conference on Information and Knowledge Management</source>
          , New Orleans, Louisiana, USA, November 2-
          <issue>8</issue>
          ,
          <year>2003</year>
          , ACM (
          <year>2003</year>
          )
          <volume>10</volume>
          {
          <fpage>16</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>He</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ounis</surname>
            ,
            <given-names>I.:</given-names>
          </string-name>
          <article-title>Query performance prediction</article-title>
          .
          <source>Information Systems</source>
          <volume>31</volume>
          (
          <issue>7</issue>
          ) (
          <year>2006</year>
          )
          <volume>585</volume>
          {
          <fpage>594</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Kwok</surname>
            ,
            <given-names>K.L.:</given-names>
          </string-name>
          <article-title>A new method of weighting query terms for ad-hoc retrieval</article-title>
          .
          <source>In: SIGIR '96: Proceedings of the 19th annual international ACM SIGIR conference on Research and development in information retrieval</source>
          , New York, NY, USA, ACM (
          <year>1996</year>
          )
          <volume>187</volume>
          {
          <fpage>195</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Jansen</surname>
            ,
            <given-names>B.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spink</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>An analysis of web searching by european alltheweb</article-title>
          .
          <source>com users. Information Processing &amp; Management</source>
          <volume>41</volume>
          (
          <issue>2</issue>
          ) (
          <year>2005</year>
          )
          <volume>361</volume>
          {
          <fpage>381</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>