<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Annotating URLs with query terms: What factors predict reliable annotations?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Suzan Verberne</string-name>
          <email>s.verberne@let.ru.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Eva D'hondt</string-name>
          <email>e.dhondt@let.ru.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Max Hinne</string-name>
          <email>mhinne@sci.ru.nl</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Wessel Kraaij</string-name>
          <email>kraaijw@acm.org</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maarten van der Heijden</string-name>
          <email>m.vanderheijden@cs.ru.nl</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Theo van der Weide</string-name>
          <email>tvdw@cs.ru.nl</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CLST, Radboud University Nijmegen</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Dept. Computer Science, Radboud University Nijmegen</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Dept. Computer Science, Radboud University Nijmegen</institution>
          ,
          <addr-line>TNO, Delft</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2009</year>
      </pub-date>
      <fpage>19</fpage>
      <lpage>23</lpage>
      <abstract>
        <p>A number of recent studies have investigated the relation between URLs and associated query terms from search engine log les. In [5], the query terms associated with the domain of a URL were used as features for a URL classi cation task. The idea is that query terms that lead to successful classication of a URL are reliable semantic descriptors of the URL content. We follow up on this work by investigating which properties of a URL and its associated query terms predict the classi cation success. We construct a number of URL and query properties as predictors and proceed to analyze these in-depth. We conclude that the classi cation success | and thus the reliability of the query terms as URL descriptors | cannot easily be predicted from properties of the URL and the queries.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Categories and Subject Descriptors</title>
      <p>H.3.3 [Information Search and Retrieval]: Information
ltering</p>
      <sec id="sec-1-1">
        <title>Click data, URL classi cation, Human factors</title>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>1. INTRODUCTION</title>
      <p>
        In previous work on the use of query log data [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], the
authors investigated the applicability of semantic
annotation of web pages by creating short document descriptions
(term lists) extracted from associated queries. The
assumption here is that, when presented to a user, these term lists
may help in the disambiguation of a URL and/or identify
whether the URL corresponds to the user's query intent [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
The term lists in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] were extracted from the (weighted) set
of query terms that are associated with an URL. In order
to nd out whether the associated term lists provided good
descriptions of the URL (and consequently, good clues in
disambiguation), a classi cation experiment was conducted
on a set of URLs, using the term lists as features.
Depending on the level of query term aggregation, a classi cation
accuracy of up to 45% was obtained [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        These classi cation results are reasonably satisfying. We
aim at a future implementation of semantic annotation of
URLs in the user interface of a web search engine. In this
implementation, the search engine not only has the query
term descriptors for a URL available, but also the indexed
content of the web page. Previous work on semantic
annotations from URL content and query logs [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] showed that the
query terms associated to a URL provide useful additional
information to terms extracted from the content of the web
page. Thus we can expect the classi cation accuracy to go
up if we not only use the query terms but also the content
terms.
      </p>
      <p>
        However, a classi cation accuracy of 45% means that URL
classi cation based on query terms only is unsuccessful for
more than half of the URLs. In order to prevent the
semantic annotation of URLs to be negatively in uenced by
unreliable query term descriptors, we aim to predict the
reliability of the query term annotations given a URL and
its associated queries. Therefore, in the current paper, we
study which properties of the URL and the associated term
list can predict the reliability of the query term annotations.
Following [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], we consider the classi cation accuracy as
indicator for the informativeness and reliability of the query
terms annotation: the better a set of query terms describes
a URL, the higher the chance that the URL is classi ed
correctly. Thus, we investigate the relation between URL
and query properties on the one hand and the classi cation
quality on the other hand.
2.
      </p>
    </sec>
    <sec id="sec-3">
      <title>RELATED WORK</title>
      <p>
        Other studies have shown that query{URL associations
and click information can serve as a means for implicit
feedback [
        <xref ref-type="bibr" rid="ref4 ref6 ref7">4, 6, 7</xref>
        ] and for learning to rank e.g.[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Several others
have investigated whether a collection of queries and click
information can be used as a model of the semantic
contents of a web page [
        <xref ref-type="bibr" rid="ref2 ref8">2, 8</xref>
        ]. A document representation based
on queries has been compared with a traditional document
content vector space representation for a clustering task [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
The query based representation resulted in a better
clustering than the document content based representation. A
similar experiment has been conducted where human assessors
were asked whether they preferred a document description
based on queries or on a vector space representation [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
The assessors tended to prefer the query-based
representation.
      </p>
      <p>
        Most related work uses site access logs of a portal page,
which means that the log les show a complete picture of all
queries leading to that site: query terms can be derived from
the referring URL in the HTTP-header. In our case, we use
a much larger collection of web pages and click data that
is based on the query log of one single search engine (See
Section 3.1). Our data do not comprise the page content of
the URLs because the page content was not available in the
query log data and recrawling would give many
inconsistencies due to web pages changing signi cantly in the course of
a few years. Our work di ers from [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] since we explicitly aim
to explain in which cases (for which types of URLs) query
log data can be used to inform the user about the semantic
contents of a web document, or for disambiguation of a URL
or identi cation of query intent.
      </p>
    </sec>
    <sec id="sec-4">
      <title>EXPERIMENTS</title>
      <p>The objective of our experiments is to identify a set of
key factors that predict whether a URL can be classi ed
correctly using query log data. With the aid of such factors,
a search engine can use terms from query log data as
document descriptors, which help the user in disambiguating
URLs or nding URLs that match the user's query intent.
3.1</p>
    </sec>
    <sec id="sec-5">
      <title>Data</title>
      <p>RFP: The Microsoft 2006 RFP1 dataset consists of
approximately 14 million queries from US users entered into
the Microsoft Live search engine in the spring of 2006. For
each query the following details are available: a query ID,
the query itself, the user session ID, a time-stamp, the URL
of the clicked document, the rank of that URL in the result
list and the number of results.</p>
      <p>DMOZ: The DMOZ Open Directory RDF Dump2 is a
set of URLs and their class labels according to the Open
Directory Project DMOZ. E.g. bikeriderstours.com |
Top/ Sports/ Cycling/ Travel/ Tour_Operators. We
restricted the data to DMOZ level 2 labels (e.g Top/ Sports).
We discarded the URLs labelled Top/Regional, since
Regional is the top node of a di erent hierarchy (a regional
classi cation). The intersection of the RFP and DMOZ
collections (with the above restriction) consists of 245.742
URLs, distributed over 15 classes.
3.2</p>
    </sec>
    <sec id="sec-6">
      <title>Properties of URL and query</title>
      <p>
        In [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] the classi cation features for URLs in the
RFPDMOZ intersection were extracted by nding the query terms
that were most strongly associated with the URL3. These
features were aggregated at the level of the URL, the domain
of the URL and the individual words in the URL. Using these
features, the URLs were classi ed with Adaboost.MH [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
The highest classi cation accuracy was achieved when the
query terms were aggregated at the level of the domain of
the URL. In the current paper, we therefore focus on query
1http://research.microsoft.com/en-us/um/people/
nickcr/wscd09/
2http://rdf.dmoz.org/
3Strength of association was calculated using
KullbackLeibler divergence with the total query collection as
background model.
terms associated with URLs on the domain level,
aggregating queries over all URLs from our data collection that
belong to the same domain4.
      </p>
      <p>In order to investigate what properties of the URL and
the associated query terms play a role in the correct
classi cation of some URLs and the incorrect classi cation of
others, we extracted the following properties for each URL
in the RFP-DMOZ intersection:</p>
      <sec id="sec-6-1">
        <title>D: The domain of the URL.</title>
        <p>DL: The number of terms the domain was compounded of
(the domain length)5
NC: The number of clicks in the RFP dataset that were
associated with the domain.</p>
        <p>NUQ: The number of unique query terms associated with
the domain.</p>
        <p>AQC: The average number of query terms per click on a</p>
        <p>URL from the domain.</p>
        <p>MCP: The position that the clicked URL had in the result
list of the search engine, averaged over all clicks that
led to the domain.</p>
        <p>PN: The proportion of navigational queries in the total
number of queries that led to the domain of this URL.
We consider a query to be navigational if the
concatenated query terms are a substring of the URL string
(e.g. \bike riders tours" is a navigational query for the
URL www.bikeriderstours.com).</p>
        <p>TWE: The token-wise entropy of the domain, as the sum
of all the terms the domain is compounded of, i.e.:
H(D) = Pt2D P (t) log P (t), with P (t) the
probability of observing term t in any domain in the
RFPDMOZ intersection.</p>
        <p>KLD: The Kullback-Leibler divergence of the associated
query term probability distribution, relative to the
distribution of all query terms in the RFP dataset, i.e.:
DKL(P kQ) = Pt2D P (t) log QP((tt)) with P (t) the
probability of observing t in all queries associated with D
and Q(t) the probability of observing t in all queries
in the RFP collection.</p>
        <p>For all but the rst of these properties, we investigated
their relation to classi cation success. Our hypothesis was
that especially NC and NUQ would have a positive
predictive value for the classi cation success. We expect that more
clicks (higher NC) and more unique query terms (higher
NUQ) for a domain result in a better representation of the
domain and therefore in a better classi cation accuracy. The
details of our analyses are in Section 4 below.
4As a consequence, we can only investigate URL properties
that generalize to the domain level. Moreover, aggregating
on the domain level has the risk of grouping together
heterogenous URLs from large domains. We come back to this
in Sections 4 and 6.
5We decompounded the domains using a script that
subsequently looks up substrings in the CELEX lexicon (http:
//www.ldc.upenn.edu/) and greedily splits the domain
string into lemmatized lexicon entries. E.g. the domain
bikeriderstours.com was decompounded into the lemmas
bike, rider and tour ).</p>
        <p>We considered three di erent strategies for nding the
relevance of each of the predictors for the success of the
classi cation: calculating the correlation coe cient in
order to get an indication of the strength and the direction of
the relation between each predictor's value and the classi
cation outcome. However, this coe cient assumes a linear
relation that is independent of other predictors. Our data
seemed more complicated than that. Therefore, we assessed
the possibility of using a logistic regression model (LRM)
for predicting the classi cation outcome based on the
predictor values (normalized to their z-score). Unfortunately,
the LRM outcome was di cult to interpret: We did get
positive and negative predictor coe cients that signi cantly
contributed to the prediction model but the model t on the
data was relatively poor.</p>
        <p>These preliminary results suggested that there is no linear
relationship between any of the predictors that we
investigated and the classi cation success. We felt however that
some tendency could be discerned from the individual
predictors' values and the classi cation accuracies for speci c
ranges of these values. In order to assess this hypothesis,
we created 10 bins for the values range of each predictor.
Subsequently, we derived the classi cation accuracy for each
bin, together with the number of domains in this range. We
plotted these numbers in bar charts in order to visualize the
relation between the value ranges of the predictors and the
classi cation accuracy.</p>
        <p>Unfortunately, we did not nd very strong tendencies for
most of the predictors that would support the idea that the
classi cation success can be predicted from these properties.
Most of the bar charts appeared to be relatively at,
conrming that the classi cation accuracy is relatively stable,
only slightly dependent of the value of the predictor. As
an example, Figure 1 shows the classi cation accuracy as a
function of the token-wise entropy of the domain. The only
bar that is rising above the others is the right-most one,
representing the domains with the maximum entropy value.
However, this bar only represents a small number of domains
(3,423) and the classi cation accuracy for this range is still
mediocre (60%).</p>
        <p>In the next sub-section, we discuss the results for the two
predictors that we had expected to give the most promising
results (see Section 3.2).
4.1</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Analysis of the NC and NUQ predictors</title>
      <p>NC: When we look at the number of domains in
relation to (a range of) the number of clicks on those domains
(Figure 2), we rst notice that most domains in our data
collection have a small number of associated clicks (1 to 4). At
the same time we see that domains with the lowest numbers
of clicks are the domains with the lowest classi cation
accuracy. This con rms our earlier assumption that many clicks
result in a better representation of the domain and therefore
a better classi cation accuracy. The maximum classi cation
accuracy is 58% (for the range of 33{64 clicks). However,
Figure 2 also shows that for a higher number of clicks, the
classi cation accuracy starts to decrease again.</p>
      <p>We suspect that this behavior can be explained from the
heterogeneity of the domains that have a large number of
associated clicks. For example, portal web sites such as
ebay.com or amazon.com contain many URLs that may be
very diverse in their semantic content. Consequently, these
URLs are harder to classify, since in the aggregated term
set for the corresponding domain there are many terms for
semantically unrelated URLs from the same domain.</p>
      <p>NUQ: When we look at Figure 3, we see that the number
of domains in a given range of unique query terms decreases
much less sharply than for the number of clicks. However, we
see a similar pattern in the classi cation accuracies for these
ranges. For domains with 17{32 unique associated terms,
the accuracy is optimal. Figure 3 shows that classi cation
accuracy sharply increases initially for an increasing number
of unique query terms, starting at 30% for domains with only
1{2 unique terms, up to 60% for domains with 17{32 unique
terms. After that point, the accuracy decreases again.</p>
      <p>Domains with very few unique terms apparently provide
a too sparse classi cation vector to be classi ed correctly.
At the other end of the spectrum, domains with too many
unique terms are hard to classify as well. We again attribute
this to the heterogeneity of the domains with a large number
of unique query terms: it is very di cult to classify them as
belonging to a single class.</p>
    </sec>
    <sec id="sec-8">
      <title>5. DISCUSSION</title>
      <p>After analyzing our predictors in detail, we found that
many of them cannot predict the classi cation success. The
two most promising predictors (number of clicks and number
of unique query terms) showed interesting tendencies but do
not provide ranges of high accuracies (optimal ranges give
60% classi cation accuracy). It is clear that the success of
classifying URLs based on query terms depends on many
di erent factors. In the previous section, we mentioned the
heterogeneity of the domain as a potentially important
factor.</p>
      <p>If we want to adapt our strategy for the heterogeneity
of domains (for example, by not providing query-based
descriptions for very heterogenous domains), the question that
rises here is how we can identify domains as being
heterogenous. Two of the factors that we saw in Section 4 are the
number of clicks and the number of unique query terms that
are associated with a domain. A third factor may be the
domain size: the more URLs a domain contains, the larger the
heterogeneity of the domain probably is. Part of our future
work will be to estimate the domain heterogeneity based on
these factors.</p>
    </sec>
    <sec id="sec-9">
      <title>6. CONCLUSION AND FURTHER WORK</title>
      <p>
        We continued the work of [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and investigated which
factors are relevant for the success of URL classi cation based
on associated query terms. We created a series of
classication success predictors and subsequently analyzed their
relation to the classi cation success. None of the predictors
we investigated can fully predict the classi cation success.
We found however that a couple of predictors show
interesting tendencies: the number of clicks on URLs (NC) and the
number of unique terms associated with a URL (NUQ).
In both cases, the predictors initially correlate positively
with the classi cation accuracy, but after a certain
saturation point this correlation becomes negative. We suggest
that this is caused by heterogeneous domains (domains that
contain URLs from di erent semantic categories). We
argue that our suggested approach of providing query terms
as document descriptors for disambiguation is particularly
useful for URLs from homogenous domains.
      </p>
      <p>An important point for further research is to determine
the heterogeneity of a domain using query log data.
Another direction is to investigate what factors predict
classication accuracy when query terms are not aggregated on
domain level, but on the level of individual URLs. As these
cannot be heterogenous, it will be worthwhile to see the
performance of the predictors in this situation.</p>
      <p>We are currently experimenting with di erent types of
classi ers in order to see whether we can improve the
classi cation accuracy of our data. We also study our data in
more detail in order to see whether removing a subset of the
click data from the training set can increase the classi
cation performance. This subset can be either category-based
(remove noisy categories), feature-based (remove instances
with too few query terms) or based on overall consistency
(remove instances that have very similar term sets but
contradictory classes).</p>
      <p>
        In the somewhat more distant future, we aim to
investigate the possibilities of implementing our URL descriptor
approach in a user interface. Following the results obtained
by [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], we will combine salient terms from the URL's
content and the queries associated with the URL into a
semantic annotation of the URLs in the result list. One challenge
that we foresee for this experiment is the evaluation: User
judgments are time-consuming but essential for this kind of
implementation.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>E.</given-names>
            <surname>Agichtein</surname>
          </string-name>
          , E. Brill, and
          <string-name>
            <given-names>S.</given-names>
            <surname>Dumais</surname>
          </string-name>
          .
          <article-title>Improving web search ranking by incorporating user behavior information</article-title>
          .
          <source>In SIGIR '06: Proceedings of the 29th annual international ACM SIGIR conference on Research and development in information retrieval</source>
          , pages
          <volume>19</volume>
          {
          <fpage>26</fpage>
          , New York, NY, USA,
          <year>2006</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>I.</given-names>
            <surname>Antonellis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Garcia-Molina</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>J.</given-names>
            <surname>Karim</surname>
          </string-name>
          .
          <article-title>Tagging with queries: How and why?</article-title>
          <source>In ACM WSDM '09</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D. J.</given-names>
            <surname>Brenes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. G.</given-names>
            <surname>Avello</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K. P.</given-names>
            <surname>Gonzalez</surname>
          </string-name>
          .
          <article-title>Survey and evaluation of query intent detection methods</article-title>
          .
          <source>In Proceedings of WSCD '09</source>
          , pages
          <fpage>1</fpage>
          <article-title>{7</article-title>
          . ACM New York, NY, USA,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Fox</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Karnawat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mydland</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dumais</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>White</surname>
          </string-name>
          .
          <article-title>Evaluating implicit measures to improve web search</article-title>
          .
          <source>ACM Trans. Inf</source>
          . Syst.,
          <volume>23</volume>
          (
          <issue>2</issue>
          ):
          <volume>147</volume>
          {
          <fpage>168</fpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hinne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Kraaij</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Raaijmakers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Verberne</surname>
          </string-name>
          , T. van der Weide, and
          <string-name>
            <surname>M. van der Heijden.</surname>
          </string-name>
          <article-title>Annotation of URLs: more than the sum of parts</article-title>
          .
          <source>In SIGIR '09: Proceedings of the 32th ACM SIGIR international conference on Information Retrieval</source>
          , New York, NY, USA,
          <year>2009</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>T.</given-names>
            <surname>Joachims</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Granka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Hembrooke</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Gay</surname>
          </string-name>
          .
          <article-title>Accurately interpreting clickthrough data as implicit feedback</article-title>
          .
          <source>In Proceedings of the 28th annual international ACM SIGIR conference</source>
          , pages
          <volume>154</volume>
          {
          <fpage>161</fpage>
          . ACM New York, NY, USA,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D.</given-names>
            <surname>Kelly</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Teevan</surname>
          </string-name>
          .
          <article-title>Implicit feedback for inferring user preference: a bibliography</article-title>
          .
          <source>SIGIR Forum</source>
          ,
          <volume>37</volume>
          (
          <issue>2</issue>
          ):
          <volume>18</volume>
          {
          <fpage>28</fpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>B.</given-names>
            <surname>Krause</surname>
          </string-name>
          , R. Jaschke,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hotho</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Stumme</surname>
          </string-name>
          .
          <article-title>Logsonomy - social information retrieval with logdata</article-title>
          .
          <source>In Hypertext</source>
          , pages
          <volume>157</volume>
          {
          <fpage>166</fpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>B.</given-names>
            <surname>Poblete</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Baeza-Yates</surname>
          </string-name>
          .
          <article-title>Query-sets: using implicit feedback and query patterns to organize web documents</article-title>
          .
          <source>In Proceedings of WWW '08</source>
          , pages
          <fpage>41</fpage>
          {
          <fpage>50</fpage>
          . ACM New York, NY, USA,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>R. E.</given-names>
            <surname>Schapire</surname>
          </string-name>
          and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Singer. Boostexter</surname>
          </string-name>
          :
          <article-title>A boosting-based system for text categorization</article-title>
          .
          <source>Machine Learning</source>
          ,
          <volume>39</volume>
          :
          <fpage>135</fpage>
          {
          <fpage>168</fpage>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>M. van der Heijden</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Hinne</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          <string-name>
            <surname>Kraaij</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Verberne</surname>
          </string-name>
          , and T. van der Weide.
          <article-title>Using query logs and click data to create improved document descriptions</article-title>
          .
          <source>In Proceedings of WSCD '09</source>
          , pages
          <fpage>64</fpage>
          {
          <fpage>67</fpage>
          . ACM New York, NY, USA,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>