<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Investigating Per-user Time Sensitivity Of Search Topics</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, University of Cape Town</institution>
          ,
          <addr-line>Private Bag X3, Rondebosch, 7701, Cape Town</addr-line>
          ,
          <country country="ZA">South Africa</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Search engines give the same results for the same query. They do not consider that a user's topics of interest may diverge at di erent times even if the query terms are the same. This paper presents the ndings of a study into how di erent topics of interest of a user are in uenced by time. The results show that most of the users have time sensitive search patterns, indicating that they have di erent topics of interest that are dominant at di erent times.</p>
      </abstract>
      <kwd-group>
        <kwd>Information retrieval</kwd>
        <kwd>Query log analysis</kwd>
        <kwd>Topic modelling</kwd>
        <kwd>Users search behaviour</kwd>
        <kwd>Time sensitive search patterns</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Search engines are used to search and retrieve information from the Web. The
Web has information on almost every topic but the search engines do not consider
diverging interests of a user and retrieve the same results for a query even if it
is issued at di erent times.</p>
      <p>An example of this could be a user who is a computer science student who
also likes sports. So, it can't be said that (s)he will search only for the topics
related to computer science. It is possible that at some point of time, (s)he will
search for sports also. So, if (s)he issues a query "tag", during study hours, (s)he
may be looking for HTML tags but, in some leisure time, the same query may
mean Tag Sports Gear, a sporting goods brand.</p>
      <p>
        Although the user queries are mostly small and ambiguous in nature [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ],
better results can be provided if a user's topics of interest and search patterns are
known. Search patterns have been studied before based on the query logs. Query
logs serve as an excellent store of knowledge as they have complete information
about what the users have searched in a given time frame. It has been observed
by previous studies that a user's search behaviour varies from workplace to home
[
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. Change in topical categories and search query volume also varies with time
[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ][
        <xref ref-type="bibr" rid="ref23">23</xref>
        ][
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. According to these studies, after analysing the query log, they found
that there is a pattern in search queries. But a general search pattern cannot be
applicable to all users.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Research Question</title>
      <p>Do the topics of interest of a user vary with time of the day?</p>
      <p>
        Motivated by the observations of past studies, this study explored the time
sensitive search pattern of a user to nd his/ her di erent topics of interest
that are dominant at di erent times. This study observed queries issued by 100
di erent users in an AOL query log [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and analysed the query set of each one of
the 100 users individually. The details and outcomes of this study are presented
in this paper. Section 2 shows the related work on user's search behaviour,
pattern identi cation and topic modelling. Section 3 covers the methodology
of our work. Section 4 presents the analysis of the data. Section 5 includes
limitations of the study. Section 6 is the conclusion of the work and future
directions.
3
3.1
      </p>
    </sec>
    <sec id="sec-3">
      <title>Related Work</title>
      <sec id="sec-3-1">
        <title>User's Search Behaviour And Pattern Identi cation</title>
        <p>
          It is crucial to analyse a user's search behaviour to provide e ective and e cient
search services. The query log data can be used to know how users use the
search engines and also about their diverse interests and preferences [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. Rose
and Levinson [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] made an attempt to understand users' search goals. They
analysed an Alta Vista query log and found that the goal of users' searches is less
navigational and more resource seeking. In a work by Tyler and Teevan [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ], the
authors analysed repeated queries and user behaviour. According to this study,
search engines can capitalize this re- nding behaviour of the user to improve
the user's search experience. In the same line of re- nding behaviour, Tyler et
al. [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] in their study found that repeated re- nding behaviour also contains
diversi cation. According to Srivastva et al. [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ], identifying the patterns in
Web usage can be helpful for marketers in placing advertisements focusing on
a certain target group. Temporal analysis of the sequence patterns can prove
useful in nding the trending topics.
        </p>
        <p>
          Temporal analysis of query logs has also been done by many researchers to
explore users' search behaviour and search patterns. A signi cant outcome of
the study by Rieh [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] is that the author was able to nd the di erence in search
behaviour of the users. According to this study, users' search behaviour di ers in
their workplace from that at home. The websites visited during working hours
were mostly related to their work while, at home, the search was of diverse
nature. In a similar work [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ], Yuye and Alistair analysed an MSN query log.
They found a general pattern in the volume of queries. There was a peak early
in the week that dropped steadily until Friday and decreased sharply over the
weekend. This pattern speaks about the weekly routine of a common working
person. They also observed an hourly pattern and found a rise in the volume
of queries from early morning, peaking at noon and decreasing steadily through
midnight. Judit et al. [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] analysed the same MSN query log for topic speci c
analysis. They found that, during weekdays, queries related to work were
dominant. In a comparatively recent work by Michael et al. [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ], the authors analysed
a Russian query log spanning one year. According to this study, queries related
to categories like "Health" and "Beauty and Style" were distributed more or
less constantly throughout the year. Some categories like "Education" observed
a drop during vacation periods. John Cosley [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] has shown that the query terms
and search patterns of users vary with time of the day and also with device
(Mobile and PC). He also compared the search patterns of weekdays and weekends.
According to this study, queries regarding task completion were dominant
during the morning on weekdays, while entertainment and shopping related queries
have shown their dominance in the evening on Mobiles and Tablets. All these
studies have analysed the query logs as a whole and suggested that there is a
pattern in users' search behaviour. Some of them [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ][
          <xref ref-type="bibr" rid="ref2">2</xref>
          ][
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] also analysed the time
dependent popularity of some topics. They did not consider and analyse each
individual user's search pattern. A common trend of topic change and popularity
cannot be applied to improve the search experience of an individual user. Each
user may have a speci c search pattern, which is di erent from the others. As
reported by Michael et al. [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ], the queries related to the category "Education"
observed a decline during the vacation period; this trend or pattern cannot be
generalized for each user. A user, who is looking for extra classes or lessons, may
search for "education" even in the vacation period.
3.2
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Topic Inference</title>
        <p>
          Every user has di erent topics of interests. To nd the di erent topics from query
logs, most of the previous studies about topic based personalized information
retrieval systems have relied on Open Directory Project (ODP) categories [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ][
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]
[
          <xref ref-type="bibr" rid="ref4">4</xref>
          ][
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]. Jansen et al. [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] used the Google Directory topical hierarchy to classify
the queries into subject categories. Arguably, Web search is not limited to these
categories because of the rich nature of the Web. It is not ideal to put the
wide range of a user's interests into prede ned categories. Unsupervised Machine
Learning may be a better tool to learn latent topics from users' search queries
[
          <xref ref-type="bibr" rid="ref14">14</xref>
          ].
4
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Methodology</title>
      <p>The goal of this study is to investigate the time sensitivity of a user's topics of
interest. We analysed queries submitted by each user separately to explore the
time sensitive search pattern of that user. This section presents the details of
our study in the following steps:
4.1</p>
      <sec id="sec-4-1">
        <title>Data Collection</title>
        <p>
          In this study, the search history of 100 users from an AOL query log [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] has been
analysed. The AOL query log is publicly available log data for research and
analysis [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ][
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. The AOL query log has been analysed before by many researchers.
Duarte et al. [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] identi ed queries with children intent from this AOL query log.
This log collection contains about 20M Web queries from 650K users issued in
three months from March 2006 to May 2006. The data is anonymized and
consists of: UserID, Query, QueryTime, ClickedRank, DestinationDomainUrl. Each
UserID represents a unique user. For this study, UserID, Query and QueryTime
elds of the query log were considered. It was assumed that each UserID is
representing a unique user. The logs of each unique user were cleaned and
preprocessed for further analysis. The details of data cleaning and pre-processing
are described below.
4.2
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>Data Cleaning</title>
        <p>
          The queries of each of the 100 users were processed individually. The entries
with empty queries were removed. The same queries that were submitted on
the same date within a time di erence of less than 10 seconds were also not
considered for analysis. According to Odjik et al. [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ], queries issued within a
few seconds of time are more likely to be spelling correction or substitution type
of formulations of the previous query issued. These queries were not considered
as di erent queries but some modi cation of the previous ones. This process of
removing incomplete, irrelevant or duplicate data is called data cleaning.
4.3
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>Pre-processing</title>
        <p>
          Pre-processing, also known as text normalization, gives a syntactical view of the
original text. Pre-processing was accomplished by using the Natural Language
Toolkit (NLTK). NLTK is a free and open source community-driven project [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ].
It is the most used platform to work with natural human language data. This
involves tokenization, stopword and punctuation marks removal, and
lemmatization.
        </p>
        <p>The process of dividing a phrase or sentence into tokens is called tokenization.
The tokens may represent words, digits or punctuation marks. After
tokenization, stopwords were detected from the data. Stopwords are common and high
frequency words that are independent of any topic like a, an, the, for and and.
Following detection of stopwords and punctuation marks, they were removed
from the query terms and the query terms were lemmatized.</p>
        <p>Lemmatization aims to remove in ectional endings and to return the base or
dictionary form of a word, which is called the lemma. This step utilizes
vocabulary along with a morphological analysis of words.</p>
        <p>Table 1 shows the aggregate number of queries of 100 users and the queries
that remained for analysis after cleaning and pre-processing.
4.4</p>
      </sec>
      <sec id="sec-4-4">
        <title>Topic Modelling</title>
        <p>
          Topic modelling is a technique that is used to identify the latent topics present in
a corpus. Topic models are the algorithms that are used to nd the main themes
or ideas in a data collection. In a number of previous works, authors have utilized
prede ned topical categories like the Open Directory Project [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ][
          <xref ref-type="bibr" rid="ref4">4</xref>
          ][
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] and the
Google Directory topical hierarchy [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] to nd the topics of interest of a user.
According to Mehrotra [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ], to learn the latent topics of interest from the user
search logs, unsupervised machine learning would be a better tool.
        </p>
        <p>
          Latent Dirichlet Allocation (LDA) is an unsupervised approach for topic
modelling. It is a generative probabilistic mode for a text corpus [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] and the most
commonly used approach for topic modelling. LDA is a three-level hierarchical
Bayesian model. In LDA, each document of a corpus is modeled as a nite
mixture over an underlying set of topics and each topic is modeled as an in nite
mixture over an underlying set of topic probabilities [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ].
        </p>
        <p>
          In this study, we assumed that each user has at least 4 di erent topics of
interest. After applying LDA on the queries of each user, topics were assigned
to the individual queries. According to LDA, a document may have more than
one topic. So, in this case, if a query fell under more than one topic, all those
topics were assigned to that query. The reason behind doing this is that most of
the Web queries are ambiguous in nature i.e. they may belong to many topics.
The aim of this study is to nd the temporal dominance of topics of interest of a
user. We analysed which topics were dominant at a particular time. This notion
of assigning multiple topics to a query along with nding temporal dominance
of topics can be utilized to disambiguate the ambiguous queries. After assigning
the topic(s), the queries were grouped according to the Time-bins.
The aim of this study is to nd time sensitive search patterns. We analyse the
relation between a topic of interest and the time when it is searched dominantly.
According to Rieh [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ], users search behaviour di ers in their workplace from
that at home. For the purpose of our study, we divided the time of a day into four
customized bins according to the common daily routine of a working person. The
four Time-bins are: Early morning (6h-8hr), Working hours (8h-18h), Evening
time (18h-24h) and Midnight (24h-6h).
5
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Results And Analysis</title>
      <p>In this section, search patterns of all of the 100 users were analysed. Keeping in
mind that this AOL query log is from 2006 i.e. more than 10 years old, people
searched less frequently because of Internet availability and usage cost. People
used to search for fewer topics. In present times, due to fast Internet connections,
easy availability and more advanced communication devices, users search more
frequently. Moreover, the number of topics of interest has also increased. In spite
of this limitation, some promising facts are revealed from its analysis. Table 2
shows the number of Time-bins used by the users to issue queries.</p>
      <p>As can be seen from Table 2, out of 100 users, the majority of the users (50)
have utilized 2 Time- bins to issue their queries. 35 users have made their queries
in 3 Time-bins while 12 users have searched in 4 Time-bins. Very few users (3)
have searched only in one Time-bin. Table 3 presents the time-sensitive search
patterns of users along with their respective numbers of dominant topics. Every
unique number in Example Pattern indicates the di erent dominant topic in a
user's search pattern. For instance, pattern 123 represents the search pattern of
a user who has searched in 3 Time-bins and in every Time-bin the dominant
topic is di erent. One of the possible pattern, 1122, when a user searches for
2 dominant topics in 4 Time-bins, was not found in any of the 100 patterns
analysed. Out of 100 users, 1 user has searched for 4 di erent dominant topics
in the 4 Time-bins, 13 users have searched 3 Time-bins with 3 distinct dominant
topics and 42 users have 2 di erent dominant topics for the 2 Time-bins they
searched in. So, 56 users have searched for di erent dominant topics in every
Time-bin. Among the users who searched in 4 Time-bins, 6 have 3 dominant
topics and 3 have 2 dominant topics. 19 users have 2 dominant topics in 3
Timebins they searched. Thus, there are 28 users who have at least either 3 or 2
dominant topics of interest. Only 13 users have shown the same dominant topic
in every searched Time-bin and, as shown in Table 2, 3 users have searched in
only 1 Time-bin.</p>
      <p>The following gures represent search patterns of some of the categories
shown in Table 3.</p>
      <p>Figure 1 shows the search pattern of Category 2 users. Each time the user
has searched, the dominant topic is di erent. Topic3 which is not searched much
at Time1 and Time3, becomes dominant at Time4.</p>
      <p>Figure 2 represents the search pattern of a Category 4 user. The user has
searched only in 2 Time-bins but, in both the Time-bins, the dominant topics
are di erent. Topic2, which has not been searched at Time3, is dominating at
Time4. Figure 3 shows the search pattern of a Category 5 user who has made
searches in 3 Time-bins. At Time1, only Topic0 is searched and it also dominates
at Time3. Topic2, which has not been searched either at Time1 or at Time3,
clearly dominates at Time4.
50
45
40
35
30
25
20
15
10
5
0
60
50
40
30
20
10
0</p>
      <p>Time1</p>
      <p>Time2</p>
      <p>Time3</p>
      <p>Time4</p>
      <p>These search patterns indicate that users have di erent topics of interest and
they prefer to search about them at di erent times.</p>
      <p>For analysing the variability in patterns of dominant topics, entropy was
calculated. Entropy is a measure of disorder or randomness and refers to the
Topic0
Topic1
Topic2
Topic3
90
80
70
60
50
40
30
20
10
0</p>
      <p>Time1</p>
      <p>Time2</p>
      <p>Time3</p>
      <p>Time4
number of possible states a variable can have. It is calculated as:
H(X) =
n
X lnp(xi):p(xi)
i=1
A greater value of entropy points to more possible states or randomness of a
variable. Table 3 shows the calculated entropy for di erent patterns of dominant
topics, suggesting an ordering and grouping of di erent topic patterns based on
variability.</p>
      <p>80 users have entropy greater than or equal to 0.636. Even if a user searched
in at least 3 Time-bins, the dominant topics were unique in 2 Time-bins. In
other words, some topics are dominantly searched in a particular time interval.
It clearly means that there is variability and uncertainty in a user's topics of
interest. A user's topics of interest di er in di erent time intervals and so we
can say that the topics of interest of a user are time-sensitive.</p>
      <p>This inference can be utilized to disambiguate the queries and provide more
useful and relevant search results.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Limitations</title>
      <p>The aim of this study is to explore the temporal dominance of topics of interest of
a user. For this purpose, we selected the queries submitted by 100 users from an
AOL query log. The time of the day was divided into four bins and it is assumed
that each user has at least four topics of interest. After processing and analysing,
it was found that most of the users search for di erent topics at di erent times.
It can be said that topics are time sensitive.</p>
      <p>This study was able to nd the time-sensitive search patterns of a user but
it has some shortcomings also.
1. The query log data was old but readily available for analysis.
2. The queries of only 100 users were analysed because the query set of every
user was cleaned manually.
3. It did not gure out the exact number of topics of interest of each user. As
we divided the time of the day into 4 Time-bins according to the routine
of a common working person, it was assumed that each user has at least 4
topics of interest for 4 time bins.
4. Although we were able to nd search patterns based on these Time-bins,
each user may have di erent search Time-bins.
7</p>
    </sec>
    <sec id="sec-7">
      <title>Conclusions And Future Work</title>
      <p>We have studied an AOL query log to explore the time sensitive search patterns
of users. We have analysed the queries of 100 di erent users and found that,
out of 100 users, 84 users have at least 2 di erent dominant topics searched at
di erent times. Only 13 users have searched for the same topic in every
Timebin and 3 users have searched only in 1 Time-bin. This study concludes that
most of the users have time sensitive topics. They search for di erent topics at
di erent times. Di erent topics are dominant at di erent time intervals. The
goal of future work is to explore and exploit the time sensitive search patterns
of a user to model a user's time sensitive search behaviour, which could prove
helpful to search engines in disambiguating the short and ambiguous queries and
also providing users with more relevant search results.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1. http://www.cim.mcgill.ca/ dudek/206/Logs/
          <article-title>AOL-user-ct-collection/user-ct-testcollection-01</article-title>
          .txt/.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Judit</given-names>
            <surname>Bar-Ilan</surname>
          </string-name>
          , Zheng Zhu, and
          <string-name>
            <given-names>Mark</given-names>
            <surname>Levene</surname>
          </string-name>
          .
          <article-title>Topic-speci c analysis of search queries</article-title>
          .
          <source>In Proceedings of the 2009 Workshop on Web Search Click Data, WSCD '09</source>
          , pages
          <fpage>35</fpage>
          {
          <fpage>42</fpage>
          , New York, NY, USA,
          <year>2009</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Steven</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Beitzel</surname>
            , Eric C. Jensen, Abdur Chowdhury, David Grossman,
            <given-names>and Ophir</given-names>
          </string-name>
          <string-name>
            <surname>Frieder</surname>
          </string-name>
          .
          <article-title>Hourly analysis of a very large topically categorized web query log</article-title>
          .
          <source>In Proceedings of the 27th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR '04</source>
          , pages
          <fpage>321</fpage>
          {
          <fpage>328</fpage>
          , New York, NY, USA,
          <year>2004</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4. Paul N Bennett, Ryen W White, Wei Chu,
          <string-name>
            <surname>Susan T Dumais</surname>
          </string-name>
          ,
          <string-name>
            <surname>Peter Bailey</surname>
            , Fedor Borisyuk, and
            <given-names>Xiaoyuan</given-names>
          </string-name>
          <string-name>
            <surname>Cui</surname>
          </string-name>
          .
          <article-title>Modeling the impact of short-and long-term behavior on search personalization</article-title>
          .
          <source>In Proceedings of the 35th international ACM SIGIR conference on Research and development in information retrieval</source>
          , pages
          <volume>185</volume>
          {
          <fpage>194</fpage>
          . ACM,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>David</surname>
            <given-names>M Blei</given-names>
          </string-name>
          , Andrew Y Ng, and
          <string-name>
            <given-names>Michael I</given-names>
            <surname>Jordan</surname>
          </string-name>
          .
          <article-title>Latent dirichlet allocation</article-title>
          .
          <source>Journal of machine Learning research</source>
          ,
          <volume>3</volume>
          (Jan):
          <volume>993</volume>
          {
          <fpage>1022</fpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>John</given-names>
            <surname>Cosley</surname>
          </string-name>
          .
          <article-title>Hearing the rhythms of human search behavior: What weve learned</article-title>
          . http://searchengineland.com
          <article-title>/human-behavior-in uences-searchmarketing-</article-title>
          <volume>197486</volume>
          /,
          <year>July 2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Sergio</given-names>
            <surname>Duarte</surname>
          </string-name>
          <string-name>
            <surname>Torres</surname>
          </string-name>
          , Djoerd Hiemstra, and
          <string-name>
            <given-names>Pavel</given-names>
            <surname>Serdyukov</surname>
          </string-name>
          .
          <article-title>Query log analysis in the context of information retrieval for children</article-title>
          .
          <source>In Proceedings of the 33rd international ACM SIGIR conference on Research and development in information retrieval</source>
          , pages
          <volume>847</volume>
          {
          <fpage>848</fpage>
          . ACM,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>Yi</given-names>
            <surname>Fang</surname>
          </string-name>
          , Naveen Somasundaram, Luo Si, Jeongwoo Ko, and
          <string-name>
            <surname>Aditya P Mathur.</surname>
          </string-name>
          <article-title>Analysis of an expert search query log</article-title>
          .
          <source>In Proceedings of the 34th international ACM SIGIR conference on Research and development in Information Retrieval</source>
          , pages
          <volume>1189</volume>
          {
          <fpage>1190</fpage>
          . ACM,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>Mansi</given-names>
            <surname>Gera</surname>
          </string-name>
          and
          <string-name>
            <given-names>Shivani</given-names>
            <surname>Goel</surname>
          </string-name>
          .
          <article-title>Data mining - techniques, methods and algorithms: A review on tools and their validity</article-title>
          .
          <source>International Journal of Computer Applications</source>
          ,
          <volume>113</volume>
          (
          <issue>18</issue>
          ),
          <year>2015</year>
          . Copyright - Copyright
          <source>Foundation of Computer Science 2015; Last updated - 2015-04-14.</source>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Bernard</surname>
            J Jansen, Zhe Liu, Courtney Weaver, Gerry Campbell, and
            <given-names>Matthew</given-names>
          </string-name>
          <string-name>
            <surname>Gregg</surname>
          </string-name>
          .
          <article-title>Real time search on the web: Queries, topics, and economic value</article-title>
          .
          <source>Information Processing &amp; Management</source>
          ,
          <volume>47</volume>
          (
          <issue>4</issue>
          ):
          <volume>491</volume>
          {
          <fpage>506</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Bernard</surname>
            J Jansen, Amanda Spink, and
            <given-names>Tefko</given-names>
          </string-name>
          <string-name>
            <surname>Saracevic</surname>
          </string-name>
          .
          <article-title>Real life, real users, and real needs: a study and analysis of user queries on the web</article-title>
          .
          <source>Information processing &amp; management</source>
          ,
          <volume>36</volume>
          (
          <issue>2</issue>
          ):
          <volume>207</volume>
          {
          <fpage>227</fpage>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Hyoung R Kim and Philip K Chan</surname>
          </string-name>
          .
          <article-title>Learning implicit user interest hierarchy for context in personalization</article-title>
          .
          <source>In Proceedings of the 8th international conference on Intelligent user interfaces</source>
          ,
          <source>pages</source>
          <volume>101</volume>
          {
          <fpage>108</fpage>
          . ACM,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13. Jin Young Kim, Kevyn Collins-Thompson, Paul N Bennett, and Susan T Dumais.
          <article-title>Characterizing web content, user interests, and search behavior by reading level and topic</article-title>
          .
          <source>In Proceedings of the fth ACM international conference on Web search and data mining</source>
          , pages
          <volume>213</volume>
          {
          <fpage>222</fpage>
          . ACM,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <given-names>Rishabh</given-names>
            <surname>Mehrotra</surname>
          </string-name>
          .
          <article-title>Topics, tasks &amp; beyond: Learning representations for personalization</article-title>
          .
          <source>In Proceedings of the Eighth ACM International Conference on Web Search and Data Mining</source>
          , pages
          <volume>459</volume>
          {
          <fpage>464</fpage>
          . ACM,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Ashish</surname>
            <given-names>Nanda</given-names>
          </string-name>
          , Rohit Omanwar, and
          <string-name>
            <given-names>Bharat</given-names>
            <surname>Deshpande</surname>
          </string-name>
          .
          <article-title>Implicitly learning a user interest pro le for personalization of web search using collaborative ltering</article-title>
          .
          <source>In Web Intelligence (WI) and Intelligent Agent Technologies (IAT)</source>
          ,
          <year>2014</year>
          IEEE/WIC/ACM International Joint Conferences on, volume
          <volume>2</volume>
          , pages
          <fpage>54</fpage>
          {
          <fpage>62</fpage>
          . IEEE,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Daan</surname>
            <given-names>Odijk</given-names>
          </string-name>
          , Ryen W White, Ahmed Hassan Awadallah, and Susan T Dumais.
          <article-title>Struggling and success in web search</article-title>
          .
          <source>In Proceedings of the 24th ACM International on Conference on Information and Knowledge Management</source>
          , pages
          <volume>1551</volume>
          {
          <fpage>1560</fpage>
          . ACM,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Greg</surname>
            <given-names>Pass</given-names>
          </string-name>
          , Abdur Chowdhury, and
          <string-name>
            <given-names>Cayley</given-names>
            <surname>Torgeson</surname>
          </string-name>
          .
          <article-title>A picture of search</article-title>
          .
          <source>In Proceedings of the 1st International Conference on Scalable Information Systems</source>
          , InfoScale '
          <fpage>06</fpage>
          , New York, NY, USA,
          <year>2006</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18. Soo Young Rieh.
          <article-title>Investigating web searching behavior in home environments</article-title>
          .
          <source>Proceedings of the American Society for Information Science and Technology</source>
          ,
          <volume>40</volume>
          (
          <issue>1</issue>
          ):
          <volume>255</volume>
          {
          <fpage>264</fpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19. Daniel E. Rose and
          <string-name>
            <given-names>Danny</given-names>
            <surname>Levinson</surname>
          </string-name>
          .
          <article-title>Understanding user goals in web search</article-title>
          .
          <source>In Proceedings of the 13th International Conference on World Wide Web, WWW '04</source>
          , pages
          <fpage>13</fpage>
          {
          <fpage>19</fpage>
          , New York, NY, USA,
          <year>2004</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Jaideep</surname>
            <given-names>Srivastava</given-names>
          </string-name>
          , Robert Cooley, Mukund Deshpande, and
          <string-name>
            <surname>Pang-Ning Tan</surname>
          </string-name>
          .
          <article-title>Web usage mining: Discovery and applications of usage patterns from web data</article-title>
          .
          <source>SIGKDD Explor</source>
          . Newsl.,
          <volume>1</volume>
          (
          <issue>2</issue>
          ):
          <volume>12</volume>
          {
          <fpage>23</fpage>
          ,
          <year>January 2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Sarah</surname>
            <given-names>K</given-names>
          </string-name>
          <string-name>
            <surname>Tyler and Jaime Teevan</surname>
          </string-name>
          .
          <article-title>Large scale query log analysis of re- nding</article-title>
          .
          <source>In Proceedings of the third ACM international conference on Web search and data mining</source>
          , pages
          <volume>191</volume>
          {
          <fpage>200</fpage>
          . ACM,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <article-title>Sarah K Tyler and Yi Zhang. Multi-session re-search: in pursuit of repetition and diversi cation</article-title>
          .
          <source>In Proceedings of the 21st ACM international conference on Information and knowledge management</source>
          , pages
          <year>2055</year>
          {
          <year>2059</year>
          . ACM,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23. Michael Volske, Pavel Braslavski, Matthias Hagen, Galina Lezina, and
          <string-name>
            <given-names>Benno</given-names>
            <surname>Stein</surname>
          </string-name>
          .
          <article-title>What users ask a search engine: Analysing one billion russian question queries</article-title>
          .
          <source>In Proceedings of the 24th ACM International on Conference on Information and Knowledge Management</source>
          ,
          <source>CIKM '15</source>
          , pages
          <fpage>1571</fpage>
          {
          <fpage>1580</fpage>
          , New York, NY, USA,
          <year>2015</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <article-title>Yuye Zhang and Alistair Mo at</article-title>
          .
          <article-title>Some observations on user search behaviour</article-title>
          .
          <source>Austr. J. Intelligent Information Processing Systems</source>
          ,
          <volume>9</volume>
          (
          <issue>2</issue>
          ):1{
          <issue>8</issue>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>