<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>April</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>A Study of User Profile Generation from Folksonomies</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nicholas Gibbins nmg@ecs.soton.ac.uk</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Algorithms</institution>
          ,
          <addr-line>Experimentation, Human Factors</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Ching-man Au Yeung</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Intelligence, Agents, Multimedia Group School of Electronics and Computer Science University of Southampton Southampton</institution>
          ,
          <addr-line>SO17 1BJ</addr-line>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Nigel Shadbolt</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2008</year>
      </pub-date>
      <volume>22</volume>
      <issue>2008</issue>
      <abstract>
        <p>Recommendation systems which aim at providing relevant information to users are becoming more and more important and desirable due to the enormous amount of information available on the Web. Crucial to the performance of a recommendation system is the accuracy of the user pro¯les used to represent the interests of the users. In recent years, popular collaborative tagging systems such as del.icio.us have aggregated an abundant amount of user-contributed metadata which provides valuable information about the interests of the users. In this paper, we present our analysis on the personal data in folksonomies, and investigate how accurate user pro¯les can be generated from this data. We reveal that the majority of users possess multiple interests, and propose an algorithm to generate user pro¯les which can accurately represent these multiple interests. We also discuss how these user pro¯les can be used for recommending Web pages and organising personal data.</p>
      </abstract>
      <kwd-group>
        <kwd>collaborative tagging</kwd>
        <kwd>folksonomy</kwd>
        <kwd>personomy</kwd>
        <kwd>user pro¯le</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Categories and Subject Descriptors</title>
      <p>H.3.4 [Information Storage and Retrieval]: Systems
and Software; H.3.5 [Information Storage and Retrieval]:
Online Information Services; H.5 [Information Interfaces
and Presentation (I.7)]:</p>
    </sec>
    <sec id="sec-2">
      <title>1. INTRODUCTION</title>
      <p>
        The amount of resources on the Web nowadays is so
enormous that retrieval of relevant information is getting more
and more di±cult. While users are desperate to obtain
information that is relevant to their needs and to avoid
information that are irrelevant, publishers of resources are also
eager to deliver their information to their targeted readers.
This has resulted in the rise of recommendation systems [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
which aim to recommend relevant and interesting resources
to users. An important aspect of user pro¯les is whether
they can truly re°ect the interests or expertise of the users.
      </p>
      <p>
        Some research works attempt to construct user pro¯les based
on the browsing history of the users [
        <xref ref-type="bibr" rid="ref22 ref9">9, 22</xref>
        ], or on the
documents collected by the users [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        Recently, the rising popularity of collaborative tagging
systems [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], such as del.icio.us1 and Flickr2, has provided
new sources for understanding the interests of Web users.
Collaborative tagging systems allow users to choose their
own words as tags to describe their favourite Web resources,
resulting in an emerging classi¯cation scheme now commonly
known as a folksonomy [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. Given that the resources and
the tags posted by Web users to these systems are
supposed to be highly dependent on their interests, folksonomies
thus provide rich information for building more accurate and
more speci¯c user pro¯les for use in various applications.
      </p>
      <p>
        Currently, only a few studies in the literature try to
construct user pro¯les from data in collaborative tagging
systems [
        <xref ref-type="bibr" rid="ref13 ref5">5, 13</xref>
        ], and usually only a single set of popular tags are
used to represent user interests. However, we observe that
tags used by users are very diverse and span across many
different domains. This implies that users usually have a wide
range of interests. Therefore, a single set of tags may not be
the most suitable representation of a user pro¯le, as it is not
able to re°ect the multiple interests of users. In this paper,
we propose a network analysis technique performed on the
personomy [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] of a user to identify the di®erent interests of
a user, and to construct a more comprehensive user pro¯le
based on the results. Evaluations show that our algorithm
is able to reveal the di®erent domains in which the users are
interested, and construct more informative and speci¯c user
pro¯les.
      </p>
      <p>This paper is structured as follows. Section 2 introduces
folksonomies and personomies. Section 3, presents the
analysis of the data collected from del.icio.us which motivated
this research. In Section 4, we describe in detail our
proposed algorithm for user pro¯le construction. Evaluations,
discussions and potential applications are presented in
Section 5. We mentioned related works in Section 6. Finally,
Section 7 concludes the paper and gives future research
directions.
2.</p>
    </sec>
    <sec id="sec-3">
      <title>FOLKSONOMIES AND PERSONOMIES</title>
      <p>
        Folksonomies [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] are user-contributed data aggregated by
collaborative tagging systems. In these systems, users are
allowed to choose terms freely to describe their favourite Web
resources. A folksonomy is generally considered to consist
1http://del.icio.us/
2http://www.°ickr.com/
of at least three sets of elements, namely users, tags and
resources. Although there can be di®erent kinds of resources,
in this article we will focus on Web documents, such as those
being bookmarked in del.icio.us. Formally, a folksonomy is
de¯ned as follows [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
      </p>
      <p>De¯nition 1. A folksonomy F is a tuple F = (U; T; D; A),
where U is a set of users, T is a set of tags, D is a set of
Web documents, and A µ U £ T £ D is a set of annotations.</p>
      <p>
        If we want to understand the interests of a single user, we
only need to concentrate on the tags and documents that
are associated with this particular user. Such set of data is
given the name personomy [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].3
      </p>
      <p>De¯nition 2. A personomy Pu of a user u is a restriction
of a folksonomy F to u: i.e. Pu = (Tu; Du; Au), where Au is
the set of annotations of the user: Au = f(t; d)j(u; t; d) 2 Ag,
Tu is the user's set of tags: Tu = ftj(t; d) 2 Aug, and Du is
the user's set of documents: Du = fdj(t; d) 2 Aug.</p>
      <p>
        This de¯nition is identical to the one mentioned in [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ],
except that we choose to exclude the sub-tag/super-tag
relation, since most collaborative tagging systems do not o®er
such functionality and we will not deal with this here.
      </p>
      <p>To perform analysis on the personomy of a user, we ¯rst
represent the personomy in the form of a network, with
nodes representing tags and documents associated with the
user. If folksonomy can be considered as a hypergraph with
three disjoint sets of nodes (user, tags and documents), a
personomy can be represented as a bipartite graph by
extracting the part that is related to the user. The bipartite
graph T Du of a personomy of a user u is de¯ned as follows.</p>
      <p>T Du = hTu [ Du; Etdi; Etd = f(t; d)j(t; d) 2 Aug
An edge exists between a tag and a document if the tag is
assigned to the document. The graph can be represented in
matrix form, which we denote as X = fxij g, xij = 1 if there
is an edge connecting ti and dj , and xij = 0 otherwise.</p>
      <p>
        To perform document clustering, we can fold the
bipartite graph into a one-mode network [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] of documents: D =
X0X. The adjacency matrix D represents the personal
repository of the user. Links between documents are weighted by
the number of tags that have been assigned to both
documents. Thus, documents with higher weights on the links
between them can be considered as more related. On the
other hand, a one-mode network of tags can be constructed
in a similar fashion: T = X0X. T represents semantic
network which shows the associations between di®erent tags.
In other words, this is the personal vocabulary or a simple
ontology used by the particular user.
      </p>
      <p>To facilitate the following discussions, we further de¯ne
several notations here. Firstly, we denote the set of
documents tagged by the tag t in the personomy of user u by
Du;t:</p>
      <p>Du;t = fdj(t; d) 2 Aug
Also, we de¯ne Cou(t1; t2) which indicates whether two tags
t1 and t2 have been used on the same document by a user:
Cou(t1; t2) =
½ 1 if (t1; d) 2 Au, (t2; d) 2 Au for some d</p>
      <p>0 otherwise
3In the blogosphere, the term personomy has also been
used in a more general sense to represent the
aggregated digit manifestation of a user on the Web. See
http://personomies.com/what-are-personomies/.</p>
      <sec id="sec-3-1">
        <title>Bookmarks</title>
      </sec>
      <sec id="sec-3-2">
        <title>Total number of users Maximum Tags Minimum Mean</title>
        <p>Maximum
Minimum
Mean</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>ANALYSIS OF PERSONOMIES</title>
      <p>To understand the characteristics of personomies in
collaborative tagging systems, we perform analysis on data
collected from del.icio.us. In particular, we want to gain insight
into the general behaviour of Web users using these systems.
We also want to understand if users are generally interested
in a rather speci¯c domain, such as we might expect when
studying the publications of a researcher, or if they are more
likely to be interested in a wide range of topics.</p>
      <p>In December 2007, we collected the bookmarking data
of 9,431 users of del.icio.us, including their bookmarks and
the tags they used, by crawling del.icio.us user names which
appeared on the page showing the recently updated
bookmarks.4 It is noted that among the 9,431 users whose data
we have collected, 246 of them apply no tags to any of their
stored bookmarks. These users are ¯ltered when performing
the following analysis. We summarise the statistics of the
data of the remaining 9,185 users in Table 1 and Figure 1.
3.1</p>
    </sec>
    <sec id="sec-5">
      <title>Number of Tags and Bookmarks of a User</title>
      <p>
        From the summary of the data in Table 1 and Figure 1,
we can see that on average a user have used 285 unique
tags and have saved 602 unique bookmarks on del.icio.us.
Although some users have over 18,000 tags and over 34,000
bookmarks, only a very small number of users have more
than a thousand tags or bookmarks. This ¯nding agrees
with what Golder and Huberman [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] report in their paper,
showing that there are a small number of users having a large
number of tags and bookmarks, and a large number of users
having a small number of tags and bookmarks, suggesting a
power-law distribution.
      </p>
      <p>In addition, we examine the correlation between the
number of tags and the number of bookmarks of the users.
Figure 2 shows a scatter plot of the data. It shows a moderate
relationship between the number of tags and the number of
bookmarks, with a correlation coe±cient of 0.55.</p>
      <p>In fact, it is natural to suggest that when there are more
bookmarks more tags are required to distinguish between
di®erent bookmarks by putting them into more speci¯c
categories. However the bookmarks and tags of the users in
the system are also highly dependent on the interests of the
users. If a user has a very speci¯c interest, a small number
of tags will be enough for even a large number of bookmarks,
as they will probably be about the same topic. On the other
hand, if a user has diverse interests, more tags may be
required to describe even a small number of bookmarks.</p>
      <p>
        A further investigation of the data reveals that the
correlation between the two numbers is stronger for users with
fewer bookmarks than those with many bookmarks. For
users with fewer than 500 bookmarks, the correlation
coe±cient is 0.43. For users with more than 5,000 bookmarks, the
4http://del.icio.us/recent
(a) Tags
(b) Bookmarks
correlation coe±cient is only 0.14. A similar result can also
be found in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. This may suggest that users with many
bookmarks can behave very di®erently: while some may
stick to using a small number of tags on new bookmarks,
others may continue to introduce new tags.
3.2
      </p>
    </sec>
    <sec id="sec-6">
      <title>Multiple Interests of Users</title>
      <p>With the average number of bookmarks signi¯cantly larger
than the average number of tags being used, it is obvious
that users are very likely to use a tag to describe more than
one bookmark. However, the usage of tags also depends on
the diversity of interests of the users. A user with only one
or two speci¯c interests is likely to use fewer tags than
another user who is interested in topics across several di®erent
domains. To understand this aspect of users in
collaborative tagging system, we propose two measures which re°ect
the diversity of interests of the users. We will give examples
based on the two ¯ctional users in Table 2, one with rather
speci¯c interests in Semantic Web related topics, while
another has more diverse interests such as cooking and sports.</p>
      <p>Firstly, we study the relations between the tags and the
bookmarks. If the tags used by a user are all assigned to
most of the bookmarks, the user is likely to have a rather
speci¯c interest, because this set of tags applies to most of
the documents that the user is interested in. On the other
hand, if most of the tags are only used on a small fraction
of bookmarks, it is likely that the user has a broader range
of interests. To quantify this characteristic, we propose a
measure called tag utilisation which is de¯ned as follows.</p>
      <p>De¯nition 3. Tag utilisation (TU) of a user u is the
average of the fractions of bookmarks on which a tag is used:
T agU til(u) =
1</p>
      <p>X jDu;tj
jTuj t2Tu jDuj
(1)</p>
      <p>In addition, the diversity of a user's interest can also be
understood by examining tag co-occurrence. If for a user
the tags are always used together with each other, it is likely
that the tags are about similar topics, and so the user should
have a rather speci¯c interest. If on the other hand the
tags are mostly used separately, they are more likely to be
about di®erent topics, and thus re°ect that the user should
have multiple interests which are quite distinctive from each
other. Such characteristic can be measure by average tag
co-occurrence ratio, which is de¯ned as follows.</p>
      <p>De¯nition 4. Average tag co-occurrence ratio (ATCR) of
a user measures how likely two tags are used together on the
same bookmark by a user:</p>
      <p>Avg T ag Co(u) =</p>
      <p>X</p>
      <p>Co(ti; tj)
ti;tj2Tu;ti6=tj 2 £ C2jTuj
(2)</p>
      <p>
        If we represent the co-occurrences between the tags as a
network (by constructing the adjacency matrix T), we can
see that the average tag co-occurrence ratio is actually
equivalent to the density of the network of tags: Co(ti; tj ) counts
the number of edges in the network, while CjTuj calculates
2
the number of possible edges based on the number of nodes.
This agrees with the formula of the density of a network:
where E is the set of edges and V is the set of nodes. Hence,
the average tag co-occurrence ratio actually re°ects the
cohesion [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] of the network of tags, which in turn re°ects
whether the tags are related to a speci¯c domain or a wide
range of topics.
      </p>
      <p>As an illustrating example, we apply these two measures
to the two users listed in Table 2. The tag utilisation of
u1 is 0:60, while that of u2 is 0:33. The average tag
cooccurrence ratio of u1 is 0:80, while that of u2 is 0:27. For
both measures, u1 scores higher than u2, this agrees with
the fact that the interests of u2 are more diverse as observed
from this user's bookmark collection.</p>
      <p>Next, we apply these two measures on the set of data
that we have collected from del.icio.us. The results are
summarised in Table 3 and Figure 3.</p>
      <p>Although the two measures are designed to measure
di®erent characteristics of personomies, the results do have very
common features. Firstly, the mean values of tag utilisation
and average tag co-occurrence ratio both very low, at 0.06
and 0.07 respectively, even though the values span across
the whole range from 0 to 1. These values mean that on
average a tag is only used on 6% of the bookmarks in a user's
collection, and that a tag is only used together with 7% of
other tags. We can see that there is a small group of points
in both graphs in Figure 3 which attain a value of 1. These
actually correspond to users who have only one bookmark
in their collection. Other than these the values drop quickly,
and the majority of personomies have values less than 0.2
(93% in both measures). Also, there is a strong correlation
between tag utilisation and average tag co-occurrence ratio,
with a correlation coe±cient of 0.71.</p>
      <p>Given these ¯gures, we reveal that for most users many
tags are used only on a small portion of their bookmarks, and
that these tags are not always used together. This suggests
that the bookmarks of the users have topics which are rather
diverse such that tags do not apply to all of them. Also,
a user's tags can be terms from di®erent domains which
are not used together very often on the bookmarks. Hence,
this indicates that users of del.icio.us have diverse interests
instead of a single interest in a very speci¯c domains.</p>
    </sec>
    <sec id="sec-7">
      <title>USER PROFILE CONSTRUCTION</title>
      <p>As the majority of users in del.icio.us are observed to be
interested in a wide range of topics from di®erent domains,
a user pro¯le in the form of a single set of tags is de¯nitely
inadequate. Hence, user pro¯les which can accommodate
the multiple interests of the users are very much desirable.</p>
      <p>Identifying the di®erent interests can be a challenging task
(a) Tag utilisation
(b) Average tag co-occurrence ratio
as tags are freely chosen by users and their actual meaning
is usually not very clear. A solution to this problem is to
exploit the associations between tags and documents in a
folksonomy. As it is obvious that documents related to the
same interest of a user would be tagged by similar tags, we
can perform clustering algorithms on the documents tagged
by a user to group documents of similar topics together,
and extract the sets of tags assigned to these documents as
indicators of the users' di®erent interests.</p>
      <p>Based on this idea, we propose a method for constructing
user pro¯les which involves constructing a network of
documents out of a personomy, applying community-discovery
algorithms to divide the nodes into clusters, and extracting
sets of tags which act as signatures of the clusters to re°ect
the interests of the users.
4.1</p>
    </sec>
    <sec id="sec-8">
      <title>Community Discovery Algorithms</title>
      <p>
        Clusters in a network are basically groups of nodes in
which nodes have more connections among each other than
with nodes in other clusters. The task of discovering clusters
of nodes in a network is usually referred to as the problem of
discovering community structures within networks [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
Approaches to this problem generally fall into one of the two
categories, namely agglomerative, which start from isolated
nodes and group nodes which are similar or close to each
other, and divisive, which operate by continuously dividing
the network into smaller clusters [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ].
      </p>
      <p>
        To quantitatively measure the `goodness' of the clusters
discovered, the measure of modularity [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] is usually used.
The modularity of a particular division of a network is
calculated based on the di®erences between the actual number of
edges within a community in the division and the expected
number of such edges if they were placed at random. Hence,
discovering the underlying community structure in a
network becomes a process to optimise the value of modularity
over all possible divisions of the network.
      </p>
      <p>
        Although modularity provides a quantitative method to
determine how good a certain division of a network is, brute
force search of the optimal value of modularity is not
always possible due to the complexity of the networks and the
large number of possible divisions. Several heuristics have
been proposed for optimizing modularity, these include
simulated annealing [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], and removing edges based on edge
betweenness [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. In addition, a faster agglomerative greedy
algorithm for optimizing modularity, in which edges which
contribute the most to the overall modularity are added one
after another, has been proposed [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. In this paper, we will
employ this fast greedy algorithm to perform clustering, as
it is e±cient and performs well on large networks.
4.2
      </p>
    </sec>
    <sec id="sec-9">
      <title>Construction of User Profiles</title>
      <p>Given a network of documents (which are bookmarks in
our case), we can apply the community-discovery algorithms
to obtain clusters of documents. As the di®erent clusters
should contain documents which are related to similar
topics, a cluster can be considered as corresponding to one of the
many interests of the user. A common way to represent user
interests is to construct a set of tags or a tag vector.
Similarly, we can obtain a set of most frequently used tags from
each of the document clusters to represent the
corresponding interest. As a summary of our method, the following list
describes the whole process of constructing a user pro¯le for
user u.</p>
      <p>1. Extract the personomy Pu of user u from the
folksonomy F, and construct the bipartite graph T Du.
2. Construct a one-mode network of documents out of
T Du, and perform modularity optimization over the
network of documents using the fast greedy algorithm.
3. For each of the clusters (communities) ci obtained in
the ¯nal division of the network, obtained a set Ki of
tags which appear on more than f % of the documents
in the cluster. The set of tags of a cluster is treated as
a signature of that cluster.
4. Finally, return a user pro¯le Pu in the form of a set of</p>
      <p>Ki0s: Pu = fKig.</p>
      <p>For the signatures of the clusters, one can include all the
tags which are used on the bookmarks in the cluster, or
include only the tags which are common to the bookmarks
in the cluster. However, the set of tags chosen for a
cluster will a®ect how accurate the pro¯le is in modelling the
user's interest. In general, for a large value of f only the
most common tags in the cluster will be included in the
signature, while a small value of f will include more tags in
the signature. We will investigate the problem of choosing a
K1
K2
K3
K4
K5
K1
K2
K3
K4</p>
      <sec id="sec-9-1">
        <title>User A</title>
        <p>webdesign, web2.0, tutorial, blog, css
linux, opensource, ubuntu, software
webhosting, ¯lesharing
grammar, english
digg, sharing, music, mp3</p>
      </sec>
      <sec id="sec-9-2">
        <title>User B</title>
        <p>webdesign, programming
interesting, art, video, funny
food, books, tobuy
lort, debate</p>
      </sec>
    </sec>
    <sec id="sec-10">
      <title>EVALUATION AND DISCUSSIONS</title>
      <p>From our data set, we select at random 1,000 users who
have over 100 bookmarks in their personomies. The
requirement of having at least 100 bookmarks is to ensure that there
are enough bookmarks for clustering so that clearer results
can be obtained. We apply our proposed method of
generating user pro¯les on these personomies, and obtain a set
of clusters of bookmarks and their signatures. We discover
that there are a substantial number of clusters with only one
bookmark. The bookmarks in these clusters are mostly not
assigned any tags. Hence, we exclude these single-bookmark
clusters in the following analysis. Figure 4 graphs the
number of clusters discovered for each of the personomies. On
average 15 clusters are discovered in each personomy.</p>
      <p>We believe that the use of multiple sets of tags in user
pro¯les should give a more accurate representation of the
interests of the users. Therefore we try to evaluate our
proposed method by asking the following question: are the sets
of tags accurate descriptions of the clusters of bookmarks
from which they are extracted? If this is the case, then the
user pro¯les should accurately represent the interests of the
users. In the following we present the evaluations which</p>
    </sec>
    <sec id="sec-11">
      <title>Precision and Recall Measures</title>
      <p>
        Our question concerns with the issue of whether the sets of
tags in the user pro¯le are accurate descriptions of the
bookmarks in the clusters. An appropriate method of evaluation
is to approach this question from an information retrieval
perspective. Given the signature of a cluster as a query,
can we retrieve all the bookmarks within that cluster and
avoid obtaining bookmarks in other clusters which are
irrelevant? In addition, how many tags should be included in
the signature in order to accurately described a cluster? To
answer such questions, we will employ the measures of
precision and recall [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] which are commonly used for evaluating
information retrieval systems.
      </p>
      <p>Precision and recall are two widely used measures for
evaluating performance of information retrieval. Precision
measures the fraction of documents in the retrieved set which
are relevant to the query, while recall measures the fraction
of relevant documents that the system is able to retrieve.</p>
      <p>To employ the precision-recall measures, we treat the
signatures of the clusters as queries, and use them to retrieve
bookmarks by comparing the tags assigned to them to those
in the queries. As for the representation of tags, we
employ a vector space model of information retrieval. In other
words, for each personomy, we construct a term vector ~e =
(e1; e2; :::; en) for each bookmark, with ei = 1 if the
bookmark is assigned the ith tag, and ei = 0 otherwise.
Similarly, the signature of a cluster is converted into a query in
the form of a term vector q~. The retrieval process is carried
out by calculating the cosine similarity between the query
vector and the bookmark vectors:</p>
      <p>
        Sim(q~; ~e) =
q~ ¢ ~e
jq~jj~ej
Those with similarity higher than a certain threshold t will
be retrieved (0 · t · 1). For a cluster c, let the set of
bookmarks in the cluster be Dc, and the set of bookmarks
retrieved by the signature of the cluster be Dx. The
precision and recall of the system on c are de¯ned as follows.
In addition, we also consider the F1 measure [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] which is a
combined measure of precision and recall.
      </p>
      <p>P recision(c) =</p>
      <p>Recall(c) =</p>
      <p>F1(c) =
jDx \ Dcj</p>
      <p>jDxj
jDx \ Dcj</p>
      <p>jDcj
2 £ P recision(c) £ Recall(c)</p>
      <p>P recision(c) + Recall(c)</p>
      <p>We calculated the three measures for the user pro¯les
generated from the 1,000 selected personomies. We control two
parameters in our evaluation, one is the value of f , the
percentage of bookmarks above which a tag is assigned to in a
cluster for it to be included in the signature, and the value
of t, the threshold of cosine similarity. The results are
presented in Figure 5.</p>
      <p>Figure 5(a) shows that for most values of similarity
threshold precision attains maximum for f in the range from 0:1
to 0:4, and thereafter it continues to decrease as f increases.
The result suggests that if only the most common tags are
included in the signatures, they will become less
representative as summaries of the clusters. This is probably due to
the fact that the most common tags are usually too general
(4)
(5)
(6)
(7)
(a) Precision</p>
      <p>(b) Recall
(c) F1 measure
and a query constructed from these tags will tend to retrieve
bookmarks from other clusters as well which are related to
a di®erent sub-topic under the common tags. On the other
hand, when one includes all the tags which appear in a
cluster (with f = 0%), the signature will include too many tags
such that it will not be similar to any of the signatures of
the bookmarks, leading again to a low precision.</p>
      <p>As for recall, we observe some di®erences for di®erent
values of similarity threshold. For small values of t (from 0.0 to
0.3), recall continues to decrease as f increases. However, for
larger values of t (from 0.4 to 1.0), recall ¯rst increases and
then decreases as t increases. This is probably due to the
reason that when the similarity threshold is low, the number of
tags in the cluster signature is less important as most of the
bookmarks will be retrieved even if their similarity with the
query is small. As f increases, fewer tags are included in the
signature and therefore it becomes more di±cult to retrieve
relevant bookmarks. On the other hand, when t becomes
higher, signatures which include all the tags in a cluster or
include only the most common tags are very dissimilar to
any of the bookmarks in the cluster, therefore recall attains
maximum somewhere between the two extremes.</p>
      <p>For common values of similarity threshold between t = 0:3
to t = 0:5, precision and recall attain maximum for values
of f between 0:1 and 0:2, with precision over 0:8 and recalls
over 0:7. F1 measures also attain maximum around these
values of t and f . This suggests that it is better to include
more tags in a cluster signature so as to make it speci¯c
enough for representing the topic of the cluster (and thus
the interest of the user represented by the cluster). Given
these results, we conclude that by choosing a suitable value
of f the tags extracted do constitute good descriptions of
the bookmarks within the clusters.
5.2</p>
    </sec>
    <sec id="sec-12">
      <title>Potential Applications</title>
      <p>Our proposed algorithm provides a new way for
constructing better user pro¯les based on the data available from
collaborative tagging. There are a number of areas in which
such algorithms can be applied to. We brie°y discuss two of
them in this section.</p>
      <p>Firstly, as the user pro¯les provide a summary of the
different interests of the users, it can be readily used to
facilitate the management and organization of personal Web
resources. For example, the sets of tags representing the
clusters of bookmarks can be used to facilitate navigation
and retrieval of a user's own bookmarks in del.icio.us. This
would be much more e±cient than navigating through the
bookmarks by a single tag.</p>
      <p>In addition, the user pro¯les can also be used to support
Web page recommendation systems. Currently, del.icio.us
provides various methods which allow users to keep track
of new bookmarks which they may ¯nd interesting, such
as subscribing to the RSS feed of a tag, or adding a user
of similar interests to one's network. However, there have
been no mechanisms which directly recommend interesting
bookmarks to the users. With the user pro¯les constructed
by our proposed method, recommendation systems will have
a better understanding of the interests of the users, and
be able to recommend more speci¯c bookmarks to users by
targeting a particular interest of the users.</p>
    </sec>
    <sec id="sec-13">
      <title>RELATED WORK</title>
      <p>
        User pro¯le representation and construction has been a
key research area in the context of personal information
agents and recommendation systems. The representation
of user pro¯les concerns with how user interests and
preferences are modelled in a structured way. Probably the
simplest form of user pro¯le is a term vector indicating which
terms are interested by the user. The weights in the
vector is usually determined by the tf-idf weighting scheme as
terms are extracted from documents interested by the user or
obtained by observing user behaviour [
        <xref ref-type="bibr" rid="ref12 ref2">2, 12</xref>
        ]. More
sophisticated representations such as the use of a weighted network
of n-grams [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] have also been proposed. However, a
single user pro¯le vector may not be enough when users have
multiple interests in diverse areas [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], and several projects
have employed multiple vectors to represent a user pro¯le.
For example, Pon et al. [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] use multiple pro¯le vectors to
represent user interests to assist recommendation of news
articles. In recent years, user-pro¯ling approaches utilizing
the knowledge contained in ontologies have been proposed.
In these approaches, a user pro¯le is represented in terms of
the concepts that the user is interested in an ontology. For
example, Middleton et al. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] propose two experimental
systems in which user pro¯les are represented in terms of a
research paper topic ontology. Similar approaches have also
been proposed to construct user pro¯les for assisting Web
searching [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] or enhancing recommendations from
collaborative ¯ltering systems [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        On the other hand, since the rise in popularity of
collaborative tagging systems, some studies have also focused on
generating user pro¯les from folksonomies. For example, in
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] a user pro¯le generator based on the annotations assigned
by the users to the documents is proposed. The user
pro¯le is represented in the form of a tag vector, which each
element in the vector indicating the number of times a tag
has been assigned to a document by the user. In [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], three
di®erent methods for constructing user pro¯les out of
folksonomy data have been proposed. The ¯rst and simplest
approach is to select the top k mostly used tags by a user
as his pro¯le. The second approach involves constructing a
weighted network of co-occurrence of tags and selecting the
top k pairs of tags which are connected by the edges with
largest weights. The third method is an adaptive approach
called the Add-A-Tag algorithm, which takes into account
the time-based nature of tagging by reducing the weights on
edges connecting two tags as time passes. In addition, [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]
discusses the issue of constructing a user pro¯le from a
folksonomy in the context of personalised Web search. In their
approach, a user pro¯le pu is represented in the form of a
weighted vector with m components (corresponding to the
m tags used by the user). The use of w d is to assign a weight
between 0 and 1 to each of the n documents. While these
attempts provide some possible methods for constructing
user pro¯les based on data in folksonomies, the possibility
of a user having multiple interests is not addressed in these
works.
7.
      </p>
    </sec>
    <sec id="sec-14">
      <title>CONCLUSIONS</title>
      <p>The emergence of collaborative tagging systems provide
valuable sources of information for understanding user
interests and constructing better user pro¯les. In this paper,
we investigated the characteristics of personomies extracted
from folksonomies, and observed that the majority of users
possess a wide range of interests, which cannot be modelled
by simple methods such as a single set of tags. A novel
method for constructing user pro¯les which take into
account the diversity of interests of the users is proposed. We
also evaluated the user pro¯les by looking at whether they
provide a good summary of the bookmarks of the users.</p>
      <p>This research work provides insight into how user
pro¯les of multiple interests can be constructed based on the
data collected from a folksonomy. From this point, we plan
to carry out further research work in two main directions.
Firstly, we will further investigate how the proposed method
can be improved. In our study, a user pro¯les constructed
treats every cluster of bookmarks and its signature as
corresponding to a distinctive interest of the user. However, it
may be true that two interests are related and are only
subtopics of a more general area. We will investigate if the
introduction of a hierarchical structure is desirable. Secondly, we
will attempt to evaluate our proposed method by applying
the user pro¯les on applications such as Web page
recommendation or personal resource management. We hope this
research will ultimately deliver useful algorithms and
applications which utilise the power of user-contributed metadata
in collaborative tagging systems.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Sarabjot</given-names>
            <surname>Singh</surname>
          </string-name>
          <string-name>
            <surname>Anand</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Patricia</given-names>
            <surname>Kearney</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Mary</given-names>
            <surname>Shapcott</surname>
          </string-name>
          .
          <article-title>Generating semantically enriched user pro¯les for web personalization</article-title>
          .
          <source>ACM Trans. Inter. Tech.</source>
          ,
          <volume>7</volume>
          (
          <issue>4</issue>
          ):
          <fpage>22</fpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Marko</given-names>
            <surname>Balabanovic</surname>
          </string-name>
          and
          <string-name>
            <given-names>Yoav</given-names>
            <surname>Shoham</surname>
          </string-name>
          .
          <article-title>Learning information retrieval agents: Experiments with automated web browsing</article-title>
          .
          <source>In Proceedings of the AAAI Spring Symposium on Information Gathering from Heterogenous</source>
          ,
          <source>Distributed Resources</source>
          , pages
          <volume>13</volume>
          {
          <fpage>18</fpage>
          ,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Robin</given-names>
            <surname>Burke</surname>
          </string-name>
          .
          <article-title>Hybrid recommender systems: Survey and experiments</article-title>
          .
          <source>User Modeling</source>
          and
          <string-name>
            <surname>User-Adapted</surname>
            <given-names>Interaction</given-names>
          </string-name>
          ,
          <volume>12</volume>
          (
          <issue>4</issue>
          ):
          <volume>331</volume>
          {
          <fpage>370</fpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Paul</given-names>
            <surname>Alexandru</surname>
          </string-name>
          <string-name>
            <surname>Chirita</surname>
          </string-name>
          , Andrei Damian, Wolfgang Nejdl, and
          <string-name>
            <given-names>Wolf</given-names>
            <surname>Siberski</surname>
          </string-name>
          .
          <article-title>Search strategies for scienti¯c collaboration networks</article-title>
          .
          <source>In P2PIR '05: Proceedings of the 2005 ACM workshop on</source>
          <article-title>Information retrieval in peer-to-peer networks</article-title>
          , pages
          <volume>33</volume>
          {
          <fpage>40</fpage>
          , New York, NY, USA,
          <year>2005</year>
          . ACM Press.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>JoÄrg</given-names>
            <surname>Diederich and Tereza Iofciu</surname>
          </string-name>
          .
          <article-title>Finding communities of practice from user pro¯les based on folksonomies</article-title>
          .
          <source>In Proceedings of the 1st International Workshop on Building Technology Enhanced Learning solutions for Communities of Practice</source>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Girvan</surname>
          </string-name>
          and
          <string-name>
            <given-names>M. E. J.</given-names>
            <surname>Newman</surname>
          </string-name>
          .
          <article-title>Community structure in social and biological networks</article-title>
          .
          <source>PROC.NATL.ACAD.SCI.USA</source>
          ,
          <volume>99</volume>
          :
          <fpage>7821</fpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Daniela</given-names>
            <surname>Godoy</surname>
          </string-name>
          and
          <string-name>
            <given-names>Analia</given-names>
            <surname>Amandi</surname>
          </string-name>
          .
          <article-title>User pro¯ling in personal information agents: a survey</article-title>
          .
          <source>Knowl. Eng. Rev.</source>
          ,
          <volume>20</volume>
          (
          <issue>4</issue>
          ):
          <volume>329</volume>
          {
          <fpage>361</fpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Scott</given-names>
            <surname>Golder</surname>
          </string-name>
          and
          <string-name>
            <given-names>Bernardo A.</given-names>
            <surname>Huberman</surname>
          </string-name>
          .
          <article-title>Usage patterns of collaborative tagging systems</article-title>
          .
          <source>Journal of Information Science</source>
          ,
          <volume>32</volume>
          (
          <issue>2</issue>
          ):
          <volume>198</volume>
          {
          <fpage>208</fpage>
          ,
          <string-name>
            <surname>April</surname>
          </string-name>
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Miha</given-names>
            <surname>Grcar</surname>
          </string-name>
          , Dunja Mladeni¶c, and
          <article-title>Marko Grobelnik. User pro¯ling for interest-focused browsing history</article-title>
          .
          <source>In SIKDD 2005 at Multiconference IS</source>
          <year>2005</year>
          , Ljubljana, Slovenia,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Roger</given-names>
            <surname>Guimera</surname>
          </string-name>
          and
          <string-name>
            <given-names>Luis A. Nunes</given-names>
            <surname>Amaral</surname>
          </string-name>
          .
          <article-title>Functional cartography of complex metabolic networks</article-title>
          .
          <source>Nature</source>
          ,
          <volume>433</volume>
          :
          <fpage>895</fpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Andreas</surname>
            <given-names>Hotho</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Robert</surname>
            <given-names>JaÄschke</given-names>
          </string-name>
          , Christoph Schmitz, and
          <string-name>
            <given-names>Gerd</given-names>
            <surname>Stumme</surname>
          </string-name>
          .
          <article-title>Information retrieval in folksonomies: Search and ranking</article-title>
          . In York Sure and John Domingue, editors,
          <source>The Semantic Web: Research and Applications</source>
          , volume
          <volume>4011</volume>
          <source>of LNCS</source>
          , pages
          <volume>411</volume>
          {
          <fpage>426</fpage>
          . Springer,
          <year>June 2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Henry</given-names>
            <surname>Lieberman</surname>
          </string-name>
          .
          <article-title>Letizia: An agent that assists web browsing</article-title>
          . In Chris S. Mellish, editor,
          <source>Proceedings of the Fourteenth International Joint Conference on Arti¯cial Intelligence (IJCAI-95)</source>
          , pages
          <fpage>924</fpage>
          {
          <fpage>929</fpage>
          , Montreal, Quebec, Canada,
          <year>1995</year>
          . Morgan Kaufmann publishers Inc.: San Mateo, CA, USA.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Elke</given-names>
            <surname>Michlmayr</surname>
          </string-name>
          and
          <string-name>
            <given-names>Steve</given-names>
            <surname>Cayzer</surname>
          </string-name>
          .
          <article-title>Learning user pro¯les from tagging data and leveraging them for personal(ized) information access</article-title>
          .
          <source>In Proceedings of the Workshop on Tagging and Metadata for Social Information Organization, 16th International World Wide Web Conference (WWW2007)</source>
          , May
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Stuart</surname>
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Middleton</surname>
          </string-name>
          ,
          <string-name>
            <surname>Nigel R. Shadbolt</surname>
          </string-name>
          , and
          <string-name>
            <surname>David C. De Roure</surname>
          </string-name>
          .
          <article-title>Ontological user pro¯ling in recommender systems</article-title>
          .
          <source>ACM Trans. Inf</source>
          . Syst.,
          <volume>22</volume>
          (
          <issue>1</issue>
          ):
          <volume>54</volume>
          {
          <fpage>88</fpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>Peter</given-names>
            <surname>Mika</surname>
          </string-name>
          .
          <article-title>Ontologies are us: A uni¯ed model of social networks and semantics</article-title>
          . Web Semant.,
          <volume>5</volume>
          (
          <issue>1</issue>
          ):5{
          <fpage>15</fpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>M. E. J.</given-names>
            <surname>Newman</surname>
          </string-name>
          .
          <article-title>Fast algorithm for detecting community structure in networks</article-title>
          .
          <source>Physical Review E</source>
          ,
          <volume>69</volume>
          :
          <fpage>066133</fpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>M. E. J.</given-names>
            <surname>Newman</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Girvan</surname>
          </string-name>
          .
          <article-title>Finding and evaluating community structure in networks</article-title>
          .
          <source>Physical Review E</source>
          ,
          <volume>69</volume>
          :
          <fpage>026113</fpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>Michael</given-names>
            <surname>Noll</surname>
          </string-name>
          and
          <string-name>
            <given-names>Christoph</given-names>
            <surname>Meinel</surname>
          </string-name>
          .
          <article-title>Web search personalization via social bookmarking and tagging</article-title>
          .
          <source>In Proceedings of the 6th International Semantic Web Conference and 2nd Asian Semantic Web Conference (ISWC/ASWC2007)</source>
          , Busan, South Korea, pages
          <volume>365</volume>
          {
          <fpage>378</fpage>
          ,
          <string-name>
            <surname>November</surname>
          </string-name>
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Raymond</surname>
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Pon</surname>
            ,
            <given-names>Alfonso F.</given-names>
          </string-name>
          <string-name>
            <surname>Cardenas</surname>
            , David Buttler,
            <given-names>and Terence</given-names>
          </string-name>
          <string-name>
            <surname>Critchlow</surname>
          </string-name>
          .
          <article-title>Tracking multiple topics for ¯nding interesting articles</article-title>
          .
          <source>In KDD '07: Proceedings of the 13th ACM SIGKDD international conference on Knowledge discovery and data mining</source>
          , pages
          <volume>560</volume>
          {
          <fpage>569</fpage>
          , New York, NY, USA,
          <year>2007</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Filippo</surname>
            <given-names>Radicchi</given-names>
          </string-name>
          , Claudio Castellano, Federico Cecconi, Vittorio Loreto, and Domenico Parisi. De¯
          <article-title>ning and identifying communities in networks</article-title>
          .
          <source>PROC.NATL.ACAD.SCI.USA</source>
          ,
          <volume>101</volume>
          :
          <fpage>2658</fpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>H.</given-names>
            <surname>Sorensen</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Mcelligot</surname>
          </string-name>
          .
          <article-title>Psun: A pro¯ling system for usenet news</article-title>
          .
          <source>In CKIM'95 Workshop on Intelligent Information Agents</source>
          ,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Kazunari</surname>
            <given-names>Sugiyama</given-names>
          </string-name>
          , Kenji Hatano, and
          <string-name>
            <given-names>Masatoshi</given-names>
            <surname>Yoshikawa</surname>
          </string-name>
          .
          <article-title>Adaptive web search based on user pro¯le constructed without any e®ort from users</article-title>
          .
          <source>In WWW '04: Proceedings of the 13th international conference on World Wide Web</source>
          , pages
          <volume>675</volume>
          {
          <fpage>684</fpage>
          , New York, NY, USA,
          <year>2004</year>
          . ACM Press.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>C. J. van Rijsbergen. Information</given-names>
            <surname>Retrieval</surname>
          </string-name>
          . Dept. of Computer Science, University of Glasgow,
          <year>1979</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Thomas</surname>
            <given-names>Vander</given-names>
          </string-name>
          <string-name>
            <surname>Wal</surname>
          </string-name>
          . Folksonomy de¯nition and wikipedia. http://www.vanderwal.net/random/ entrysel.php?blog=1750,
          <string-name>
            <surname>November</surname>
            <given-names>2</given-names>
          </string-name>
          ,
          <year>2005</year>
          . Accessed 13
          <string-name>
            <surname>Feb</surname>
          </string-name>
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>S.</given-names>
            <surname>Wasserman</surname>
          </string-name>
          and
          <string-name>
            <given-names>K.</given-names>
            <surname>Faust</surname>
          </string-name>
          .
          <article-title>Social network analysis</article-title>
          . Cambridge University Press, Cambridge,
          <year>1994</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Xujuan</surname>
            <given-names>Zhou</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sheng-Tang</surname>
            <given-names>Wu</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Yuefeng</given-names>
            <surname>Li</surname>
          </string-name>
          , Yue Xu,
          <string-name>
            <surname>Raymond Y. K. Lau</surname>
            , and
            <given-names>Peter D.</given-names>
          </string-name>
          <string-name>
            <surname>Bruza</surname>
          </string-name>
          .
          <article-title>Utilizing search intent in topic ontology-based user pro¯le for web mining</article-title>
          .
          <source>In WI '06: Proceedings of the 2006 IEEE/WIC/ACM International Conference on Web Intelligence</source>
          , pages
          <fpage>558</fpage>
          {
          <fpage>564</fpage>
          , Washington, DC, USA,
          <year>2006</year>
          . IEEE Computer Society.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>