<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Explicit vs. Implicit Tagging for User Modeling</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Enrique Frias-Martinez</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Manuel Cebrian</string-name>
          <email>manuelc@tid.es</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>J. Moises Pascual</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nuria Oliver</string-name>
          <email>nuriao@tid.es</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>D</institution>
          ,
          <addr-line>Emilio Vargas 6, 28043, Madrid</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Data Mining and User Modeling Group, Telefonica Research Emilio Vargas 6</institution>
          ,
          <addr-line>28043, Madrid</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Telefonica I</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Tagging has been popularized by Web 2.0 sites as a way to describe resources. Typically, tagging has been done in an explicit way in which users directly describe with tags the resources in which they are interested. However in today's ubiquitous computing environments, it is possible to implicitly tag resources. This paper: (1) introduces the concept of explicit and implicit tagging for user models in two domains: Web 2.0 and mobile phone usage, respectively; and (2) compares the characteristics of both tagging mechanisms. Results indicate that the use of tags in both approaches is very similar, whereas the statistical characteristics of the common-interest networks are different.</p>
      </abstract>
      <kwd-group>
        <kwd>Tagging</kwd>
        <kwd>User Modeling</kwd>
        <kwd>Cell Phone Usage</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Tagging has become a de facto method for assigning a set of descriptors (or
keywords) to Internet digital content. The use of keywords for describing content was
already in use before the Internet. However, with the advent of Web 2.0 technologies,
a collaborative dimension was added. The main characteristics of a collaborative
tagging system are [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]: (1) the free-nature of the tags, (2) tags are bottom-up
nonhierarchical classifications and (3) there are no specific rules on how to annotate the
resources and annotations are not necessarily done by experts. These characteristics
are different from traditional classification hierarchies (taxonomies) where there are a
limited number of tags which define a top-down hierarchy.
      </p>
      <p>Most Web 2.0 tagging systems, such as Flickr (www.flickr.com) or Bibsonomy
(www.bibsonomy.com), can be characterized as explicit or user-guided tagging, due
to the involvement of the users in describing the resources. These systems, however,
suffer from some drawbacks: (1) the semantics given to a tag is not necessarily the
same for all users; (2) users may not be sure about what they are tagging; (3) users
may tag different resources with the same set of tags to save time; etc. These
limitations could be partially solved by implicit tagging systems. The concept of
implicit or automatic tagging consists of assigning tags to a given resource without the
intervention of a user. In today’s ubiquitous computing environments, there are a lot
of sources of information that can be used to automatically tag resources, including:
geographical information, mobility and communication patterns, information about
resources explicitly tagged, etc.</p>
      <p>
        In this paper, we present two tagging systems, one explicit (Bibsonomy) and the
other implicit (derived from mobile data usage) and compare their statistical
properties. The goal is to test to which extent what it is known about explicit tagging
environments can be applied to implicit tagging systems. The concept of explicit vs.
implicit tagging is to some extent inherited from the adaptive vs. adaptable [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] or
implicit vs. explicit [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] concepts used in user modeling.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2 User Model Generation from Explicit &amp; Implicit Tagging</title>
      <p>In this section, we present the two tagging environments used in our analysis. The
explicit tagging system used is Bibsonomy. The implicit tagging system is
constructed using the calling behavior of mobile phone users to businesses.</p>
      <sec id="sec-2-1">
        <title>2.1 Bibsonomy</title>
        <p>
          Bibsonomy [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] is a social bookmarking system in which users describe the
resources added to their shared personal library by means of tags. The data considered
for this study is freely available at [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. We use the TAS (Tagging Association) file
from [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] which contains 816,197 entries. Each entry consists of a user ID, one tag,
and a resource (bookmark or publication) tagged by that user. The TAS file contains
only non-spammers. No tag semantics were used in the experiments. Each user
typically has multiple entries for a given resource, one for each tag introduced. The
final dataset used in our experiments contains 2,467 unique users who assigned
69,902 unique tags to 268,692 resources. The set of tags explicitly introduced by a
user are considered the user model that describes the interest of that particular user.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2 Cell Phone Usage (SME)</title>
        <p>Data from a major cell phone carrier was obtained for a number of users close to
3,000. The original data set only contained the originating phone number and the
destination phone number, both encrypted. In this context, the resource being tagged
is each user and the tagging mechanism assigns tags that describe his/her interests. In
order to automatically assign tags to users, an encrypted directory of businesses and
services was considered. The directory contains a set of predefined tags for each
encrypted business phone number: (1) the type of business (e.g. hotel, restaurant, car
dealer, etc.); (2) a categorization of the business – only in the case of some
businesses, (e.g. number of stars in the case of hotels, price range in the case of
restaurants) and (3) an identifier of the town where the business is located.</p>
        <p>The tagging interest model for each user was obtained by only considering the
phone calls made by the user to the phone numbers included in the directory. A
composite tag was generated for each one of these calls by concatenating the set of
predefined tags for that particular business. In our study, we considered 57 business
types, 3 of which also had a categorization (hotels, restaurants and academies), and 51
towns. From a total of 3366 possibilities after tag concatenation, 2044 unique tags
were generated. We refer to this system as SME for “small and medium enterprises”
because the directory of business used presented these types of companies.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Comparative Analysis</title>
      <p>This section compares the statistical characteristics of both tagging systems. The
comparative analysis focuses on: (1) tag frequency; (2) total number of tags per user
and (3) characteristics of the common-interest social networks created.</p>
      <sec id="sec-3-1">
        <title>3.1 Tag Frequency</title>
        <p>
          The number of unique tags used in both systems is significantly very different: 69902
tags in Bibsonomy and 2044 tags in SME. Figure 1 (left) presents in a log-log
representation the survival function of the tag frequency for Bibsonomy (bib,
Bibsonomy, in circles) and SME (SME, in squares). The x-axis represents the
probability of each tag, ordered by decreasing probability, while the y-axis represents
the tag frequency. The software used was Clauset’s et al. [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] algorithm for power
law fitting and the MATLAB statistical toolbox for the lognormal and exponential fit.
The best fit was identified in terms of root mean squared error. Both tagging systems
show a similar power law distribution with similar slope values (1.75 –explicit-- and
1.8 --implicit--), which create almost two parallel distributions. The head of both
distributions represents tags that are heavily used, such as “software” in the case of
Bibsonomy or “Hotel4Stars28079” in the case of SME. Similar statistical behavior
has been shown in other explicit tagging communities [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ][
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. Preferential attachment
behavior, which is typical in explicit systems, seems to hold true also in our implicit
system, i.e. there are a few core businesses that receive the core of the calls.
        </p>
        <p>Intuitively, one of the biggest advantages of an explicit tagging system over an
implicit one is the free nature of the vocabulary. Nevertheless, our statistical results
indicate that the use of the tags is fairly similar in both cases. The results also
highlight a well known behavior in explicit tagging systems: users may have an
unlimited number of tags at their disposal. However, they end up using a reduced
number of tags. Therefore and in the context of this experiment, the tagging behavior
in the explicit system is similar to that of an implicit tagging system</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2 Total Number of Tags per User</title>
        <p>
          Figure 1 (right) presents in a log-log scale the survival function of the total number of
tags used by each user in the explicit (squares) and implicit (circles) tagging systems.
The total number of tags is given by the total number of tags introduced by a user
(including repetitions) in Bibsonomy and by the total number of calls in SME. In this
case, while SME has a straight forward power law distribution with α=2.93, the
explicit system is better modeled by two distributions: a lognormal distribution
(µ =2.87, σ=2.14) for its head, and a power-law distribution with α=2.13 for its tail. In
the literature of explicit tagging systems, the tail of the distribution of tags per
resource is usually approximated with a power law fitting [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ][
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. This difference in
the behavior of both tagging systems is probably not relevant, as the power law
distribution of SME is somewhat defined by the size of the data collected.
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3 Characteristics of the Interest-based Social Network</title>
        <p>Next, we build a network of users that share common interests as measured by the
number of common tags that are either introduced (Bibsonomy) or inferred (SME).
The network is composed of two types of nodes: (1) user nodes and (2) interest nodes.
The structure of the network is created by linking the user nodes with the interest
nodes, as specified by each user model. Two elements can be studied to characterize
the structure of each network: the number of unique tags per user and the number of
unique users per tag. Figure 2 (left) depicts the number of unique tags per user for
Bibsonomy (circles) and SME (squares). The x-axis represents the probability of the
number of different tags per user while the y-axis presents the number of tags.
Bibsonomy follows a power law distribution (α=3.44) while SME is better fitted by
an exponential distribution with coefficient 1.67. Figure 2 (right) displays the survival
function of the number of unique users per tag for Bibsonomy (circles) and SME
(squares). In this case, both systems are better fitted by a power law distribution with
α=2.37 (Bibsonomy) and α=1.83 (SME).</p>
        <p>
          Power law distributions have been observed in a great variety of networks,
including the WWW [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] and citation graphs [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. These observations seem to hold true
in a network based on common interests. Our results also indicate that the networks
built from implicit and explicit tagging systems have a similar architecture, with the
difference that the role of user nodes in the explicit social network is played by the
interest nodes in the implicit network and vice versa. This inversion of roles is
probably related to the different number of interests represented in each system.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusions &amp; Future Work</title>
      <p>In this paper, we have compared two tagging mechanisms: an explicit system
provided by Bibsonomy and an implicit one generated from cell phone usage
behavior. Our results show that the statistical tag frequency of the explicit and
implicit tagging systems is similar. While users in Bibsonomy may have the
opportunity to use any tag, they end up using a reduced set of tags. This behavior is
probably due to users having a limited set of interests that can be described by a
reduced number of tags. The interest networks built from explicit and implicit tagging
display different statistical behavior in the degree distribution of user and interest
nodes. This difference might be caused by the limited number of interest tags used in
the implicit case.</p>
      <p>
        In future work, we plan to: (a) test if the statistical behavior of the tagging systems
is better characterized by a Double Pareto log normal distribution [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]; (b) study how
the size of the tagging vocabulary affects its statistical properties; (c) explore how
well our results generalize to other domains and (d) develop mobile applications that
include the user models learned from the implicit system.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Golder</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>The structure of collaborative tagging systems</article-title>
          .
          <source>J. Inf. Science</source>
          (
          <year>2006</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>2. http://www.kde.cs.uni-assel.de/ws/rsdc08/.</mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Jäschke</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hotho</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schmitz</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stumme</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>Analysis of the Publication Sharing Behaviour in BibSonomy, Proc. Knowledge Architectures for Smart Applications (</article-title>
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Sigurbjörnsson</surname>
          </string-name>
          , B.,
          <string-name>
            <surname>van</surname>
            <given-names>Zwol</given-names>
          </string-name>
          ,
          <string-name>
            <surname>R.</surname>
          </string-name>
          :
          <source>Flicker Tag Recommendation based on Collective Knowledge</source>
          ,
          <string-name>
            <surname>WWW</surname>
          </string-name>
          <year>2008</year>
          (
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Angelova</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lipczak</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Milios</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pralat</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Characterizing a social bookmarking and tagging network</article-title>
          ,
          <source>Mining Social Data Workshop</source>
          , 18th Europ. Conf.
          <source>AI</source>
          (
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Seshadri</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Machiraju</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sridharan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bolot</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Faloutsos</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leskovec</surname>
          </string-name>
          , J.:
          <source>Mobile Call Graphs: Beyond Power-Law and Lognormal Distributions</source>
          ,
          <fpage>596</fpage>
          -
          <lpage>604</lpage>
          ,
          <string-name>
            <surname>KDD</surname>
          </string-name>
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Broder</surname>
            ,
            <given-names>A.Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kumar</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maghoul</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Raghavan</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rajagopalan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stata</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tomkins</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wiener</surname>
          </string-name>
          , J.:
          <article-title>Graph Structure in the web: experiment and models</article-title>
          ,
          <source>WWW</source>
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Redner</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>How popular is your paper? An empirical study of citation distribution</article-title>
          .
          <source>European Physics Journal B</source>
          (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Brusilovsky</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schwarz</surname>
            <given-names>E.</given-names>
          </string-name>
          <article-title>User as Student: Towards an Adaptive Interface for Advanced Web-Based Applications</article-title>
          .
          <source>User Modeling: Proc.6th Int. Conf. UM97</source>
          ,
          <fpage>177</fpage>
          -
          <lpage>188</lpage>
          (
          <year>1997</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Quiroga</surname>
            ,
            <given-names>L.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mostafa</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <article-title>Empirical evaluation of explicit versus implicit acquisition of user profiles</article-title>
          .
          <source>In Proc.fourth ACM conference on Digital libraries</source>
          ,
          <fpage>238</fpage>
          -
          <lpage>239</lpage>
          (
          <year>1999</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Cluset</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rohilla</surname>
          </string-name>
          , C, in press.
          <article-title>Power-law distributions in empirical data</article-title>
          .
          <source>SIAM Review.</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>