<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Making Sense of Users' Web Activity</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Knowledge Media Institute, The Open University</institution>
          ,
          <addr-line>Milton Keynes</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>1. William Jones and Jaime Teevan (editors), Personal Information Management, University of Washington Press, 2007</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Personal information management (PIM), as described by [1], is \the practice
and study of the activities people perform to acquire, organise, maintain, retrieve,
use, and control distribution of information items". More and more services rely
on the Web to communicate with their users. The way users can control the
distribution of personal information exchanged daily through various Web channels
therefore appears as a crucial task for PIM. However, while the de nition above
clearly covers such activities, PIM has traditionally been focusing more on the
aspects of supporting information organisation and integration for the purpose
retrieval. Indeed, the types of personal information mentioned in [1] include
elements such as \information about a person but kept by and under the control
of others", but ignore one of the most di cult type of information to manage:
information about a person which is being shared and exposed to others.</p>
      <p>The related issues not only concern the ways to monitor, store and retrieve
this speci c type of information, but also the ways for users to make sense of
the huge amounts of information they are exchanging on the Web, knowingly
or unknowingly. Indeed, as a rst building block in this area, we developed a
tool dedicated to tracking the activity of an individual user on the Web. In
practice, this tool takes the form of a `local proxy' intercepting and storing
(using Semantic Web standards) the HTTP tra c on the user's computer. At
a higher level, we can see this tool as a `Web Li elogger', dedicated to the
undiscriminating collection of information concerning the user's online activity.
While relatively basic in principle, experimenting with this tool over a period
of time generates huge amounts of data (100 Million Triples for a single user in
2.5 months) which, when studied, allows us to unveil interesting, and sometimes
surprising aspects of the users Web life.</p>
      <p>The use of semantic technologies o ers the right level of exibility for the
management of such large, heterogeneous data, but more importantly, provides
us with the data integration and modelling approaches necessary to making
sense of the data. For example, mapping the collected semantic logs with a
representation of the user pro le allows us to construct models of the perceived
trust the user gives to various websites regarding the handling of his/her personal
information, and of the sensitivity of this information. Going a step further, by
applying di erent ontologies over the data, and linking it to the Web of Data,
we can build di erent perspectives on the traces of Web activity produced by
the user, providing as many \interpretations" of the user's interaction with the
Web, in addition to tools supporting him/her in managing this interaction.</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>