<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Exploiting the Annotation Practice for Personal and Collective Information Management</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Guillaume Cabanac</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Max Chevalier</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Claude Chrisment</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christine Julien</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>IRIT-SIG, UMR 5505, Universite de Toulouse</institution>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>LGC</institution>
          ,
          <addr-line>EA 2043, Universite de Toulouse</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2008</year>
      </pub-date>
      <fpage>67</fpage>
      <lpage>78</lpage>
      <abstract>
        <p>Information nowadays is a capital for any organization intending to be reactive and aware of its environment. Unfortunately most modern organizations overdose on information as almost every member daily accesses, extracts and stores a growing amount of documents, i.e. vehicles for information. This situation is even deteriorating as individual e orts to organize and search for information yield poorly from the organization standpoint since di usion mechanisms are limited. We propose a personal and collective information management architecture in order to take advantage of individual e orts, and to manage documents in a collective and sustainable way. This is based on the integration of the document lifecycle activities depicting the way people manage information and documents. The proposed architecture exploits individual e orts through interdependent processes designed on a mutual bene t scheme. These processes rely on the annotation practice, considered as a representative evidence of the way individuals interact with information.</p>
      </abstract>
      <kwd-group>
        <kwd>Collective IS</kwd>
        <kwd>Annotation</kwd>
        <kwd>Information-related Activities</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Modern organizations such as companies, R&amp;D labs, or communities
increasingly rely on Information Systems (IS) as vehicles for the information relevant
to their activities. In the same time IS cause informational overdoses: the
organization is seldom capable of optimally handling the collected information as a
whole. This issue is twofold as it arises at individual level and a collective level.
Firstly, knowledge workers|people who mainly produce and work with
information in the workplace; representing 31% of the US workforce in 1995, their
proportion \will continue to increase signi cantly into the new millennium" [1,
p. 51]|have a hard time identifying, nding, and keeping relevant information
related to their activities. Secondly, they rarely distribute the information they
introduced into the organization, although co-workers would bene t from it since
one may assume that their needs are close, or even similar regarding their
activities. The cost of this twofold issue is valued at a minimum of $33,000,000 a year
for a hundred knowledge workers organization [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. A solution to these problems
should capitalize on collective information which remains static, unproductive,
and scattered in the maze of the organization otherwise. The purpose of this
paper is to illuminate the issues involved in collective activities related to
information and its medium, namely documents. These activities coming from the
document lifecycle are presented in Sect. 2 according to individual and collective
levels, thus covering the issues stated beforehand. Then Sect. 3 introduces our
proposition: an original architecture integrating and exploiting the document
lifecycle activities. It is based on the personal and collective annotation practice
as an inter-activity vehicle for information. The main idea is that annotating a
document re ects individuals' cognitive e orts (e.g. learning, arguing,
correcting) while interacting with documents. In addition, the architecture provides
knowledge workers with personalized assistance thanks to the processes de ned
in this section. Moreover, the results of an activity improves the performance
of the other ones, for the individual and for the group. Section 4 outlines
experiments related to the proposed information management architecture and
processes; it also describes the prototype system that implements this
architecture. The prototype is demonstrated with Web documents, as a common source
of information for any organization nowadays.
2
      </p>
      <p>Current Issues of Common Document-related</p>
      <p>
        Organizational Activities
Within an organization, managing collective documents properly is a
performance factor: it relies on the optimization of the various activities that facilitate
the access to documents, and by extension to the corresponding information.
The document lifecycle [1, p. 203] gives a comprehensive view of six major
document-related organizational activities|noted from À to Å|that
knowledge workers achieve individually or collectively, supported by the appropriate
software. Marshalling and Extracting Information À relies on \pull systems"
such as search engines and social bookmarking [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Creating, Authoring Á and
Finalizing Documents Â is powered by word processors with marking and
annotation capabilities. Distribution and Work Flow Ã may be performed
manually via emails, mailing-lists, and posts on the intranet or automatically when
exploiting work ows, recommender systems [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], and social networks [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Using
Documents Ä mainly refers to active reading, i.e. critical thinking supported by
informal annotations stemming from the common paper-based annotation
practice [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], e.g. remarks, comments, summaries. In the digital world, readers may
use a software called annotation system, such as Annotea [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Finally, Filing and
Archiving Documents Å consists in storing documents in a Personal
Information Space (PIS), mostly for nding them later, building a legacy, and sharing
them [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. People commonly classify documents into thematic folder hierarchies,
e.g. le system, email client, bookmark hierarchy. This latter feature enables
Web users to keep and organize interesting documents in a hierarchical
structure that quickly evolves as people add, on average, three to four bookmarks per
navigation session [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
      <p>The study of document-related activities À to Å reveals a plethora of existing
systems, may they be individually or collectively targeted. They intend to exploit
collective documents by taking into account the experience, the expertise, the
activities, the contacts, etc. of knowledge workers. Since individuals bene t from
the group's activities, and as the group reciprocally bene ts from individuals'
activities, such systems seem to be in line with a mutual pro t scheme. However,
we notice that each system is highly specialized: it considers only one or two|at
most|partitioned activities among the six activities of the document lifecycle
depicted in [1, p. 203]. Moreover, these activities seem to be linear, although we
think that a person does not carry on document-related activities that way. On
the contrary, people can obviously search for information À, start to author a
document Á, then go back browsing again À for going into the subject in depth.
As systems are specialized, each one only considers a small part of the real work
of users while neglecting the four or ve other activities. As a consequence, in
our view, systems do not capitalize enough on knowledge workers' daily
activities, thus leading to both limited sustainability and limited long term e ciency.
In order to overcome the aforementioned issues, this paper puts forward an
approach based on a federated and multiuser architecture. This intends to cover
the whole activities of the document lifecycle. Unlike the linear and partitioned
lifecycle proposed in [1, p. 203], we intend to integrate activities and to exploit
their outcome. The ultimate aim is to help each individual, which is in turn
bene cial to the group.
3</p>
      <p>Personal and Collective Annotation-based Information
Management
The study of document-related activities revealed two main issues. Firstly, the
document lifecycle currently involves too many systems: people need to master at
least six distinct applications, thus leading to cognitive overload. Secondly, each
system is highly specialized since it is designed for a unique activity. This design
results in scattered and partial user pro les, leading to poor adaptation and
support. To tackle these problems, we propose an original architecture providing:
{ Personal support. The architecture relies on a uni ed model that federates
users and their six document-related activities, thus avoiding information
and user pro les scattering. Along with dedicated processes described later
in this section, this federated architecture helps individuals to nd À, to
exploit Ä, to organize Å, to author ÁÂ, and to distribute Ã information.
This design actually enables each knowledge worker to build up his own
sustainable Personal Information Space (PIS) day by day.
{ Collective support. Our proposal is a multiuser architecture that models
knowledge workers within their organization. It exploits the constituted
capital (knowledge workers' PISs) by implementing automated processes based
on a mutual bene t scheme. The basic idea is that people implicitly
contribute while achieving their common activities. They bene t in return by
receiving information relevant to their work, which is automatically extracted
from the activities of the group. This approach makes the organization-wide
available information|extracted or contributed by the members|pro table
for the whole group, whereas it is too often unexploited in people's desk
drawers and computer le systems. As a result, our approach provides any
organization with a sustainable information management.</p>
      <p>The study of document-related activities in Sect. 2 showed that the
annotation practice is already part of three individual and collective activities:
authoring Á, nalizing Â and exploiting Ä digital documents. This motivated our
choice: placing the annotation practice at the heart of our architecture, as it
can also cover the three remaining activities. Next sections de ne the \collective
annotation" concept, and show how it also federates information retrieval À,
di usion Ã, and organization Å through dedicated processes depicted in Fig. 1.
s
e
c
r
u
o
S
n
o
i
t
a
m
r
o
f
n
I
s
u
o
i
r
a
V</p>
    </sec>
    <sec id="sec-2">
      <title>NAVIm</title>
    </sec>
    <sec id="sec-3">
      <title>SOCIALVALIDATIONm</title>
    </sec>
    <sec id="sec-4">
      <title>Navigation of Userm</title>
    </sec>
    <sec id="sec-5">
      <title>Navigation of Usern</title>
    </sec>
    <sec id="sec-6">
      <title>SOCIALVALIDATIONn</title>
    </sec>
    <sec id="sec-7">
      <title>NAVIn</title>
    </sec>
    <sec id="sec-8">
      <title>Userm</title>
    </sec>
    <sec id="sec-9">
      <title>RECO</title>
    </sec>
    <sec id="sec-10">
      <title>Usern</title>
    </sec>
    <sec id="sec-11">
      <title>PROTODOCm</title>
    </sec>
    <sec id="sec-12">
      <title>REORGm</title>
    </sec>
    <sec id="sec-13">
      <title>PISm</title>
      <p>...</p>
    </sec>
    <sec id="sec-14">
      <title>PISn</title>
    </sec>
    <sec id="sec-15">
      <title>REORGn</title>
    </sec>
    <sec id="sec-16">
      <title>PROTODOCn</title>
    </sec>
    <sec id="sec-17">
      <title>UNIFIED</title>
    </sec>
    <sec id="sec-18">
      <title>VIEW</title>
      <p>
        the organization
Creating a bookmark is a common way of keeping track of an interesting
document [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Many concrete work situations require keeping not only the document
but also the passages of interest (e.g. a few sentences, a de nition, a picture, a
schedule) along with notes. Though, creating a bookmark is rather unsuitable
for this task as it can neither point to parts of documents, nor keep any reader's
notes. To overcome these limits we opted for the \collective annotation" concept
as de ned in [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. The UML Class Diagram in Fig. 2 formalizes this fundamental
concept. It depicts a low level of detail as we intentionally hide the attribute and
operation compartments of the classes for brevity concerns. We shall explain
the class diagram using a concrete scenario where a User visualizes a Resource
such as a Web page, a schedule on the intranet of his company, a picture in his
le system. Whenever he wants to keep a piece of interesting information, he
creates an AnchoredAnnotation. Two situations may arise. On the one hand, if
the user wants to keep the entire resource, the system needs to store its location
with a GlobalAnchoring. This mimics a classical bookmark storing the URLs of
documents. On the other hand, the user may want to keep parts of a resource
only, e.g. two non contiguous sentences. To handle this case, an alternative is to
modify the resource by adding markers referencing the beginning and the end of
the user's selection. This only works when facing modi able resources, which is
a quite unusual case with public documents. That is why we opted for another
alternative which consists of storing an anchoring point (LocalAnchoring ) that
locates the user's selection unambiguously. To handle the various resource
formats, the LocalAnchoring abstract class must be re ned; for HTML documents
it can be an XPointer that expresses the selection path in the document object
model (DOM) for instance.
      </p>
      <p>Resource</p>
      <p>Comment 0..1 &lt; contains</p>
      <p>&lt; describes * Tag
1
location</p>
      <p>Type</p>
      <p>1 *
* describes &gt; * Annotation * &lt; creates 1 User
&lt; cites</p>
      <p>*
* *
Anchoring 1..* &lt; annotates * AnchoredAnnotation
*
Reply</p>
      <p>1
replies
*rep1lies
1</p>
      <p>ArgumentativeAnnotation
GlobalAnchoring</p>
      <p>BookmarkingAnnotation
LocalAnchoring
*</p>
      <p>StandardAnnotation
&lt; contains
sub folders
0..1 1
*
*
0..1
Folder</p>
      <p>
        Previous classes model objective data that an annotation system infers from
the document part selected by the user. The remaining classes model the
subjective information introduced by the user who creates an annotation. Any
Annotation may contain a Comment without restriction on the media or on the format,
e.g. rich text, an audio recording. Being a subclass of Resource, a Comment or any
part of it can be annotated in turn: this design allows the creation of recursive
annotations. In addition to a comment, the annotator (i.e. the user creating the
annotation) may indicate references to other (parts of) resources. Moreover, he
can associate such citations as well as the annotation itself with Types, thanks to
the \&lt; cites" and \describes &gt;" relationships. The Type abstract class describes
the annotator's intent by giving an overview of its meaning: subclasses may cover
taxonomies of objectives (e.g. comment, example, question), of actions (e.g. to
do, to read), of opinions (e.g. refutation, neutral, con rmation), or any
domainspeci c concepts (e.g. business intelligence: partner, product, competitor, etc.).
Providing organizational members with such taxonomies may help them to
describe encountered information with a common ground, so as to improve their
understandability. As opposed to assigning a prede ned Type coming from a
xed vocabulary, the Tag class enables users to describe an annotation with
their own words, as promoted by the social bookmarking approach, cf. Sect. 2.
When it comes to annotating a resource, we discern three main user objectives
that the proposed model takes into account: keeping and organizing information,
note-taking, and discussing. Firstly, a BookmarkingAnnotation enables knowledge
workers to keep and organize information. This kind of Annotation refers to the
common bookmarking practice [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] while allowing a ner-grained anchoring: it
is anchored to parts of a resource whereas a classical bookmark concerns the
entire resource. In order to get access to his kept information, each User owns a
PIS structured as classical bookmarks, i.e. a hierarchy of Folders. One reason to
provide a hierarchy comes from the study [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] which underlines that knowledge
workers' need to classify information into hierarchies \to get things done." The
second purpose of an Annotation is achieved by the StandardAnnotation class that
enables to take notes on a resource without necessarily requiring its classi
cation. Proofreading during activity Â generates many annotations (corrections,
misprints, etc.) that the annotator does not need to classify in his PIS; what is
really essential for him and his co-workers instead is to view these annotations
while re-reading the annotated document. Finally, the third purpose of
collective annotation is to discuss in the context of the documents. Such a debate is
initiated by an ArgumentativeAnnotation; later other readers may express their
standpoints by formulating a Reply to this annotation, or to Replies recursively,
thus forming a discussion thread similar to the Usenet ones.
      </p>
      <p>
        Summing up the architecture design, a User creates an AnchoredAnnotation
to keep encountered information. He may organize it the way he wants (e.g. by
topic, by project, by date) within his personal Folder hierarchy (PIS). To achieve
the automated processes mentioned earlier, we endow the proposed architecture
with the additional model depicted in Fig. 3 (note that the two models
complement each other; the latter is commented throughout this section). The proposed
architecture is designed along with the six processes depicted in Fig. 1.
Concerning the activity Å, the Reorg process aims to reduce the high cognitive load
involved when reorganizing a PIS [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. It takes as input the user's PIS for
suggesting him a thematic classi cation, then the user can accept it partially or
entirely. The proposed algorithm [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] is based on a Hierarchical Agglomerative
Clustering [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] which requires the Indexation of the Resource contents.
      </p>
      <p>
        Recommendation *
Index
The previous section depicted the twofold role of an annotation. For individual
activities its main purposes are to support critical thinking in-context (i.e. not on
a separate sheet) while reading documents, and to keep interesting information in
the reader's PIS. This is respectively achieved by StandardAnnotation and
BookmarkingAnnotation. Regarding collective activities, annotations can be shared so
as to support collaborative work: readers get previous readers' comments and
feedback, they can participate in in-context debates through discussion threads
as well (ArgumentativeAnnotation). One essential issue about annotation
systems in general concerns their scalability. When displayed within documents
at the exact place they were authored, they might disturb the reader.
Empirical evidence shows that the di culty in reading a document increases with the
number of displayed collective annotations. For instance, the video http://www.
irit.fr/~Guillaume.Cabanac/annotation/demoAmaya.wmv demonstrates this issue
with the W3C Annotea/Amaya annotation system [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. We propose two
complementary ways of reducing readers' e orts. As a rst adaptation to User needs,
our architecture hides any Comment that is not expressed in a Language they
chose (Fig. 3). This avoids displaying utterly incomprehensible information to
users. The second adaptation regards ArgumentativeAnnotations and the debates
they may spark o . When many debates are anchored to a given document,
the reader may consult each one in turn. Given a debate, deducing its
participants' global opinion mentally enables to evaluate the \social validity" of the
ArgumentativeAnnotation. Although necessary for critical judgment, this
evaluation requires cognitive e orts while rst identifying argument opinions, and then
synthesizing opinions recursively upwards in the discussion thread. We intend to
relieve readers of this burden thanks to the SocialValidation process that we
de ned in [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. It mimics individuals by synthesizing Reply opinions to obtain
the social validity of the ArgumentativeAnnotation. This value is gradual as it
ranges from \refuted" to \con rmed," it represents the global opinion expressed
in the considered discussion thread. Given this process, the second adaptation we
mentioned consists of informing users about the degree of consensus (resp.
controversy) of each debate. As a result, they can focus on stabilized information,
or on ongoing discussions where people have not found a consensus yet.
3.3
      </p>
      <sec id="sec-18-1">
        <title>Creating and Finalizing Documents Á Â</title>
        <p>
          Much of the explicit knowledge of an organization is held in the documents
it produces, since knowledge workers spend a great amount of time
authoring documents. A common task consists of extracting the essential facts from
several documents, so as to synthesize them in a report [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. Regarding the
proposed architecture, BookmarkingAnnotations are a sure way for a reader to
collect nuggets of information along with his Comments and interpretations. For
a speci c project (e.g. Stock Exchange daily analysis) the user may create a
dedicated Folder in his PIS gathering all the relevant annotations he wants to
keep. When dealing with collaborative search and analysis, the implicated
Entities (either Users or Groups of users) can create annotations and retrieve them
from a shared Folder provided they obtained the appropriate Grant. In order
to assist knowledge workers in harnessing the collectively collected information,
the ProtoDoc process drafts a document from any selected Folder of the PIS.
This proto-document encompasses each annotation, the contents of its Anchoring
within the annotated Resource, its social validity, and the information provided
by the annotator (Comment, citations, etc.). The user can use his favorite word
processor to complete and rework this draft afterwards.
3.4
        </p>
      </sec>
      <sec id="sec-18-2">
        <title>Improving Collective Information Retrieval À</title>
        <p>
          The proposed architecture empowers the two classical modalities at the heart
of information retrieval: searching and browsing. Regarding the search
modality, we proposed to complement it by taking into account collective annotations
in [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. The basic idea is to use readers' contributions, namely their
annotations as a \social feedback" to improve IR recall (by retrieving more documents
relevant to the query) as well as IR precision (by retrieving relevant documents
only). Concerning IR recall, \the vocabulary problem" states that a user's query
rarely (&lt; 20%) contains the same terms as the required documents [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. We
suggest matching the query with annotations in order to indirectly nd documents,
and even passages when dealing with contextual IR|aiming to answer a query
with passages instead of complete documents. Concerning IR precision, readers'
Comments allow the disambiguation of annotated documents, and the
integration of complementary terms in the indexing process. In addition, the social
validity of discussion threads sparked o by ArgumentativeAnnotations may lead
to characterize a Resource as trustworthy, controversial, popular, alive, outdated,
abandoned . . . Taking these indicators into account allows the adaptation of the
search engine to users' preferences. As regards to the browsing modality, the
Navi process recommends documents to the user, provided they are relevant to
his current navigation, see Fig. 1. This recommendation process is original in
many respects. Firstly, since recommendations come from each organizational
member's PIS, it exploits co-workers' ability to nd and classify interesting
documents. By doing so, long-time retrieved documents and then forgotten in
individuals' folders are automatically proposed to other people. We put forward the
hypotheses that i ) a document inserted in a folder was considered interesting by
his owner, since he achieved a cognitive e ort for selecting the most appropriate
folder in his PIS. Moreover, ii ) other knowledge workers that share similar
interests may be interested in such a document. A second original feature concerns
the algorithm that matches recommendable documents with the one that the
user is currently viewing. Classical methods are content-based: they recommend
documents according to their text only. Conversely we proposed usage-based
similarity metrics [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] to evaluate how closely knowledge workers use two given
documents. Provided that people store together the documents they nd similar
(for any reason: their topic, their author, etc.), we stated that the more two
documents get closely classi ed by the most people, the more they are usage-based
similar. The proposed metrics is a tree walking algorithm that accesses
knowledge workers' PISs. The Navi process exploits it to recommend usage-based
similar documents coming from co-workers PISs to each user during his
navigation stage. The user may also view recommendations on a map representing the
knowledge workers' documents organized according to the usage-based metrics,
thanks to the UnifiedView process detailed in [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]. This allows the exploration
and discovering of the capitalized documents that knowledge workers introduce
daily into the organization.
3.5
        </p>
        <p>
          Distributing Collective Information to Knowledge Workers Ã
The study [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] reports how a poor collective information di usion leads
organizations to a terri cally counterproductive and costly outcome: waste of time,
information recreation, etc. We propose to capitalize on the information kept by
knowledge workers, which often stays dormant in their computers otherwise. To
do that, we o er Users the capability to send manual recommendations to the
other Entities he knows. A dual feature enables Users to register for noti cations
concerning other Entities' documents. By doing so, one can proactively specify
whose documents shall interest him, akin to Web syndication via RSS feeds. As
underlined in Sect. 2 manual di usion is limited by various human factors: the
sender's social network, his willingness to share when information is commonly
perceived as power, the implied cognitive e orts . . . To overcome these limits,
we propose to complement manual di usion with automatic di usion thanks to
the Reco process as depicted in Fig. 1. It considers each document that enters
the organization (i.e. retrieved by any member) as a candidate one for
recommendation. As a result, it exploits collective information retrieval to help each
member. In a word, the Reco process fully explained in [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] works as follows.
Each candidate document is rst indexed, cf. the Term and Indexation classes.
Then its thematic similarity with Users' PIS Folders is evaluated: each folder
is represented as a classi er built by extracting features from its documents.
Finally the candidate document is recommended in the folder having the best
similarity value, provided that it exceeds a dynamic threshold. This threshold
ensures that recommendations don't overload knowledge workers.
        </p>
        <p>
          Ongoing Experiments and Prototype Development
As a rst step towards validating the proposed architecture, we experimented
with two processes among the six proposed ones in an organization of 14
researchers whose PISs contained 4,079 documents (resp. 486 folders) with an
average of 291 documents (resp. 34 folders) by user. Five users were asked to
evaluate the information recommended by the Navi process as they browsed
the Web, following a xed navigation. This experiment detailed in [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] showed
that the more a user browses the Web, the more the recommendations he gets
are accurate. We also experimented with the Reco process through the TREC
2001 OSHUMED/MeSH collection in order to compare di erent strategies for
recommending information in a folder hierarchy like a PIS. Finally, we are
currently experimenting with the SocialValidation process depicted in Sect. 3.1.
The purpose of this experimentation is to evaluate how close the proposed
algorithm [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] is to human perception of consensus in argumentative discussions.
We designed a protocol in compliance with the experimental psychology
standards for Internet-based experimenting [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. An online Java WebStart software
(cf. http://www.irit.fr/~Guillaume.Cabanac/expe) allows the participation of
worldwide volunteers informed via mailing lists, e.g. the ACM SIGCHI. This
experiment in progress launched in April 2007 has been raising 179 inscriptions,
118 of whom started and 51 nished the experiment. The contributed evaluations
are currently under study.
        </p>
        <p>
          A second ongoing work concerns the development of a prototype which
implements the personal and collective information management architecture
introduced in this paper. This proof of concept software called TafAnnote (cf. http:
//www.irit.fr/~Guillaume.Cabanac/TafAnnote) is designed on a two-tier model.
The client side is a toolbar for the Mozilla Firefox Web browser that gives access
to the supported features: annotation creation, discussion thread support, PIS
management and annotation search. The client side communicates through a
HTTP connection with the server side that stores annotations. Whenever a user
requires a Web page, the browser retrieves its contents and sends its URL to the
annotation server. Finally the client side merges each fetched annotation with the
retrieved document model (DOM): annotations are displayed in-context
according to their respective anchoring points. TafAnnote currently supports the HTML
format by storing anchoring points as XPointers. Annotating di erent document
formats is a challenge as speci c anchoring techniques, i.e. subclasses of the
LocalAnchoring are required. We addressed this problem in the decisional systems
context by de ning a dedicated anchoring technique [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] for the datawarehouse
resources, namely multidimensional schemata and tables.
5
        </p>
        <p>Conclusion and Future Works
This paper investigated how knowledge workers achieve the six activities of the
document lifecycle [1, p. 203]. We reviewed both individual and collective
prominent approaches and systems dealing with each activity. This state-of-the-art
study revealed two main issues. Firstly, the great diversity of systems implies
data scattering and cognitive overload for the user, who has to master six
systems in all, i.e. one per activity. Secondly, the activities are represented as linear
and partitioned, although people do not behave this way. As a matter of fact
each system is highly specialized in a unique activity, this leads to partial user
modeling and limited support by extension. In order to overcome these issues, we
proposed to federate the lifecycle activities into a uni ed multiuser architecture.
Its ultimate purpose is to support each user for his daily document-related tasks.
An original aspect of our proposal is that the organization is the real source of
support. Indeed we exploit each knowledge worker's Personal Information Space
(PIS) to help any user; such an individual assistance improves the performance
of every user, which in turn improves the organization itself as a whole. The
architecture we modeled is based on a key concept that we related to each activity:
collective annotation as an evidence of knowledge workers' intellectual work. In
addition we de ned six automated processes represented in Fig. 1. They help
each user to reorganize his PIS (Reorg), to evaluate the social validity of
argumentative annotations (SocialValidation), to draft a proto-document from
encountered nuggets of information kept thanks to annotations (ProtoDoc), to
discover collective documents relevant to his navigation (Navi) or to his longterm
interests (Reco), and to get a UnifiedView of the knowledge workers'
documents. This collective information management architecture is currently under
experiment as the SocialValidation process is the object of an Internet-based
experiment rallying worldwide participants.</p>
        <p>
          Perspectives for this work mainly concern further validating the proposed
architecture and processes. We shall use methods from the cognitive sciences to
investigate whether knowledge workers improve their e ciency in daily
activities, and evaluate the trade-o between added constraints and bene ts as
proposed in [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]. Less HCI-related issues concerning digital annotation must also be
addressed. One of them concerns the anchoring on various document formats,
another deals with \robust" anchoring on evolving resources (e.g. modi ed
documents), and a third one refers to scalability issues of the current client-server
architecture as well as annotation visualization.
        </p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Sellen</surname>
            ,
            <given-names>A.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harper</surname>
            ,
            <given-names>R.H.</given-names>
          </string-name>
          :
          <article-title>The Myth of the Paperless O ce</article-title>
          . MIT Press, Cambridge, USA (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Feldman</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>The high cost of not nding information</article-title>
          .
          <source>KM World magazine 13(3) (March</source>
          <year>2004</year>
          ) available online http://www.kmworld.com/Articles/ PrintArticle.aspx?ArticleID=
          <fpage>9534</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Cabanac</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chevalier</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chrisment</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Julien</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soule-Dupuy</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tchienehom</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          : Web Information Retrieval:
          <article-title>Towards Social Information Search Assistants</article-title>
          . In Kidd, T.,
          <string-name>
            <surname>Chen</surname>
          </string-name>
          , I., eds.
          <source>: Social Information Technology: Connecting Society and Cultural Issues. Information Science Reference (March</source>
          <year>2008</year>
          )
          <volume>217</volume>
          {
          <fpage>251</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Montaner</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lopez</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>de la Rosa</surname>
            ,
            <given-names>J.L.</given-names>
          </string-name>
          :
          <article-title>A Taxonomy of Recommender Agents on the Internet</article-title>
          .
          <source>Artif. Intell. Rev</source>
          .
          <volume>19</volume>
          (
          <issue>4</issue>
          ) (
          <year>2003</year>
          )
          <volume>285</volume>
          {
          <fpage>330</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Zhang</surname>
          </string-name>
          , J.,
          <string-name>
            <surname>Ackerman</surname>
            ,
            <given-names>M.S.</given-names>
          </string-name>
          :
          <article-title>Searching For Expertise in Social Networks: A Simulation of Potential Strategies</article-title>
          . In: Group'05: Proceedings of the international conference on Supporting group work, New York, NY, USA, ACM Press (
          <year>2005</year>
          )
          <volume>71</volume>
          {
          <fpage>80</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Marshall</surname>
            ,
            <given-names>C.C.</given-names>
          </string-name>
          :
          <article-title>Toward an ecology of hypertext annotation</article-title>
          .
          <source>In: Hypertext'98: Proceedings of the 9th conference on Hypertext and hypermedia</source>
          , New York, NY, USA, ACM Press (
          <year>1998</year>
          )
          <volume>40</volume>
          {
          <fpage>49</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Kahan</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koivunen</surname>
            ,
            <given-names>M.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prud'Hommeaux</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Swick</surname>
            ,
            <given-names>R.R.</given-names>
          </string-name>
          :
          <article-title>Annotea: an open RDF infrastructure for shared Web annotations</article-title>
          .
          <source>Comp. Netw</source>
          .
          <volume>32</volume>
          (
          <issue>5</issue>
          ) (
          <year>August 2002</year>
          )
          <volume>589</volume>
          {
          <fpage>608</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Kaye</surname>
            ,
            <given-names>J.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vertesi</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Avery</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dafoe</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>David</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Onaga</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosero</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pinch</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>To Have and to Hold: Exploring the Personal Archive</article-title>
          .
          <source>In: CHI'06: Proceedings of the conference on Human Factors in computing systems</source>
          , New York, NY, USA, ACM Press (
          <year>2006</year>
          )
          <volume>275</volume>
          {
          <fpage>284</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Abrams</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baecker</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chignell</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Information Archiving with Bookmarks: Personal Web Space Construction and Organization</article-title>
          .
          <source>In: CHI'98: Proceedings of the conference on Human factors in computing systems</source>
          , New York, NY, USA, ACM Press (
          <year>1998</year>
          )
          <volume>41</volume>
          {
          <fpage>48</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Cabanac</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chevalier</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chrisment</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Julien</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          : Collective Annotation:
          <article-title>Perspectives for Information Retrieval Improvement</article-title>
          .
          <source>In: RIAO'07: Proceedings of the 8th conference on Information Retrieval and its Applications</source>
          , CID (May
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Phuwanartnurak</surname>
            ,
            <given-names>A.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gill</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bruce</surname>
          </string-name>
          , H.:
          <article-title>Don't Take My Folders Away!: Organizing Personal Information to Get Things Done</article-title>
          . In: CHI'
          <article-title>05 extended abstracts on Human factors in computing systems</article-title>
          , New York, NY, USA, ACM Press (
          <year>2005</year>
          )
          <volume>1505</volume>
          {
          <fpage>1508</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Chevalier</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chrisment</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Julien</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Helping People Searching the Web: Towards an Adaptive and a Social System</article-title>
          .
          <source>In: ICWI'04: Proceedings of the 3rd International Conference WWW/Internet</source>
          , IADIS (
          <year>2004</year>
          )
          <volume>405</volume>
          {
          <fpage>412</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Jardine</surname>
          </string-name>
          , N.,
          <string-name>
            <surname>van Rijsbergen</surname>
            ,
            <given-names>C.J.:</given-names>
          </string-name>
          <article-title>The Use of Hierarchic Clustering in Information Retrieval</article-title>
          .
          <source>Information Storage and Retrieval</source>
          <volume>7</volume>
          (
          <issue>5</issue>
          ) (
          <year>1971</year>
          )
          <volume>217</volume>
          {
          <fpage>240</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Furnas</surname>
            ,
            <given-names>G.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Landauer</surname>
            ,
            <given-names>T.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gomez</surname>
            ,
            <given-names>L.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dumais</surname>
          </string-name>
          , S.T.:
          <article-title>The Vocabulary Problem in Human-System Communication</article-title>
          .
          <source>Commun. ACM</source>
          <volume>30</volume>
          (
          <issue>11</issue>
          ) (
          <year>1987</year>
          )
          <volume>964</volume>
          {
          <fpage>971</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Cabanac</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chevalier</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chrisment</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Julien</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>An Original Usage-based Metrics for Building a Uni ed View of Corporate Documents</article-title>
          . In Wagner, R.,
          <string-name>
            <surname>Revell</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pernul</surname>
          </string-name>
          , G., eds.
          <source>: DEXA'07: Proceedings of the 18th International Conference on Database and Expert Systems Applications</source>
          . Volume
          <volume>4653</volume>
          of LNCS., Springer (
          <year>September 2007</year>
          )
          <volume>202</volume>
          {
          <fpage>212</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Reips</surname>
          </string-name>
          , U.D.:
          <article-title>Standards for Internet-Based Experimenting</article-title>
          .
          <source>Experimental Psychology</source>
          <volume>49</volume>
          (
          <issue>4</issue>
          ) (
          <year>2002</year>
          )
          <volume>243</volume>
          {
          <fpage>256</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Cabanac</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chevalier</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ravat</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Teste</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>An Annotation Management System for Multidimensional Databases</article-title>
          . In Song, I.Y.,
          <string-name>
            <surname>Eder</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nguyen</surname>
          </string-name>
          , T.M., eds.:
          <source>DaWaK'07: Proceedings of the 9th International Conference on Data Warehousing and Knowledge Discovery</source>
          . Volume
          <volume>4654</volume>
          of LNCS., Springer (
          <year>September 2007</year>
          )
          <volume>89</volume>
          {
          <fpage>98</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Millen</surname>
            ,
            <given-names>D.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fontaine</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          :
          <article-title>Improving Individual and Organizational Performance through Communities of Practice</article-title>
          . In: GROUP'03: Proceedings of the international conference on Supporting group work, New York, NY, USA, ACM Press (
          <year>2003</year>
          )
          <volume>205</volume>
          {
          <fpage>211</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>