=Paper= {{Paper |id=Vol-101/paper-11 |storemode=property |title=Egocentric Search Method for Authoring Support in Semantic Weblog |pdfUrl=https://ceur-ws.org/Vol-101/Ikki_Ohmukai-et-al.pdf |volume=Vol-101 }} ==Egocentric Search Method for Authoring Support in Semantic Weblog== https://ceur-ws.org/Vol-101/Ikki_Ohmukai-et-al.pdf
          Egocentric Search Method for Authoring Support in
                          Semantic Weblog

                        Ikki Ohmukai                                     Kosuke Numa                          Hideaki Takeda
               The Graduate University for                     Yokohama National University                  National Institute of
                   Advanced Studies                            79-1 Tokiwadai, Hodogaya-ku,                      Informatics
                   2-1-2 Hitotsubashi,                                Yokohama-shi                           2-1-2 Hitotsubashi,
                       Chiyoda-ku                                    Kanagawa, Japan                             Chiyoda-ku
                      Tokyo, Japan                                                                              Tokyo, Japan
                                                                  d02hc038@ynu.ac.jp
                   i2k@grad.nii.ac.jp                                                                        takeda@nii.ac.jp


ABSTRACT                                                                             mission for development of technologies, i.e., technologies
In this paper we propose egocentric search methods based                             just for us. Shneiderman pointed out that we should shift
on the concept of ”Information and Communicate Activities                            our vision from ”old computing” to ”new computing”. He
Navigation (ICAN)” and an authoring support system for                               explained it in his recent book as follows[7]; ”The old com-
Weblog (blog). ICAN regulates the human activities from a                            puting was about what computers could do; the new com-
viewpoint of information and communication support. We                               puting is about what users can do. Successful technologies
introduce the idea of ”Collect” and ”Relate” in the ICAN                             are those that are in harmony with users’ needs. They must
table into the information retrieval and the search method                           support relationships and activities that enrich the users’
which uses contents and human relationship produced by                               experiences.”
daily blogging. Our egocentric methods provide more sub-                                We should shift our focus from information and commu-
jective search result than the conventional engines. We ap-                          nication technologies (ICT) to information and communica-
ply the methods to improve the quality of the small contents                         tion activities (ICA). We should investigate what are human
made with Weblog tools.                                                              activities on information and communication and how we
                                                                                     can assist people in these activities.

Keywords
Social network, egocentric search, Weblog                                            1.2    Information and Communication Activi-
                                                                                            ties
1.     INTRODUCTION                                                                     Human activities on information such as collecting infor-
                                                                                     mation and communication such as contacting to people are
1.1      From ICT to ICA                                                             only a part of human activities but they become to play an
   Computers and networks enrich and facilitate our life so                          important role more and more in modern life.
that they now become indispensable for our life. They some-                             They include various kinds of activities. Shneiderman
times enhance our traditional daily activities with their in-                        shows a simple and therefore understandable model called
creasing computing and networking power like documenting                             ART (Activities and Relationships Table) for them[7]. One
and communicating with other people, and sometime of-                                axis of the table is activity category, i.e., Collect (informa-
fer new ways for our activities with new technologies like                           tion), Relate (Communication), Create (Innovation), and
WWW.                                                                                 Donate (Dissemination). The other is category of relation-
   On the other hand, most people become to live with worry                          ship, i.e., Self, Family and friends, Colleagues and neighbors,
that unceasing improvement of computers and networks and                             and Citizens and markets. We agree with relationship cat-
installation of new software technologies would change their                         egories, while we think that activity categories should be
life and business.                                                                   elaborated more because information handling and commu-
   It is not because of such technologies themselves but be-                         nication among people are mixed.
cause of our vision to technologies. We are so eager to de-                             To explicate the difference, we propose two-layered model
velop new technologies that we almost loose the original                             as an extension of his model shown in Fig.1. The first layer
                                                                                     has three elements that concern information handling, i.e.,
                                                                                     Collect (information)
Permission to make digital or hard copies of all or part of this work for            Create (information)
personal or classroom use is granted without fee provided that copies are            and
not made or distributed for profit or commercial advantage and that copies           Donate (information).
bear this notice and the full citation on the first page. To copy otherwise, to         It shows user-centered view of life cycle of information. In-
republish, to post on servers or to redistribute to lists, requires prior specific   formation is collected, then new information is created based
permission and/or a fee.                                                             on the collected information, and finally created information
K-CAP2003 Knowledge Markup & Semantic Annotation Workshop ’03
Sanibel Island, Florida USA                                                          is donated to the society for future creation. It should be
Copyright 200X ACM X-XXXXX-XX-X/XX/XX ...$5.00.                                      noted that new information is seldom created from scratch
but created based on existing information.1                        ics is the introductions and comments of the web sites that
  The second layer has also three elements that concerns           include from news sites to the other small contents.
communication handling, i.e.,                                         Some Weblog sites attract the attention with their own
Relate (people)                                                    editorial policy. The authors of Weblog sites reedit the ex-
Collaborate (with people)                                          isting web contents by quoting them. Moreover there are
and                                                                new types of Weblogs that criticize the other Weblogs so
Present (people).                                                  that these Weblogs are regarded to organize the ”Weblog
                                                                   community”. There are more than 100,000 Weblogs in the
                                                                   United States Weblogs make people to change from infor-
                                                                   mation receiver into information sender and distributor.
                                                                      Most of Weblog site uses the contents management system
                                                                   (CMS) called Weblog tool. Weblog tools enable the author
                                                                   to describe and edit the small contents via a web browser
                                                                   and transform the contents form text format to HTML files.
                                                                   These tools are implemented based on MVC (Model / View
                                                                   / Controller) model which is the fundamental concept of web
                                                                   applications. The author defines a view template once then
                                                                   do not have to decorate the contents with various HTML
                                                                   tags. This model decreases the cost of publication remark-
                                                                   ably comparing with traditional style which requires local
Figure 1: Information and Communication Activi-                    text editor and FTP. This feature contributes abundant pro-
ties                                                               duction of the small contents. Fig.2 shows typical site with
                                                                   Weblog tool.

   It is communication-centered view of the above process.
People establish relationship with other people, then col-
laborate with them to create new information, and finally
present themselves as donor of new information. Having
both information and communication layers is not redun-
dant. What we refer as ”information” in the context of com-
puter technologies is stored data in computers, while human
is the source of ”information” in the broader sense, i.e., hu-
man can offer information dynamically. We should consider
communication in order to include the function ”human as
information source”. This parallel view of information and
communication activities has thus six categories as activi-
ties. Ideally all categories should be supported by comput-
ers. Some categories like Collect is well investigated, but
others are not. In particular, the three categories in the
communication layer should be investigated more.
   We aim to investigate information and communication ac-
tivities and support people in the all categories of the activi-               Figure 2: Typical Weblog Site
ties. We call such support ”information and communication
activity navigation (ICAN)”. It helps people to create new
information by guiding information space and human net-              A huge number of the small contents and citations among
work.                                                              Weblog communities are increasing day by day. Some efforts
                                                                   such as topic discovery, trend analysis and content ranking
                                                                   are applied to these large amount of information.
2.    WEBLOG AND SEMANTIC WEB                                        Weblog facilitates to publish the small contents, however,
                                                                   the cost of contents arrangement and classification remains
2.1    Weblog and Small Contents                                   extremely high. Most of Weblog tools are specialized to en-
   Recently Weblog (blog) or blogging has come into the            hance the convenience of publishing by transforming from
spotlight in the World Wide Web[3]. There is no strict defi-       text to normal HTML files. There are already billions of
nition about Weblog but it is recognized as a web site which       HTML files on the Internet so that people are facing trou-
consists of miscellaneous notes updated daily[1]. In such          bles both to discover her/his objective articles and to use
sites the authors do not make efforts to knit up these con-        information effectively. Therefore it is skeptically consid-
tents and just align them in chronological order. We call          ered that Weblog will just accelerate this trend.
these frequently-posted contents as small contents in this         2.2 Semantic Web
paper. Small contents include various subjects including
journal, expertise and critique. One of most popular top-            There are great hopes that the Semantic Web technolo-
                                                                   gies will resolve our current condition of information over-
1
 We do not claim that information creation is just combi-          load. According to the manifest[2], the Semantic Web is an
nation of existing information. Rather creativity arises with      environment, which consists of the contents with machine-
understanding and interpretation of existing information.          readable (semantic) tags and the software agents, to realize
autonomous information distribution and syndication. Re-         the authoring content and polish the content out with the
source Description Framework (RDF)[13] and other ontol-          search results. Iteration of these processes may improve the
ogy definition languages[14] are recommended by W3C as           quality of each content in Weblogs.
elemental technologies of the Semantic Web and these are
now in practical use.
   However it is difficult to produce contents with semantic     4. IMPLEMENTATION
tags because of their complicated syntax and vocabulary.
Ordinary people hardly find a merit of semantic annotation       4.1 System Architecture
because it is a time-consuming task. It is also impossible         We implemented authoring support system for Weblogs
to annotate the semantic tags to existing enormous infor-        with our proposed method as shown in Fig.3. The system
mation on the Internet. There are some researches about          consists of three modules as follows.
automatic annotation with AI techniques and natural lan-
guage processing[4] however their effects are still unclear.
   In this research we aim to integrate both technologies,          • Weblog tool
Weblog and Semantic Web, to achieve the platform which                We use ready-made Weblog tool as an infrastructure
enables to share, reuse and reedit our small contents. We             of our system. In a number of Weblog tools released
provide a new function to the contents management sys-                recently we introduce Movable Type system[8] which
tems like Weblog tools and allow a semantic annotation to             is one of the most popular tools. Movable Type can
existing contents semiautomatically. Hereby it is possible            communicate CGI (Common Gateway Interface) pro-
to apply the effects of the Semantic Web to all of the con-           grams via MetaWeblog API[11] based on XML-RPC
tents on the Internet. As a result, links in Weblogs are              protocol[10].
transformed into semantic annotations so that the Weblog
contents and the web contents refered by Weblogs can be
                                                                    • Editor
worked as Semantic Web.
   In this paper we propose the egocentric search methods             We developed an Editor interface as Web application.
with relational annotation and description support system             It can connect the Weblog tools and the Cache Database
for Weblogging as a first stop of our project.                        and execute the egocentric search methods. We will
                                                                      explain it in detail in following section.
3.   CONCEPT OF EGOCENTRIC SEARCH
   As mentioned above Weblog tools contribute to increase           • Cache Database (DB)
the amount of small contents. However these tools do not              Cache DB stores all contents which is linked in the
help improve the quality of contents. There is fear that flood        user’s Weblog.
of ”junk” contents sweeps the Internet as a result.
   We explore new ways of authoring support for blogging.
Most of the contents on Weblogs are closely related to the
other contents from different sites. Therefore it is important
for Weblog authors to know similar or related contents to
currently describing contents.
   We illustrate the procedures of authoring support accord-
ing to the ICAN table discussed before2 . We assume the
user usually refers and comments on several contents on the
Internet in her/his Weblog site. These activities associate
the user’s contents with the other contents. We define these
activities as ”Collect” of information. We also consider these
linkages not only as relationships between each contents but
relations between the authors who have these contents. This
fact corresponds to ”Relate” in the ICAN table.
   Thus the contents and human network are built around
the user with these Collect and Relate activities.
   In case of describing a new content, that is the ”Create”
activity, the authoring support system will retrieve the con-
tents and human network close to the new content. The
closeness among the small contents is calculated not as the
score of semantic similarity but as the distance from the
user on her/his network. We call these search methods as
”Egocentric search”3 . The user will add a new link onto
2
                                                                       Figure 3: Snapshot of Proposed System
  Donate and Present are achieved as the original functions
of Weblog that are one of most excellent functions among
other information publishing tools.
3
  The word ”egocentric” is borrowed from Social Network
Analysis[12]. In Social Network Analysis, sociocentric and       network and the latter is focused in network of individuals.
egocentric network analysis provide two distinctive views for    Search engines like google shares the same view with the
network where the former concerns the nature of the whole        former, and our approach with the latter.
4.2    Extension of RSS
   Most of Weblog tools generate a RSS (RDF Site Sum-
mary4 )[5][6] file automatically. RSS is an XML-based meta-
data format for describing an abstract of Web pages. Basic
elements of RSS are shown as follows:

   • 
       element contains information of entire Web
      site such as name of Web site and author.

   • 
       element describes metadata of a content in a
      Web site. In a single Weblog site there are multiple
      contents (entries) so that their titles and update times
      are described in the item elements correspond to them.

   The RSS, which is the unified regular description format
for Web sites, is now propagating from Weblogs to enter-
prise sites so that people can incorporate various contents
into her/his Weblog using the RSS called ”Content Syndi-
cation”. However the RSS is for describing content relation
in single site, not for inter-site relationship. Therefore we
extend the concept of the RSS to describe metadata like
inter-site relation. We call this metadata ”RDF Content
Summary (RCS)” that is annotated to every entry in We-
blog. The RCS uses following modules in addition to the
elements of the RSS.

   • 
       module represents an ordinary hy-
      perlink. ”semblog” indicates the XML namespace we             Figure 4: Example of RDF Content Summary
      originally defined. Instance URI of the element is ex-
      tracted from  tag in a HTML document.

   •                                             with these link information. This network indicates not only
       module shows the URI which is pro-        relations of contents but also human relationships because
      vided as a reverse link (also called ”TrackBack”[9]) by    all entries on Weblog are owned by an author. Each path of
      several Weblog tools. For example, the author of We-       the human networks is weighted relatively to the frequency
      blog B publishes an entry 1 and pings to the entry         of citation.
      X in Weblog A, then the system of Weblog A recog-             Once the user cites some site as a topic in the new con-
      nizes this message and appends the URI of entry 1 to       tent (entry), the Editor program performs three types of
      the entry X. This type of reverse link is regarded as a    egocentric search and shows the result.
      metadata or an annotation of entry X.
                                                                    • Relative Chain Search
   As just described, the RCS maintains both link informa-
tion of related contents and metadata of entry itself. Fur-           Relative chain search returns the contents which is di-
thermore the Movable Type system can generate equivalent              rectly linked with the entry cited by the authoring con-
RCSs for each entry simply with template. Fig.4 shows an              tent. This model is based on a simple model but con-
example of RCS file.                                                  sequently it seems most trustful. (Fig.5(a))

4.3    Search Methods                                               • Relative Co-citation Search
   In this section we explain the search methods in the cir-          Relative co-citation search discovers the entries that
cumstance described previously.                                       link same contents as the authoring entry links to.
   The users daily write and post the small contents to their         Co-citation entries are retrieved from the Cache DB
Weblog sites with the Editor program. The Editor scans the            and the search result contains the weight of authors.
text strings of these contents each times they are posted.            (Fig.5(b))
If content contains a hyperlink, the Editor acquires whole
content and RCS of the link and store it in the Cache DB.           • Relative Keyword Search
Then the Editor extracts hyperlinks from stored contents
                                                                      Relative keyword search picks up the entries by key-
and constructs a entry network around the user’s contents
                                                                      word matching from the Cache DB. Different from the
4
  RSS is also a acronym of ”Rich Site Summary” or ”Really             conventional search engines, our method targets only
Simple Syndication”.                                                  related sites around the user’s Weblog. (Fig.5(c))
  The user read search results by these methods and can          method based on a pure P2P model, which does not depend
append the link of some helpful contents to describing con-      on the cache DB. Future model may create a foothold of the
tents. This process may enrich the user’s content and change     Semantic Web for ”the rest of us”. We will take an experi-
the search result of the system.                                 mental proof of our system with large Weblog communities
                                                                 in the near future.

                                                                 6. REFERENCES
                                                                  [1] E. Aimeur, G. Brassard, and S. Paquet. Using
                                                                      Personal Knowledge Publishing to Facilitate Sharing
                                                                      Across Communities. Workshop on (Virtual)
                                                                      Community Informatics, Held in conjunction with the
                                                                      Twelfth International World Wide Web Conference
                                                                      (WWW2003), 2003.
                                                                  [2] T. Berners-Lee. A roadmap to the Semantic Web.
                                                                      http://www.w3.org/DesignIssues/Semantic.html,
                                                                      1998.
                                                                  [3] R. Blood. We’ve Got Blog: How Weblogs are
                                                                      Changing Our Culture. Perseus Publishing, 2002.
                                                                  [4] S. Dill, N. Eiron, D. Gibson, and et al. SemTag and
                                                                      Seeker: Bootstrapping the Semantic Web via
                                                                      Automated Semantic Annotation. Proceedings of the
                                                                      Twelfth International World Wide Web Conference
                                                                      (WWW2003), 2003.
                                                                  [5] B. Hammersley. Content Syndication with RSS.
                                                                      O’Reilly & Associates, 2003.
                                                                  [6] RDF Site Summary 1.0 Specification Working Group.
                                                                      RDF Site Summary (RSS) 1.0.
                                                                      http://web.resource.org/rss/1.0/spec, 2001.
                                                                  [7] B. Shneiderman. Leonardo’s Laptop: Human Needs
                                                                      and the New Computing Technologies. MIT Press,
                                                                      2002.
                                                                  [8] Six Apart. Movable Type.
                                                                      http://www.movabletype.org/, 2003.
                                                                  [9] B. Trott and M. Trott. TrackBack Technical
                                                                      Specification. http:
                                                                      //www.movabletype.org/docs/mttrackback.html,
                                                                      2002.
                                                                 [10] UserLand Software. XML-RPC Specification.
                                                                      http://www.xmlrpc.com/spec, 1999.
                                                                 [11] UserLand Software. MetaWeblog API.
                                                                      http://www.xmlrpc.com/metaWeblogApi, 2002.
                                                                 [12] B. Wellman. An Egocentric Network Tale: Comment
        Figure 5: Egocentric Search Methods                           on Bien et al. Social Networks, 15:423–436, 1993.
                                                                 [13] World Wide Web Consortium (W3C). Resource
                                                                      Description Framework (RDF) Model and Syntax
                                                                      Specification.
5.   CONCLUSIONS                                                      http://www.w3.org/TR/REC-rdf-syntax, 1999.
   In this paper we propose egocentric search methods based      [14] World Wide Web Consortium (W3C). OWL Web
on the concept of ”Information and Communicate Activities             Ontology Language Overview.
Navigation (ICAN)” and an authoring support system for                http://www.w3.org/TR/owl-features/, 2003.
Weblog (blog). ICAN regulates the human activities from a
viewpoint of information and communication support. We
introduce the idea of ”Collect” and ”Relate” in the ICAN
table into the information retrieval and the search method
which uses contents and human relationship produced by
daily blogging. Our egocentric methods provide more sub-
jective search result than the conventional engines. We ap-
ply the methods to improve the quality of the small contents
made with Weblog tools.
   We will develop the metadata format and the extensions
of RSS moreover to represent the relationships among the
contents explicitly. We will also provide an egocentric search