<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Using Named Entities as a basis to share associative trails between Semantic Desktops1</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Pat Croke</string-name>
          <email>Pat.Croke@hp.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ann Johnston</string-name>
          <email>Ann.Johnston@hp.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kim Tighe</string-name>
          <email>Kim.Tighe@hp.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Hewlett-Packard Galway Ltd</institution>
          ,
          <addr-line>Ballybrit Business Park, Galway</addr-line>
          ,
          <country country="IE">Ireland</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2005</year>
      </pub-date>
      <abstract>
        <p>In this paper we illustrate the use of Named Entities as a basis to share associative trails [1] between Semantic Desktops using a Firefox [6] extension called Trailblazer. Trailblazer can automatically detect the context of a Web page. It identifies the most relevant category for the Web page based on its context along with relevant links to additional information and notes. Named Entities belonging to that category are highlighted. Trailblazer facilitates the building and augmentation of categories with links from Named Entities to Web pages, documents, mail messages, personal notes, etc. It enables the import/export of categories and their related associative trails between users. It facilitates the distribution of discovered information through the sharing of associative trails in order to support community based knowledge. This allows faster initiation of new members to a group and information noted by one member as important can be highlighted to the rest of the group.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
    </sec>
    <sec id="sec-2">
      <title>Scenario</title>
      <p>We describe Trailblazer in terms of how a user, Ted, might experience it for the task of collecting
information on horse racing. Then we extend the scenario further to include how he shares his collected
information with other enthusiasts. He also uses Trailblazer to create a separate category of trails about
wine to share with his friends.</p>
      <p>Ted enjoys horse racing and to further his interest joins a horse racing club. The club consists of a group
of knowledgeable racing enthusiasts who share their knowledge between themselves. They maintain this
knowledge using a Firefox extension called Trailblazer which allows them to identify Horses, Trainers,
Jockeys, Racecourses, etc. as Named Entities. They add associative trails to these Named Entities in the
form of links to relevant Web pages and notes that they write. When Ted joins the club he is given
Trailblazer and access to the club’s associative trails. The Firefox browser has been altered to include the
Trailblazer panel at the bottom of the screen as seen in Figure 1. He follows the club’s instructions to
import their horse racing trail and visits his favorite online racing news Web site, ‘Irish Racing’. He can see
from the Trailblazer panel that the page has been categorized as ‘Horse Racing’ as shown in Figure 1.</p>
      <sec id="sec-2-1">
        <title>Using Named Entities as a basis to share associative trails between Semantic Desktops 3</title>
        <p>As he is studying the information Ted recalls how he overheard a Trainer at the Cork Racecourse
talking about ‘Chardonnay’ and how the horse is recovering from a cough and will not be back to full
fitness for the race on Thursday. He adds this information as a note to the Named Entity ‘Chardonnay’ as
shown in Figure 3.</p>
        <p>Trailblazer enabled Ted to import the horse racing associative trail. It also allows him to export a copy
of the horse racing associative trail with his added information back to the club’s Web site. Ted’s
information on Chardonnay’s cough is now available for sharing with the other club members. Other
members of the club can now import this horse racing trail from the shared account and merge the
additional information gathered by Ted with their horse racing trail. The local racing enthusiasts at the club
can now make an informed decision with regards to the horse Chardonnay’s chances of winning his race.</p>
        <p>Ted is also a keen wine connoisseur and has many friends who share this hobby. Having seen how
powerful Trailblazer is for sharing horse racing information he decides to use Trailblazer to create
associative trails for wine. He creates a wine category and goes to a Web page about wine to show
Trailblazer the words and phrases which typically occur on a page about wine. This is used to train
Trailblazer as shown in Figure 4.</p>
        <p>Ted continues to train Trailblazer about wine using a number of typical wine pages. He then adds a
number of Named Entity types to his wine category, for example ‘Grape Type’ and ‘Region’. When he
selects a Named Entity on a wine page, for example ‘Chardonnay’, he adds it to the category as a ‘Grape
Type’. He continues to add a number of Named Entities with bookmarks and notes from his own
knowledge of wine as shown in Figure 5.</p>
        <p>Using his new category, when Ted opens a Web page about wine and presses the ‘Find Entities’ button
any of the Named Entities that are on the page will get highlighted. When he passes a mouse over the
entity, a tool tip informs him of its type. When he clicks on an entity for example ‘Chardonnay’, the
associated bookmarks and notes are displayed as shown in Figure 6.</p>
      </sec>
      <sec id="sec-2-2">
        <title>Using Named Entities as a basis to share associative trails between Semantic Desktops 5</title>
        <p>Ted links the selected Named Entities with bookmarks and notes information. He adds his first note on
the best foods to pair with Chardonnay. His second note is a reminder that his girlfriend’s favorite wine is
Chablis which is made from Chardonnay grapes. However, Ted does not want to share this information
with anyone so he marks it as a private note to himself and therefore only he can see the note about his
girlfriend as shown in Figure 7.</p>
        <p>His selected Named Entities will form the basis for his wine trail. He will then browse several other
wine Web sites from which he will follow the same steps to collect more information about wine. When
trained Trailblazer will then be able to automatically detect if a Web page applies to wine and will load that
domain specific category with its additional information.</p>
        <p>Although Ted has the Named Entity ‘Chardonnay’ as an instance of ‘Horse’ in his ‘Horse Racing’
category, he can also add ‘Chardonnay’ to his ‘Wine’ category as an instance of ‘Grape Type’. Trailblazer
differentiates between the two as shown in Figures 8 and 9.</p>
        <p>Trailblazer categorizes the context of the page in Figure 8 as a horse racing Web page so the Named
Entity ‘Chardonnay’ is recognized as a horse. In Figure 9, the Web page has a context of wine so using the
context of this page ‘Chardonnay’ is a type of grape. Ted can now share his associative trails with his wine
connoisseur friends and together they can build and enhance their knowledge of wine.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Architecture</title>
      <p>
        Trailblazer is an extension for the Mozilla Firefox [
        <xref ref-type="bibr" rid="ref5">9</xref>
        ] Web browser. Extensions are applications that can
be downloaded and installed into a Firefox browser to add new functionality [6]. They can add anything
from a toolbar button to a completely new feature. Extensions allow Firefox to be customized to fit the
personal needs of each user if they require additional features.
      </p>
      <p>
        Interfaces for Mozilla Firefox extensions are built using the XML User-interface Language (XUL) [10]
and Cascading Style Sheets (CSS). Extensions are coded in JavaScript [12] and can link to native
components using the cross-platform Component Object Model (XPCOM) [
        <xref ref-type="bibr" rid="ref6">11</xref>
        ]. JavaScript links the XUL
and XPCOM components together. Firefox uses RDF as its data storage format.
      </p>
      <p>JavaScript is an open, platform-independent, event-driven, interpreted programming language. RDF is
an open, general-purpose data format for representing information in the Web. It is a common framework
therefore application designers can leverage the availability of common RDF parsers and processing tools.
This ability to exchange information between different applications means that the information may be
made available to applications other than those for which it was originally created. XUL (pronounced zool,
rhymes with cool) is an open, cross-platform language designed specifically for building portable user
interfaces. A XUL overlay is a XUL file containing elements to be inserted into another XUL file. This
insertion occurs when the other XUL file is rendered into an application interface. The browser window of
Firefox is a XUL window. Extensions can be made to the browser window using an overlay file. XUL is
used to overlay the browser to add the Trailblazer features. The Trailblazer overlay is rendered into the
Firefox browser, generating the main Trailblazer panel as part of the browser window.</p>
      <sec id="sec-3-1">
        <title>Using Named Entities as a basis to share associative trails between Semantic Desktops 7</title>
        <p>When Trailblazer is installed and the user launches Firefox, the standard browser window opens with the
addition of a Trailblazer panel. In order to facilitate the addition of Named Entities and their bookmarks
and notes, extra windows created using XUL dialog files are popped up by Trailblazer to request
information from the user.</p>
        <p>
          Firefox allows an installed extension to have additional privileges, such as reading local files and
modifying user preferences. The Input/Output component of Trailblazer uses another Firefox extension
called JSLib [
          <xref ref-type="bibr" rid="ref7">13</xref>
          ], which contains a library of JavaScript functions to access RDF files. This provides many
types of error checking, as well as a friendly abstraction away from RDF/XML interfaces.
        </p>
        <sec id="sec-3-1-1">
          <title>Page Recognition</title>
          <p>The Trailblazer Vocabulary store contains the information to calculate the most likely match for the
context of the current Web page. Each Category stored within Trailblazer contains the Named Entity types,
Named Entities, and their bookmarks and notes related to that category.</p>
          <p>When a user chooses the ‘Train’ option on the currently viewed page, the Trainer component extracts
the text from the page to supplement the information held against the current category within the
Vocabulary. Trailblazer uses a Vector Space Model (VSM) to model the collected training information.</p>
          <p>The VSM is an algebraic model used for information filtering and information retrieval [14]. It
represents natural language documents in a formal manner by the use of vectors in a multi-dimensional
space. Each dimension represents a term or concept found in the documents. The model creates a space in
which documents are represented by vectors and allows decisions to be made about which documents are
similar to each other. Spatial proximity is the physical distance between two items in space, for example the
distance between two vectors or lines. Semantic proximity is the difference between two items in meaning,
for example the context of the ‘Irish Racing’ Web page and the context of the ‘Wine Connoisseur’ Web
page. The VSM uses spatial proximity to calculate semantic proximity. The Trainer uses the words and
their frequency of occurrence from the current page to augment the existing training data.</p>
          <p>When a page is loaded into the browser a comparison is made between the text on the page and the
training data contained in the Trailblazer Vocabulary. A similarity measure is used to compare a vector
representation of each category Trailblazer has had training on, and a vector representation of the currently
viewed page. This calculates the most likely match, therefore deriving the context of the page in relation to
the information obtained from training. Currently, Trailblazer is using the Cosine Similarity Measure [7] to
compare the vectors of information. This is derived from the Law of Cosines, which is an extension of the
Pythagorean Theorem.</p>
          <p>The Cosine Similarity Measure is the cosine (in radians) of the angle between two vectors. For two
vectors a and b the cosine similarity between a and b is given by:
a ⋅ b
a × b
i.e. the dot product of a and b divided by the modulus of a multiplied by the modulus of b.</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>Named Entity Recognition (NER)</title>
          <p>
            When a user presses the ‘Find Entities’ button, the NER functionality processes the currently viewed
Web page and highlights any discovered Named Entities from the previously derived category. NER [
            <xref ref-type="bibr" rid="ref8">15</xref>
            ] is
a subtask of information extraction that seeks to locate and classify the elements in text into predefined
categories such as the names of people, organizations, locations, expressions of times, quantities, monetary
values, percentages, etc. The Named Entities which have been stored for the category are grouped together
on a word length basis. Each word from the Web page is processed in turn. A check is made to determine if
the word could be part of a Named Entity. If it is, an attempt is made to match it by comparing it to the
Named Entities beginning with the longest occurrences first. If it is possible for more than one match to be
made, the longest match will always be used. For example take the two Named Entities ‘Ascot’ and
‘Alisha’s Ascot’. ‘Ascot’ is the name of a racetrack and ‘Alisha’s Ascot’ is the name of a horse. When the
text ‘Ascot’ is detected in a page, an attempt is always made to match the Named Entity ‘Alisha’s Ascot’
first. ‘Ascot’ will only be annotated as a racetrack when the word ‘Alisha’s’ does not occur before it. For
efficiency the NER routine is a simple gazetteer-based process. The rare event of two different Named
Entities having the same name and same type within the same context is managed by adding a comment to
the Named Entity. The Trailblazer recognition process appears to be instant and does not delay the user’s
browsing experience. Heavier natural language processing to resolve this rare type of conflict would affect
this performance. When a Named Entity is discovered, span tags are used to add a highlight to the html
page. The page is reloaded with no change to its structure, only the highlight on the text is added. When the
user clicks on a highlighted Named Entity, its associated bookmarks and notes are displayed in the
Trailblazer panel.
          </p>
        </sec>
        <sec id="sec-3-1-3">
          <title>Export Process</title>
          <p>Information held in Trailblazer can be exported on a category by category basis. Export files are named
using the category name and an extension of EXP. These are generated in an RDF format. Each export file
stores the words and their frequency of occurrence in relation to the export category. It also stores Named
Entity types and Named Entities along with their related bookmarks and notes. Private notes are not
exported.</p>
          <p>Using Named Entities as a basis to share associative trails between Semantic Desktops 9
&lt;/RDF:RDF&gt;</p>
        </sec>
        <sec id="sec-3-1-4">
          <title>Import Process</title>
          <p>The import facility allows the user to import a file. Named Entities and their associated information are
extracted from the file and appended into the users existing category file. Named Entity duplicates are not
added, but their additional bookmarks and notes are merged. If a category file doesn’t exist a new one is
created. The words and their weights for that category are then extracted from the export file and integrated
into the user’s vocabulary file. This ensures that not only the annotation information is transferred between
users, but also the training information to allow Trailblazer to detect the context of Web pages.</p>
          <p>
            Annotea [
            <xref ref-type="bibr" rid="ref3">3</xref>
            ] is a W3C LEAD (Live Early Adoption and Demonstration) project under Semantic Web
Advanced Development (SWAD). It is currently supported by two Web browsers Amaya [4] and Annozilla
[5] which, like Trailblazer, is a Firefox extension. Annotea enhances collaboration via shared
metadatabased Web annotations, bookmarks [16], and their combinations. Annotations are comments, notes,
explanations, or other types of external remarks that can be attached to a Web document or a selected part
of the document without needing to touch the document. When the user gets the document he or she can
also load the annotations attached to it from a selected annotation server or several servers and see what his
peer group thinks [
            <xref ref-type="bibr" rid="ref9">17</xref>
            ]. Similarly shared bookmarks can be attached to Web documents, to help organize
them under different topics, to easily find them later, to help find related material and to collaboratively
filter bookmarked material. When a user accesses a Web page that has been referred to by shared
bookmarks this is highlighted on the page. They are then able to see and access other pages bookmarked by
other people as relevant to the topic. Trailblazer differs from Annotea in the respect that Trailblazer’s
bookmarks and notes are linked to a Named Entity, whereas Annotea’s are linked to a document or part of
a document using XPointers. This makes Trailblazer useful for pages that have not been viewed by the
community before, but contain Named Entities that are known to that community. Both Trailblazer and
Annotea allow annotations to be either shared or private. They also both use RDF for storing their
annotations. Work [18] is in progress to extend the Annotea address scheme to beyond XPointers. This
could allow Trailblazer in the future to use the Annotea schema.
          </p>
          <p>Magpie [19] is a Web browser extension which uses NER based on a supplied ontology to add links to
Named Entities on a Web page. Each ontology contains metadata associated with the Named Entities
which, when selected by the user, are resolved by a Web server and an appropriate Web page is returned.
Based on the user’s interpretation of what the page is about, they choose an appropriate ontology to use to
identify the Named Entities on that page and their associated links. Trailblazer differs from Magpie here in
that it uses its cosine similarity logic to automatically identify the correct category to use for Named Entity
identification. Magpie’s links are supplied in the ontology and cannot be added to by the user. Additionally
it does not have the facility to allow users to add annotations.</p>
          <p>
            XP Smart Tags [
            <xref ref-type="bibr" rid="ref10">20</xref>
            ] allow Smart Tag action buttons to be attached to a document in Microsoft’s Office
XP. It uses a DLL (Dynamic Link Library) called a Recognizer to identify items of interest on a page and
add Smart Tag action buttons. Out of the box it can recognize people’s names and add a Smart Tag to allow
the user to send them a mail or set up a meeting for example. Developers can create custom recognizer and
action DLL’s. This functionality would enable the building of read-only Trailblazer-like capability based
on ontologies across the whole of the Office XP desktop environment.
5
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusions and Future Work</title>
      <p>In this paper we have described the Firefox extension Trailblazer and how it can be used in a
collaborative environment. We have shown how Named Entities can be used as a basis for sharing
associative trails between Semantic Desktops. Future work will extend these capabilities across the whole
desktop infrastructure. Leveraging the XUL overlay technology we will apply the Trailblazer extension to
Thunderbird [21], which is the Mozilla mail client. This will allow the making and sharing of associative
trails across Web pages, RSS [22] and Mail. By building the capability in behind Smart Tags we should get
the same coverage across Windows XP. The Annotea schema appears to have good potential to be
extended to meet Trailblazer’s requirements. We also intend to build Annotea XPointer based bookmark
and annotation capability into Trailblazer as this will allow the support of document specific associative
trails. An additional area of exploration is the potential for using the Annotea server for sharing Trailblazer
Named Entity based associative trails. Security needs to be addressed to ensure that links only come from
trusted sources and that links are not to dangerous or undesirable Web resources.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Vannevar</given-names>
            <surname>Bush</surname>
          </string-name>
          , As we may think,
          <source>The Atlantic Monthly</source>
          <year>1946</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>N.</given-names>
            <surname>Chinchor</surname>
          </string-name>
          , editor (
          <year>1997</year>
          ). MUC-7
          <source>Named Entity Task Definition, Version 3.5. September 17th DARPA.</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3] Annotea http://www.annotea.org/ [4] Amaya http://www.w3.org/Amaya/ [5] Annozilla http://annozilla.mozdev.org/ [6]
          <string-name>
            <given-names>Firefox</given-names>
            <surname>Extensions</surname>
          </string-name>
          https://addons.mozilla.org/extensions/?application=Firefox [7]
          <string-name>
            <given-names>C. J. van Rijsbergen. Information</given-names>
            <surname>Retrieval</surname>
          </string-name>
          ,
          <year>1979</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>D.</given-names>
            <surname>Brickley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. V.</given-names>
            <surname>Guha</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <source>RDF Vocabulary Description Language 1</source>
          .0:
          <string-name>
            <given-names>RDF</given-names>
            <surname>Schema</surname>
          </string-name>
          .
          <source>W3C Recommendation 10 February</source>
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [9] Firefox http://www.mozilla.org/products/Firefox/ [10]
          <string-name>
            <surname>Vaughn</surname>
            <given-names>Bullard</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kevin</surname>
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>Michael C.</given-names>
          </string-name>
          <string-name>
            <surname>Daconta</surname>
          </string-name>
          .
          <source>Essential XUL Programming</source>
          . Wiley,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [11]
          <article-title>XPCOM</article-title>
          and XPConnect http://www.xulplanet.com/tutorials/xultu/xpcom.html [12]
          <string-name>
            <given-names>David</given-names>
            <surname>Flanagan. JavaScript - The Definitive Guide. O'Reilly</surname>
          </string-name>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [13] JSLib http://jslib.mozdev.org [14]
          <string-name>
            <given-names>G.</given-names>
            <surname>Salton and M. McGill</surname>
          </string-name>
          .
          <article-title>Introduction to Modern Information Retrieval</article-title>
          .
          <string-name>
            <surname>McGraw-Hill</surname>
          </string-name>
          ,
          <year>1983</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [15] http://en.wikipedia.org/wiki/Named_Entity_Recognition [16]
          <string-name>
            <given-names>M.</given-names>
            <surname>Koivunen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Swick</surname>
          </string-name>
          and
          <string-name>
            <given-names>E.</given-names>
            <surname>Prud'Hommeaux. Annotea Shared</surname>
          </string-name>
          <article-title>Bookmarks</article-title>
          .
          <source>In Proc. Of KCAP</source>
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>J.</given-names>
            <surname>Kahan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Koivunen</surname>
          </string-name>
          ,
          <string-name>
            <surname>E.</surname>
          </string-name>
          <article-title>Prud'Hommeaux, and</article-title>
          <string-name>
            <given-names>R.</given-names>
            <surname>Swick</surname>
          </string-name>
          .
          <article-title>Annotea: An Open RDF Infrastructure for Shared Web Annotations</article-title>
          .
          <source>In Proc. of the WWW10 International Conference. Hong Kong</source>
          ,
          <year>2001</year>
          [18] http://www.w3.org/2001/Annotea/Plan/context/newcontext.html [19]
          <string-name>
            <surname>John</surname>
            <given-names>Domingue</given-names>
          </string-name>
          , Martin Dzbor and
          <string-name>
            <given-names>Enrico</given-names>
            <surname>Motta</surname>
          </string-name>
          .
          <article-title>Collaborative Semantic Web Browsing with Magpie</article-title>
          .
          <source>In Proc. of the 1st European Semantic Web Symposium (ESWS)</source>
          ,
          <year>May 2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [20] Smart Tags http://office.microsoft.com/en-gb/assistance/HA010347451033.aspx [21] Thunderbird http://www.mozilla.org/products/thunderbird/ [22]
          <string-name>
            <given-names>Ben</given-names>
            <surname>Hammersley</surname>
          </string-name>
          .
          <article-title>Developing Feeds with RSS</article-title>
          and
          <string-name>
            <surname>Atom. O'Reilly</surname>
          </string-name>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>