<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Web Information Extraction systems for Web Semantization?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jan Dedek</string-name>
          <email>dedek@ksi.mff.cuni.cz</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Software Engineering, Faculty of Mathematics and Physics Charles University in Prague, Czech Republic Institute of Computer Science, Academy of Science of the</institution>
          <country country="CZ">Czech Republic</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper we present a survey of web information extraction systems and semantic annotation platforms. The survey is concentrated on the problem of employment of these tools in the process of web semantization. We compare the approaches with our own solutions and propose some future directions in the development of the web semantization idea.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>There exist many extraction tools that can process
web pages and produce structured machine
understandable data (or information) that corresponds with
the content of a web page. This process is often called
Web Information Extraction (WIE). In this paper we
present a survey of web information extraction systems
and we connect these systems with the problem of web
semantization.</p>
      <p>The paper is structured as follows. First we sketch
the basic ideas of semantic web and web
semantization. In the next two sections methods of web
information extraction will presented. Then description of
our solutions (work in progress) will continue. And
finally just before the conclusion we will discuss the
connection of WIE systems with the problem of web
semantization.</p>
      <sec id="sec-1-1">
        <title>The Semantic Web in use</title>
        <p>The idea of the Semantic Web [4] (World Wide Web
dedicated not only to human but also to machine –
software agents) is very well known today. Let us just
shortly demonstrate its use with respect to the idea of
Web Semantization (see in next section).</p>
        <p>The Fig. 1 shows a human user using the
(Semantic) Web in three possible manners: a keyword query,
a semantic query and by using a software agent. The
difference between the first two manners (keyword and
semantic query) can be illustrated with the question:
“Give me a list of the names of E.U. heads of state.”
This example from interesting article [16] by Ian
Horrocks shows the big difference between use of a
semantic query language instead of keywords. In the
semantic case you should be given exactly the list of names
you were requesting without having to pore through
results of (probably more then one) keyword queries.
Of course the user have to know the syntax of the
semantic query language or have a special GUI1 at
hand.</p>
        <p>The last and the most important possibility (in the
semantic or semantized setting) is to use some
(personalized) software agent that is specialized to tasks of
some kind like planning a business trip or finding the
most optimal choice from all the relevant job offers,
flats for rent, cars for sale, etc.</p>
        <p>Both the semantic querying and software agents
engagement is actually impossible to realize without
any kind of adaptation of the web of today in the
semantic direction.
1.2</p>
      </sec>
      <sec id="sec-1-2">
        <title>Web Semantization</title>
        <p>The idea of Web Semantization [9] consist in
gradual enrichment of the current web content as an
automated process of third party annotation for
mak1 Such handy GUI can be found for example in the KIM
project [20].
General
Applicable</p>
        <p>Domain Specific</p>
        <p>Structure of</p>
        <p>Document
e.g. HTML tables
Specific</p>
        <p>Form Specific</p>
        <p>Text</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Web information extraction</title>
      <p>ing at least a part of today’s web more suitable for The distinguishing between general applicable
machine processing and hence enabling it intelligent methods and the others that have meaningful
applitools for searching and recommending things on the cation only in some specific setting (specific domain,
web (see [3]). specific form of input) is very important for Web
Se</p>
      <p>The most strait forward idea is to fill a seman- mantization because when we try to produce
annotic repository with some information that is automat- tations in large scale, we have to control which web
ically extracted from the web and make it available to resource is suitable for which processing method (see
software agents so they could access to the web of to- in Sect. 5).
day in semantic manner (e.g. through semantic search
engine).</p>
      <p>The idea of a semantic repository and a public ser- 2.1 General applicable
vice providing semantic annotations was
experimentally realized in the very recognized work of IBM
Almaden Research Center: the SemTag [13]. This work
demonstrated that an automated semantic annotation
can be applied in a large scale. In their experiment
they annotated about 264 million web pages and
generated about 434 millions of semantic tags. They also
provided the annotations as a Semantic Label Bureau
– a HTTP server providing annotations for web
documents of 3rd parties.</p>
      <p>The most significant (and probably the only one)
generally applicable IE task is so called Instance
Resolution Task. The task can be described as follows: Given
a general ontology, find all the instances from the
ontology that are present in the processed resource. This
task is usually realized in two steps: (1) Named
Entity Recognition (see in Sect. 3.1), (2) Disambiguation
of ontology instances that can be connected with the
found named entities. Success of the method can be
strongly improved with coreference resolution (see in
Sect. 3.1).</p>
      <p>Let us mention several good representatives of this
approach: the SemTag application [13], the KIM
project [20] and the PANKOW annotation method [7]
based on smart formulation of Google API queries.</p>
      <p>The task of a web information extraction system is to
transform the web pages into program-friendly
structures such as a relational database. There exists a rich
variety of Web Information Extraction systems. The
results generated by distinct tools usually can not be 2.2 Domain specific
directly compared since the addressed extraction tasks
are different. The extraction tasks can be distinguished Domain and from specific IE approaches are the
typaccording several dimensions: the task domain, the au- ical cases. More specific information is more precise,
tomation degree, the techniques used, etc. These di- more complex and so more useful and interesting. But
mensions are analyzed in detail in the recent publica- the extraction method has to be trained to each new
tions [6] and [18]. Here we will concentrate on a lit- domain separately. This usually means indispensable
tle bit more specific division of WIE according to the effort.
needs of the Web Semantization (see in Sect. 5). The A good example of domain specific information
exdivision is demonstrated on the Fig. 2 and should traction system is SOBA [5]. This complex system is
not be considered as disjoint division of the methods capable to integrate different IE approaches and
exbut rather as emphasization of different aspects of the tract information from heterogeneous data resources,
methods. For example many extraction methods are including plain text, tables and image captions but
domain and form specific at the same time. the whole system is concentrated on the single domain
of football. Next similarly complex system is
ArtEquAKT [1], which is entirely concentrated on the
domain of art.
2.3</p>
      <sec id="sec-2-1">
        <title>Form specific</title>
        <p>Beyond general applicable extraction methods there
exist many methods that exploit specific form of the
input resource. The linguistic approaches usually
process text consisting of natural language sentences.
The structure-oriented approaches can be strictly
oriented on tables [19] or exploit repetitions of structural
patterns on the web page [21] (such algorithm can be
only applicable to pages that contain more than one
data record), and there are also approaches that use
the structure of whole site (e.g. site of single web shop
with summary pages with products connected with
links to pages with details about single product) [17].
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Information extraction from text-based resources</title>
      <p>In this section we will discuss the information
extraction from textual resources.
3.1</p>
      <sec id="sec-3-1">
        <title>Tasks of information extraction</title>
        <p>There are classical tasks of text preprocessing and
linguistic analysis like
Text Extraction – e.g from HTML, PDF or DOC,
Tokenization – detection of words, spaces,
punctuations, etc.,
Segmentation – sentence and paragraph detection,
POS Tagging – part of speech assignment, often
including lemmatization and morphological
analysis,
Syntactic Analysis (often called linguistic parsing )
– assignment of the grammatical structure to given
sentence with respect to given linguistic formalism
(e.g. formal grammar),</p>
      </sec>
      <sec id="sec-3-2">
        <title>Coreference Resolution (or anaphora resolution) –</title>
        <p>resolving what a pronoun, or a noun phrase refers
to. These references often cross boundaries of
a single sentence.</p>
        <p>Besides these classical general applicable tasks, there
are further well defined tasks, which are more closely
related to the information extraction. These tasks are
domain dependent. These tasks were widely developed
in the MUC-6 conference 1995 [15] and considered as
semantic evaluation in the first place. These
information extraction tasks are:
WIE systems for Web Semantization</p>
      </sec>
      <sec id="sec-3-3">
        <title>Named Entity Recognition: This task recognizes</title>
        <p>and classifies named entities such as persons,
locations, date or time expression, or measuring units.
More complex patterns may also be recognized as
structured entities such as addresses.</p>
      </sec>
      <sec id="sec-3-4">
        <title>Template Element Construction: Populates tem</title>
        <p>plates describing entities with extracted roles (or
attributes) about one single entity. This task
is often performed stepwise sentence by sentence,
which results in a huge set of partially filled
templates.</p>
      </sec>
      <sec id="sec-3-5">
        <title>Template Relation Construction: As each temp</title>
        <p>late describes information about one single entity,
this tasks identifies semantic relations between
entities.</p>
        <p>Template Unification: Merges multiple elementary
templates that are filled with information about
identical entities.</p>
      </sec>
      <sec id="sec-3-6">
        <title>Scenario Template Production: Fits the results</title>
        <p>of Template Element Construction and Template
Relation Construction into templates describing
pre-specified event scenarios (pre-specified“queries
on the extracted data”).</p>
        <sec id="sec-3-6-1">
          <title>Appelt and Israel [2] wrote an excellent tutorial summarizing these traditional IE tasks and systems built on them.</title>
          <p>3.2</p>
        </sec>
      </sec>
      <sec id="sec-3-7">
        <title>Information extraction benchmarks</title>
        <p>Contrary to the WIE methods based on the web page
structure, where we (the authors) do not know about
any well established benchmark for these methods2,
the situation in the domain of text based IE is fairly
different. There are several conferences and events
concentrated on the support of automatic machine
processing and understanding of human language in text
form. Different research topics as text (or information)
retrieval3, text summarization4 are involved.</p>
        <p>On the filed of information extraction, we have to
mention the long tradition of the Message
Understanding Conference5 [15] starting in 1987. In 1999 the event
of Automatic Content Extraction (ACE) Evaluation6
started, which is becoming a track in the Text Analysis
Conference (TAC)7 this year (in 2009).
2 It is probably at least partially caused by the vital
development of the presentation techniques on the web that
is still well in progress.
3 e.g. Text REtrieval Conference (TREC)</p>
        <p>http://trec.nist.gov/
4 e.g. Document Understanding Conferences</p>
        <p>http://duc.nist.gov/
5 Briefly summarized in http://en.wikipedia.org/
wiki/Message Understanding Conference.
6 http://www.itl.nist.gov/iad/mig/tests/ace/
7 http://www.nist.gov/tac</p>
        <sec id="sec-3-7-1">
          <title>All these events prepare several specialized datasets together with information extraction tasks and play an important role as information extraction benchmarks.</title>
          <p>4
4.1</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Our solutions</title>
      <sec id="sec-4-1">
        <title>Extraction based on structural similarity</title>
        <p>and stores the data in an ontology. We have made
initial experiments in the domain of reports of traffic
accidents. The results showed that this method can e.g.
aid summarization of the number of injured people.</p>
        <p>To avoid the need of manual design of extraction
rules we focused on the data extraction phase and
made some promising experiments [8] with the
machine learning procedure of Inductive Logic
Programming for automated learning of the extraction rules.</p>
        <p>This solution is directed to extraction of
information which is closely connected with the meaning of
text or meaning of a sentence.</p>
        <p>Our first approach for the web information extraction
is to use the structural similarity in web pages
containing large number of table cells and for each cell
a link to detailed pages. This is often presented in web
shops and on pages that presents more than one object
(product offer). Each object is presented in a similar 5 The Web Semantization setting
way and this fact can be exploited.</p>
        <p>As web pages of web shops are intended for hu- In this section we will discuss possibilities and
obstrucman usage creators have to make their comprehension tions connected with the employment of web
informaeasier. Acquaintance with several years of web shops tion extraction systems in the process of web
semanhas converged to a more or less similar design fashion. tization.</p>
        <p>There are often cumulative pages with many products One aspect of the realization of the web
semanin a form of a table with cells containing a brief de- tization idea is the problem of integration of all the
scription and a link to a page with details about each components and technologies starting with web
crawlparticular product. ing, going through numerous complex analyses
(docu</p>
        <p>Our main idea is to use a DOM tree representation ment preprocessing, document classification, different
of the summary web page and by breadth first search extraction procedures), output data integration and
encounter similar subtrees. The similarity of these sub- indexing, and finally implementation of query and
pretrees is used to determine the data region – a place sentation interface. This elaborate task is neither easy
where all the objects are stored. It is represented as nor simple but today it is solved in all the extensive
a node in the DOM tree, underneath it there are the projects and systems mentioned above.
similar sub-trees, which are called data records. The novelty that web semantization brings into
ac</p>
        <p>We8 have developed and implemented this idea [14] count is the cross domain aspect. If we do not want to
on the top of Mozilla Firefox API and experimentally stay with just general ontologies and general
applicatested on table pages from several domains (cars, note- ble extraction methods then we need a methodology
books, hotels). Similarity between subtrees was Lev- how to deal with different domains. The system has to
enshtein editing distance (for a subtree considered as support extension to a new domain in generic way. So
a linear string), learning thresholds for decision were we need a methodology and software to support this
trained. action. This can for example mean: to add a new
ontology for the new domain, to select and train proper
extractors and classifiers for the suitable input pages.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Linguistic information extraction 5.1</title>
        <p>Our second approach [11, 12, 10] for the web
information extraction is based on deep linguistic analysis. We
have developed a rule-based method for extraction of
information from text-based web resources in Czech
and now we are working on its adaptation to
English. The extraction rules correspond to tree queries on
linguistic (syntactic) trees made form particular
sentences. We have experimented with several linguistic
tools for Czech, namely Tools for machine annotation
– PDT 2.0 and the Czech WordNet.</p>
        <p>Our present system captures text of web-pages,
annotates it linguistically by PDT tools, extracts data
An interesting point is the question: Whose effort will
be used in the process of supporting new domain in
the web semantization process? How skilled such user
has to be? There are two possibilities (demonstrated
on the Fig 3). The easier one is that we have to
employ very experienced expert who will decide about
the new domain and who will also realize the support
needed for the new domain. In the Fig 3 this situation
is labeled as Provider Initiated and Provider Trained
because the expert works on the side of the system
8 Thanks go mainly to Duˇsan Maruˇsˇc´ak and Peter Vojt´aˇs. that provides the semantics.</p>
      </sec>
      <sec id="sec-4-3">
        <title>User initiative and effort</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <article-title>mentation of tables</article-title>
          .
          <source>In SIGMOD '04: Proceedings</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <source>of the 2004 ACM SIGMOD International Conference</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <year>2004</year>
          ,
          <fpage>119</fpage>
          -
          <lpage>130</lpage>
          . 18. B. Liu: Web Data Mining. Springer-Verlag,
          <year>2007</year>
          . 19.
          <string-name>
            <surname>D. Pinto</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Mccallum</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Wei</surname>
            , and
            <given-names>B.W.</given-names>
          </string-name>
          <string-name>
            <surname>Croft</surname>
          </string-name>
          : Table
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <source>'03: Proceedings of the 26th annual international ACM</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <year>2003</year>
          ,
          <fpage>235</fpage>
          -
          <lpage>242</lpage>
          . 20. B.
          <string-name>
            <surname>Popov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Kiryakov</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Ognyanoff</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Manov</surname>
          </string-name>
          , and
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <article-title>tion extraction and retrieval</article-title>
          .
          <source>Nat. Lang</source>
          . Eng.,
          <volume>10</volume>
          ,
          <fpage>3</fpage>
          -
          <lpage>4</lpage>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <year>2004</year>
          ,
          <fpage>375</fpage>
          -
          <lpage>392</lpage>
          . 21. H.
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          <string-name>
            <surname>Meng</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Raghavan</surname>
          </string-name>
          , and C. Yu:
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <source>In WWW Conference</source>
          ,
          <year>2005</year>
          ,
          <fpage>66</fpage>
          -
          <lpage>75</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>