<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Extraction of Semantic XML DTDs from Texts Using Data Mining Techniques</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Karsten Winkler</string-name>
          <email>kwinkler@ebusiness.hhl.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Myra Spiliopoulou</string-name>
          <email>myra@ebusiness.hhl.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Leipzig Graduate School of Management Department of E-Business Jahnallee 59</institution>
          ,
          <addr-line>D-04109 Leipzig</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Although composed of unstructured texts, documents contained in textual archives such as public announcements, patient records and annual reports to shareholders often share an inherent though undocumented structure. In order to facilitate efficient, structure-based search in archives and to enable information integration of text collections with related data sources, this inherent structure should be made explicit as detailed as possible. Inferring a semantic and structured XML document type definition (DTD) for an archive and subsequently transforming the corresponding texts into XML documents is a successful method to achieve this objective. The main contribution of this paper is a new method to derive structured XML DTDs in order to extend previously derived flat DTDs. We use the DIAsDEM framework to derive a preliminary, unstructured XML DTD whose components are supported by a large number of documents. However, all XML tags contained in this preliminary DTD cannot a priori be assumed to be mandatory. Additionally, there is no fixed order of XML tags and automatically tagging an archive using a derived DTD always implicates tagging errors. Hence, we introduce the notion of probabilistic XML DTDs whose components are assigned probabilities of being semantically and structurally correct. Our method for establishing a probabilistic XML DTD is based on discovering associations between, resp. frequent sequences of XML tags.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;semantic annotation</kwd>
        <kwd>XML</kwd>
        <kwd>DTD derivation</kwd>
        <kwd>knowledge discovery</kwd>
        <kwd>data mining</kwd>
        <kwd>clustering</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        Most organizations are not only “drowning” in data, they are
also “struggling” to cope with huge amounts of text
documents. Tan points out that up to 80% of a company’s
information is stored in unstructured textual documents [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ].
      </p>
      <p>Hence, capturing interesting and actionable knowledge from
textual databases is a major challenge for the data mining
community. Creating semantic markup is one form of
providing explicit knowledge about text archives to facilitate
searching and browsing or to enable information integration</p>
      <p>
        The work of this author is funded by the German Research Society
(DFG grant no. SP 572/4-1).
with related data sources. Unfortunately, most users are not
willing to manually create metadata due to the efforts and
costs involved [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Thus, text mining techniques are required
that (semi-) automatically create semantic markup and tag
documents accordingly.
      </p>
      <p>In this paper, we present the KDD approach pursued in
the research project DIAsDEM whose German acronym
stands for “Data Integration for Legacy Systems and
SemiStructured Documents by Means of Data Mining
Techniques”. Our goal is semantic tagging of textual content
with meta-data to facilitate searching, querying,
identification of and integration with associated texts and relational
data. Hence, we aim at deriving a structured XML DTD
that serves as a quasi-schema for the document collection
and enables the provision of database-like querying services
on textual data. DIAsDEM focuses on text collections with
domain-specific vocabulary and syntax that frequently share
an inherent, but undocumented structure.</p>
      <p>
        The DIAsDEM framework for semantic tagging of
domainspecific texts was introduced in [
        <xref ref-type="bibr" rid="ref11 ref12">12, 11</xref>
        ]. However,
applying the Java-based DIAsDEM Workbench to a text archive
currently results in a collection of semantically tagged XML
documents that are described by the extracted flat,
unstructured XML DTD. However, we ultimately aim at
integrating the resulting XML documents with other related data
sources. In this context, the derived unstructured, rather
preliminary DTD should be transformed into more
structured DTD that reflects both ordering and optionality of tags.
      </p>
      <p>Given that all XML tags are derived by data mining
techniques (i.e. iterative clustering as explained in section 3),
they are not crisp due to tagging errors. Taking this critical
fact into account, we introduce the notion of a
probabilistic DTD that describes the most likely orderings of XML
tags and that contains statistical properties for each tag. The
structured DTD will be the basis for future information
integration efforts that involve XML archives generated by the
DIAsDEM Workbench. We introduce two algorithms for
inferring a probabilistic DTD that utilize association rule
discovery algorithms and sequence mining techniques.</p>
      <p>The rest of this paper is organized as follows: The next
&lt;?xml version="1.0" encoding="ISO-8859-1"?&gt;
&lt;!DOCTYPE CommercialRegisterEntry SYSTEM ’CommercialRegisterEntry.dtd’&gt;
&lt;CommercialRegisterEntry&gt; &lt;BusinessPurpose&gt; Der Betrieb von Spielhallen in Teltow und das
Aufstellen von Geldspiel- und Unterhaltungsautomaten. &lt;/BusinessPurpose&gt; &lt;ShareCapital
AmoutOfMoney="25000 EUR"&gt; Stammkapital: 25.000 EUR. &lt;/ShareCapital&gt;
&lt;LimitedLiabilityCompany&gt; Gesellschaft mit beschränkter Haftung. &lt;/LimitedLiabilityCompany&gt;
&lt;ConclusionArticles Date="12.11.1998; 19.04.1999"&gt; Der Gesellschaftsvertrag ist am 12.11.1998
abgeschlossen und am 19.04.1999 abgeändert. &lt;/ConclusionArticles&gt; (...) Einzelvertretungsbefugnis kann erteilt
werden. &lt;AppointmentManagingDirector Person="Balski; Pawel; Berlin; 14.04.1965"&gt;
Pawel Balski, 14.04.1965, Berlin, ist zum Geschäftsführer bestellt. &lt;/AppointmentManagingDirector&gt; (...)
&lt;PublicationMedia&gt; Nicht eingetragen: Die Bekanntmachungen der Gesellschaft erfolgen im Bundesanzeiger.
&lt;/PublicationMedia&gt; &lt;/CommercialRegisterEntry&gt;
section briefly discusses related work. Section 3 gives an
overview of our framework for semantic tagging of
domainspecific text collections. Section 4 introduces the notion of
probabilistic DTDs for textual archives and develops two
methods for deriving them. Finally, we conclude and give
directions for future research in section 5.</p>
    </sec>
    <sec id="sec-2">
      <title>RELATED WORK</title>
      <p>
        Nahn and Mooney propose the combination of methods from
KDD and information extraction to perform text mining tasks
[
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. They apply standard KDD techniques to a
collection of structured records that contain previously extracted,
application-specific features from texts. Feldman et al.
propose text mining at the term level instead of focusing on
linguistically tagged words [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. The authors represent each
document by a set of terms and additionally construct a
taxonomy of terms. The resulting dataset is input to KDD
algorithms such as association rule discovery. Our DIAsDEM
framework adopts the idea of representing texts by terms and
concepts. However, our goal is the semantic tagging of
structural text units (e.g., sentences or paragraphs) within the
document according to a global DTD and not the
characterization of the entire document’s content. Loh et al. suggest to
extract concepts rather than individual words for subsequent
use in KDD efforts at the document level. [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Similarly to
our framework, the authors suggest to exploit existing
vocabularies such as thesauri for concept extraction. Mikheev and
Finch describe a workbench to acquire domain knowledge
from texts [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. Similar to the DIAsDEM Workbench, their
approach combines methods from different fields of research
in a unifying framework.
      </p>
      <p>Our approach shares with this research thread the objective of
extracting semantic concepts from texts. However, concepts
to be extracted in DIAsDEM must be appropriate to serve as
elements of the XML DTD. Among other implications,
discovering a concept that is peculiar to a single text unit is not
sufficient for our purposes, although it may perfectly reflect
the corresponding content. In order to derive a DTD, we need
to discover groups of text units that share some semantic
concepts. Moreover, we concentrate on domain-specific texts,
which significantly differ from average texts with respect to
word frequency statistics. These collections can hardly be
processed using standard text mining software because the
integration of relevant domain knowledge is a prerequisite
for successful knowledge discovery.</p>
      <p>
        There are only a few research activities aiming at the
transformation of texts into semantically annotated XML
documents: Becker et al. introduce the search engine
GETESS that supports query processing on texts by deriving and
processing XML text abstracts [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. These abstracts
contain language-independent, content-weighted summaries of
domain-specific texts. In DIAsDEM, we do not separate
meta-data from original texts but rather provide a
semantic annotation, keeping the texts intact for later processing
or visualization. Given the aforementioned linguistic
particularities of the application domains we investigate, a DTD
characterizing the content of the documents is more
appropriate than inferences on their content. In order to transform
existing content into XML documents, Sengupta and Purao
propose a method that infers DTDs by using already tagged
documents as input [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. In contrast, we propose a method
that tags plain text documents and derives a DTD for them.
Closer to our approach is the work of Lumera, who uses
keywords and rules to semi-automatically convert legacy data
into XML documents [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. However, his approach relies on
establishing a rule base that drives the conversion, while we
use a KDD methodology that reduces human effort.
Semi-structured data is another topic of related research
within the database community [
        <xref ref-type="bibr" rid="ref1 ref6">6, 1</xref>
        ]. A lot of effort has
recently been put into methods inferring and representing
structure in similar semi-structured documents [
        <xref ref-type="bibr" rid="ref14 ref21 ref27">21, 27, 14</xref>
        ].
However, these approaches only derive a schema for a given
set of semi-structured documents. In DIAsDEM, we have to
simultaneously solve the problems of both semi-structuring
text documents by semantic tagging and inferring an
appropriately structured XML DTD that describes the related
archive. We are not aware of any scientific or commercial
approaches employing probabilistic document type definitions
as introduced in this paper for describing text archives or
integrating texts with related data sources.
      </p>
    </sec>
    <sec id="sec-3">
      <title>THE DIAsDEM FRAMEWORK</title>
      <p>In this paper, the notion of semantic tagging refers to the
activity of annotating texts with domain-specific XML tags
that might contain additional attributes. Rather than
classifying entire documents or tagging single terms, we aim at
semantically tagging text units such as sentences or
paragraphs. Table 1 illustrates this concept of semantic tagging,
whereas each sentence of this German Commercial Register
entry is a text unit. In this example, the semantics of most
sentences are made explicit by XML tags that partly
contain additional attributes describing extracted named entities
(e.g., names of persons and amounts of money). The XML
document depicted in Table 1 was created by applying the
DIAsDEM framework to a collection of 1,145 textual
Commercial Register entries containing 10,785 text units. This
collection includes all entries related to foundations of
companies in the district of the German city Potsdam in 1999.
In Germany, companies are obliged by law to submit
various information about business affairs to local Commercial
Registers. Although Commercial Registers are an important
source of information in daily business transactions, their
textual content can only be searched using full-text queries at
the moment. Hence, semantically semi-structuring these
textual archives provides the basis for information integration
and creation of value-adding services related to information
brokerage. XML query languages could be employed to
submit both both content- and structure-based queries against
semantically tagged XML archives.</p>
      <p>
        Our framework pursues two objectives for a given archive
of text documents: All text documents should be
semantically tagged and an appropriate, preliminary flat XML DTD
should be derived for the archive. Semantic tagging in
DIAsDEM is a two-phase process. We have designed a knowledge
discovery in textual databases (KDT) process that constitutes
the first phase in order to build clusters of semantically
similar text units, to tag documents in XML according to the
results and to derive an XML DTD describing the archive.
The KDT process that was introduced in [
        <xref ref-type="bibr" rid="ref11 ref12">12, 11</xref>
        ] results in a
final set of clusters whose labels serve as XML tags and DTD
elements. Huge amounts of new documents can be converted
into XML documents in the second, batch-oriented and
productive phase of the DIAsDEM framework. All text units
contained in new documents are clustered by the previously
built text unit clusterer and are subsequently tagged with the
corresponding cluster labels.
      </p>
      <p>In DIAsDEM we concentrate on the semantic tagging of
similar text documents originating from a common domain.
Nevertheless, the DIAsDEM approach is appropriate for
semantically tagging various kinds of archives such as
public announcements of courts and administrative authorities,
quarterly and annual reports to shareholders, textual patient
records in health care applications as well as product and
service descriptions published on electronic marketplaces.
=======
==========
==========
==============</p>
      <p>=======
==============
==============
==================</p>
      <p>============
==================================================
============
=====================================
========
===================</p>
      <sec id="sec-3-1">
        <title>Text Documents</title>
        <p>=============== == ========== ==================
=== === ===</p>
      </sec>
      <sec id="sec-3-2">
        <title>UML Schema</title>
        <p>========
======== ======== ========
In the remainder of this section, we briefly introduce the first
phase of the DIAsDEM framework whose iterative and
interactive KDT process is depicted in Figure 1. This process
is termed “iterative” because the clustering algorithm is
invoked repeatedly. Our notion of iterative clustering should
not be confused with the fact that most clustering algorithms
perform multiple passes over the data before converging.</p>
        <p>This process is also “interactive”, because a knowledge
engineer is consulted for cluster evaluation and final cluster
naming decisions at the end of each iteration.</p>
        <p>Besides the initial text documents to be tagged, the
following domain knowledge constitutes input to our KDT process:
A thesaurus containing a domain-specific taxonomy of terms
and concepts, a preliminary UML schema of the domain and
descriptions of specific named entities of importance, e.g.
persons and companies. The UML schema reflects the
semantics of named entities and the relationships among them,
as they are initially conceived by application experts. This
schema serves as a reference for the DTD to be derived from
discovered semantic tags, but there is no guarantee that the
&gt; SupervisoryBoard | (...) | Owner | FoundationPartnership )*
&lt;!ELEMENT CommercialRegisterEntry ( #PCDATA | BusinessPurpose | ShareCapital |
ModificationMainOffice | FullyLiablePartner | AppointmentManagingDirector |
GeneralPartnership | InitialShareholders | NonCashCapitalContribution |
LimitedLiabilityCompany | ConclusionArticles | ModificationRegisteredName |
&lt;?xml version="1.0" encoding="ISO-8859-1"?&gt;
&lt;!ELEMENT BusinessPurpose (#PCDATA)&gt;
&lt;!ELEMENT ShareCapital (#PCDATA)&gt; (...)
&lt;!ELEMENT FoundationPartnership (#PCDATA)&gt;
final DTD will be contained in or will contain this schema.
clusters containing approx. 85% of text units.</p>
        <p>
          Similarly to a conventional KDD process, our process starts
with a preprocessing phase: After setting the level of
granularity by determining the size of text units to be tagged,
the Java- and Perl-based DIAsDEM Workbench performs
basic NLP preprocessing such as tokenization, normalization
and word stemming using TreeTagger [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ]. Instead of
removing stop words, we establish a drastically reduced
feature space by selecting a limited set of terms and concepts
(so-called text unit descriptors) from the thesaurus and the
UML schema. Text unit descriptors are currently chosen
by the knowledge engineer because they must reflect
important concepts of the application domain. All text units are
mapped into Boolean vectors of this feature space.
Additionally, named entities of interest are extracted from text units by
a separate module of the DIAsDEM Workbench. In our case
study, we created a small thesaurus and selected 70 relevant
descriptors and 109 non-descriptors pointing to descriptors.
        </p>
        <p>In the pattern discovery phase, all text unit vectors contained
in the initial archive are clustered based on similarity of their
content. The objective is to discover dense and homogeneous
text unit clusters. Clustering is performed in multiple
iterations. Each iteration outputs a set of clusters, which the
DIAsDEM Workbench partitions into ”acceptable” and
”unacceptable” ones according to our quality criteria. A cluster
of text unit vectors is ”acceptable”, if and only if (i) its
cardinality is large and the corresponding text units are (ii)
homogeneous and (iii) can be semantically described by a small
number of text unit descriptors. Members of “acceptable”
cluster are subsequently removed from the dataset for later
labeling, whereas the remaining text unit vectors are input
data to the clustering algorithm in the next iteration. In each
iteration, the cluster similarity threshold value is stepwise
decreased such that “acceptable” clusters become
progressively less specific in content. The KTD process is based on a
plug-in concept that allows the execution of different
clustering algorithms within the DIAsDEM Workbench. In the case
study, we employed the demographic clustering function
included in the IBM Intelligent Miner for Data that maximizes
the value of Condorcet’s criterion. After three iterations, the
DIAsDEM Workbench discovered altogether 73 “acceptable”
The postmining phase consists of a labeling step, in which
“acceptable” clusters are semi-automatically assigned a
label. Ultimately, cluster labels are determined by the
knowledge engineer. However, the DIAsDEM Workbench performs
both a pre-selection and a ranking of candidate cluster
labels for the expert to choose from. All default cluster labels
are derived from feature space dimensions (i.e. from text
unit descriptors) that are prevailing in each “acceptable”
cluster. Cluster labels actually correspond to XML tags that are
subsequently used to annotate cluster members. Finally, all
original documents are tagged using valid XML tags.
Additionally, XML tags are enhanced by attributes reflecting
previously extracted named entities and their values. Table 2
contains an excerpt of the flat, unstructured XML DTD that
was automatically derived from XML tags in the case study.</p>
        <p>It coarsely describes the semantic structure of the resulting
XML collection. Currently, named entities that serve as
additional attributes of XML tags are not fully evaluated by the
DIAsDEM Workbench.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>ESTABLISHING A PROBABILISTIC DTD</title>
      <p>The output of the DIAsDEM Workbench is a set of semantic
XML tags which should be used as XML tags to describe the
content of the archive documents. To reflect the content of
the archive at an abstract level, it is essential to compose the
tags into a DTD. Since the semantic annotations are derived
with data mining techniques, they are not crisp. Thus, it is
essential that the validity of each tag is expressed in
quantitative terms and is estimated properly. Furthermore, an
ordering should be imposed upon the tags. Hence, after deriving
semantic XML tags, we combine them into a probabilistic
DTD by (i) deriving the most likely ordering of the tags and
(ii) computing the statistical properties of each tag inside the
document type definition.</p>
      <p>The reader may recall that a semantic annotation is
actually the label of a cluster discovered by the DIAsDEM
Workbench. The underlying clustering mechanism produces
nonoverlapping clusters. This implies that a text unit belongs to
exactly one cluster, to the effect that it can be annotated with
the label of this cluster only. Hence, the tags/labels derived
the DIAsDEM Workbench cannot be nested. An extension of
the DIAsDEM Workbench by a hierarchical clustering
algorithm would allow for the establishment of subclusters and
thus for the nesting of (sub)cluster labels. However, this is
planned as future work.</p>
      <p>The objectives of the DTD establishment method are the
specification of the most appropriate ordering of tags, the
identification of correlated or mutually exclusive tags and the
adornment of each tag and each correlation among tags with
statistical properties. These properties form the basis for
reliable query processing, because they determine the expected
precision and recall of the query results. In the following, we
first introduce the statistical properties we consider for the
DTD tags and their associations and describe the
methodology for computing these statistics. To model the complete
statistical information pertinent in these tags and their
relationships, we use a hypergraph structure. We then introduce a
mechanism that derives a probabilistic DTD from this graph.
TagSupport The tags of the DTD are cluster labels derived
by a statistical approach. Thus, in terms of XML, they are
observed as optional per se. In many application areas, a
domain expert can provide suggestions as to which tags should
be observed as mandatory. Despite this, there is no
guarantee that the expert’s suggestions hold true in the archive: The
text unit containing this information may have been
misclassified by the DIAsDEM Workbench, or the information may
be simply absent from the document. For example, although
one would expect that each movie has a regisseur, there are
movies whose regisseur is unknown or inapplicable, due to
the nature of the movie. The property TagSupport offers an
indicator of whether a tag may be considered as potentially
mandatory. We define it as the ratio of documents where this
tag appears to the total number of documents in the archive.
of error type II text units is higher, indicating that some text
units were not placed in the cluster they semantically belong
to. With 0.95 confidence, the overall error rate in the entire
dataset is in the interval [2.6%, 5.9%] which is a promising
result.</p>
      <p>GroupSupport In most of the above statistics, we juxtapose
the frequence of appearance of a tag with the frequency of
a group of tags, be it a set or a sequence. We use the term</p>
    </sec>
    <sec id="sec-5">
      <title>Statistical Properties of Semantic XML Tags</title>
      <p>The statistical properties of DTD tags are depicted in Table 3
and described in the following paragraphs. The first column
contains the names of the properties. The second column
reflects whether the property is peculiar to the whole set of
tags as cluster labels (i.e. the whole ”model”), to each tag
or to a group of associated tags. The last column names the
mechanism to be applied to derive the value of each property
for each tag.</p>
      <p>Accuracy The DIAsDEM Workbench derives semantic
XML tags as labels of clusters. These clusters constitute a
model over the data, in the conventional statistical sense. In
terms of data classification, such models are subject to
misclassification errors. We identify two types of
misclassification:</p>
      <p>Error type I: A text unit is assigned to the wrong cluster,
i.e. the cluster label does not reflect the content of the text
unit.</p>
      <p>Error type II: A text unit is not assigned to any cluster,
although there is a cluster with a label reflecting the content
of the text unit.</p>
      <p>For the envisaged DTD, only the error type I is relevant. We
use the term accuracy of the model as the probability that
cluster labels reflect the content of cluster members. The
accuracy value affects the DTD as a whole, it is not peculiar to
individual tags. Therefore, we do not incorporate this value
in the statistical adornment of the individual tags.</p>
      <p>In order to evaluate the quality of out approach in absence of
pre-tagged documents, we drew a random sample containing
5% out of 10,785 text units and asked a domain specialist to
verify the annotations of these text units with respect to both
error types. Within the sample, error type I (error type II)
occured in 0.4% (3.6%) of text units. Hence, tagged text units
are most likely to be correctly processed. The percentage
Property
Accuracy
TagSupport
AssociationConfidence
AssociationLift
LocationConfidence
GroupSupport</p>
      <p>Radius
model
tag
set of tags
set of tags
sequence of tags
set or sequence of tags
y 1 : : : ! An edge represents a relationship of the form yn x;
x however, we use the convention that is the source node and
We represent the tags and their associations in a directed
graph. Its nodes are individual tags, sequences of adjacent
tags or sets of co-occuring tags. Each node is adorned with
the statistical properties pertinent to a tag, resp. tag group.
the group of nodes in the rule’s LHS is the target. Similarly
to nodes, an edge is adorned with the statistics of the
orderinsensitive or order-sensitive association it represents.
GroupSupport as the ratio of the number of documents
containing a group of tags to the total number of documents. In
fact, for any set of at least two tags, this property assumes one
value for the set and as many values as are the perturbations
of set members. In the following subsection, we show how
we model the statistical information pertinent to individual
tags, to tag groups (i.e. sets or sequences) and to
relationships among them in a seamless way.</p>
    </sec>
    <sec id="sec-6">
      <title>Modeling Statistics of Associated XML Tags</title>
      <p>Some of the values of the statistical properties depicted in
Table 3 are already made available as part of the DIAsDEM
Workbench output, while the remaining ones can be
computed by data mining algorithms. To exploit these values for
the establishment of a probabilistic DTD, we need a
representation model and an algorithm that builds the DTD when
processing this model. We introduce here a generic graph
structure, in which all statistical information on tags, groups
of tags and tag relationships is depicted. This structure is
appropriate for the establishment of a DTD or an XMLschema
with rich statistical adornments. In the next subsection, we
discuss two algorithms for DTD establishment.
Computation method
DIAsDEM Workbench
simple statistics
association rule discovery
association rule discovery (ARD)
sequence mining (SeqM)</p>
      <p>ARD/SeqM
lists. Second, we annotate each list with a flag that indicates
whether the group depicted by the list is order-sensitive or
order-insensitive; in the latter case, the ordering of the list is
irrelevant but must be unique. Third, we guarantee
uniqueness, i.e. that all permutations of the same group of tags are
mapped into the same order-insensitive list, by requiring that
order-insensitive list are lexicographically ordered.
V A
signature:
By this definition, a tag can only appear in the graph if its
TagSupport is more than zero. This is consistent with the
fact that XML tags are derived with a KDD method.</p>
      <p>For the groups of tags, we must distinguish among
ordersensitive and order-insensitive groups. To do so, we
perform three steps. First, we model tag groups as ordered</p>
      <p>V 0 where contains only those groups of annotations, for
which the GroupSupport value is above a given threshold.</p>
      <p>This threshold can be specified as input to the mining
software, as is usual in KDD applications, or may be set as low
as 0. Of course, the threshold value affects the size of the
graph and the execution time of the algorithm that traverses
it to build the DTD.
X := (0; 1] [ fN U where LLg, with signature:</p>
      <p>In this signature, the statistical properties refer to the edge’s
source given the group of nodes in the edge’s target. If the
target is a sequence of adjacent tags, then the location
confidence is the only valid statistical property, while the
association confidence and lift are inapplicable. If the target is a set
of tags, then the location confidence is inapplicable. When
a statistical property is inapplicable, it assumes the NULL
value.</p>
      <p>Graph Properties The components of our
“DTDestablishment graph” are tags, groups of tags and
relationships among them, all adorned with statistical values.</p>
      <p>All tags discovered by the DIAsDEM Workbench are present
in this graph. Which groups of tags are present depends on
the threshold value for the group support. Conceivable are
both a minimalistic approach with a high threshold, by which
only very frequent groups are present, and a maximalistic
approach with a zero-value threshold, by which all tag
combinations occuring in the documents are present.</p>
      <p>If we opt for the minimalistic approach, the graph will not
be connected in the general case. It will contain only the
groups of tags being more frequent than a threshold, and the
frequent relationships among them. Certain tags may be
isolated, because they only rarely appear in combination with
other tags. Contrary to it, the maximalistic approach ensures
that all combinations of tags appearing together in documents
are depicted in the graph, and that the graph is connected,
except of the unlikely case that some documents contain a
single tag not occuring in any other documents.</p>
      <p>The upper limit to the graph size indicates that threshold
values for the statistical properties are essential for obtaining a
manageable graph. On the other hand, each cutoff value
implies an information loss. Therefore, we observe the
DTDestablishment graph under the maximalistic approach as a
reference structure and introduce two algorithms that derive
a probabilistic DTD by constructing only a part of this graph.
In the following, we present two algorithms that derive a
DTD by constructing part of the DTD-establishment graph.</p>
      <p>Each tag of this DTD is adorned by only two (derived)
probabilistic values, one referring to the tag itself and one to its
location inside the DTD. The algorithms are using different
heuristics to derive this DTD: the first one concentrates on
the pairs of tags appearing most frequently together, while
the second one gives preference to maximal sequences of
tags. The reader may recall that the computation of the
statistics for the relationships among the tags require the activation
of data mining software. Hence, each of the algorithms is
backed by a miner that returns the desired statistics.</p>
    </sec>
    <sec id="sec-7">
      <title>DTD Derivation</title>
      <p>The DTD-establishment graph in its maximalistic version
captures all relationships among the semantic tags found
by the DIAsDEM Workbench. Similarly to the process of
schema establishment for a conventional database
application, the designer must decide which relationships among
the real-world entities are worth capturing and which are not.
In our context, “worth capturing” refers to statistical values,
presuming that a DTD should reflect the relationships usually
present in the documents rather than the rare ones. However,
the DTD-establishment graph contains relationships among
sets and among sequences of tags, each one adorned with
different (and only partially comparable) statistics.</p>
    </sec>
    <sec id="sec-8">
      <title>Backward Construction of DTD Sequences</title>
      <p>This algorithm observes a DTD as a set of alternative
sequences and builds each sequence backwards, starting at
each last tag and proceeding until the first one. Concretely,
the algorithm builds “maximal” sequences, where
maximality means that the first tag of the sequence is the first tag in
most of the documents supporting the sequence.
: : : ; The maximum among c0; cu determines the rest of the
procedure: If c0 is maximum, is the first element of the
sequence. In this case, the sequence is marked as “done” and
as “maximal” according to the maximality criterion already
mentioned.
where c0 is the ratio of documents where has no
predecessor divided by the total number of documents.
k 1 confidences of the other tags is less than some small
k k values. In other words, there are tags with u, such
k k ". Then, all tags are acceptable alternatives, resulting to
If there is a ci larger than the other elements, then xi is the
predecessor of in the sequence. However, it can be the case
that the maximum is only marginally larger than the other
that (i) one of them has shows the maximum location
confidence but (ii) the difference of this value from the location
alternative subsequences.
s = 1 : : : s a subsequence k, then is “done” but it must
x number of documents and comparing this value to the tag
1 puting the ratio of documents starting with over the whole
s and thus is maximal. Otherwise, documents starting with
jx cj s 1 ", then most documents containing start with
1 mostly adhere to a different maximal sequence.</p>
      <p>Identifying maximal tag-sequences. In each iteration, the
algorithm considers longer frequent sequences returned by
the sequence miner, namely those ending with each
subsequence already built. If no frequent sequence is found for
also be checked whether it is maximal. This implies
comsupport of 1, say c. The comparison is performed across
the same guidelines as for alternative tag predecessors: if
s side each maximal sequence containing it:</p>
      <p>Statistics of maximal tag-sequences. At a final step, the
algorithm filters out all sequences that are done but are not
maximal. It then assigns probability values to each tag
ins this tag with respect to the subsequence of leading to it.</p>
      <p>The TagConfidence is the tag’s TagSupport multiplied by
the accuracy of the model output by the DIAsDEM
Workbench.</p>
      <p>The TagPositionConfidence is the location confidence of
Backward versus forward sequence construction. The
backward-sequence-construction method generates
alternative sequences of DTD tags by pruning the frequent
sequences of adjacent tags produced by a sequence miner. An
equivalent method can be devised by forward-sequence
construction. This would have the advantage of being
appropriate for incorporation to a sequence miner’s core as well, since
most miners of this category perform forward sequence
construction: In that case, the mining kernel would be modified
to expand a sequence by the most likely successor tag only.</p>
    </sec>
    <sec id="sec-9">
      <title>A DTD as a Tree of Alternatives</title>
      <p>This algorithm observes a DTD as a tree of alternative
subsequences and adorns each tag with its support with respect
to the subsequence leading to it inside the tree: this is the
number of documents starting with this subsequence of tags.
Similarly to the sequence-construction algorithm described
above, a tag may appear in more than one subsequences,
having different predecessors in each one.</p>
      <p>
        Observing the DTD as a tree implies a common root. In
the general case, each document of the archive may start at
a different tag. We assume a dummy root, the children of
which are those tags that appear first in documents. In
general, a tree node refers to a tag , and its children refer to
the tags appearing after in the context of ’s own
predecessors. In a sense, the DTD as a tree of alternatives resembles
a DataGuide as proposed in [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], although the latter contains
no statistical adornments.
      </p>
      <p>The tree-of-alternatives differs from the
sequenceconstruction algorithm in two ways: Firstly, it considers
all sequences of tags that appear in documents instead
of frequent ones only. Secondly, it only observes
complete sequences, while a sequence miner returns arbitrary
subsequences of tags.</p>
      <p>
        The tree-of-alternatives method is realized by the
preprocessor module of the Web usage miner WUM [
        <xref ref-type="bibr" rid="ref24 ref25">25, 24</xref>
        ]. This
module is responsible for coercing sequences of events by
common prefix and placing them in a tree structure, called
“aggregated tree”. This tree is input to the navigation
pattern discovery process performed by the WUM core. The
sequences of tags in documents can be observed as sequences
of events, to the effect that the WUM preprocessor can also
be used to build a DTD over an archive as a tree of
alternative tag sequences. Figure 2 depicts an example of such a tree
that related to our case study. Note that the XML document
depicted in Table 1 is partly described by this DTD excerpt.
      </p>
    </sec>
    <sec id="sec-10">
      <title>CONCLUSION</title>
      <p>Most of the knowledge hidden in electronic media of an
organization is encapsulated in documents. Acquiring this
knowledge implies effective querying of the documents as well as
the combination of information pieces from different textual
assets. This functionality is usually confined to
databaselike query processors, while text search engines scan
individual assets and return ranked results. In this study, we
have presented a methodology that enables query processing
and joining of text sources by structuring them. We propose
the derivation of an XML DTD over a domain-specific text
archive by means of data mining techniques.</p>
      <p>The semantic characterization of text units is the core of our
approach as well as the derivation of XML tags from these
characterisations. This is undertaken by the DIAsDEM
Workbench which is concisely described in the first part of this
study. Our main emphasis is on combining these tags that
reflect the semantics of many text units across the archive into
a single DTD that reflects the semantics of the archive as a
whole. We have shown that this DTD is a probabilistic
ap(...)</p>
      <sec id="sec-10-1">
        <title>BusinessPurpose, 950 (...)</title>
      </sec>
      <sec id="sec-10-2">
        <title>FullyLiablePartner, 95</title>
      </sec>
      <sec id="sec-10-3">
        <title>LimitedLiabilityCompany, 123</title>
      </sec>
      <sec id="sec-10-4">
        <title>Procuration, 39</title>
      </sec>
      <sec id="sec-10-5">
        <title>ModificationArticles_ShareCapital, 6</title>
      </sec>
      <sec id="sec-10-6">
        <title>ConclusionArticlesOfAssociation,13 (...) ShareCapital, 129 Procuration, 5</title>
        <p>proximation of the archive content and have derived a set of
statistical properties that reflect the quality of this
approximation, for the whole DTD, for tags inside the DTD and for
relationships among these tags.</p>
        <p>The statistical properties of tags and of their relationships
form the basis for combining them into a complete DTD
in the XML sense or even into an XMLschema. We use a
graph structure to depict all statistics that can serve as a
basis for this operation and propose two mechanisms that derive
DTDs by employing a mining algorithm and a set of heuristic
rules. We have tested our methodology on an archive of
documents from a regional Commercial Register in Germany:
We have derived a set of tags with the DIAsDEM Workbench
and then implemented one of the proposed mechanisms to
derive a DTD for it.</p>
        <p>Our future work includes the implementation of the second
mechanism for DTD derivement and the establishment of a
framework for the comparison of derived DTDs in terms of
expressiveness and accuracy. Of course, the ultimate goal
of our work is the establishment of a full-fledged querying
mechanism over the text archives. To this purpose, we intend
to couple our DTD derivation methods with a query
mechanism for semi-structured data. Since the DTDs we derive
are of probabilistic nature, this implies also the design of a
model that evaluates the quality of the query results.</p>
      </sec>
    </sec>
    <sec id="sec-11">
      <title>ACKNOWLEDGMENTS</title>
      <p>We thank the German Research Society for funding the
project DIAsDEM, the Bundesanzeiger Verlagsgesellschaft
mbH for providing data and our project collaborators
Evguenia Altareva and Stefan Conrad for helpful discussions. The
IBM Intelligent Miner for Data is kindly provided by IBM in
terms of the IBM DB2 Scholars Program.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>S.</given-names>
            <surname>Abiteboul</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Buneman</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Suciu</surname>
          </string-name>
          .
          <article-title>Data on the Web: From Relations to Semistructured Data and XML</article-title>
          . Morgan Kaufman Publishers, San Francisco,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>R.</given-names>
            <surname>Agrawal</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Srikant</surname>
          </string-name>
          .
          <article-title>Mining sequential patterns</article-title>
          .
          <source>In Proc. of Int. Conf. on Data Engineering</source>
          , Taipei, Taiwan, Mar.
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>M.</given-names>
            <surname>Baumgarten</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. G.</given-names>
            <surname>Büchner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. S.</given-names>
            <surname>Anand</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. D.</given-names>
            <surname>Mulvenna</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J. G.</given-names>
            <surname>Hughes</surname>
          </string-name>
          .
          <article-title>Navigation pattern discovery from internet data</article-title>
          .
          <source>In [17]</source>
          , pages
          <fpage>70</fpage>
          -
          <lpage>87</lpage>
          .
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>M.</given-names>
            <surname>Becker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bedersdorfer</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Bruder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Düsterhöft</surname>
          </string-name>
          , and
          <string-name>
            <surname>G. Neumann. GETESS</surname>
          </string-name>
          :
          <article-title>Constructing a linguistic search index for an Internet search engine</article-title>
          .
          <source>In Proceedings of the 5th International Conference on Applications of Natural Language to Information Systems</source>
          , Versailles, France,
          <year>June 2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Berry</surname>
          </string-name>
          and
          <string-name>
            <given-names>G.</given-names>
            <surname>Linoff. Data Mining</surname>
          </string-name>
          <article-title>Techniques: For Marketing, Sales</article-title>
          and
          <string-name>
            <given-names>Customer</given-names>
            <surname>Support</surname>
          </string-name>
          . John Wiley &amp; Sons, Inc.,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>P.</given-names>
            <surname>Buneman</surname>
          </string-name>
          .
          <article-title>Semistructured data</article-title>
          .
          <source>In Proceedings of the Sixteenth ACM SIGACT-SIGMOD-SIGART Symposium on Principles of Database Systems</source>
          , pages
          <fpage>117</fpage>
          -
          <lpage>121</lpage>
          , Tucson,
          <string-name>
            <surname>AZ</surname>
          </string-name>
          , USA, May
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>M.</given-names>
            <surname>Erdmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Maedche</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.-P.</given-names>
            <surname>Schnurr</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Staab</surname>
          </string-name>
          .
          <article-title>From manual to semi-automatic semantic annotation: About ontology-based text annotation tools</article-title>
          .
          <source>ETAI Journal - Section on Semantic Web</source>
          ,
          <volume>6</volume>
          ,
          <year>2001</year>
          . To appear.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>R.</given-names>
            <surname>Feldman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fresko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kinar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lindell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Liphstat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rajman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Schler</surname>
          </string-name>
          , and
          <string-name>
            <given-names>O.</given-names>
            <surname>Zamir</surname>
          </string-name>
          .
          <article-title>Text mining at the term level</article-title>
          .
          <source>In Proceedings of the Second European Symposium on Principles of Data Mining and Knowledge Discovery</source>
          , pages
          <fpage>65</fpage>
          -
          <lpage>73</lpage>
          , Nantes, France,
          <year>September 1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>W.</given-names>
            <surname>Gaul</surname>
          </string-name>
          and
          <string-name>
            <given-names>L.</given-names>
            <surname>Schmidt-Thieme</surname>
          </string-name>
          .
          <article-title>Mining web navigation path fragments</article-title>
          .
          <source>In [13]</source>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>R.</given-names>
            <surname>Goldman</surname>
          </string-name>
          and
          <string-name>
            <surname>J. Widom.</surname>
          </string-name>
          <article-title>DataGuides: Enabling query formulation and optimization in semistructured databases</article-title>
          .
          <source>In VLDB'97</source>
          , pages
          <fpage>436</fpage>
          -
          <lpage>445</lpage>
          , Athens, Greece, Aug.
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11. H.
          <string-name>
            <surname>Graubitz</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Spiliopoulou</surname>
            , and
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Winkler</surname>
          </string-name>
          .
          <article-title>The DIAsDEM framework for converting domain-specific texts into XML documents with data mining techniques</article-title>
          .
          <source>In Proceedings of the First IEEE International Conference on Data Mining</source>
          , San Jose, CA, USA, November/December 2001. To appear.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12. H.
          <string-name>
            <surname>Graubitz</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Winkler</surname>
            , and
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Spiliopoulou</surname>
          </string-name>
          .
          <article-title>Semantic tagging of domain-specific text documents with DIAsDEM</article-title>
          .
          <source>In Proceeding of the 1st International Workshop on Databases, Documents, and Information Fusion (DBFusion</source>
          <year>2001</year>
          ), pages
          <fpage>61</fpage>
          -
          <lpage>72</lpage>
          , Magdeburg, Germany, May
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <given-names>R.</given-names>
            <surname>Kohavi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Spiliopoulou</surname>
          </string-name>
          , and J. Srivastava, editors.
          <source>KDD'2000 Workshop WEBKDD'2000 on Web Mining for E-Commerce - Challenges and Opportunities</source>
          , Boston, MA, Aug.
          <year>2000</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <given-names>P. A.</given-names>
            <surname>Laur</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Masseglia</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Poncelet</surname>
          </string-name>
          .
          <article-title>Schema mining: Finding regularity among semistructured data</article-title>
          . In D. A.
          <string-name>
            <surname>Zighed</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Komorowski</surname>
          </string-name>
          , and J. Z˙ ytkow, editors,
          <source>Principles of Data Mining and Knowledge Discovery: 4th European Conference, PKDD</source>
          <year>2000</year>
          , volume
          <volume>1910</volume>
          <source>of Lecture Notes in Artificial Intelligence</source>
          , pages
          <fpage>498</fpage>
          -
          <lpage>503</lpage>
          , Lyon, France,
          <year>September 2000</year>
          . Springer, Berlin, Heidelberg.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <given-names>S.</given-names>
            <surname>Loh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. K.</given-names>
            <surname>Wives</surname>
          </string-name>
          , and
          <string-name>
            <surname>J. P. M.</surname>
          </string-name>
          <article-title>d</article-title>
          . Oliveira.
          <article-title>Conceptbased knowledge discovery in texts extracted from the Web</article-title>
          .
          <source>ACM SIGKDD Explorations</source>
          ,
          <volume>2</volume>
          (
          <issue>1</issue>
          ):
          <fpage>29</fpage>
          -
          <lpage>39</lpage>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <given-names>J.</given-names>
            <surname>Lumera</surname>
          </string-name>
          .
          <article-title>Große Mengen an Altdaten stehen XMLUmstieg im Weg</article-title>
          . Computerwoche,
          <volume>27</volume>
          (
          <issue>16</issue>
          ):
          <fpage>52</fpage>
          -
          <lpage>53</lpage>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <given-names>B.</given-names>
            <surname>Masand</surname>
          </string-name>
          and M. Spiliopoulou, editors.
          <source>Advances in Web Usage Mining and User Profiling: Proceedings of the WEBKDD'99 Workshop</source>
          ,
          <string-name>
            <surname>LNAI</surname>
          </string-name>
          <year>1836</year>
          . Springer Verlag,
          <year>July 2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <given-names>A.</given-names>
            <surname>Mikheev</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Finch</surname>
          </string-name>
          .
          <article-title>A workbench for acquisition of ontological knowledge from natural language</article-title>
          .
          <source>In Proceedings of the Seventh conference of the European Chapter for Computational Linguistics</source>
          , pages
          <fpage>194</fpage>
          -
          <lpage>201</lpage>
          , Dublin, Ireland,
          <year>March 1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19. U. Y. Nahm and
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Mooney</surname>
          </string-name>
          .
          <article-title>Using information extraction to aid the discovery of prediction rules from text</article-title>
          .
          <source>In Proceedings of the Sixth International Conference on Knowledge Discovery and Data Mining (KDD2000) Workshop on Text Mining</source>
          , pages
          <fpage>51</fpage>
          -
          <lpage>58</lpage>
          , Boston, MA, USA,
          <year>August 2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <given-names>A.</given-names>
            <surname>Nanopoulos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Katsaros</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Manolopoulos</surname>
          </string-name>
          .
          <article-title>Effective prediction of web-user accesses: A data mining approach</article-title>
          .
          <source>In Proceeding of the Workshop WEBKDD</source>
          <year>2001</year>
          :
          <article-title>Mining Log Data Across All Customer TouchPoints</article-title>
          , San Francisco, CA, USA,
          <year>August 2001</year>
          . To appear.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <given-names>S.</given-names>
            <surname>Nestrov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Abiteboul</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Motwani</surname>
          </string-name>
          .
          <article-title>Inferring structure in semi-structured data</article-title>
          .
          <source>SIGMOD Record</source>
          ,
          <volume>26</volume>
          (
          <issue>4</issue>
          ):
          <fpage>39</fpage>
          -
          <lpage>43</lpage>
          ,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <given-names>H.</given-names>
            <surname>Schmid</surname>
          </string-name>
          .
          <article-title>Probabilistic part-of-speech tagging using decision trees</article-title>
          .
          <source>In Proceedings of International Conference on New Methods in Language Processing</source>
          , pages
          <fpage>44</fpage>
          -
          <lpage>49</lpage>
          , Manchester, UK,
          <year>September 1994</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <given-names>A.</given-names>
            <surname>Sengupta</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Purao</surname>
          </string-name>
          .
          <article-title>Transitioning existing content: Inferring organization-spezific document structures. In K. Turowski and</article-title>
          K. J. Fellner, editors,
          <source>Tagungsband der 1. Deutschen Tagung XML</source>
          <year>2000</year>
          ,
          <article-title>XML Meets Business</article-title>
          , pages
          <fpage>130</fpage>
          -
          <lpage>135</lpage>
          , Heidelberg, Germany, May
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <given-names>M.</given-names>
            <surname>Spiliopoulou</surname>
          </string-name>
          .
          <article-title>The laborious way from data mining to web mining</article-title>
          .
          <source>Int. Journal of Comp</source>
          . Sys.,
          <string-name>
            <surname>Sci</surname>
          </string-name>
          . &amp;
          <string-name>
            <surname>Eng</surname>
          </string-name>
          ., Special Issue on “
          <source>Semantics of the Web”</source>
          ,
          <volume>14</volume>
          :
          <fpage>113</fpage>
          -
          <lpage>126</lpage>
          , Mar.
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <given-names>M.</given-names>
            <surname>Spiliopoulou</surname>
          </string-name>
          and
          <string-name>
            <given-names>L. C.</given-names>
            <surname>Faulstich</surname>
          </string-name>
          .
          <article-title>WUM: A Tool for Web Utilization Analysis</article-title>
          .
          <source>In extended version of Proc. EDBT Workshop WebDB'98, LNCS 1590</source>
          , pages
          <fpage>184</fpage>
          -
          <lpage>203</lpage>
          . Springer Verlag,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>A</surname>
          </string-name>
          .
          <string-name>
            <surname>-H. Tan</surname>
          </string-name>
          .
          <article-title>Text mining: The state of the art and the challenges</article-title>
          .
          <source>In Proceedings of the PAKDD 1999 Workshop on Knowledge Disocovery from Advanced Databases</source>
          , pages
          <fpage>65</fpage>
          -
          <lpage>70</lpage>
          , Beijing, China,
          <year>April 1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <given-names>K.</given-names>
            <surname>Wang</surname>
          </string-name>
          and
          <string-name>
            <given-names>H.</given-names>
            <surname>Liu</surname>
          </string-name>
          .
          <article-title>Discovering structural association of semistructured data</article-title>
          .
          <source>IEEE Transactions on Knowledge and Data Engineering</source>
          ,
          <volume>12</volume>
          (
          <issue>3</issue>
          ):
          <fpage>353</fpage>
          -
          <lpage>371</lpage>
          , May/June 2000.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>