=Paper= {{Paper |id=Vol-99/paper-7 |storemode=property |title=Extraction of Semantic XML DTDs from Texts Using Data Mining Techniques |pdfUrl=https://ceur-ws.org/Vol-99/Karsten_Winkler-et-al.pdf |volume=Vol-99 |dblpUrl=https://dblp.org/rec/conf/kcap/WinklerS01 }} ==Extraction of Semantic XML DTDs from Texts Using Data Mining Techniques== https://ceur-ws.org/Vol-99/Karsten_Winkler-et-al.pdf
                Extraction of Semantic XML DTDs from Texts
                       Using Data Mining Techniques
                                              Karsten Winkler and Myra Spiliopoulou
                                                 Leipzig Graduate School of Management
                                                        Department of E-Business
                                                 Jahnallee 59, D-04109 Leipzig, Germany
                                                    {kwinkler,myra}@ebusiness.hhl.de




Abstract                                                               with related data sources. Unfortunately, most users are not
Although composed of unstructured texts, documents con-                willing to manually create metadata due to the efforts and
tained in textual archives such as public announcements, pa-           costs involved [7]. Thus, text mining techniques are required
tient records and annual reports to shareholders often share           that (semi-) automatically create semantic markup and tag
an inherent though undocumented structure. In order to fa-             documents accordingly.
cilitate efficient, structure-based search in archives and to en-
able information integration of text collections with related          In this paper, we present the KDD approach pursued in
data sources, this inherent structure should be made explicit          the research project DIAsDEM whose German acronym
as detailed as possible. Inferring a semantic and structured           stands for “Data Integration for Legacy Systems and Semi-
XML document type definition (DTD) for an archive and                  Structured Documents by Means of Data Mining Tech-
subsequently transforming the corresponding texts into XML             niques”. Our goal is semantic tagging of textual content
documents is a successful method to achieve this objective.            with meta-data to facilitate searching, querying, identifica-
The main contribution of this paper is a new method to de-             tion of and integration with associated texts and relational
rive structured XML DTDs in order to extend previously de-             data. Hence, we aim at deriving a structured XML DTD
rived flat DTDs. We use the DIAsDEM framework to derive                that serves as a quasi-schema for the document collection
a preliminary, unstructured XML DTD whose components                   and enables the provision of database-like querying services
are supported by a large number of documents. However, all             on textual data. DIAsDEM focuses on text collections with
XML tags contained in this preliminary DTD cannot a priori             domain-specific vocabulary and syntax that frequently share
be assumed to be mandatory. Additionally, there is no fixed            an inherent, but undocumented structure.
order of XML tags and automatically tagging an archive us-
ing a derived DTD always implicates tagging errors. Hence,             The DIAsDEM framework for semantic tagging of domain-
we introduce the notion of probabilistic XML DTDs whose                specific texts was introduced in [12, 11]. However, apply-
components are assigned probabilities of being semantically            ing the Java-based DIAsDEM Workbench to a text archive
and structurally correct. Our method for establishing a prob-          currently results in a collection of semantically tagged XML
abilistic XML DTD is based on discovering associations be-             documents that are described by the extracted flat, unstruc-
tween, resp. frequent sequences of XML tags.                           tured XML DTD. However, we ultimately aim at integrat-
                                                                       ing the resulting XML documents with other related data
Keywords                                                               sources. In this context, the derived unstructured, rather
semantic annotation, XML, DTD derivation, knowledge dis-               preliminary DTD should be transformed into more struc-
covery, data mining, clustering                                        tured DTD that reflects both ordering and optionality of tags.
                                                                       Given that all XML tags are derived by data mining tech-
INTRODUCTION                                                           niques (i.e. iterative clustering as explained in section 3),
Most organizations are not only “drowning” in data, they are           they are not crisp due to tagging errors. Taking this critical
also “struggling” to cope with huge amounts of text docu-              fact into account, we introduce the notion of a probabilis-
ments. Tan points out that up to 80% of a company’s in-                tic DTD that describes the most likely orderings of XML
formation is stored in unstructured textual documents [26].            tags and that contains statistical properties for each tag. The
Hence, capturing interesting and actionable knowledge from             structured DTD will be the basis for future information in-
textual databases is a major challenge for the data mining             tegration efforts that involve XML archives generated by the
community. Creating semantic markup is one form of pro-                DIAsDEM Workbench. We introduce two algorithms for in-
viding explicit knowledge about text archives to facilitate            ferring a probabilistic DTD that utilize association rule dis-
searching and browsing or to enable information integration            covery algorithms and sequence mining techniques.
   
   The  work of this author is funded by the German Research Society
(DFG grant no. SP 572/4-1).                                            The rest of this paper is organized as follows: The next
                     
                     

                       Der Betrieb von Spielhallen in Teltow und das
                     Aufstellen von Geldspiel- und Unterhaltungsautomaten.   Stammkapital: 25.000 EUR. 
                      Gesellschaft mit beschränkter Haftung. 
                      Der Gesellschaftsvertrag ist am 12.11.1998
                     abgeschlossen und am 19.04.1999 abgeändert.  (...) Einzelvertretungsbefugnis kann erteilt
                     werden. 
                     Pawel Balski, 14.04.1965, Berlin, ist zum Geschäftsführer bestellt.  (...)
                      Nicht eingetragen: Die Bekanntmachungen der Gesellschaft erfolgen im Bundesanzeiger.
                      




                       Table 1: XML document containing an annotated Commercial Register entry



section briefly discusses related work. Section 3 gives an                      word frequency statistics. These collections can hardly be
overview of our framework for semantic tagging of domain-                       processed using standard text mining software because the
specific text collections. Section 4 introduces the notion of                   integration of relevant domain knowledge is a prerequisite
probabilistic DTDs for textual archives and develops two                        for successful knowledge discovery.
methods for deriving them. Finally, we conclude and give
directions for future research in section 5.                                    There are only a few research activities aiming at the trans-
                                                                                formation of texts into semantically annotated XML doc-
RELATED WORK
                                                                                uments: Becker et al. introduce the search engine GET-
Nahn and Mooney propose the combination of methods from
                                                                                ESS that supports query processing on texts by deriving and
KDD and information extraction to perform text mining tasks
                                                                                processing XML text abstracts [4]. These abstracts con-
[19]. They apply standard KDD techniques to a collec-
                                                                                tain language-independent, content-weighted summaries of
tion of structured records that contain previously extracted,
                                                                                domain-specific texts. In DIAsDEM, we do not separate
application-specific features from texts. Feldman et al. pro-
                                                                                meta-data from original texts but rather provide a seman-
pose text mining at the term level instead of focusing on lin-
                                                                                tic annotation, keeping the texts intact for later processing
guistically tagged words [8]. The authors represent each doc-
                                                                                or visualization. Given the aforementioned linguistic partic-
ument by a set of terms and additionally construct a taxon-
                                                                                ularities of the application domains we investigate, a DTD
omy of terms. The resulting dataset is input to KDD algo-
                                                                                characterizing the content of the documents is more appro-
rithms such as association rule discovery. Our DIAsDEM
                                                                                priate than inferences on their content. In order to transform
framework adopts the idea of representing texts by terms and
                                                                                existing content into XML documents, Sengupta and Purao
concepts. However, our goal is the semantic tagging of struc-
                                                                                propose a method that infers DTDs by using already tagged
tural text units (e.g., sentences or paragraphs) within the doc-
                                                                                documents as input [23]. In contrast, we propose a method
ument according to a global DTD and not the characteriza-
                                                                                that tags plain text documents and derives a DTD for them.
tion of the entire document’s content. Loh et al. suggest to
                                                                                Closer to our approach is the work of Lumera, who uses key-
extract concepts rather than individual words for subsequent
                                                                                words and rules to semi-automatically convert legacy data
use in KDD efforts at the document level. [15]. Similarly to
                                                                                into XML documents [16]. However, his approach relies on
our framework, the authors suggest to exploit existing vocab-
                                                                                establishing a rule base that drives the conversion, while we
ularies such as thesauri for concept extraction. Mikheev and
                                                                                use a KDD methodology that reduces human effort.
Finch describe a workbench to acquire domain knowledge
from texts [18]. Similar to the DIAsDEM Workbench, their
approach combines methods from different fields of research                     Semi-structured data is another topic of related research
in a unifying framework.                                                        within the database community [6, 1]. A lot of effort has
                                                                                recently been put into methods inferring and representing
Our approach shares with this research thread the objective of                  structure in similar semi-structured documents [21, 27, 14].
extracting semantic concepts from texts. However, concepts                      However, these approaches only derive a schema for a given
to be extracted in DIAsDEM must be appropriate to serve as                      set of semi-structured documents. In DIAsDEM, we have to
elements of the XML DTD. Among other implications, dis-                         simultaneously solve the problems of both semi-structuring
covering a concept that is peculiar to a single text unit is not                text documents by semantic tagging and inferring an ap-
sufficient for our purposes, although it may perfectly reflect                  propriately structured XML DTD that describes the related
the corresponding content. In order to derive a DTD, we need                    archive. We are not aware of any scientific or commercial ap-
to discover groups of text units that share some semantic con-                  proaches employing probabilistic document type definitions
cepts. Moreover, we concentrate on domain-specific texts,                       as introduced in this paper for describing text archives or in-
which significantly differ from average texts with respect to                   tegrating texts with related data sources.
                                                                                              ===         ====         ===
                                                                                =======                                ===                        ====
THE DIAsDEM FRAMEWORK                                                           ==========
                                                                                ==========
                                                                           =======
                                                                                ==========
                                                                                =======
                                                                                              ===
                                                                                              ===
                                                                                              ===    ==    ===
                                                                                                                ===    ===
                                                                                                                       ===
                                                                                                                       ===
                                                                                                                                                  ====                     Person = ==========

                                                                                                                                                                           Date =
                                                                                                                                                                                       ==========
                                                                                                                                                                                       ==========
                                                                                                                                                                                    \==\==\==
                                                                           ==========
                                                                                 =========    ===                      ===                                                          \==\==\====
                                                                           ==========
                                                                                ==========                                                                                          \==\=====\==
                                                                       =======
                                                                           ==========
                                                                                ==========
                                                                           =======
                                                                                ========                                               ====       ====    ====
                                                                       ==========
                                                                            =========
                                                                                 =========                                             ====       ====    ====             Corporation = =====
In this paper, the notion of semantic tagging refers to the            ==========
                                                                           ==========
                                                                                ==========
                                                                       ==========
                                                                           ==========
                                                                       =======
                                                                           ========
                                                                        =========
                                                                            =========
                                                                       ==========
                                                                           ==========
                                                                       ==========             ===         ===         ===
                                                                                                                                                                                            =====
                                                                                                                                                                                            =====
                                                                                                                                                                           Place = ====, ====,
                                                                                                                                                                                    ===, ===, ===,
                                                                       ========                                                                                                     ======, ====

activity of annotating texts with domain-specific XML tags              =========
                                                                       ==========                                                     ==== ====
                                                                                                                                      ==== ====
                                                                                                                                                     ==== ====
                                                                                                                                                     ==== ====             Currency =    \====\==
                                                                                                                                                                                         \===\==


                                                                     Text Documents           UML Schema                                 Thesaurus                  Entity Descriptions
that might contain additional attributes. Rather than classi-
fying entire documents or tagging single terms, we aim at
semantically tagging text units such as sentences or para-            Preprocessing:                NLP Preprocessing and Creation of Text Units
graphs. Table 1 illustrates this concept of semantic tagging,                                       Extraction and Replacement of Named Entities
                                                                                                    Selection of Features (Text Unit Descriptors)
whereas each sentence of this German Commercial Register                                            Mapping of Text Units into Feature Vectors

entry is a text unit. In this example, the semantics of most
sentences are made explicit by XML tags that partly con-
tain additional attributes describing extracted named entities                                Clustering:                                 Setting of Parameters
(e.g., names of persons and amounts of money). The XML                                                                                    Execution of Algorithm
                                                                                                                                          Evaluation of Cluster Quality
document depicted in Table 1 was created by applying the                                                                                  Cluster Inspection

DIAsDEM framework to a collection of 1,145 textual Com-
mercial Register entries containing 10,785 text units. This           Persons:
                                                                                                           + ====             +
collection includes all entries related to foundations of com-          ==== =============
                                                                        ==== ============
                                                                        ==== =============
                                                                        ==== ========
                                                                                                             ====
                                                                                                          ==== ==
                                                                                                          ==== ==
                                                                                                                             ===
                                                                                                                             ===
                                                                                                                                                                  _ ====
                                                                                                                                                                    ====
                                                                                                                                                                 ==== ==
                                                                                                                                                                 ==== ==
                                                                                                                                                                               _
                                                                                                                                                                              ===
                                                                                                                                                                              ===
                                                                        ==== =============                                   ======
                                                                        ==== =============
panies in the district of the German city Potsdam in 1999.              ==== ===========
                                                                        ==== ============

                                                                      Dates:
                                                                                                            + ====
                                                                                                               ====
                                                                                                                             ======
                                                                                                                               ====
                                                                                                                               ====
                                                                                                                             ====
                                                                                                                                                                  _ ====
                                                                                                                                                                     ====
                                                                                                                                                                              ======
                                                                                                                                                                              ======
                                                                                                                                                                                ====
                                                                                                                                                                                ====
                                                                                                                             ====                                             ====

In Germany, companies are obliged by law to submit vari-                ==== ==.======.===
                                                                        ==== ==.==.===
                                                                        ==== ==.=======.===
                                                                        ==== ==.==.===
                                                                                                           ===== ==
                                                                                                           ===== ==
                                                                                                             =====
                                                                                                                                                                 ===== ==
                                                                                                                                                                 ===== ==
                                                                                                                                                                   =====
                                                                                                                                                                              ====




ous information about business affairs to local Commercial           Named Entities                 Acceptable Clusters                              Unacceptable Clusters

Registers. Although Commercial Registers are an important
source of information in daily business transactions, their
textual content can only be searched using full-text queries at       Postprocessing:                 Cleansing and Refinement of Clusters
                                                                                                      Semantic Labeling of Acceptable Clusters
the moment. Hence, semantically semi-structuring these tex-                                           XML Tagging of Text Units
                                                                                                      Creation of DTD and XML Documents
tual archives provides the basis for information integration
and creation of value-adding services related to information
brokerage. XML query languages could be employed to sub-                                                                                                                   <−>====<\>
                                                                                                                                                                           <−>=======
                                                                        1  ====
                                                                           ====       2                         <========>                                                 ==========
                                                                                                                                                                      <−>====<\>
                                                                                                                                                                           ==========

mit both both content- and structure-based queries against              ==== ==
                                                                        ==== ==
                                                                                     ===
                                                                                     ===
                                                                                     ======
                                                                                     ======
                                                                                                                <=====>
                                                                                                                <=======>
                                                                                                                <=======>
                                                                                                                <====>
                                                                                                                <======>
                                                                                                                                                                           =====<\>
                                                                                                                                                                      <−>=======
                                                                                                                                                                             <−>======
                                                                                                                                                                      ==========
                                                                                                                                                                           ==========
                                                                                                                                                                  <−>====<\>
                                                                                                                                                                      ==========
                                                                                                                                                                           ==========
                                                                                                                                                                      =====<\>
                                                                                                                                                                           =======<\>
                                                                                                                                                                  <−>=======
                                                                                                                                                                       <−>======
                                                                                                                                                                  ========== <−>======
                                                                                                                                                                      ==========
                                                                                                                                                                           =======<\>
                                                                                                                                                                  ==========
                                                                                                                <======>                                              ==========
semantically tagged XML archives.                                        3  ====
                                                                            ====
                                                                        ===== ==
                                                                                       ====
                                                                                       ====
                                                                                     ====
                                                                                     ====
                                                                                                                <========>
                                                                                                                <=======>
                                                                                                                                                                  =====<\>
                                                                                                                                                                      =======<\>
                                                                                                                                                                   <−>======
                                                                                                                                                                       <−>======
                                                                                                                                                                  ==========
                                                                                                                                                                      =======<\>
                                                                                                                                                                  ==========
                                                                                                                                                                  =======<\>
                                                                        ===== ==                                                                                   <−>======
                                                                                                                                                                  =======<\>
                                                                          =====
                                                                                                    XML Document
                                                                    Text Unit Clusterer             Type Definition                                        XML Documents
Our framework pursues two objectives for a given archive
of text documents: All text documents should be semanti-
                                                                       Figure 1: Iterative and interactive KDT process
cally tagged and an appropriate, preliminary flat XML DTD
should be derived for the archive. Semantic tagging in DIAs-
DEM is a two-phase process. We have designed a knowledge
discovery in textual databases (KDT) process that constitutes     In the remainder of this section, we briefly introduce the first
the first phase in order to build clusters of semantically sim-   phase of the DIAsDEM framework whose iterative and in-
ilar text units, to tag documents in XML according to the         teractive KDT process is depicted in Figure 1. This process
results and to derive an XML DTD describing the archive.          is termed “iterative” because the clustering algorithm is in-
The KDT process that was introduced in [12, 11] results in a      voked repeatedly. Our notion of iterative clustering should
final set of clusters whose labels serve as XML tags and DTD      not be confused with the fact that most clustering algorithms
elements. Huge amounts of new documents can be converted          perform multiple passes over the data before converging.
into XML documents in the second, batch-oriented and pro-         This process is also “interactive”, because a knowledge engi-
ductive phase of the DIAsDEM framework. All text units            neer is consulted for cluster evaluation and final cluster nam-
contained in new documents are clustered by the previously        ing decisions at the end of each iteration.
built text unit clusterer and are subsequently tagged with the
corresponding cluster labels.                                     Besides the initial text documents to be tagged, the follow-
                                                                  ing domain knowledge constitutes input to our KDT process:
In DIAsDEM we concentrate on the semantic tagging of              A thesaurus containing a domain-specific taxonomy of terms
similar text documents originating from a common domain.          and concepts, a preliminary UML schema of the domain and
Nevertheless, the DIAsDEM approach is appropriate for se-         descriptions of specific named entities of importance, e.g.
mantically tagging various kinds of archives such as pub-         persons and companies. The UML schema reflects the se-
lic announcements of courts and administrative authorities,       mantics of named entities and the relationships among them,
quarterly and annual reports to shareholders, textual patient     as they are initially conceived by application experts. This
records in health care applications as well as product and ser-   schema serves as a reference for the DTD to be derived from
vice descriptions published on electronic marketplaces.           discovered semantic tags, but there is no guarantee that the
                    

                    

                    
                     (...)
                    




                     Table 2: Preliminary flat, unstructured XML DTD of Commercial Register entries



final DTD will be contained in or will contain this schema.           clusters containing approx. 85% of text units.

Similarly to a conventional KDD process, our process starts           The postmining phase consists of a labeling step, in which
with a preprocessing phase: After setting the level of gran-          “acceptable” clusters are semi-automatically assigned a la-
ularity by determining the size of text units to be tagged,           bel. Ultimately, cluster labels are determined by the knowl-
the Java- and Perl-based DIAsDEM Workbench performs ba-               edge engineer. However, the DIAsDEM Workbench performs
sic NLP preprocessing such as tokenization, normalization             both a pre-selection and a ranking of candidate cluster la-
and word stemming using TreeTagger [22]. Instead of re-               bels for the expert to choose from. All default cluster labels
moving stop words, we establish a drastically reduced fea-            are derived from feature space dimensions (i.e. from text
ture space by selecting a limited set of terms and concepts           unit descriptors) that are prevailing in each “acceptable” clus-
(so-called text unit descriptors) from the thesaurus and the          ter. Cluster labels actually correspond to XML tags that are
UML schema. Text unit descriptors are currently chosen                subsequently used to annotate cluster members. Finally, all
by the knowledge engineer because they must reflect impor-            original documents are tagged using valid XML tags. Addi-
tant concepts of the application domain. All text units are           tionally, XML tags are enhanced by attributes reflecting pre-
mapped into Boolean vectors of this feature space. Addition-          viously extracted named entities and their values. Table 2
ally, named entities of interest are extracted from text units by     contains an excerpt of the flat, unstructured XML DTD that
a separate module of the DIAsDEM Workbench. In our case               was automatically derived from XML tags in the case study.
study, we created a small thesaurus and selected 70 relevant          It coarsely describes the semantic structure of the resulting
descriptors and 109 non-descriptors pointing to descriptors.          XML collection. Currently, named entities that serve as ad-
                                                                      ditional attributes of XML tags are not fully evaluated by the
In the pattern discovery phase, all text unit vectors contained       DIAsDEM Workbench.
in the initial archive are clustered based on similarity of their
content. The objective is to discover dense and homogeneous           ESTABLISHING A PROBABILISTIC DTD
text unit clusters. Clustering is performed in multiple iter-         The output of the DIAsDEM Workbench is a set of semantic
ations. Each iteration outputs a set of clusters, which the           XML tags which should be used as XML tags to describe the
DIAsDEM Workbench partitions into ”acceptable” and ”un-               content of the archive documents. To reflect the content of
acceptable” ones according to our quality criteria. A cluster         the archive at an abstract level, it is essential to compose the
of text unit vectors is ”acceptable”, if and only if (i) its cardi-   tags into a DTD. Since the semantic annotations are derived
nality is large and the corresponding text units are (ii) homo-       with data mining techniques, they are not crisp. Thus, it is
geneous and (iii) can be semantically described by a small            essential that the validity of each tag is expressed in quanti-
number of text unit descriptors. Members of “acceptable”              tative terms and is estimated properly. Furthermore, an order-
cluster are subsequently removed from the dataset for later           ing should be imposed upon the tags. Hence, after deriving
labeling, whereas the remaining text unit vectors are input           semantic XML tags, we combine them into a probabilistic
data to the clustering algorithm in the next iteration. In each       DTD by (i) deriving the most likely ordering of the tags and
iteration, the cluster similarity threshold value is stepwise         (ii) computing the statistical properties of each tag inside the
decreased such that “acceptable” clusters become progres-             document type definition.
sively less specific in content. The KTD process is based on a
plug-in concept that allows the execution of different cluster-       The reader may recall that a semantic annotation is actu-
ing algorithms within the DIAsDEM Workbench. In the case              ally the label of a cluster discovered by the DIAsDEM Work-
study, we employed the demographic clustering function in-            bench. The underlying clustering mechanism produces non-
cluded in the IBM Intelligent Miner for Data that maximizes           overlapping clusters. This implies that a text unit belongs to
the value of Condorcet’s criterion. After three iterations, the       exactly one cluster, to the effect that it can be annotated with
DIAsDEM Workbench discovered altogether 73 “acceptable”               the label of this cluster only. Hence, the tags/labels derived
the DIAsDEM Workbench cannot be nested. An extension of             of error type II text units is higher, indicating that some text
the DIAsDEM Workbench by a hierarchical clustering algo-            units were not placed in the cluster they semantically belong
rithm would allow for the establishment of subclusters and          to. With 0.95 confidence, the overall error rate in the entire
thus for the nesting of (sub)cluster labels. However, this is       dataset is in the interval [2.6%, 5.9%] which is a promising
planned as future work.                                             result.

The objectives of the DTD establishment method are the              TagSupport     The tags of the DTD are cluster labels derived
specification of the most appropriate ordering of tags, the         by a statistical approach. Thus, in terms of XML, they are
identification of correlated or mutually exclusive tags and the     observed as optional per se. In many application areas, a do-
adornment of each tag and each correlation among tags with          main expert can provide suggestions as to which tags should
statistical properties. These properties form the basis for re-     be observed as mandatory. Despite this, there is no guaran-
liable query processing, because they determine the expected        tee that the expert’s suggestions hold true in the archive: The
precision and recall of the query results. In the following, we     text unit containing this information may have been misclas-
first introduce the statistical properties we consider for the      sified by the DIAsDEM Workbench, or the information may
DTD tags and their associations and describe the methodol-          be simply absent from the document. For example, although
ogy for computing these statistics. To model the complete           one would expect that each movie has a regisseur, there are
statistical information pertinent in these tags and their rela-     movies whose regisseur is unknown or inapplicable, due to
tionships, we use a hypergraph structure. We then introduce a       the nature of the movie. The property TagSupport offers an
mechanism that derives a probabilistic DTD from this graph.         indicator of whether a tag may be considered as potentially
                                                                    mandatory. We define it as the ratio of documents where this
Statistical Properties of Semantic XML Tags                         tag appears to the total number of documents in the archive.
The statistical properties of DTD tags are depicted in Table 3
and described in the following paragraphs. The first column         AssociationConfidence In association rules’ discovery, the
contains the names of the properties. The second column             miner identifies items occuring together. Equivalently, we
reflects whether the property is peculiar to the whole set of       are interested in tags that affect the appearance of other tags.
tags as cluster labels (i.e. the whole ”model”), to each tag        We use the term AssociationConfidence for a tag x given the
or to a group of associated tags. The last column names the         tags y1 ; : : : ; yn in much the same way as confidence is de-
mechanism to be applied to derive the value of each property        fined for association rules [5]: It is the ratio of documents,
for each tag.                                                       where the tags y1 ; : : : ; yn and x appear to the documents con-
                                                                    taining y1 ; : : : ; yn .
Accuracy     The DIAsDEM Workbench derives semantic
XML tags as labels of clusters. These clusters constitute a         AssociationLift      Similarly to the association rules’
model over the data, in the conventional statistical sense. In      paradigm, the correlation among a tag x and a set of
terms of data classification, such models are subject to mis-       tags y1 ; : : : ; yn can be spurious, caused by a very high
classification errors. We identify two types of misclassifica-      support of x in the whole population. The statistic called lift
tion:                                                               or improvement is defined to alleviate this problem: it is the
 Error type I: A text unit is assigned to the wrong cluster,       ratio of the AssociationConfidence of x given y 1 ; : : : ; yn to
                                                                    the TagSupport of x in the whole population [5]. In our case,
  i.e. the cluster label does not reflect the content of the text                                          onf idence(x;y1 ;:::;yn )

  unit.                                                             this would be the ratio AssociationC
                                                                                                       T agSupport(x)
                                                                                                                                     .
 Error type II: A text unit is not assigned to any cluster,        LocationConfidence          The aforementioned statistical proper-
  although there is a cluster with a label reflecting the content   ties on associated tags do not take the ordering of tags into
  of the text unit.                                                 account. In a DTD, the ordering of tags is essential. We use
For the envisaged DTD, only the error type I is relevant. We        the term LocationConfidence of a tag x given the sequence of
use the term accuracy of the model as the probability that          adjacent tags y1  y2  : : :  yn as the number of documents con-
cluster labels reflect the content of cluster members. The ac-      taining the sequence y 1  y2  : : :  yn  x to the number of docu-
curacy value affects the DTD as a whole, it is not peculiar to      ments containing y 1  y2  : : :  yn . This definition differs from
individual tags. Therefore, we do not incorporate this value        the conventional statistics known for sequence mining [2],
in the statistical adornment of the individual tags.                because we are concentrating on adjacent tags, disallowing
                                                                    the occurrence of arbitrary tags in-between. Conventional
In order to evaluate the quality of out approach in absence of      sequence mining do not satisfy this requirement. However,
pre-tagged documents, we drew a random sample containing            some Web usage miners have been designed to distinguish
5% out of 10,785 text units and asked a domain specialist to        between adjacent and non-adjacent events [3, 9, 24, 20].
verify the annotations of these text units with respect to both
error types. Within the sample, error type I (error type II) oc-    GroupSupport     In most of the above statistics, we juxtapose
cured in 0.4% (3.6%) of text units. Hence, tagged text units        the frequence of appearance of a tag with the frequency of
are most likely to be correctly processed. The percentage           a group of tags, be it a set or a sequence. We use the term
                      Property                    Radius                    Computation method
                      Accuracy                    model                     DIAsDEM Workbench
                      TagSupport                  tag                       simple statistics
                      AssociationConfidence       set of tags               association rule discovery
                      AssociationLift             set of tags               association rule discovery (ARD)
                      LocationConfidence          sequence of tags          sequence mining (SeqM)
                      GroupSupport                set or sequence of tags   ARD/SeqM

                                             Table 3: Statistics for derived XML tags



GroupSupport as the ratio of the number of documents con-            lists. Second, we annotate each list with a flag that indicates
taining a group of tags to the total number of documents. In         whether the group depicted by the list is order-sensitive or
fact, for any set of at least two tags, this property assumes one    order-insensitive; in the latter case, the ordering of the list is
value for the set and as many values as are the perturbations        irrelevant but must be unique. Third, we guarantee unique-
of set members. In the following subsection, we show how             ness, i.e. that all permutations of the same group of tags are
we model the statistical information pertinent to individual         mapped into the same order-insensitive list, by requiring that
tags, to tag groups (i.e. sets or sequences) and to relation-        order-insensitive list are lexicographically ordered.
ships among them in a seamless way.
                                                                     More formally, let P (V ) be the set of all lists of elements in
Modeling Statistics of Associated XML Tags                           V , i.e. (TagName,TagSupport)-pairs. An x 2 P (V )  f0; 1g
Some of the values of the statistical properties depicted in         has the form (< v1 ; : : : ; vk >; 1), where < v1 ; : : : ; vk >
Table 3 are already made available as part of the DIAsDEM            is a list of elements from V and the value 1 indicates
Workbench output, while the remaining ones can be com-               that this list represents an order-sensitive group. Similarly,
puted by data mining algorithms. To exploit these values for         x0 = (< v1 ; : : : ; vk >; 0) would represent the unique order-
the establishment of a probabilistic DTD, we need a repre-           insensitive group composed of v 1 ; : : : ; vk 2 V .
sentation model and an algorithm that builds the DTD when
processing this model. We introduce here a generic graph             For example, let a; b 2 V be two tags annotated with their
structure, in which all statistical information on tags, groups      TagSupport, whereby a precedes b lexicographically. The
of tags and tag relationships is depicted. This structure is ap-     groups (< a; b >; 1) and (< b; a >; 1) are two distinct order-
propriate for the establishment of a DTD or an XMLschema             sensitive groups of the two elements. The group (< a; b >
with rich statistical adornments. In the next subsection, we         ; 0) is the order-insensitive group of the two elements. Fi-
discuss two algorithms for DTD establishment.                        nally, the group (< b; a >; 0) is not permitted, because the
                                                                     group is order-insensitive but the list violates the (default)
We represent the tags and their associations in a directed
                                                                     lexicographical ordering of list elements.
graph. Its nodes are individual tags, sequences of adjacent
tags or sets of co-occuring tags. Each node is adorned with          Using P (V )  f0; 1g, we define V 0  (P (V )  f0; 1g) 
the statistical properties pertinent to a tag, resp. tag group.      (0; 1] with signature:
An edge represents a relationship of the form y 1 : : : yn ! x;
however, we use the convention that x is the source node and
the group of nodes in the rule’s LHS is the target. Similarly                   < GroupOf T ags; GroupSupport >
to nodes, an edge is adorned with the statistics of the order-
insensitive or order-sensitive association it represents.            where V 0 contains only those groups of annotations, for
                                                                     which the GroupSupport value is above a given threshold.
Semantic Tags as Graph Nodes     Let A be the set of seman-
                                                                     This threshold can be specified as input to the mining soft-
tic XML tags derived by the DIAsDEM Workbench, and let
V  A  (0; 1] be the set of graph nodes conforming to the
                                                                     ware, as is usual in KDD applications, or may be set as low
                                                                     as 0. Of course, the threshold value affects the size of the
signature:
                                                                     graph and the execution time of the algorithm that traverses
               < T agN ame; T agSupport >                            it to build the DTD.
By this definition, a tag can only appear in the graph if its
                                                                     Tag Relationships as Graph Edges The set of nodes consti-
                                                                     tuting our graph is = V V 0 , indicating that a node may be
                                                                                         V       [
TagSupport is more than zero. This is consistent with the
fact that XML tags are derived with a KDD method.
                                                                     a singleton tag or a group of tags with its/their statistics. An
For the groups of tags, we must distinguish among order-             edge emanates from an element of V and points to an element
sensitive and order-insensitive groups. To do so, we per-            of V 0 , i.e. from a tag to an associated group of tags. Formally,
form three steps. First, we model tag groups as ordered              the set of edges E is a subset of (V  V 0 )  X  X  X ,
where X := (0; 1] [ fN U LLg, with signature:                       DTD Derivation
                                                                    The DTD-establishment graph in its maximalistic version
           < Edge; AssociationConf idence;                          captures all relationships among the semantic tags found
        AssociationLif t; LocationConf idence >                     by the DIAsDEM Workbench. Similarly to the process of
                                                                    schema establishment for a conventional database applica-
In this signature, the statistical properties refer to the edge’s   tion, the designer must decide which relationships among
source given the group of nodes in the edge’s target. If the        the real-world entities are worth capturing and which are not.
target is a sequence of adjacent tags, then the location confi-     In our context, “worth capturing” refers to statistical values,
dence is the only valid statistical property, while the associa-    presuming that a DTD should reflect the relationships usually
tion confidence and lift are inapplicable. If the target is a set   present in the documents rather than the rare ones. However,
of tags, then the location confidence is inapplicable. When         the DTD-establishment graph contains relationships among
a statistical property is inapplicable, it assumes the NULL         sets and among sequences of tags, each one adorned with
value.                                                              different (and only partially comparable) statistics.

Graph Properties The components of our “DTD-                        In the following, we present two algorithms that derive a
establishment graph” are tags, groups of tags and rela-             DTD by constructing part of the DTD-establishment graph.
tionships among them, all adorned with statistical values.          Each tag of this DTD is adorned by only two (derived) prob-
All tags discovered by the DIAsDEM Workbench are present            abilistic values, one referring to the tag itself and one to its
in this graph. Which groups of tags are present depends on          location inside the DTD. The algorithms are using different
the threshold value for the group support. Conceivable are          heuristics to derive this DTD: the first one concentrates on
both a minimalistic approach with a high threshold, by which        the pairs of tags appearing most frequently together, while
only very frequent groups are present, and a maximalistic           the second one gives preference to maximal sequences of
approach with a zero-value threshold, by which all tag              tags. The reader may recall that the computation of the statis-
combinations occuring in the documents are present.                 tics for the relationships among the tags require the activation
                                                                    of data mining software. Hence, each of the algorithms is
If we opt for the minimalistic approach, the graph will not         backed by a miner that returns the desired statistics.
be connected in the general case. It will contain only the
groups of tags being more frequent than a threshold, and the        Backward Construction of DTD Sequences
frequent relationships among them. Certain tags may be iso-         This algorithm observes a DTD as a set of alternative se-
lated, because they only rarely appear in combination with          quences and builds each sequence backwards, starting at
other tags. Contrary to it, the maximalistic approach ensures       each last tag and proceeding until the first one. Concretely,
that all combinations of tags appearing together in documents       the algorithm builds “maximal” sequences, where maximal-
are depicted in the graph, and that the graph is connected,         ity means that the first tag of the sequence is the first tag in
except of the unlikely case that some documents contain a           most of the documents supporting the sequence.
single tag not occuring in any other documents.
                                                                    Backward expansion of tag-subsequences. The algorithm
The size of the graph depends on the threshold value for                                             2
                                                                    starts with an arbitrary tag  A and then identifies the tag
group support and for the confidence and lift values. The           most likely to appear before  :
maximalistic approach delivers an upper limit. To com-               If no such tag exists, then the sequence cannot be expanded
pute it, let m be the number of tags/cluster labels output            anymore. It is marked as done and the algorithm shifts
by the DIAsDEM Workbench and let n  m be the largest                 to the next sequence that is not done yet, or to the next
                             P  in any document. For
number of distinct tags appearing                                     arbitrary tag from A, until all tags are processed.
each tag, there are  1 :=
                              n
                              i=1
                                  1   n
                                            perturbations to be      If there is a most likely predecessor of  , say  0 , it is
                                      i                               prepended to the sequence. The next iteration starts, in
considered, resulting in an equal number of order-sensitive           which the most likely predecessor of  0   (in general: of
groups and in  2 := n(n2+1) order-insensitive ones. There            the subsequence built thus far) must be found.
is one edge per (Tag,Group)-pair. Moreover, a tag partic-            If there are k predecessor tags, none of which is more
ipates in a maximum of  1 + 2 groups, thus resulting in             likely than the others, k alternative incomplete sequences
m + m  (1 + 2 ) graph nodes and m  ( 1 + 2 ) edges.             are produced by duplicating the sequence built thus far.
The upper limit to the graph size indicates that threshold val-       The algorithm processes them iteratively.
ues for the statistical properties are essential for obtaining a    The predecessors of a tag  can be found by invoking a se-
manageable graph. On the other hand, each cutoff value im-          quence miner that returns all frequent sequences of adjacent
plies an information loss. Therefore, we observe the DTD-           tags. For the first iteration, the algorithm uses the frequent
establishment graph under the maximalistic approach as a            pairs x1  ; : : : ; xu   , each one leading to  with a (location)
reference structure and introduce two algorithms that derive        confidence c1 ; : : : ; cu respectively. SinceP     these tags are the
                                                                                                                          u
a probabilistic DTD by constructing only a part of this graph.      immediate predecessors of  it holds that i=1 ci = 1 c0 ,
where c0 is the ratio of documents where  has no predeces-        to the subsequence leading to it inside the tree: this is the
sor divided by the total number of documents.                      number of documents starting with this subsequence of tags.
                                                                   Similarly to the sequence-construction algorithm described
The maximum among c 0 ; : : : ; cu determines the rest of the      above, a tag may appear in more than one subsequences, hav-
procedure: If c 0 is maximum,  is the first element of the        ing different predecessors in each one.
sequence. In this case, the sequence is marked as “done” and
as “maximal” according to the maximality criterion already         Observing the DTD as a tree implies a common root. In
mentioned.                                                         the general case, each document of the archive may start at
                                                                   a different tag. We assume a dummy root, the children of
If there is a ci larger than the other elements, then x i is the   which are those tags that appear first in documents. In gen-
predecessor of  in the sequence. However, it can be the case      eral, a tree node refers to a tag  , and its children refer to
that the maximum is only marginally larger than the other          the tags appearing after  in the context of  ’s own predeces-
values. In other words, there are k tags with k  u, such          sors. In a sense, the DTD as a tree of alternatives resembles
that (i) one of them has shows the maximum location confi-         a DataGuide as proposed in [10], although the latter contains
dence but (ii) the difference of this value from the location      no statistical adornments.
confidences of the other k 1 tags is less than some small
". Then, all k tags are acceptable alternatives, resulting to k    The tree-of-alternatives differs from the sequence-
alternative subsequences.                                          construction algorithm in two ways: Firstly, it considers
                                                                   all sequences of tags that appear in documents instead
Identifying maximal tag-sequences.      In each iteration, the     of frequent ones only. Secondly, it only observes com-
algorithm considers longer frequent sequences returned by          plete sequences, while a sequence miner returns arbitrary
the sequence miner, namely those ending with each subse-           subsequences of tags.
quence already built. If no frequent sequence is found for
a subsequence s =  1 : : : k , then s is “done” but it must      The tree-of-alternatives method is realized by the preproces-
also be checked whether it is maximal. This implies com-           sor module of the Web usage miner WUM [25, 24]. This
puting the ratio of documents starting with  1 over the whole     module is responsible for coercing sequences of events by
number of documents and comparing this value x to the tag          common prefix and placing them in a tree structure, called
support of  1 , say c. The comparison is performed across         “aggregated tree”. This tree is input to the navigation pat-
the same guidelines as for alternative tag predecessors: if        tern discovery process performed by the WUM core. The se-
jx cj  ", then most documents containing s start with  1         quences of tags in documents can be observed as sequences
and thus s is maximal. Otherwise, documents starting with          of events, to the effect that the WUM preprocessor can also
1 mostly adhere to a different maximal sequence.                  be used to build a DTD over an archive as a tree of alterna-
                                                                   tive tag sequences. Figure 2 depicts an example of such a tree
Statistics of maximal tag-sequences.   At a final step, the al-    that related to our case study. Note that the XML document
gorithm filters out all sequences that are done but are not        depicted in Table 1 is partly described by this DTD excerpt.
maximal. It then assigns probability values to each tag  in-
side each maximal sequence s containing it:                        CONCLUSION
 The TagConfidence is the tag’s TagSupport multiplied by          Most of the knowledge hidden in electronic media of an orga-
  the accuracy of the model output by the DIAsDEM Work-            nization is encapsulated in documents. Acquiring this knowl-
  bench.                                                           edge implies effective querying of the documents as well as
 The TagPositionConfidence is the location confidence of          the combination of information pieces from different textual
  this tag with respect to the subsequence of s leading to it.     assets. This functionality is usually confined to database-
                                                                   like query processors, while text search engines scan indi-
Backward versus forward sequence construction.              The    vidual assets and return ranked results. In this study, we
backward-sequence-construction method generates alterna-           have presented a methodology that enables query processing
tive sequences of DTD tags by pruning the frequent se-             and joining of text sources by structuring them. We propose
quences of adjacent tags produced by a sequence miner. An          the derivation of an XML DTD over a domain-specific text
equivalent method can be devised by forward-sequence con-          archive by means of data mining techniques.
struction. This would have the advantage of being appropri-
ate for incorporation to a sequence miner’s core as well, since    The semantic characterization of text units is the core of our
most miners of this category perform forward sequence con-         approach as well as the derivation of XML tags from these
struction: In that case, the mining kernel would be modified       characterisations. This is undertaken by the DIAsDEM Work-
to expand a sequence by the most likely successor tag only.        bench which is concisely described in the first part of this
                                                                   study. Our main emphasis is on combining these tags that re-
A DTD as a Tree of Alternatives                                    flect the semantics of many text units across the archive into
This algorithm observes a DTD as a tree of alternative sub-        a single DTD that reflects the semantics of the archive as a
sequences and adorns each tag with its support with respect        whole. We have shown that this DTD is a probabilistic ap-
                                    BusinessPurpose, 97            ShareCapital, 661                      LimitedLiableCompany, 605            (...)



                                                                                   (...)                                           ModificationArticles_ShareCapital, 6
                                                                                                     Procuration, 39
                       BusinessPurpose, 950                 FullyLiablePartner, 95
                                                                                                                                 ConclusionArticlesOfAssociation,13
                                                                             LimitedLiabilityCompany, 123
                                                           (...)
                                                                                                   Procuration, 5             ConclusionArticles, 575
                                                        ShareCapital, 129


                                                                                                                    (...)
                                                                                           (...)

                                                                          PartnershipLimitedByShares, 20                              ResolutionByShareholders 11



                      Root, 1134                     FullyLiablePartner, 31                        Procuration, 7
                                                                                                                                ModificationArticles_MainOffice, 129


                                                                                           (...)
                                                                                                       AuthorityRepresentation_ManagingDirector, 401


                            (...)        Owner, 11          LimitedLiableCompany, 3




                                            Figure 2: A DTD as a tree of alternative tag sequences



proximation of the archive content and have derived a set of                                               REFERENCES
statistical properties that reflect the quality of this approxi-                                               1. S. Abiteboul, P. Buneman, and D. Suciu. Data on the
mation, for the whole DTD, for tags inside the DTD and for                                                        Web: From Relations to Semistructured Data and XML.
relationships among these tags.                                                                                   Morgan Kaufman Publishers, San Francisco, 2000.

The statistical properties of tags and of their relationships                                                  2. R. Agrawal and R. Srikant. Mining sequential patterns.
form the basis for combining them into a complete DTD                                                             In Proc. of Int. Conf. on Data Engineering, Taipei, Tai-
in the XML sense or even into an XMLschema. We use a                                                              wan, Mar. 1995.
graph structure to depict all statistics that can serve as a ba-                                               3. M. Baumgarten, A. G. Büchner, S. S. Anand, M. D.
sis for this operation and propose two mechanisms that derive                                                     Mulvenna, and J. G. Hughes. Navigation pattern dis-
DTDs by employing a mining algorithm and a set of heuristic                                                       covery from internet data. In [17], pages 70–87. 2000.
rules. We have tested our methodology on an archive of doc-
uments from a regional Commercial Register in Germany:                                                         4. M. Becker, J. Bedersdorfer, I. Bruder, A. Düsterhöft,
We have derived a set of tags with the DIAsDEM Workbench                                                          and G. Neumann. GETESS: Constructing a linguistic
and then implemented one of the proposed mechanisms to                                                            search index for an Internet search engine. In Proceed-
derive a DTD for it.                                                                                              ings of the 5th International Conference on Applica-
                                                                                                                  tions of Natural Language to Information Systems, Ver-
Our future work includes the implementation of the second                                                         sailles, France, June 2000.
mechanism for DTD derivement and the establishment of a
framework for the comparison of derived DTDs in terms of                                                       5. M. J. Berry and G. Linoff. Data Mining Techniques:
expressiveness and accuracy. Of course, the ultimate goal                                                         For Marketing, Sales and Customer Support. John Wi-
of our work is the establishment of a full-fledged querying                                                       ley & Sons, Inc., 1997.
mechanism over the text archives. To this purpose, we intend
                                                                                                               6. P. Buneman. Semistructured data. In Proceedings of
to couple our DTD derivation methods with a query mech-
                                                                                                                  the Sixteenth ACM SIGACT-SIGMOD-SIGART Sympo-
anism for semi-structured data. Since the DTDs we derive
                                                                                                                  sium on Principles of Database Systems, pages 117–
are of probabilistic nature, this implies also the design of a
                                                                                                                  121, Tucson, AZ, USA, May 1997.
model that evaluates the quality of the query results.
                                                                                                               7. M. Erdmann, A. Maedche, H.-P. Schnurr, and S. Staab.
ACKNOWLEDGMENTS                                                                                                   From manual to semi-automatic semantic annotation:
                                                                                                                  About ontology-based text annotation tools. ETAI Jour-
We thank the German Research Society for funding the                                                              nal - Section on Semantic Web, 6, 2001. To appear.
project DIAsDEM, the Bundesanzeiger Verlagsgesellschaft
mbH for providing data and our project collaborators Evgue-                                                    8. R. Feldman, M. Fresko, Y. Kinar, Y. Lindell, O. Liph-
nia Altareva and Stefan Conrad for helpful discussions. The                                                       stat, M. Rajman, Y. Schler, and O. Zamir. Text mining
IBM Intelligent Miner for Data is kindly provided by IBM in                                                       at the term level. In Proceedings of the Second Eu-
terms of the IBM DB2 Scholars Program.                                                                            ropean Symposium on Principles of Data Mining and
    Knowledge Discovery, pages 65–73, Nantes, France,         19. U. Y. Nahm and R. J. Mooney. Using information ex-
    September 1998.                                               traction to aid the discovery of prediction rules from
                                                                  text. In Proceedings of the Sixth International Confer-
 9. W. Gaul and L. Schmidt-Thieme. Mining web naviga-             ence on Knowledge Discovery and Data Mining (KDD-
    tion path fragments. In [13], 2000.                           2000) Workshop on Text Mining, pages 51–58, Boston,
                                                                  MA, USA, August 2000.
10. R. Goldman and J. Widom. DataGuides: Enabling
    query formulation and optimization in semistructured      20. A. Nanopoulos, D. Katsaros, and Y. Manolopoulos. Ef-
    databases. In VLDB’97, pages 436–445, Athens,                 fective prediction of web-user accesses: A data mining
    Greece, Aug. 1997.                                            approach. In Proceeding of the Workshop WEBKDD
                                                                  2001: Mining Log Data Across All Customer Touch-
11. H. Graubitz, M. Spiliopoulou, and K. Winkler. The             Points, San Francisco, CA, USA, August 2001. To ap-
    DIAsDEM framework for converting domain-specific              pear.
    texts into XML documents with data mining tech-
    niques. In Proceedings of the First IEEE Interna-         21. S. Nestrov, S. Abiteboul, and R. Motwani. Inferring
    tional Conference on Data Mining, San Jose, CA, USA,          structure in semi-structured data. SIGMOD Record,
    November/December 2001. To appear.                            26(4):39–43, 1997.
                                                              22. H. Schmid. Probabilistic part–of–speech tagging using
12. H. Graubitz, K. Winkler, and M. Spiliopoulou. Seman-
                                                                  decision trees. In Proceedings of International Confer-
    tic tagging of domain-specific text documents with DI-
                                                                  ence on New Methods in Language Processing, pages
    AsDEM. In Proceeding of the 1st International Work-
                                                                  44–49, Manchester, UK, September 1994.
    shop on Databases, Documents, and Information Fu-
    sion (DBFusion 2001), pages 61–72, Magdeburg, Ger-        23. A. Sengupta and S. Purao. Transitioning existing con-
    many, May 2001.                                               tent: Inferring organization-spezific document struc-
                                                                  tures. In K. Turowski and K. J. Fellner, editors,
13. R. Kohavi, M. Spiliopoulou, and J. Srivastava, edi-           Tagungsband der 1. Deutschen Tagung XML 2000,
    tors. KDD’2000 Workshop WEBKDD’2000 on Web                    XML Meets Business, pages 130–135, Heidelberg, Ger-
    Mining for E-Commerce — Challenges and Opportu-               many, May 2000.
    nities, Boston, MA, Aug. 2000. ACM.
                                                              24. M. Spiliopoulou. The laborious way from data mining
14. P. A. Laur, F. Masseglia, and P. Poncelet. Schema min-        to web mining. Int. Journal of Comp. Sys., Sci. & Eng.,
    ing: Finding regularity among semistructured data. In         Special Issue on “Semantics of the Web”, 14:113–126,
    D. A. Zighed, J. Komorowski, and J. Żytkow, editors,         Mar. 1999.
    Principles of Data Mining and Knowledge Discovery:
    4th European Conference, PKDD 2000, volume 1910           25. M. Spiliopoulou and L. C. Faulstich. WUM: A Tool
    of Lecture Notes in Artificial Intelligence, pages 498–       for Web Utilization Analysis. In extended version of
    503, Lyon, France, September 2000. Springer, Berlin,          Proc. EDBT Workshop WebDB’98, LNCS 1590, pages
    Heidelberg.                                                   184–203. Springer Verlag, 1999.
                                                              26. A.-H. Tan. Text mining: The state of the art and
15. S. Loh, L. K. Wives, and J. P. M. d. Oliveira. Concept-
                                                                  the challenges. In Proceedings of the PAKDD 1999
    based knowledge discovery in texts extracted from the
                                                                  Workshop on Knowledge Disocovery from Advanced
    Web. ACM SIGKDD Explorations, 2(1):29–39, 2000.
                                                                  Databases, pages 65–70, Beijing, China, April 1999.
16. J. Lumera. Große Mengen an Altdaten stehen XML-           27. K. Wang and H. Liu. Discovering structural association
    Umstieg im Weg. Computerwoche, 27(16):52–53,                  of semistructured data. IEEE Transactions on Knowl-
    2000.                                                         edge and Data Engineering, 12(3):353–371, May/June
                                                                  2000.
17. B. Masand and M. Spiliopoulou, editors. Advances in
    Web Usage Mining and User Profiling: Proceedings of
    the WEBKDD’99 Workshop, LNAI 1836. Springer Ver-
    lag, July 2000.

18. A. Mikheev and S. Finch. A workbench for acquisi-
    tion of ontological knowledge from natural language.
    In Proceedings of the Seventh conference of the Eu-
    ropean Chapter for Computational Linguistics, pages
    194–201, Dublin, Ireland, March 1995.