<?xml version="1.0" encoding="UTF-8"?>
<TEI xml:space="preserve" xmlns="http://www.tei-c.org/ns/1.0" 
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" 
xsi:schemaLocation="http://www.tei-c.org/ns/1.0 https://raw.githubusercontent.com/kermitt2/grobid/master/grobid-home/schemas/xsd/Grobid.xsd"
 xmlns:xlink="http://www.w3.org/1999/xlink">
	<teiHeader xml:lang="en">
		<fileDesc>
			<titleStmt>
				<title level="a" type="main">Extraction of Semantic XML DTDs from Texts Using Data Mining Techniques</title>
			</titleStmt>
			<publicationStmt>
				<publisher/>
				<availability status="unknown"><licence/></availability>
			</publicationStmt>
			<sourceDesc>
				<biblStruct>
					<analytic>
						<author role="corresp">
							<persName><forename type="first">Karsten</forename><surname>Winkler</surname></persName>
							<email>kwinkler@ebusiness.hhl.de</email>
							<affiliation key="aff0">
								<orgName type="department">Graduate School of Management Department of E-Business</orgName>
								<address>
									<addrLine>Jahnallee 59</addrLine>
									<postCode>D-04109</postCode>
									<settlement>Leipzig, Leipzig</settlement>
									<country key="DE">Germany</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Myra</forename><surname>Spiliopoulou</surname></persName>
							<affiliation key="aff0">
								<orgName type="department">Graduate School of Management Department of E-Business</orgName>
								<address>
									<addrLine>Jahnallee 59</addrLine>
									<postCode>D-04109</postCode>
									<settlement>Leipzig, Leipzig</settlement>
									<country key="DE">Germany</country>
								</address>
							</affiliation>
						</author>
						<title level="a" type="main">Extraction of Semantic XML DTDs from Texts Using Data Mining Techniques</title>
					</analytic>
					<monogr>
						<imprint>
							<date/>
						</imprint>
					</monogr>
					<idno type="MD5">586B6B39E6B8D6473A1C96C84CEA0A23</idno>
				</biblStruct>
			</sourceDesc>
		</fileDesc>
		<encodingDesc>
			<appInfo>
				<application version="0.7.2" ident="GROBID" when="2023-03-19T16:11+0000">
					<desc>GROBID - A machine learning software for extracting information from scholarly documents</desc>
					<ref target="https://github.com/kermitt2/grobid"/>
				</application>
			</appInfo>
		</encodingDesc>
		<profileDesc>
			<textClass>
				<keywords>
					<term>semantic annotation</term>
					<term>XML</term>
					<term>DTD derivation</term>
					<term>knowledge discovery</term>
					<term>data mining</term>
					<term>clustering</term>
				</keywords>
			</textClass>
			<abstract>
<div xmlns="http://www.tei-c.org/ns/1.0"><p>Although composed of unstructured texts, documents contained in textual archives such as public announcements, patient records and annual reports to shareholders often share an inherent though undocumented structure. In order to facilitate efficient, structure-based search in archives and to enable information integration of text collections with related data sources, this inherent structure should be made explicit as detailed as possible. Inferring a semantic and structured XML document type definition (DTD) for an archive and subsequently transforming the corresponding texts into XML documents is a successful method to achieve this objective. The main contribution of this paper is a new method to derive structured XML DTDs in order to extend previously derived flat DTDs. We use the DIAsDEM framework to derive a preliminary, unstructured XML DTD whose components are supported by a large number of documents. However, all XML tags contained in this preliminary DTD cannot a priori be assumed to be mandatory. Additionally, there is no fixed order of XML tags and automatically tagging an archive using a derived DTD always implicates tagging errors. Hence, we introduce the notion of probabilistic XML DTDs whose components are assigned probabilities of being semantically and structurally correct. Our method for establishing a probabilistic XML DTD is based on discovering associations between, resp. frequent sequences of XML tags.</p></div>
			</abstract>
		</profileDesc>
	</teiHeader>
	<text xml:lang="en">
		<body>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>INTRODUCTION</head><p>Most organizations are not only "drowning" in data, they are also "struggling" to cope with huge amounts of text documents. Tan points out that up to 80% of a company's information is stored in unstructured textual documents <ref type="bibr" target="#b25">[26]</ref>. Hence, capturing interesting and actionable knowledge from textual databases is a major challenge for the data mining community. Creating semantic markup is one form of providing explicit knowledge about text archives to facilitate searching and browsing or to enable information integration £ The work of this author is funded by the German Research Society (DFG grant no. SP 572/4-1).</p><p>with related data sources. Unfortunately, most users are not willing to manually create metadata due to the efforts and costs involved <ref type="bibr" target="#b6">[7]</ref>. Thus, text mining techniques are required that (semi-) automatically create semantic markup and tag documents accordingly.</p><p>In this paper, we present the KDD approach pursued in the research project DIAsDEM whose German acronym stands for "Data Integration for Legacy Systems and Semi-Structured Documents by Means of Data Mining Techniques". Our goal is semantic tagging of textual content with meta-data to facilitate searching, querying, identification of and integration with associated texts and relational data. Hence, we aim at deriving a structured XML DTD that serves as a quasi-schema for the document collection and enables the provision of database-like querying services on textual data. DIAsDEM focuses on text collections with domain-specific vocabulary and syntax that frequently share an inherent, but undocumented structure.</p><p>The DIAsDEM framework for semantic tagging of domainspecific texts was introduced in <ref type="bibr" target="#b11">[12,</ref><ref type="bibr" target="#b10">11]</ref>. However, applying the Java-based DIAsDEM Workbench to a text archive currently results in a collection of semantically tagged XML documents that are described by the extracted flat, unstructured XML DTD. However, we ultimately aim at integrating the resulting XML documents with other related data sources. In this context, the derived unstructured, rather preliminary DTD should be transformed into more structured DTD that reflects both ordering and optionality of tags. Given that all XML tags are derived by data mining techniques (i.e. iterative clustering as explained in section 3), they are not crisp due to tagging errors. Taking this critical fact into account, we introduce the notion of a probabilistic DTD that describes the most likely orderings of XML tags and that contains statistical properties for each tag. The structured DTD will be the basis for future information integration efforts that involve XML archives generated by the DIAsDEM Workbench. We introduce two algorithms for inferring a probabilistic DTD that utilize association rule discovery algorithms and sequence mining techniques.</p><p>The rest of this paper is organized as follows: The next ?xml version="1.0" encoding="ISO-8859-1"? !DOCTYPE CommercialRegisterEntry SYSTEM 'CommercialRegisterEntry.dtd'</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>CommercialRegisterEntry</head><p>BusinessPurpose Der Betrieb von Spielhallen in Teltow und das Aufstellen von Geldspiel-und Unterhaltungsautomaten. /BusinessPurpose ShareCapital AmoutOfMoney="25000 EUR" Stammkapital: <ref type="bibr" target="#b24">25</ref> </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>RELATED WORK</head><p>Nahn and Mooney propose the combination of methods from KDD and information extraction to perform text mining tasks <ref type="bibr" target="#b18">[19]</ref>. They apply standard KDD techniques to a collection of structured records that contain previously extracted, application-specific features from texts. Feldman et al. propose text mining at the term level instead of focusing on linguistically tagged words <ref type="bibr" target="#b7">[8]</ref>. The authors represent each document by a set of terms and additionally construct a taxonomy of terms. The resulting dataset is input to KDD algorithms such as association rule discovery. Our DIAsDEM framework adopts the idea of representing texts by terms and concepts. However, our goal is the semantic tagging of structural text units (e.g., sentences or paragraphs) within the document according to a global DTD and not the characterization of the entire document's content. Loh et al. suggest to extract concepts rather than individual words for subsequent use in KDD efforts at the document level. <ref type="bibr" target="#b14">[15]</ref>. Similarly to our framework, the authors suggest to exploit existing vocabularies such as thesauri for concept extraction. Mikheev and Finch describe a workbench to acquire domain knowledge from texts <ref type="bibr" target="#b17">[18]</ref>. Similar to the DIAsDEM Workbench, their approach combines methods from different fields of research in a unifying framework.</p><p>Our approach shares with this research thread the objective of extracting semantic concepts from texts. However, concepts to be extracted in DIAsDEM must be appropriate to serve as elements of the XML DTD. Among other implications, discovering a concept that is peculiar to a single text unit is not sufficient for our purposes, although it may perfectly reflect the corresponding content. In order to derive a DTD, we need to discover groups of text units that share some semantic concepts. Moreover, we concentrate on domain-specific texts, which significantly differ from average texts with respect to word frequency statistics. These collections can hardly be processed using standard text mining software because the integration of relevant domain knowledge is a prerequisite for successful knowledge discovery.</p><p>There are only a few research activities aiming at the transformation of texts into semantically annotated XML documents: Becker et al. introduce the search engine GET-ESS that supports query processing on texts by deriving and processing XML text abstracts <ref type="bibr" target="#b3">[4]</ref>. These abstracts contain language-independent, content-weighted summaries of domain-specific texts. In DIAsDEM, we do not separate meta-data from original texts but rather provide a semantic annotation, keeping the texts intact for later processing or visualization. Given the aforementioned linguistic particularities of the application domains we investigate, a DTD characterizing the content of the documents is more appropriate than inferences on their content. In order to transform existing content into XML documents, Sengupta and Purao propose a method that infers DTDs by using already tagged documents as input <ref type="bibr" target="#b22">[23]</ref>. In contrast, we propose a method that tags plain text documents and derives a DTD for them. Closer to our approach is the work of Lumera, who uses keywords and rules to semi-automatically convert legacy data into XML documents <ref type="bibr" target="#b15">[16]</ref>. However, his approach relies on establishing a rule base that drives the conversion, while we use a KDD methodology that reduces human effort.</p><p>Semi-structured data is another topic of related research within the database community <ref type="bibr" target="#b5">[6,</ref><ref type="bibr" target="#b0">1]</ref>. A lot of effort has recently been put into methods inferring and representing structure in similar semi-structured documents <ref type="bibr" target="#b20">[21,</ref><ref type="bibr" target="#b26">27,</ref><ref type="bibr" target="#b13">14]</ref>. However, these approaches only derive a schema for a given set of semi-structured documents. In DIAsDEM, we have to simultaneously solve the problems of both semi-structuring text documents by semantic tagging and inferring an appropriately structured XML DTD that describes the related archive. We are not aware of any scientific or commercial approaches employing probabilistic document type definitions as introduced in this paper for describing text archives or integrating texts with related data sources.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>THE DIAsDEM FRAMEWORK</head><p>In this paper, the notion of semantic tagging refers to the activity of annotating texts with domain-specific XML tags that might contain additional attributes. Rather than classifying entire documents or tagging single terms, we aim at semantically tagging text units such as sentences or paragraphs. In Germany, companies are obliged by law to submit various information about business affairs to local Commercial Registers. Although Commercial Registers are an important source of information in daily business transactions, their textual content can only be searched using full-text queries at the moment. Hence, semantically semi-structuring these textual archives provides the basis for information integration and creation of value-adding services related to information brokerage. XML query languages could be employed to submit both both content-and structure-based queries against semantically tagged XML archives.</p><p>Our framework pursues two objectives for a given archive of text documents: All text documents should be semantically tagged and an appropriate, preliminary flat XML DTD should be derived for the archive. Semantic tagging in DIAs-DEM is a two-phase process. We have designed a knowledge discovery in textual databases (KDT) process that constitutes the first phase in order to build clusters of semantically similar text units, to tag documents in XML according to the results and to derive an XML DTD describing the archive. The KDT process that was introduced in <ref type="bibr" target="#b11">[12,</ref><ref type="bibr" target="#b10">11]</ref> results in a final set of clusters whose labels serve as XML tags and DTD elements. Huge amounts of new documents can be converted into XML documents in the second, batch-oriented and productive phase of the DIAsDEM framework. All text units contained in new documents are clustered by the previously built text unit clusterer and are subsequently tagged with the corresponding cluster labels.</p><p>In DIAsDEM we concentrate on the semantic tagging of similar text documents originating from a common domain. Nevertheless, the DIAsDEM approach is appropriate for semantically tagging various kinds of archives such as public announcements of courts and administrative authorities, quarterly and annual reports to shareholders, textual patient records in health care applications as well as product and service descriptions published on electronic marketplaces.</p><formula xml:id="formula_0">========== ========== ========== ========= ========== ========== ========= ========== ======= ======= ======== ========== ========== ========== ========= ========== ========== ========= ========== ======= ======= ======== ========== ========== ========== ========= ========== ========== ========= ========== ======= ======= ======== === === === === === === === === === === === === === === === == ==== === ==== ==== ==== ==== ==== ==== ==== ==== ==== ==== ==== ==== ==== ==== ==== ==== \==\==\== \==\==\==== \==\=====\== \===\== \====\== Date = Place = Corporation = Currency = Person = ========== ========== ========== ===== ===== ===== ====, ====, ===, ===, ===, ======, ==== ==== ==== ==== ==== == == ==== ==== ===== ===== === === == == ===== ====== ====== ==== ==== ==== ==== _ _ _ Unacceptable Clusters + + + ==== ==== ==== ==== == == ==== ==== ===== ===== === === == == ===== ====== ====== ==== ==== ==== ====</formula></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Acceptable Clusters</head><p>Persons: In the remainder of this section, we briefly introduce the first phase of the DIAsDEM framework whose iterative and interactive KDT process is depicted in Figure <ref type="figure" target="#fig_0">1</ref>. This process is termed "iterative" because the clustering algorithm is invoked repeatedly. Our notion of iterative clustering should not be confused with the fact that most clustering algorithms perform multiple passes over the data before converging. This process is also "interactive", because a knowledge engineer is consulted for cluster evaluation and final cluster naming decisions at the end of each iteration.</p><formula xml:id="formula_1">Dates: ==== ============ ==== ============= ==== ======== ==== ============= ==== =========== ==== ============ ==== ==.======.=== ==== ==.==.=== ==== ==.=======.=== ==== ==.==.=== ==== ============= ==== ============= Named Entities &lt;========&gt; &lt;=======&gt; &lt;=======&gt; &lt;====&gt; &lt;======&gt; &lt;======&gt; &lt;========&gt; &lt;=====&gt; &lt;=======&gt; Type Definition XML Document ========== ========== ========== ========== &lt;−&gt;====&lt;\&gt; &lt;−&gt;======= =====&lt;\&gt; &lt;−&gt;====== =======&lt;\&gt; &lt;−&gt;====== =======&lt;\&gt; ========== ========== ========== ========== &lt;−&gt;====&lt;\&gt; &lt;−&gt;======= =====&lt;\&gt; &lt;−&gt;====== =======&lt;\&gt; &lt;−&gt;====== =======&lt;\&gt; ========== ========== ========== ========== &lt;−&gt;====&lt;\&gt; &lt;−&gt;======= =====&lt;\&gt; &lt;−&gt;====== =======&lt;\&gt; &lt;−&gt;====== =======&lt;\&gt; XML Documents ==== ==== ==== ==== == == ==== ==== ===== ===== === === == == ===== ====== ====== ==== ==== ==== ==== Text Unit Clusterer</formula><p>Besides the initial text documents to be tagged, the following domain knowledge constitutes input to our KDT process: A thesaurus containing a domain-specific taxonomy of terms and concepts, a preliminary UML schema of the domain and descriptions of specific named entities of importance, e.g. persons and companies. The UML schema reflects the semantics of named entities and the relationships among them, as they are initially conceived by application experts. This schema serves as a reference for the DTD to be derived from discovered semantic tags, but there is no guarantee that the  Similarly to a conventional KDD process, our process starts with a preprocessing phase: After setting the level of granularity by determining the size of text units to be tagged, the Java-and Perl-based DIAsDEM Workbench performs basic NLP preprocessing such as tokenization, normalization and word stemming using TreeTagger <ref type="bibr" target="#b21">[22]</ref>. Instead of removing stop words, we establish a drastically reduced feature space by selecting a limited set of terms and concepts (so-called text unit descriptors) from the thesaurus and the UML schema. Text unit descriptors are currently chosen by the knowledge engineer because they must reflect important concepts of the application domain. All text units are mapped into Boolean vectors of this feature space. Additionally, named entities of interest are extracted from text units by a separate module of the DIAsDEM Workbench. In our case study, we created a small thesaurus and selected 70 relevant descriptors and 109 non-descriptors pointing to descriptors.</p><p>In the pattern discovery phase, all text unit vectors contained in the initial archive are clustered based on similarity of their content. The objective is to discover dense and homogeneous text unit clusters. Clustering is performed in multiple iterations. Each iteration outputs a set of clusters, which the DIAsDEM Workbench partitions into "acceptable" and "unacceptable" ones according to our quality criteria. A cluster of text unit vectors is "acceptable", if and only if (i) its cardinality is large and the corresponding text units are (ii) homogeneous and (iii) can be semantically described by a small number of text unit descriptors. Members of "acceptable" cluster are subsequently removed from the dataset for later labeling, whereas the remaining text unit vectors are input data to the clustering algorithm in the next iteration. In each iteration, the cluster similarity threshold value is stepwise decreased such that "acceptable" clusters become progressively less specific in content. The KTD process is based on a plug-in concept that allows the execution of different clustering algorithms within the DIAsDEM Workbench. In the case study, we employed the demographic clustering function included in the IBM Intelligent Miner for Data that maximizes the value of Condorcet's criterion. After three iterations, the DIAsDEM Workbench discovered altogether 73 "acceptable" clusters containing approx. 85% of text units.</p><p>The postmining phase consists of a labeling step, in which "acceptable" clusters are semi-automatically assigned a label. Ultimately, cluster labels are determined by the knowledge engineer. However, the DIAsDEM Workbench performs both a pre-selection and a ranking of candidate cluster labels for the expert to choose from. All default cluster labels are derived from feature space dimensions (i.e. from text unit descriptors) that are prevailing in each "acceptable" cluster. Cluster labels actually correspond to XML tags that are subsequently used to annotate cluster members. Finally, all original documents are tagged using valid XML tags. Additionally, XML tags are enhanced by attributes reflecting previously extracted named entities and their values. Table <ref type="table" target="#tab_2">2</ref> contains an excerpt of the flat, unstructured XML DTD that was automatically derived from XML tags in the case study. It coarsely describes the semantic structure of the resulting XML collection. Currently, named entities that serve as additional attributes of XML tags are not fully evaluated by the DIAsDEM Workbench.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>ESTABLISHING A PROBABILISTIC DTD</head><p>The output of the DIAsDEM Workbench is a set of semantic XML tags which should be used as XML tags to describe the content of the archive documents. To reflect the content of the archive at an abstract level, it is essential to compose the tags into a DTD. Since the semantic annotations are derived with data mining techniques, they are not crisp. Thus, it is essential that the validity of each tag is expressed in quantitative terms and is estimated properly. Furthermore, an ordering should be imposed upon the tags. Hence, after deriving semantic XML tags, we combine them into a probabilistic DTD by (i) deriving the most likely ordering of the tags and (ii) computing the statistical properties of each tag inside the document type definition.</p><p>The reader may recall that a semantic annotation is actually the label of a cluster discovered by the DIAsDEM Workbench. The underlying clustering mechanism produces nonoverlapping clusters. This implies that a text unit belongs to exactly one cluster, to the effect that it can be annotated with the label of this cluster only. Hence, the tags/labels derived the DIAsDEM Workbench cannot be nested. An extension of the DIAsDEM Workbench by a hierarchical clustering algorithm would allow for the establishment of subclusters and thus for the nesting of (sub)cluster labels. However, this is planned as future work.</p><p>The objectives of the DTD establishment method are the specification of the most appropriate ordering of tags, the identification of correlated or mutually exclusive tags and the adornment of each tag and each correlation among tags with statistical properties. These properties form the basis for reliable query processing, because they determine the expected precision and recall of the query results. In the following, we first introduce the statistical properties we consider for the DTD tags and their associations and describe the methodology for computing these statistics. To model the complete statistical information pertinent in these tags and their relationships, we use a hypergraph structure. We then introduce a mechanism that derives a probabilistic DTD from this graph.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Statistical Properties of Semantic XML Tags</head><p>The statistical properties of DTD tags are depicted in Table <ref type="table" target="#tab_3">3</ref> and described in the following paragraphs. The first column contains the names of the properties. The second column reflects whether the property is peculiar to the whole set of tags as cluster labels (i.e. the whole "model"), to each tag or to a group of associated tags. The last column names the mechanism to be applied to derive the value of each property for each tag.</p><p>Accuracy The DIAsDEM Workbench derives semantic XML tags as labels of clusters. These clusters constitute a model over the data, in the conventional statistical sense. In terms of data classification, such models are subject to misclassification errors. We identify two types of misclassification:</p><p>¯Error type I: A text unit is assigned to the wrong cluster, i.e. the cluster label does not reflect the content of the text unit.</p><p>¯Error type II: A text unit is not assigned to any cluster, although there is a cluster with a label reflecting the content of the text unit.</p><p>For the envisaged DTD, only the error type I is relevant. We use the term accuracy of the model as the probability that cluster labels reflect the content of cluster members. The accuracy value affects the DTD as a whole, it is not peculiar to individual tags. Therefore, we do not incorporate this value in the statistical adornment of the individual tags.</p><p>In order to evaluate the quality of out approach in absence of pre-tagged documents, we drew a random sample containing 5% out of 10,785 text units and asked a domain specialist to verify the annotations of these text units with respect to both error types. Within the sample, error type I (error type II) occured in 0.4% (3.6%) of text units. Hence, tagged text units are most likely to be correctly processed. The percentage of error type II text units is higher, indicating that some text units were not placed in the cluster they semantically belong to. With 0.95 confidence, the overall error rate in the entire dataset is in the interval [2.6%, 5.9%] which is a promising result.</p><p>TagSupport The tags of the DTD are cluster labels derived by a statistical approach. Thus, in terms of XML, they are observed as optional per se. In many application areas, a domain expert can provide suggestions as to which tags should be observed as mandatory. Despite this, there is no guarantee that the expert's suggestions hold true in the archive: The text unit containing this information may have been misclassified by the DIAsDEM Workbench, or the information may be simply absent from the document. For example, although one would expect that each movie has a regisseur, there are movies whose regisseur is unknown or inapplicable, due to the nature of the movie. The property TagSupport offers an indicator of whether a tag may be considered as potentially mandatory. We define it as the ratio of documents where this tag appears to the total number of documents in the archive.</p><p>AssociationConfidence In association rules' discovery, the miner identifies items occuring together. Equivalently, we are interested in tags that affect the appearance of other tags. We use the term AssociationConfidence for a tag Ü given the tags Ý ½ Ý Ò in much the same way as confidence is defined for association rules <ref type="bibr" target="#b4">[5]</ref>: It is the ratio of documents, where the tags Ý ½ Ý Ò and Ü appear to the documents containing Ý ½ Ý Ò .</p><p>AssociationLift Similarly to the association rules' paradigm, the correlation among a tag Ü and a set of tags Ý ½ Ý Ò can be spurious, caused by a very high support of Ü in the whole population. The statistic called lift or improvement is defined to alleviate this problem: it is the ratio of the AssociationConfidence of Ü given Ý ½ Ý Ò to the TagSupport of Ü in the whole population <ref type="bibr" target="#b4">[5]</ref>. In our case, this would be the ratio ××Ó Ø ÓÒ ÓÒ Ò ´Ü Ý½ ÝÒµ Ì ËÙÔÔÓÖØ´Üµ</p><p>.</p><p>LocationConfidence The aforementioned statistical properties on associated tags do not take the ordering of tags into account. In a DTD, the ordering of tags is essential. We use the term LocationConfidence of a tag Ü given the sequence of adjacent tags Ý ½ ¡Ý ¾ ¡ ¡Ý Ò as the number of documents containing the sequence Ý ½ ¡Ý ¾ ¡ ¡Ý Ò ¡Ü to the number of documents containing Ý ½ ¡Ý ¾ ¡ ¡Ý Ò . This definition differs from the conventional statistics known for sequence mining <ref type="bibr" target="#b1">[2]</ref>, because we are concentrating on adjacent tags, disallowing the occurrence of arbitrary tags in-between. Conventional sequence mining do not satisfy this requirement. However, some Web usage miners have been designed to distinguish between adjacent and non-adjacent events <ref type="bibr" target="#b2">[3,</ref><ref type="bibr" target="#b8">9,</ref><ref type="bibr" target="#b23">24,</ref><ref type="bibr" target="#b19">20]</ref>.</p><p>GroupSupport In most of the above statistics, we juxtapose the frequence of appearance of a tag with the frequency of a group of tags, be it a set or a sequence. We use the term GroupSupport as the ratio of the number of documents containing a group of tags to the total number of documents. In fact, for any set of at least two tags, this property assumes one value for the set and as many values as are the perturbations of set members. In the following subsection, we show how we model the statistical information pertinent to individual tags, to tag groups (i.e. sets or sequences) and to relationships among them in a seamless way.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Modeling Statistics of Associated XML Tags</head><p>Some of the values of the statistical properties depicted in Table <ref type="table" target="#tab_3">3</ref> are already made available as part of the DIAsDEM Workbench output, while the remaining ones can be computed by data mining algorithms. To exploit these values for the establishment of a probabilistic DTD, we need a representation model and an algorithm that builds the DTD when processing this model. We introduce here a generic graph structure, in which all statistical information on tags, groups of tags and tag relationships is depicted. This structure is appropriate for the establishment of a DTD or an XMLschema with rich statistical adornments. In the next subsection, we discuss two algorithms for DTD establishment.</p><p>We represent the tags and their associations in a directed graph. Its nodes are individual tags, sequences of adjacent tags or sets of co-occuring tags. Each node is adorned with the statistical properties pertinent to a tag, resp. tag group. An edge represents a relationship of the form Ý ½ Ý Ò Ü;</p><p>however, we use the convention that Ü is the source node and the group of nodes in the rule's LHS is the target. Similarly to nodes, an edge is adorned with the statistics of the orderinsensitive or order-sensitive association it represents.</p><p>Semantic Tags as Graph Nodes Let be the set of semantic XML tags derived by the DIAsDEM Workbench, and let Î ¢ ´¼ ½ be the set of graph nodes conforming to the signature:</p><formula xml:id="formula_2">Ì AE Ñ Ì ËÙÔÔÓÖØ</formula></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>By this definition, a tag can only appear in the graph if its</head><p>TagSupport is more than zero. This is consistent with the fact that XML tags are derived with a KDD method.</p><p>For the groups of tags, we must distinguish among ordersensitive and order-insensitive groups. To do so, we perform three steps. First, we model tag groups as ordered lists. Second, we annotate each list with a flag that indicates whether the group depicted by the list is order-sensitive or order-insensitive; in the latter case, the ordering of the list is irrelevant but must be unique. Third, we guarantee uniqueness, i.e. that all permutations of the same group of tags are mapped into the same order-insensitive list, by requiring that order-insensitive list are lexicographically ordered.</p><p>More formally, let È ´Î µ be the set of all lists of elements in</p><formula xml:id="formula_3">Î , i.e. (TagName,TagSupport)-pairs. An Ü ¾ È ´Î µ ¢ ¼ ½ has the form ´ Ú ½ Ú ½µ,</formula><p>where</p><formula xml:id="formula_4">Ú ½ Ú</formula><p>is a list of elements from Î and the value 1 indicates that this list represents an order-sensitive group. Similarly,</p><formula xml:id="formula_5">Ü ¼ ´ Ú ½ Ú ¼µ would represent the unique order- insensitive group composed of Ú ½ Ú ¾ Î .</formula><p>For example, let ¾ Î be two tags annotated with their TagSupport, whereby precedes lexicographically. The groups ´ ½µ and ´ ½µ are two distinct ordersensitive groups of the two elements. The group ´ ¼µ is the order-insensitive group of the two elements. Finally, the group ´ ¼µ is not permitted, because the group is order-insensitive but the list violates the (default) lexicographical ordering of list elements.</p><formula xml:id="formula_6">Using È ´Î µ ¢ ¼ ½ , we define Î ¼ ´È ´Î µ ¢ ¼ ½ µ ¢ ´¼ ½ with signature: ÖÓÙÔÇ Ì × ÖÓÙÔËÙÔÔÓÖØ</formula><p>where Î ¼ contains only those groups of annotations, for which the GroupSupport value is above a given threshold. This threshold can be specified as input to the mining software, as is usual in KDD applications, or may be set as low as 0. Of course, the threshold value affects the size of the graph and the execution time of the algorithm that traverses it to build the DTD.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Tag Relationships as Graph Edges</head><p>The set of nodes constituting our graph is Î Î Î ¼ , indicating that a node may be a singleton tag or a group of tags with its/their statistics. An edge emanates from an element of Î and points to an element of Î ¼ , i.e. from a tag to an associated group of tags. Formally, the set of edges is a subset of ´Î ¢ Î ¼ µ ¢ ¢ ¢ , where ´¼ ½ AE Í Ä Ä , with signature:</p><formula xml:id="formula_7">××Ó Ø ÓÒ ÓÒ Ò ××Ó Ø ÓÒÄ Ø ÄÓ Ø ÓÒ ÓÒ Ò</formula><p>In this signature, the statistical properties refer to the edge's source given the group of nodes in the edge's target. If the target is a sequence of adjacent tags, then the location confidence is the only valid statistical property, while the association confidence and lift are inapplicable. If the target is a set of tags, then the location confidence is inapplicable. When a statistical property is inapplicable, it assumes the NULL value.</p><p>Graph Properties The components of our "DTDestablishment graph" are tags, groups of tags and relationships among them, all adorned with statistical values. All tags discovered by the DIAsDEM Workbench are present in this graph. Which groups of tags are present depends on the threshold value for the group support. Conceivable are both a minimalistic approach with a high threshold, by which only very frequent groups are present, and a maximalistic approach with a zero-value threshold, by which all tag combinations occuring in the documents are present.</p><p>If we opt for the minimalistic approach, the graph will not be connected in the general case. It will contain only the groups of tags being more frequent than a threshold, and the frequent relationships among them. Certain tags may be isolated, because they only rarely appear in combination with other tags. Contrary to it, the maximalistic approach ensures that all combinations of tags appearing together in documents are depicted in the graph, and that the graph is connected, except of the unlikely case that some documents contain a single tag not occuring in any other documents.</p><p>The size of the graph depends on the threshold value for group support and for the confidence and lift values. The maximalistic approach delivers an upper limit. To compute it, let Ñ be the number of tags/cluster labels output by the DIAsDEM Workbench and let Ò Ñ be the largest number of distinct tags appearing in any document. For each tag, there are ½ È Ò ½ ½ Ò perturbations to be considered, resulting in an equal number of order-sensitive groups and in ¾ Ò´Ò•½µ ¾ order-insensitive ones. There is one edge per (Tag,Group)-pair. Moreover, a tag participates in a maximum of ½ • ¾ groups, thus resulting in</p><formula xml:id="formula_8">Ñ • Ñ ¢ ´ ½ • ¾ µ graph nodes and Ñ ¢ ´ ½ • ¾ µ edges.</formula><p>The upper limit to the graph size indicates that threshold values for the statistical properties are essential for obtaining a manageable graph. On the other hand, each cutoff value implies an information loss. Therefore, we observe the DTDestablishment graph under the maximalistic approach as a reference structure and introduce two algorithms that derive a probabilistic DTD by constructing only a part of this graph.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>DTD Derivation</head><p>The DTD-establishment graph in its maximalistic version captures all relationships among the semantic tags found by the DIAsDEM Workbench. Similarly to the process of schema establishment for a conventional database application, the designer must decide which relationships among the real-world entities are worth capturing and which are not. In our context, "worth capturing" refers to statistical values, presuming that a DTD should reflect the relationships usually present in the documents rather than the rare ones. However, the DTD-establishment graph contains relationships among sets and among sequences of tags, each one adorned with different (and only partially comparable) statistics.</p><p>In the following, we present two algorithms that derive a DTD by constructing part of the DTD-establishment graph. Each tag of this DTD is adorned by only two (derived) probabilistic values, one referring to the tag itself and one to its location inside the DTD. The algorithms are using different heuristics to derive this DTD: the first one concentrates on the pairs of tags appearing most frequently together, while the second one gives preference to maximal sequences of tags. The reader may recall that the computation of the statistics for the relationships among the tags require the activation of data mining software. Hence, each of the algorithms is backed by a miner that returns the desired statistics.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Backward Construction of DTD Sequences</head><p>This algorithm observes a DTD as a set of alternative sequences and builds each sequence backwards, starting at each last tag and proceeding until the first one. Concretely, the algorithm builds "maximal" sequences, where maximality means that the first tag of the sequence is the first tag in most of the documents supporting the sequence.</p><p>Backward expansion of tag-subsequences. The algorithm starts with an arbitrary tag ¾ and then identifies the tag most likely to appear before : ¯If no such tag exists, then the sequence cannot be expanded anymore. It is marked as done and the algorithm shifts to the next sequence that is not done yet, or to the next arbitrary tag from , until all tags are processed.</p><p>¯If there is a most likely predecessor of , say ¼ , it is prepended to the sequence. The next iteration starts, in which the most likely predecessor of ¼ ¡ (in general: of the subsequence built thus far) must be found.</p><p>¯If there are predecessor tags, none of which is more likely than the others, alternative incomplete sequences are produced by duplicating the sequence built thus far. The algorithm processes them iteratively.</p><p>The predecessors of a tag can be found by invoking a sequence miner that returns all frequent sequences of adjacent tags. For the first iteration, the algorithm uses the frequent pairs Ü ½ ¡ Ü Ù ¡ , each one leading to with a (location) confidence ½ Ù respectively. Since these tags are the immediate predecessors of it holds that</p><formula xml:id="formula_9">È Ù ½ ½ ¼ ,</formula><p>where ¼ is the ratio of documents where has no predecessor divided by the total number of documents.</p><p>The maximum among ¼ Ù determines the rest of the procedure: If ¼ is maximum, is the first element of the sequence. In this case, the sequence is marked as "done" and as "maximal" according to the maximality criterion already mentioned.</p><p>If there is a larger than the other elements, then Ü is the predecessor of in the sequence. However, it can be the case that the maximum is only marginally larger than the other values. In other words, there are tags with Ù, such that (i) one of them has shows the maximum location confidence but (ii) the difference of this value from the location confidences of the other ½ tags is less than some small . Then, all tags are acceptable alternatives, resulting to alternative subsequences.</p><p>Identifying maximal tag-sequences. In each iteration, the algorithm considers longer frequent sequences returned by the sequence miner, namely those ending with each subsequence already built. If no frequent sequence is found for a subsequence × ½ , then × is "done" but it must also be checked whether it is maximal. This implies computing the ratio of documents starting with ½ over the whole number of documents and comparing this value Ü to the tag support of ½ , say . The comparison is performed across the same guidelines as for alternative tag predecessors: if Ü , then most documents containing × start with ½ and thus × is maximal. Otherwise, documents starting with ½ mostly adhere to a different maximal sequence. ¯The TagPositionConfidence is the location confidence of this tag with respect to the subsequence of × leading to it.</p><p>Backward versus forward sequence construction. The backward-sequence-construction method generates alternative sequences of DTD tags by pruning the frequent sequences of adjacent tags produced by a sequence miner. An equivalent method can be devised by forward-sequence construction. This would have the advantage of being appropriate for incorporation to a sequence miner's core as well, since most miners of this category perform forward sequence construction: In that case, the mining kernel would be modified to expand a sequence by the most likely successor tag only.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>A DTD as a Tree of Alternatives</head><p>This algorithm observes a DTD as a tree of alternative subsequences and adorns each tag with its support with respect to the subsequence leading to it inside the tree: this is the number of documents starting with this subsequence of tags. Similarly to the sequence-construction algorithm described above, a tag may appear in more than one subsequences, having different predecessors in each one.</p><p>Observing the DTD as a tree implies a common root. In the general case, each document of the archive may start at a different tag. We assume a dummy root, the children of which are those tags that appear first in documents. In general, a tree node refers to a tag , and its children refer to the tags appearing after in the context of 's own predecessors. In a sense, the DTD as a tree of alternatives resembles a DataGuide as proposed in <ref type="bibr" target="#b9">[10]</ref>, although the latter contains no statistical adornments.</p><p>The tree-of-alternatives differs from the sequenceconstruction algorithm in two ways: Firstly, it considers all sequences of tags that appear in documents instead of frequent ones only. Secondly, it only observes complete sequences, while a sequence miner returns arbitrary subsequences of tags.</p><p>The tree-of-alternatives method is realized by the preprocessor module of the Web usage miner WUM <ref type="bibr" target="#b24">[25,</ref><ref type="bibr" target="#b23">24]</ref>. This module is responsible for coercing sequences of events by common prefix and placing them in a tree structure, called "aggregated tree". This tree is input to the navigation pattern discovery process performed by the WUM core. The sequences of tags in documents can be observed as sequences of events, to the effect that the WUM preprocessor can also be used to build a DTD over an archive as a tree of alternative tag sequences. Figure <ref type="figure" target="#fig_3">2</ref> depicts an example of such a tree that related to our case study. Note that the XML document depicted in Table <ref type="table" target="#tab_0">1</ref> is partly described by this DTD excerpt.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>CONCLUSION</head><p>Most of the knowledge hidden in electronic media of an organization is encapsulated in documents. Acquiring this knowledge implies effective querying of the documents as well as the combination of information pieces from different textual assets. This functionality is usually confined to databaselike query processors, while text search engines scan individual assets and return ranked results. In this study, we have presented a methodology that enables query processing and joining of text sources by structuring them. We propose the derivation of an XML DTD over a domain-specific text archive by means of data mining techniques.</p><p>The semantic characterization of text units is the core of our approach as well as the derivation of XML tags from these characterisations. This is undertaken by the DIAsDEM Workbench which is concisely described in the first part of this study. Our main emphasis is on combining these tags that reflect the semantics of many text units across the archive into a single DTD that reflects the semantics of the archive as a whole. We have shown that this DTD is a probabilistic ap- proximation of the archive content and have derived a set of statistical properties that reflect the quality of this approximation, for the whole DTD, for tags inside the DTD and for relationships among these tags.</p><p>The statistical properties of tags and of their relationships form the basis for combining them into a complete DTD in the XML sense or even into an XMLschema. We use a graph structure to depict all statistics that can serve as a basis for this operation and propose two mechanisms that derive DTDs by employing a mining algorithm and a set of heuristic rules. We have tested our methodology on an archive of documents from a regional Commercial Register in Germany:</p><p>We have derived a set of tags with the DIAsDEM Workbench and then implemented one of the proposed mechanisms to derive a DTD for it.</p><p>Our future work includes the implementation of the second mechanism for DTD derivement and the establishment of a framework for the comparison of derived DTDs in terms of expressiveness and accuracy. Of course, the ultimate goal of our work is the establishment of a full-fledged querying mechanism over the text archives. To this purpose, we intend to couple our DTD derivation methods with a query mechanism for semi-structured data. Since the DTDs we derive are of probabilistic nature, this implies also the design of a model that evaluates the quality of the query results.</p></div><figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_0"><head>Figure 1 :</head><label>1</label><figDesc>Figure 1: Iterative and interactive KDT process</figDesc></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_1"><head>?</head><label></label><figDesc>xml version="1.0" encoding="ISO-8859-1"? !ELEMENT CommercialRegisterEntry ( #PCDATA | BusinessPurpose | ShareCapital | ModificationMainOffice | FullyLiablePartner | AppointmentManagingDirector | GeneralPartnership | InitialShareholders | NonCashCapitalContribution | LimitedLiabilityCompany | ConclusionArticles | ModificationRegisteredName | SupervisoryBoard | (...) | Owner | FoundationPartnership )* !ELEMENT BusinessPurpose (#PCDATA) !ELEMENT ShareCapital (#PCDATA) (...) !ELEMENT FoundationPartnership (#PCDATA)</figDesc></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_2"><head></head><label></label><figDesc>Statistics of maximal tag-sequences. At a final step, the algorithm filters out all sequences that are done but are not maximal. It then assigns probability values to each tag inside each maximal sequence × containing it:¯The TagConfidence is the tag's TagSupport multiplied by the accuracy of the model output by the DIAsDEM Workbench.</figDesc></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_3"><head>Figure 2 :</head><label>2</label><figDesc>Figure 2: A DTD as a tree of alternative tag sequences</figDesc></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" type="table" xml:id="tab_0"><head>Table 1 : XML document containing an annotated Commercial Register entry</head><label>1</label><figDesc>.000 EUR. /ShareCapital LimitedLiabilityCompany Gesellschaft mit beschränkter Haftung. /LimitedLiabilityCompany ConclusionArticles Date="12.11.1998; 19.04.1999" Der Gesellschaftsvertrag ist am 12.11.1998 abgeschlossen und am 19.04.1999 abgeändert. /ConclusionArticles (...) Einzelvertretungsbefugnis kann erteilt werden. AppointmentManagingDirector Person="Balski; Pawel; Berlin; 14.04.1965" Pawel Balski, 14.04.1965, Berlin, ist zum Geschäftsführer bestellt. /AppointmentManagingDirector (...) PublicationMedia Nicht eingetragen: Die Bekanntmachungen der Gesellschaft erfolgen im Bundesanzeiger.</figDesc><table><row><cell>/PublicationMedia</cell><cell>/CommercialRegisterEntry</cell></row><row><cell cols="2">section briefly discusses related work. Section 3 gives an</cell></row><row><cell cols="2">overview of our framework for semantic tagging of domain-</cell></row><row><cell cols="2">specific text collections. Section 4 introduces the notion of</cell></row><row><cell cols="2">probabilistic DTDs for textual archives and develops two</cell></row><row><cell cols="2">methods for deriving them. Finally, we conclude and give</cell></row><row><cell>directions for future research in section 5.</cell><cell></cell></row></table></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" type="table" xml:id="tab_1"><head>Table 1</head><label>1</label><figDesc></figDesc><table><row><cell>illustrates this concept of semantic tagging,</cell></row><row><cell>whereas each sentence of this German Commercial Register</cell></row><row><cell>entry is a text unit. In this example, the semantics of most</cell></row><row><cell>sentences are made explicit by XML tags that partly con-</cell></row><row><cell>tain additional attributes describing extracted named entities</cell></row><row><cell>(e.g., names of persons and amounts of money). The XML</cell></row><row><cell>document depicted in Table 1 was created by applying the</cell></row><row><cell>DIAsDEM framework to a collection of 1,145 textual Com-</cell></row><row><cell>mercial Register entries containing 10,785 text units. This</cell></row><row><cell>collection includes all entries related to foundations of com-</cell></row><row><cell>panies in the district of the German city Potsdam in 1999.</cell></row></table></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" type="table" xml:id="tab_2"><head>Table 2 : Preliminary flat, unstructured XML DTD of Commercial Register entries final</head><label>2</label><figDesc>DTD will be contained in or will contain this schema.</figDesc><table /></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" type="table" xml:id="tab_3"><head>Table 3 : Statistics for derived XML tags</head><label>3</label><figDesc></figDesc><table><row><cell>Property</cell><cell>Radius</cell><cell>Computation method</cell></row><row><cell>Accuracy</cell><cell>model</cell><cell>DIAsDEM Workbench</cell></row><row><cell>TagSupport</cell><cell>tag</cell><cell>simple statistics</cell></row><row><cell cols="2">AssociationConfidence set of tags</cell><cell>association rule discovery</cell></row><row><cell>AssociationLift</cell><cell>set of tags</cell><cell>association rule discovery (ARD)</cell></row><row><cell>LocationConfidence</cell><cell>sequence of tags</cell><cell>sequence mining (SeqM)</cell></row><row><cell>GroupSupport</cell><cell cols="2">set or sequence of tags ARD/SeqM</cell></row></table></figure>
		</body>
		<back>

			<div type="acknowledgement">
<div xmlns="http://www.tei-c.org/ns/1.0"><head>ACKNOWLEDGMENTS</head><p>We thank the German Research Society for funding the project DIAsDEM, the Bundesanzeiger Verlagsgesellschaft mbH for providing data and our project collaborators Evguenia Altareva and Stefan Conrad for helpful discussions. The IBM Intelligent Miner for Data is kindly provided by IBM in terms of the IBM DB2 Scholars Program.</p></div>
			</div>

			<div type="references">

				<listBibl>

<biblStruct xml:id="b0">
	<monogr>
		<title level="m" type="main">Data on the Web: From Relations to Semistructured Data and XML</title>
		<author>
			<persName><forename type="first">S</forename><surname>Abiteboul</surname></persName>
		</author>
		<author>
			<persName><forename type="first">P</forename><surname>Buneman</surname></persName>
		</author>
		<author>
			<persName><forename type="first">D</forename><surname>Suciu</surname></persName>
		</author>
		<imprint>
			<date type="published" when="2000">2000</date>
			<publisher>Morgan Kaufman Publishers</publisher>
			<pubPlace>San Francisco</pubPlace>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b1">
	<analytic>
		<title level="a" type="main">Mining sequential patterns</title>
		<author>
			<persName><forename type="first">R</forename><surname>Agrawal</surname></persName>
		</author>
		<author>
			<persName><forename type="first">R</forename><surname>Srikant</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proc. of Int. Conf. on Data Engineering</title>
				<meeting>of Int. Conf. on Data Engineering<address><addrLine>Taipei, Taiwan</addrLine></address></meeting>
		<imprint>
			<date type="published" when="1995-03">Mar. 1995</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b2">
	<monogr>
		<title level="m" type="main">Navigation pattern discovery from internet data</title>
		<author>
			<persName><forename type="first">M</forename><surname>Baumgarten</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><forename type="middle">G</forename><surname>Büchner</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><forename type="middle">S</forename><surname>Anand</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><forename type="middle">D</forename><surname>Mulvenna</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><forename type="middle">G</forename><surname>Hughes</surname></persName>
		</author>
		<imprint>
			<date type="published" when="2000">2000</date>
			<biblScope unit="page" from="70" to="87" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b3">
	<analytic>
		<title level="a" type="main">GETESS: Constructing a linguistic search index for an Internet search engine</title>
		<author>
			<persName><forename type="first">M</forename><surname>Becker</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Bedersdorfer</surname></persName>
		</author>
		<author>
			<persName><forename type="first">I</forename><surname>Bruder</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><surname>Düsterhöft</surname></persName>
		</author>
		<author>
			<persName><forename type="first">G</forename><surname>Neumann</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceedings of the 5th International Conference on Applications of Natural Language to Information Systems</title>
				<meeting>the 5th International Conference on Applications of Natural Language to Information Systems<address><addrLine>Versailles, France</addrLine></address></meeting>
		<imprint>
			<date type="published" when="2000-06">June 2000</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b4">
	<monogr>
		<title level="m" type="main">Data Mining Techniques: For Marketing, Sales and Customer Support</title>
		<author>
			<persName><forename type="first">M</forename><forename type="middle">J</forename><surname>Berry</surname></persName>
		</author>
		<author>
			<persName><forename type="first">G</forename><surname>Linoff</surname></persName>
		</author>
		<imprint>
			<date type="published" when="1997">1997</date>
			<publisher>John Wiley &amp; Sons, Inc</publisher>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b5">
	<analytic>
		<title level="a" type="main">Semistructured data</title>
		<author>
			<persName><forename type="first">P</forename><surname>Buneman</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceedings of the Sixteenth ACM SIGACT-SIGMOD-SIGART Symposium on Principles of Database Systems</title>
				<meeting>the Sixteenth ACM SIGACT-SIGMOD-SIGART Symposium on Principles of Database Systems<address><addrLine>Tucson, AZ, USA</addrLine></address></meeting>
		<imprint>
			<date type="published" when="1997-05">May 1997</date>
			<biblScope unit="page" from="117" to="121" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b6">
	<analytic>
		<title level="a" type="main">From manual to semi-automatic semantic annotation: About ontology-based text annotation tools</title>
		<author>
			<persName><forename type="first">M</forename><surname>Erdmann</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><surname>Maedche</surname></persName>
		</author>
		<author>
			<persName><forename type="first">H.-P</forename><surname>Schnurr</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Staab</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">ETAI Journal -Section on Semantic Web</title>
		<imprint>
			<biblScope unit="volume">6</biblScope>
			<date type="published" when="2001">2001</date>
		</imprint>
	</monogr>
	<note>To appear</note>
</biblStruct>

<biblStruct xml:id="b7">
	<analytic>
		<title level="a" type="main">Text mining at the term level</title>
		<author>
			<persName><forename type="first">R</forename><surname>Feldman</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Fresko</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Y</forename><surname>Kinar</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Y</forename><surname>Lindell</surname></persName>
		</author>
		<author>
			<persName><forename type="first">O</forename><surname>Liphstat</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Rajman</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Y</forename><surname>Schler</surname></persName>
		</author>
		<author>
			<persName><forename type="first">O</forename><surname>Zamir</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceedings of the Second European Symposium on Principles of Data Mining and Knowledge Discovery</title>
				<meeting>the Second European Symposium on Principles of Data Mining and Knowledge Discovery<address><addrLine>Nantes, France</addrLine></address></meeting>
		<imprint>
			<date type="published" when="1998-09">September 1998</date>
			<biblScope unit="page" from="65" to="73" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b8">
	<monogr>
		<title level="m" type="main">Mining web navigation path fragments</title>
		<author>
			<persName><forename type="first">W</forename><surname>Gaul</surname></persName>
		</author>
		<author>
			<persName><forename type="first">L</forename><surname>Schmidt-Thieme</surname></persName>
		</author>
		<imprint>
			<date type="published" when="2000">2000</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b9">
	<analytic>
		<title level="a" type="main">DataGuides: Enabling query formulation and optimization in semistructured databases</title>
		<author>
			<persName><forename type="first">R</forename><surname>Goldman</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Widom</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">VLDB&apos;97</title>
				<meeting><address><addrLine>Athens, Greece</addrLine></address></meeting>
		<imprint>
			<date type="published" when="1997-08">Aug. 1997</date>
			<biblScope unit="page" from="436" to="445" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b10">
	<analytic>
		<title level="a" type="main">The DIAsDEM framework for converting domain-specific texts into XML documents with data mining techniques</title>
		<author>
			<persName><forename type="first">H</forename><surname>Graubitz</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Spiliopoulou</surname></persName>
		</author>
		<author>
			<persName><forename type="first">K</forename><surname>Winkler</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceedings of the First IEEE International Conference on Data Mining</title>
				<meeting>the First IEEE International Conference on Data Mining<address><addrLine>San Jose, CA, USA</addrLine></address></meeting>
		<imprint>
			<date type="published" when="2001-12">November/December 2001</date>
		</imprint>
	</monogr>
	<note>To appear</note>
</biblStruct>

<biblStruct xml:id="b11">
	<analytic>
		<title level="a" type="main">Semantic tagging of domain-specific text documents with DI-AsDEM</title>
		<author>
			<persName><forename type="first">H</forename><surname>Graubitz</surname></persName>
		</author>
		<author>
			<persName><forename type="first">K</forename><surname>Winkler</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Spiliopoulou</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceeding of the 1st International Workshop on Databases, Documents, and Information Fusion (DBFusion 2001)</title>
				<meeting>eeding of the 1st International Workshop on Databases, Documents, and Information Fusion (DBFusion 2001)<address><addrLine>Magdeburg, Germany</addrLine></address></meeting>
		<imprint>
			<date type="published" when="2001-05">May 2001</date>
			<biblScope unit="page" from="61" to="72" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b12">
	<analytic>
		<title level="a" type="main">editors</title>
		<author>
			<persName><forename type="first">R</forename><surname>Kohavi</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Spiliopoulou</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Srivastava</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">KDD&apos;2000 Workshop WEBKDD&apos;2000 on Web Mining for E-Commerce -Challenges and Opportunities</title>
				<meeting><address><addrLine>Boston, MA</addrLine></address></meeting>
		<imprint>
			<publisher>ACM</publisher>
			<date type="published" when="2000-08">Aug. 2000</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b13">
	<analytic>
		<title level="a" type="main">Schema mining: Finding regularity among semistructured data</title>
		<author>
			<persName><forename type="first">P</forename><forename type="middle">A</forename><surname>Laur</surname></persName>
		</author>
		<author>
			<persName><forename type="first">F</forename><surname>Masseglia</surname></persName>
		</author>
		<author>
			<persName><forename type="first">P</forename><surname>Poncelet</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Principles of Data Mining and Knowledge Discovery: 4th European Conference, PKDD 2000</title>
		<title level="s">Lecture Notes in Artificial Intelligence</title>
		<editor>
			<persName><forename type="first">D</forename><forename type="middle">A</forename><surname>Zighed</surname></persName>
		</editor>
		<editor>
			<persName><forename type="first">J</forename><surname>Komorowski</surname></persName>
		</editor>
		<editor>
			<persName><forename type="first">J</forename><surname>Żytkow</surname></persName>
		</editor>
		<meeting><address><addrLine>Lyon, France; Berlin, Heidelberg</addrLine></address></meeting>
		<imprint>
			<publisher>Springer</publisher>
			<date type="published" when="1910-09">1910. September 2000</date>
			<biblScope unit="page" from="498" to="503" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b14">
	<analytic>
		<title level="a" type="main">Conceptbased knowledge discovery in texts extracted from the Web</title>
		<author>
			<persName><forename type="first">S</forename><surname>Loh</surname></persName>
		</author>
		<author>
			<persName><forename type="first">L</forename><forename type="middle">K</forename><surname>Wives</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><forename type="middle">P M D</forename><surname>Oliveira</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">ACM SIGKDD Explorations</title>
		<imprint>
			<biblScope unit="volume">2</biblScope>
			<biblScope unit="issue">1</biblScope>
			<biblScope unit="page" from="29" to="39" />
			<date type="published" when="2000">2000</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b15">
	<analytic>
		<title level="a" type="main">Große Mengen an Altdaten stehen XML-Umstieg im Weg</title>
		<author>
			<persName><forename type="first">J</forename><surname>Lumera</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Computerwoche</title>
		<imprint>
			<biblScope unit="volume">27</biblScope>
			<biblScope unit="issue">16</biblScope>
			<biblScope unit="page" from="52" to="53" />
			<date type="published" when="2000">2000</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b16">
	<analytic>
		<author>
			<persName><forename type="first">B</forename><surname>Masand</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Spiliopoulou</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Advances in Web Usage Mining and User Profiling: Proceedings of the WEBKDD&apos;99 Workshop</title>
		<title level="s">LNAI</title>
		<imprint>
			<publisher>Springer Verlag</publisher>
			<date type="published" when="2000-07">July 2000</date>
			<biblScope unit="volume">1836</biblScope>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b17">
	<analytic>
		<title level="a" type="main">A workbench for acquisition of ontological knowledge from natural language</title>
		<author>
			<persName><forename type="first">A</forename><surname>Mikheev</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Finch</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceedings of the Seventh conference of the European Chapter for Computational Linguistics</title>
				<meeting>the Seventh conference of the European Chapter for Computational Linguistics<address><addrLine>Dublin, Ireland</addrLine></address></meeting>
		<imprint>
			<date type="published" when="1995-03">March 1995</date>
			<biblScope unit="page" from="194" to="201" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b18">
	<analytic>
		<title level="a" type="main">Using information extraction to aid the discovery of prediction rules from text</title>
		<author>
			<persName><forename type="first">U</forename><forename type="middle">Y</forename><surname>Nahm</surname></persName>
		</author>
		<author>
			<persName><forename type="first">R</forename><forename type="middle">J</forename><surname>Mooney</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceedings of the Sixth International Conference on Knowledge Discovery and Data Mining (KDD-2000) Workshop on Text Mining</title>
				<meeting>the Sixth International Conference on Knowledge Discovery and Data Mining (KDD-2000) Workshop on Text Mining<address><addrLine>Boston, MA, USA</addrLine></address></meeting>
		<imprint>
			<date type="published" when="2000-08">August 2000</date>
			<biblScope unit="page" from="51" to="58" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b19">
	<analytic>
		<title level="a" type="main">Effective prediction of web-user accesses: A data mining approach</title>
		<author>
			<persName><forename type="first">A</forename><surname>Nanopoulos</surname></persName>
		</author>
		<author>
			<persName><forename type="first">D</forename><surname>Katsaros</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Y</forename><surname>Manolopoulos</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceeding of the Workshop WEBKDD 2001: Mining Log Data Across All Customer Touch-Points</title>
				<meeting>eeding of the Workshop WEBKDD 2001: Mining Log Data Across All Customer Touch-Points<address><addrLine>San Francisco, CA, USA</addrLine></address></meeting>
		<imprint>
			<date type="published" when="2001-08">August 2001</date>
		</imprint>
	</monogr>
	<note>To appear</note>
</biblStruct>

<biblStruct xml:id="b20">
	<analytic>
		<title level="a" type="main">Inferring structure in semi-structured data</title>
		<author>
			<persName><forename type="first">S</forename><surname>Nestrov</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Abiteboul</surname></persName>
		</author>
		<author>
			<persName><forename type="first">R</forename><surname>Motwani</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">SIGMOD Record</title>
		<imprint>
			<biblScope unit="volume">26</biblScope>
			<biblScope unit="issue">4</biblScope>
			<biblScope unit="page" from="39" to="43" />
			<date type="published" when="1997">1997</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b21">
	<analytic>
		<title level="a" type="main">Probabilistic part-of-speech tagging using decision trees</title>
		<author>
			<persName><forename type="first">H</forename><surname>Schmid</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceedings of International Conference on New Methods in Language Processing</title>
				<meeting>International Conference on New Methods in Language Processing<address><addrLine>Manchester, UK</addrLine></address></meeting>
		<imprint>
			<date type="published" when="1994-09">September 1994</date>
			<biblScope unit="page" from="44" to="49" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b22">
	<analytic>
		<title level="a" type="main">Transitioning existing content: Inferring organization-spezific document structures</title>
		<author>
			<persName><forename type="first">A</forename><surname>Sengupta</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Purao</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Tagungsband der 1. Deutschen Tagung XML 2000, XML Meets Business</title>
				<editor>
			<persName><forename type="first">K</forename><surname>Turowski</surname></persName>
		</editor>
		<editor>
			<persName><forename type="first">K</forename><forename type="middle">J</forename><surname>Fellner</surname></persName>
		</editor>
		<meeting><address><addrLine>Heidelberg, Germany</addrLine></address></meeting>
		<imprint>
			<date type="published" when="2000-05">May 2000</date>
			<biblScope unit="page" from="130" to="135" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b23">
	<analytic>
		<title level="a" type="main">The laborious way from data mining to web mining</title>
		<author>
			<persName><forename type="first">M</forename><surname>Spiliopoulou</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Int. Journal of Comp. Sys., Sci. &amp; Eng., Special Issue on &quot;Semantics of the Web</title>
		<imprint>
			<biblScope unit="volume">14</biblScope>
			<biblScope unit="page" from="113" to="126" />
			<date type="published" when="1999-03">Mar. 1999</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b24">
	<analytic>
		<title level="a" type="main">WUM: A Tool for Web Utilization Analysis</title>
		<author>
			<persName><forename type="first">M</forename><surname>Spiliopoulou</surname></persName>
		</author>
		<author>
			<persName><forename type="first">L</forename><forename type="middle">C</forename><surname>Faulstich</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">extended version of Proc. EDBT Workshop WebDB&apos;98</title>
		<title level="s">LNCS</title>
		<imprint>
			<publisher>Springer Verlag</publisher>
			<date type="published" when="1999">1999</date>
			<biblScope unit="volume">1590</biblScope>
			<biblScope unit="page" from="184" to="203" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b25">
	<analytic>
		<title level="a" type="main">Text mining: The state of the art and the challenges</title>
		<author>
			<persName><forename type="first">A.-H</forename><surname>Tan</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceedings of the PAKDD 1999 Workshop on Knowledge Disocovery from Advanced Databases</title>
				<meeting>the PAKDD 1999 Workshop on Knowledge Disocovery from Advanced Databases<address><addrLine>Beijing, China</addrLine></address></meeting>
		<imprint>
			<date type="published" when="1999-04">April 1999</date>
			<biblScope unit="page" from="65" to="70" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b26">
	<analytic>
		<title level="a" type="main">Discovering structural association of semistructured data</title>
		<author>
			<persName><forename type="first">K</forename><surname>Wang</surname></persName>
		</author>
		<author>
			<persName><forename type="first">H</forename><surname>Liu</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">IEEE Transactions on Knowledge and Data Engineering</title>
		<imprint>
			<biblScope unit="volume">12</biblScope>
			<biblScope unit="issue">3</biblScope>
			<biblScope unit="page" from="353" to="371" />
			<date type="published" when="2000-06">May/June 2000</date>
		</imprint>
	</monogr>
</biblStruct>

				</listBibl>
			</div>
		</back>
	</text>
</TEI>
