<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Learning Relations using Collocations</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Analysis of Large Text Corpora</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Gerhard Heyer, Martin Läuter, Uwe Quasthoff, Thomas Wittig, Christian Wolff Leipzig University Computer Science Institute, Natural Language Processing Department Augustusplatz 10 / 11 D-04109 Leipzig</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>German English Dutch French word tokens 300</institution>
          <addr-line>M 250 M 22 M 15 M sentences 13.4 M 13 M 1.5 M 860,000 word types 6 M 1.2 M 600,000 230,000 Table 1: Basic Characteristics of the Corpora</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes the application of statistical analysis of large corpora to the problem of extracting semantic relations from unstructured text. We regard this approach as a viable method for generating input for the construction of ontologies as ontologies use well-defined semantic relations as building blocks (cf. van der Vet &amp; Mars 1998). Starting from a short description of our corpora as well as our language analysis tools, we discuss in depth the automatic generation of collocation sets. We further give examples of different types of relations that may be found in collocation sets for arbitrary terms. The central question we deal with here is how to postprocess statistically generated collocation sets in order to extract named relations. We show that for different types of relations like cohyponyms or instance-of-relations, different extraction methods as well as additional sources of information can be applied to the basic collocation sets in order to verify the existence of a specific type of semantic relation for a given set of terms. Corpus Linguistics is generally understood as a branch of computational linguistics dealing with large text corpora for the purpose of statistical processing of language data (cf. Armstrong 1993, Manning &amp; Schütze 1999). With the availability of large text corpora and the success of robust corpus processing in the nineties, this approach has recently become increasingly popular among computational linguists (cf. Sinclair 1991, Svartvik 1992). Since 1995 a German text corpus of more than 300 million words has been collected (cf. Quasthoff 1998B, Quasthoff &amp; Wolff 2000), containing approx. 6 million different word forms in approx. 13 million sentences, which serves as input for the analysis methods described below. Similarly structured corpora have recently been set up for other European languages as well (English, French, Dutch), with more languages to follow in the near future (see table 1).</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The basic goal of this corpus-based approach is to collect
large amounts of textual data as input for semantic
processing. Starting off from a rather simple data model
tailored for large amounts of data and efficient processing
using a relational data base system at storage level we
employ a simple yet powerful technical infrastructure for
processing texts to be included in the corpus. Beside basic
procedures for text integration into the corpus various
tools have been developed for post-processing linguistic
data. Among them the automatic calculation of
sentencebased word collocations stands out as an especially
valuable tool for corpus-based language technology
applications
        <xref ref-type="bibr" rid="ref11 ref12 ref13 ref6">(see Quasthoff 1998A, Quasthoff &amp; Wolff 2000)</xref>
        .
Additional, application oriented tools exist for search
engine optimization as well as automatic document
classification
        <xref ref-type="bibr" rid="ref13 ref6">(see Heyer, Quasthoff &amp; Wolff 2000)</xref>
        . The
corpora are available on the WWW (http://www. wortschatz.
uni-leipzig.de) and may be used as a large online
dictionary.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Collocations</title>
      <p>
        The occurrence of two or more words within a
welldefined unit of information (sentence, document) is called
a collocation. For the selection of meaningful and
significant collocations, an adequate collocation measure has to
be defined. In the literature, quite a number of different
collocation measures can be found; for an in-depth
discussion of various collocation measures and their
application cf.
        <xref ref-type="bibr" rid="ref15">Smadja 1993</xref>
        ,
        <xref ref-type="bibr" rid="ref9">Lemnitzer 1998</xref>
        , Krenn 2000.
2.1
      </p>
      <sec id="sec-2-1">
        <title>The Collocation Measure</title>
        <p>In the following, our approach towards measuring the
significance of the joint occurrence of two words A and B
in a sentence is discussed. Let
a, b be the number of sentences containing A and B,
k be the number of sentences containing both A
and B,
n be the total number of sentences.</p>
        <p>
          Our significance measure calculates the probability of
joint occurrence of rare events. The results of this
measure are quite similar to the well-known
log-likelihoodmeasure
          <xref ref-type="bibr" rid="ref7">(cf. Krenn 2000)</xref>
          :
Let x = ab/n and define:
        </p>
        <p>- log ⎛⎜1- e- x ∑k-1⋅ 1 xi ⎞⎟
sig(A, B) = ⎝ i=0 i! ⎠ .</p>
        <p>log n
For 2x &lt; k, we get the following approximation which is
much easier to calculate:</p>
        <p>sig(A,B) = (x – k log x + log k!) / log n
In the case of next neighbor collocations we replace the
definition of the above variables by the following. Instead
of a sentence we consider pairs (A, B) of words which are
next neighbors in this sentence. Hence, instead of one
sentence of n words we have n - 1 pairs. For right
neighbor collocations (A, B) let
a, b be the number of pairs of type (A, ?) and (?, B)
resp.,
k be the number of pairs (A, B),
n be the total number of pairs. This equals the total
number of running words minus the number of
sentences.</p>
        <p>Given these variables, the significance measure is
calculated as shown above. In general, this measure yields
semantically acceptable collocation sets for values above
an empirically determined positive threshold (see
examples in section 3 below).
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Properties of the Collocation Measure</title>
        <p>In order to describe basic properties of this measure, we
write sig(n, k, a, b) instead of sig(A, B) where n, k, a, and
b are defined as above.</p>
        <p>Simple co-occurance: A and B occur only once, and they
occur together:</p>
        <p>sig(n,1,1,1) 1 (for n ).</p>
        <p>Independence: A and B occur statistically independently
with probabilities p and q:</p>
        <p>sig(n,npq,np,nq) (for n ).</p>
        <p>Additivity: The unification of the words B and B‘ just adds
the corresponding significances. For k/b we have
sig(n,k,a,b) + sig(n,k‘,a,b‘) sig(n,k+k‘,a,b+b‘)
Enlarging the corpus by a factor m:</p>
        <p>sig(mn, mk, ma, mb) = m sig(n, k, a, b)
1.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Finding Collocations</title>
        <p>
          For calculating the collocation measure for any reasonable
pairs we first count the joint occurrences of each pair.
This problem is complex both in time and storage.
Nevertheless, we managed to calculate the collocation
measure for any pair with total frequency of at least 3 for each
component. Our approach is based on extensible ternary
search trees
          <xref ref-type="bibr" rid="ref2">(cf. Bentley &amp; Sedgewick 1998)</xref>
          where a
count can be associated to a pair of word numbers. The
memory overhead from the original implementation could
be reduced by allocating the space for chunks of 100,000
nodes at once. Even when using this technique on a large
memory computer more than one run through the corpus
may be necessary, taking care that every pair is only
counted once. The resulting word pairs above a threshold
significance are put into a database where they can be
accessed and grouped in many different ways. As
collocations are calculated for different language corpora, our
examples will be taken from the English as well as the
German database.
1.4
        </p>
      </sec>
      <sec id="sec-2-4">
        <title>Visualization of Collocations</title>
        <p>
          Beside textual output of collocation sets, visualizing them
as graphs is an additional type of representation: We
choose a word and arrange its collocates in the plane so
that collocations between collocates are taken into
account. This results in graphs that show homogeneity
where words are interconnected and they show separation
where collocates have little in common. Linguistically
speaking, polysemy is made visible (see fig. 1 below).
Technically speaking, we use simulated annealing to
position the words
          <xref ref-type="bibr" rid="ref4">(see Davidson &amp; Harel 1996)</xref>
          . Line
thickness represents the significance of the collocation. Of
course, all words in the graph are linked to the central
word, the rest of the picture is automatically computed,
but represents semantic connectedness surprisingly well.
Unfortunately the relations between the words are just
presented, but not yet named. Fig. 1 shows the collocation
graph for space. Three different meaning contexts can be
recognized in the graph:
• real estate,
• computer hardware, and
• astronautics.
        </p>
        <p>The connection between address and memory results from
the fact that address is another polysemous concept.
If we fix one word and look at its set of collocates, then
some semantic relations appear more often than others.
The following example shows the most significant
collocations for king ordered by significance:
queen (90), mackerel (83), hill (49), Milken (47), royal (44),
monarch (33), King (30), crowned (30), migratory (30), rook
(29), throne (29), Jordanian (26), junk-bond (26), Hussein (25),
Saudi (25), monarchy (25), crab (23), Jordan (22), Lekhanya
(21), Prince (21), Michael (20), Jordan's (19), palace (19),
undisputed (18), Elvis (17), Shah (17), deposed (17), Panchayat
(16), Zahir (16), fishery (16), former (16), junk (16), constitution
(15), exiled (15), Bhattarai (14), Presley (14), Queen (14),
crown (14), dethroned (14), him (14), Arab (13), Moshoeshoe
(13), himself (13), pawns (13), reigning (13), Fahd (12), Nepali
(12), Rome (12), Saddam (12), once (12), pawn (12), prince
(12), reign (12), [...] government (10) [...]
The following types of relations can be identified:
• Cohyponymy (e. g. Shah, queen, rook, pawn),
• top-level syntactic relations, which translate to
semantic ‘actor-verb’ and often used properties of a
noun (reign; royal, crowned, dethroned),
• instance-of (Fahd, Hussein, Moshoeshoe),
• special relations given by multiwords (A prep/det/
conj B, e. g. king of Jordan), and
• unstructured set of words describing some subject
area, e. g. constitution, government.</p>
        <p>Note that synonymy rarely occurs in the lists. The
relations may be classified according to the properties
symmetry, anti-symmetry, and transitivity.
3.1</p>
      </sec>
      <sec id="sec-2-5">
        <title>Symmetric Relations</title>
        <p>Let us call a relation r symmetric if r(A, B) always implies
r(B, A). Examples of symmetric relations are
• synonymy,
• cohyponomy (or similarity),
• elements of a certain subject area, and
• relations of unknown type.</p>
        <p>Usually, sentence collocations express symmetric
relations.
3.2</p>
      </sec>
      <sec id="sec-2-6">
        <title>Anti-symmetric Relations</title>
        <p>Let us call a relation r anti-symmetric if r(A, B) never
implies r(B, A). Examples of anti-symmetric relations are
• hyponymy and
• relations between properties and its owners like
action and actor or class and instance.</p>
        <p>
          Usually, next neighbor collocations of two words express
anti-symmetric relations. In the case of next neighbor
collocations consisting of more than two words (like A
prep/det/conj B e. g. Samson and Delilah), the relation
might be symmetric, for instance in the case of
conjunctions like and or or
          <xref ref-type="bibr" rid="ref8">(cf. Läuter &amp; Quasthoff 1999)</xref>
          .
3.3
        </p>
      </sec>
      <sec id="sec-2-7">
        <title>Transitivity</title>
        <p>Transitivity of a relation means that r(A, B) and r(B, C)
always implies r(A, C). In general, a relation found
experimentally will not be transitive, of course. But there
may be a part where transitivity holds.</p>
        <p>Some of the most prominent transitive relations are the
cohyponymy, hyponymy, synonymy, and is-a relations.
Note that our graphical representation mainly shows
transitive relations per construction. This kind of relation is
also able to give further results in the combination
procedures described below.
4</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Other Sources for Relations</title>
      <p>While we may intellectually identify types of semantic
relations in collocations sets, additional information and /
or analysis is needed for automatically naming these
relations. In the following, we give different examples for
such complementary information.
4.1</p>
      <sec id="sec-3-1">
        <title>Pattern Based Relations</title>
        <p>Simple pattern-based relations can be extracted from text
if knowledge about information categories like proper
names is used as input. As our corpora include several
large lists of classified terms like names of professions
and last names, extraction rules may be defined:
i. Extraction of first names:</p>
        <p>A pattern like (profession) ? (last name) implies
(with high probability) that the unknown category
? is in fact a first name. Examples are
actress Julia Roberts
hockey hero Wayne Gretzky</p>
        <p>Senator Jesse Helms
ii. Extraction of instance-of-relations given the class
name: The pattern (class name) like ? implies (with
high probability) that the unknown category ? is in
fact a instance name. Examples are:
metals like nickel, arsenic and lead
rivers like the Ganges
newspapers like Pravda
The applicability of patterns like these may heavily
depend on language characteristics like preposition usage.
This type of extraction method is simple and well known;
in our approach it is combined with collocation analysis,
thus yielding better results both in quality and in quantity
(see section 5).
4.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Compounds</title>
        <p>German compounds consist of two (or more) words glued
together by varying mechanisms. The head word (coming
second) is further determined by the first part of the
compound (modifier), which may originally be an adjective,
another noun or a verb stem. In almost all cases a
semantic relation between both parts and the compound can be
found. In section 5.3 we show how the combination of
compound segmentation with collocation analysis can be
used for identifying named relations in compounds.
4.3</p>
      </sec>
      <sec id="sec-3-3">
        <title>Feature Vectors Given by Collocations and</title>
      </sec>
      <sec id="sec-3-4">
        <title>Clustering</title>
        <p>To investigate the meaning of a word A, its contexts in the
texts have to be examined because they reflect the use of
A. If two words A and B have similar contexts, that is,
they are alike in their use, this indicates that there is a
semantic relation between A and B of some kind.
A kind of average context for every word A is formed by
all collocations for A with a significance above a certain
threshold.</p>
        <p>This average context of A is transferred into a feature
vector of A using all words as features as usual. This
results in sparse vectors used for description. The feature
vector of word A is indeed a description of the meaning of
A, because the most important words of the contexts of A
are included.</p>
        <p>Clustering of feature vectors can be used to investigate the
relations between a group of similar words and to figure
out whether or not all the relations are of the same kind.
The following HACM algorithm has an additional natural
reason to stop. It works bottom up like this:
• All words are treated as (basic) items. Each item has
a description (feature vector).
• In each step of the clustering process the two items A
and B with the most similar description vectors are
searched and fitted together to create a new complex
item C combining the words in A and B. The scalar
product is used for determining similarity between
vectors.</p>
        <p>Each step of the clustering algorithm reduces the
number of items by one.
• The feature vector for C is constructed from the
feature vectors of A and B. Therefore we calculate a
combined significance for C with respect to all words
Xi as follows:
sig (C, X i ) = na sig (C, X i ) + nb sig (C, X i )
na + nb na + nb
for all i, 1 ≤ i ≤ n with
n total number of words in the corpus,
na number of words combined in item A, and
nb number of words combined in item B.
• The algorithm stops if only one item is left or if all
remaining feature vectors are orthogonal. This results
usually in a very natural clustering if the threshold for
constructing the feature vectors is suitably chosen.
A cluster of words with probably the same semantic
relation between each of them can be found in the
analysis tree by comparing the similarity between items inside
the items A and B (if these items are complex) with the
calculated similarity between A and B, when fitting them
together to C. If there is a large difference between them,
this is an indication for a different relation between words
combined in item A and words combined in item B. In the
appendix, some examples for this type of semantic
clustering are given.</p>
      </sec>
      <sec id="sec-3-5">
        <title>Symmetric clustering</title>
        <p>If we assume that a cluster represents a semantic relation,
the cluster should represent the possible symmetry and
transitivity of the underlying semantic relation.
Symmetry and transitivity ensure that the terms to be
clustered will themselves be responsible for the
clustering. This in turn implies that the terms found in the cluster
will also be found in the feature vector in prominent
positions.</p>
        <p>In example 1 (Appendix) the clustering result for January
is shown. In the first column we find the terms to be
clustered, on the right hand side there are the components
of the feature vectors ordered by significance.</p>
        <p>The clustered items both appear together and share a
certain aspect. The names of the months or weekdays as
names for periods of time cluster together, just because
they are collocates with one another. The same can be
shown to be true for teammates, metals, colors or fruit.</p>
      </sec>
      <sec id="sec-3-6">
        <title>Anti-symmetric clustering</title>
        <p>For anti-symmetric relations the situation is different.
Again the elements of the original set to be clustered
share a certain aspect, but this aspect is described by a
distinct set of words. Presumably this second set of words
will also cluster. Moreover, it will use the original set as
clustering terms.</p>
        <p>This is shown in example 2 (Appendix). Here we show
that the set given by Präsident, Vorsitzender, Vorsitzende,
Sprecher, Sprecherin properly clusters using words like
sagte, erklärte, teilte (German verbs of utterance).
Conversely, in example 3 (Appendix) we find the set
verwies, mitteilte, meinte, bestätigte, betonte properly
clusters using terms from the above cluster.
4.4</p>
      </sec>
      <sec id="sec-3-7">
        <title>Homogeneous Relations: Iterating the</title>
      </sec>
      <sec id="sec-3-8">
        <title>Collocation Process</title>
        <p>The extraction of collocation sets from plain text can be
viewed as some kind of information condensation. This
process can be iterated if collocation sets themselves are
subjected to the collocation analysis again and again. We
might expect that some of the collocational relations are
strengthened while others will vanish from the iterated
sets of collocations which we will call higher order
collocations. We describe two experiments for the iteration
process: Instead of plain text we start with collocation
sets, using sentence collocations for experiment 1 and
next neighbor collocations for experiment 2. In the case of
a symmetric relation we observe a strengthening while
iterating sentence collocations. In the case of an
antisymmetric relation we observe the same when iterating
next neighbor collocations.</p>
      </sec>
      <sec id="sec-3-9">
        <title>Experiment 1: Iterating Sentence Collocations</title>
        <p>The production of collocations is applied to sets of
sentence collocations instead of sentences. E.g., the
collection of 500,000 sentence collocations has the following
‘sentence‘ (collocation set) for Hemd (shirt): Hemd
Krawatte Hose weißes Anzug weißem Jeans trägt trug
bekleidet weißen Jacke schwarze Jackett schwarzen Weste
kariertes Schlips Mann</p>
      </sec>
      <sec id="sec-3-10">
        <title>Experiment 2: Iterating Next Neighbor Collocations</title>
        <p>In this experiment, the production of collocations is
applied to sets of next neighbor collocations instead of
sentences. The collection of 250,000 next neighbor
collocations has the following two ‘sentences‘ for Hemd (shirt):
weißes weißem weißen blaues kariertes kariertem offenem
aufs karierten gestreiftes letztes [...] (left neighbors)
näher bekleidet ausgezogen spannt trägt aufknöpft
ausgeplündert auszieht wechseln aufgeknöpft ausziehen [...]
(right neighbors)
Example for iterated neighbor collocations of Auto (car):
Original collocations: fahren, Wagen, prallte, Fahrer,
seinem, fuhr, fährt, Polizei, erfaßt, gefahren
Both, experiment 1 and experiment 2 result in collocation
sets carrying a homogeneous semantic relation.
5</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Combining Non-contradictory Partial</title>
    </sec>
    <sec id="sec-5">
      <title>Results</title>
      <p>In section 3 we have given evidence that collocation sets
contain various types of semantic relations without
explicitly naming them while section 4 has introduced a
number of methods for relation extraction. This section
shows different ways of combining results of these
extraction approaches. The results of these combination give
more and / or better results.
5.1</p>
      <sec id="sec-5-1">
        <title>Identical Results</title>
        <p>Two or more of the above algorithms may suggest a
certain relation between two words, for instance,
cohyponymy.</p>
        <p>Example: If both the second order collocations introduced
in section 4.4, and clustering by feature vectors
(section 4.3) independently yield similar sets of words as a
result, this may be taken as an indication of cohyponymy
between the words, e. g. sagte, betonte, kündigte, wies,
nannte, warnte, bekräftigte, meinte […] (German verbs of
utterance).
5.2</p>
      </sec>
      <sec id="sec-5-2">
        <title>Supporting Second Results</title>
        <p>In the second combination type a known relation given by
one method of extraction is verified by an identical but
unnamed second result as follows:
Result 1: There is certain relation r between A and B
Result 2: There is some strong (but unknown) relation
between A and B (e. g. given by a collocation set)
Conclusion: Result 1 holds with more evidence.
One can use this support of orthogonal tests in many
ways: Without knowing anything about deeper language
structure or parsing we can filter out verbs just by testing
if a string accepts at least two of the endings –(e)s, -ing
and –ed/t. The recall is remarkably high. In German we
tested only one mechanism of noun formation from a verb
and got 70% of all verbs with a precision of 83%.
Word formation mechanisms can be explored further. In
German compound nouns are joint together to form one
word. There are several (highly irregular) patterns of
gluing letters between the words. Testing all available
word tokens whether they could be the compound of two
stemmed words from word lists of 93,000 current nouns
reveals just under a million compounds in their stemmed
form. Here stemming accuracy is supported by the
existence of both compounds in the basic list. When
eliminating a hundred words which are prone to generate
wrong separations this algorithm achieves an accuracy of
90%.</p>
        <p>Example:
Result 2: The German compound Entschädigungsgesetz
can be divided into Gesetz and Entschädigung with an
unknown relation.</p>
        <p>Result 1 is given by the four word next neighbor
collocation Gesetz über die Entschädigung. Similarly
Stundenkilometer is analyzed as Kilometer pro Stunde.
In these examples, result 1 is not enough because there are
collocations like Woche auf dem Tisch which do not
describe a meaningful semantic relation.
5.3</p>
      </sec>
      <sec id="sec-5-3">
        <title>Combining Three Results</title>
        <p>Result 1: There is relation r between A and B
Result 2: B is similar to B’ (cohyponymy)
Result 3: There is some strong but unknown relation
between A and B’
Conclusion: There is a relation r between A and B’
Example: As result 1 we might know that Schwanz (tail)
is part of Pferd (horse). Similar terms to Pferd are both
Kuh (cow) and Hund (dog) (result 2). Both of them have
the term Schwanz in their set of significant collocations
(result 3). Hence we might correctly conjecture that both
Kuh and Hund have a tail (Schwanz) as part of their body.
In contrast, Reiter (rider) is a strong collocation to Pferd
and might (incorrectly) be conjectured to be another
similar concept, but Reiter is no collocation with respect
to Schwanz. Hence, the absence of result 3 prevents us
from making an incorrect conclusion.
5.4</p>
      </sec>
      <sec id="sec-5-4">
        <title>Similarity Used to Infer a Strong Property</title>
        <p>Let us call an property p important, if it is preserved under
similarity. This strong feature can be used as follows:
Result 1: A has a certain important property p
Result 2: B is similar to A (i. e., B is a cohyponym of A)
Conclusion: B has the same property p
Example: We consider A and B as similar if they are in the
set of right neighbor collocations of Hafenstadt (port
town) (result 2). If we know that Hafenstadt is a property
of its typical right neighbors (result 1) we may infer this
property for more then 200 cities like Split, Sidon,
Durban, Kismayo, Tyrus, Vlora, Karachi, Durres, […].
5.5</p>
      </sec>
      <sec id="sec-5-5">
        <title>Subject Area Inferred from Collocation Sets</title>
        <p>Result 1: A, B, C, ... are collocates of a certain term.
Result 2: Some of them belong to a certain subject area.
Conclusion: All of them belong to this subject area.
Example: Consider the following top entries in the
collocation set of carcinoma: patients, cell, squamous,
radiotherapy, lung, thyroid, treated, hepatocellular,
metastases, adenocarcinoma, cervix, irradiation, breast,
treatment, CT, therapy, renal, cases, bladder, cervical,
tumor, cancer, metastatic, radiation, uterine, ovarian,
chemotherapy, […]
If we know that some of them belong to the subject area
Medicine, we can add this subject area to the other
members of the collocation set as well.
6</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>
        In this paper, we described different approaches for the
extraction of named semantic relations from large text
corpora. The types of relations are compatible with
relations typically used for constructing ontologies
        <xref ref-type="bibr" rid="ref3">(cf.
Chandrasekaran 1999:22)</xref>
        . The combination of different
types of input information as well as the application of
robust statistical analysis methods guarantees that this
approach may be applied to texts from arbitrary domains
and different languages. Especially, our results may be
used for the automatic generation of semantic relations in
order to fill and expand ontology hierarchies.
7
      </p>
    </sec>
    <sec id="sec-7">
      <title>Appendix: Clustering Examples</title>
      <sec id="sec-7-1">
        <title>Example (1): Clustering Months and Days</title>
        <p>8.2</p>
      </sec>
      <sec id="sec-7-2">
        <title>Example (2): Clustering Leaders</title>
        <p>Präsident _________ sagte, Boris Jelzin, erklärte, stellvertretende, Bill Clinton, stellvertretender, Richter
Vorsitzender _______ | sagte, erklärte, stellvertretende, stellvertretender, Richter, Abteilung, bestätigte
Vorsitzende ___ | | sagte, erklärte, stellvertretende, Richter, bestätigte, Außenministeriums, teilte, gestern
Sprecher _ | | | sagte, erklärte, Außenministeriums, bestätigte, teilte, gestern, mitteilte, Anfrage
Sprecherin _|_|_ | | sagte, erklärte, stellvertretende, Richter, Abteilung, bestätigte, Außenministeriums, sagt
Chef _ | | | Abteilung, Instituts, sagte, sagt, stellvertretender, Professor, Staatskanzlei, Dr.
Leiter _|___|_|_|_
8.3</p>
        <p>Example (3): Clustering Verbs of Utterance
verwies _____________ Sprecher, werde, gestern, Vorsitzende, Polizei, Sprecherin, Anfrage, Präsident, gebe
mitteilte ___________ | Sprecher, werde, gestern, Vorsitzende, Polizei, Sprecherin, Anfrage, Präsident, Montag
meinte _______ | | Sprecher, werde, gestern, Vorsitzende, Sprecherin, Anfrage, Präsident, gebe, Interview
bestätigte_____ | | | Sprecher, werde, gestern, Vorsitzende, Sprecherin, Anfrage, Präsident, gebe, Interview
betonte ___ | | | | Sprecher, werde, gestern, Vorsitzende, Sprecherin, Präsident, gebe, Interview, würden, Bonn
sagte _ | | | | | Sprecher, werde, gestern, Vorsitzende, Sprecherin, Präsident, gebe, Interview, würden
erklärte _|_|_|_|_ | | Sprecher, werde, gestern, Vorsitzende, Sprecherin, Präsident, Anfrage, gebe, Interview
warnte _ | | | Präsident, Vorsitzende, SPD, eindringlich, Ministerpräsident, CDU, Außenminister, Zugleich
sprach _|_______|_|_|_</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Armstrong</surname>
            ,
            <given-names>S</given-names>
          </string-name>
          . (ed.) (
          <year>1993</year>
          ).
          <article-title>Using Large Corpora</article-title>
          .
          <source>Computational Linguistics</source>
          <volume>19</volume>
          (
          <issue>1</issue>
          /2) (
          <year>1993</year>
          )
          <article-title>[Special Issue on Corpus Processing, repr</article-title>
          . MIT Press 1994].
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Bentley</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ; Sedgewick,
          <string-name>
            <surname>R.</surname>
          </string-name>
          (
          <year>1998</year>
          ).
          <article-title>“Ternary Search Trees</article-title>
          .” In: Dr.
          <source>Dobbs Journal</source>
          ,
          <year>April 1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Chandrasekaran</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          et al. (
          <year>1999</year>
          ).
          <article-title>“What are Ontologies, and Why Do We Need Them?”</article-title>
          <source>In: Intelligent Systems</source>
          <volume>14</volume>
          (
          <issue>1</issue>
          ) (
          <year>1999</year>
          ),
          <fpage>20</fpage>
          -
          <lpage>26</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Davidson</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harel</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          (
          <year>1996</year>
          ).
          <article-title>“Drawing Graphs Nicely Using Simulated Annealing</article-title>
          .”
          <source>In: ACM Transactions on Graphics</source>
          <volume>15</volume>
          (
          <issue>4</issue>
          ),
          <fpage>301</fpage>
          -
          <lpage>331</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Francis</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ; Kucera,
          <string-name>
            <surname>H.</surname>
          </string-name>
          (
          <year>1982</year>
          ).
          <article-title>Frequency Analysis of English Language</article-title>
          . Boston: Houghton Mifflin.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Heyer</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ; Quasthoff,
          <string-name>
            <surname>U.</surname>
          </string-name>
          ; Wolff,
          <string-name>
            <surname>Ch.</surname>
          </string-name>
          (
          <year>2000</year>
          ).
          <article-title>“Aiding Web Searches by Statistical Classification Tools</article-title>
          .“ In: Knorz,
          <string-name>
            <surname>G.</surname>
          </string-name>
          ; Kuhlen,
          <string-name>
            <surname>R</surname>
          </string-name>
          . (eds.) (
          <year>2000</year>
          ).
          <article-title>Informationskompetenz - Basiskompetenz in der Informationsgesellschaft</article-title>
          .
          <source>Proc. 7</source>
          .
          <string-name>
            <surname>Intern</surname>
            . Symposium f. Informationswissenschaft,
            <given-names>ISI</given-names>
          </string-name>
          <year>2000</year>
          ,
          <article-title>Darmstadt</article-title>
          . Konstanz: UVK,
          <fpage>163</fpage>
          -
          <lpage>177</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Krenn</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          (
          <year>2000</year>
          ).
          <article-title>“Distributional and Linguistic Implications of Collocation Identification</article-title>
          .”
          <source>In: Proc. Collocations Workshop</source>
          , DGfS Conference, Marburg,
          <year>March 2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Läuter</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Quasthoff</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          (
          <year>1999</year>
          ).
          <article-title>“Kollokationen und semantisches Clustering</article-title>
          .” In: Gippert,
          <string-name>
            <surname>J</surname>
          </string-name>
          . (ed.) (
          <year>1999</year>
          ). Multilinguale Corpora. Codierung, Strukturierung,
          <source>Analyse. Proc. 11</source>
          .
          <string-name>
            <surname>GLDV-Jahrestagung</surname>
          </string-name>
          . Prague: Enigma Corporation,
          <fpage>34</fpage>
          -
          <lpage>41</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Lemnitzer</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          (
          <year>1998</year>
          ).
          <article-title>“Komplexe lexikalische Einheiten in Text und Lexikon</article-title>
          .” In: Heyer,
          <string-name>
            <surname>G.</surname>
          </string-name>
          ; Wolff, Ch. (eds.).
          <source>Linguistik und neue Medien. Wiesbaden: Dt. Universitätsverlag</source>
          ,
          <volume>85</volume>
          -
          <fpage>91</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Manning</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ch</surname>
            . D.; Schütze,
            <given-names>H.</given-names>
          </string-name>
          (
          <year>1999</year>
          ).
          <source>Foundations of Statistical Language Processing</source>
          . Cambridge/MA, London: The MIT Press.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Quasthoff</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          (
          <year>1998A</year>
          ).
          <article-title>“Tools for Automatic Lexicon Maintenance: Acquisition, Error Correction, and the Generation of Missing Values</article-title>
          .“
          <source>In: Proc. First International Conference on Language Resources &amp; Evaluation [LREC]</source>
          , Granada, May
          <year>1998</year>
          , Vol. II,
          <fpage>853</fpage>
          -
          <lpage>856</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Quasthoff</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          (
          <year>1998B</year>
          ). “Projekt der deutsche Wortschatz.” In: Heyer,
          <string-name>
            <given-names>G.</given-names>
            ,
            <surname>Wolff</surname>
          </string-name>
          , Ch. (eds.).
          <source>Linguistik und neue Medien. Wiesbaden: Dt. Universitätsverlag</source>
          ,
          <volume>93</volume>
          -
          <fpage>99</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Quasthoff</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          ; Wolff,
          <string-name>
            <surname>Ch.</surname>
          </string-name>
          (
          <year>2000</year>
          ).
          <article-title>“An Infrastructure for Corpus-Based Monolingual Dictionaries</article-title>
          .”
          <source>In: Proc. LREC-2000. Second International Conference On Language Resources and Evaluation</source>
          . Athens, May/June 2000, Vol. I,
          <fpage>241</fpage>
          -
          <lpage>246</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Sinclair</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>1991</year>
          ).
          <source>Corpus Concordance Collocation</source>
          . Oxford: Oxford University Press.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Smadja F.</surname>
          </string-name>
          (
          <year>1993</year>
          ). “Retrieving Collocations from Text: Xtract.”
          <source>In: Computational Linguistics</source>
          <volume>19</volume>
          (
          <issue>1</issue>
          ) (
          <year>1993</year>
          ),
          <fpage>143</fpage>
          -
          <lpage>177</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>Svartvik</surname>
            ,
            <given-names>J</given-names>
          </string-name>
          . (ed.) (
          <year>1992</year>
          ).
          <source>Directions in Corpus Linguistics: Proc. Nobel Symposium 82</source>
          , Stockholm, 4-
          <fpage>8</fpage>
          August
          <year>1991</year>
          . Berlin: Mouton de Gruyter [=Trends in Linguistics 65].
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>van der Vet</surname>
          </string-name>
          , P. E.;
          <string-name>
            <surname>Mars</surname>
            ,
            <given-names>N. J. I.</given-names>
          </string-name>
          (
          <year>1998</year>
          ). “
          <string-name>
            <surname>Bottom-Up Construction</surname>
          </string-name>
          of Ontologies.”
          <source>In: IEEE Transactions on Knowledge and Data Engineering</source>
          <volume>10</volume>
          (
          <issue>4</issue>
          ) (
          <year>1998</year>
          ),
          <fpage>513</fpage>
          -
          <lpage>526</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>