<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Using a distributional neighbourhood graph to enrich semantic frames in the field of the environment</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Gabriel Bernier-Colborne Marie-Claude L'Homme Observatoire de linguistique Sens-Texte (OLST) Universite ́ de Montre ́al C.P. 6128, succ. Centre-Ville Montre ́al (QC) Canada</institution>
          ,
          <addr-line>H3C 3J7</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2015</year>
      </pub-date>
      <fpage>9</fpage>
      <lpage>16</lpage>
      <abstract>
        <p>This paper presents a semi-automatic method for identifying terms that evoke semantic frames (Fillmore, 1982). The method is tested as a means of identifying lexical units that can be added to existing frames or to new, related frames, using a large corpus on the environment. It is hypothesized that a method based on distributional semantics, which exploits the assumption that words that appear in similar contexts have similar meanings, can help unveil lexical units that evoke the same frame or related frames. The method employs a distributional neighbourhood graph, in which each word is connected to its nearest neighbours according to a distributional semantic model. Results show that most lexical units identified using this method can in fact be assigned to frames related to the field of the environment.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Recent work has shown that Frame
Semantics
        <xref ref-type="bibr" rid="ref1 ref6 ref7">(Fillmore, 1982; Fillmore and Baker, 2010)</xref>
        is an extremely useful framework to account for
the lexical structure of specialized fields of
knowledge
        <xref ref-type="bibr" rid="ref12 ref13 ref20 ref3 ref4 ref4">(Dolbey et al., 2006; Faber et al., 2006;
Schmidt, 2009; L’Homme et al., 2014)</xref>
        . It is
especially attractive in terminology since it provides an
apparatus to connect linguistic properties of terms
to a more abstract conceptual representation level.
      </p>
      <p>1The work reported in this paper is carried out within a
larger project entitled “Understanding the environment
linguistically and textually”, whose objective is to develop
methods for characterizing the contents of texts on two
different levels: 1. textual (using methods and techniques derived
from corpus linguistics and text mining); and 2. linguistic
(based on lexical semantic models).</p>
      <p>
        Frame Semantics has proved especially useful
to represent predicative units (verbs such as
deforest, recycle, warm; predicative nouns such as
impact, pollution, salinization; adjectives such as
clean, green, sustainable), units that are often
ignored in terminological resources. L’Homme
et al. (2014) showed that the framework and
more specifically the methodology devised within
the FrameNet Project
        <xref ref-type="bibr" rid="ref18">(Ruppenhofer et al., 2010)</xref>
        could be used to represent various lexico-semantic
properties of predicative terms (in English and in
French). L’Homme and Robichaud (2014) showed
that frames could be connected via a series of
relations and contribute to help us understand how
terms are used to express environmental
knowledge. However, as will be seen below, the work
that led to the definition of frames and relations
between frames mentioned above was done
manually and turns out to be quite time-consuming.
In this paper, we explore the potential of a
semiautomatic, graph-based method to discover
framerelevant lexical units based on corpus evidence.
      </p>
      <p>This paper is structured as follows. Section 2
explains how semantic frames help reveal part
of the lexical structure of a specialized field of
knowledge. Section 3 describes the graph-based
method used to identify frame-relevant lexical
units. Section 4 discusses how the model used
in the manual evaluation of this method was
selected. Section 5 presents the evaluation
methodology and the results of the evaluation.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Frame Semantics applied to the field of the environment</title>
      <p>
        In a specialized field such as the environment,
many concepts correspond to processes, events
and properties which are typically expressed
linguistically by predicative terms (verbs, predicative
nouns and adjectives). However, traditional
terminological models (and even less traditional ones,
such as ontologies) are not properly equipped
to describe the terms that denote these concepts
and account for their specific linguistic
properties, namely the fact that they require arguments
(X changes Y; impact of X on Y). Frame
Semantics
        <xref ref-type="bibr" rid="ref1 ref6 ref7">(Fillmore, 1982; Fillmore and Baker, 2010)</xref>
        presents itself as a suitable alternative to these
models since it is designed to connect linguistic
properties to an abstract conceptual structure. In
addition, it is well equipped to represent
predicative lexical units and their argument structure.
2.1
      </p>
      <sec id="sec-2-1">
        <title>Discovering frames in the field of the environment</title>
        <p>L’Homme et al. (2014) describe a method to
discover semantic frames based on an existing
terminological resource called DiCoEnviro2, that
contains English and French terms related to the field
of the environment. Each entry in DiCoEnviro is
devoted to a lexical unit (LU), i.e. a lexical item
that conveys a specific meaning, and states the
argument structure of the LU, as in the following
examples:
warm1a, vi: climate[Patient] warms
warm1b, vt: gas[Agent] or change[Cause] warms
climate[Patient]
warm, adj.: warm climate[Patient]</p>
        <p>Argument structures state the number of
obligatory participants, and two different systems are
used to label them: the first one accounts for
the semantic roles of arguments (Agent, Patient,
Cause); the second one gives a typical term, i.e.
a term that is representative of what can appear in
that position.</p>
        <p>Many entries – especially entries that describe
predicative terms – come with annotated contexts
that show how arguments3 are realized in
sentences extracted from an environmental corpus.
For example, annotated contexts for warm1b are
shown in Table 1.</p>
        <p>2See http://olst.ling.umontreal.ca/
cgi-bin/dicoenviro/search_enviro.cgi.</p>
        <p>3Non-obligatory participants are also annotated, as shown
in the last sentence in Table 1, in which the phrase since 1750,
which expresses Time, is annotated.</p>
        <p>The primary radiative effect of CO2 and
water vapour[CAUSE] is to WARM the surface
climate[PATIENT] but cool the stratosphere.
As increases in other greenhouse
gases[CAUSE] WARM the atmosphere and
surface[PATIENT], the amount of water vapour
also increases, amplifying the initial warming
effect of the other greenhouse gases.</p>
        <p>The simulations of this assessment report (for
example, Figure 5) indicate that the estimated
net effect of these perturbations[CAUSE] is to
HAVE WARMED the global climate[PATIENT]
since 1750[TIME].</p>
        <p>
          Argument structures and annotations were used
to discover frames using two different methods.
A semantic frame is a knowledge structure that
represents specific situations (e.g. a teaching
situation, a selling situation, a driving situation). A
frame includes participants (called frame elements
or FEs), some of which are obligatory (core FEs)
and some of which are optional (non-core FEs).
For instance, the Operate vehicle frame describes
a situation in which a Vehicle is set in motion by a
Driver and includes the following core FEs: Area,
Driver, Goal, Path, Source, and Vehicle. Lexical
units such as cycle, cruise, drive, pedal, and ride
evoke this frame
          <xref ref-type="bibr" rid="ref8">(FrameNet, 2015)</xref>
          . In this
previous work, it was assumed that terms that share
similarities with regard to their argument
structures (number and semantic roles of arguments)
and that share similarities with regard to the
nonobligatory participants annotated in contexts are
likely to evoke the same frame.
        </p>
        <p>The first method consisted in comparing the
argument structures and non-obligatory participants
of terms already encoded in the terminological
resource. This method shows that the verbs cool1a,
warm1a and the nouns cooling1 and warming1
share many features. They all have a single
argument (a Patient) and share some non-obligatory
participants (Degree, Duration, Location).</p>
        <p>The second method – which was applied only
to the English terms – consisted in comparing the
contents of the terminological resource to that of
FrameNet.4 Relevant data were extracted from the
FrameNet database for terms that were recorded
in the terminological resource, as shown in
Figure 1. This figure shows an example in which
a correspondence between FrameNet and the
terminological database could be established.
However, in many instances, matches could not be
made as nicely. In various cases, specific frames
needed to be defined for the environmental terms
(for instance, a new frame was created to
capture adjectives such as clean, environmental and
green, whose meaning can be loosely described as
“that does not harm the environment”). In other
cases, existing frames in FrameNet needed to be
adapted to the data extracted from the
terminological database for different reasons (slightly more
specific meanings, different number of arguments,
etc.).
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>A “framed” representation of the terminology of the environment</title>
        <p>It soon became obvious that some of the frames
identified based on the methods described in
Section 2.1 could be linked. For instance, all
processes related to changes affecting the
environment appeared to be somehow related.</p>
        <p>
          Again, using
          <xref ref-type="bibr" rid="ref8">FrameNet (2015)</xref>
          as a reference,
relations were established between some of the
frames defined for environmental terms. Two
relations not found in
          <xref ref-type="bibr" rid="ref8">FrameNet (2015)</xref>
          were
added (Is opposed to and Is a property of). This
work led to the development of a resource called
the Framed DiCoEnviro,5 in which users can
navigate through frames and relations between
frames, and access the terms that evoke these
frames along with their annotations. Figure 2
shows some of the relations identified between the
frame Change of temperature (COT) (that
contains verbs such as cool1a, warm1a and the nouns
cooling1 and warming1) and other frames.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Method for discovering related LUs</title>
      <p>The methods described above allowed us to
define a first subset of frames that are relevant for the
field of the environment, link part of these frames
and assign lexical units (LUs) to them. Based on
this preliminary data, we explored the potential of
a semi-automatic method to enrich our resource by
adding new LUs to existing frames or discovering
new frames. This method exploited distributional
information obtained from a much larger corpus
than the one used in the work described above.</p>
      <p>
        The method we tested to discover related LUs
is based on the neighbourhood graph induced by a
distributional model of semantics. Distributional
semantic models are commonly used to estimate
semantic similarity, the underlying hypothesis
be4The FrameNet team releases an XML version of the
database
        <xref ref-type="bibr" rid="ref1 ref7">(Baker and Hong, 2010)</xref>
        .
      </p>
      <p>5See http://olst.ling.umontreal.ca/
dicoenviro/framed/index.php (in development).
Cause temperature change
Change of temperature
Cause change of position on a scale
Cause balance
Ambient temperature</p>
      <p>Weather event</p>
      <p>Change natural feature</p>
      <p>Change of phase</p>
      <p>Change position on a scale</p>
      <p>Balancing
Cause change natural feature</p>
      <p>Ceasing to be</p>
      <p>Undergo change of state</p>
      <p>Water emanating</p>
      <p>Change of impact
Damaging</p>
      <p>Being at risk</p>
      <p>Progress
Cause change of impact</p>
      <p>
        Cause change into organized society
ing that words that appear in similar contexts tend
to be semantically related
        <xref ref-type="bibr" rid="ref10">(Harris, 1954)</xref>
        . The
usual method of querying a distributional model
is simply to compute, given a particular word, a
sorted list of similar words. This method has
several drawbacks, as has been pointed out recently
by Gyllensten and Sahlgren (2015), who use a
relative neighbourhood graph to query distributional
models in a way that accounts for the fact that the
query can have multiple senses. The method used
here is similar in that it exploits a distributional
neighbourhood graph. This allows us to take a list
of terms and visualize their semantic
neighbourhood, in order to identify related terms that can be
encoded as frame-evoking LUs, either in existing
frames or in new ones.
      </p>
      <p>
        Various kinds of graphs could be used to
compute and visualize the distributional
neighbourhood of a particular word or set of words. We use
a k-nearest-neighbour (k-NN) graph, two
examples of which are the symmetric k-NN graph and
the mutual k-NN graph
        <xref ref-type="bibr" rid="ref15">(Maier et al., 2007)</xref>
        . In
a symmetric k-NN graph, two words wi and wj
are connected if wi is among the k nearest
neighbours (NNs) of wj or if wj is among the k NNs
of wi. In a mutual k-NN graph, the two words are
connected only if both conditions are true: wi is
among the k NNs of wj and wj is among the k
NNs of wi. In this work, we chose to use a
mutual k-NN graph6, the intuition behind this
decision being that if two words are mutual NNs, there
is a better chance that they actually do have
similar meanings. This principle has been exploited
elsewhere
        <xref ref-type="bibr" rid="ref2 ref5">(Ferret, 2012; Claveau et al., 2014)</xref>
        .
      </p>
      <p>The graph construction procedure can be
sum6We also tested the symmetric k-NN graph, but we only
report results obtained with the mutual graph. We achieved
higher F-scores using the mutual graph.
marized as follows. Given a distributional
semantic model, we compute the pairwise similarity
between all words. For each word, we compute its k
NNs by sorting all other words in decreasing order
of similarity to that word and keeping the k most
similar. Then, for each word wi and each
neighbour wj in the k NNs of wi, we add an edge in
the graph between wi and wj if wi is also among
the k NNs of wj . The resulting graph can be used
to visualize the distributional neighbourhood of a
term or set of terms.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Model selection</title>
      <p>Any model that allows us to estimate the
semantic similarity of two words can be used to build a
semantic neighbourhood graph such as the one
described in Section 3. We tested two different
distributional semantic models for this purpose. Both
models have several parameters which must be set
and which can have a significant impact on the
accuracy of the model in a given application. We
therefore used an automatic evaluation procedure
to tune the models’ parameters and select a model
for manual evaluation.
4.1</p>
      <sec id="sec-4-1">
        <title>Corpus and reference data</title>
        <p>
          The corpus used to build the models is the
PANACEA Environment English monolingual
corpus (Catalog Reference ELRA-W0063), a
corpus containing 28071 web pages related to the
environment (approximately 50 million tokens).
The corpus was compiled automatically using
a focused web crawler developed within the
PANACEA project, and is freely distributed by
ELDA for research purposes.7 The corpus was
7See http://catalog.elra.info/product_
info.php?products_id=1184.
converted from XML to raw text and lemmatized
using TreeTagger
          <xref ref-type="bibr" rid="ref19">(Schmid, 1994)</xref>
          .
        </p>
        <p>Reference data were extracted from the Framed
DiCoEnviro.8 The reference data are sets of
LUs that evoke the same semantic frame. The
list of English LUs was extracted from each of
the frames included in the Change of temperature
(COT) scenario9 (cf. Figure 2). Two LUs
(thawing and thinning) were excluded because they
were not in the vocabulary used to construct the
models, which contains the 10,000 most frequent
lemmatized words in the corpus, excluding stop
words. We obtained 13 sets containing a total of
53 LUs, each frame containing between 2 and 7
LUs. The number of unique LUs is 45, several
LUs evoking more than one frame.10
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Models tested</title>
        <p>
          Two different distributional semantic models were
tested. The first is a bag-of-words (BOW)
model
          <xref ref-type="bibr" rid="ref14 ref21">(Schu¨tze, 1992; Lund et al., 1995)</xref>
          , which is
based on a word-word cooccurrence matrix
computed using a sliding context window. The second
is word2vec
          <xref ref-type="bibr" rid="ref16 ref16 ref17 ref17">(Mikolov et al., 2013a; Mikolov et
al., 2013b)</xref>
          , a neural language model that has been
used in many NLP applications in the past few
years. Word2vec (W2V) learns distributed word
representations that can be used in the same way
as BOW vectors to estimate semantic similarity.
        </p>
        <p>The models’ parameters were tuned by testing
various combinations of parameter values,
building neighbourhood graphs from each resulting
model, and computing evaluation metrics on these
graphs based on the reference data described in
Section 4.1.</p>
        <p>Some of the main choices that must be made
when training a model using word2vec pertain to
the architecture of the model (continuous
skipgram or continuous bag-of-words), the training
algorithm (hierarchical softmax or negative
sampling), the use of subsampling of frequent words,
the size (dimensionality) of the word vectors and
8Data extracted on 2015-05-22. Data has been added
since then, as the resource is in development.</p>
        <p>9Frames related to the scenario only through a See also
relation were excluded.</p>
        <p>10Polysemous LUs evoke different frames. For
instance, warm1a (intransitive verb) evokes the
Change of temperature frame; warm1b (transitive verb)
evokes the Cause temperature change frame; and warm2
(adjective) evokes the Ambient temperature frame.
the size of the context window. We tested
various values for each of these parameters, including
the recommended values11 when available. A
total of 160 models were tested. In the case of the
BOW model, important parameters12 include the
type, shape and size of the context window, the
weighting scheme applied to the cooccurrence
frequencies, and the use of dimensionality reduction.
Again, we tested different values for these
parameters. Each model was tested with and without
dimensionality reduction, for which we used
singular value decomposition (SVD). A total of 320
BOW models were built and evaluated (160
unreduced and 160 reduced using SVD).</p>
        <p>For both models, we used the cosine similarity
to estimate the similarity between words.
4.3</p>
      </sec>
      <sec id="sec-4-3">
        <title>Evaluation metrics for model selection</title>
        <p>For each model tested, we constructed multiple
kNN graphs, using different values of k. For each
of these graphs, we computed evaluation metrics
using the reference data described in Section 4.1.
We used precision and recall to check to what
extent LUs belonging to the same frame were
connected in the graph. These metrics are computed
for each of the 45 unique LUs in the reference
data. Let wi be an LU, R(wi) the set of related
LUs that evoke at least one of the frames evoked
by wi, and NN (wi) the set of words that are
adjacent to wi in the graph. Furthermore, let TP i (true
positives) be the number of words in NN (wi) that
are one of the related LUs in R(wi), FP i (false
positives) the number of words in NN (wi) that are
not in R(wi) and FN i (false negatives) the
number of words in R(wi) that are not in NN (wi). The
evaluation metrics are then calculated as usual:
precisioni =
recalli =</p>
        <p>TP i
TP i + FP i</p>
        <p>TP i</p>
        <p>TP i + FN i
F-scorei =
2</p>
        <p>precisioni recalli
precisioni + recalli
11See https://code.google.com/p/
word2vec/#Performance.</p>
        <p>12Several studies have assessed the influence of this
model’s parameters. The relative importance of several
parameters was quantified using analysis of variance by Lapesa
et al. (2014).
0.2120
0.1681
0.1429
0.1253
0.1108</p>
        <p>By analyzing how precision and recall varied
with respect to the BOW model’s parameters, we
determined the optimal parameter values for this
application. For example, the optimal window
size was determined to be 3 words. The
corresponding graph was then evaluated manually.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Evaluation</title>
      <p>Once the model had been selected, the
corresponding neighbourhood graph was evaluated
manually. The evaluation was carried out by one
of the co-authors of this paper, who is responsible
for the development of the Framed DiCoEnviro.
The 45 unique LUs in the reference data had 137
unique neighbours (adjacent nodes in the graph).
These 137 words were evaluated manually in
order to determine to what extent the graph can serve
to discover frame-evoking LUs that can be added
to the database.</p>
      <p>The evaluation was carried out one frame
at a time by observing the subgraph
corresponding to that frame’s LUs and their
neighbours (adjacent nodes in the neighbourhood
graph). For example, the subgraph for the frame
Cause change of impact is shown in Figure 3. In
each subgraph, the LUs already encoded in that
frame were highlighted in green, and those
encoded in other frames in the COT scenario were
highlighted in blue. One or more numbers were
appended to the label of each LU to indicate which
frame(s) it evokes.</p>
      <p>For each word that was not already encoded as
an LU in the COT scenario (i.e. for each white
node), the evaluator was asked to choose one of
the following categories:
1. The word should be encoded as an LU in the
COT scenario
(a) in an existing frame;
(b) in a new frame.
2. The word should be encoded as an LU in
another scenario
(a) in an existing frame that is related to the</p>
      <p>COT scenario (by a See also relation);
(b) in an existing frame that is not related to
the COT scenario;
(c) in a new frame.
3. The word should not be encoded as an LU in
the database, but it is the realization of a core
FE of one of the frames in the COT scenario.
4. The word should not be encoded as an LU in
the database, nor is it the realization of a core
FE of one of the frames in the COT scenario.</p>
      <p>Table 4 shows the results of this evaluation. As
these results show, most lexical items identified by
the method (105 out of 137) can be encoded in a
relevant frame in the field of the environment and
should be added to our resource. Among these,
88 would be frame-evoking LUs (categories 1 and
2) and 17 would be encoded as realizations of
FEs (category 3). Interestingly, 48 lexical items
are related to the COT scenario (categories 1a and
1b). The method allowed us to identify: 1. new
frame-evoking LUs (such as amplification, drop,
and scarcity) that had not been encoded in
existing frames (category 1a); 2. LUs (such as
alteration and eliminate) that evoke frames that had not
been created (category 1b); and 3. variants (such
as cooler for cool and stabilise for stabilize).</p>
      <p>Category</p>
      <p>Nb cases
1a
1b
1 (total)
2a
2b
2c
2 (total)
3
4
Total</p>
      <p>The method also identified 40 items that would
be encoded in environmentally relevant frames,
but in a different scenario (category 2). It is worth
pointing out that among these, 7 items correspond
to LUs that would evoke a frame that is linked to
the COT scenario (category 2a).</p>
      <p>Finally, although 32 lexical items identified by
the method would not be encoded in the resource
and are thus considered false positives from the
point of view of our application, further
explanations are required. Some lexical items could evoke
more general frames. For instance, rapid and slow
would appear in the same frame if the general
lexicon were considered. Other items identified are
acronyms. GW, for instance, is the acronym for
global warming. Technically, it could be defined
as an LU evoking the COT frame, but multi-word
terms and acronyms are not considered in the
resource.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Concluding remarks</title>
      <p>All in all the results obtained are quite interesting
and show that the method can be used to assist
lexicographers when defining frames and their
lexical content, as the distributional neighbourhood of
frame-evoking LUs often contain LUs that evoke
the same frame or related frames. Distributional
neighbourhood graphs provide information about
the content of a specialized corpus that would be
impossible to extract manually from such a large
corpus. They are a very useful complement to
other corpus tools, such as term extractors and
concordancers, as they help lexicographers save
time and locate relevant lexical units (near
synonyms, variants) that they would otherwise miss.</p>
      <p>In future work, we plan to integrate this
methodology to assist lexicographers when
defining new frames related to the field of the
environment. It could be particularly useful to
obtain a view on corpora that deal with new or more
specific topics and unveil the lexical units used
to convey the knowledge related to these topics.
It would also be interesting to test the potential
of the method in other fields of knowledge.
Extensions of this work could also involve using a
graph-based clustering method to discover sets of
lexical units that evoke the same frame without
using existing frames.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>This work was supported by the Social Sciences
and Humanities Research Council (SSHRC) of
Canada.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Collin</given-names>
            <surname>Baker</surname>
          </string-name>
          and
          <string-name>
            <given-names>Jisup</given-names>
            <surname>Hong</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Release 1.5 of the FrameNet data</article-title>
          . International Computer Science Institute. Berkeley.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Vincent</given-names>
            <surname>Claveau</surname>
          </string-name>
          , Ewa Kijak, and
          <string-name>
            <given-names>Olivier</given-names>
            <surname>Ferret</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Explorer le graphe de voisinage pour ame´liorer les the´saurus distributionnels</article-title>
          . In Actes de la 21e confe
          <article-title>´rence sur le traitement automatique des langues naturelles (TALN)</article-title>
          , p.
          <fpage>220</fpage>
          -
          <lpage>231</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Dolbey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Michael</given-names>
            <surname>Ellsworth</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Jan</given-names>
            <surname>Scheffczyk</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>BioFrameNet: A Domain-specific FrameNet Extension with Links to Biomedical Ontologies</article-title>
          .
          <source>Proceedings of KR-MED</source>
          <year>2006</year>
          : Biomedical Ontology in Action, Baltimore, Maryland.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Pamela</given-names>
            <surname>Faber</surname>
          </string-name>
          et al.
          <year>2006</year>
          .
          <article-title>Process-oriented terminology management in the domain of Coastal Engineering</article-title>
          .
          <source>Terminology</source>
          <volume>12</volume>
          (
          <issue>2</issue>
          ):
          <fpage>189</fpage>
          -
          <lpage>213</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Olivier</given-names>
            <surname>Ferret</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Combining bootstrapping and feature selection for improving a distributional thesaurus</article-title>
          .
          <source>In Proceeding of the 20th European Conference on Artificial Intelligence (ECAI)</source>
          , p.
          <fpage>336</fpage>
          -
          <lpage>341</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Charles J.</given-names>
            <surname>Fillmore</surname>
          </string-name>
          .
          <year>1982</year>
          .
          <article-title>Frame Semantics</article-title>
          .
          <source>In Linguistics in the Morning Calm</source>
          , p.
          <fpage>111</fpage>
          -
          <lpage>137</lpage>
          . Seoul: Hanshin Publishing Co.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Charles J.</given-names>
            <surname>Fillmore</surname>
          </string-name>
          and
          <string-name>
            <given-names>Collin</given-names>
            <surname>Baker</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>A frames approach to semantic analysis</article-title>
          .
          <source>In B. Heine and H</source>
          . Narrog (ed.),
          <source>The Oxford Handbook of Linguistic Analysis</source>
          , p.
          <fpage>313</fpage>
          -
          <lpage>339</lpage>
          . Oxford: Oxford University Press.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>FrameNet.</surname>
          </string-name>
          <year>2015</year>
          . https://framenet.icsi. berkeley.edu/fndrupal/. Accessed:
          <fpage>2015</fpage>
          - 09-24.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Amaru</given-names>
            <surname>Cuba</surname>
          </string-name>
          Gyllensten and
          <string-name>
            <given-names>Magnus</given-names>
            <surname>Sahlgren</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Navigating the semantic horizon using relative neighborhood graphs</article-title>
          .
          <source>CoRR, abs/1501</source>
          .02670.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Zellig S.</given-names>
            <surname>Harris</surname>
          </string-name>
          .
          <year>1954</year>
          .
          <article-title>Distributional structure</article-title>
          .
          <source>Word</source>
          ,
          <volume>10</volume>
          (
          <issue>2-3</issue>
          ):
          <fpage>146</fpage>
          -
          <lpage>162</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Gabriella</given-names>
            <surname>Lapesa</surname>
          </string-name>
          ,
          <source>Stefan Evert, and Sabine Schulte im Walde</source>
          .
          <year>2014</year>
          .
          <article-title>Contrasting syntagmatic and paradigmatic relations: Insights from distributional semantic models</article-title>
          .
          <source>In Proceedings of the Third Joint Conference on Lexical and Computational Semantics (*SEM</source>
          <year>2014</year>
          ), p.
          <fpage>160</fpage>
          -
          <lpage>170</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Marie-Claude L'Homme and Benoˆıt Robichaud</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Frames and terminology: Representing predicative units in the field of the environment</article-title>
          .
          <source>In Proceedings of the 4th Workshop on Cognitive Aspects of the Lexicon (CogALex-IV)</source>
          , p.
          <fpage>186</fpage>
          -
          <lpage>197</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Marie-Claude L'Homme</surname>
            , Benoˆıt Robichaud, and
            <given-names>Carlos</given-names>
          </string-name>
          <string-name>
            <surname>Subirats</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Discovering frames in specialized domains</article-title>
          .
          <source>In Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14)</source>
          , p.
          <fpage>1364</fpage>
          -
          <lpage>1371</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Kevin</given-names>
            <surname>Lund</surname>
          </string-name>
          , Curt Burgess, and Ruth Ann Atchley.
          <year>1995</year>
          .
          <article-title>Semantic and associative priming in highdimensional semantic space</article-title>
          .
          <source>In Proceedings of the 17th Annual Conference of the Cognitive Science Society</source>
          , p.
          <fpage>660</fpage>
          -
          <lpage>665</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Markus</given-names>
            <surname>Maier</surname>
          </string-name>
          , Matthias Hein, and Ulrike Von Luxburg.
          <year>2007</year>
          .
          <article-title>Cluster identification in nearest-neighbor graphs</article-title>
          .
          <source>In Algorithmic Learning Theory</source>
          , p.
          <fpage>196</fpage>
          -
          <lpage>210</lpage>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          , Kai Chen, Greg Corrado, and
          <string-name>
            <given-names>Jeffrey</given-names>
            <surname>Dean</surname>
          </string-name>
          .
          <year>2013a</year>
          .
          <article-title>Efficient estimation of word representations in vector space</article-title>
          .
          <source>In Proceedings of ICLR.</source>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          , Ilya Sutskever, Kai Chen, Greg S. Corrado, and
          <string-name>
            <given-names>Jeff</given-names>
            <surname>Dean</surname>
          </string-name>
          .
          <year>2013b</year>
          .
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          .
          <source>In Proceedings of NIPS</source>
          , p.
          <fpage>3111</fpage>
          -
          <lpage>3119</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Josef</given-names>
            <surname>Ruppenhofer</surname>
          </string-name>
          , Michael Ellsworth,
          <string-name>
            <surname>Miriam R. L. Petruck</surname>
          </string-name>
          ,
          <string-name>
            <surname>Christopher R. Johnson</surname>
            , and
            <given-names>Jan</given-names>
          </string-name>
          <string-name>
            <surname>Scheffczyk</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>FrameNet II: Extended theory and practice</article-title>
          . http://framenet2.icsi. berkeley.edu/docs/r1.5/book.pdf. Accessed:
          <fpage>2015</fpage>
          -09-24.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Helmut</given-names>
            <surname>Schmid</surname>
          </string-name>
          .
          <year>1994</year>
          .
          <article-title>Probabilistic part-of-speech tagging using decision trees</article-title>
          .
          <source>In Proceedings of the International Conference on New Methods in Language Processing.</source>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Schmidt</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>The Kicktionary: A Multilingual Lexical Resource of Football Language</article-title>
          . In H.C. Boas (ed.),
          <source>Multilingual FrameNets in Computational Lexicography. Methods and Applications</source>
          , p.
          <fpage>101</fpage>
          -
          <lpage>134</lpage>
          . Berlin/NewYork: Mouton de Gruyter.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Hinrich</given-names>
            <surname>Schu</surname>
          </string-name>
          ¨tze.
          <year>1992</year>
          .
          <article-title>Dimensions of meaning</article-title>
          .
          <source>In Proceedings of the 1992 ACM/IEEE Conference on Supercomputing (Supercomputing'92)</source>
          , p.
          <fpage>787</fpage>
          -
          <lpage>796</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>