<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Legal Jargon in an Environmental TKB: Pollution Phraseology (Short Paper)</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Arianne Reimerink</string-name>
          <email>arianne@ugr.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pamela Faber</string-name>
          <email>pfaber@ugr.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Melania Cabezas-García</string-name>
          <email>melaniacabezas@ugr.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pilar León-Araúz</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Granada, Departamento de Traducción e Interpretación, C/ Buensuceso</institution>
          ,
          <addr-line>11, 18071 Granada</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Despite its importance, Environmental Law has largely been ignored in environmental knowledge bases. EcoLexicon (ecolexicon.ugr.es) has recently begun to include information on the domain. This paper takes the methodological perspective of Frame-based Terminology [1, 2, 3] to analyze typical verb collocations in Environmental Law that will be added to the phraseology module of EcoLexicon. Corpus analysis was used to compare the behaviour of verbs collocating with pollution in Environmental Science and Environmental Law. Verbs were classified based on the lexical domains and semantic classes in Faber and Mairal [4]. The differences were mostly based on the specificity of the other arguments and the emphasis on the polluter in Environmental Law. This resulted in a proposal for the inclusion and configuration of legal information in EcoLexicon.</p>
      </abstract>
      <kwd-group>
        <kwd>1 Environmental Law</kwd>
        <kwd>TKB</kwd>
        <kwd>phraseology</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>POLLUTION are more prominent in this subdomain as compared to the environmental domain as a
whole: time and origin (see examples 3 and 4).</p>
      <p>1. The pollutants disperse in a downward direction causing substantial air pollution at ground
level but cannot escape upwards because of the inversion.
2. …the polluter pays principle, the person responsible for the pollution cannot be identified or
cannot be held liable under Community or national legislation…
3. Indeed, the phenomenon of historical pollution represents the result of the convergence and
interaction of a number of different factors…
4. Historically the regulation of vessel-source pollution has engendered conflict between coastal
States…</p>
      <p>These results entailed changes in the conceptual networks and the definitions of EcoLexicon.
Since these differences at the conceptual level also affect the linguistic level, namely, the choice of
verbs, terms and phraseological structures, this study analyzed verb collocations in Environmental
Law to add to the phraseology module of EcoLexicon, which is currently under construction. In this
pilot study, we focus on phraseology in English. Future research will also address the topic in
Spanish, one of the other major languages of EcoLexicon.</p>
      <p>The rest of this paper is organized as follows: Section 2 explains the phraseology extraction
method and the results; Section 3 provides a proposal for the representation of these results in the
phraseology module; and Sections 4 summarizes the conclusions that can be derived from this
research.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Phraseology extraction</title>
    </sec>
    <sec id="sec-3">
      <title>2.1. Extraction method</title>
      <p>
        When completed, the phraseology module of EcoLexicon will be one of the most important for the
representation of Environmental Law data because of its legal terminology. The phraseology module
is based on a wide interpretation of the concept of collocation and at its core are verb collocations. In
FBT, verb collocations are frequent combinations of two or more lexical units composed of a noun +
verb or a verb + noun where the meaning of the verb is limited by the meaning of the noun. However,
at the same time, the verb restricts the type of noun with which it can combine [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. For example, in
the collocation “the fire burns”, the verb only allows for arguments that can be on fire, whereas the
argument “fire” needs a verb that refers to the process of combustion [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. In this module, verbs will
be classified based on their meaning in combination with the terms with which they collocate. Verbs
will not have their own entry in EcoLexicon but will be included as additional information in the term
entries. The inclusion of a phraseme in EcoLexicon is essentially based on frequency of occurrence in
the corpus. However, as will be shown, frequency changes when comparing different subdomains.
Therefore, different phrasemes and examples will be shown depending on the context the end user is
focussing on in EcoLexicon.
      </p>
      <p>
        To compare the collocational behaviour of pollution in Environmental Science and the subdomain
of Environmental Law, Sketch Engine (https://www.sketchengine.eu/) [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] was used. As a reference
corpus, we used the EcoLexicon Environmental Corpus (EEC, 23 million words) available in the
Open Corpora section of Sketch Engine and compared it to a corpus specifically created for this
purpose: the Environmental Law corpus, composed of EEC texts, tagged with the domain of
Environmental Law, as well as additional texts from the same domain harvested from the Internet
(enLaw, 9.7 million words). The EEC and enLaw were both compiled in Sketch Engine with the Penn
Treebank tagset and the EcoLexicon Semantic Sketch Grammar (ESSG) [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>
        The ESSG is a Corpus Query Language (CQL)-based grammar [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] as is the default grammar used
for word sketches in Sketch Engine. Whereas Sketch Engine’s default grammar provides grammatical
relations, such as verb-object, modifiers, and prepositional phrases, the ESSG was developed for the
extraction of semantic word sketches based on some of the most common semantic relations in
terminology: generic-specific, part-whole, location, cause, and function. The Sketch Engine
functionalities used to compare the two corpora were the following: Word Sketch and Concordance.
For the Word Sketch function the default settings provided by Sketch Engine were used.
      </p>
      <p>
        After extraction, verbs were categorized according to the lexical domains in Faber and Mairal [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
They analyzed and categorized the semantic and syntactic structure of 12,000 general language verbs
through definition factorization, as described in the Lexical Grammar Model (LGM), and validated by
corpus analysis. This resulted in the following general lexical domains: EXISTENCE (be, happen),
CHANGE (become, change), POSSESSION (have), SPEECH (say, talk), EMOTION (feel), ACTION
(do, make), MENTAL PERCEPTION (know, think), MOVEMENT (move, go, come), PHYSICAL
PERCEPTION (see, hear, taste, smell, touch), MANIPULATION (use), CONTACT/IMPACT (hit,
break), and POSITION (put, be). Other smaller classes included LIGHT, SOUND, BODY
FUNCTIONS, WEATHER, etc.
2.2.
      </p>
    </sec>
    <sec id="sec-4">
      <title>Extraction results and discussion</title>
      <p>The data extracted are in Tables 1-4 in the Annex. Table 1 shows that the verbs that collocate with
pollution in both corpora mostly belong to the domain of CAUSATIVE EXISTENCE, more
specifically to cause something to exist (cause), to cause something to cease to exist (eliminate), and
to cause something to not happen (prevent, avoid). Other important lexical domains are CHANGE,
more specifically, to cause something to change by decreasing it (abate, reduce, minimize, mitigate,
decrease, limit) and MANIPULATION (control, monitor). Finally, the lexical domains of VISUAL
PERCEPTION, COGNITION, and SPEECH are present with verbs such as consider, define, regard.</p>
      <p>In both word sketches, air is high up on the list. However this is a tagging mistake as in these
cases air is a noun in an adjectival position and not a verb. The tagging mistake occurs in
constructions where the tagger is incapable of interpreting to as a preposition.</p>
      <p>In the word sketch of verbs with pollution as subject (see Table 2), there are fewer results for the
EEC because the numbers of collocations with pollution did not exceed a certain threshold. This
makes sense because the EEC is a corpus on the environment. Pollution is thus only one of the aspects
to be considered. In contrast, in the enLaw corpus, pollution is a central concept, and that is why
collocations with pollution are statistically more relevant. The lexical domain of the verbs that
predominate in both corpora is EXISTENCE: originate, occur, arise, be, emanate, become, include.
Another lexical domain present in both corpora is CHANGE (reduce, increase), to cause something to
change by making it worse (destroy, damage, harm, threaten) and more general causative verbs such
as cause, affect, derive, result.</p>
      <p>
        The verb flush in the EEC word sketch of pollution is the result of the term pollution flushing,
which is a process through which pollution is removed from a water body through natural or artificial
currents or tides. It can be classified as to cause something to cease to exist (EXISTENCE) or as
MOVEMENT [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>After analysing pollution, we also analysed the verb pollute and the noun polluter in Word Sketch.
When looking at the results for the word sketch object_of, there were no obvious differences between
the verb’s behaviour in enLaw and EEC, apart from the difference in the number of results. Table 3
shows polluter as the object of verbs. Once again, the enLaw corpus provides more results, some of
which are directly related to the legal domain: prosecute, sue. Another important lexical domain is
MANIPULATION: implement, regulate, oblige, force, compel, deter, require, etc. Finally, the word
sketch polluter subject_of showed the verb pay as the very first result for both corpora. This is of
course because one of the most important principles of Environmental Law is the polluter-pays
principle.</p>
      <p>Apart from the fact that there are more results for pollution in enLaw, the lexical domains of the
verbs collocating with pollution were very similar in both corpora. The differences pertained to the
arguments of the verbs.</p>
      <p>Figure 1 (see Annex) shows an extract of the concordances of the CQL abate + pollution in
enLaw. The second argument that collocates with this combination is an institutional body (state,
UK), a company (industries, firms), measure (measures) or cost (expenditures, costs). The second
argument for the CQL minimise + pollution (Figure 2) is mostly measure (requirements, directive,
measures). The second argument for the CQL control + pollution (Figure 3) includes institutional
body (state, administration, agencies) and measure (strategies, measures, regulations, laws).</p>
      <p>In Environmental Law, the verbs abate, minimise and control would be included in the
phraseology module under the term pollution in the following phrasemes: INSTITUTIONAL
BODY/COMPANY/MEASURE/COST + CHANGE [decrease] + POLLUTION; INSTITUTIONAL
BODY/MEASURE + MANIPULATION + POLLUTION.</p>
      <p>One of the participants that is specific to the POLLUTION frame in Environmental Law is
evidently the POLLUTER. Figure 4 shows an extract of the concordances of the CQL pollution
caused_by in enLaw. The cause is evidently the polluting industry (ship, operational discharges,
activities) or the person or entity responsible (polluters, manufacturers, persons, parties,
corporation).</p>
    </sec>
    <sec id="sec-5">
      <title>3. Phraseology representation</title>
      <p>The results showed that the lexical domains of the verbs that collocate with pollution were quite
similar in the EEC and enLaw corpora. The differences are mostly based on the specificity of the
other arguments and the emphasis on the POLLUTER in the Environmental Law subdomain. To
represent this in the phraseology module, under the term pollution, the choice of example sentences
provided for the subdomain would be the following (Figure 5).</p>
    </sec>
    <sec id="sec-6">
      <title>4. Conclusions</title>
      <p>The results described in this paper show that Frame-based Terminology provides the
methodological underpinnings to extract the subtle differences between Environmental Science and
its subdomains at the linguistic level. Specifically, verbal collocations in the Environmental Law
domain differ from those in the Environmental Science domain in regard to the specificity of the
arguments. These differences must be included in terminological knowledge bases in order to provide
an accurate representation of environmental knowledge. Differences at the conceptual level pervade
the linguistic level because of the choice of verbs and their arguments. Representing this
phraseological knowledge for all the terms in EcoLexicon in English and in Spanish will be one of the
challenges for the future development of EcoLexicon.</p>
    </sec>
    <sec id="sec-7">
      <title>5. Acknowledgements</title>
      <p>This research was carried out as part of the projects PID2020-118369GBI00 and
A-HUM-600UGR20, funded by the Spanish Ministry of Science and Innovation and the Regional Government of
Andalusia.</p>
    </sec>
    <sec id="sec-8">
      <title>6. References</title>
      <p>2997
control
cause
prevent
combat
reduce
eliminate
air
address
regulate
avoid
minimise
abate
concern
emit
limit
regard
produce
generate
tackle
mitigate
minimize
include
define
increase
cover</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>P.</given-names>
            <surname>Faber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A Cognitive</given-names>
            <surname>Linguistics</surname>
          </string-name>
          <article-title>View of Terminology and Specialized Language</article-title>
          , De Gruyter Mouton, Berlin/Boston,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>P.</given-names>
            <surname>Faber</surname>
          </string-name>
          ,
          <article-title>Frames as a Framework for Terminology</article-title>
          , in: H.
          <string-name>
            <surname>J. Kockaert</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Steurs</surname>
          </string-name>
          (Eds.),
          <source>Handbook of Terminology</source>
          ,
          <volume>1</volume>
          , John Benjamins Publishing Company,
          <year>2015</year>
          , pp.
          <fpage>14</fpage>
          -
          <lpage>33</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>P.</given-names>
            <surname>Faber</surname>
          </string-name>
          ,
          <article-title>Frame-based Terminology</article-title>
          , in P. Faber, M.C. L´Homme (Eds.), Theoretical Perspectives on Terminology:
          <article-title>Explaining terms, concepts and specialized knowledge</article-title>
          , volume
          <volume>23</volume>
          <source>of Terminology and Lexicography Research and Practice</source>
          , John Benjamins, Amsterdam,
          <year>2022</year>
          , pp.
          <fpage>353</fpage>
          -
          <lpage>376</lpage>
          . doi:https://doi.org/10.1075/tlrp.23.
          <year>16fab</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>P.</given-names>
            <surname>Faber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. Mairal</given-names>
            <surname>Usón</surname>
          </string-name>
          ,
          <article-title>Constructing a lexicon of English verbs</article-title>
          . Berlin: Mouton de Gruyter,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>P.</given-names>
            <surname>Faber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Reimerink</surname>
          </string-name>
          , Framing Terminology in Legal Translation,
          <source>International Journal of Legal Discourse</source>
          ,
          <volume>4</volume>
          .1 (
          <year>2019</year>
          )
          <fpage>15</fpage>
          -
          <lpage>46</lpage>
          . doi:https://doi.org/10.1515/ijld-2019-
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Reimerink</surname>
          </string-name>
          , Pollution in Environmental Law:
          <article-title>Comparative Corpus Analysis</article-title>
          ,
          <source>International Journal of Lexicography</source>
          ,
          <year>ecab027</year>
          ,
          <year>2021</year>
          . doi:https://doi.org/10.1093/ijl/ecab027.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>P</surname>
          </string-name>
          , Faber,
          <string-name>
            <given-names>P.</given-names>
            <surname>León-Araúz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A</given-names>
            . Reimerink, EcoLexicon: New Features and Challenges, in: I.
            <surname>Kernerman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. Kosem</given-names>
            <surname>Trojina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Krek</surname>
          </string-name>
          , L. Trap-Jensen (Eds.), GLOBALEX 2016:
          <article-title>Lexicographic Resources for Human Language Technology in conjunction with the 10th edition of the Language Resources</article-title>
          and Evaluation Conference, Portoroz,
          <year>2016</year>
          , pp.
          <fpage>73</fpage>
          -
          <lpage>80</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>P.</given-names>
            <surname>León</surname>
          </string-name>
          <string-name>
            <surname>Araúz</surname>
          </string-name>
          ,
          <source>Representación Multidimensional del Conocimiento Especializado: El Uso de Marcos desde la Macroestructura hasta la Microestructura</source>
          ,
          <source>Ph.D. thesis</source>
          , Universidad de Granada, Granada, Spain,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Buendía Castro</surname>
          </string-name>
          ,
          <article-title>Phraseology in Specialized Language and its Representation in Environmental Knowledge Resources</article-title>
          ,
          <source>Ph.D. thesis</source>
          , Universidad de Granada, Granada, Spain,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S.</given-names>
            <surname>Montero Martínez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. Buendía</given-names>
            <surname>Castro</surname>
          </string-name>
          , Clasificación semántica de colocaciones verbales para la adquisición y codificación de conocimiento experto: El caso de los riesgos naturales, Revista Española de Lingüística Aplicada,
          <volume>30</volume>
          .1 (
          <year>2017</year>
          )
          <fpage>240</fpage>
          -
          <lpage>272</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>A.</given-names>
            <surname>Kilgarriff</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Baisa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bušta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jakubíček</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Kovář</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Michelfeit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rychlý</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Suchomel</surname>
          </string-name>
          ,
          <source>The Sketch Engine: Ten Years on, Lexicography 1.1</source>
          (
          <issue>2014</issue>
          )
          <fpage>7</fpage>
          -
          <lpage>36</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>P.</given-names>
            <surname>León-Araúz</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . San Martín, P. Faber,
          <article-title>Pattern-based Word Sketches for the Extraction of Semantic Relations</article-title>
          ,
          <source>in: Proceedings of the 5th International Workshop on Computational Terminology (Computerm2016)</source>
          ,
          <source>COLING</source>
          <year>2016</year>
          , Osaka, Japan,
          <year>2016</year>
          , pp.
          <fpage>73</fpage>
          -
          <lpage>82</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>M.</given-names>
            <surname>Jakubíček</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kilgarriff</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Kovářr</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rychlý</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Suchomel</surname>
          </string-name>
          ,
          <article-title>The TenTen Corpus Family</article-title>
          , in: 7th
          <source>International Corpus Linguistics Conference CL</source>
          <year>2013</year>
          , Lancaster,
          <year>2013</year>
          , pp.
          <fpage>125</fpage>
          -
          <lpage>127</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>