<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Standardizing Language with Word Embeddings and Language Modeling in Reports of Near Misses in Seveso Industries</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Simone Bruno?</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Silvia Maria Ansaldi</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Patrizia Agnello</string-name>
          <email>p.agnellog@inail.it</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fabio Massimo Zanzotto?</string-name>
          <email>fabio.massimo.zanzotto@uniroma2.it</email>
        </contrib>
      </contrib-group>
      <abstract>
        <p>Standardizing technical language has always been a strong necessity of the technological society. Today, Natural Language Processing as well as the widespread use of computerized document writing can give a tremendous boost in reaching the goal of standardizing technical language. In this paper, we propose two methods for standardizing language. These methods have been applied to the dataset of near misses, collected during the inspections at Major-Accident Hazard (MAH) Industries.1</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Standardizing technical language has always been
a strong necessity of the technological society.
Artifacts, objects, measures and so on should have
a clear name and a clear description in order to
assure mutual understanding, which leads to the
reach of important goals in building and
controlling machines. However, language
standardization has always the same problem: language is a
social phenomenon
        <xref ref-type="bibr" rid="ref6">(de Saussure, 1916)</xref>
        . Hence,
whenever a group gather for designing or using
a technical object, this group can develop a
specific sub-language or just adapt the shared
technical language. This adapted sub-language can be
then effectively used to refer to parts of this
technical object. It is sufficient that group members
agree upon this language and the mutual
understanding occur. Yet, the language used by the
specific group may prevent the others to understand
what is written.
      </p>
      <p>Nowadays, Natural Language Processing as
well as the widespread use of computerized
document writing can give a tremendous boost in
reaching the goal of standardizing technical
language. Language in use can be captured and, then,
1Copyright c 2019 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).
analyzed. Technical people can be invited to use a
standardized dictionary with writing suggestions.</p>
      <p>
        This paper discusses two different methods of
standardizing technical languages, which have
been applied to a dataset of near misses coming
from the inspections at Major-Accident Hazard
(MAH) industries, named also “Seveso”
industries. . The first method aims to help a
standardization agency to propose the standard language for
writing these reports. We proposed to analyze
language in use by word embedding similarity such
that the standardization agency can propose a
language that is close to the one used. The second
method aims to reduce the use of unnecessary
synonyms in compiling reports of near misses. In fact,
using unnecessary synonyms may result in
confusing the report. For this problem, we propose to use
a combination of language modeling derived from
the CBOW model of the word2vec
        <xref ref-type="bibr" rid="ref5">(Mikolov et al.,
2013)</xref>
        along with a classical cosine similarity
using word embeddings. We experimented with a
dataset of anonymized reports of near misses from
Seveso Industries, which INAIL has institutionally
collected.
      </p>
      <p>The rest of the paper is organized as follows.
Section 2 describes the application scenario and
the dataset. Section 3 shortly reports on the
models used in this study and proposes the two tasks.
Section 4 reports on a preliminary analysis of the
possible results of the system. Finally, Section 5
draws some conclusions and proposes further
investigations.
2
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>Background</title>
      <sec id="sec-2-1">
        <title>Scenario</title>
        <p>The European “Seveso” Directive deals with the
control of major-accident hazards involving
dangerous substances, which can cause toxic clouds,
fire, or explosion with consequences to people,
as</p>
        <p>Data (Date): 2007-02-15
Trasudamento OCD da serbatoio di stoccaggio OCD
Durante le operazioni di riempimento del serbatoio K2 da nave cisterna, si e` notato
un leggero trasudamento diOCD per corrosione del mantello (sottospessore localizzato
mantello serbatoio) a quota 6 metri circa lungo il latoovest. Uno degli operatori addetto
ai controlli durante la discarica della nave ha evidenziato l’evento. L’operazionedi
discarica della nave cisterna e` stata fermata. Non si sono avuti rilasci, a meno del leggero
trasudamento.
serbatoio
olio combustibile (ocd)
Descrizione (Description)
Fallimento procedure di manutenzione e
controllo.</p>
        <p>Azioni pianificate (Planned Actions)
Fuori servizio e bonifica del serbatoio.
sets and environment, also outside the
establishments. All European Member States apply this
Directive, which foresees periodical inspections by
National Competent Authorities; in Italy, Inail is
one of these authorities. During the inspection,
the operator has to provide the inspectors with the
list of near-misses, minor incidents, and accidents
occurred in the last ten years. Near misses and
minor incidents are events of losses of
containment, involving dangerous substances with none
or minor consequences, respectively. In Seveso
industries, the registration and the analysis of near
misses is strongly recommended, as they can be
considered as precursors of incidents with serious
consequences.</p>
        <p>In Italy, under Seveso legislation, there are
about a thousand industries, including refineries,
petrochemical, and chemical. One of the pillars
of the Seveso Directive is the Safety Management
System SMS, whose adoption is mandatory for the
establishments’ operators, in order to control
major accident hazards.</p>
        <p>The Safety Management System (SMS),
implemented by the establishment’s operator, addresses
technical measures and organizational procedures
in order to guarantee human, asset and
environmental safety, with a view to the prevention of
major accident or the mitigation of their
consequences.</p>
        <p>In the recent inspections, the focus is often
toward the study of the incidents and near misses
(see Figure 1). The approach based on near-miss
discussion is considered more “risk based” as it is
able to single out the critical issues of the safety
system.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Corpus</title>
        <p>The dataset refers to the near misses reports
provided by the operators of “Seveso” establishments.
The collection of reports on near misses, hereafter
referred as REP corpus, consists of 1300
documents called ”operative experiences”. These
operative experiences span the period from 2006 to
2017 and are related to 320 plants.</p>
        <p>Each “operative experience” tells about the
events occurred in the recent past (see Figure 1
for an example). Each event is registered by the
operators filling in a pre-defined form. The
document contains information including the date, a
title summarizing the event, a short description, the
reference to failed, missing or misapplied
technical or procedural barriers, those that stopped the
escalation and the recovering actions, and
eventually the planned actions for improving the safety.</p>
        <p>
          It is out of scope of this paper to discuss the
different methods used in the literature to
manage near miss information for improving the safety
management system. However, the common
objective is to exploit the valuable information
contained.
          <xref ref-type="bibr" rid="ref2">(Ansaldi et al., 2018)</xref>
          describe a method
to extract knowledge from this collection of
documents, and to support foresights or intuitions about
the safety of process industries. Another
application has been developed for understanding if
the lessons from major accidents have been fully
learnt and implemented
          <xref ref-type="bibr" rid="ref1">(Ansaldi et al., 2016)</xref>
          . The
issue has been addressed by looking for
similarities between near misses and accident
characteristics, and by evaluating their semantic distance.
        </p>
        <p>Although the form of the document is the same
adopted for all operators, the compiling mode
varies by the establishments and by the type of
event recorded. The accuracy of the documents is
not homogeneous and the interpretation of
operative experience concept changes from one
establishment to another; their carefulness varies on the
sector activities, and often reveals the safety
culture of the establishment. At a few establishments,
just the releases of hazardous substance without
consequences are registered. In other cases,
reports include anomalies, unsafe situations,
failures, and trivial errors; that is, events not directly
related to major accident hazard. The documents
are various, but represent truthful pictures of
deviations occurred inside the establishment.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Methods</title>
      <p>The overall goal is to show that existing
methodologies can help in standardizing language in the
specific case of reports on near misses on Seveso
industries and we aim to perform this
standardization with two tools: (1) analyzing similarities
among words in current reports; (2) propose a
methodology to help in writing these reports.
3.1</p>
      <sec id="sec-3-1">
        <title>Challenges</title>
        <p>The specific case of reports on near misses is
particular for several compelling reasons. The first
compelling reason is that reports are written by
operators belonging to sub-communities of
speakers. In fact, people working in each plant can be
considered a sub-community, which shares a
particular language. Hence, standardizing language
of reports means also harmonize sub-languages of
different sub-communities, which do not interact.
This problem is particularly severe when the aim
is to standardize language across the whole Seveso
industries. The second compelling reason is the
different background of reports’ writers. Reports
are in fact written by operators, which may have
different knowledge, different school degree, and
different cultural background. This reason makes
particularly relevant the goal to help writers in
compiling reports on near misses.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Enabling Tools and Methodologies</title>
        <p>
          To meet the overall goal , we here experiment with
standard and well-assessed models and
methodologies: the notion of word embedding. In fact,
the long tradition of representing word meaning in
vectors is what is needed to: (1) help the
standardization organism to develop a common and
acceptable language; (2) devise ways to suggest more
appropriate words to writers of reports. In this study,
we used two different word embeddings:
General Language Word Embeddings
(GLwe)
          <xref ref-type="bibr" rid="ref4">(Cimino et al., 2018)</xref>
          : these are
word embeddings pre-trained with word2vec
          <xref ref-type="bibr" rid="ref5">(Mikolov et al., 2013)</xref>
          on a general purpose
corpus of the Italian language, that is, itWaC
          <xref ref-type="bibr" rid="ref3">(Baroni et al., 2009)</xref>
          Domain-adapted Word Embeddings (Dawe):
these are word embeddings obtained training
word2vec
          <xref ref-type="bibr" rid="ref5">(Mikolov et al., 2013)</xref>
          using GLwe
as initialization and the REP training corpus
3.3
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>Task 1: Understanding Language of</title>
      </sec>
      <sec id="sec-3-4">
        <title>Near-Miss Reports</title>
        <p>We aim to provide the standardization organism,
that is, INAIL, the possibility to investigate the
language used in these reports on near misses. The
possibility we explored is to provide a visual
representation of similarity computed using similarity
among word embeddings. Giving this visual
representation, researchers in INAIL can devise the
definition of a standard language that is built on a
common and shared language. This idea is similar
to what has been done in the past for terminology
extraction. The real added value is that similarity
among terms is computed according to word
embeddings.
3.4</p>
      </sec>
      <sec id="sec-3-5">
        <title>Task 2: Standardizing Report with</title>
      </sec>
      <sec id="sec-3-6">
        <title>Assisted Writing</title>
        <p>We aim to provide a tool to assist operators while
writing reports. We explored the first capability of
this tool, that is, avoiding unnecessary use of
synonyms while writing. In the Italian tradition,
using repeating words is seen as bad writing. Hence,
when writing, synonyms are used to introduce a
variation. However, for technical documents,
unnecessary use of synonyms in core concepts may
introduce misunderstanding. Hence, we envisage
a tool that helps in reducing use of synonyms.</p>
        <p>The algorithm governing the tool works as
follows. While writing a report, the algorithm
accumulate words in a set W . Whenever a new
content word w is added, the algorithms compute
the similarity with the words in the set W . If
there is a word w0 2 W for which the
similarity sim(w; w0) = wT w0 is above a threshold ,
the algorithm suggests w0 as a possible
substitution of w. In this way, the operator is forced to
(a) General Language Word Embeddings</p>
        <p>Text
La perdita non si era evidenziata al controllo dell’area effettuato preliminarmente all’inizio attivita`,
ne´ rilevata dal CTM presente in zona area (sim = 0:64)</p>
        <p>attivita` (sim = 0:31)</p>
        <p>Necessita` di prevedere un piu` elevato grado di protezione contro la perdita di contenimento da
fondo serbatoi. La fuoriuscita perdita (sim = 0:54)</p>
        <p>contenimento (sim = 0:42)
think whether the word w0 that s/he already used
is similar to the word s/he is using now. In this
case, w0 can be used to replace w and an
unnecessary synonym is avoided.
For the first task, we experimented with the two
dictionaries: the General Language word
embeddings (GLwe) and the Domain-adpated word
embeddings (Dawe). Similarity spaces for the two
word embeddings (see Figure 2) may help in
understanding whether unnecessary synonyms are
used and, hence, suggest a standardized word that
should be used for a group of words.</p>
        <p>Using the two dictionaries, we built two
similarity spaces (Figure 2) obtained as follows. We
selected 10 frequent words in the REP training
corpus and, then, we presented in the two figures the
top 15 words that are more similar to the 10
selected frequent words. The similarity spaces are
built according to GLwe (Figure 2a) and
according to Dawe (Figure 2b).</p>
        <p>The Dawe similarity space (Figure 2b) gives
apparently better hints on how words are used. The
dictionary seems to be more tailored to the specific
domain. In fact, there is an interesting groups of
words such as favvenuto, accaduto, occorso,
verificatosi g and f causato, provocatog. These gropus
are missing in the GLwe similarity space (Figure
2a).
For the second task, we experimented with some
sample reports. The algorithm in action is reported
in Figure 3. This test has been carried out on
existing reports and aimed to show that some words
can be replaced with previously used words. In the
report #174, the word zona can be replaced with
the word area, which has been previously used.</p>
        <p>In the report #175, the word focolaio could be
replaced with the word incendio. Finally, in the
report #109, the word fuoriuscita can be replaced
with the word perdita. However, the operator is
free to accept or refuse the suggestion if this is not
satisfactory.
5</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusion and Future Work</title>
      <p>
        Standardizing language is a need of our
technological society. In this paper, we investigated the
possibility of using modern NLP techniques to reach
this goal in the specific scenario of near misses
in Seveso Industries. Initial results on the corpus
provided by Inail are interesting and leave room
for improvement. Future model should include
the treatment of multi-word expressions by
using compositional distributional semantic models
        <xref ref-type="bibr" rid="ref8 ref9">(Zesch et al., 2013; Zanzotto et al., 2015)</xref>
        , should
merge distributional and ontological models, and
should include a clear model for repaying
knowledge producers
        <xref ref-type="bibr" rid="ref7">(Zanzotto, 2019)</xref>
        .
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Silvia</given-names>
            <surname>Maria</surname>
          </string-name>
          <string-name>
            <surname>Ansaldi</surname>
          </string-name>
          , Patrizia Agnello, and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Bragatto</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Incidents triggered by failures of level sensors</article-title>
          .
          <source>Chemical Engineering Transactions</source>
          ,
          <volume>53</volume>
          :
          <fpage>223</fpage>
          -
          <lpage>228</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Silvia</given-names>
            <surname>Maria</surname>
          </string-name>
          <string-name>
            <surname>Ansaldi</surname>
          </string-name>
          , Annalisa Pirone, Rosaria Vallerotonda Maria, Paolo Bragatto, Patrizia Agnello, and Corrado Delle Site.
          <year>2018</year>
          .
          <article-title>How inspections outcomes may improve the foresight of operators and regulators in seveso industries</article-title>
          .
          <source>Chemical Engineering Transactions</source>
          ,
          <volume>67</volume>
          :
          <fpage>367</fpage>
          -
          <lpage>372</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Marco</given-names>
            <surname>Baroni</surname>
          </string-name>
          , Silvia Bernardini, Adriano Ferraresi, and
          <string-name>
            <given-names>Eros</given-names>
            <surname>Zanchetta</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>The wacky wide web: a collection of very large linguistically processed web-crawled corpora</article-title>
          .
          <source>Language Resources and Evaluation</source>
          ,
          <volume>43</volume>
          (
          <issue>3</issue>
          ):
          <fpage>209</fpage>
          -
          <lpage>226</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Andrea</given-names>
            <surname>Cimino</surname>
          </string-name>
          , Lorenzo De Mattei, and Felice Dell'Orletta.
          <year>2018</year>
          .
          <article-title>Multi-task learning in deep neural networks at EVALITA 2018</article-title>
          .
          <source>In Proceedings of the Sixth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2018</year>
          )
          <article-title>co-located with the Fifth Italian Conference on Computational Linguistics (CLiC-it</article-title>
          <year>2018</year>
          ), Turin, Italy,
          <source>December 12-13</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          , Kai Chen, Greg Corrado, and
          <string-name>
            <given-names>Jeffrey</given-names>
            <surname>Dean</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Efficient estimation of word representations in vector space</article-title>
          .
          <source>CoRR, abs/1301</source>
          .3781.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Ferdinand de Saussure</surname>
          </string-name>
          .
          <year>1916</year>
          .
          <article-title>Cours de linguistique ge´ne´rale</article-title>
          . Payot, Paris.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Massimo Zanzotto</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Viewpoint: Humanin-the-loop Artificial Intelligence</article-title>
          .
          <source>Journal of Artificial Intelligence Research</source>
          ,
          <volume>64</volume>
          :
          <fpage>243</fpage>
          -
          <lpage>252</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Massimo</surname>
          </string-name>
          <string-name>
            <surname>Zanzotto</surname>
          </string-name>
          , Lorenzo Ferrone, and
          <string-name>
            <given-names>Marco</given-names>
            <surname>Baroni</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>When the whole is not greater than the combination of its parts: A ”decompositional” look at compositional distributional semantics</article-title>
          .
          <source>Comput. Linguist.</source>
          ,
          <volume>41</volume>
          (
          <issue>1</issue>
          ):
          <fpage>165</fpage>
          -
          <lpage>173</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>T.</given-names>
            <surname>Zesch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Korkontzelos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.M.</given-names>
            <surname>Zanzotto</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Biemann</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Semeval-2013 task 5: Evaluating phrasal semantics</article-title>
          . volume
          <volume>2</volume>
          , pages
          <fpage>39</fpage>
          -
          <lpage>47</lpage>
          . Cited By 12.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>