<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Descriptors for the detection of the chemical risk</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Natalia Grabar</string-name>
          <email>natalia.grabar@univ-lille3.fr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thierry Hamon</string-name>
          <email>hamon@limsi.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>LIMSI-CNRS, Orsay, Universite ́ Paris 13</institution>
          ,
          <addr-line>Sorbonne Paris Cite ́</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>UMR8163 STL, CNRS, Universite ́ Lille 3</institution>
          ,
          <addr-line>Villeneuve d'Ascq</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2015</year>
      </pub-date>
      <abstract>
        <p>We propose an experience on the automatic detection of sentences conveying the notion of chemical risk. Our objective is to study which resources are useful for the automatic detection of such sentences. Lexical, semantic and opinion-oriented content of the sentences is studied. Our results indicate that not only lexical and semantic content must be taken into account, but also markers related to the modality, opinion and polarity.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Chemical risk is relative to situations in which
chemical products are dangerous for human or
animal health and consumption, and for
environment. The automatization of the process can help
the experts to control and manage large amounts
of scientific literature, that have to be analyzed
to support the decision making process
        <xref ref-type="bibr" rid="ref6">(van der
Sluijs et al., 2008)</xref>
        . The sentences that must be
recognized are for instance: The Panel concluded that
the current NOAEL for BPA (5 mg/kg b.w./day)
would be sufficiently low to exclude any concern
for this effect, or Despite this lack of evidence,
the possibility of poultry and egg consumption as
an exposure route to HPAIV remains a concern to
food safety experts. Such sentences are to be
assigned in categories related to the chemical risk:
the first sentence is related to the significance of
the results, while the second is related to the
quality of the scientific hypothesis. If such sentences
are detected in scientific publications or reports,
it means that these publications or reports contain
information not fully reliable and can possibly
indicate the insufficiency of the corresponding
studies and the presence of the risk.
      </p>
      <p>
        The chemical risk is poorly studied, although
the notion of the risk is addressed by other works:
building of the dedicated resources
        <xref ref-type="bibr" rid="ref2">(Makki et al.,
2008)</xref>
        , exploring of known industrial incidents
        <xref ref-type="bibr" rid="ref5 ref7">(Tulechki and Tanguy, 2012)</xref>
        , computing the
exposition to the risk
        <xref ref-type="bibr" rid="ref3">(Marre et al., 2010)</xref>
        . Our
objective is to study which resources are useful for the
automatic detection of the sentences which convey
the notion of the chemical risk.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Material and Methods</title>
      <p>In addition to the lexical and semantic content of
the text, we use several kinds of resources in order
to favour one aspect or another. These resources
contain markers oriented on modality, opinion and
polarity expressed by the authors on the proposed
experiements: (1) uncertainty (possible, should,
may, usually) indicates that there are doubts on
the results presented, their interpretation, etc.; (2)
negation (no, neither, lack, absent, missing)
indicates that the results have not been observed, that
the study does not respect the expected norms,
etc.; (3) limitations (only, shortcoming,
insufficient) indicates that there are some limits of the
work, such as unsufficient sample size, small
number of tests or doses explored, etc.; (4)
approximation (approximately, commonly, estimated)
indicates other kinds of insufficiency related to
imprecise values of substances, samples, dosage, etc.</p>
      <p>
        The work is done with the corpus on
chemical risk reporting on several chemical experiments
with bisphenol A
        <xref ref-type="bibr" rid="ref1">(EFSA Panel, 2010)</xref>
        . It contains
over 80,000 occurrences. The reference data are
obtained through a manual categorization of the
corpus sentences: 425 sentences are assigned to
55 classes of the chemical risk.
      </p>
      <p>1
0.8
0.6
0.4
0.2
0
192
freq
norm
tfidf
freq
norm
tfidf
all
form lemm lf
descripteurs
lft
stag
tag
all
form lemm lf
descripteurs
lft
stag
tag
(a) Significance of the results
(b) Natural variability</p>
      <p>
        We tackle the problem through the supervized
categorization with the Weka platform
        <xref ref-type="bibr" rid="ref8">(Witten
and Frank, 2005)</xref>
        . Sentences correspond to the
units, while 7 classes (most frequent) of the
chemical risk are the categories to which the sentences
have to be assigned. The resources and the
linguistic annotation of corpus
        <xref ref-type="bibr" rid="ref4">(Schmid, 1994)</xref>
        provide
several descriptors. These are used to build
several sets of descriptors. They represent the
semantic and linguistic content of the sentences: forms
(the forms such as they occur in the corpus),
lemmas (lemmatized forms), lf (combination of forms
and lemmas), tag (POS tags, such as nouns, verbs,
adjectives), lft (combination of forms, lemmas and
POS-tags), stag (semantic tags of words, such as
uncertainty, negation, limitations), all
(combination of all the descriptors available). The
descriptors are weighted with various methods (freq raw
frequency, norm normalization by the length of the
sentences, and tfidf tf-idf normalization).
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <p>
        Figure 1 presents some results obtained for two
categories: Significance of the results and Natural
variability of the results. We can observe some
difference according to the descriptors: the
exploitation of forms, semantic tags (with
Significance of the results) and various combinations of
descriptors provide results that are often better for
these two categories and for other categories. We
assume that these two kinds of descriptors (lexical
and semantic content of corpus and the descriptors
related to modality, polarity and opinion
        <xref ref-type="bibr" rid="ref5 ref7">(Vinodhini and Chandrasekaran, 2012)</xref>
        ) provide
complementary views on the content and should be
combined. These results also indicate that chemical
risk is not fully conceptual category but is also
related to subjective and contextual values.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>EFSA</given-names>
            <surname>Panel</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Scientific opinion on Bisphenol A: evaluation of a study investigating its neurodevelopmental toxicity, review of recent scientific literature on its toxicity and advice on the danish risk assessment of Bisphenol A</article-title>
          .
          <source>EFSA journal</source>
          ,
          <volume>8</volume>
          (
          <issue>9</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>110</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>J</given-names>
            <surname>Makki</surname>
          </string-name>
          , AM Alquier, and
          <string-name>
            <given-names>V</given-names>
            <surname>Prince</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Ontology population via NLP techniques in risk management</article-title>
          .
          <source>In Proceedings of ICSWE.</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>A</given-names>
            <surname>Marre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S</given-names>
            <surname>Biver</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M</given-names>
            <surname>Baies</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C</given-names>
            <surname>Defreneix</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C</given-names>
            <surname>Aventin</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Gestion des risques en radiothe´rapie</article-title>
          . Radiothe´rapie,
          <volume>724</volume>
          :
          <fpage>55</fpage>
          -
          <lpage>61</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>H</given-names>
            <surname>Schmid</surname>
          </string-name>
          .
          <year>1994</year>
          .
          <article-title>Probabilistic part-of-speech tagging using decision trees</article-title>
          .
          <source>In International Conference on New Methods in Language Processing</source>
          , pages
          <fpage>44</fpage>
          -
          <lpage>49</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>N</given-names>
            <surname>Tulechki</surname>
          </string-name>
          and
          <string-name>
            <given-names>L</given-names>
            <surname>Tanguy</surname>
          </string-name>
          .
          <year>2012</year>
          . Effacement de dimensions de similarite´
          <article-title>textuelle pour l'exploration de collections de rapports d'incidents ae´ronautiques</article-title>
          .
          <source>In TALN</source>
          , pages
          <fpage>439</fpage>
          -
          <lpage>446</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Jeroen P van der Sluijs</surname>
          </string-name>
          , Arthur C Petersen,
          <string-name>
            <surname>Peter H M Janssen</surname>
          </string-name>
          ,
          <string-name>
            <surname>James S Risbey</surname>
          </string-name>
          , and
          <string-name>
            <surname>Jerome R Ravetz</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Exploring the quality of evidence for complex and contested policy decisions</article-title>
          .
          <source>Environ. Res. Lett.</source>
          ,
          <volume>3</volume>
          (
          <issue>2</issue>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>G</given-names>
            <surname>Vinodhini and RM Chandrasekaran</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Sentiment analysis and opinion mining: A survey</article-title>
          .
          <source>International Journal of Advanced Research in Computer Science and Software Engineering</source>
          ,
          <volume>2</volume>
          (
          <issue>6</issue>
          ):
          <fpage>282</fpage>
          -
          <lpage>292</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>I.H.</given-names>
            <surname>Witten</surname>
          </string-name>
          and
          <string-name>
            <given-names>E.</given-names>
            <surname>Frank</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Data mining: Practical machine learning tools and techniques</article-title>
          . Morgan Kaufmann, San Francisco.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>