<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>[Demo] A webtool for analyzing land-use planning documents</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mohammad Amin Farvardin</string-name>
          <email>amin.farvardin@teledetection.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Eric Kergosien</string-name>
          <email>eric.kergosien@univ-lille3.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mathieu Roche</string-name>
          <email>mathieu.roche@cirad.fr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maguelonne Teisseire</string-name>
          <email>maguelonne.teisseire@teledetection.fr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>GERiiCO</institution>
          ,
          <addr-line>Univ. Lille 3</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>LIRMM</institution>
          ,
          <addr-line>CNRS, Univ. Montpellier</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>UMR TETIS (Irstea</institution>
          ,
          <addr-line>Cirad, AgroParisTech), Montpellier</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In previous work, di erent methods have been proposed in order to semi-automatically mine geospatial information and opinions in documents [3]. In this paper, we present the Web application, SentiAnnotator, based on NLP methods to extract and visualize geospatial information with the associated entities. The evaluation of our application shows good results on a French corpus, i.e. F-measure of 0.74 and 0.75 respectively for the identi cation of spatial features and organizations.</p>
      </abstract>
      <kwd-group>
        <kwd>Land-use planning</kwd>
        <kwd>Web application</kwd>
        <kwd>Geospatial features</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Researchers and experts of land-use planning are looking for decisional tools
for helping them to have an overview of user's awareness on territories. In this
context, we de ned the Opiland method that enables to semi-automatically
analyze sentiments related to land-use planning documents [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. In this paper, we
present the developed software and the associated web services for discovering
and for integrating meaning in free texts available on the Web. This kind of
textual data (e.g. blogs, newspapers, and so on) is generally complex but useful
for public policy dialogue and decision-making. We thus propose an approach
that enables (i) to automatically extract features related to land-use planning,
and (ii) to give to experts the possibility of evaluating sentiments related to
geospatial features, with the ultimate objective of evaluating the policy impact
for adapting their decisions. The main originality of our software concerns the
integration of di erent levels of semantics present in a document. This is really
crucial in order to improve the analysis of information, specially for land-use
planning domain. Generally in the opinion mining eld, the connection between
opinion and topic is studied. Actually in the land-use planning domain, it is
necessary to take into account a larger number of relevant elements like spatial
features and organizations. Our software o ers this possibility for helping the
experts to do a ner analysis of Web data.
      </p>
      <p>In this demo paper, we present the SentiAnnotator web application4 to
extract di erent features related to land-use planning. Hereafter, in Section 2, we
focus more precisely on the deployment of natural language processing methods
to extract geospatial information, i.e. spatial features and organizations. The web
application for uploading, indexing, and marking textual documents is detailed
in Section 3. Finally, after a quick look to our system evaluation in Section 4,
future work related to our project is drawn in Section 5.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Geospatial information extraction</title>
      <p>
        Named Entity Recognition (NER) methods identify di erent types of Named
Entities (NE): dates, people, organizations, themes, numeric values, as well as
locations. There is a signi cant number of available systems, such as OpenNLP5,
OpenCalais6, and CasEN [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. To recognize Named Entities several approaches
are based on supervised learning methods. In this context, a bag-of-words
representation is often used [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. But this kind of statistical approach is not adapted
for small data sets we are faced with in the land-use planning domain. Other
approaches based on symbolic methods concern geoparsing [
        <xref ref-type="bibr" rid="ref2 ref4 ref5">2, 4, 5</xref>
        ]. The work
of [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] proposes linguistic patterns to extract Spatial Features (SF) from texts.
These patterns are based on a cognitive model where SF is composed of at least
one NE and one variable number of spatial indicators specifying its location.
Five spatial relation types are considered: orientation, distance, adjacency,
inclusion, and geometric which de nes union or intersection linking two SF. In
our proposal, we add new patterns to improve the automatic identi cation of
SF (absolute spatial features (A SF) and relative spatial features (R SF) [
        <xref ref-type="bibr" rid="ref3 ref5">3, 5</xref>
        ]).
The SF annotation is based on the classical typology of the domain and more
precisely on the sub-types of locations. Locations can be polysemous: human
constructions (e.g. buildings) and addresses (e.g. streets). To take into account
all these language speci cities, some rules (patterns) have been added.
      </p>
      <p>Moreover we propose a new type of patterns to identify Organizations (OE)
which is a speci c NE useful for land-use planning domain. The addition of
speci c rules enables to identify OE which could be confused with SF in documents.
Such rules are: (1) an OE is followed by an action verb; (2) an OE is proceeded
by prepositions: with, by, for, on behalf of, etc.</p>
      <p>In order to manage these geospatial information, we developed the web
application SentiAnnotator (http://siso.teledetection.fr/viewer.jsp). A screenshot
is presented Figure 1. The web services use the Gate system7. After uploading
a corpus (in French for this current version), Spatial Features and
Organizations are extracted using the implemented rules. Moreover other concepts are</p>
      <sec id="sec-2-1">
        <title>4 http://siso.teledetection.fr/ 5 https://opennlp.apache.org/ 6 http://www.opencalais.com/ 7 https://gate.ac.uk/</title>
        <p>
          A webtool for analyzing land-use planning documents
extracted using Gate: (i) thematics based on a lexicon using Agrovoc thesaurus8,
(ii) Opinions related to land-use planning domain [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>The SentiAnnotator Web application</title>
      <p>The web application (Figure 1) allows users to upload corpora, to index
documents with speci c web services in order to mark di erent kinds of information
(spatial features, organizations, opinions, and themes), to visualize, to correct
the results, and to download validated results in XML format. More speci cally,
it is possible to upload corpora (frame 1), each marked corpus is saved on the
server and automatically available in the web application (frame 2). After having
downloaded documents, users can select the marked features (frame 5), see the
results on the selected documents in frame 3. In this frame, spatial features are
in blue color, organizations in purple color, the positive opinions in green color,
negative ones in red color and neutral in yellow color. By selecting di erent
categories from frame 5, the related marked information will be kindles in frame 3
and listed by type in frame 4. In case of nding any mistakes, users can unselect
marked information (frame 4). Finally expert can export the selected corrected
documents by clicking the top right bottom. The downloaded corpus consists
of selected documents with the marked information except those were removed
by the user. The administration page allows users to upload, edit, and delete
pipelines de ned in the Gate format. It is also possible to remove processed
corpora and to edit the uploaded pipeline rules and the available lexicons.</p>
      <sec id="sec-3-1">
        <title>8 http://aims.fao.org/fr/agrovoc</title>
        <p>Farvardin et al.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Evaluation</title>
      <p>Three experts of the project evaluated the process for extracting geospatial
information by using SentiAnnotator application. We use a French corpus
composed of 4328 words (71 spatial features and 117 organizations). The evaluations
(with classical measure, i.e. Precision, Recall, and F-measure) have been
investigated by comparing the manual extraction done by experts with the web service
results. For SF, we obtain an excellent recall (0.91) and an acceptable precision
(0.62), the F-measure is 0.74. We extract the great majority of SF but the rules
still return some errors. The rules to identify OE are very e cient and return
high precision (0.85) but the value of recall is lower (0.67). The F-measure for
organization identi cation is 0.74. The rules for organization extraction seem
well-adapted to the domain but they have to be extended in order to improve
the recall that remains low.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion and Future Work</title>
      <p>In this paper, we have presented a Web application called SentiAnnotator
including web services (1) to annotate corpora with features related to
landuse planning, and (2) to evaluate achieved approaches with experts. Experts are
using this tool for analyzing the construction project of a road around Villeveyrac
(France). Future work will be dedicated to the improvements of the de ned
linguistics patterns for discovering NE in order to tackle the issues related to
the land-use planning speci cities and the multilingual aspects. We also plan to
extend our approach to di erent types of textual contents such as tweets.</p>
      <p>Acknowledgments: The authors thank Midi Libre (French newspaper) for its
expertise on the corpus and all partners of the Senterritoire project for their
involvement (MSH-M, Geosud Equipex, Numev Labex, and Tectoniq PEPS project).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>X.</given-names>
            <surname>Carreras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Marquez</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L. S.</given-names>
            <surname>Padro</surname>
          </string-name>
          .
          <article-title>A simple named entity extractor using adaboost</article-title>
          .
          <source>In In Proceedings of CoNLL-2003</source>
          , pages
          <fpage>152</fpage>
          {
          <fpage>155</fpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>M.</given-names>
            <surname>Gaio</surname>
          </string-name>
          and
          <string-name>
            <given-names>V.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          .
          <article-title>Towards heterogeneous resources-based ambiguity reduction of sub-typed geographic named entities</article-title>
          .
          <source>In Int. Conf. of GeoSpatial Semantics</source>
          , pages
          <volume>217</volume>
          {
          <fpage>234</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>E.</given-names>
            <surname>Kergosien</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Laval</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Roche</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Teisseire</surname>
          </string-name>
          .
          <article-title>Are opinions expressed in land-use planning documents?</article-title>
          <source>International Journal of Geographical Information Science</source>
          ,
          <volume>28</volume>
          (
          <issue>4</issue>
          ):
          <volume>739</volume>
          {
          <fpage>762</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Leidner</surname>
          </string-name>
          and
          <string-name>
            <given-names>M. D.</given-names>
            <surname>Lieberman</surname>
          </string-name>
          .
          <article-title>Detecting geographical references in the form of place names and associated spatial natural language</article-title>
          .
          <source>SIGSPATIAL Special</source>
          ,
          <volume>3</volume>
          (
          <issue>2</issue>
          ):5{
          <fpage>11</fpage>
          ,
          <year>July 2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>J.</given-names>
            <surname>Lesbegueries</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Sallaberry</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Gaio</surname>
          </string-name>
          .
          <article-title>Associating spatial patterns to textunits for summarizing geographic information</article-title>
          .
          <source>In Proceedings of ACM SIGIR</source>
          <year>2006</year>
          .
          <article-title>Geographic Information Retrieval</article-title>
          , Workshop, pages
          <volume>40</volume>
          {
          <fpage>43</fpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>D.</given-names>
            <surname>Maurel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Friburger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.-Y.</given-names>
            <surname>Antoine</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Eshkol-Taravella</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Nouvel</surname>
          </string-name>
          .
          <article-title>Casen: a transducer cascade to recognize french named entities</article-title>
          .
          <source>TAL</source>
          ,
          <volume>52</volume>
          (
          <issue>1</issue>
          ):
          <volume>69</volume>
          {
          <fpage>96</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>