<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>How to deal with heterogeneous data?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mathieu Roche UMR TETIS (Cirad</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Irstea</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>AgroParisTech) - France mathieu.roche@cirad.fr</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>LIRMM (CNRS, University of Montpellier) - France Web:</institution>
        </aff>
      </contrib-group>
      <fpage>19</fpage>
      <lpage>20</lpage>
      <abstract>
        <p>Data heterogeneity and content In the context of Web 2.0, text content is often heterogenous (i.e. lexical heterogeneity). For instance, some words may be shortened or lengthened with the use of specific graphics (e.g. emoticons) or hashtags. Specific processing is necessary in this context. For instance, with an opinion classification task based on the message SimBig is an aaaaattractive conference!, the results are generally improved by removing repeated characters (i.e. a). But information on the sentiment intensity identified by the character elongation is lost with this normalization. This example highlights the difficulty of dealing with heterogeneous textual data content. The following sub-section describes the heterogeneity according the document types (e.g. images and texts).</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <sec id="sec-1-1">
        <title>Heterogeneity and document types</title>
        <p>
          Impressive amounts of high spatial resolution
satellite data are currently available. This raises
the issue of fast and effective satellite image
analysis as costly human involvement is still required.
Meanwhile, large amounts of textual data are
available via the Web and many research
communities are interested in the issue of
knowledge extraction, including spatial information. In
this context, image-text matching improves
information retrieval and image annotation techniques
          <xref ref-type="bibr" rid="ref4">(Forestier et al., 2012)</xref>
          . This provides users with
a more global data context that may be useful
for experts involved in land-use planning
          <xref ref-type="bibr" rid="ref1">(Alatrista Salas et al., 2014)</xref>
          .
2
        </p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Text-mining method for matching heterogenous data</title>
      <p>
        A generic approach to address the heterogeneity
issue consists of extracting relevant features in
documents. In our work, we focus on 3 types
of features: thematic, spatial, and temporal
features. These are extracted in textual documents
using natural language processing (NLP) techniques
based on linguistic and statistic information
        <xref ref-type="bibr" rid="ref7">(Manning and Sch u¨tze, 1999)</xref>
        :
• The extraction of thematic information is
based on the recognition of relevant terms in
texts. For instance, terminology extraction
techniques enable extraction of single-word
terms (e.g. irrigation) or phrases (e.g.
rice crops). The most efficient
state-ofthe-art term recognition systems are based
on both statistical and linguistic information
        <xref ref-type="bibr" rid="ref6">(Lossio-Ventura et al., 2015)</xref>
        .
• Extracting spatial information from
documents is still challenging. In our work, we
use patterns to detect these specific named
entities. Moreover, a hybrid method enables
disambiguation of spatial entities and
organizations. This method combines symbolic
approaches (i.e. patterns) and machine learning
techniques
        <xref ref-type="bibr" rid="ref9">(Tahrat et al., 2013)</xref>
        .
• In order to extract temporal expressions in
texts, we use rule-based systems like
HeidelTime
        <xref ref-type="bibr" rid="ref8">(Stro¨ tgen and Gertz, 2010)</xref>
        . This
multilingual system extracts temporal
expressions from documents and normalizes them.
HeidelTime applies different normalization
strategies depending on the text types, e.g.
news, narrative, or scientific documents.
      </p>
      <p>These different methods are partially used in
the projects summarized in the following section.
More precisely, Section 3 presents two projects
that investigate heterogeneous data in agricultural
the domain.1
3</p>
    </sec>
    <sec id="sec-3">
      <title>Applications in the agricultural domain</title>
      <p>3.1</p>
      <sec id="sec-3-1">
        <title>Animal disease surveillance</title>
        <p>
          New and emerging infectious diseases are an
increasing threat to countries. Many of these
diseases are related to globalization, travel and
international trade. Disease outbreaks are
conventionally reported through an organized multilevel
health infrastructure, which can lead to delays
from the time cases are first detected, their
laboratory confirmation and finally public
communication. In collaboration with the CMAEE2 lab, our
project proposes a new method in the epidemic
intelligence domain that is designed to discover
knowledge in heterogenous web documents
dealing with animal disease outbreaks. The proposed
method consists of four stages: data acquisition,
information retrieval (i.e. identification of
relevant documents), information extraction (i.e.
extraction of symptoms, locations, dates, diseases,
affected animals, etc.), and evaluation by different
epidemiology experts
          <xref ref-type="bibr" rid="ref2">(Arsevska et al., 2014)</xref>
          .
3.2
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Information extraction from experimental data</title>
        <p>
          Our joint work with the IATE3 lab and
AgroParisTech4 deals with knowledge engineering issues
regarding the extraction of experimental data from
scientific papers to be subsequently reused in
decision support systems. Experimental data can be
represented by n-ary relations, which link a
studied topic (e.g. food packaging, transformation
process) with its features (e.g. oxygen permeability
in packaging, biomass grinding). This knowledge
is capitalized in an ontological and
terminological resource (OTR). Part of this work consists of
recognizing specialized terms (e.g. units of
measures) that have many lexical variations in
scientific documents in order to enrich an OTR
          <xref ref-type="bibr" rid="ref3">(Berrahou et al., 2013)</xref>
          .
        </p>
        <p>1http://www.textmining.biz/agroNLP.
html</p>
        <p>2Joint research unit (JRU) regarding the control of exotic
and emerging animal diseases – http://umr-cmaee.
cirad.fr</p>
        <p>3JRU in the area of agro-polymers and emerging
technologies – http://umr-iate.cirad.fr
4http://www.agroparistech.fr</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusion</title>
      <p>
        Heterogenous data processing enables us to
address several text-mining issues. Note that we
integrated the knowledge of experts in the core
of research applications summarized in Section 3.
In future work, we plan to investigate other
techniques dealing with heterogeneous data, such as
visual analytics approaches
        <xref ref-type="bibr" rid="ref5">(Keim et al., 2008)</xref>
        .
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>H.</given-names>
            <surname>Alatrista Salas</surname>
          </string-name>
          , E. Kergosien,
          <string-name>
            <given-names>M.</given-names>
            <surname>Roche</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Teisseire</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>ANIMITEX project: Image analysis based on textual information</article-title>
          .
          <source>In Proc. of Symposium on Information Management and Big Data (SimBig)</source>
          , Vol-
          <volume>1318</volume>
          , CEUR, pages
          <fpage>49</fpage>
          -
          <lpage>52</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>E.</given-names>
            <surname>Arsevska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Roche</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Lancelot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Hendrikx</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Dufour</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Exploiting textual source information for epidemio-surveillance</article-title>
          .
          <source>In Proc. of Metadata and Semantics Research - 8th Research Conference (MTSR) - Communications in Computer and Information Science</source>
          , Volume
          <volume>478</volume>
          , pages
          <fpage>359</fpage>
          -
          <lpage>361</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>S.L.</given-names>
            <surname>Berrahou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Buche</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dibie-Barthelemy</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Roche</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>How to extract unit of measure in scientific documents?</article-title>
          <source>In Proc. of International Conference on Knowledge Discovery and Information Retrieval (KDIR)</source>
          ,
          <source>Text Mining Session</source>
          , pages
          <fpage>249</fpage>
          -
          <lpage>256</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>G.</given-names>
            <surname>Forestier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Puissant</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wemmert</surname>
          </string-name>
          , and P. Ganc¸arski.
          <year>2012</year>
          .
          <article-title>Knowledge-based region labeling for remote sensing image interpretation</article-title>
          .
          <source>Computers, Environment and Urban Systems</source>
          ,
          <volume>36</volume>
          (
          <issue>5</issue>
          ):
          <fpage>470</fpage>
          -
          <lpage>480</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>D.A.</given-names>
            <surname>Keim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Mansmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Schneidewind</surname>
          </string-name>
          , J. Thomas, and
          <string-name>
            <given-names>H.</given-names>
            <surname>Ziegler</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Visual analytics: Scope and challenges</article-title>
          .
          <source>In Visual Data Mining</source>
          , pages
          <fpage>76</fpage>
          -
          <lpage>90</lpage>
          . Springer-Verlag.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>J.A.</given-names>
            <surname>Lossio-Ventura</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Jonquet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Roche</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Teisseire</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Biomedical term extraction: overview and a new methodology</article-title>
          .
          <source>Information Retrieval Journal (IRJ</source>
          )
          <article-title>- special issue ”Medical Information Retrieval”</article-title>
          , to appear.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>C.D. Manning</surname>
          </string-name>
          and H. Schu¨tze.
          <year>1999</year>
          .
          <article-title>Foundations of Statistical Natural Language Processing</article-title>
          . MIT Press, Cambridge, MA, USA.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>J.</given-names>
            <surname>Stro</surname>
          </string-name>
          <article-title>¨tgen and</article-title>
          <string-name>
            <given-names>M.</given-names>
            <surname>Gertz</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Heideltime: High quality rule-based extraction and normalization of temporal expressions</article-title>
          .
          <source>In Proc. of International Workshop on Semantic Evaluation</source>
          , pages
          <fpage>321</fpage>
          -
          <lpage>324</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>S.</given-names>
            <surname>Tahrat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Kergosien</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bringay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Roche</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Teisseire</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Text2geo: from textual data to geospatial information</article-title>
          .
          <source>In Proc. of International Conference on Web Intelligence, Mining and Semantics (WIMS).</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>