<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Species recommendation using intensity models and sampling bias correction (GeoLifeCLEF 2019: Lof of Lof team)</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Monestiez Pascal</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christophe Botella</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>BioSP</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Avignon</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>France</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>UMR AMAP</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Montpellier</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>France</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>This paper presents three algorithms for species spatial recommendation in the context of the GeoLifeCLEF 2019 challenge. We submitted three runs to this task, all based on the estimation of species environmental intensities through Poisson processes models: The rst is directly derived from MAXENT method used for species distribution models. The second method is a modi cation that uses sites were species observed as background points in MAXENT to correct for spatial sampling bias due to heterogeneous sampling in the training occurrences. The last method jointly estimates species and sampling intensities to correct for sampling bias. The best run was the MAXENT method which was ranked 14 over 44 runs with a top30 accuracy of 0.111 on the test set while the worst performing method was LOF with an accuracy of 0.086 (ranked 19).</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Predicting the species most likely to be observed from a location participate to
build better biodiversity identi cation systems by reducing the list of candidate
species that are observable at a given location. This may unlock the
participation of citizen masses in biodiversity monitoring. It also helps experts with the
burden of data quality check. Last but not least, it might serve educational
purposes thanks to biodiversity discovery applications providing innovative features
such as contextualized educational pathways. For this purpose, the GeoLifeCLEF
2019
        <xref ref-type="bibr" rid="ref2 ref3">([Botella et al., 2019b])</xref>
        task aims at evaluating species spatial
recommendation algorithms to stimulate their improvement. It is one of the three tasks of
the LifeCLEF 2019 evaluation campaign
        <xref ref-type="bibr" rid="ref1">([Alexis Joly, 2019])</xref>
        .
      </p>
      <p>
        This challenge is highly related to the problem known as Species Distribution
Modeling (SDM) in ecology [
        <xref ref-type="bibr" rid="ref5">Elith and Leathwick, 2009</xref>
        ]. SDM have become
increasingly important in the last few decades for the study of biodiversity, macro
ecology, community ecology and the ecology of conservation. Concretely, the
goal of SDM is to infer the spatial distribution of a given species, and they are
often based on a set of geo-localized occurrences of that species (collected by
naturalists, eld ecologists, nature observers, citizen sciences project, etc.). No
standard methods for species distribution method have been implemented in last
year edition of GeoLifeCLEF. We implemented three methods based on Poisson
point processes models
        <xref ref-type="bibr" rid="ref4">([Diggle, 2003], [Renner et al., 2015])</xref>
        that estimate the
relative environmental intensity of species from geolocated occurrences. In all
our approaches, we estimate an absolute species intensity over space, and
derive, for any location, the relative probability of any species from normalization
of its intensity by the sum of intensities over all species.
      </p>
      <p>
        The rst run (MAXENT) submitted is a simple loglinear Poisson process
implemented with maxnet R package
        <xref ref-type="bibr" rid="ref7">([Phillips et al., 2017])</xref>
        for 141 species of the test
set with a selection of environmental variables globally suited for modeling plant
species distribution. We then t a modi ed version with Target-Group
background points [
        <xref ref-type="bibr" rid="ref8">Phillips et al., 2009</xref>
        ] (called TGB), a way of correcting for
sampling selection bias. Finally, we used another sampling bias correction method
(LOF). We implemented a joint estimation of all species intensities along with
the spatial sampling e ort over a regular mesh of squares, that is then set
constant for prediction.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Data pre-processing</title>
      <p>The dataset description may be found on the challenge overview
[Botella et al., 2019b].</p>
      <p>Species selection - We rst ltered species occurrences whose identi cation
certainty score ( eld FirstResPLv2Score) was above 0.85 in the PL complete
dataset. Then, we kept only the 300 species with highest number of occurrences
to prevent over- tting (list of species L0). We then made the intersection of those
species with species of GeoLifeCLEF test set, which gave the list of 141 species
(L1) included in our predictions, so our predictions lack 703 test set species. L1
is shown in the table SpeciesTable.csv of the Github repository 3. MAXENT
and TGB runs used only occurrences with identi cation score superior to 0:98,
while LOF run used occurrences with identi cation score superior to 0:85, but
they were tted on the same list of species L1. In all cases, we nally kept only
the occurrences which had valid values for the selected environmental variables
(described below).</p>
      <p>
        Environmental features selection We selected a set of 9 environmental variables
to model the environmental intensity of species included in the model.
Following the recommendations of [
        <xref ref-type="bibr" rid="ref6">Mod et al., 2016</xref>
        ] on environmental variables for
3 https://github.com/ChrisBotella/GLC19runs
modeling macro ecological species niches, we included mean and annual
variation of temperature (chbio 1, chbio 5), annual precipitations (chbio 12),
potential evapo-transpiration (etp), elevation (alti), slope (slope), available
water capacity of the soil (awc top), a soil pH proxy (bs top) and a simpli ed
plant habitat type descriptor (based on clc). Even though, the land cover
category and elevation are not directly linked to species eco-physiological
requirements, they have strong empirical links with species distributions as described
by [
        <xref ref-type="bibr" rid="ref6">Mod et al., 2016</xref>
        ] and have a much sharper spatial grain, with a resolution
around 100 meters. Then, we de ned features derived from those
environmental variables that would constitute the linear predictor of the species intensity.
For continuous environmental variables we chose to model the intensity response
with a Gaussian density function, which means that we kept the original
variable and added a quadratic transformation of it to the linear predictor. This
concerns variables: chbio 1, chbio 5, chbio 12-etp, etp, alti and slope. We
combined annual precipitations chbio 12 and potential evapotranspiration etp
into chbio 12-etp, called the water balance, which is commonly used in plants
SDM [
        <xref ref-type="bibr" rid="ref6">Mod et al., 2016</xref>
        ]. We included categorical pedologic variables
representing physico-chemical properties categories. However, for clc variable, we
aggregated the 48 initial land cover categories into 5 to avoid in ating the number of
parameters for the land cover e ect. Indeed, we de ned a Simpli ed Habitat
Typology (spht) with types: cultivated, forest, grasslands, urban and other.
Each type includes several CORINE Land Cover 2012 categories as shown in
Table 1. We also included an interaction e ect between water balance and slope in
the model. As a summary, equation 1 shows the R formula of the linear predictor
of any species intensity. It yields 18 features parameters.
      </p>
      <p>etp + I(etp2) + I(chbio 12-etp) + I((chbio 12-etp)2)
+chbio 1 + I(chbio 12) + chbio 5 + I(chbio 52) + alti + I(alti2)
+slope + I(slope2) + awc top + I(awc top2) + bs top + I(bs top2)
+spht + slope:I(chbio 12-etp)
(1)
3</p>
    </sec>
    <sec id="sec-3">
      <title>Methodology</title>
      <p>
        The R code for tting models and producing the runs can be found on the
dedicated github repository https://github.com/ChrisBotella/GLC19runs/.
run 27124: MAXENT - We used the R package maxnet
        <xref ref-type="bibr" rid="ref7">([Phillips et al., 2017])</xref>
        to
t independently each species intensity from its occurrences. This package
implements the method MAXENT. We constrained the features to be those of the
previous paragraph Environmental features selection. This method requires to
provide quadrature points, which are meant to represent the distribution of
environmental variables in the spatial domain where the species is observed. We drew
a set Z0 of background points uniformly over the French territory until there was
at least 3 points per 4x4km square cells of a regular grid. With this setting, the
method approximately ts a L1 penalized Poisson Process for each species with
the given loglinear intensity model over environmental features. For each species
of L1, we ran the maxnet function with the occurrences of the species (score
identi cation 0:98) and background points Z0. For more implementation details on
this part, one must refer to the le make maxent and tgb models.R of the github
repository. Once the maxnet model was tted for each species i 2 [1; 141], its
features parameters i 2 Rp were stored. Then, we estimated a posteriori the
intercept i of the linear predictor because maxnet package doesn't provides it. We
can compute it with the following formula i = log(ni= Pz2Ztest exp( iT x(z))),
where ni is the number of training occurrences for i and Ztest is the set of test
occurrences locations. Then we compute the probability of species i at location
z as exp( i + iT x(z))= Pj1=411 exp( j + jT x(z)). This second step is implemented
in the script le make maxent and tgb runs.R of the github repository. We note
that this two steps procedure is equivalent to a one step estimation of features
and intercept parameters.
run 27123: TGB - This run implements the Target-Group Background method
introduced in [
        <xref ref-type="bibr" rid="ref8">Phillips et al., 2009</xref>
        ]. It is the same as the MAXENT run (27124)
except that the background points are selected di erently. We rst de ned
elementary sites where it is assume that the sampling e ort is constant across
sampled sites. We used the cells of the raster grid of the alti variable as sites
because it is the highest resolution variable of our model, so the environment
is roughly constant in those sites. We took a background point for each site
where lies at least an occurrence of L0. We removed points were non-valid
environmental features were found. Once the model were tted, we used the same
procedure as for MAXENT to derive run predictions. The model tting and run
building are implemented in the script les make maxent and tgb models.R and
make maxent and tgb runs.R of the github repository.
run 27063: LOF - We tted a marked point process where the marks are the
species identi ers. For any species i, its occurrences intensity is decomposed as
z ! exp(PQ
      </p>
      <p>
        j=1 j1z2cj ) exp( iT x(z)). The rst factor models the sampling e ort
which is shared between all species, while the second is the species environmental
intensity, i.e. its abundance. Each species intensity has exactly the same
structure has the ones estimated in previous runs. The sampling e ort is constructed
as a step function constant over units of a spatial mesh. The mesh is a regular
spatial grid of squares of 4km side over the French metropolitan territory
including Corsica, and restricted it to squares whose center was closer than 4km from
the border or coast. We kept only the squares that had more than 5 occurrences
inside and excluded the occurrences from the other squares. Thus, we ended
up for this model with a set of 475,138 occurrences from L0, distributed over
15,556 spatial squares covering around 40% of the French territory. This model
and tting procedure is entirely described in details in the submitted article
[Botella et al., 2019a] which preprint is already downloadable 4
        <xref ref-type="bibr" rid="ref2 ref3">(the submitted
4 https://filedn.com/ltHjVlFrSfSJL1S4hXqqTzy/botella_MEE_2019_Main.pdf
run are directly extracted from the real data illustration)</xref>
        and the R
implementation of this method can be found in the script le plantnet effort.R
at https://github.com/ChrisBotella/SamplingEffort. Once we tted this
model, we extracted for each species of L1 its environmental intensity
component, and then built the run predictions as for other runs. This last step is
implemented in the script le make lof run.R of the repository.
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>Results and discussion</title>
      <p>The top30 results of our three runs on the test set are represented in the graph
of Figure, which was pulled from the GeoLifeCLEF 2019 overview working note
[Botella et al., 2019b], with dark yellow bars.</p>
      <p>The best run was the MAXENT method which was ranked 14 over 44 runs
with a top30 accuracy of 0.111 on the test set while the worst performing method
was LOF with an accuracy of 0.086 (ranked 19). TGB made a score of 0.098.
Thus, both bias correction methods failed to improve the model performance
compared to the standard MAXENT approach. Theoretically, it was expectable
that bias correction would not change the performance. Indeed, global sampling
bias in the training data, which similarly multiplies the intensity of all species,
doesn't a ect the relative probabilities across species at a given place. However,
the loss of performance is surprising. In the case of LOF, this performance gap
might be due to the training data that include more occurrences of each species
with less identi cation reliability, or to a model variance problem due to the
high number of sampling e ort parameters (around 15,000). The performance
gap between MAXENT (27124) and TGB (27123) shows that the background
points selection scheme have an impact on model predictive power.</p>
      <p>Despite the small number of features in the MAXENT model and the fact
that we only included 141 of the 844 test species in our models prediction,
MAXENT method had better performances than purely spatial machine
learning algorithms (runs 26988, 27102), arti cial neural networks (runs 26875) and
some CNN methods (runs 27004, 27005) learnt on all species. The few
environmental features selected for the species model based on expert knowledge might
have enabled to capture an important part of species abundance variance while
avoiding to fall in the trap of model over- tting.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Perspectives</title>
      <p>Further work include investigating why bias correction methods failed compared
to MAXENT, and study the generalisation power, at long spatial range, of
MAXENT predictions compared to other more complex models like ANN and CNN,
more likely to over t.
distribution models: implications for background and pseudo-absence data. Ecological
Applications, 19(1):181{197.</p>
      <p>Renner et al., 2015. Renner, I. W., Elith, J., Baddeley, A., Fithian, W., Hastie, T.,
Phillips, S. J., Popovic, G., and Warton, D. I. (2015). Point process models for
presence-only analysis. Methods in Ecology and Evolution, 6(4):366{379.
CLC category description spht category name Raster code
Non-irrigated arable land cultivated 12
Permanently irrigated land cultivated 13
Vineyards cultivated 15
Fruit trees and berry plantations cultivated 16
Complex cultivation patterns cultivated 20
agriculture, with areas of natural vegetation cultivated 21
Agro-forestry areas cultivated 22
Pastures grasslands 18
Natural grasslands grasslands 26
Moors and heathland grasslands 27
Sclerophyllous vegetation grasslands 28
Broad-leaved forest forest 23
Coniferous forest forest 24
Mixed forest forest 25
Transitional woodland-shrub forest 29
Continuous urban fabric urban 1
Discontinuous urban fabric urban 2
Industrial or commercial units urban 3
Road and rail networks and associated land urban 4
Airports urban 6
Green urban areas urban 10
Sport and leisure facilities urban 11
Port areas other 5
Mineral extraction sites other 7
Dump sites other 8
Construction sites other 9
Rice elds other 14
Olive groves other 17
Annual crops associated with permanent crops other 19
Beaches, dunes, sands other 30
Bare rocks other 31
Sparsely vegetated areas other 32
Burnt areas other 33
Glaciers and perpetual snow other 34
Inland marshes other 35
Peat bogs other 36
Salt marshes other 37
Salines other 38
Intertidal ats other 39
Water courses other 40
Water bodies other 41
Coastal lagoons other 42
Estuaries other 43
Sea and ocean other 44
No data other 48
Unclassi ed land surface other 49
Unclassi ed water bodies other 50
Table 1. spht (Aggregated land cover) categories correspondance with Corine Land
Cover 2012.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Alexis</given-names>
            <surname>Joly</surname>
          </string-name>
          ,
          <year>2019</year>
          . Alexis Joly, Herve Goeau,
          <string-name>
            <surname>C. B. S. K. M. S. H. G. P. B. W.-P. V. R. P. F.-R. S. H. M.</surname>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>Overview of lifeclef 2019: a new snapshot of the performance of species identi cation and prediction algorithms</article-title>
          .
          <source>In Proceedings of CLEF</source>
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Botella</surname>
          </string-name>
          et al.,
          <year>2019a</year>
          . Botella,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Joly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Bonnet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Munoz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            , and
            <surname>Monestiez</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.</surname>
          </string-name>
          (
          <year>2019a</year>
          ).
          <article-title>[under review] jointly estimating ecological niches and spatial sampling e ort from multiple species occurrences</article-title>
          .
          <source>In Methods in Ecology and Evolution.</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Botella</surname>
          </string-name>
          et al.,
          <year>2019b</year>
          . Botella,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Servajean</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Bonnet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            , and
            <surname>Joly</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          (
          <year>2019b</year>
          ).
          <article-title>Overview of geolifeclef 2019: plant species prediction using environment and animal occurrences</article-title>
          .
          <source>In CLEF working notes</source>
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Diggle</surname>
          </string-name>
          ,
          <year>2003</year>
          . Diggle,
          <string-name>
            <surname>P.</surname>
          </string-name>
          (
          <year>2003</year>
          ).
          <article-title>Statistical analysis of spatial point patterns</article-title>
          : Oxford university press. New York.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Elith</surname>
          </string-name>
          and Leathwick,
          <year>2009</year>
          . Elith,
          <string-name>
            <given-names>J.</given-names>
            and
            <surname>Leathwick</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. R.</surname>
          </string-name>
          (
          <year>2009</year>
          ).
          <article-title>Species Distribution Models: Ecological Explanation and Prediction Across Space and Time</article-title>
          . Annual Review of Ecology, Evolution, and Systematics,
          <volume>40</volume>
          :
          <fpage>677</fpage>
          {
          <fpage>697</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Mod</surname>
          </string-name>
          et al.,
          <year>2016</year>
          . Mod,
          <string-name>
            <given-names>H. K.</given-names>
            ,
            <surname>Scherrer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Luoto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            , and
            <surname>Guisan</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          (
          <year>2016</year>
          ).
          <article-title>What we use is not what we know: environmental predictors in plant distribution models</article-title>
          .
          <source>Journal of Vegetation Science</source>
          ,
          <volume>27</volume>
          (
          <issue>6</issue>
          ):
          <volume>1308</volume>
          {
          <fpage>1322</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Phillips</surname>
          </string-name>
          et al.,
          <year>2017</year>
          . Phillips,
          <string-name>
            <given-names>S. J.</given-names>
            ,
            <surname>Anderson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. P.</given-names>
            ,
            <surname>Dud</surname>
          </string-name>
          <string-name>
            <given-names>k</given-names>
            , M.,
            <surname>Schapire</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. E.</given-names>
            , and
            <surname>Blair</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. E.</surname>
          </string-name>
          (
          <year>2017</year>
          ).
          <article-title>Opening the black box: an open-source release of maxent</article-title>
          . Ecography.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Phillips</surname>
          </string-name>
          et al.,
          <year>2009</year>
          . Phillips,
          <string-name>
            <surname>S. J.</surname>
          </string-name>
          ,
          <source>Dud</source>
          k, M.,
          <string-name>
            <surname>Elith</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Graham</surname>
            ,
            <given-names>C. H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lehmann</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leathwick</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Ferrier</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          (
          <year>2009</year>
          ).
          <article-title>Sample selection bias and presence-only</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>