<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>An aggregate learning approach for interpretable semi-supervised population prediction and disaggregation using ancillary data?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Guillaume Derval</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Frederic Docquier</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pierre S</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>ICTEAM, UCLouvain</institution>
          ,
          <addr-line>Louvain-la-Neuve</addr-line>
          ,
          <country country="BE">Belgium</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>IRES, UCLouvain</institution>
          ,
          <addr-line>Louvain-la-Neuve</addr-line>
          ,
          <country country="BE">Belgium</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Most countries periodically organize rounds of censuses of their population at a granularity that di ers from country to country. The level of disaggregation is often governed by the administrative division of the country. Although census data are usually considered as accurate in terms of population counts and characteristics, the spatial granularity, that is sometimes in the order of hundreds of square kilometers, is too coarse for evaluating local policy reforms or for making informed decisions about health and well-being of people, economic and environmental interventions, security, etc. For example, ne-grained, high-resolution mappings of the distribution of the population are required to assess the number of people living at or near sea level, near hospitals, in the vicinity or airports and highways, in con ict areas, etc. They are also needed to understand how population movements react to various types of shocks such as natural disasters, con icts, plant creation and closures, etc. Multiple methods can be used to produce gridded data sets (also called rasters), with pixels of a relatively small scale compared to the administrative units of the countries. Gridded Population of the World (GPW) [1] provides a gridded dataset of the whole world by (mostly) redistributing the population in a given census unit uniformly on the census unit surface. More advanced and successful models rely on ancillary and remotely sensed data, and are trained using machine learning techniques. These data can include information sensed by satellite or data provided by NGOs and governments. All methods in the literature are converted into standard supervised regression learning tasks. Supervised regression learning aims to predict an output value associated with a particular input vector. In its standard form, the training set contains an individual output value for each input vector. Unfortunately, the disaggregation problem does not directly t into this supervised regression learning framework since the prediction function is not directly available for input vectors. In a disaggregation problem, the input consists of a partition of the training set (the pixels of each unit) and for each partition, the sum of the output values is constrained by the census count. This framework is exactly the one</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        introduced as the aggregated output learning problem by [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The PCD method
introduced in this paper conceives the formulation of the disaggregation problem
as an aggregated output learning problem. PCD is able to train the model based
on a much larger training set composed of pixels. This approach is parameterized
by the error function that the user seeks to minimize.
      </p>
      <p>
        As a case study, we experiment it on Cambodia using various sets of remotely
sensed/ancillary data. Our main result, the disaggregated map of Cambodia, is
depicted on the left panel of Fig 1 (on the right panel is shown the state of the
art, the RF method [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]). The paper also discusses methodological issues raised by
existing approaches. In particular, we demonstrate that the previously used error
metrics are biased when available census data involves administrative units with
highly heterogeneous surfaces and population densities. We propose alternative
metrics that better re ect the accuracy properties that should be ful lled by a
sound disaggregation approach. We then present the results for Cambodia and
compare methods using various error metrics, providing statistical evidence that
PCD-LinExp generates the most accurate results.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1. Center for International Earth Science Information Network - CIESIN - Columbia University:
          <article-title>Gridded population of the world, version 4 (gpwv4): Population density</article-title>
          , revision
          <volume>10</volume>
          (20180711
          <year>2017</year>
          ), https://doi.org/10.7927/H4DZ068D
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Musicant</surname>
            ,
            <given-names>D.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Christensen</surname>
            ,
            <given-names>J.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Olson</surname>
            ,
            <given-names>J.F.</given-names>
          </string-name>
          :
          <article-title>Supervised learning by training on aggregate outputs</article-title>
          .
          <source>In: Seventh IEEE International Conference on Data Mining (ICDM</source>
          <year>2007</year>
          ). pp.
          <volume>252</volume>
          {
          <fpage>261</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Stevens</surname>
            ,
            <given-names>F.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gaughan</surname>
            ,
            <given-names>A.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Linard</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tatem</surname>
            ,
            <given-names>A.J.:</given-names>
          </string-name>
          <article-title>Disaggregating census data for population mapping using random forests with remotely-sensed and ancillary data</article-title>
          .
          <source>PLOS ONE 10(2)</source>
          ,
          <volume>1</volume>
          {
          <fpage>22</fpage>
          (02
          <year>2015</year>
          ). https://doi.org/10.1371/journal.pone.0107042
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>