<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Predicting child labor in Peru: A comparison of logistic regression and neural networks techniques</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Christian Fernando Libaque-Saenz</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Juan Lazo</string-name>
          <email>jg.lazol@up.edu.pe</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Karla Gabriela Lopez-Yucra</string-name>
          <email>karla.lopez@pucp.pe</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Edgardo R. Bravo</string-name>
          <email>er.bravoo@up.edu.pe</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Pontificia Universidad Cat o ́lica del Peru ́ Av.</institution>
          <addr-line>Universitaria 1801, San Miguel, Lima 32</addr-line>
          ,
          <country country="PE">Peru</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Universidad del Pac ́ıfico Avenida Salaverry 2020, Jesu ́ s Mar ́ıa</institution>
          ,
          <addr-line>Lima 11</addr-line>
          ,
          <country country="PE">Peru</country>
        </aff>
      </contrib-group>
      <fpage>69</fpage>
      <lpage>79</lpage>
      <abstract>
        <p>Child labor is a relevant problem in developing countries because it may have a negative impact on economic growth. Policy makers and government agencies need information to correctly allocate their scarce resources to deal with this problem. Although there is research attempting to predict the causes of child labor, previous studies have used only linear statistical models. Non-linear models may improve predictive capacity and thus optimize resource allocation. However, the use of these techniques in this field remains unexplored. Using data from Peru, our study compares the prediction capability of the traditional logit model with artificial neural networks. Our results show that neural networks could provide better predictions than the logit model. Findings suggest that geographical indicators, income levels, gender, family composition and educational levels significantly predict child labor. Moreover, the neural network suggests the relevance of each factor which could be useful to prioritize strategies. As a whole, the neural network could help government agencies to tailor their strategies and allocate resources more efficiently.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Child labor is a critical problem in developing
countries because it could negatively affect
economic growth
        <xref ref-type="bibr" rid="ref12">(Hanushek, 2013)</xref>
        . Child labor has
a negative effect on human capital, which is
defined as the stock of skills that the labor force
possesses
        <xref ref-type="bibr" rid="ref9">(Goldin, 2016)</xref>
        . Children who work have
a high probability of becoming individuals with
a low stock of skills in both quantity and
quality
        <xref ref-type="bibr" rid="ref4">(Becker, 1962)</xref>
        . In fact, these children (who
work) usually do not dedicate their efforts to study
and sometimes they do not even attend school at
all. In turn, this low level of human capital and
the associated lack of skills have a negative impact
on individuals earnings and income
        <xref ref-type="bibr" rid="ref12">(Hanushek,
2013)</xref>
        . Therefore, as a countrys human capital
decreases, its economy decreases as well.
      </p>
      <p>
        According to the International Labor
Organization (ILO), in Latin America this phenomenon
reached 12.5 million children and teenagers
between 5 and 17 years old in 2014
        <xref ref-type="bibr" rid="ref20">(Lopez, 2016)</xref>
        .
Although this number has decreased from 20
million in 2010, an important fact is that the
number of children working in dangerous activities has
increased from 9 million in 2010 to 9.6 million
in 2014
        <xref ref-type="bibr" rid="ref20">(Lopez, 2016)</xref>
        . As for the case of Peru,
the National Housing Survey (ENAHO in
Spanish) shows that 21% of teenagers between 12 and
17 years old had been working in 2014
        <xref ref-type="bibr" rid="ref20">(Lopez,
2016)</xref>
        . In other words, 1 out of 5 teenagers works
in Peru.
      </p>
      <p>
        Child labor can not only lead to gaps among
countries but also within a country. In Peru, for
example, the child labor rate in rural areas is twice
as high as in urban areas
        <xref ref-type="bibr" rid="ref28">(Sausa, 2016)</xref>
        . By
assessing child labor by region, Huancavelica presents
the highest rate of child labor (58%), which is
more than 10 times that for Tumbes (5%) the
latter is the region with the lowest rate of child labor
        <xref ref-type="bibr" rid="ref28">(Sausa, 2016)</xref>
        . Therefore, this phenomenon could
negatively impact social and economic inclusion
by increasing socioeconomic differences. It is
important that governments formulate adequate
programs and policies to reduce child labor. It is also
important that they identify those children with a
high probability of becoming workers in order to
allocate resources in the correct place. There are
various techniques to achieve this goal. We have
traditional techniques such as logit models, and
modern techniques such as neural networks. The
principal difference is that the former capture
linear effects, while the latter can capture non-linear
relationships. It is important to have a model with
high predictive capability, and therefore it is
necessary to compare the predictive power of the
different models.
      </p>
      <p>Table 1 shows a summary of the issues covered
by previous research in this field. All these
studies used traditional techniques; to the best of our
knowledge, in this field there are few studies
using modern techniques such as neural networks.
For example, Rodrigues, Prata, and Silva (2015)
used data from Brazil and decision trees to search
for patterns in the variables explaining child
labor. The objective of the present study is to
compare the predictive power of traditional and
modern models in regard to child labor (i.e., correctly
identify those children who work). It is expected
that our results will shed light on the difference
between models in terms of predictive power. By
identifying the antecedents to child labor and the
technique with the best predictive power, we will
be able to provide recommendations to the
Peruvian government.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Theoretical background</title>
      <p>Classification problems such as the child labor
issue can be addressed by several techniques,
both parametric and non-parametric. Parametric
techniques (e.g., discriminant analysis, the logit
model) require the prior specification of a function
(or model) that relates the independent variables
(Xi) with the dependent variable (Y ). In practical
terms, this function may be known grounded in
theory or assumed. These techniques use
observations of Y and Xi to estimate the parameters of
the function. Once the parameters have been
estimated, they can be used for prediction with new
participants. One disadvantage of the
parametric techniques is that they have a rigid structure
(the mathematical function does not change and it
only allows for estimating the parameters). Thus,
these techniques may not be appropriate to
represent phenomena that do not follow well-known
mathematical functions.</p>
      <p>
        In contrast, non-parametric techniques (e.g.,
artificial neural networks) do not assume a function a
priori but instead approximate the function based
on observation. Once the function has been
approximated, it can be used to predict new cases.
One relative advantage of these techniques is that
they can represent complex non-linear
mathematical functions. In other research arenas, this
flexibility of non-parametric techniques has, under
certain conditions, demonstrated the superiority of
its predictive power over that of parametric
techniques
        <xref ref-type="bibr" rid="ref1 ref3">(e.g., Abdou et al., 2008; Altman et al.,
1994)</xref>
        .
      </p>
      <p>Our research compares the logit model
(parametric technique) with artificial neural networks
(non-parametric technique) in the field of child
labor. The application of these models for predictive
purposes involves the following steps:
• The sample is randomly divided into two
subsamples.
• The parameters of the model are estimated
with one of the subsamples.
• The predictive capacity of the model (number
of hits over total observations) is assessed.
• With these estimated parameters, prediction
of the dependent variable for the other
subsample is conducted.
• The predictive capacity of the model (with
the test data) is assessed.
2.1</p>
      <sec id="sec-2-1">
        <title>Logit model</title>
        <p>
          The logit model is a method that uses
independent variables to estimate the probability of
occurrence of a discrete outcome in the dependent
variable
          <xref ref-type="bibr" rid="ref16">(Lattin et al., 2003)</xref>
          . According to the
number of discrete outcomes, this technique can
be divided into binary logit or multinomial logit
models
          <xref ref-type="bibr" rid="ref15 ref16">(Hosmer et al., 2013; Lattin et al., 2003)</xref>
          .
The former defines a dependent variable with two
discrete outcomes whereas the latter represents a
logit model with more than two discrete outcomes
for the dependent variable
          <xref ref-type="bibr" rid="ref15 ref16">(Hosmer et al., 2013;
Lattin et al., 2003)</xref>
          . In both cases, the discrete
outcomes for the dependent variable should be
mutually exclusive
          <xref ref-type="bibr" rid="ref16">(Lattin et al., 2003)</xref>
          .
        </p>
        <p>The logit model has a straightforward and
closed functional form that is easily estimated
using maximum likelihood methods (Lattin et al.,</p>
        <sec id="sec-2-1-1">
          <title>Author</title>
          <p>Rodr´ıguez (2002)
Emerson and Souza (2002)
Sapelli and Torche (2004)
Lavado and Gallegos (2005)
Garc´ıa (2006)
Gunnarsson, Orazem, and Snchez (2006)
Alca´zar (2008)
Rodr´ıguez and Vargas (2008)
Rodr´ıguez and Vargas (2009)</p>
        </sec>
        <sec id="sec-2-1-2">
          <title>Lima, Mesquita ,and Wanamaker (2015) Le and Homel (2015) He (2016) Table 1: Literature review</title>
          <p>
            2003, p. 475). The logit technique does not
assume restrictions on the normality of the
distribution of variables
            <xref ref-type="bibr" rid="ref21">(Press and Wilson, 1978)</xref>
            . Also,
independent variables can be both continuous and
categorical variables
            <xref ref-type="bibr" rid="ref16">(Lattin et al., 2003)</xref>
            . This
technique is a special case of regression, which
uses a transformation of the discrete dependent
variable. This model assumes: 1) a categorical
dependent variable with mutually exclusive
outcomes, 2) independent variables can be continuous
or categorical, 3) independence of observations, 4)
absence of multicollinearity between independent
variables, 5) a linear relationship between the
continuous independent variables and the logit
transformation of the dependent variable, and 6)
absence of outliers.
          </p>
          <p>The logit model is defined by the following
function:</p>
          <p>pi
1
pi
!
Logit(pi) = Ln
= ↵ + XiT
+ "i (1)
where pi is the probability that an observation
takes a specific outcome of the dependent variable,
↵ is the constant term; is the corresponding
vector of the coefficients; and "i is the error term.
2.2</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>Artificial neural networks</title>
        <p>
          A neural network is, in a general sense, a
machine designed to model the way in which the
brain performs a particular task or function of
interest
          <xref ref-type="bibr" rid="ref13">(Haykin, 1998)</xref>
          . The functioning of the brain
is applied in this design because of its “(. . . )
capability to organize its structural constituents, known
as neurons, so as to perform certain computations
(e.g., pattern recognition, perception, and motor
control) many times faster than the fastest
digital computer in existence today”
          <xref ref-type="bibr" rid="ref13">(Haykin, 1998,
p. 23)</xref>
          . Therefore, a neural network
resembles the brain mainly in two aspects: 1) the way
knowledge is acquired by the network from its
environment (i.e., learning process); and 2) the
strength of interneuron connections (i.e.,
synaptic weights), which are used to store the acquired
knowledge
          <xref ref-type="bibr" rid="ref13">(Haykin, 1998)</xref>
          . Accordingly, an
artificial neural network is a physical cellular network
that is able to acquire, store, and utilize
experiential knowledge
          <xref ref-type="bibr" rid="ref30">(Zurada, 1992)</xref>
          . A fundamental
unit in the operation of a neural network is the
neuron. It is an information-processing unit which has
three basic elements: a set of synapses or
connecting links, each one with a weight or strength of its
own; an adder for summing the input signals; and
an activation function for limiting the amplitude of
the output of a neuron
          <xref ref-type="bibr" rid="ref13">(Haykin, 1998, p. 32)</xref>
          . The
neurons perform simple operations, transmitting
their results to neighboring processors. Hence, the
ability of a neural network to perform non-linear
relationships between its inputs and outputs makes
it a useful technique for pattern recognition and
modeling of complex systems
          <xref ref-type="bibr" rid="ref5">(Bishop, 1995)</xref>
          .
        </p>
        <p>
          According to their topology, neural networks
can be feedforward or feedback networks. In the
former, the mapping goes from an input to an
output layer instantaneously since there is no delay
between them. This type of network is
characterized by its lack of feedback which implies that the
neural network has no explicit connection between
layers
          <xref ref-type="bibr" rid="ref30">(Zurada, 1992)</xref>
          . In contrast, the latter has
a connection between the output and input layers
          <xref ref-type="bibr" rid="ref30">(Zurada, 1992)</xref>
          .
        </p>
        <p>
          Another typology of neural networks is
related to the learning paradigm which distinguishes
between supervised learning and non-supervised
learning. The first implies that the knowledge of
the environment available to the teacher is
transferred to the neural network through training as
fully as possible
          <xref ref-type="bibr" rid="ref13">(Haykin, 1998)</xref>
          . Also, it implies
an error-correction learning in which the network
parameters are adjusted under the combined
influence of the training vector (i.e., example) and
the error signal (i.e., difference between the
desired response and the actual response of the
network). This adjustment is carried out step by step
in order to make the neural network emulate the
teacher
          <xref ref-type="bibr" rid="ref13">(Haykin, 1998)</xref>
          . On the other hand, the
second does not consider a teacher to oversee the
learning process. In this case, there are no labeled
examples of the function to be learned by the
network. The learning of an input-output mapping is
performed through continued interaction with the
environment or based on the optimization of its
parameters in order to develop the ability to form
internal representations
          <xref ref-type="bibr" rid="ref13">(Haykin, 1998)</xref>
          .
        </p>
        <p>
          This research uses a Multilayer Perceptron
neural network with a back-propagation algorithm
which consists of applying a family of
gradientbased optimization methods to find the optimal
value of the weights based on minimizing the error
norm between the desired output and the output
calculated by the neural network
          <xref ref-type="bibr" rid="ref26">(Rumelhart et al.,
1986)</xref>
          . In this type of network, the processing is
performed by the inputs. The output obtained is
compared to the expected output. From the
obtained error, a process of adjustment of weights is
applied, attempting to minimize the error.
2.3
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>Child labor in Peru</title>
        <p>The concept of child labor varies from country to
country depending on the cultural context.
According to the ILO, child labor refers to a work that
is dangerous and harmful to the physical,
mental, or moral wellness of the child, interfering with
his/her education.</p>
        <p>
          In the case of Peru, the minimum age for a child
to be allowed to legally work is 14 years old, as
long as these activities do not harm their integrity
nor negatively impact their studies
          <xref ref-type="bibr" rid="ref20">(Lopez, 2016)</xref>
          .
Also, they must have the permission of their
parents or legal guardians to engage in these
activities. In exceptional cases, children between 12 and
14 years old could also work as long as the work
meets the same requirements
          <xref ref-type="bibr" rid="ref20">(Lopez, 2016)</xref>
          . In
the present research, a child was considered to be
a worker if he/she helps in the family business, in
domestic tasks in a house that is not his or her own,
in producing products to be sold, in agriculture
activities, in selling products or providing services.
        </p>
        <p>According to the National Housing Survey,
child labor between 6 and 13 years old in rural
areas (67.5%) is twice as prevalent as child labor
in urban areas (32.5%). However, in the range
from 14 to 17 years old, the values are similar
(49.7% and 50.3% for rural and urban areas,
respectively). Another important issue is that child
labor rates significantly differ between cities. For
example, Huancavelica is the city with the highest
rate of child labor with 79.0%, followed by Puno,
Huanuco, and Amazonas with 69.0%, 65.0%, and
64.0% respectively. Trujillo has the lowest child
labor rate, at about 5.0%, which is significantly
lower than the others. Not surprisingly, the cities
with the highest rates of child labor are also those
with the lowest incomes per capita. Furthermore,
according to the National Institute of Statistics and
Informatics (INEI in Spanish), economic activity
for females (63.3%) is considerable lower than for
males (81.4%).</p>
        <p>
          Based on the above paragraph, we included
variables capturing: 1) age and gender; 2) type of
residence area such as urban/rural, region, stratum,
and schooling available; and 3) socioeconomic
variables such as expenses, education of the
family head, type of housing, housing ownership, and
housing status (adequacy, coverage of basic needs,
sanitation). In addition, following
          <xref ref-type="bibr" rid="ref20">(Lopez, 2016)</xref>
          ,
we included family characteristics as potential
antecedents to child labor. Indeed, families where
both parents work are less likely to have their
children working, while the number of children could
increase the probability that one or more children
work. In these cases, the oldest child is the one
with the highest probability of engaging in
economic activities. Finally, current schooling status
could also be a potential factor for child labor
because those children who are behind in their
studies are potentially engaged in other activities.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Research method</title>
      <sec id="sec-3-1">
        <title>Measurement model</title>
        <p>Data were collected from the Peruvian National
Housing Survey (ENAHO) for the year 2014. We
eliminated the data for the months of January,
February and March to eliminate seasonality. The
rationale is that those months are holidays in
Peruvian schools and thus the probability of child labor
is high but does not imply that children stop
studying to carry it out. Data include children between
12 and 17 years old at the national level who meet
the following criteria: 1) is the son/daughter of the
head of the family, and 2) he/she has not yet
finished school.</p>
        <p>For analysis, we used logit and neural networks
techniques to find the antecedents to child labor
and to classify children according to the
probability of becoming a worker. We used these two
techniques to compare predictive power because a
correct prediction may allow governments to
correctly allocate resources to deal with this
problem. The first technique is based on linear
relationships, while the latter can manage non-linear
effects. Thus, differences in their results are
expected. In the case of the logit model, we
randomly divided the full sample into 2 subsamples:
1) a training subsample consisting of 85% of the
full sample, and 2) a test subsample made up of
the remaining 15%. We used the training
subsample to calibrate the model (i.e., estimate the
parameters of the function), and the test subsample
to assess the predictive power of these results. In
the case of neural networks, we randomly divided
the sample into 3 subsamples: 1) a training
subsample (70% of the total data), 2) a validation
subsample (15% of the total data), and 3) a test
subsample (15% of the total data). We used the
training and validation subsamples together to estimate
the parameters of the model. To avoid overfitting
and guarantee that the results of this stage could
be generalized, we validated the predictive
quality of the model with only the validation
subsample every 1000 interactions. This process allows
a better estimation of the weights of the network.
Finally, we assessed the predictive power of the
model with the test subsample.
4
4.1</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <sec id="sec-4-1">
        <title>Logit results</title>
        <p>We conducted a preliminary analysis including
all 17 independent variables. Results show that
only 9 variables were statistically significant
(variables with coefficients with p-value less than 0.05)
in explaining the variance of our dependent
variable (WORK). The other 8 variables (p-values
higher than 0.05) were not considered in the
subsequent analysis given that they do not have any
impact on the dependent variable. Retained
variables are divided into 6 categorical variables:
URBAN, AREA, STRATUM, OWN, ADEQ, and
UNMET; and 3 continuous variables: EXPENSE,
EDU HEAD, and SIBLINGS. We calculated the
coefficients of the model using equation (1), where
pi is the probability that child i becomes a worker.</p>
        <p>
          We assessed whether assumptions of logistic
regression were met. Assumptions 1, 2, and 3 were
determined by the model and data collection. For
assumption 4, we conducted a linear regression
to obtain VIF values. All VIF values were lower
than 5 (the independent variable URBAN has the
highest VIF value at 2.274). Therefore, there is
no evidence of multicollinearity problems in our
model
          <xref ref-type="bibr" rid="ref11">(Hair et al., 2011)</xref>
          . For the fifth
assumption, we used the Box and Tidwell (1962)
procedure. This procedure establishes that if the
interaction between an independent continuous
variable and its natural logarithm transformation is
found to be significant, this variable is not
linearly related to the logit of the dependent
variable. In addition, following Tabachnick and
Fidells (2007) recommendation, we used a
Bonferroni correction for the statistical significance
level by dividing it by the number of independent
variables running this test including the constant
term. This correction provided a significance level
of 0.0038 (i.e., 0.05/13, where 0.05 is the
original significance level and 13 is the sum of
variables including the constant term: 1 constant term,
6 categorical independent variables, 3 continuous
independent variables, and 3 interaction terms).
P-values for the interaction terms were 0.688 for
EDU HEAD, 0.999 for SIBLINGS, and 0.0041
for EXPENSE. Based on this assessment, all
p
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>Variable</title>
        <sec id="sec-4-2-1">
          <title>Worker (WORK)</title>
          <p>Age (AGE)
Education of the family head
(EDU HEAD)
Younger siblings (SIBLINGS)</p>
        </sec>
        <sec id="sec-4-2-2">
          <title>Family composition (COMPO)</title>
        </sec>
        <sec id="sec-4-2-3">
          <title>Education centers (CENTER)</title>
        </sec>
        <sec id="sec-4-2-4">
          <title>Monthly expense (EXPENSE)</title>
        </sec>
        <sec id="sec-4-2-5">
          <title>Maleness (MALE)</title>
        </sec>
        <sec id="sec-4-2-6">
          <title>Urban (URBAN)</title>
        </sec>
        <sec id="sec-4-2-7">
          <title>Oldest child (OLD CHI)</title>
        </sec>
        <sec id="sec-4-2-8">
          <title>School backwardness (DELAY)</title>
        </sec>
        <sec id="sec-4-2-9">
          <title>Geographic area (AREA)</title>
        </sec>
        <sec id="sec-4-2-10">
          <title>Geographic stratum (STRATUM)</title>
        </sec>
        <sec id="sec-4-2-11">
          <title>Type of housing (TYPE)</title>
          <p>Level of schooling of the head of the family (in years)
Number of children under 5 years old in the family
Ratio of the number of adults (18 years old or older) to the number
of children (younger than 18 years old) in the family
Ratio of the number of education centers to the number of
schoolage children in the province of residence of the family
Natural logarithm of the total monthly expense per family
member
Categorical Independent Variables
1 = If the child is male
0 = If the child is female
1 = If the residence of the family is located in the urban area
0 = If the residence of the family is located in a non-urban area
1 = If the child is the oldest in the family
0 = If the child is not the oldest in the family
1 = If the child presents school backwardness
0 = If the child does not present school backwardness
1 = North Coast
2 = Center Coast
3 = South Coast
4 = North Highlands
5 = Center Highlands
6 = South Highlands
7 = Jungle
8 = Lima Metropolitan Area
1 = More than 100,000 dwellings
2 = From 20,001 to 100,000 dwellings
3 = From 10,001 to 20,000 dwellings
4 = From 4,001 to 10,000 dwellings
5 = From 401 to 4,000 dwellings
6 = 400 dwellings or fewer
7 = Composite rural area
8 = Simple rural area
1 = Independent house
2 = Apartment in building
3 = Chalet
4 = Neighborhood house
5 = Shack or cottage
6 = Improvised housing
7 = Non-housing premises
8 = Other</p>
        </sec>
        <sec id="sec-4-2-12">
          <title>Housing ownership (OWN)</title>
        </sec>
        <sec id="sec-4-2-13">
          <title>Housing inadequacy (ADEQ)</title>
        </sec>
        <sec id="sec-4-2-14">
          <title>Uncovered basic needs (UNMET) Absence of sanitation (HYGIENIC)</title>
          <p>1 = Rented
2 = Owned by the family, totally paid
3 = Owned by the family, as result of squatting
4 = Owned by the family, paying off a loan
5 = Given by the workplace of one of the members
6 = Given by other family or institution
7 = Other
1 = If the housing is inadequate
0 = If the housing is adequate
1 = If the housing has unmet basic needs
0 = If the housing has not unmet basic needs
1 = If the house does not have sanitation
0 = If the housing has sanitation
values were over the value of 0.0038 and thus our
model satisfied the linearity assumption. For the
sixth assumption, we found 4 outliers of concern
which were not considered in subsequent
analysis. Results of the logistic model are presented
in Table 3. Our model is statistically significant
( 2 = 2300.885, df = 25, p = 0.000), and
explains between 33.2% and 47.8% of the variance
in child labor.</p>
          <p>In terms of predictive value, our model correctly
predicted 82.61% of cases, with 55.80% of correct
positive classifications (sensitivity) and 93.02% of
correct negative classifications (specificity).
Accordingly, our model has an efficiency (average of
sensitivity and specificity) of 74.41% and a mean
absolute percentage error (MAPE) of 17.39%.
Although our model has an adequate overall
predictive power, the Hosmer and Lemeshow goodness
of fit test was significant ( 2 = 39.889, df =
8, p = 0.000) showing that it is poor at predicting
the categorical outcomes. The reason for this
finding may be the difference between sensitivity and
specificity. Finally, coefficients (B) were found
to be significant based on the Wald test. Table 3
also shows the standard error (SE) of the
coefficients and their odd ratio (OR). We then assessed
the model with our test subsample. Our model
correctly predicted 80.64% of all the cases in this
sample, with a sensitivity of 52.78%, a specificity
of 91.88%, an efficiency of 72.33%, and a MAPE
of 19.36%.
4.2</p>
        </sec>
      </sec>
      <sec id="sec-4-3">
        <title>Neural network results</title>
        <p>
          For purposes of comparison, we chose a simple
neural network. Accordingly, we used a hidden
layer with activation functions Hyperbolic
tangentsigmoidy, and an output layer with activation
functions Log-sigmoid. The value of weights and
bias are updated according to gradient descent
momentum and an adaptive learning rate. The
training parameters of the neural network were:
Maximum number of epochs to train: 40000, learning
rate: 0.01, momentum constant: 0.7, performance
goal: 10-5. These values were set following
current literature
          <xref ref-type="bibr" rid="ref13 ref30">(Haykin, 1998; Zurada, 1992)</xref>
          . They
were also adjusted during the training process
using an adaptive algorithm to find better
parameters.
        </p>
        <p>The first neural network used the 17 proposed
independent variables (inputs). With the training
and validation subsamples we obtained the best
neural network made up of 38 neurons in the
hidden layer and 1 neuron in the output layer. This
model predicted 88.26% of all the cases, with a
sensitivity of 90.97% and specificity of 87.21%.
The efficiency of the model was thus 89.09%,
and the MAPE was 11.74%. When applying this
model to the test subsample, it predicted 85.11%
of all the cases, with a sensitivity of 90.02%, a
specificity of 79.42%, an efficiency of 84.72%,
and a MAPE of 14.89%.</p>
        <p>In addition, by analyzing the weight of the
inputs of the neural network, we ranked the
independent variables from the highest to the lowest
effect: AREA (7.7), EXPENSE (7.4), HYGIENIC
(7.4), STRATUM (7.0), MALE (6.2), OWN (6.0),
SIBLINGS (5.9), TYPE (5.9), OLD CHI (5.9),
AGE (5.7), EDU HEAD (5.5), COMPO (5.5),
ADEQ (5.1), DELAY (5.0), CENTER (4.9),
URBAN (4.6), and UNMET (4.4).</p>
        <p>The second neural network used only the 9
variables that were statistically significant in the logit</p>
      </sec>
      <sec id="sec-4-4">
        <title>Variables</title>
        <p>*p &lt; 0.05, **p &lt; 0.01, ***p &lt; 0.001
B=Coefficients; SE=Standard error; OR=Odds ratio
SS=Skipped for simplicity. (For categorical variables with more than 2 categories, there is a coefficient for each category.
We are choosing not to report them all because our focus is the predictive power of the model.)
model for a straight comparison. In this model,
with the training and validation subsamples we
obtained the best neural network made up of 30
neurons in the hidden layer and 1 neuron in the output
layer. Our model achieved 84.45% of correct total
predictions, with a sensitivity of 79.61%, a
specificity of 86.34%, an efficiency of 82.97%, and a
MAPE of 15.55%. When using our models
parameters on the test subsample, it predicted 81.69%
of all the cases, with a sensitivity of 78.86%,
specificity of 84.23%, efficiency of 81.55%, and
a MAPE of 18.31%.</p>
        <p>For this model, the ranking of the inputs
according to their weights is: AREA (12.9), STRATUM
(12.6), EDU HEAD (12.4), SIBLINGS (12.3),
URBAN (10.9), ADEQ (10.6), UNMET (9.7),
OWN (9.5), and EXPENSE (9.1).
4.3</p>
      </sec>
      <sec id="sec-4-5">
        <title>Technique comparison</title>
        <p>The results of the previous section are summarized
in Table 4. Considering that the logit model used
9 variables (8 were not considered because they
have no significant impact on the dependent
variable), the neural network used these same 9
variables and the same instances to ensure a fair
comparison. In addition, Table 4 shows the results
of the neural network technique with the
complete 17 variables to assess if this non-linear model
could extract important information from those 8
variables without a linear impact on the
dependent variable. This table shows that overall neural
network technique performed better than the logit
model. In fact, the neural network obtained the
highest values of accuracy (correct total - positive
and negative - predictions). Also, considering that
it is more important to predict when a child has
high probabilities of becoming a worker than to
predict that a child will be non-worker, sensitivity
stands as our most important metric when
comparing models. By an inspection of Table 4,
sensitivity of the neural network technique was superior to
the values obtained from the logit model. In spite
of these results, the logit model was superior in
terms of specificity. However, specificity is a
metric for correct predictions of non-workers, which
is not relevant in our case. In addition, Figure 1
shows the ROC curve of prediction for these
techniques.
Overall, the results show that the neural network
technique surpasses the logit model in predictive
capacity of child labor (sensitivity). Indeed, this
phenomenon may have a more complex structure
than is assumed by the logit model. In
consequence, the neural network (which adopts
nonlinear relationships) could capture sources of
variation that are not identified by the logit
technique. An accurate prediction of this phenomenon
could be used by policy makers and government
agencies to design adequate strategies or to invest
scarce resources efficiently to deal with this
problem.</p>
        <p>Also, our findings show that the neural network
model with 17 variables performed better than the
9-variables models (logit or neural network). This
result suggests that this additional set of variables
capture an important variability in explaining child
labor. In other words, the neural network model
with 17 variables does not ignore information that
is relevant to the prediction. This result could be
used by decision makers to avoid discarding
relevant factors when dealing with this phenomenon.</p>
        <p>Another important result is that the neural
network model shows that geographical indicators,
income levels, gender, family composition and
educational levels significantly predict child labor.
These results are aligned with those of the logit
model showing that stratum, geographic area, and
housing conditions have a significant impact on
our dependent variable. These results can be used
to determine the relevance of each factor. In turn,
this relevance-based ranking of factors could
further help government agencies to better allocate
their resources and implement their strategies to
reduce child labor.</p>
        <p>Finally, previous studies in this field have used
linear statistical models to predict child labor. Our
study shows that the use of computational
intelligence techniques, such as the neural network,
could provide better predictions, which leads to
better decision making.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Hussein</given-names>
            <surname>Abdou</surname>
          </string-name>
          , John Pointon, and
          <string-name>
            <surname>Ahmed</surname>
          </string-name>
          El-Masry.
          <year>2008</year>
          .
          <article-title>Neural nets versus conventional techniques in credit scoring in Egyptian banking</article-title>
          .
          <source>Expert Systems with Applications</source>
          <volume>35</volume>
          (
          <issue>3</issue>
          ):
          <fpage>1275</fpage>
          -
          <lpage>1292</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Lorena</given-names>
            <surname>Alca</surname>
          </string-name>
          ´zar.
          <year>2008</year>
          .
          <article-title>Asistencia y desercio´n en escuelas secundarias rurales del Peru´</article-title>
          , pages
          <fpage>41</fpage>
          -
          <lpage>82</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Edward</surname>
            <given-names>I Altman</given-names>
          </string-name>
          , Giancarlo Marco, and
          <string-name>
            <given-names>Franco</given-names>
            <surname>Varetto</surname>
          </string-name>
          .
          <year>1994</year>
          .
          <article-title>Corporate distress diagnosis: Comparisons using linear discriminant analysis and neural networks (the Italian experience)</article-title>
          .
          <source>Journal of Banking &amp; Finance</source>
          <volume>18</volume>
          (
          <issue>3</issue>
          ):
          <fpage>505</fpage>
          -
          <lpage>529</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Gary</given-names>
            <surname>Becker</surname>
          </string-name>
          .
          <year>1962</year>
          .
          <article-title>Investment in Human Capital: A Theoretical Analysis</article-title>
          .
          <source>Journal of Political Economy</source>
          <volume>70</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Christopher M Bishop</surname>
          </string-name>
          .
          <year>1995</year>
          .
          <article-title>Neural Networks for Pattern Recognition</article-title>
          . Oxford University Press, Inc., New York, NY, USA.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>George EP Box</surname>
          </string-name>
          and Paul W Tidwell.
          <year>1962</year>
          .
          <article-title>Transformation of the Independent Variables</article-title>
          .
          <source>Technometrics</source>
          <volume>4</volume>
          (
          <issue>4</issue>
          ):
          <fpage>531</fpage>
          -
          <lpage>550</lpage>
          . https://doi.org/10.1080/00401706.
          <year>1962</year>
          .
          <volume>10490038</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Patrick M. Emerson</surname>
          </string-name>
          and Andre´ Portela Souza.
          <year>2002</year>
          .
          <article-title>Bargaining over sons and daughters: Child labor, school attendance and intra-household gender bias in brazil.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Luis</given-names>
            <surname>Garc</surname>
          </string-name>
          ´ıa.
          <year>2006</year>
          .
          <article-title>The supply of child labor and household work</article-title>
          .
          <source>MPRA Paper 31402</source>
          , University Library of Munich, Germany.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Claudia</given-names>
            <surname>Goldin</surname>
          </string-name>
          .
          <year>2016</year>
          . Human Capital, Springer Verlag, Heidelberg, Germany.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Victor´ıa Orazem Peter</surname>
            <given-names>F</given-names>
          </string-name>
          &amp;
          <article-title>Sa´nchez Mario A Gunnarsson</article-title>
          .
          <year>2006</year>
          .
          <article-title>Child labor and school achievement in latin america</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Joe F Hair</surname>
          </string-name>
          ,
          <string-name>
            <surname>Christian M Ringle</surname>
            ,
            <given-names>and Marko</given-names>
          </string-name>
          <string-name>
            <surname>Sarstedt</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>PLS-SEM: Indeed a Silver Bullet</article-title>
          .
          <source>Journal of Marketing Theory and Practice</source>
          <volume>19</volume>
          (
          <issue>2</issue>
          ):
          <fpage>139</fpage>
          -
          <lpage>152</lpage>
          . https://doi.org/10.2753/MTP1069-6679190202.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Eric A</given-names>
            <surname>Hanushek</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Economic growth in developing countries: The role of human capital</article-title>
          .
          <source>Economics of Education Review</source>
          <volume>37</volume>
          :
          <fpage>204</fpage>
          -
          <lpage>212</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Simon</given-names>
            <surname>Haykin</surname>
          </string-name>
          .
          <year>1998</year>
          . Neural Networks:
          <string-name>
            <given-names>A Comprehensive</given-names>
            <surname>Foundation. Prentice Hall</surname>
          </string-name>
          <string-name>
            <surname>PTR</surname>
          </string-name>
          , Upper Saddle River, NJ, USA, 2nd edition.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Huajing</given-names>
            <surname>He</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Child labour and academic achievement: Evidence from gansu province in china</article-title>
          .
          <source>China Economic Review</source>
          <volume>38</volume>
          (C):
          <fpage>130</fpage>
          -
          <lpage>150</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>D W</given-names>
            <surname>Hosmer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S</given-names>
            <surname>Lemeshow</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R X</given-names>
            <surname>Sturdivant</surname>
          </string-name>
          .
          <year>2013</year>
          . Applied Logistic Regression. Wiley Series in Probability and Statistics. Wiley.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>James M Lattin</surname>
            ,
            <given-names>J D</given-names>
          </string-name>
          <string-name>
            <surname>Carroll</surname>
            , and
            <given-names>P E</given-names>
          </string-name>
          <string-name>
            <surname>Green</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>Analyzing Multivariate Data. Number v. 1 in Analyzing Multivariate Data</article-title>
          . Thomson Brooks/Cole.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Pablo</given-names>
            <surname>Lavado</surname>
          </string-name>
          and Jose´ Gallegos.
          <year>2005</year>
          .
          <article-title>La dina´mica de la desercio´n escolar en el Peru´: un enfoque usando modelos de duracio´n</article-title>
          .
          <source>Working Papers 05-08</source>
          , Departamento de Econom´
          <article-title>ıa, Universidad del Pac´ıfico</article-title>
          . https://ideas.repec.org/p/pai/wpaper/05-
          <fpage>08</fpage>
          .html.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>Huong Thu Le and Ross Homel</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>The impact of child labor on children's educational performance: Evidence from rural Vietnam</article-title>
          .
          <source>Journal of Asian Economics</source>
          <volume>36</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>13</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Luiz</given-names>
            <surname>Renato</surname>
          </string-name>
          <string-name>
            <surname>Lima</surname>
          </string-name>
          , Shirley Mesquita, and
          <string-name>
            <given-names>Marianne</given-names>
            <surname>Wanamaker</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Child labor and the wealth paradox: The role of altruistic parents</article-title>
          .
          <source>Economics Letters</source>
          <volume>130</volume>
          :
          <fpage>80</fpage>
          -
          <lpage>82</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Karla</given-names>
            <surname>Lopez</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Determinantes del trabajo infantil y la desercio´n escolar en menores de 12 a 17 an˜os en el Peru´ para los an˜os 2006 y 2014</article-title>
          .
          <article-title>Bachelor's degree</article-title>
          , PUCP, Peru´.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>S</given-names>
            <surname>James</surname>
          </string-name>
          Press and Sandra Wilson.
          <year>1978</year>
          .
          <article-title>Choosing between Logistic Regression and Discriminant Analysis</article-title>
          .
          <source>Journal of the American Statistical Association</source>
          <volume>73</volume>
          (
          <issue>364</issue>
          ):
          <fpage>699</fpage>
          -
          <lpage>705</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <surname>Diego C Rodrigues</surname>
          </string-name>
          ,
          <string-name>
            <surname>David N Prata</surname>
          </string-name>
          , and Michel A Silva.
          <year>2015</year>
          .
          <article-title>Exploring social data to understand child labor</article-title>
          .
          <source>International Journal of Social Science and Humanity</source>
          <volume>5</volume>
          (
          <issue>1</issue>
          ):
          <fpage>29</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <given-names>Jose</given-names>
            <surname>Rodriguez</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>Adquisicio´n de educacio´n escolar ba´sica en el Peru´: uso del tiempo de los menores en edad escolar</article-title>
          . Departamento de Econom´ıa - Pontificia
          <source>Universidad Cato´lica del Peru´</source>
          . http://EconPapers.repec.org/RePEc:pcp:pucotr:
          <fpage>otr2002</fpage>
          -
          <lpage>02</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          Jose Rodr´ıguez and
          <string-name>
            <given-names>Silvana</given-names>
            <surname>Vargas</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Trabajo infantil en el Peru´. Magnitud y perfiles vulnerables</article-title>
          .
          <source>Informe Nacional</source>
          <year>2007</year>
          -2008. Departamento de Econom´ıa - Pontificia
          <source>Universidad Cato´lica del Peru´.</source>
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          Jose´ Rodr´ıguez and
          <string-name>
            <given-names>Silvia</given-names>
            <surname>Vargas</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Escolaridad y trabajo infantil: patrones y determinantes de la asignacio´n del tiempo de nin˜os y adolescentes en Lima Metropolitana</article-title>
          .
          <source>Technical report</source>
          , PUCP, Peru.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <given-names>David E.</given-names>
            <surname>Rumelhart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Geoffrey E.</given-names>
            <surname>Hinton</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Ronald J.</given-names>
            <surname>Williams</surname>
          </string-name>
          .
          <year>1986</year>
          .
          <article-title>Learning internal representations by error propagation</article-title>
          .
          <source>In Parallel Distributed Processing: Explorations in the Microstructure of Cognition</source>
          , Volume
          <volume>1</volume>
          : Foundations, MIT Press, pages
          <fpage>318</fpage>
          -
          <lpage>362</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <given-names>Claudio</given-names>
            <surname>Sapelli</surname>
          </string-name>
          and Ar´ıstides Torche.
          <year>2004</year>
          . Desercio´n Escolar y Trabajo Juvenil: ¿Dos Caras de Una Misma Decisio´n? Cuadernos de econom´ıa pages
          <fpage>173</fpage>
          -
          <lpage>198</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <string-name>
            <given-names>Mariela</given-names>
            <surname>Sausa</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>El trabajo infantil es ma´s alto y ma´s penoso en las zonas rurales</article-title>
          .
          <source>Peru´ 21.</source>
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <source>BG Tabachnick and LS Fidell</source>
          .
          <year>2007</year>
          .
          <article-title>Multivariate analysis of variance and covariance</article-title>
          . In Using Multivariate Statistics, Allyn and Bacon Boston.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <surname>Jacek M Zurada</surname>
          </string-name>
          .
          <year>1992</year>
          .
          <article-title>Introduction to artificial neural systems</article-title>
          .
          <source>West St</source>
          . Paul.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>