<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Astrophysical Data Analytics based on Neural Gas Models, using the Classification of Globular Clusters as Playground</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>© Giuseppe Angora</string-name>
          <email>gius.angora@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>© Massimo Brescia</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>© Maurizio Paolillo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Physics “E. Pancini”, University Federico II</institution>
          ,
          <addr-line>Via Cinthia 6, 80126 Napoli</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>INAF Astronomical Observatory of Capodimonte</institution>
          ,
          <addr-line>Via Moiariello 16, 80131 Napoli</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Institute of Astrophysics, Pontificia Universidad Católica de Chile</institution>
          ,
          <addr-line>Av. Vicuña Mackenna 4860, Macul, Santiago</addr-line>
          ,
          <country country="CL">Chile</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Proceedings of the XIX International Conference “Data Analytics and Management in Data Intensive Domains” (DAMDID/RCDL'2017)</institution>
          ,
          <addr-line>Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Giuseppe Riccio</institution>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>Stefano Cavuoti</institution>
        </aff>
      </contrib-group>
      <fpage>381</fpage>
      <lpage>388</lpage>
      <abstract>
        <p>In Astrophysics, the identification of candidate Globular Clusters through deep, wide-field, single band HST images, is a typical data analytics problem, where methods based on Machine Learning have revealed a high efficiency and reliability, demonstrating the capability to improve the traditional approaches. Here we experimented some variants of the known Neural Gas model, exploring both supervised and unsupervised paradigms of Machine Learning, on the classification of Globular Clusters, extracted from the NGC1399 HST data. Main focus of this work was to use a well-tested playground to scientifically validate such kind of models for further extended experiments in astrophysics and using other standard Machine Learning methods (for instance Random Forest and Multi Layer Perceptron neural network) for a comparison of performances in terms of purity and completeness.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        The current and incoming astronomical synoptic
surveys require efficient and automatic data analytics
solutions to cope with the explosion of scientific data
amounts to be processed and analyzed. This scenario,
quite similar to other scientific and social contexts,
pushed all communities involved in data-driven
disciplines to explore data mining techniques and
methodologies, most of which connected to the Machine
Learning (hereafter ML) paradigms, i. e.
supervised/unsupervised self-adaptive learning and
parameter space optimization[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ],[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ],[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] .
      </p>
      <p>
        Following this premise, this paper is focused on the
investigation about the use of a particular kind of ML
methods, known as Neural Gas (NG) models[
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], to
solve classification problems within the astrophysical
context, characterized by a complex multi-dimensional
parameter space. In order to scientifically validate such
models, we decided to approach a typical astrophysical
playground, already solved with ML methods [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]
and to use in parallel other two ML techniques, chosen
among the most standard, respectively, Random Forest
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and Multi Layer Perceptron Neural Network[
        <xref ref-type="bibr" rid="ref23">23</xref>
        ], as
comparison baseline.
      </p>
      <p>
        The astrophysical case is related to the identification
of Globular Clusters (GCs) in the galaxy NGC1399 using
single band photometric data obtained through
observations with the Hubble Space Telescope (HST)
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ],[
        <xref ref-type="bibr" rid="ref27">27</xref>
        ].
      </p>
      <p>
        The physical identification and characterization of a
Globular Cluster (GC) in external galaxies is considered
important for a variety of astrophysical problems, from
the dynamical evolution of binary systems, to the
analysis of star clusters, galaxies and cosmological
phenomena [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ].
      </p>
      <p>
        Here, the capability of ML methods to learn and
recognize peculiar classes of objects, in a complex and
noising parameter space and by learning the hidden
correlation among object’s parameters, has been
demonstrated particularly suitable in the problem of GC
classification[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. In fact, multi-band wide-field
photometric data (colours and luminosities) are usually
required to recognize GCs within external galaxies, due
to the high risk of contamination of background galaxies,
which appear indistinguishable from galaxies located
few Mpc away, when observed by ground-based
instruments. Furthermore, in order to minimize the
contamination, high-resolution space-borne data are also
required, since they are able to provide particular
physical and structural features (such as concentration,
core radius, etc.), thus improving the GC classification
performance [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ].
      </p>
      <p>
        In[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] we demonstrated the capability of ML methods
to classify GCs using only single band images from
Hubble Space Telescope with a classification accuracy
of 98.3%, a completeness of 97.8% and only 1.6% of
residual contamination. Thus confirming that ML
methods may yield low contamination by minimizing the
observing requirements and extending the investigation
to the outskirts of nearby galaxies.
      </p>
      <p>These results gave us an optimal playground where
to train NG models and to validate their potential to solve
classification problems characterized by complex data
with a noising parameter space.</p>
      <p>The paper is structured as follows: in Sect. 2 we
describe the data used to test of the various methods. In
Sect. 3 we provide a short methodological and technical
description of the models. In Sect. 4 we describe the
experiments and results about the parameter space
analysis and classification experiments, while in Sect. 5
we discuss the results and draw our conclusions.</p>
    </sec>
    <sec id="sec-2">
      <title>2 The Astrophysical Playground</title>
      <p>
        As introduced, the HST single band data use dare very
suitable to investigate the classification of GCs. They, in
fact, are deep and complete in terms of wide-field
coverage, i. e. able to sample the GC population, to
ensure a high S/N ratio required to measure structural
parameters [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Furthermore, they provide the
possibility to study the overall properties of the GC
populations, which usually may differ from those of the
central region of a galaxy.
      </p>
      <p>With such data we intend to verify that Neural Gas
based models could be able to identify GCs with low
contamination even with single band photometric
information. Throughout the confirmation of such
behavior, we are confident that these models could solve
other astrophysical problems as well as in other
datadriven problem contexts.</p>
      <sec id="sec-2-1">
        <title>2.1 The data</title>
        <p>
          The data used in the described experiment consist of
wide field single band HST observations of the giant
elliptical NGC1399 galaxy, located in the core of the
Fornax cluster[
          <xref ref-type="bibr" rid="ref27">27</xref>
          ]. Due to its distance (D=20.130 Mpc,
see[
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]), it is considered an optimal case where to cover
a large fraction of its GC system with a restricted number
of observations. This dataset was used by[
          <xref ref-type="bibr" rid="ref25">25</xref>
          ] to study
the GC-LMXB connection and the structural properties
of the GC population. The optical data were taken with
the HST Advanced Camera for Surveys, in the broad V
band filter, with 2108 seconds of integration time for
each field. The observations were arranged in a 3x3 ACS
mosaic with a scale of 0.03 arcsec/pix, and combined into
a single image using the MultiDrizzle routine[
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]. The
field of view of the ACS mosaic covers ~100 square
arcmin (Figure 1), extending out to a projected
galactocentric distance of ~55 kpc.
        </p>
        <p>
          The source catalog was generated using Sextractor
[
          <xref ref-type="bibr" rid="ref4">4</xref>
          ],[
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], by imposing a minimum area of 20 pixels: it
contains 12915 sources and reaches 7σ detection at
        </p>
        <p>Figure 1 The FoV covered by the HST/ACS mosaic in
the broad V band</p>
        <p>The source subsample used to build our Knowledge
Base (KB) to train the ML models, is composed by 2100
sources with 11features (7 photometric and 4
morphological parameters).</p>
        <p>Such parameter space includes three aperture
magnitudes within 2, 6 and 20 pixels (mag_aper1,
mag_aper2, mag_aper3), is ophotal magnitude
(mag_iso), kron radius (kron_rad), central surface
brightness (mu0), FWHM (fwhm_im),and the four
structural parameters, respectively, ellipticity, King's
tidal, effective and core radii (calr_t, calr_h, calr_c). The
target values of the KB required as ground truth for
training and validation, i.e. the binary column indicating
the source as GC or not GC, is provided through the
typical selection based on multi-band magnitude and
colour cuts. The original 2100 sources having a target
assigned have been randomly shuffled and split into a
training (70%) and a blind test set (30%).</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3 The Machine Learning Models</title>
      <p>In our work we tested three different variants of the
Neural Gas model, using two additional machine
learning methods, respectively feed-forward neural
network and Random Forest, as comparison benchmarks.
In the following all main features of these models are
described.</p>
      <sec id="sec-3-1">
        <title>3.1 Growing Neural Gas</title>
        <p>
          Growing Neural Gas (GNG) is presented by[
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] as a
variant of the Neural Gas algorithm (introduced by[
          <xref ref-type="bibr" rid="ref21">21</xref>
          ]),
which combines the Competitive Hebbian Learning
(CHL, [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ]) with a vector quantization technique to
achieve a learning that retains the topology of the dataset.
        </p>
        <p>
          Vector quantization techniques[
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] encode a data
manifold, e.g. , using a finite set of reference
vectors
vector
vector
        </p>
        <p>
          , . Every data
is described by the best matching reference
for which the distortion error
is
minimal. This procedure divides the manifold
into a
number
, called Voronoi
polyhedra[
          <xref ref-type="bibr" rid="ref24">24</xref>
          ], within which each data vector is
described by the corresponding reference vector .
        </p>
        <p>
          The Neural Gas network is a vector quantization
model characterized by N neural units, each one
associated to a reference vector, connected to each other.
When an input is extracted, it induces a synaptic
excitation detected by all the neurons in the graph and
causes its adaptation. As shown in[
          <xref ref-type="bibr" rid="ref21">21</xref>
          ], the adaptation
rule can be described as a “winner-takes-most” instead
of “winner-takes-all” rule:
. (1)
The step size describes the overall extent of the
adaptation. While is a function in
which is the “neighborhood-ranking” of the reference
vectors. Simultaneously, the first and second Best
Matching Units (BMUs) develop connections between
each other[
          <xref ref-type="bibr" rid="ref21">21</xref>
          ].
        </p>
        <p>
          Each connection has an “age”; when the age of a
connection exceeds a pre-specified lifetime T, it is
removed[
          <xref ref-type="bibr" rid="ref21">21</xref>
          ]. Martinez's reasoning is interesting[
          <xref ref-type="bibr" rid="ref22">22</xref>
          ]:
they demonstrate how the dynamics of neural units can
be compared to a gaseous system. Let’s define the
density of vector reference at location
through
, where is the volume of
Voronoi polyhedra. Hence, is a step function on
each Voronoi polyhedra, but we can still imagine that
their volumes change slowly from one polyhedra to the
next, with continuous. In this way, it is possible to
derive an expression for the average change:
(2)
where
        </p>
        <p>is the data point distribution.</p>
        <p>The equation suggests the name Neural Gas: the average
change of the reference vectors corresponds to a motion
of particles in a potential . Superimposed on
the gradient of this potential there is a force proportional
to , which points toward the direction of the
space where the particle density is low.</p>
        <p>
          Main idea behind the GNG network is to
successively add new units to an initially small network,
by evaluating local statistical measures collected during
previous adaptation steps[
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. Therefore, each neural
unit in the graph has associated a local reconstruction
error, updated for the BMU at each iteration (i. e. each
time an input is extracted): .
        </p>
        <p>Unlike the Neural Gas network, in the GNG the
synaptic excitation is limited to the receptive fields
related to the Best Matching Unit and its topological
neighbors:
It is no longer necessary to calculate the ranking for all
neural units, but it is sufficient to determine the first and
the second BMU.</p>
        <p>
          The increment of the number of units is performed
periodically: during the adaptation steps the error
accumulation allows to identify the regions in the input
space where the signal mapping causes major errors.
Therefore, to reduce this error, new units are inserted in
such regions[
          <xref ref-type="bibr" rid="ref14">14</xref>
          ].
        </p>
        <p>
          An elimination mechanism is also provided: once
the connections, whose age is greater than a certain
threshold, have been removed, if their connected units
remain isolated (i.e. without emanating edges), those
units are removed[
          <xref ref-type="bibr" rid="ref14">14</xref>
          ].
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2 GNG with Radial Basis Function</title>
        <p>
          Fritzke describes an incremental Radial Basis Function
(RBF) network suitable for classification and regression
problems [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ].
        </p>
        <p>
          The network can be figured out as a standard RBF
network [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], with a GNG algorithm as embedded
clustering method, used to handle the hidden layer.
        </p>
        <p>Each unit of this hybrid model (hereafter GNGRBF)
is a single perceptron with an associated reference vector
and a standard deviation. For a given input-output pair
, the activation of the i-th unit is
described by
.</p>
        <p>Each of the single perceptron computes a weighted
sum of the activations:</p>
        <p>The adaptation rule applies to both reference vectors
forming the hidden layer and the RBF weights. For the
first, the adaptation rule is the same of the updating rule
for the GNG network, while for the weights:</p>
        <p>Similarly to the GNG network, new units are inserted
where the prediction error is high, updating only the Best
Matching Unit at each iteration:
(3)</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3 Supervised Growing Neural Gas</title>
        <p>
          The Supervised Growing Neural Gas (SGNG) algorithm
is a modification of the GNG algorithm that uses class
labels of data to guide the partitioning of data into
optimal clusters[
          <xref ref-type="bibr" rid="ref15">15</xref>
          ],[
          <xref ref-type="bibr" rid="ref20">20</xref>
          ]. Each of the initial neurons is
labelled with a unique class label. To reduce the class
impurity inside the cluster, the original learning rule (1)
is reformulated by considering the case where the BMU
belongs or not to the same class of the neuron whose
reference vector is the closest to the current input.
Depending on such situation the SGNG learning rule is
expressed alternatively as:
        </p>
        <p>
          Where
(4)
is the nearest class neuron and
is a function specifically introduced to
maintain neurons sufficiently distant one each other. For
the neuron which is topologically close to the neuron ,
the rule intends to increase the clustering accuracy[
          <xref ref-type="bibr" rid="ref20">20</xref>
          ].
The insertion mechanism has to reduce not only the
intradistances between data in a cluster, but also the impurity
of the cluster. Each unit has associated two kinds of error:
an aggregated and a class error. A new neuron is inserted
close to the neuron having a highest class error
accumulated, while the label is the same as the neuron
label with the greater aggregated error.
        </p>
      </sec>
      <sec id="sec-3-4">
        <title>3.4 Multi Layer Perceptron</title>
        <p>
          The Multi Layer Perceptron (MLP) architecture is
one of the most typical feed-forward neural
networks[
          <xref ref-type="bibr" rid="ref23">23</xref>
          ]. The term feed-forward is used to identify
basic behavior of such neural models, in which the
impulse is propagated always in the same direction, e.g.
from neuron input layer towards output layer, through
one or more hidden layers (the network brain), by
combining the sum of weights associated to all neurons.
        </p>
        <p>As easy to understand, the neurons are organized in
layers, with proper own role. The input signal, simply
propagated throughout the neurons of the input layer, is
used to stimulate next hidden and output neuron layers.
The output of each neuron is obtained by means of an
activation function, applied to the weighted sum of its
inputs.</p>
        <p>
          The weights adaptation is obtained by the Logistic
Regression rule[
          <xref ref-type="bibr" rid="ref17">17</xref>
          ], by estimating the gradient of the
cost function, the latter being equal to the logarithm of
the likelihood function between the target and the
prediction of the model. In this work, our implementation
of the MLP is based on the public library Theano[
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].
        </p>
      </sec>
      <sec id="sec-3-5">
        <title>3.5 Random Forest</title>
        <p>
          Random Forest (RF) is one of the most widely known
machine learning ensemble methods [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], since it uses a
random subset of candidate data features to build an
ensemble of decision trees. Our implementation makes
use of the public library scikit-learn[
          <xref ref-type="bibr" rid="ref26">26</xref>
          ]. This method has
been chosen mainly because it provides for each input
feature as core of importance (rank) measured in terms
of its informative contribution percentage to the
classification results. From the architectural point of
view, a RF is a collection (forest) of tree-structured
classifiers , where the are independent,
identically distributed random vectors and each tree casts
a unit vote for the most popular class at input. Moreover,
a fundamental property of the RF is the intrinsic absence
of training over fitting[
          <xref ref-type="bibr" rid="ref5">5</xref>
          ].
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4 The experiments</title>
      <p>The five models previously introduced have been
applied to the dataset described in Sec. 2.1 and their
performances have been compared to verify the
capability of NG models to solve particularly complex
classification problems, like the astrophysical
identification of GCs from single-band observed data.</p>
      <p>
        In order to evaluate the performances of the selected
classifiers, we decided to use three among the classical
and widely used statistical estimators, respectively,
average efficiency, purity, completeness and F1-score,
which can be directly derived from the confusion
matrix[
        <xref ref-type="bibr" rid="ref28">28</xref>
        ], showed in Figure 2. The average
efficiency(also known as accuracy, hereafter AE), is the
ratio between the sum of correctly classified objects on
both classes (true positives for both classes, hereafter tp)
and the total amount of objects in the test set. The purity
(als known as precision, hereafter pur) of a class
measures the ratio between the correctly classified
objects and the sum of all objects assigned to that class
(i.e. tp/ [tp+fp], where fp indicates the false
positives).While the completeness (also known as recall,
hereafter comp) of a class is the ratio tp/ [tp+fn], where
fn is the number of false negatives of that class. The
quantity tp+fn corresponds to the total amount of objects
belonging to that class. The F1-score is a statistical test
that considers both the purity and completeness of the
test to compute the score (i. e. 2 [pur*comp]/
[pur+comp]).
      </p>
      <p>By definition, the dual quantity of the purity is the
contamination, another important measure which
indicates the amount of misclassified objects for each
class.</p>
      <p>
        In statistical terms, it is well known the classical
tradeoff between purity and completeness in any
classification problem, particularly accentuated in
astrophysical problems[
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. In the specific case of the
GC identification, from the astrophysical point of view,
we were mostly interested to the purity, i. e. to ensure the
highest level of true GCs correctly identified by the
classifiers[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. However, within the comparison
experiments described in this work, our main goal was to
evaluate the performances of the classifiers mostly
related to the best tradeoff between purity and
completeness.
      </p>
      <sec id="sec-4-1">
        <title>4.2 Analysis of the Data Parameter Space</title>
        <p>
          Before to perform the classification experiments, we
preliminarily investigated the parameter space, defined
by the 11 features defined in Sec. 2.1, identifying each
object within the KB dataset of 2100 objects. Main goal
of this phase was to measure the importance of any
feature, i.e. its relevance in terms of informative
contribution to the solution of the problem. In the ML
context, this analysis is usually called feature
selection[
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. Its main role is to identify the most
relevant features of the parameter space, trying to
minimize the impact of the well known problem of the
curse of dimensionality, i.e. the fact that ML models
exhibit a decrease of performance accuracy when the
number of features is significantly higher than
optimal[
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]. This problem is mainly addressed to cases
with a huge amount of data and dimensions. However,
its effects may also impact contexts with a limited
amount of data and parameter space dimension.
        </p>
        <p>The Random Forest model resulted particularly
suitable for such analysis, since it is intrinsically able to
provide a feature importance ranking during the training
phase. The feature importance of the parameter space,
representing the dataset used in this work, is shown in
Figure 3.</p>
        <p>From the astrophysical point of view, this ranking is
in accordance with the physics of the problem. In fact, as
expected, among the five most important features there
are the four magnitudes, i. e. the photometric log-scale
measures of the observed object’s photonic flux through
different apertures of the detector. Furthermore, almost
all photometric features resulted as the most relevant.
Finally, by looking at the Figure 3, there is an interesting
gap between the first six and the last five features, whose
cumulative contribution is just ~11% of the total. Finally,
a very weak joined contribution (~3%) is carried by the
two worst features (kron_rad and calr_c), which can be
considered as the most noising/redundant features for the
problem domain.</p>
        <p>Based on such considerations, the analysis of the
parameter space provides a list of most interesting
classification experiments to be performed with the
selected five ML models. This list is reported in Table 1.</p>
        <p>The experiment E1 is useful to verify the efficiency
by considering the four magnitudes.</p>
        <p>The experiment E2 is based on the direct evaluation
of the best group of features as derived from the
importance results.</p>
        <p>The classification efficiency of the full photometric
subset of features is evaluated through the experiment
E3.</p>
        <p>Finally, the experiment E4 is performed to verify the
results by removing only the two worst features.</p>
        <p>Following the results of the parameter space analysis,
the original domain of features has been reduced, by
varying the number and types of included features.
Therefore, the classification experiments have been
performed on the dataset, described in Sec. 2.1,
composed by 2100 objects and represented by a
parameter space with up to a maximum of 9 features
(Table 1).</p>
        <p>Table 2 Statistical analysis of the classification
performances obtained by the five ML models on the
blind test set for the four selected experiments. All
quantities are expressed in percentage and related to
average efficiency (AE), purity for each class (purGC,
purNotGC), completeness for each class (compGC,
compNotGC) and the F1-score for GC class. The
contamination is the dual value of the purity</p>
        <sec id="sec-4-1-1">
          <title>ID Estimator R%F M%LP SG%NG GNG%RBF G%NG</title>
          <p>E1
E2
E3
E4</p>
          <p>AE
purGC
compGC
AE
purGC
compGC
AE
purGC</p>
          <p>The dataset has been randomly shuffled and split into
a training set of 1470 objects (70% of the whole KB) and
a blind test set of 630 objects (the residual 30% of the
KB).</p>
          <p>These datasets have been used to train and test the
selected five ML classifiers. The analysis of results,
reported in Table 2, has been performed on the blind test
set, in terms of the statistical estimators defined in
Sec. 4.2.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5 Discussion and Conclusions</title>
      <p>As already underlined, main goal of this work is the
validation of NG models as efficient classifiers in noising
and multi-dimensional problems, with performances at
least comparable to other ML methods, considered
“traditional” in terms of their use in such kind of
problems.</p>
      <p>By looking at Table 2 and focusing on the statistics
for the three NG models, it is evident that their result is
able to identify GCs from other background objects,
reaching a satisfying tradeoff between purity and
completeness in all experiments and for both classes. The
occurrence of statistical fluctuations is mostly due to the
different parameter space used in the four experiments.
Nevertheless, none of the three NG models overcome the
others in terms of the measured statistics.</p>
      <p>If we compare the NG models with the two additional
ML methods (Random Forest and MLP neural network),
their performances appears almost the same. This implies
that NG methods show classification capabilities fully
comparable to other ML methods.</p>
      <p>Another interesting aspect is the analysis of the
degree of coherence among the NG models in terms of
commonalities within classified objects. Table 3 reports
the percentages of common predictions for the objects
correctly classified by considering, respectively both and
single classes. On average, the three NG models are in
agreement among them for about 80% of the objects
correctly classified.</p>
      <p>Table 3 Statistics for the three NG models related to the
common predictions of the correctly classified objects.
Second column is referred to both classes, while the
third and fourth columns report, respectively, the
statistics for single classes</p>
      <sec id="sec-5-1">
        <title>EXP ID GC+%notGC G%C no%tGC</title>
        <p>E1
E2
E3
86.0
79.8
81.1
85.4
79.8
82.5
86.9
79.8
79.2</p>
        <p>E4 77.8 77.4 78.4</p>
        <p>This is also confirmed by looking at the Figure 4,
where the tabular results of Table 3 are showed through
the Venn diagrams, reporting also more details about
their classification commonalities.</p>
        <p>Finally, from the computational efficiency point of
view, the NG models have theoretically a higher
complexity than Random Forest and neural networks.
But, since they are based on a dynamic evolution of the
internal structure, their complexity strongly depends on
the nature of the problem and its parameter space.</p>
        <p>
          Nevertheless, all the presented ML models have a
variable architectural attitude to be compliant with the
parallel computing paradigms. Besides the
embarrassingly parallel architecture of the Random
Forest, the use of optimized libraries, like Theano[
          <xref ref-type="bibr" rid="ref1">1</xref>
          ],
make also models like MLP highly efficient. From this
point of view NG models have a high potentiality to be
parallelized. By optimizing GNG, the GNGRBF would
automatically benefit, since both share the same search
space, except for the RBF training additional cost. In
practice, the hidden layer of the supervised network
behaves just like a GNG network whose neurons act as
inputs for the RBF network. Consequently, with the same
number of iterations, the GNGRBF network performs a
major number of operations.
        </p>
        <p>On the other hand, the SGNG network is similar to
the GNG network, although characterized by a neural
insertion mechanism over a long period, thus avoiding
too rapid changes in the number of neurons and excessive
oscillations of reference vectors. Therefore, on average,
the SGNG network computational costs are higher than
the models based on the standard Neural Gas mechanism.</p>
        <p>In conclusion, although a more intensive test
campaign on these models is still ongoing, we can assert
that Neural Gas based models are very promising as
problem-solving methods, also in presence of complex
and multi-dimensional classification and clustering
problems, especially if preceded by an accurate analysis
and optimization of the parameter space within the
problem domain.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements</title>
      <p>MB acknowledges the PRIN-INAF 2014 Glittering
kaleidoscopes in the sky: the multifaceted nature and
role of Galaxy Clusters, and the PRIN-MIUR 2015
Cosmology and Fundamental Physics: illuminating the
Dark Universe with Euclid.</p>
      <p>MB, GL and MP acknowledge the
H2020-MSCAITN-2016 SUNDIAL (SUrvey Network for Deep
Imaging Analysis and Learning), financed within the
Call H2020-EU.1.3.1.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Al-Rfou</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alain</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Almahairi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          et al.:
          <article-title>Theano: A {Python} Framework for Fast Computation of Mathematical Expressions</article-title>
          . arXiv e-printsabs/1605.02688 (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Annunziatella</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mercurio</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brescia</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cavuoti</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Longo</surname>
          </string-name>
          , G.:
          <article-title>Inside Catalogs: A Comparison of Source Extraction Software</article-title>
          .
          <source>PASP 125</source>
          ,
          <issue>923</issue>
          (
          <year>2013</year>
          ).
          <source>doi: 10.1086/669333</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Astroinformatics</surname>
            . In: Brescia,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Djorgovski</surname>
            ,
            <given-names>S.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Feigelson</surname>
            ,
            <given-names>E.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Longo</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cavuoti</surname>
          </string-name>
          , S. (eds.)
          <source>International Astronomical Union Symposium</source>
          ,
          <volume>325</volume>
          (
          <year>2017</year>
          ).
          <source>ISBN: 9781107169951</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Bertin</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arnouts</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>SExtractor: Software for Source Txtraction</article-title>
          .
          <source>A&amp;A Suppl. Series</source>
          ,
          <volume>117</volume>
          , pp.
          <fpage>393</fpage>
          -
          <lpage>404</lpage>
          (
          <year>1996</year>
          ). doi:
          <volume>10</volume>
          .1051/aas:1996164
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Breiman</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <source>Machine Learning</source>
          ,
          <volume>45</volume>
          . Springer Eds., pp.
          <fpage>25</fpage>
          -
          <lpage>32</lpage>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Brescia</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cavuoti</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Longo</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nocella</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garofalo</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , et al.:
          <article-title>DAMEWARE: A Web Cyberinfrastructure for Astrophysical Data Mining</article-title>
          .
          <source>PASP</source>
          .
          <volume>126</volume>
          ,
          <issue>942</issue>
          (
          <year>2014</year>
          ).
          <source>doi: 10.1086/677725</source>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Brescia</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Longo</surname>
          </string-name>
          , G.:
          <article-title>Astroinformatics, Data Mining and the Future of Astronomical Research</article-title>
          .
          <source>Nuclear Instruments and Methods in Physics Research A</source>
          ,
          <volume>720</volume>
          , pp.
          <fpage>92</fpage>
          -
          <lpage>94</lpage>
          ,
          <string-name>
            <surname>Elsevier</surname>
          </string-name>
          (
          <year>2013</year>
          ). doi:
          <volume>10</volume>
          .1016/j.nima.
          <year>2012</year>
          .
          <volume>12</volume>
          .027
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Brescia</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cavuoti</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paolillo</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Longo</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Puzia</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>The Detection of Globular Clusters in Galaxies as a Data Mining Problem</article-title>
          .
          <source>MNRAS 421</source>
          ,
          <issue>2</issue>
          , pp.
          <fpage>1155</fpage>
          -
          <lpage>1165</lpage>
          (
          <year>2012</year>
          ). doi:
          <volume>10</volume>
          .1111/j.1365-
          <fpage>2966</fpage>
          .
          <year>2011</year>
          .
          <volume>20375</volume>
          .x
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Broomhead</surname>
            ,
            <given-names>D.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lowe</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Radial Basis Functions, Multi-Variable Functional Interpolation and Adaptive Networks</article-title>
          .
          <source>Technical report. RSRE 4148</source>
          (
          <year>1988</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Carlson</surname>
            ,
            <given-names>M.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Holtzman</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          :
          <article-title>Measuring Sizes of Marginally Resolved Young Globular Clusters with the Hubble Space Telescope</article-title>
          .
          <source>PASP 113</source>
          ,
          <issue>790</issue>
          , pp.
          <fpage>1522</fpage>
          -
          <lpage>1540</lpage>
          (
          <year>2001</year>
          ).
          <source>doi: 10.1086/324417</source>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Cavuoti</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garofalo</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brescia</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paolillo</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pescapè</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Longo</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ventre</surname>
          </string-name>
          , G.:
          <article-title>Astrophysical Data Mining with GPU. A Case Study: Genetic Classification of Globular Clusters</article-title>
          . New Astronomy,
          <volume>26</volume>
          , pp.
          <fpage>12</fpage>
          -
          <lpage>22</lpage>
          (
          <year>2014</year>
          ). doi:
          <volume>10</volume>
          .1016/j.newast.
          <year>2013</year>
          .
          <volume>04</volume>
          .004
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>D</given-names>
            <surname>'Isanto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Cavuoti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Brescia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Donalek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Longo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            ,
            <surname>Riccio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            ,
            <surname>Djorgovski</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.G.</surname>
          </string-name>
          :
          <article-title>An Analysis of Feature Relevance in the Classification of Astronomical Transients with Machine Learning Methods</article-title>
          .
          <source>MNRAS</source>
          <volume>457</volume>
          (
          <issue>3</issue>
          ), pp.
          <fpage>3119</fpage>
          -
          <lpage>3132</lpage>
          (
          <year>2016</year>
          ). doi:
          <volume>10</volume>
          .1093/mnras/stw157
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Dunn</surname>
            ,
            <given-names>L.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jerjen</surname>
          </string-name>
          , H.:
          <article-title>First Results from SAPAC: Toward a Three-dimensional Picture of the Fornax Cluster Core</article-title>
          .
          <source>AJ</source>
          <volume>132</volume>
          (
          <issue>3</issue>
          ), pp.
          <fpage>1384</fpage>
          -
          <lpage>1395</lpage>
          (
          <year>2006</year>
          ).
          <source>doi: 10.1086/506562</source>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Fritzke</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>A Growing Neural Gas Network Learns Topologies</article-title>
          .
          <source>In: Advances in Neural Information Processing System</source>
          , 7,
          <string-name>
            <given-names>G.</given-names>
            <surname>Tesauro</surname>
          </string-name>
          , D.S. Touretzky and T.K. Leen (eds.), MIT Press, Cambridge MA (
          <year>1995</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Fritzke</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Supervised Learning with Growing Cell Structures</article-title>
          .
          <source>In: Advances in Neural Information Processing System</source>
          ,
          <volume>6</volume>
          ,
          <string-name>
            <surname>Cowan</surname>
            ,
            <given-names>J.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tesauro</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Alspector</surname>
            ,
            <given-names>J</given-names>
          </string-name>
          . (eds.), Morgan-Kaufmann, pp.
          <fpage>255</fpage>
          -
          <lpage>262</lpage>
          (
          <year>1994</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Guyon</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elisseeff</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>An Introduction to Variable and Feature Selection</article-title>
          .
          <source>JMLR 3</source>
          , pp.
          <fpage>1157</fpage>
          -
          <lpage>1182</lpage>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Harrell</surname>
            ,
            <given-names>F.E.</given-names>
          </string-name>
          : Regression Modeling Strategies. Springer-Verlag (
          <year>2001</year>
          ).
          <source>ISBN 0-387-95232-2</source>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Hughes</surname>
            ,
            <given-names>G.F.</given-names>
          </string-name>
          :
          <article-title>On the Mean Accuracy of Statistical Pattern Recognizers</article-title>
          .
          <source>IEEE Transactions on Information Theory</source>
          ,
          <volume>14</volume>
          (
          <issue>1</issue>
          ), pp.
          <fpage>55</fpage>
          -
          <lpage>63</lpage>
          (
          <year>1968</year>
          ). doi:
          <volume>10</volume>
          .1109/TIT.
          <year>1968</year>
          .1054102
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Koekemoer</surname>
            ,
            <given-names>A.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fruchter</surname>
            ,
            <given-names>A.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hook</surname>
            ,
            <given-names>R.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hack</surname>
            ,
            <given-names>W.:</given-names>
          </string-name>
          <article-title>MultiDrizzle: An Integrated Pyraf Script for Registering, Cleaning and Combining Images</article-title>
          . In: The 2002
          <string-name>
            <given-names>HST</given-names>
            <surname>Calibration</surname>
          </string-name>
          <article-title>Workshop</article-title>
          . Santiago Arribas, Anton Koekemoer, and Brad Whitmore (eds.). Baltimore, MD: Space Telescope Science Institute (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Jirayusakul</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aryuwattanamongkol</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>A Supervised Growing Neural Gas Algorithm for Cluster Analysis</article-title>
          . Springer-Verlag (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Martinez</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schulten</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>A Neural-Gas Network Learns Topologies</article-title>
          .
          <source>In: Artificial Neural Networks. T.</source>
          <string-name>
            <surname>Kohonen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Makisara</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          <string-name>
            <surname>Simula</surname>
          </string-name>
          , and J.
          <source>Kangas (eds.)</source>
          , Amsterdam, The Netherlands, Elsevier, pp.
          <fpage>397</fpage>
          -
          <lpage>402</lpage>
          (
          <year>1991</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Martinez</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berkovich</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schulten</surname>
            ,
            <given-names>K.J.</given-names>
          </string-name>
          :
          <article-title>Neural Gas Network for Vector Quantization and its Application to Time-Series Prediction</article-title>
          .
          <source>In: IEEE Transactions on Neural Networks</source>
          ,
          <volume>4</volume>
          (
          <issue>4</issue>
          ), pp.
          <fpage>558</fpage>
          -
          <lpage>569</lpage>
          (
          <year>1993</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>McCulloch</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pitts</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <source>Bulletin of Mathematical Biophysics</source>
          ,
          <volume>5</volume>
          (
          <issue>4</issue>
          ), pp.
          <fpage>115</fpage>
          -
          <lpage>133</lpage>
          (
          <year>1943</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Montoro</surname>
            ,
            <given-names>J.C.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Abascal</surname>
            ,
            <given-names>J.L.F.</given-names>
          </string-name>
          :
          <article-title>The Voronoi Polyhedra as Tools for Structure Determination in Simple Disordered Systems</article-title>
          .
          <source>J. Phys. Chem</source>
          .,
          <volume>97</volume>
          (
          <issue>16</issue>
          ), pp.
          <fpage>4211</fpage>
          -
          <lpage>4215</lpage>
          (
          <year>1993</year>
          ). doi:
          <volume>10</volume>
          .1021/j100118a044
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Paolillo</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Puzia</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goudfrooij</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          et al.:
          <article-title>Probing the GC-LMXB Connection in NGC 1399: A Wide-field Study with the Hubble Space Telescope and Chandra</article-title>
          .
          <source>ApJ</source>
          ,
          <volume>736</volume>
          (
          <issue>2</issue>
          ), p.
          <volume>90</volume>
          (
          <year>2011</year>
          ). doi:
          <volume>10</volume>
          .1088/
          <fpage>0004</fpage>
          -637X/736/2/90
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Pedregosa</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Varoquaux</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gramfort</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          et al.:
          <article-title>Scikit-learn: Machine Learning in Python</article-title>
          . JMLR,
          <volume>12</volume>
          , pp.
          <fpage>2825</fpage>
          -
          <lpage>2830</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <surname>Puzia</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paolillo</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goudfrooij</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maccarone</surname>
            ,
            <given-names>T.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fabbiano</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Angelini</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Wide-field Hubble Space Telescope Observations of the Globular Cluster System in NGC 1399</article-title>
          . ApJ,
          <volume>786</volume>
          (
          <issue>2</issue>
          ), p.
          <volume>78</volume>
          (
          <year>2014</year>
          ). doi:
          <volume>10</volume>
          .1088/
          <fpage>0004</fpage>
          -637X/786/2/78
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <surname>Stehman</surname>
            ,
            <given-names>S.V.</given-names>
          </string-name>
          :
          <article-title>Selecting and Interpreting Measures of Thematic Classification Accuracy</article-title>
          .
          <source>Remote Sensing of Environment</source>
          ,
          <volume>62</volume>
          (
          <issue>1</issue>
          ), pp.
          <fpage>77</fpage>
          -
          <lpage>89</lpage>
          (
          <year>1997</year>
          ). doi:
          <volume>10</volume>
          .1016/S0034-
          <volume>4257</volume>
          (
          <issue>97</issue>
          )
          <fpage>00083</fpage>
          -
          <lpage>7</lpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>