<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Stacked Ensemble Model based on RUSBoost and a Cost-Sensitive Convolutional Neural Network for Class Imbalance in Big Data Analytics</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Lawrence Nderu</string-name>
          <email>lawrence_nderu@live.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dorcas Gicuku Mwigereri</string-name>
          <email>dorcausgicuku@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Peter Irungu Mwangi</string-name>
          <email>peterirungu16@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>School of Computing and, Information Technology, Jomo Kenyatta University of, Agriculture and Technology</institution>
          ,
          <addr-line>Nairobi</addr-line>
          ,
          <country country="KE">Kenya</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>- Machine learning algorithms have been designed such that they focus more on attaining high accuracy levels and as such they tend to misclassify instances that belong to the minority class which are often the instances of interest. In big data analytics, the issue of handling class imbalance has been understudied and this is mainly attributed to the problems of small disjuncts of an imbalanced dataset. This study proposes a stacked ensemble model based on RUSBoost and cost sensitive convolution neural network for tackling class imbalance issues in big data analytics, The RUSBoost algorithm handles the class imbalance at the data level by intelligently and randomly removing samples from the majority class while the costsensitive convolutional neural network handles the imbalance at the algorithm level by enabling the convolutional algorithm automatically learn the cost during the training period. The ImageNet and WHOI-Plankton datasets were used to evaluate the proposed model as they met the imbalance ration standard set of 100:1 in addition to containing more than 100,000 records. The results posted showed the proposed stacked ensemble model outperformed existing ensemble techniques such as the SMOTEBoost and AdaBoost which were seen to immensely improve on the classification performance given an imbalanced dataset. The recall, precision and gmean value posted by the proposed algorithm were 80.2%, 95% and 87.3%respectively. Therefore the proposed model is seen as a better, faster and less complex alternative for handling class imbalance in big dataset.</p>
      </abstract>
      <kwd-group>
        <kwd>class imbalance</kwd>
        <kwd>big data analytics</kwd>
        <kwd>random under sampling</kwd>
        <kwd>cost</kwd>
        <kwd>Convolutional neural network</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>I. INTRODUCTION</title>
      <p>
        Big data, a term with no formal definition is commonly
characterized with 5 V’s. that is, velocity, volume, veracity,
variety and value whereby the volume and variety
characteristics are coined from the aspect that big data
involves massive amount of various structured,
semistructure and unstructured data while the velocity
characteristic implies that the data is collected with a very
high rate. The veracity characteristic is used to indicate the
quality of the data collected while the value characteristic
explains whether the data collected is of importance,
depending on the problem [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. With the availability of these
massive data, most organizations aim at analyzing the data
so as to identify hidden patterns and insights that may not
be easily identified by a human, for making informed
decisions and support their strategies.
      </p>
      <p>
        In data analytics, deep learning methods have gained
greater attention as compared to other machine learning
(ML) classification algorithms they have been associated
with improved performance levels especially on domains
that involve complex datasets [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. This has mainly been
attributed to factors such as availability of hardware and
software components, availability of data, improvements of
the algorithms that help to speed up the time taken to train a
model and generalize new data. Despite these
advancements, [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]states that these deep learning algorithms
have been seen to perform poorly given an imbalanced
dataset whereby the samples from the majority class are
seen to dominate the gradient value which updates the
overall model’s weight and as a result, the errors from the
majority class are reduced while error obtained from the
minority class are increased leading to a slow convergence
of the network.
      </p>
      <p>
        Class imbalance, a problem common to classification
algorithms, results from having fewer samples from the
minority group as compared to the sample amount retrieved
from the majority group in the same dataset [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The
underlying assumption of both DL and ML classification
algorithms is that the classes, present in a dataset, have been
represented in almost equal proportions [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        Nonetheless, this is contrary to what happens in real
world scenarios whereby the aspect being measured is often
less represented as compared to its counterpart. Depending
on factors such as the imbalance ratio, complexity of the
concept represented, classifier involved and overall size of
training set, the degree of class imbalance is seen to vary
from minor to severe class imbalance and class rarity[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        Currently, techniques identified for dealing will class
imbalance issues can be broadly classified into data-level,
algorithm level and hybrid techniques [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Data level
techniques for handling class imbalance have been designed
such that they eliminate the class imbalance in the training
dataset either through random under-sampling(RUS) that
involves removing samples randomly from the majority
class or through random over-sampling(ROS) that involves
a process of randomly duplicating values from the minority
class. The algorithm level techniques aim at modifying the
algorithm such that it will not be biased towards the majority
class instances while the hybrid techniques ,which are also
referred to as ensemble techniques, combine both data level
and algorithm modification techniques [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] did a survey paper that aimed at summarizing the
research works conducted from 2010 to 2018 to address the
issue of class imbalance in big data. To ensure they only
considered works that involved big data only, they selected
works that utilized datasets consisting of a minimum of
100,000 records.
      </p>
      <p>
        From their work it was noted that there exists very little
work on class imbalance when dealing with big data.
Additionally, tackling the issues of class imbalance on the
data-level, the ROS gives better performance but this is
exclusively on smaller datasets as using this technique on a
large dataset may result to generating a very huge amount
of dataset that may lead to an increase in the training time
thus lowering efficiency. For big data analytics, they
suggested use of RUS which is assumed will perform better
by removing noise and redundant samples from the dataset.
Additionally, the algorithm level techniques that were
identified included the mean false error(MFE) loss and
mean squared false error(MSFE) proposed by [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] , focal
loss[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] that aims at reducing the impact samples that are
easily classified have on the loss of the classifier and use of
a cost-sensitive convolutional neural network (CNN) [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]
that is capable of learning the weights and cost during the
training process. Of the identified algorithm level
techniques, the cost-sensitive CNN and focal loss technique
have been seen to generalize and adapt to various domains
that involve complex datasets as they are not dependent on
the knowledge of a specific domain. Ensemble algorithms
such as AdaBoost[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], UnderBagging[
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], RUSBoost[
        <xref ref-type="bibr" rid="ref13">13</xref>
        ],
SmoteBoost[
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], and overbagging[
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] were identified. Of
these algorithms, the SmoteBoost, RUSBoost and
Underbagging were seen to post high performance levels.
Additionally the RUSBoost was seen to have more
superiority in terms of performance and being the least
complex algorithm. Despite the great performance levels
being posted by the ensemble, it was noted that there still
exists a problem of cost definition.
      </p>
      <p>To handle the class imbalance problem in big data
analytics, this study proposes an ensemble technique based
on RUSBoost and cost-sensitive CNN. Using this approach,
it is expected that the deep learning algorithm will be able
to perform better even with an imbalanced dataset and the
cost associated with the training process will be
automatically defined.</p>
      <p>II.</p>
      <p>
        RELATED WORKS
[
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] did an analysis on data level techniques for
handling class imbalance in big data analytics, that is, the
ROS, RUS and SMOTE in the MapReduce framework and
the impact these techniques have when evaluated using
decision trees(DT) and random forest(RF) in Spark and
Hadoop. To train the model, the evolutionary computation
for big data and big learning competition (ECBDL14)
dataset was used. To further evaluate the impact on these
techniques given different number of maps in a MapReduce
environment, 1,8,16, 32 and 64 maps over both Spark and
Hadoop were evaluated. The results obtained showed that
SMOTE and RF performed better than SMOTE and DT.
Also it was observed that the RUS and ROS performed
better than SMOTE in the big data scenario and that the
performance level obtained greatly depended on the
behavior of the classifier used. To sum it up, it was
concluded that the ROS performed independent of the
number of maps being used while the data distribution was
greatly affected when the RUS was applied on different
partitions. As a result, increasing the number of partitions
was seen to have severe effect for the RUS and this was
mainly attributed to the lack of data.
      </p>
      <p>
        [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] Described a Random over Sampling and
Evolutionary Feature Weighting for Random Forest
(ROSEFW-RF) which won the ECBDL14 big data
challenge held for the big data challenge in bioinformatics
domain. The ROSEFW-RF technique, based on
MapReduce environment was used to tackle the class
imbalance problem using random oversampling then the
evolutionary feature weighting was used to identify relevant
features which were then selected using a threshold. A RF
model was then trained using the preprocessed dataset. To
evaluate the performance measure of the ROSEFW-RF,
64,192 and 256 mappers were used with 100 trees defined
for the RF model. They further did an experiment of using
RF classifier with a dataset that used 100% oversampling
ratio. The results obtained during this first experiment
showed that a very low true positive rate (TPR) was
obtained and compared to the true negative rate (TNR) and
that this difference tends to increase with a decrease in the
number of mappers used. Experiments done using the
ROSEFW-RF showed that this technique outperformed
other strategies in the competition and was capable of
balancing the TNR and TPR which had been a considered
to be a difficult task during the competition. For future
research, they proposed analysis of the effects of the number
of maps and other classifiers in addition to utilizing a
strategy that would combine over sampling and under
sampling or even use of instance reduction techniques for
handling the class imbalance challenge on big data
analytics.
      </p>
      <p>
        In [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] the authors proposed a supervised technique that
would handle the class imbalance between nonortholog and
ortholog classes observed on ortholog detection in different
yeast species. The proposed methodology was structured
such that it involved three steps that involved calculation of
the different gene pair features that were supposed to be
combined, building the ML classifier and classification of
the obtained gene pair features. The proposed technique was
then compared using various proposed models such as the
RF for Big Data with Cost-Sensitive (RF-BDCS) described
in [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ],the Random Oversampling with RF for Big
Data(ROS+RF-BD) described in [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] and the Support
Vector Machines for Big Data(SVM-BD) in the Apache
Spark environment. During the experiment, the authors
selected datasets with various genome yeast pairs, that is:
Saccharomyces-Klutveromyces lactis, Saccharomyces
cerevisiae-Candida glabrata and Saccharomyces
cerevisiaeSchizosaccharomyces pombe. To evaluate the performance
measure, the true positive rate (TPR), true negative rate
(TNR), the area under curve (AUC) and G-Mean
performance metrics were used. The results obtained
indicated that the proposed supervised technique for gene
pairwise feature combination gave the best for pairwise
ortholog detection in big data scenarios. Additionally, using
ROS with the SVM-BD classifier gave better results as
compared to the other tested techniques. For future works,
they recommended use of new gene pair features for the
supervised algorithm for pairwise ortholog detection.
      </p>
    </sec>
    <sec id="sec-2">
      <title>III. PROPOSED METHOD</title>
      <p>A. Dataset Description</p>
      <p>
        [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] Defined class imbalance in data as data having an
imbalance ratio between the ranges of 100:1 to
10000:1.Consequetly,to train and evaluate the proposed
model, datasets with a class imbalance ratio of 100:1 and
above were considered in this study. Additionally, to
incorporate the big data aspect, the WHOI-Plankton[
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]
and ImageNet[
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] datasets with a minimum of 100,000
records, as described in in table 1, were obtained.
      </p>
      <p>The obtained dataset was partitioned to training and test
dataset using k-cross fold validation whereby k=10. Nine
(9) folds were used to train the model while the other one
fold was used to test the performance of the model.</p>
      <p>
        To harness the computational power and great
performance of RUSBoost as reported by [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] and address
the issue of cost definition in convolutional neural networks
while handling datasets with class imbalance, this study
proposes a stacked ensemble based on the RUSBoost and a
cost-sensitive CNN. The RUSBoost was used to sample the
dataset at the data level by under sampling instances from
the majority class and the obtained result fed to the
costsensitive CNN for classification as illustrated in figure 1.
      </p>
      <p>
        The residual network (ResNet) architecture for training
deep neural networks was used to develop the proposed
model. The decision to adopt the ResNet architecture was
motivated by the high level of performance reported when
using the ResNet architecture for deep neural networks as a
result of its skip connections [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ].
      </p>
      <p>Fig 1: The RUSBoost, Cost Sensitive CNN Stacked Ensemble Model</p>
      <p>To enable the CNN automatically determine the cost
associated with each misclassification, an additional cost
layer was added to manipulate the output of the
convolutional layer before it is fed to the softmax layer for
classification. The cost layer was designed such that it was
capable of automatically updating the cost matrix in
equation (1) using an empirical risk value illustrated in
equation (2).</p>
    </sec>
    <sec id="sec-3">
      <title>Where: R1 (o) is the empirical risk.</title>
      <p>Y is the class labels while n is the total number of
instances in the dataset.</p>
      <p>C is the cost matrix whereby the cost value is set as an
imbalanced ratio (IR) when the class predicted q matches
the actual class p</p>
      <p>The o(i) is the predicted output while the desired output
is represented by d(i) .</p>
      <p>The RUSBoost was implemented such that the instances
from the majority class are under sampled using an
intelligent random under sampling technique as illustrated
in the table 2 which gives the RUSBoost algorithm.
C. Performance Evaluation</p>
      <p>
        To evaluate the performance measure of the developed
ensemble model, a confusion matrix shown in table 3, was
used to show the true positive, true negative, false positive
and false negative values. These values from the confusion
matrix were then used to calculate the precision, recall and
geometric mean of the true negative rate (TNR) and the true
positive rate (TPR) metrics as shown in equation(3),(4) and
(5) respectively for evaluation of the proposed model.
According [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] to the accuracy level is not a favorable
performance evaluation metric in an imbalanced data as it
is not sensitive to class imbalance hence might give
misleading results. Consequently it was not considered in
this study.
(1)
(2)
(3)
(4)
(5)
1. Initialize all wi=eights to 1/m for all instances
      </p>
      <p>
        ; D1(i) = 1/m for all i
Repeat for t = 1,2,3,… T
[
        <xref ref-type="bibr" rid="ref1">0,1</xref>
        ]
Dt)
Create temporary training set S`t
with a D`t distribution using random
undersampling
D`t weights
Call WL with S`t dataset with their
Get back a hypothesis ht : X*Y -&gt;
Calculate the pseudo-loss ( for S and
  ( )(1 − ℎ (  ,
      </p>
      <p>)ℎ (  ,  ))</p>
    </sec>
    <sec id="sec-4">
      <title>Calculate</title>
      <p>parameter :
Update Dt:
the
weight
update
  =
є

1 − є
1(1+ℎ (  ,  )−ℎ (  , : ≠  ))
2
Normalize Dt+1 : Let
  = ∑   +1( )</p>
      <p>+1( ) =   +1( )
3. Output the final hypothesis
 ⋲
 ( ) = argmax ∑ ℎ ( ,  ) log
 
1
 

 =1</p>
    </sec>
    <sec id="sec-5">
      <title>Actual Positive True</title>
    </sec>
    <sec id="sec-6">
      <title>False</title>
      <p>the goal of enabling the classifier to automatically learn the
cost during the training process and as a result improve on
the performance level of the model.
classifier was capable of minimizing the errors as a result of
misclassifications.
positives and false negatives. According to the literature
reviewed, there exists limited studies that focus on class
imbalance in big data analytics. Consequently, it has been
suggested various techniques for handling class imbalance
in traditional machine learning algorithms can be extend to
big data analytics using deep learning techniques. In this
work an ensemble model based on RUSBoost and
costsensitive CNN for handling class imbalance issue in
datasets is proposed. The performance measure of the
proposed model was compared with the performance posted
when using RUS with CNN and RUSBoost with cost
sensitive CNN.</p>
      <p>According to the obtained results, the ensemble model
based on RUSBoost and a cost sensitive CNN model has
proved superiority in terms of performance measure as
compared to the using the RUSBoost . As a result, this
technique can be adopted to a variety of big-data analytics
applications for handling imbalanced datasets with a CNN
classifier.</p>
      <p>For future works, we recommend evaluation of the
proposed model on the different big data analytics platforms
such as Hadoop and Apache Spark so as to evaluate if it is
independent of the big data platform used.</p>
    </sec>
    <sec id="sec-7">
      <title>ACKNOWLEDGEMENT</title>
      <p>We would like to thank the Africa-ai Japan Project 2020
for the financial support offered for this project</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>I.</given-names>
            <surname>Kalbandi</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Anuradha</surname>
          </string-name>
          , “
          <article-title>A Brief Introduction on Big Data 5Vs Characteristics and Hadoop Technology,” Procedia Comput</article-title>
          . Sci., vol.
          <volume>48</volume>
          , pp.
          <fpage>319</fpage>
          -
          <lpage>324</lpage>
          , Dec.
          <year>2015</year>
          , doi: 10.1016/j.procs.
          <year>2015</year>
          .
          <volume>04</volume>
          .188.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Muniasamy</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Alasiry</surname>
          </string-name>
          , “
          <article-title>Deep Learning: The Impact on Future eLearning,”</article-title>
          <string-name>
            <given-names>Int. J.</given-names>
            <surname>Emerg</surname>
          </string-name>
          . Technol. Learn., vol.
          <volume>15</volume>
          , p.
          <fpage>188</fpage>
          ,
          <string-name>
            <surname>Jan</surname>
          </string-name>
          .
          <year>2020</year>
          , doi: 10.3991/ijet.v15i01.
          <fpage>11435</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Johnson</surname>
          </string-name>
          and
          <string-name>
            <surname>T. M. Khoshgoftaar</surname>
          </string-name>
          , “
          <article-title>Survey on deep learning with class imbalance</article-title>
          ,
          <source>” J. Big Data</source>
          , vol.
          <volume>6</volume>
          , no.
          <issue>1</issue>
          , p.
          <fpage>27</fpage>
          ,
          <year>2019</year>
          , doi: 10.1186/s40537-019-0192-5.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Leevy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. M.</given-names>
            <surname>Khoshgoftaar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. A.</given-names>
            <surname>Bauder</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N.</given-names>
            <surname>Seliya</surname>
          </string-name>
          , “
          <article-title>A survey on addressing high-class imbalance in big data</article-title>
          ,
          <source>” J. Big Data</source>
          , vol.
          <volume>5</volume>
          , no.
          <issue>1</issue>
          , p.
          <fpage>42</fpage>
          ,
          <year>2018</year>
          , doi: 10.1186/s40537-018-0151-6.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Iram</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.-B.</given-names>
            <surname>Vialatte</surname>
          </string-name>
          , and
          <string-name>
            <surname>M. I. Qamar</surname>
          </string-name>
          , “Chapter 1
          <article-title>- Early Diagnosis of Neurodegenerative Diseases from Gait Discrimination to Neural Synchronization,” in Emerging Topics in Computer Science</article-title>
          and Applied Computing,
          <string-name>
            <given-names>D.</given-names>
            <surname>Al-Jumeily</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hussain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Mallucci</surname>
          </string-name>
          , and
          <string-name>
            <surname>C. B. T</surname>
            .-A. C. in M. and
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Oliver</surname>
          </string-name>
          , Eds. Boston: Morgan Kaufmann,
          <year>2016</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>26</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>R. A.</given-names>
            <surname>Bauder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. M.</given-names>
            <surname>Khoshgoftaar</surname>
          </string-name>
          , and T. Hasanin, “
          <article-title>An Empirical Study on Class Rarity in Big Data,”</article-title>
          <source>Proc. - 17th IEEE Int. Conf. Mach. Learn. Appl. ICMLA</source>
          <year>2018</year>
          , pp.
          <fpage>785</fpage>
          -
          <lpage>790</lpage>
          ,
          <year>2019</year>
          , doi: 10.1109/ICMLA.
          <year>2018</year>
          .
          <volume>00125</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>K.</given-names>
            <surname>Yang</surname>
          </string-name>
          et al.,
          <article-title>“Hybrid Classifier Ensemble for Imbalanced Data,”</article-title>
          <source>IEEE Trans. Neural Networks Learn. Syst.</source>
          , vol.
          <volume>31</volume>
          , no.
          <issue>4</issue>
          , pp.
          <fpage>1387</fpage>
          -
          <lpage>1400</lpage>
          ,
          <year>2020</year>
          , doi: 10.1109/TNNLS.
          <year>2019</year>
          .
          <volume>2920246</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          , W. Liu,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Cao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Meng</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P. J.</given-names>
            <surname>Kennedy</surname>
          </string-name>
          , “
          <article-title>Training deep neural networks on imbalanced data sets</article-title>
          ,” in
          <source>2016 International Joint Conference on Neural Networks (IJCNN)</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>4368</fpage>
          -
          <lpage>4374</lpage>
          , doi: 10.1109/IJCNN.
          <year>2016</year>
          .
          <volume>7727770</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>T.-Y.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. B.</given-names>
            <surname>Girshick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Dollár</surname>
          </string-name>
          , “
          <article-title>Focal Loss for Dense Object Detection</article-title>
          ,”
          <source>2017 IEEE Int. Conf. Comput. Vis.</source>
          , pp.
          <fpage>2999</fpage>
          -
          <lpage>3007</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S. H.</given-names>
            <surname>Khan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hayat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bennamoun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. A.</given-names>
            <surname>Sohel</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Togneri</surname>
          </string-name>
          , “
          <article-title>Cost-Sensitive Learning of Deep Feature Representations From Imbalanced Data</article-title>
          .,
          <source>” IEEE Trans. neural networks Learn. Syst.</source>
          , vol.
          <volume>29</volume>
          , no.
          <issue>8</issue>
          , pp.
          <fpage>3573</fpage>
          -
          <lpage>3587</lpage>
          , Aug.
          <year>2018</year>
          , doi: 10.1109/TNNLS.
          <year>2017</year>
          .
          <volume>2732482</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>K.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Li</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Shao</surname>
          </string-name>
          , “
          <string-name>
            <surname>Improved</surname>
            <given-names>PSO</given-names>
          </string-name>
          _
          <article-title>AdaBoost ensemble algorithm for imbalanced data,” Sensors (Switzerland)</article-title>
          , vol.
          <volume>19</volume>
          , no.
          <issue>6</issue>
          ,
          <year>2019</year>
          , doi: 10.3390/s19061476.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>B.</given-names>
            <surname>Raghuwanshi</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Shukla</surname>
          </string-name>
          , “
          <article-title>Class imbalance learning using UnderBagging based Kernelized Extreme Learning Machine,”</article-title>
          <source>Neurocomputing</source>
          , vol.
          <volume>329</volume>
          , pp.
          <fpage>172</fpage>
          -
          <lpage>187</lpage>
          ,
          <year>2018</year>
          , doi: 10.1016/j.neucom.
          <year>2018</year>
          .
          <volume>10</volume>
          .056.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>S. R.</given-names>
            <surname>Mounce</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ellis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Edwards</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. L.</given-names>
            <surname>Speight</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Jakomis</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J. B.</given-names>
            <surname>Boxall</surname>
          </string-name>
          , “
          <article-title>Ensemble Decision Tree Models Using RUSBoost for Estimating Risk of Iron Failure in Drinking Water Distribution Systems</article-title>
          ,” Water Resour. Manag., vol.
          <volume>31</volume>
          , no.
          <issue>5</issue>
          , pp.
          <fpage>1575</fpage>
          -
          <lpage>1589</lpage>
          ,
          <year>2017</year>
          , doi: 10.1007/s11269-017-1595-8.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>W.</given-names>
            <surname>Feng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Huang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Ren</surname>
          </string-name>
          , “
          <article-title>Class imbalance ensemble learning based on the margin theory</article-title>
          ,
          <source>” Appl. Sci.</source>
          , vol.
          <volume>8</volume>
          , no.
          <issue>5</issue>
          ,
          <year>2018</year>
          , doi: 10.3390/app8050815.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>L.</given-names>
            <surname>Nanni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Fantozzi</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N.</given-names>
            <surname>Lazzarini</surname>
          </string-name>
          , “
          <article-title>Coupling different methods for overcoming the class imbalance problem</article-title>
          ,
          <source>” Neurocomputing</source>
          , vol.
          <volume>158</volume>
          ,
          <year>2015</year>
          , doi: 10.1016/j.neucom.
          <year>2015</year>
          .
          <volume>01</volume>
          .068.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>A.</given-names>
            <surname>Fernández</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          <article-title>del</article-title>
          <string-name>
            <surname>Río</surname>
            ,
            <given-names>N. V.</given-names>
          </string-name>
          <string-name>
            <surname>Chawla</surname>
            , and
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Herrera</surname>
          </string-name>
          , “
          <article-title>An insight into imbalanced Big Data classification: outcomes and challenges,” Complex Intell</article-title>
          . Syst., vol.
          <volume>3</volume>
          , no.
          <issue>2</issue>
          , pp.
          <fpage>105</fpage>
          -
          <lpage>120</lpage>
          ,
          <year>2017</year>
          , doi: 10.1007/s40747-017-0037-9.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>I.</given-names>
            <surname>Triguero</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          <article-title>del</article-title>
          <string-name>
            <surname>Río</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>López</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Bacardit</surname>
            ,
            <given-names>J. M.</given-names>
          </string-name>
          <string-name>
            <surname>Benítez</surname>
            , and
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Herrera</surname>
          </string-name>
          , “
          <article-title>ROSEFW-RF: The winner algorithm for the ECBDL'14 big data competition: An extremely imbalanced big data bioinformatics problem,” Knowledge-Based Syst</article-title>
          ., vol.
          <volume>87</volume>
          , pp.
          <fpage>69</fpage>
          -
          <lpage>79</lpage>
          ,
          <year>2015</year>
          , doi: https://doi.org/10.1016/j.knosys.
          <year>2015</year>
          .
          <volume>05</volume>
          .027.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>D.</given-names>
            <surname>Galpert</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          <article-title>del</article-title>
          <string-name>
            <surname>Río</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Herrera</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Ancede-Gallardo</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Antunes</surname>
            , and
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Agüero-Chapin</surname>
          </string-name>
          ,
          <article-title>“An Effective Big Data Supervised Imbalanced Classification Approach for Ortholog Detection in Related Yeast Species</article-title>
          ,”
          <source>Biomed Res. Int.</source>
          , vol.
          <year>2015</year>
          , p.
          <fpage>748681</fpage>
          ,
          <year>2015</year>
          , doi: 10.1155/
          <year>2015</year>
          /748681.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>A.</given-names>
            <surname>Haldankar</surname>
          </string-name>
          and
          <string-name>
            <given-names>K.</given-names>
            <surname>Bhowmick</surname>
          </string-name>
          , “
          <article-title>A cost sensitive classifier for Big Data</article-title>
          ,” in IEEE International Conference on Advances in Electronics, Communication and Computer Technology, Dec.
          <year>2016</year>
          , pp.
          <fpage>122</fpage>
          -
          <lpage>127</lpage>
          , doi: 10.1109/ICAECCT.
          <year>2016</year>
          .
          <volume>7942567</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>S.</given-names>
            <surname>Del Rio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Benitez</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F.</given-names>
            <surname>Herrera</surname>
          </string-name>
          , “
          <article-title>Analysis of Data Preprocessing Increasing the Oversampling Ratio for Extremely Imbalanced Big Data Classification</article-title>
          ,” in 2015 IEEE Trustcom/BigDataSE/ISPA, Aug.
          <year>2015</year>
          , pp.
          <fpage>180</fpage>
          -
          <lpage>185</lpage>
          , doi: 10.1109/Trustcom.
          <year>2015</year>
          .
          <volume>579</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>H.</given-names>
            <surname>He</surname>
          </string-name>
          and
          <string-name>
            <given-names>E. A.</given-names>
            <surname>Garcia</surname>
          </string-name>
          , “
          <article-title>Learning from Imbalanced Data,”</article-title>
          <source>IEEE Trans. Knowl. Data Eng.</source>
          , vol.
          <volume>21</volume>
          , no.
          <issue>9</issue>
          , pp.
          <fpage>1263</fpage>
          -
          <lpage>1284</lpage>
          ,
          <year>2009</year>
          , doi: 10.1109/TKDE.
          <year>2008</year>
          .
          <volume>239</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>B. E. F.</given-names>
            <surname>Sosik Heidi</surname>
          </string-name>
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Peacock</surname>
          </string-name>
          Emily E.,
          <string-name>
            <surname>“</surname>
          </string-name>
          WHOI-PLANKTON,” WHOAS:
          <string-name>
            <surname>Woods Hole Open Acess Server</surname>
          </string-name>
          ,
          <year>2014</year>
          . .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>J.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Socher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Li</surname>
          </string-name>
          , and
          <string-name>
            <surname>L.</surname>
          </string-name>
          <article-title>Fei-fei, “Imagenet: A large-scale hierarchical image database</article-title>
          ,” in In CVPR,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>C.</given-names>
            <surname>Seiffert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. M.</given-names>
            <surname>Khoshgoftaar</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Van Hulse</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Napolitano</surname>
          </string-name>
          , “
          <article-title>RUSBoost: Improving classification performance when training data is skewed</article-title>
          ,
          <source>” Proc. - Int. Conf. Pattern Recognit</source>
          ., pp.
          <fpage>8</fpage>
          -
          <lpage>11</lpage>
          ,
          <year>2008</year>
          , doi: 10.1109/icpr.
          <year>2008</year>
          .
          <volume>4761297</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>A.</given-names>
            <surname>Khan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sohail</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.</given-names>
            <surname>Zahoora</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A. S.</given-names>
            <surname>Qureshi</surname>
          </string-name>
          , “
          <article-title>A survey of the recent architectures of deep convolutional neural networks</article-title>
          ,
          <source>” Artif. Intell. Rev.</source>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>70</lpage>
          ,
          <year>2020</year>
          , doi: 10.1007/s10462-020-09825-6.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>C.</given-names>
            <surname>Seiffert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. M.</given-names>
            <surname>Khoshgoftaar</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Van Hulse</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Napolitano</surname>
          </string-name>
          , “
          <article-title>RUSBoost: A hybrid approach to alleviating class imbalance</article-title>
          ,
          <source>” IEEE Trans. Syst. Man, Cybern. Part ASystems Humans</source>
          , vol.
          <volume>40</volume>
          , no.
          <issue>1</issue>
          , pp.
          <fpage>185</fpage>
          -
          <lpage>197</lpage>
          ,
          <year>2010</year>
          , doi: 10.1109/TSMCA.
          <year>2009</year>
          .
          <volume>2029559</volume>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>