<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A p pli c ati o n -b a s e d s p a m d et e cti o n wit h m a c hi n e l e ar ni n g al g orit h m s</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>A li Er b e y</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>N e c a a tti n B arı ş çı</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <fpage>198</fpage>
      <lpage>206</lpage>
      <abstract>
        <p>1 Uş a k U ni versit y, Dist a nc e E d uc ati o n V oc ati o n al Sc h o ol , De p art me nt of C o m p uter Pr o gr a m mi n g , Uş a k, T ur k e y 2 G a zi U ni versit y , F ac ult y of Tec h n ol o g y , D e p art m e nt of C o m p uter E n gi neeri n g , A n k ar a, T ur k e y A b str a ct T o d a y, t h e us e of s o ci al m e di a sit es s u c h as F a c e b o o k, I nst a gr a m, T witt er is i n cr e asi n g d a y b y d a y. S h ar es o n s o ci al m e di a c a n e v e n c h a n g e t h e a g e n d a b y r e a c hi n g h u g e m ass es. O n T witt er, a s o ci al m e di a sit e, t h e a g e n d a t o pi cs c a n b e f oll o w e d t hr o u g h t h e s e cti o n c all e d Tr e n d -T o pi c. T his Tr e n d - T o pi c s e cti o n m a y b e m a ni p ul at e d b y s p a m m ers fr o m ti m e t o ti m e. I n or d er t o a v oi d s u c h u n w a nt e d sit u ati o ns, it is n e c ess ar y t o d et er mi n e w h et h er t h e us er is s p a m or n ot. M a c hi n e l e ar ni n g al g orit h ms c a n cl assif y w h et h er a us er is s p a m or n ot. Wit h m a c hi n e l e ar ni n g al g orit h ms, s u c c essf ul r es ults ar e als o o bt ai n e d i n sit u ati o ns s u c h as i m a g e pr o c essi n g, s p e e c h, v oi c e r e c o g niti o n a n d m al w ar e d et e ct i o n. I n t his st u d y, m a c hi n e l e ar ni n g al g orit h ms N aiv e B a y es, K N e ar est N ei g h b or s, R a n d o m F or est, j 4 8, M ultil a y er P er c e ptr o n w er e us e d t o cl assif y us ers. As a r es ult of t h e e v al u ati o ns, R a n d o m F or est al g orit h m, o n e of t h e m a c hi n e l e ar ni n g al g orit h ms us e d, m a d e t h e m ost s u c c essf ul cl assifi c ati o n wit h a n a c c ur a c y r at e of 8 8 % . K e y w or d s 1 T witt er, s p a m d et e cti o n, m a c hi n e l e ar ni n g 1. I ntr o d u cti o n I V U S 2 0 2 2: 2 7t h I nt er n ati o n al C o nf er e n c e o n I nf or m ati o n T e c h n ol o g y E M AI L: ali er b e y @ g m ail . c o m ( A. Er b e y); n b aris ci @ g a zi . ed u.tr ( N. B arış çı) O R CI D: 0 0 0 0 - 00 0 2 - 09 3 0 - 40 8 1 ( A. Er b e y); 0 0 0 0 - 00 0 2 - 87 6 2 - 50 9 1 ( N. B arış çı) © 2 0 2 2 C o p yri g ht f or t his p a per b y its a ut h ors. Use per mitte d u n der Creati ve CPWErooUrckResehdoinpgs hIStSpN:/c1e6u1r3-w-0s.o7r3g CCoEm mUo nRs LiWcoernskesAthtriobputi o nPr4.o0 cI neteer ndaitino ngasl(C(C CB YE 4.U0).R-W S. or g)</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        T o d a y, wit h t h e wi d e s pr e a d u s e of t h e i nt er n et
a n d t h e i n cr e as e i n t h e us e of m o bil e d e vi c es,
o nli n e s o ci al n et w or ki n g sit es, s o ci al n et w or ks
s u c h as F a c e b o o k, T witt er a n d Li n k e dI n ar e
b e c o mi n g m or e a n d m or e p o p ul ar [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. T h es e sit es
ar e f oll o w e d b y milli o ns of p e o pl e; I n a d diti o n t o
b ei n g sit es w h er e fri e n ds, f a mil y or a c q u ai nt a n c es
c a n b e c o nt a ct e d, t h e y ar e al s o us e d as
mi cr o bl o g gi n g s er vi c e s, r e c o m m e n d ati o n
s er vi c es, r e al -ti m e n e ws s o ur c es a n d c o nt e nt
s h ari n g pl a c es [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Us er s c a n s h ar e b y cr e ati n g
st at u s m es s a g es o n T witt er, o n e of t h es e sit es. I n
T witt er, w hi c h i s a p o p ul ar mi cr o bl o g gi n g sit e i n
t er ms of s h ari n g, t h es e st at us m es s a g es cr e at e d ar e
c all e d t w e et s [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Wit h t h es e t w e et s s e nt b y t h e
us er s, t h e Tr e n d T o pi c s e cti o n, w hi c h c o nstit ut es
t h e e xi sti n g a g e n d a t o pi cs, i s f or m e d.
      </p>
      <p>
        Tr e n d T o pi c s e cti o n c a n b e dir e ct e d t o t h e
a g e n d a i n a n u n d esir a bl e w a y wit h m es s a g es s e nt
fr o m ti m e t o ti m e, o ut of p ur p os e. T h e h e a v y us e
of s o ci al m e di a h as f a cilit at e d t h e n e gl e ct of t h es e
e n vir o n m e nt s b y m ali ci o us p e o pl e [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. W a nti n g t o
c h a n g e t h e a g e n d a i s al s o a m et h o d t h at c a n b e
n e gl e ct e d b y m ali ci o us p e o pl e. I n or d er t o pr e v e nt
s u c h o mi s si o ns, m a n y st u di es h a v e b e e n c arri e d
o ut i n ar e as s u c h as n at ur al l a n g u a g e pr o c es si n g
a n d d at a mi ni n g wit h t h e d at a c oll e ct e d fr o m
T witt er [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>W h e n w e l o o k at t h e e xisti n g st u di es i n
d et e cti n g s p a m wit h T witt er p ost s, w e s e e t h at
t h er e i s a l ot of w or k. T h es e st u di es ar e cl ust er e d
i n c ert ai n ar e as. T witt er s p a m d et e cti o n st u di es ar e
m ai nl y h a n dl e d i n t hr e e gr o u ps. T h es e ar e: a)
t h os e w h o o nl y e x a mi n e t h e t w e et s b y t e xt
mi ni n g, b) t h os e w h o a n al y z e t h e t w e et t e xt b y
as s o ci ati n g wit h t h e us er w h o s e nt t h e t w e et, c)
t h os e w h o e x a mi n e t h e r el ati o ns of us er s wit h
s p a m.</p>
      <p>
        T e xt mi ni n g -b as e d r es e ar c h m ostl y f o c us es o n
t w e et t e xt. I n t h es e st u di e s, r es e ar c h er s fir st
e xtr a ct f e at ur es a n d t h e n cl as sif y t h e m wit h
al g orit h ms s u c h as N ai v e B a y es a n d j 4 8. F e at ur e
extraction sometimes follows feature selection to
improve classification accuracy and reduce
training time. Gupta and Kumar [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] used multiple
linear regression to select important features.
Other features such as the tweet's character count,
word count, like or retweet count are used in most
research.
      </p>
      <p>Some research takes user characteristics into
account when deciding whether a user is spam.
These features can be account age, number of
followers/followers, follower/followers’ rate,
format of the profile page. However, since these
features can be easily changed by the user, they
are considered to be less reliable.</p>
      <p>Because user characteristics can change easily,
some researchers have studied the relationships
between spammers and real users. By examining
their following / following relationships, they
created a network for each user. Setting up these
networks can be costly in terms of computation
time, power and data collection time.</p>
      <p>In the following sections of this study,
obtaining the spammy dataset, classification,
detection of spammy users and feature selection
processes are carried out. In the last section, the
obtained results are evaluated.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Material and method</title>
      <p>We collect the dataset before making the
classification. The data collection process has an
important role in the classification process.
2.1.</p>
    </sec>
    <sec id="sec-3">
      <title>Spamming Twitter dataset</title>
      <p>In this study, user characteristics and the
method of evaluating users' tweet attributes were
chosen in order to classify spam. The reason for
choosing this method is that it is less costly in
terms of data collection and it is seen to give better
results regardless of the tweet content, as it
depends on user characteristics.</p>
      <p>For training, a topic was selected from the
Turkey Trend topic list, since a data set with spam
users should be obtained. 15000 tweets were
collected from this trending topic with the Twitter
public API. Then, repetitive data, news content,
tweets containing URL only were removed from
this dataset and the remaining 3798 tweets were
classified as spam and not spam. As a result, 3798
tweets were classified as 1666 spam and 2132
non-spam users.</p>
      <p>In order to evaluate which users can send spam
tweets after classification, the last maximum of
100 tweets of each user were collected. Since
some users did not have 100 tweets, a total of
305.604 tweets were reached. Then, these tweets
were processed for each user, and some
unnecessary data such as smileys were removed
from the text of the tweet. Because tweets are
unofficial texts, some autocorrect libraries were
used and typos were corrected. Then, as a more
complex process, some new features are obtained
from this tweet data, such as how often the user
tweets, the average number of characters of the
user's tweets, or the unique words tweeted in those
100 tweets, represented as columns in the final
dataset has been done.</p>
      <p>As a result of the data collection and
preprocessing stage, a dataset consisting of 3798
rows representing each user and 22 columns in
total, including features such as the age of the
user, whether he is a verified account, whether he
entered a URL on the profile page, was obtained.</p>
    </sec>
    <sec id="sec-4">
      <title>2.2. Classification</title>
      <p>The next part after data collection is to
determine whether the user is a spam user with
different machine learning algorithms.</p>
      <p>
        Weka software was used to classify the
obtained data as spam or not spam. Weka is a
program developed for machine learning and text
mining intended to assist in the application of
machine learning techniques [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>In this study, Naive Bayes algorithm, k Nearest
Neighbor algorithm, Random Forest algorithm,
j48 algorithm and Multilayer Perceptron
classification algorithms were used in Weka
software.</p>
      <p>Considering the studies in the literature, the
algorithms used in other studies are shown in
Table 1.</p>
      <p>The reason for choosing the NB, KNN, RF ,
j48 and MLP algorithms used in this study is to
try to create a combination of algorithms that are
widely used in the literature and in addition to
them, less used algorithms. The reason for
choosing the most used algorithms is to make
comparisons with previous studies. The reason for
choosing the less used algorithms is to create an
alternative to the frequently used algorithms.</p>
    </sec>
    <sec id="sec-5">
      <title>2.2.1. Naive Bayes</title>
      <p>
        The Naive Bayes (NB) algorithm is a simple
probabilistic classifier that calculates a probability
set by counting the frequency and combinations
of values in a given data set. It is a classification
algorithm that classifies data by calculating it with
probability principles. Naive Bayes is a popular
algorithm used commercially or open source for
email spam filtering [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
    </sec>
    <sec id="sec-6">
      <title>2.2.2. k Nearest Neighbors</title>
      <p>
        In 1968, Cover and Hart proposed the k
Nearest Neighbor (KNN) algorithm, which they
have been working on for a long time [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. The
intuition underlying the K Nearest Neighbor
Classification is quite simple, samples are
classified according to the class of their nearest
neighbors [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Having an efficient algorithm for
performing nearest neighbor operations on large
datasets can provide rapid improvements for
many applications [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. KNN is one of the useful
algorithms in terms of speed.
      </p>
    </sec>
    <sec id="sec-7">
      <title>2.2.3. Random Forest</title>
      <p>
        Breiman [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] developed the Random Forest
(RF) method as an extension of classification
trees. In the RF algorithm, each node has a
random feature selection [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. It is an algorithm
that aims to increase the classification value by
using more than one decision tree.
2.2.4. j48
      </p>
      <p>
        The purpose of the Decision Tree Algorithm is
to determine how the feature vector behaves for a
few samples [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. In the WEKA data mining tool,
J48 is an open-source Java implementation of the
C4.5 algorithm [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
      </p>
    </sec>
    <sec id="sec-8">
      <title>2.2.5. Multilayer Perceptron</title>
      <p>
        Multilayer perceptron (MLP) is an algorithm
that can be effectively used for classification
purposes and has been used a lot recently. In
general, back propagation algorithm learning
technique based on slope drop method is used in
MLP. With this technique, the error between the
desired output and the produced output is
minimized [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ].
      </p>
    </sec>
    <sec id="sec-9">
      <title>2.3. Determination of spam users</title>
      <p>The data obtained during the classification of
the data was divided into 80% training data and
20% test data and evaluated in NB, KNN, RB, j48
and MLP algorithms. It is known that 1666 of
3798 users are spam users and 2132 of them are
non-spam users in the dataset. The number of
users to test is 760 people, which is 20% of the
data. Looking at whether users are spam with the
Naive Bayes algorithm, the Naive Bayes
algorithm classified users with an accuracy rate of
76%.</p>
      <p>When the complexity matrix of the NB
algorithm is examined, the data are shown in
Table 2.</p>
      <p>According to the complexity matrix of the
algorithm in Table 2; Of the 760 people in the
20% test data, 382 people who were spam were
classified as spam, and 60 people who were spam
were classified as non-spam. 122 non-spam were
classified as spam, while 196 non-spam were
classified as non-spam. When we look at the KNN
algorithm, one of the machine learning
algorithms, it is seen that it classifies users with
an accuracy rate of 74%.</p>
      <p>The complexity matrix of the KNN algorithm
is as shown in Table 3.</p>
      <p>According to the complexity matrix of the
KNN algorithm in Table 3; Of the 760 people in
the 20% test data, 357 people who were spam
were classified as spam, and 85 people who were
spam were classified as non-spam. 107 non-spam
were classified as spam, while 211 non-spam were
classified as non-spam.</p>
      <p>When the Random Forest algorithm was used
in the study, an accuracy rate of 88% was
achieved.</p>
      <p>According to the RF algorithm complexity
matrix in Table 4; Of the 760 people in the 20%
test data, 406 people who were spam were
classified as spam, and 36 people who were spam
were classified as non-spam. 53 non-spam
classified as spam, 265 non-spam classified as
non-spam.</p>
      <p>85% accuracy rate was observed with the J48
algorithm. The complexity matrix of the J48
algorithm is as shown in Table 5.</p>
      <p>According to the complexity matrix of the j48
algorithm in Table 5; Of the 760 people in the
20% test data, 382 people who were spam were
classified as spam, and 60 people who were spam
were classified as non-spam. 50 non-spam
classified as spam, 268 non-spam classified as
non-spam.</p>
      <p>The MLP algorithm found an accuracy rate of
82%. The complexity matrix of the MLP
algorithm is as shown in Table 6.</p>
      <p>According to the complexity matrix of the
MLP algorithm in Table 6; Of the 760 people in
the 20% test data, 390 people who were spam
were classified as spam, and 52 people who were
spam were classified as non-spam. 82 non-spam
people were classified as spam, while 236
nonspam were classified as not spam.</p>
    </sec>
    <sec id="sec-10">
      <title>2.4. Feature selection</title>
      <p>
        Feature selection is one of the important steps
of pattern recognition, machine learning and data
mining. Its purpose is to eliminate irrelevant and
redundant variables in order to understand the
data, reduce the computational requirement,
reduce the dimensionality effect, and improve the
performance of the predictor [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. The sections
selected by the Weka software after feature
selection are shown in Table 7.
      </p>
      <p>When the data is re-evaluated after the feature
selection, the performances of the algorithms are
seen in Figure 1.</p>
      <p>As shown in Figure 1, it has obtained similar
results with feature selection and without feature
selection.</p>
    </sec>
    <sec id="sec-11">
      <title>3. Conclusion and discussion</title>
      <p>In this study, SPAM users on Twitter were
tried to be detected. In this study, the data set
consists of 3798 user information and different
machine learning algorithms are used to classify
users. The Random Forest algorithm achieved the
highest accuracy rate of 88%. The NB algorithm
achieved 76%, KNN 74%, j48 83% and MLP 80%
correct classification rates. When the algorithms
were applied again after the feature selection was
made, it was observed that the accuracy rate
decreased in other algorithms except the NB
algorithm.</p>
      <p>In addition to the accuracy rates in the
algorithms, the number of users whose real class
is negative but classified as positive in the
complexity matrix is also important. Since these
users are classified as spam even though they are
not spam, they will suffer if the algorithm is
trusted. This will create an undesirable situation.
When we look at the results, it is seen that the j48
algorithm gives the lowest rate with 50 users.</p>
      <p>As a suggestion for future research, different
optimizations of feature extraction and different
machine learning algorithms methods can be tried
on the collected data and more successful
classification results can be achieved.</p>
    </sec>
    <sec id="sec-12">
      <title>4. References</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A. H.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>"Detecting spam bots in online social networking sites: a machine learning approach,"</article-title>
          <source>in IFIP Annual Conference on Data and Applications Security and Privacy</source>
          . Springer. Berlin, Heidelberg,
          <year>2010</year>
          . pp.
          <fpage>335</fpage>
          -
          <lpage>342</lpage>
          doi:10.1007/978-3-
          <fpage>642</fpage>
          -13739-6_
          <fpage>25</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Pennacchiotti</surname>
          </string-name>
          and
          <string-name>
            <surname>A.M. Popescu,</surname>
          </string-name>
          <article-title>A machine learning approach to twitter user classification</article-title>
          .
          <source>in Fifth International AAAI Conference on Weblogs and Social Media</source>
          .
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Go</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Bhayani</surname>
          </string-name>
          and
          <string-name>
            <given-names>L.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <article-title>Twitter sentiment classification using distant supervision</article-title>
          .
          <source>CS224N Project Report</source>
          , Stanford,
          <year>2009</year>
          .
          <volume>1</volume>
          (
          <issue>12</issue>
          ): pp.
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>İ.</given-names>
            <surname>Aydın</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sevi</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.U.</given-names>
            <surname>Salur</surname>
          </string-name>
          ,
          <article-title>"Detection of Fake Twitter Accounts with Machine Learning Algorithms,"</article-title>
          <source>2018 International Conference on Artificial Intelligence and Data Processing (IDAP)</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>4</lpage>
          , doi: 10.1109/IDAP.
          <year>2018</year>
          .
          <volume>8620830</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>E.S</given-names>
            <surname>Akgül</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Ertano</surname>
          </string-name>
          and
          <string-name>
            <given-names>B.</given-names>
            <surname>Diri</surname>
          </string-name>
          ,
          <article-title>Twitter verileri ile duygu analizi</article-title>
          . Pamukkale University Journal of Engineering Sciences.
          <year>2016</year>
          , Vol.
          <volume>22</volume>
          <issue>Issue 2</issue>
          ,
          <fpage>p106</fpage>
          -
          <lpage>110</lpage>
          .
          <year>5p</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>D.K</given-names>
            <surname>Gupta</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Kumar</surname>
          </string-name>
          .
          <article-title>Spam And Sentiment Analysis Model For Twitter Data Using Statistical Learning</article-title>
          .
          <source>in Proceedings of the Third International Symposium on Computer Vision and the Internet</source>
          .
          <year>2016</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>G.</given-names>
            <surname>Holmes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Donkin</surname>
          </string-name>
          and
          <string-name>
            <given-names>I.H.</given-names>
            <surname>Witten</surname>
          </string-name>
          ,
          <string-name>
            <surname>Weka:</surname>
          </string-name>
          <article-title>A machine learning workbench</article-title>
          .
          <year>1994</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Diale</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Celik</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Van Der Walt</surname>
          </string-name>
          ,
          <article-title>Unsupervised feature learning for spam email filtering</article-title>
          .
          <source>Computers &amp; Electrical Engineering</source>
          ,
          <year>2019</year>
          . 74: pp.
          <fpage>89</fpage>
          -
          <lpage>104</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Mccord</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Chuah</surname>
          </string-name>
          .
          <article-title>Spam detection on twitter using traditional classifiers</article-title>
          .
          <source>in international conference on Autonomic and trusted computing</source>
          .
          <source>2011</source>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>T.R.</given-names>
            <surname>Patil</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Sherekar</surname>
          </string-name>
          ,
          <article-title>Performance analysis of Naive Bayes and J48 classification algorithm for data classification</article-title>
          .
          <source>International journal of computer science and applications</source>
          ,
          <year>2013</year>
          .
          <volume>6</volume>
          (
          <issue>2</issue>
          ): pp.
          <fpage>256</fpage>
          -
          <lpage>261</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>V.</given-names>
            <surname>Metsis</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Androutsopoulos</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Paliouras</surname>
          </string-name>
          .
          <article-title>Spam filtering with naive bayeswhich naive bayes? in CEAS</article-title>
          .
          <year>2006</year>
          . Mountain View, CA.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>A.</given-names>
            <surname>Kataria</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <article-title>A review of data classification using k- nearest neighbour algorithm</article-title>
          .
          <source>International Journal of Emerging Technology and Advanced Engineering</source>
          ,
          <year>2013</year>
          .
          <volume>3</volume>
          (
          <issue>6</issue>
          ): pp.
          <fpage>354</fpage>
          -
          <lpage>360</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>P.</given-names>
            <surname>Cunningham</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.J.</given-names>
            <surname>Delany</surname>
          </string-name>
          ,
          <article-title>k-Nearest neighbour classifiers</article-title>
          .
          <source>Multiple Classifier Systems</source>
          ,
          <year>2007</year>
          .
          <volume>34</volume>
          (
          <issue>8</issue>
          ): pp.
          <fpage>1</fpage>
          -
          <lpage>17</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>M.Mujaand D.G.Lowe</surname>
          </string-name>
          ,
          <article-title>Scalable nearest neighbor algorithms for high dimensional data</article-title>
          .
          <source>IEEE transactions on pattern analysis and machine intelligence</source>
          ,
          <year>2014</year>
          .
          <volume>36</volume>
          (
          <issue>11</issue>
          ): pp.
          <fpage>2227</fpage>
          -
          <lpage>2240</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>L.</given-names>
            <surname>Breiman</surname>
          </string-name>
          ,
          <article-title>Random forests</article-title>
          .
          <source>Machine learning</source>
          ,
          <year>2001</year>
          .
          <volume>45</volume>
          (
          <issue>1</issue>
          ): pp.
          <fpage>5</fpage>
          -
          <lpage>32</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>K.J. Archer</surname>
            , and
            <given-names>R.V.</given-names>
          </string-name>
          <string-name>
            <surname>Kimes</surname>
          </string-name>
          ,
          <article-title>Empirical characterization of random forest variable importance measures</article-title>
          .
          <source>Computational Statistics &amp; Data Analysis</source>
          ,
          <year>2008</year>
          .
          <volume>52</volume>
          (
          <issue>4</issue>
          ): pp.
          <fpage>2249</fpage>
          -
          <lpage>2260</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>G.</given-names>
            <surname>Kaur</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Chhabra</surname>
          </string-name>
          ,
          <article-title>Improved J48 classification algorithm for the prediction of diabetes</article-title>
          .
          <source>International Journal of Computer Applications</source>
          ,
          <year>2014</year>
          .
          <volume>98</volume>
          (
          <issue>22</issue>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>A.R.</given-names>
            <surname>Yılmaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Yavuz</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Erkmen</surname>
          </string-name>
          .
          <article-title>Training multilayer perceptron using differential evolution algorithm for signature recognition application</article-title>
          .
          <source>in 2013 21st Signal Processing and Communications Applications Conference (SIU)</source>
          .
          <year>2013</year>
          . IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>F.</given-names>
            <surname>Asdaghi</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Soleimani</surname>
          </string-name>
          ,
          <article-title>An effective feature selection method for web spam</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>