<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards Multi-Class Smishing Detection: A Novel Feature Vector Approach and the Smishing-4C Dataset</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alicia Martínez-Mendoza</string-name>
          <email>alicia.martinez@unileon.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francisco Jáñez-Martino</string-name>
          <email>francisco.janez@unileon.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrés Carofilis</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Laura Fernández-Robles</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Enrique Alegre</string-name>
          <email>enrique.alegre@unileon.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Eduardo Fidalgo</string-name>
          <email>eduardo.fidalgo@unileon.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Electrical, Systems and Automation Engineering, Universidad de León</institution>
          ,
          <addr-line>ES</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Researcher at INCIBE (Spanish National Cybersecurity Institute)</institution>
          ,
          <addr-line>León, ES</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>Every day, more mobile phones are being hit by smishing, the phishing messages that we receive via Short Message Services. The diferent smishing messages could be classified according to the type of fraud, which could help in identifying target entities, specific victims, and even in detecting campaigns. Multi-class classification of smishing is still largely unexplored in the research community. Therefore, in this paper, we propose a feature vector to describe smishing messages that helps to distinguish them from regular messages. Our proposal presents six features: length value, number of spelling errors, phone, URL, slang, and company name. To demonstrate the discriminative capacity to classify smishing messages into diferent categories, we created Smishing-4C, a dataset built using samples from Kaggle and Mendeley datasets labeled in four types of smishing: Bank/Finance, Rewards, Dating, and Short Message Service. Using Smishing-4C, we trained several Machine and Deep Learning models to establish baseline results to detect diferent types of smishing using short text classification methods. We found that, in Smishing-4C, the combination of Bag of Words and the proposed 6-feature vector obtains an F1 score of 0.788, outperforming transformer-based models.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;text classification</kwd>
        <kwd>smishing classification</kwd>
        <kwd>multiclass classification</kwd>
        <kwd>Smishing-4C dataset</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Smishing describes a phishing technique in which an
attacker uses the Short Message Service (SMS) as the
medium to deliver a phishing attack. The objective of
the phisher is to either obtain personal information or
credentials from the user, or to distribute malware,
usually aiming to obtain a financial benefit [ 1]. Phishers
exploit social engineering techniques to mislead victims
into believing that the SMS comes from a trusted source.</p>
      <p>This message usually requests the victim performs an
action such as clicking on a link, calling a phone
number, or sending an email [1]. The link will redirect the
user to a fake login website where users introduce their
credentials, which will then be sent to the attacker [2].</p>
      <p>Smishing attacks have experienced remarkable growth
in the last years, with 500 million smishing messages
reported in 2023 [3], which supposes a global financial
cost of $800 per person, according to Carnegie Mellon</p>
      <p>University [4]. Due to the impact of these cyberattacks,
recently, researchers have developed smishing detection
models [5, 6, 7]. Many authors have used traditional
machine learning models and have considered a variety
of smishing features, such as the length of the message
or the appearance of an URL [8]. Nevertheless, the vast
majority of authors only focus on the task of smishing
detection and do not classify it into diferent types of scams.</p>
      <p>A Computer Emergency Response Team (CERT) receives
reports of smishing1 that need to be processed and cate- learning and deep learning methods have been applied,
gorised, to identify the target of the attack. A model based and some authors have adopted additional stages in the
on multi-class smishing classification would automate detection process, such as using regular expressions to
the classification of these reports, enabling faster report iflter messages containing keywords commonly found
processing and early response to attacks. Consequently, in smishing [10], or adding a phishing URL classification
messages could be grouped based on the type of scam, step [11]. In addition, Akande et al. [5] have gone beyond
allowing for the identification of the targeted entity or proposing smishing detection methods and have
develuser profile and helping in the detection of campaigns. oped a mobile application capable of detecting incoming
This would facilitate tasks such as informing targets to SMS smishing messages on a smartphone.
reduce the number of future victims and it would be use- Numerous authors have studied the nature of smishing
ful for identifying the objective of the smisher, such as messages to identify unique characteristics that
diferentiattacking a bank account, receiving a payment or access- ate them from legitimate messages. Mishra and Soni [12]
ing personal information through a social media account. presented a smishing dataset built from samples from
To the best of our knowledge, only Zhang et al. [9] have the Almeida SMS collection [13] and Pinterest smishing
performed classification into 14 diferent types of smish- screenshots [14]. The authors identified the presence of
ing, which they group under three categories, i.e., illegal phone numbers, email addresses, and URLs as relevant
promotion, fraud, and advertisement, while the rest of features, as these are the channels through which the
atthe authors focus on smishing detection. tacker expects the victim to send their personal
informa</p>
      <p>Aiming to progress in the task of smishing classifica- tion or credentials, as supported by the study performed
tion, we propose the following contributions: by Timko et al. [15]. In a later work, Mishra and Soni
[8] used a vector with five smishing features (misspelled
• Smishing-4C, a text-based smishing dataset con- words, leet words, symbols, special characters, and
smishtaining 120 smishing samples labeled in four cat- ing keywords) with machine learning methods, obtaining
egories of smishing: Bank/Finance, Rewards, Dat- an accuracy score of 97.93% using a Backpropagation
ing, and SMS service. The combination of these approach for smishing detection. Other features have
four categories is used for the first time for the been proposed by [16]. On their Smishtank website [17],
task of multi-class smishing classification. they extract the most important features of each
submit• A novel feature vector with six features based on ted smishing sample: URL, named entities, Virus Total
text, which has not been used before in smishing score, and domain history.
detection or classification, comprising the length The use of feature vectors has also been proven
adof the SMS, number of writing errors, phone, URL, vantageous in the work of Sonowal [18]. In this case,
slang, and the company name. the authors used a combination of BOW representation
• The results obtained for Smishing-4C with four and a vector containing 13 features: size, email, URL,
machine learning and four deep learning mod- phone, number of the alphabet, special characters,
misels, demonstrating that the proposed 6-feature spellings, readability, and number of uppercase
charvector, combined with a BOW representation, ob- acters, digits, spaces, punctuation marks and parts of
tains the highest F1-score in multi-class smishing speech. The performance of their method for smishing
classification on the Smishing-4C dataset. detection achieved 98.40% accuracy on the Almeida SMS
dataset. In contrast with BOW representation, Awumee</p>
      <p>This paper is organized as follows. Section 2 presents et al. [19] evaluated the performance of machine
learnthe literature review of smishing classification meth- ing models, such as Logistic Regression (LR), Decision
ods. After that, Section 3 describes the creation of the Tree (DT), Support Vector Machines (SVM), and Random
Smishing-4C dataset, the selection of the smishing fea- Forest (RF) with Term Frequency-Inverse Document
Fretures, and the evaluation method. The details of the quency (TF-IDF) vectorization. The best result, 99.47%
experimentation are included in Section 4, and the re- accuracy, was obtained with RF.
sults are presented in Section 5 and discussed in Section 6. Lee et al. [6] proposed a method for multilingual
smishFinally, Section 7 presents the conclusions of this work. ing detection, capable of detecting smishing in English
and Korean messages. Additionally, they included in
their work not only the use of text-processing models
2. Related work but also image processing. Also using a Korean dataset,
Previous research on smishing detection reveals that Seo et al. [7] presented a lightweight on-device classifier
most authors have studied the classification into legiti- resistant to text-evasion attacks.
mate and fraudulent messages. For this purpose, machine Although many of the aforementioned works focused
on traditional methods, other researchers, such as
Mam1https://www.incibe.es/ciudadania/ayuda/reporte-de-fraude bina et al. [20], compared the performance of five deep
3. Methodology
learning models, namely Convolutional Neural Networks
(CNN), Long Short-Term Memory (LSTM), Gated
Recurrent Units (GRU), BiLSTM and BERT, obtainig a 3.1. Smishing-4C creation
98.38 accuracy score on the Kaggle Smishing dataset
[21].Transformer-based models often underperform on We create the Smishing 4 Classes (Smishing-4C) dataset
short unstructured texts, due to the dificulty in obtain- taking and labeling samples from two publicly available
ing suficient context, which also applies for texts shorter smishing datasets. Kaggle Smishing [21] and Mendeley
than SMS, such as file names [ 22]. The use of trans- Smishing [26], which contain samples of SMS in English,
former models also appears in the work of Ghourabi and labeled for smishing binary classification (legitimate or
Alohaly [23], who proposed using GPT-3 transformer smishing messages). This dataset is made publicly
availembeddings and ensemble learning, reaching and a 99.91 able for multi-class classification 2.
accuracy score on the Almeida SMS dataset. In fact, Karl We select only the smishing samples and label them
and Scherp [24] demonstrated that transformer-based for the four types of smishing that we identified as the
models can achieve a performance comparable to that of most abundant in our dataset and are referenced in
premodels designed specifically for short text, reaching the vious work related to smishing [9, 25, 15]: Bank/Finance,
highest value of 99.88 accuracy score with ERNIE on the Rewards, Dating, SMS service.</p>
      <p>Short Texts of Products and Services dataset. The dataset was labeled manually for smishing classes</p>
      <p>Previous work related to smishing has focused on by three annotators. The interannotator agreement on
smishing detection. To the best of our knowledge, only 50 samples showed a Fleiss’s Kappa coeficient of 0.671.
Zhang et al. [9] have carried out multi-class classification The classes were redefined to avoid overlaps between
of smishing messages. They performed agglomerative them. The definitions of the classes are given below:
clustering on the Fake Base Station (FBS) dataset and Bank/Finance: messages coming from a bank or
fiidentified four main types of messages: illegal promo- nancial services entity. The topic of the message
mentions, fraud, advertisement, and others, plus 14 subcat- tions a bank online account, credit card, transactions,
egories. After clustering, they used the categories for taxes or other financial operations.
classification. However, the FBS dataset is created from SMS Service: messages indicating that the user has
SMS in Chinese and it has not been validated whether unread messages or new voicemails, messages regarding
this classes are present in English datasets. subscriptions to a service or online account or customer</p>
      <p>Authors who worked in smishing detection have also service announcements, messages from internet service
identified diferent types of smishing, although they providers.
have not used them for classification. For example, after Dating: messages related to dating, friendship,
peranalysing the data they retrieved from Twitter reports sonal relationships, secret admirer messages or sexual
of smishing, Tang et al. [25] proposed a division into content.
eight smishing categories: Account alert, Finance, Prize, Rewards: messages indicating that the user has
reDelivery, Credit card, Tax fraud, COVID-19, and Others. ceived an award, free product or prize.</p>
      <p>In addition, the Smishtank reports have been analyzed Furthermore, we consider this division is also relevant
by Timko and Rahman [16], who have recognized 10 cat- for identifying the sender’s profile, as each category can
egories of smishing: Account alert, Prize/Contest, Scams be linked to a type of entity, such as banking institutions,
(undelivered package), Payday loan/credit, Wrong num- e-commerce brands, dating services, and internet
serber/romance, Job advertisements, Link only messages, vice providers. Table 1 shows examples of each type of
Finance/crypto, Lawsuits/settlement, Advertisement. smishing.</p>
      <p>It can be observed that some types of fraud, partic- The smishing-4C dataset contains 30 manually labeled
ularly those related to finance, deliveries, and prizes, samples for each smishing category and 120 samples.
are common across these works. The advantage of the Even if it is a small number of labeled samples, we
conSmishing-4C dataset is that it is labeled for the most com- sider that it is suficient to develop a proof of concept
mon types of fraud, which has not been done before in model [27, 28].</p>
      <p>English smishing datasets. Additionally, these categories
are closely related to diferent types of senders: bank- 3.2. Smishing features
ing entities, delivery service companies and e-commerce
brands.</p>
      <sec id="sec-1-1">
        <title>In Section 2, we mention the characteristics that other</title>
        <p>authors identified as smishing features. We select those
features that are common to several papers, considering
them to be the most significant and general to any
smishing dataset [12, 8, 16, 18]. Although they have been used
in previous work on smishing detection, the smishing
features have only been evaluated in the binary
classification case, and have not yet been tested for the multiclass
classification task. We propose a combination of features
specifically designed for the classification of diferent
categories of smishing. The experiments described in this
paper demonstrate that smishing features, in addition 0 50 100 150 200 250
to being applicable to binary classification, enable the Length value
difeFreeanttuiaretisotnhaotf aspmpiesahrinognlcylainssoesn.e paper may be specific BaRnkew/Fainrdasnce SMDSasteinrvgice
to a particular dataset. Thus, we select the following
features: Length value, Number of writing errors, Phone, Figure 2: Length of the messages in each smishing class. Most
and URL. Additionally, we consider other important fea- messages contain between 100 and 150 characters, although
tures for smishing detection because they can help to they tend to be shorter for SMS service than for the rest of
distinguish between smishing classes: Company names the classes.
and Slang. Company name was mentioned as an
important characteristic by [29] and will difer depending on SMS service 39
the class (Bank/Finance messages will probably include
names of bank entities, while SMS service will likely
contain names of internet service providers). Slang is related Rewards 66
to the type of text the attacker aims to create in each type
of scam, which can be more formal for Bank/Finance and Dating 38
more informal for Rewards. The description and analysis
of the features are shown below.</p>
        <p>Length value: number of characters in the SMS. Fig- Bank/Finance 49
ure 2 represents the length of the messages for each class. 0 20 40 60</p>
        <p>Number of writing errors: number of writing errors
found in the SMS. It is common to observe a high number Total number of writing errors
of writing errors in Rewards messages, where abbrevia- Figure 3: Number of writing errors in each smishing class.
tions, slang, and misspellings often appear. The number Rewards is the class with the highest number of errors.
of writing errors per class can be observed in Figure 3.</p>
        <p>Phone: indicates whether the message contains a
phone number. This feature was automatically extracted
using regular expressions, by selecting numbers between Company name: name of a known entity or
com5 and 11 digits. Figure 4 presents the number of messages pany, for example, bank entities, delivery companies or
in Smishing-4C that contain phone numbers. internet service providers. This feature was manually</p>
        <p>URL: indicates whether the message contains a labeled by an annotator, and it contains the string
corURL, making a distinction between long and shortened responding to a company name. We consider that this
URL. A shortened URL is formed by a shortened do- information is relevant for the task of smishing detection
main (bit.ly, tinyurl.com, ow.ly, t.co) and an identifier and classification. As shown in Figure 6, each company
(https://tinyurl.com/m3q2xt). The proportion of URL and usually corresponds to a specific smishing class. In
addishortened URL is shown in Figure 5. tion, having the information about the company name
as a string could aid in other smishing-related tasks such
as campaign detection.</p>
        <p>Slang: indicates whether the SMS contains slang, also
known as internet language or abbreviations commonly
found in short texts. These are common in informal
SMS because of the limited length of the messages. For
example, slang expressions are “U” (=you), “4” (=for), or
“2” (=to). If the SMS contains slang, this features is labeled
as “YES”, and if it does not contain slang, it is labeled as
“NO”. As shown in Figure 7, it is unusual to find this in
messages from Bank/Finance, as attackers aim to recreate
the formal language used by this type of entities.</p>
        <p>A summary of the features, their possible values, and
23</p>
        <p>28
26
30
an example is shown in Table 2.</p>
        <p>After this analysis, we select 6 features that we use
as a feature vector for multi-class classification: Length
value, Number of writing errors, Phone, URL, Slang,
and Company name. We hypothesize that adding this
feature vector to another feature representation such as
BOW or N-grams will improve the performance of
multiclass smishing classification because the feature vector
contains information that is not represented by BOW
4. Experimentation
or N-grams and need to be extracted from the text in a
diferent way. BOW and N-grams are techniques used to
analyse the vocabulary and writing style of a text. BOW For this experimentation, we use the Scikit-learn3 library
can identify frequently occurring words in a particular for SVM, LR, RF, and DT models, and maintain the default
class, while N-grams can provide insight into the writing parameters for an initial result. After that, we identify
style and the frequency appearance of word sets in a class. the model that provides the best F1-score, which is in our
Therefore, the vector adds relevant characteristics of the case LR, and use Grid Search to determine the optimal
messages that would aid classiefirs distinguish between hyperparameter settings for this model. We obtain that
the smishing classes. We expect that the combination of the optimal hyperparameters are multi_class=
‘multinoboth will yield a higher performance. mial’, penalty= None, and solver= ‘saga’. Then, we test if
this setting improves the performance of the model.
3.3. Evaluation method In relation to the BOW representation, we used the
Scikit-learn CountVectorizer function to retrieve the term
To validate the efectiveness of the proposed 6-feature frequency. The text was lowercased before tokenization,
vector, we selected four machine learning models and the resulting dictionary has a size of 1077 tokens.
(SVM [30], LR [31], RF [32], DT [33], [19]) and compared Regarding the deep learning models, we used the
patheir performance with four state-of-the-art deep learn- rameters of Karl and Scherp [24] for fine-tuning. The
ing models (MLP [34], LSTM [35], BERT [36], ERNIE [37]) learning rate for MLP is set to 1 · 10− 3, and to 2 · 10− 3 for
typically used in short text classification [24]. LSTM. Both are trained for 100 epochs. BERT’s learning</p>
        <p>As input for the machine learning models, we use two rate is set to 5 · 10− 5 and trained for 10 epochs. Finally,
diferent representations: BOW (as in [ 18]) and N-grams. ERNIE uses a learning rate of 25 · 10− 6 and is trained for
Then, we concatenate our proposed feature vector to each 3 epochs.
of these representations and test its efect on performance. For all models, we evaluate performance using the
In addition, we add the comparison with the method precision, recall, and F1-scores and we use 5-fold
crossproposed by Zhang et al. [9]. Although we do not use the validation.
same dataset, we use their proposed method (N-grams
with TF-IDF) because, unlike other authors in related
work, they perform multi-class smishing classification. 5. Results</p>
      </sec>
      <sec id="sec-1-2">
        <title>In this section, we evaluate the performance of diferent</title>
        <p>classification methods: the method proposed by Zhang
et al. [9], transformer models, and our proposal adding
the 6-feature vector. The performance results are shown
in Table 3.</p>
        <p>First, we present in the first rows of Table 3 the
performance of the traditional models using only the 6-feature
vector. The optimal F1-score is 0.452, obtained with RF,
which indicates that the features, when considered in
isolation, are insuficient for the models to make
accurate predictions. We include the performance per class
in Table 4 to highlight that the features can however be
helpful in the classification of Bank/Finance samples.</p>
        <p>Then, considering the state-of-the-art performance of
transformer models for text classification tasks and their
suitability for short text classification [ 24], we present in
Table 3 their performance on Smishing-4C. The highest
F1-score, 0.701, is obtained with ERNIE.</p>
        <p>After that, Table 3 shows the results of the method
proposed by Zhang et al. [9], using word unigrams and
bi-grams with TF-IDF for multi-class classification of
smishing categories. In Smishing-4C, the best result is a
0.682 F1-score, obtained with LR.</p>
        <p>Next, Table 3 shows the performance of four
traditional Machine Learning classifiers (SVM, LR, RF, and
DT) on Smishing-4C, combined with a BOW
representation standalone and, after that, concatenating it with our
proposed 6-feature vector. In both instances, LR achieves
the highest performance. For BOW, we obtain a 0.691
F1score, which is lower than the result for ERNIE. However,
the addition of the proposed 6-feature vector boosts the
F1-Score of BOW to 0.763, a value higher than the ones
obtained with the transformer models. This suggests
that the feature vector contains relevant information to
distinguish between the smishing classes of Smishing-4C.</p>
        <p>Finally, Table 3 presents the performance obtained
using unigrams and 4-grams. Although the highest score
is a 0.677 F1-score, lower than ERNIE’s performance,
the addition of the feature vector increases again the
performance for all the models.</p>
        <p>In tables 6 and 5, we can observe the performance per
class for the best transformer model (ERNIE) and the best
traditional model (LR). SMS service is the class with the
lowest F1-score in both instances, while Bank/Finance
and Dating obtain higher performance. These results are
also reflected in the confusion matrices of Figures 8 and
9.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>6. Discussion</title>
      <sec id="sec-2-1">
        <title>The results presented in Section 5 report the performance</title>
        <p>of diferent models trained on Smishing-4C dataset,
which was labeled into four smishing categories. This
dataset is used to validate our proposal and test the
hypothesis that the proposed smishing features help to
distinguish between diferent types of smishing.</p>
      </sec>
      <sec id="sec-2-2">
        <title>In the first place, the performance results of the 6</title>
        <p>feature vector indicate that the features do not
contain suficient information for multi-class classification.</p>
        <p>However, upon examination of the performance per
class in Table 4, we observe favorable results for the
Bank/Finance class. As we could infer from the feature
analysis in Section 3.2, Bank/Finance is the class which is
easier to diferentiate from the rest by using features, and
this intuition is strengthened by these results. On the
contrary, the class Dating is more challenging to predict
using only the feature vector. For this reason, we need
to add the information from the text, and combine the
6-feature vector with the word representations (N-grams,
BOW or TF-IDF). We conclude that the combination of
Table 5 tion used by Zhang et al. [9], we observe that the
perforPerformance per class for LR model with BOW + 6 features. mance of this proposed method is lower on Smishing-4C
Label Precision Recall F1-score than on the FBS dataset. Nevertheless, we included a
Bank/Finance 0.754 0.667 0.705 comparison with this method because, to the best of our
Dating 0.867 0.867 0.867 knowledge, it is the only one in the related work which
Rewards 0.667 0.667 0.667 applies to multi-class smishing classification.
SMS service 0.537 0.600 0.565 About the deep learning models, transformer models
typically work better on bigger datasets and
moderatelength texts that allow them to extract context and
both the feature vector and the word representation re- achieve a higher performance in text classification [ 38].
trieves superior results to those obtained when using However, other approaches, such as Sentence-BERT [39]
them separately. can extract context from short texts. In this work, we
con</p>
        <p>Regarding the comparison with the TF-IDF representa- sider the findings of Karl and Scherp [24] and evaluate
deep learning models on the Smishing-4C dataset, leav- that Bank/Finance is misclassified as belonging to other
ing the application of Sentence-BERT for future research. classes more often than samples from other classes
We observe that BERT and ERNIE achieve a higher per- are predicted as Bank/Finance. This may be due to
formance than MLP, LSTM, and TF-IDF (up to 3 points other classes having features diferent to the ones in
higher). Bank/Finance. For instance, concerning the feature Slang,</p>
        <p>However, transformer models are outperformed by it is noted that Bank/Finance rarely includes slang.
Howmachine learning models when using BOW and the 6- ever, the frequency of slang in the other classes is similar,
feature vector on Smishing-4C. This might change when indicating that the classifier will have an easier time
disthe Smishing-4C dataset increases its size. For that rea- tinguishing Bank/Finance from the other classes, but will
son, in future works, we will add more samples from the not perform well distinguishing the other three classes.
Kaggle and Mendeley datasets, and we will test again
the performance of these models. Then, we will
compare it with our feature vector proposal, considering the 7. Conclusions
possibility of combining the vector with transformers.</p>
        <p>As seen for BOW and (1,4)-grams, the 6-feature vector In this paper, we have presented Smishing-4C, a dataset
always enhances the performance. For LR and BOW, labeled for 4 classes of smishing: Bank/Finance, Rewards,
performance increases 5 points, for LR and N-grams, 2 Dating, and SMS service. We have evaluated four
tradipoints. In both cases, LR is the classifier that achieves the tional models (SVM, LR, RF, and DT) and four deep
learnbest performance. Furthermore, we observe that BOW ing models (MLP, LSTM, ERNIE, and BERT) on
Smishingis superior to N-grams. This could be attributed to the 4C. To increase the performance of the traditional models,
nature of the data, since they comprise unstructured text, we have proposed a vector with 6 features (Length value,
it may be easy to identify common keywords within a Number of writing errors, Phone, URL, Slang, and
Comclass (represented by BOW), while it may be less common pany name), and we have evaluated its combination with
to find a specific set of words that appear in the same BOW and (1,4)-gram representation. We have observed
order every time (represented by N-grams). that this feature vector contributes positively to the
multi</p>
        <p>Overall, the models obtain higher precision values than class classification of four smishing categories. Results
recall, indicating a lower number of false positives. The show an F1-score of 0.701 with ERNIE, 0.691 with LR and
TF-IDF and N-grams representations show the greatest BOW, which increases to 0.764 when adding the 6-feature
diference between these metrics (up to 10 points), while vector. Therefore, we prove that feature vectors are not
this diference is less pronounced in deep learning mod- only useful for binary smishing classification, as it has
els (up to 4 points) and in BOW, especially in BOW + 6 been studied in previous work, but it can also be used in
features (up to 3 points). Additionally, we note that for multi-class classification of smishing.
BOW, the most significant diference between these met- In future work, we will increase the size of the
rics appears for DT. The behavior of this model may be Smishing-4C dataset by including samples from publicly
discriminating a class that is easier to distinguish from the available smishing datasets. This will allow us to test the
rest. For instance, we have observed that Bank/Finance performance of the models using the combination of the
is more easily distinguishable based on features such as 6-feature vector and BOW representation when trained
Company name or Slang. LR might perform better due on a higher number of samples.
to the linear relationship between classes and certain
features such as Length value, BOW representation, and Acknowledgments
the correlation between Slang and Company names and
the class Bank/Finance. This work has been funded by the Recovery,
Transfor</p>
        <p>The comparison of the performance per class for mation, and Resilience Plan, financed by the European
ERNIE and LR indicates that SMS service is the class Union (Next Generation), thanks to the LUCIA project
with the lowest F1-score for both models. This denotes (Fight against Cybercrime by applying Artificial
Intellithat SMS service might be more dificult to distinguish gence) granted by INCIBE to the University of León.
from other categories because of the definition of this
class. The content of the SMS service is more diverse
than that of the other classes and may partially overlap References
with them. While the Bank/Finance category always
pertains to activities related to bank accounts or finance, the [1] S. Mishra, D. Soni, Implementation of ‘Smishing
Rewards category exclusively discusses prizes, and the Detector’: an eficient model for smishing
detecDating category is solely focused on dating content. tion using neural network, SN Computer Science 3</p>
        <p>In addition, the classes in which precision is higher (2022) 189.
than recall, such as the case of Bank/Finance, indicate
[2] D. Amato, How a Simple Text Message Can [16] D. Timko, M. L. Rahman, Commercial anti-smishing
Lead to Fraud, https://www.rbcroyalbank.com/ tools and their comparative efectiveness against
en-ca/my-money-matters/money-academy/ modern threats, in: Proceedings of the 16th ACM
cyber-security/cyber-security-for-business/ Conference on Security and Privacy in Wireless
how-a-simple-text-message-can-lead-to-fraud/, and Mobile Networks, 2023, pp. 1–12.
2024. [17] M. Lutfor Rahman, D. Timko, Smishtank, https://
[3] V. Dovgopoliuk, 90+ Smishing Statistics: Phish- smishtank.com/, 2023.</p>
        <p>ing, SMS &amp; Cybercrime, https://marketsplash.com/ [18] G. Sonowal, Detecting phishing SMS based on
mulsmishing-statistics/, 2023. tiple correlation algorithms, SN computer science
[4] Carnegie Mellon University, Stay Alert For Fraud- 1 (2020) 361.</p>
        <p>ulent Text Messages, https://www.cmu.edu/iso/ [19] G. S. Awumee, J. O. Agyemang, S. S. Boakye, D.
Benews/2024/smishing-news-article1.html, 2024. mpong, SmishShield: A Machine Learning-Based
[5] O. N. Akande, O. Gbenle, O. C. Abikoye, R. G. Jimoh, Smishing Detection System, in: International
ConH. B. Akande, A. O. Balogun, A. Fatokun, SMSPRO- ference on Wireless Intelligent and Distributed
EnTECT: An automatic smishing detection mobile ap- vironment for Communication, Springer, 2023, pp.
plication, ICT Express 9 (2023) 168–176. 205–221.
[6] H. Lee, S. Jeong, S. Cho, E. Choi, Visualization [20] I. S. Mambina, J. D. Ndibwile, D. Uwimpuhwe, K. F.</p>
        <p>Technology and Deep-Learning for Multilingual Michael, Uncovering SMS Spam in Swahili Text
Spam Message Detection, Electronics 12 (2023) 582. Using Deep Learning Approaches, IEEE Access
[7] J. W. Seo, J. S. Lee, H. Kim, J. Lee, S. Han, J. Cho, (2024).</p>
        <p>C.-H. Lee, On-Device Smishing Classifier Resistant [21] Kaggle, SMS Smishing collection dataset,
to Text Evasion Attack, IEEE Access (2024). https://www.kaggle.com/datasets/galactus007/
[8] S. Mishra, D. Soni, DSmishSMS-A System to Detect sms-smishing-collection-data-set, 2022.</p>
        <p>Smishing SMS, Neural Computing and Applications [22] M. W. Al-Nabki, E. Fidalgo, E. Alegre, R.
Alaiz35 (2023) 4975–4992. Rodriguez, Short text classification approach to
[9] Y. Zhang, B. Liu, C. Lu, Z. Li, H. Duan, S. Hao, M. Liu, identify child sexual exploitation material,
ScienY. Liu, D. Wang, Q. Li, Lies in the Air: Characteriz- tific Reports 13 (2023) 16108.
ing Fake-base-station Spam Ecosystem in China, in: [23] A. Ghourabi, M. Alohaly, Enhancing Spam Message
Proceedings of the 2020 ACM SIGSAC Conference Classification and Detection Using
Transformeron Computer and Communications Security, 2020, Based Embedding and Ensemble Learning, Sensors
pp. 521–534. 23 (2023) 3861.
[10] A. Sharaf, V. Pathak, S. S. Paul, Deep learning- [24] F. Karl, A. Scherp, Transformers are Short-Text
Clasbased smishing message identification using regular sifiers, in: International Cross-Domain Conference
expression feature generation, Expert Systems 40 for Machine Learning and Knowledge Extraction,
(2023) e13153. Springer, 2023, pp. 103–122.
[11] A. K. Jain, B. B. Gupta, K. Kaur, P. Bhutani, W. Al- [25] S. Tang, X. Mi, Y. Li, X. Wang, K. Chen, Clues
halabi, A. Almomani, A content and URL analysis- in tweets: Twitter-guided discovery and analysis
based eficient approach to detect smishing SMS in of SMS spam, in: Proceedings of the 2022 ACM
intelligent systems, International Journal of Intelli- SIGSAC Conference on Computer and
Communigent Systems 37 (2022) 11117–11141. cations Security, 2022, pp. 2751–2764.
[12] S. Mishra, D. Soni, SMS Phishing Dataset for Ma- [26] S. Mishra, D. Soni, SMS Phihsing Dataset for
Machine Learning and Pattern Recognition, in: Inter- chine Learning and Pattern Recognition, DOI: 10.
national Conference on Soft Computing and Pattern 17632/f45bkkt8pr.1, 2022.</p>
        <p>Recognition, Springer, 2022, pp. 597–604. [27] T. Stupak, How Much Data Is Required To Train
[13] T. A. Almeida, J. M. G. Hidalgo, A. Yamakami, Con- ML Models in 2024?, https://www.akkio.com/post/
tributions to the study of SMS spam filtering: new how-much-data-is-required-to-train-ml, 2023.
collection and results, in: Proceedings of the 11th [28] D. Rajput, W.-J. Wang, C.-C. Chen, Evaluation of
ACM symposium on Document engineering, 2011, a decided sample size in machine learning
applicapp. 259–262. tions, BMC bioinformatics 24 (2023) 48.
[14] Pinterest, Smishing Dataset, https://in.pinterest. [29] D. Timko, M. L. Rahman, Smishing dataset i:
Phishcom/seceduau/smishing-dataset/?lp=true, 2023. ing sms dataset from smishtank.com, arXiv preprint
[15] D. Timko, D. H. Castillo, M. L. Rahman, More Than arXiv:2402.18430 (2024).</p>
        <p>50% Of The Time, Users Detect Real SMS as Fake: A [30] C. Cortes, V. Vapnik, Support-vector networks,
Smishing Detection Study Of US Population, arXiv Machine learning 20 (1995) 273–297.
preprint arXiv:2311.06911 (2023). [31] D. R. Cox, The regression analysis of binary
sequences, Journal of the Royal Statistical Society</p>
        <p>Series B: Statistical Methodology 20 (1958) 215–232.
[32] L. Breiman, Random forests, Machine learning 45</p>
        <p>(2001) 5–32.
[33] B. De Ville, Decision trees, Wiley Interdisciplinary</p>
        <p>Reviews: Computational Statistics 5 (2013) 448–455.
[34] L. Galke, A. Scherp, Bag-of-words vs. graph vs.</p>
        <p>sequence in text classification: questioning the
necessity of text-graphs and the surprising strength
of a wide MLP, arXiv preprint arXiv:2109.03777
(2021).
[35] S. Hochreiter, J. Schmidhuber, Long short-term</p>
        <p>memory, Neural computation 9 (1997) 1735–1780.
[36] J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, BERT:</p>
        <p>Pre-training of deep bidirectional transformers for
language understanding, in: J. Burstein, C.
Doran, T. Solorio (Eds.), Proceedings of the 2019
Conference of the North American Chapter of the
Association for Computational Linguistics: Human
Language Technologies, Volume 1 (Long and Short
Papers), Association for Computational Linguistics,
Minneapolis, Minnesota, 2019, pp. 4171–4186. URL:
https://aclanthology.org/N19-1423. doi:10.18653/
v1/N19-1423.
[37] Y. Sun, S. Wang, Y. Li, S. Feng, H. Tian, H. Wu,</p>
        <p>H. Wang, Ernie 2.0: A continual pre-training
framework for language understanding, in: Proceedings
of the AAAI conference on artificial intelligence,
volume 34, 2020, pp. 8968–8975.
[38] Y. Wan, B. Yang, D. F. Wong, L. S. Chao, L. Yao,</p>
        <p>H. Zhang, B. Chen, Challenges of neural machine
translation for short texts, Computational
Linguistics 48 (2022) 321–342.
[39] N. Reimers, I. Gurevych, Sentence-bert: Sentence
embeddings using siamese bert-networks, arXiv
preprint arXiv:1908.10084 (2019).</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>