<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Jaén, Spain
* Corresponding author.
$ mario.graf@infotec.mx (M. Graf); dmoctezuma@centrogeo.edu.mx (D. Moctezuma); eric.tellez@infotec.mx
(E. Tellez); sabino.miranda@infotec.mx (S. Miranda)
 https://mgrafg.github.io/ (M. Graf); https://www.centrogeo.org.mx/areas-profile/dmoctezuma (D. Moctezuma);
https://sadit.github.io/ (E. Tellez); https://www.infotec.mx/es_mx/Infotec/sabino-miranda-jimenez (S. Miranda)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Ingeotec at DA-VINCIS: Bag-of-Words Classifiers</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mario Graf</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Daniela Moctezuma</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Eric Tellez</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sabino Miranda</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CICESE Centro de Investigación Científica y de Educación Superior de Ensenada.</institution>
          <addr-line>Carretera Ensenada - Tijuana No. 3918, Zona Playitas, CP. 22860, Ensenada, B.C.</addr-line>
          <country country="MX">México</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>CONAHCYT Consejo Nacional de Humanidades, Ciencia y Tecnología</institution>
          ,
          <addr-line>1582 Insurgentes Sur 1582, Crédito Constructor, Ciudad de México, 03940</addr-line>
          <country country="MX">México</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>CentroGEO Centro de Investigación en Ciencias de Información Geoespacial, Circuito Tecnopolo Norte No. 117, Col. Tecnopolo Pocitos II</institution>
          ,
          <addr-line>C.P., Aguascalientes, Ags 20313</addr-line>
          <country country="MX">México</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>INFOTEC Centro Público de Investigación en Tecnologías de la Información y Comunicación, 112 Circuito Tecnopolo Sur, Parque Industrial Tecnopolo 2</institution>
          ,
          <addr-line>Aguascalientes, 20326</addr-line>
          ,
          <country country="MX">México</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>Violence is a serious problem that can have a devastating impact on individuals and communities. In some cases, the virtual world is prone to aggressive expressions and insults, among other violent expressions. Also, social media platforms provide a valuable source of information for detecting and monitoring violent events, as people often share posts about them in real time. This information can be used to improve responses to violence in the world and to design better crime prevention policies. In this sense, the DA-VINCIS task at IberLEF 2023 is a competition to develop multimodal models to detect violent incidents on Twitter. The task has two tracks: i) violent event identification and ii) violent event category recognition. This manuscript describes our solution system for the (i) track using only text-based features. More particularly, we used EvoMSA, our multilingual text classification framework based on stacked generalization, to develop competitive models against more complex approaches achieving an F1 score of 0.89, just 3 hundredths behind the best model in the final ranking.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Identification of violent events</kwd>
        <kwd>Models based on a bag of words</kwd>
        <kwd>Pre-trained vocabularies</kwd>
        <kwd>Stacked generalization</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Physical and psychological violence victims can sufer depression, anxiety, and post-traumatic
stress disorder. Social media platforms provide a valuable source of information for detecting
and monitoring violent events, as people often share posts about them in real-time. It is possible
to create models to detect automatically violent incidents in social media for better responses in
the real world and improve the design of crime prevention policies. This problem is particularly
dificult since, in some cases, even human beings disagree about what is and what is not an
insult or violent message in text [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        As an efort to attend this problem, the DA-VINCIS task [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] at IberLEF 2023 asks for models to
detect violent incidents on Twitter using images and text. It asks for methods that classify tweets
as mentioning a violent event. The shared task has two tracks: i) violent event identification
and ii) violent event category recognition. Our solution focuses on the first sub-task using the
provided text, i.e., we provide a model using natural language processing features.
      </p>
      <p>
        Arellano et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] presented the previous edition of the challenge focusing on Twitter messages
with two sub-tasks: binary and multi-label classification tasks. DA-VINCIS 2022 solutions were
all Transformer-based.
      </p>
      <p>
        Nevertheless, there is more extensive literature related to violence detection on Twitter; for
instance, in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] is proposed a model based on natural language processing to identify intimate
partner violence (IPV) on Twitter where achieved scores of F1-score of 0.76 on IPV class and 0.97
on the non-IPV class. This issue is not tackled in either English or Spanish languages; it has been
addressed in other languages such as Arabic; in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] sentiment analysis is used based on lexicon
to classify violence on social networks. They used two social media platforms, Facebook and
Twitter, to acquire messages to be classified. The results get a performance metric of F1-score
between 0.75 and 0.8 values.
      </p>
      <p>
        This manuscript describes our system solution to the DA-VINCIS 2023, which is based on our
EvoMSA framework (evomsa.readthedocs.io) [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. EvoMSA is a multilingual text classification
framework built on stacked generalization [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], which efectively combines the output of multiple
models into a single prediction. Our approach uses lexical and semantic features, all computed
as bags of words. Compared to more complex deep learning approaches, the resulting system is
competitive for violent event identification, even only using features from pure text messages.
      </p>
      <p>This manuscript is organized as follows, Section 2 details our proposed system sent to the
DAVINCIS challenge. Section 3 presents our results, and some conclusions are given in Section 4.
Finally, the appendix has the EvoMSA code to use our models.</p>
    </sec>
    <sec id="sec-2">
      <title>2. System description</title>
      <p>
        Since we focus on text messages, violent event identification can be seen as a text classification
task - in fact, many of the tasks encountered at IberLEF2023 [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] can be posed as text
categorization problems. Given the content of a tweet written in natural language, our model must
identify whenever a violent event is mentioned.
      </p>
      <p>A standard approach to tackle text classification problems is to pose it as a supervised learning
problem. In supervised learning, everything starts with a dataset composed of pairs of inputs
and outputs; in this case, the inputs are texts, and the outputs correspond to the associated
labels or categories. The aim is that the developed algorithm can automatically assign a label to
any given text independently, whether it was in the original dataset. The feasible categories are
only those found on the original dataset. In some circumstances, the method can also inform
the confidence it has in its prediction so the user can decide whether to use or discard it.</p>
      <p>Following a supervised learning approach requires that the input is in amenable representation
for the learning algorithm; usually, this could be a vector. One of the most common methods to
represent a text into a vector is to use a Bag of Words (BoW) model, which works by having
a fixed vocabulary where each component represents an element in the vocabulary and the
presence of it in the text is given by a non-zero value.</p>
      <p>The core idea of a BoW is that after the text is normalized and tokenized, each token  is
associated with a vector vt ∈ R where the -th component, i.e., vt, contains the
InverseDocument-Frequency (IDF) value of the token  and ∀̸=vt = 0. The set of vectors v
corresponds to the vocabulary, there are  diferent tokens in the vocabulary, and by definition
∀̸= vi · vj = 0, where vi ∈ R, vj ∈ R, and (· ) is the dot product. It is worth mentioning that
any token outside the vocabulary is discarded.</p>
      <p>Using this notation, a text  is represented by the sequence of its tokens, i.e., (1, 2, . . .);
the sequence can have repeated tokens, e.g.,  = . Then each token is associated with its
respective vector v (keeping the repetitions), i.e., (vt1 , vt2 , . . .). Finally, the text  is represented
as:
x =</p>
      <p>∑︀ vt ,
‖∑︀ vt‖
(1)
where the sum goes for all the elements of the sequence, x ∈ R, and ‖w‖ is the Euclidean
norm of vector w. The term frequency is implicitly computed in the sum because the process
allows token repetitions.</p>
      <p>The second representation developed relies on using a dense representation based on BoW.
The idea is to represent a text in a vector space where the components have a more complex
meaning than the BoW model. In BoW, each component’s meaning corresponds to the associated
token, and the IDF value gives its importance. The complex behavior comes from associating
each component to the decision value of a text classifier (e.g., BoW) trained on a labeled dataset
that is diferent from the task at hand, albeit nothing forbids to be related to it. The datasets
from which these decision functions come can be built using a self-supervised approach or
annotating texts.</p>
      <p>Without loss of generality, it is assumed that there are  labeled datasets, each one contains
a binary text classification problem; noting that if a dataset has  labels, then this dataset can
be represented as  binary classification problems following the one versus the rest approach,
i.e., it is transformed to  datasets.</p>
      <p>For each of these  binary text classification problems, a BoW classifier is built using the
default parameters (a pre-trained BoW representation and a linear Support Vector Machine
(SVM) as the classifier). Consequently, there are  binary text classifiers, i.e., (1, 2, . . . ,  ).
Additionally, the decision function of  is a value where the sign indicates the class. The
text representation is the vector obtained by concatenating the decision functions of the 
classifiers and then normalizing the vector to have length 1.</p>
      <p>A text  is represented with vector x′ ∈ R where the value x′  corresponds to the decision
function of . Given that the classifier  is a linear SVM, the decision function corresponds to
the dot product between the input vector and the weight vector w plus the bias w0 , where the
′
weight vector and the bias are the parameters of the classifier. That is, the value x  corresponds
to</p>
      <p>∑︀ vt
x′  = w · ‖∑︀ vk‖ + w0 ,
x′ = W · ‖∑∑︀︀ vvtk‖ + w0,
x′ = ∑︁ ut + w0,</p>
      <p>where vt is the IDF vector associated to the token  of the text . In matrix notation, vector x′ is
where matrix W ∈ R×  contains the weights, and w0 ∈ R is the bias. Another way to
see the previous formulation is by defining a vector ut = ‖∑︀1vk‖ Wvt. Consequently, x′ is
defined as:
(2)
(3)
(4)
vectors u ∈ R correspond to the tokens; this is the reason we refer to this model as a dense
′
BoW. Finally, the vector representing the text  is the normalized x′ , i.e., x = ‖xx′ ‖ .</p>
      <sec id="sec-2-1">
        <title>2.1. BoW parameters</title>
        <p>
          Diferent BoW representations were created and implemented following the approach described
in [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. The first step was to set all the characters to lowercase and remove diacritics and
punctuation symbols. Additionally, the users and the URLs were removed from the text. Once the
text is normalized, it is split into bigrams, words, and q-grams of characters with  = {2, 3, 4}.
        </p>
        <p>The pre-trained BoW is estimated from 4,194,304 (222) tweets randomly selected in a larger
collection of messages. The IDF values were estimated from the collections, and some tokens
were selected from all the available ones found in the collection. Two procedures were used to
select the tokens; the first corresponds to selecting the  tokens with the highest frequency,
and the other to normalize the frequency w.r.t. their type, i.e., bigrams, words, and q-grams
of characters. Once the frequency is normalized, one selects the  tokens with the highest
normalized frequency. The value of  is 217; however, one can also find in the library models
for 213, 214, . . . , 217.</p>
        <p>It is also possible to train the BoW model using the training set; in this case, we used the
default parameters. The only diference is that vocabulary size  is unlimited, containing all the
tokens as found in the training set.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Dense representation parameters</title>
        <p>The dense representations start by defining the labeled datasets used to create them. These
datasets are organized in three groups. The first one is composed of human-annotated datasets;
we refer to them as dataset. The second groups contain a set of self-supervised dataset where the
objective is to predict the presence of an emoji as expected; these models are referred to as emoji.
The final group is also a set of self-supervised datasets where the task is to predict the presence
of a particular word, namely keyword. This final group includes a set of dense representations
where the words are selected from the training set; we refer to these representations as tailored,
given that these were designed based on the problem at hand. The words were the most
discriminative based on the BoW classifier; these were the ones with the highest absolute value
between the product of the coeficient computed by the linear SVM and the IDF coeficient.</p>
        <p>Following an equivalent approach used in the development of the pre-trained BoW, diferent
dense representations were created; these correspond to varying the size of the vocabulary
and the two procedures used to select the tokens. The vector space create by the dataset
representation is R57, for the emoji models is R567, and, finally, for the keyword is R2048.</p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Configurations</title>
        <p>We tested 13 diferent algorithms for each task. The configuration having the best performance
was submitted to the contest. The best performance was computed using k-fold cross-validation
( = 5).</p>
        <p>The diferent configurations tested in this competition are described below. These
configurations include BoW and a combination of BoW with dense representations. Stack generalization
combines the diferent text classifiers, and the top classifier was a Naive Bayes algorithm. The
specific implementation of this configuration can be seen in EvoMSA’s documentation,
particularly in the Section Competition. The implementation of the best configuration for each
problem is also described in the appendix.
bow Pre-trained BoW where the tokens are selected based on a normalized frequency w.r.t. its
type, i.e., bigrams, words, and q-grams of characters.
bow_voc_selection Pre-trained BoW where the tokens correspond to the most frequent ones.
bow_training_set BoW trained with the training set; the number of tokens corresponds to
all the tokens in the set.
stack_bow_keywords_emojis Stack generalization approach where the base classifiers are
the BoW, the emojis, and the keywords dense BoW.
stack_bow_keywords_emojis_voc_selection Stack generalization approach where the
base classifiers are the BoW, the emojis, and the keywords dense BoW. The tokens in
these models were selected based on a normalized frequency w.r.t. its type, i.e., bigrams,
words, and q-grams of characters.
stack_bows Stack generalization approach where the base classifiers are BoW with the two
token selection procedures described previously (i.e., bow and bow_voc_selection).
stack_2_bow_keywords Stack generalization approach where with four base classifiers.</p>
        <p>These correspond to two BoW and two dense BoW (emojis and keywords), where the
diference in each is the procedure used to select the tokens, i.e., the most frequent or
normalized frequency.
stack_2_bow_tailored_keywords Stack generalization approach with four base classifiers.</p>
        <p>These correspond to two BoW and two dense BoW (emojis and keywords), where the
diference in each is the procedure used to select the tokens, i.e., the most frequent
or normalized frequency. The second diference is that the dense representation with
normalized frequency also includes models for the most discriminant words selected by
a BoW classifier in the training set. We refer to these latter representations as tailored
keywords.
stack_2_bow_all_keywords Stack generalization approach with four base classifiers
equivalent to stack_2_bow_keywords where the diference is that the dense representations
include the models created with the human-annotated datasets.
stack_2_bow_tailored_all_keywords Stack generalization approach with four base
classiifers equivalent to stack_2_bow_all_keywords, where the diference is that the dense
representation with normalized frequency also includes the tailored keywords.
stack_3_bows Stack generalization approach with three base classifiers. All of them are
BoW; the first two correspond pre-trained BoW with the two token selection procedures
described previously (i.e., bow and bow_voc_selection), and the latest is a BoW trained
on the training set (i.e., bow_training_set).
stack_3_bows_tailored_keywords Stack generalization approach with five base classifiers.</p>
        <p>The first corresponds to a BoW trained on the training set, and the rest are used in
stack_2_bow_tailored_keywords.
stack_3_bow_tailored_all_keywords Stack generalization approach with five base
classiifers. It is comparable to stack_3_bows_tailored_keywords being the diference in the use
of the tailored keywords.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Results</title>
      <p>Table 1 presents the performance, in terms of F1, of the diferent configurations in the k-fold
cross-validation ( = 5). It also includes the performance of our submission in the competition
and the competition’s winner. The table includes two columns Tailored and Dense; these provide
an overview of the systems. The tailored column indicates whether the training set was used to
define the words in the self-supervised datasets or the BoW vocabulary was obtained from it.
The dense column indicates the use of a dense representation.</p>
      <p>It can be observed that, for edition 2023, the best configuration is a stack generalization
algorithm with four base classifiers; two BoW and two dense representations. One of the
dense models also includes tailored keywords. The best configuration for edition 2022 is quite
similar to the one for 2023; the diference is that the 2022 configuration does not use tailored
keywords. The configuration used in 2023 is the second-best configuration of 2022, according
to the performance obtained in the k-fold cross-validation approach.</p>
      <p>Another comparison that one can make with the Table’s data is computing the diference
between the best performance in k-fold cross-validation and the worst; it can be observed that
for the 2023 edition, the diference is 1.4%, and for 2022 is 13.3%. The diference of 1.4% indicates
that following this approach will be complicated to improve. On the other hand, for the edition
2022, there might be room for improvement following the presented approach.</p>
      <p>The second block in the table presents the performance of our participation as INGEOTEC
in the competition and the performance of the competition’s winner. The last row shows the
diference in percentage between these two values. As can be seen, the diference in both
competitions is 4.1%, indicating that the approach presented is competitive with the winner.
However, there is more viability in the 2022 competition than in 2023, as seen in the diference
of performance on the diferent configurations tested.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusions</title>
      <p>We have presented our system solution and results for detecting violent incidents on Twitter
using only text-based features in the context of the DA-VINCIS 2023 challenge. Our system is
based on the ensemble of multiple models using stacked generalization, more precisely with
our EvoMSA framework. Our approach uses lexical and semantic features, all computed as bags
of words. We achieved an F1 score of 0.89, just 3 hundredths behind the best model in the final
ranking.</p>
      <p>Our results show that developing competitive models for violent event identification is
possible using only text-based features and, even more, bag-of-words-based models. This is
important because it means that we can use similar methodologies for languages without large
language models or when it is impossible to aford the prediction costs of transformer-based
solutions.</p>
    </sec>
    <sec id="sec-5">
      <title>A. Library Usage</title>
      <p>
        This appendix aims to illustrate how the best configurations were implemented using EvoMSA
(evomsa.readthedocs.io) [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. The first step is to install the library, which can be done using the
Anaconda package manager with the following instruction.
      </p>
      <p>conda install -c conda-forge EvoMSA</p>
      <p>A more general approach to installing EvoMSA is through the use of the command pip, as
illustrated in the following instruction.</p>
      <p>pip install EvoMSA</p>
      <p>Once EvoMSA is installed, one must load a few libraries; the libraries used in the best
configurations are the following.</p>
      <p>from EvoMSA import BoW, DenseBoW
from EvoMSA import StackGeneralization as Stack
from EvoMSA.utils import Linear
from microtc.utils import tweet_iterator</p>
      <p>The best configuration found for DAVINCIS 2023 corresponds to a stack generalization with
four base classifiers, two BoW, and two dense BoW, where one dense BoW contains models
identified in the training set. The following function implements it. It can be observed that the
second and eighth lines create instances of BoW. The third line creates the instance of one of
the dense BoW; the fifth line has the name of the tailored keywords, which are downloaded and
loaded on line sixth line. The ninth line initializes the second dense BoW corresponding to the
vocabulary with the most frequent tokens. The thirteenth line creates the stack generalization,
and the last line performs the predictions.</p>
      <p>def stack_2_bow_tailored_all_keywords(lang, tr, vs):
bow = BoW(lang=lang)
keywords = DenseBoW(lang=lang, n_jobs=-1,</p>
      <p>skip_dataset=set([’davincis2022_1’]))
tailored = ’IberLEF2022_DAVINCIS_task1_Es.json.gz’
keywords.text_representations_extend(tailored)
keywords.select(D=tr)
bow2 = BoW(lang=lang, voc_selection=’most_common’)
keywords2 = DenseBoW(lang=lang, n_jobs=-1,
skip_dataset=set([’davincis2022_1’]),
voc_selection=’most_common’)
keywords2.select(D=tr)
stack = Stack(decision_function_models=[bow, bow2, keywords,
keywords2]).fit(tr)
return stack.predict(vs)</p>
      <p>DAVINCIS 2022 best configuration is almost equivalent to the 2023 best
configuration. The only diference is that the former edition did not use the tailored keywords.
The following function implements the 2022 configuration; the absence of the line
keywords.text_representations_extend can be observed, which is responsible for incorporating the
tailored keywords.</p>
      <p>def stack_2_bow_all_keywords(lang, tr, vs):
bow = BoW(lang=lang)
keywords = DenseBoW(lang=lang, n_jobs=-1,</p>
      <p>skip_dataset=set([’davincis2022_1’]))
keywords.select(D=tr)
bow2 = BoW(lang=lang, voc_selection=’most_common’)
keywords2 = DenseBoW(lang=lang, n_jobs=-1,
skip_dataset=set([’davincis2022_1’]),
voc_selection=’most_common’)
keywords2.select(D=tr)
stack = Stack(decision_function_models=[bow, bow2, keywords,
keywords2]).fit(tr)</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J.</given-names>
            <surname>Guberman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Schmitz</surname>
          </string-name>
          , L. Hemphill,
          <article-title>Quantifying toxicity and verbal violence on twitter</article-title>
          ,
          <source>in: Proceedings of the 19th ACM Conference on Computer Supported Cooperative Work and Social Computing Companion, CSCW '16 Companion</source>
          , Association for Computing Machinery, New York, NY, USA,
          <year>2016</year>
          , p.
          <fpage>277</fpage>
          -
          <lpage>280</lpage>
          . URL: https://doi.org/10.1145/2818052. 2869107. doi:
          <volume>10</volume>
          .1145/2818052.2869107.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>H.</given-names>
            <surname>Jarquín-Vásquez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. I. Hernández</given-names>
            <surname>Farías</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Arellano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. J.</given-names>
            <surname>Escalante</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Villaseñor-Pineda</surname>
          </string-name>
          , M. Montes y Gómez,
          <string-name>
            <given-names>F.</given-names>
            <surname>Sanchez-Vega</surname>
          </string-name>
          ,
          <article-title>Overview of da-vincis at iberlef 2023: Detection of aggressive and violent incidents from social media in spanish</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>71</volume>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>L. J.</given-names>
            <surname>Arellano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. J.</given-names>
            <surname>Escalante</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. Villaseñor</given-names>
            <surname>Pineda</surname>
          </string-name>
          , M. Montes y Gomez,
          <string-name>
            <surname>F.</surname>
          </string-name>
          <article-title>Sanchez Vega, Overview of DA-VINCIS at IberLEF 2022: Detection of Aggressive and Violent Incidents from Social Media in Spanish</article-title>
          , Procesamiento de Lenguaje Natural (
          <year>2022</year>
          )
          <fpage>207</fpage>
          -
          <lpage>215</lpage>
          . URL: https://scholar.google.com/scholar?hl
          <article-title>=es&amp;as_sdt=0%2C5&amp;q=Overview+of+DA-VINCIS+ at</article-title>
          +IberLEF+
          <year>2022</year>
          %3A+Detection+of+Aggressive+and+Violent+Incidents+from+Social+ Media+in+Spanish&amp;btnG=.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Al-Garadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Guo</surname>
          </string-name>
          , E. Warren,
          <string-name>
            <given-names>Y.-C.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lakamana</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sarker</surname>
          </string-name>
          ,
          <article-title>Natural language model for automatic identification of intimate partner violence reports from twitter</article-title>
          ,
          <source>Array</source>
          <volume>15</volume>
          (
          <year>2022</year>
          )
          <fpage>100217</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Khalafat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Ja</surname>
          </string-name>
          <article-title>'far</article-title>
          , R. Al-Sayyed,
          <string-name>
            <given-names>M.</given-names>
            <surname>Eshtay</surname>
          </string-name>
          , T. Kobbaey,
          <article-title>Violence detection over online social networks: An arabic sentiment analysis approach</article-title>
          , iJIM
          <volume>15</volume>
          (
          <year>2021</year>
          )
          <fpage>91</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Graf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Miranda-Jiménez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. S.</given-names>
            <surname>Tellez</surname>
          </string-name>
          , D. Moctezuma,
          <article-title>EvoMSA: A Multilingual Evolutionary Approach for Sentiment Analysis</article-title>
          ,
          <source>Computational Intelligence Magazine</source>
          <volume>15</volume>
          (
          <year>2020</year>
          )
          <fpage>76</fpage>
          -
          <lpage>88</lpage>
          . URL: http://arxiv.org/abs/
          <year>1812</year>
          .02307.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D. H.</given-names>
            <surname>Wolpert</surname>
          </string-name>
          , Stacked generalization,
          <source>Neural Networks</source>
          <volume>5</volume>
          (
          <year>1992</year>
          )
          <fpage>241</fpage>
          -
          <lpage>259</lpage>
          . URL: https://www.sciencedirect.com/science/article/pii/S0893608005800231. doi:
          <volume>10</volume>
          . 1016/S0893-
          <volume>6080</volume>
          (
          <issue>05</issue>
          )
          <fpage>80023</fpage>
          -
          <lpage>1</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Jiménez-Zafra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Rangel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Montes-y</surname>
          </string-name>
          <string-name>
            <surname>Gómez</surname>
          </string-name>
          ,
          <source>Overview of IberLEF 2023: Natural Language Processing Challenges for Spanish and other Iberian Languages, Procesamiento del Lenguaje Natural</source>
          <volume>71</volume>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>E. S.</given-names>
            <surname>Tellez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Miranda-Jiménez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Graf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Moctezuma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. R.</given-names>
            <surname>Suárez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O. S.</given-names>
            <surname>Siordia</surname>
          </string-name>
          ,
          <article-title>A simple approach to multilingual polarity classification in Twitter</article-title>
          ,
          <source>Pattern Recognition Letters</source>
          <volume>94</volume>
          (
          <year>2017</year>
          )
          <fpage>68</fpage>
          -
          <lpage>74</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.patrec.
          <year>2017</year>
          .
          <volume>05</volume>
          .024.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>