<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>G. Y. Bade);</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Hope Speech in Social Media Texts using Transformer</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Girma Yohannis Bade</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Olga Kolesnikova</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>José Luis Oropeza</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Grigori Sidorov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kidist Feleke Bergene</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Instituto Politécnico Nacional (IPN), Centro de Investigación en Computación (CIC)</institution>
          ,
          <addr-line>Mexico City</addr-line>
          ,
          <country country="MX">Mexico</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>Systems to regulate and remove hateful, abusive, and ofensive content from the internet have been developed in the last several years. But occasionally, those in positions of authority abuse this type of censorship to thwart the democratic right to free speech. Consequently, studies must address online content that is uplifting, encouraging, and supporting from a positive reinforcement perspective. In this regard, HOPE_ IberLEF 2024 created a dataset to recognize positivity in social media comments to encourage those who need mindset treatments. It consisted of two tasks with main two aims. Task-1 is about hope speech for equality, diversity,and inclusion, and task-2 is about hope for an expectation for future desire. Then we have been involved in two tasks and propose the tasks with three algorithms, including Logistic regression, Word2Vec, and Transformer-base. Among these three algorithms, the model with Transformer-based outperformed all others. For task-1, our model achieved a 0.55 macro F1-score. For Polyhope binary data, the model achieved 0.75 and 0.82 macro F1- scores for Spanish and English respectively. Similarly, for polyhope multiclass data, our model achieved 0.48 macro F1-score and 0.55 F1-score in Spanish and English datasets respectively.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Hope speech</kwd>
        <kwd>social media</kwd>
        <kwd>machine learning</kwd>
        <kwd>NLP</kwd>
        <kwd>transformer</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Social integration increases the importance of knowing and being known through social networking
platforms. It is essential for everyone’s overall well-being, especially those more susceptible to social
isolation [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. One of the most important aspects of social media messages is positive ideas,
encouragements, thankfulness, appreciation, and acknowledgments provided to participants by their peers in
times of illness, stress, and isolation [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. These elements significantly impact people’s mental, physical,
and psychological well-being [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Despite these benefits, social media content also contains a significant
number of spiteful or unpleasant posts like hate and ofensive [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ], abusive comments [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] fake news
[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], etc. In this regard, there has a significant number of studies have been taking place to tackle these
and cyber-bullying in general. However, in a real scenario, some people heal the heartbreak of people
and extend their sense of living. These people always speak or wish positivity for themselves and others
as optimists. Any concepts associated with positive ideas, encouragements, aspirations, support, and
wishes that might refresh the mind of the victim are known as hope speech.
      </p>
      <p>
        As noted by Palakodety et al [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], hope speech is a kind of speech that can help people feel less worried
in an unfriendly setting. In this paper, hope speech can be seen in two ways. a) in terms of Equality,
diversity, and inclusion [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], b) in terms of expectation [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. The aim of the former is to ensure hope
in the aspect of equality, diversity, and inclusion whereas the later concerns hope as an expectation
in a future [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ]. People who have high hopes do not respond to obstacles the same way as people
who have low hopes; rather, they see obstacles as challenges to be overcome and use their pathway
ideas to map out a diferent course to their objectives [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Moreover, it has been discovered that a lot of
positive factors, like academic success and reduced despair, are correlated with high hope. Low hope,
on the other hand, is linked to unfavorable consequences like diminished well-being [
        <xref ref-type="bibr" rid="ref10 ref11 ref12">10, 11, 12</xref>
        ]. The
greater hope is consistently related to better academic, athletic, physical health [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], psychological
adjustment, and psychotherapy outcomes. Hope speech identification is the process of identifying
whether a comment is Hope Speech or Non-Hope Speech.
      </p>
      <p>
        The advancements in natural language processing (NLP) techniques and machine learning algorithms
have opened a golden opportunity to leverage this desirability [
        <xref ref-type="bibr" rid="ref14">14, 15</xref>
        ]. They enable us to build a
language model in these two dimensions by using valuable social media opinions [16]. We use negative
opinions to moderate social media content and positive opinions to boost equality, diversity, inclusion,
and expectations. As part of the (HOPE@IberLEF 2024) shared task, this paper discusses whether those
social opinions are hope speeches or not [17]. In this regard, gold-standard Spanish and English datasets
are provided by the competition organizers [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Related works</title>
      <p>
        It’s important to refer to the universal feelings that result from the collection of beliefs about the
caliber of our lives as well-being, thriving, happiness, and life satisfaction [18]. Generally, favorable
life outcomes can multiply personal achievement and performance by extending from one stage of life
to another [19]. To uplift and boost those merits, a significant number of research works are being
developed [
        <xref ref-type="bibr" rid="ref1 ref2 ref3 ref8">1, 2, 3, 8, 17</xref>
        ]. Since 2021, for instance, Hope Speech Detection for Equality, Diversity, and
Inclusion(EDI) is taking place for various languages including English, Tamil, Malayalam, Kannada,
Spanish, and so on [
        <xref ref-type="bibr" rid="ref1 ref8">1, 8, 15</xref>
        ].
      </p>
      <p>
        Roy et al. [20] presented a dataset of multilingual hope speeches in English, Tamil, Malayalam, and
Kannada that encourage EDI. It was gathered to guarantee EDI in language technology and to spread
optimism. Since it includes information from the LGBTQIA+ community, people with disabilities, and
women working in science, engineering, technology, and management (STEM) have been participated
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. To develop benchmark systems, they tested the Hope Speech dataset for Equality, Diversity, and
Inclusion (HopeEDI) using various cutting-edge deep learning and machine learning models.
      </p>
      <p>Puranik et al. [21] have attempted to find and promote good and supportive information on various
sites. They employed diferent transformer-based models to categorize social media remarks as hope
speech or not hope speech in English, Malayalam, and Tamil dialects. They have traversed through
transfer learning of several state-of-the-art transformer models for languages such as English, Tamil,
and Malayalam. From their experimentation, ULMFiT achieved an F1-score of 0.9356 on English data.
On the other hand, mBERT achieved a 0.8545 F1-score on the Malayalam test dataset and distilmBERT
achieved a 0.5926 weighted F1-score on Tamil test dataset.</p>
      <p>A hope-speech identification work such as [ 22, 23] was introduced to automatically identify online
content that could help to temper hate on social media. Pre-processing and transfer-learning models
have been utilized during an experiment. According to the result, Convolution neural networks in
the pre-trained multilingual-BERT model produced the greatest results, 0.54 macro-F1 over other
benchmarks.</p>
      <p>Junaida and Ajees [24] demonstrated deep learning methods for word representation using
contextaware string embedding and text representation utilizing pooled document embedding and recurrent
neural networks (RNN). It has assessed and contrasted the three models using various methodologies for
every language. Based on the observation, the suggested method performed better than the baselines
and operated as intended. The languages with the greatest weighted average F-scores are Malayalam,
Tamil, and English, with values of 0.84, 0.93, and 0.58, respectively.</p>
    </sec>
    <sec id="sec-3">
      <title>3. System Description</title>
      <p>In this section, we ofer thorough information regarding the dataset and details of the experiments
carried out in our study. Moreover, it dives into the dataset statistics, preprocessing, algorithm selection,
and the workflow of the proposed model.</p>
      <sec id="sec-3-1">
        <title>3.1. Datasets</title>
        <p>
          For any machine learning task, the dataset is crucial and mandatory. The real world has various data
types like text, video, audio, time series, etc. Preparing datasets for a specific task is labor-intensive
and time-consuming, especially in the NLP domain. Opportunely, for this study, we have been given
a gold-standard dataset by (HOPE@IberLEF 2024) workshop [
          <xref ref-type="bibr" rid="ref9">9, 25, 26</xref>
          ]. It consisted of two datasets
for two tasks. The first dataset is a dataset in the Spanish language aiming to identify hope speech in
terms of quality, diversity, and inclusion [25]. This dataset was collected between 2020 and 2023. It is
an improved and extended version of the previous Spanish_HopeEDI dataset [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. The corpus consisted
of training, development, and test datasets. The training and development datasets are from tweets
and a test set from tweets related to the LGTBI collective and other topics. The data of the first task is
annotated the datasets as Hope Speech(HS) [27] if the text of the tweet:
• explicitly supports the social integration of minorities.
• it’s a positive inspiration.
• explicitly encourages people who might find themselves in a situation.
        </p>
        <p>• unconditionally promotes tolerance.</p>
        <p>On the contrary, it is marked as Not Hope Speech(NHS) [27] if the text of the tweet:
• expresses a negative sentiment towards a community.
• explicitly seeks violence or
• uses gender-based insults.</p>
        <p>
          On the other hand, the second task’s data aims to hope speech in terms of expectations [
          <xref ref-type="bibr" rid="ref8">8, 27</xref>
          ]. In
this case, the most recent tweets were retrieved in the period from January 2022 to June 2022 [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. The
dataset consisted of both English and Spanish tweets in binary and multiclass labels. The following
table 1 shows the data statistics of both tasks.
        </p>
        <p>
          Training Data: The portion of the data that our model is trained on [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. This is the real data both
input and output that our model sees and gains knowledge from.
        </p>
        <p>Validation Data: The portion of data used to fit the model on the training dataset, perform regular
evaluations of the model, and make necessary adjustments to the hyperparameters (first set parameters
before the model starts learning). When the model is training, this data is useful. Moreover, this dataset
gives you insight into how your model would perform in a test dataset. If your model performs better
in the development dataset, it’s some sort of clue that your model would work on the test datasets too,
even if it’s not completely guaranteed [28].</p>
        <p>Testing Data: Testing data ofers an objective assessment once the model has been fully trained and
developed [29]. In our case, the model forecasts or classifies some values when we supply the testing
data inputs (that were unseen and not labeled). We assess our model by contrasting its output with the
manually transcripted real output present on the organizer’s side. The following table 2 shows sample
instances of data. In table 2 second column, the abbreviations in the brackets ’G’, ’R’, and’UR’ represent
generalized, realistic, and unrealistic hope respectively.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Preprocessing</title>
        <p>Preprocessing is preparing raw data for machine learning algorithms by cleaning, converting text into
numeric mode, and organizing it(annotating) [30]. It is the vital stage that fills in the gaps between
raw data and useful insights because raw data is rarely in an ideal state [31, 32]. During the data
preparation phase of machine learning tasks, there are typical or standard activities that we should use.
The following are some among others.</p>
        <p>Handling Missing Data: In real-world datasets, handling missing data is a dificult task because
some users did not fill out the required fields [ 33]. The pre-processing methods like imputation, removal
of missing data, and removing null values ensure that the model is fed accurate and comprehensive
data [34]. It must be handled to prevent the model from performing worse [35]. In addition, we used
“raw[’category’].fillna(0, inplace=True)” to handle empty strings of class labels. Where ’raw’ is the
data-frame object variable and ’category’ is the column name of class variable.</p>
        <p>Data Cleaning: Finding and fixing inaccuracies or flaws in the data is known as data cleaning [ 36].
In this regard, we explored our datasets and corrected the inaccuracies encountered.</p>
        <p>Data Encoding: Since machine learning algorithms usually operate on numerical data, it is necessary
to properly encode categorical or text data variables into numerical [36]. based on the algorithm we
have been implementing, two types of techniques are employed. For traditional machine learning
algorithms, we have employed the ’TF-IDF’ text vectorization approach. Because it understands the
context using the instance positions. In addition, the ‘class label’ attributes are converted by using the
function known as to_numeric(). In the case of deep learning algorithms, they have their vectorizer
known as "keras Tokenizer". After the tokenization process, again ’keras Embedding’ function is used
to convert the text into numeric form.</p>
        <p>Apart from these built-in preprocessing tasks, we have also used user-defined way of methods to
remove unwanted characters like numbers, special symbols, extra white spaces, and so on from the
texts. In the case of English language texts, there is a chance of throwing the stop word because their
presence in the text did not bear any further relevance rather consuming the computational cost. This
also helps to increase the relevance of the datasets and then the performance of the model too [37].</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Model selection</title>
        <p>To train the model, we picked and used the algorithms from three levels: machine learning, deep
learning, and transformer-based. From machine learning, we picked up logistic regression. Logistic
regression is most commonly used for classification problems. In logistic regression, the logistic function
(sigmoid) plays a vital role in modeling the relationship between the input features and the probability
of outputs. The model trains by using maximum likelihood estimation to optimize the parameters [15].</p>
        <p>We have used Word2Vec with the LSTM algorithm from deep learning. In deep learning, Neural
Networks (ANNs) are used and it’s a relatively new class of machine learning techniques to address
complicated problems. ANN is made up of two or more layers of processing units called neurons
that are intended to resemble the human nervous system. Layers and activation factors enable these
models to acquire hierarchical non-linear characteristics for class distinction. Neural network type, the
number of layers, the number of neurons, and activation functions at each layer are all specified by this
architecture [38].</p>
        <p>The novel deep learning (DL) architecture that has gained popularity and enabled recent advances
in the field of natural language processing is known as Transformer. It was first introduced in the
paper, Vaswani et al. [39] to improve the caliber of DL and NLP research tasks. It has shown remarkable
eficiency and has a lot of potential for general usage in artificial intelligence applications. It mainly
uses the self-attention process to extract inherent properties. The table 3 shows three algorithms from
common state-of-the-arts.</p>
        <p>The parameters details for the table 3 are as follows. In the case of logistic regression, we accepted
all default parameters except employing the TF-IDF vectorization approach.</p>
        <p>For deep learning algorithms, we chose ’Word2Vec &amp; LSTM’ algorithms to make a comparison between
traditional machine learning and transformer-based. Actually, Word2Vec came from gensim models
and LSTM from the Keras framework. The Word2Vec algorithm is parameterized as (vector_size=100,
window=5, min_count=1). These are some of the most common parameters among others. The ’window’
here is the size of the token before and after the target token to grasp the sense of semantics from the
context. In the case of LSTM, it used a sequential mode of learning with a sequence size of 128. The
activation function is set to the sigmoid and loss is ’binary_crossentropy’. Particularly, we set the batch
size to 32 and epochs to 10.</p>
        <p>In the case of the transformer, Input_layer is Input of max_length, whereas embedding_layers are the
embedding of input_dimension in Word2Vec embedding. The attention layers are the embedding layer
of self-attention, and finally, it utilizes the ’sigmoid’ activation function for the binary classification.
However, in the case of multiclass classification, it utilized ’softmax’ as an activation function because
it takes the maximum values to determine the outputs. The model tries to optimize the error using the
’adam’ optimizer and the loss function set to the ’binary_crossentropy’ value. Likewise, the model has
trained in 5 epochs for 32 sample batch-size.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Result and Discussion</title>
      <p>The developed models were based on the Logistic regression(LR), Word2Vec with LSTM, and
Transformer-base classifier and they have been evaluated in terms of precision, recall,macro-F1, and
accuracy scores. However, the macro-F1 mainly determines the model’s performance because it evaluates
the model by harmonizing the class labels. According to the result published by the task organizer, the
transformer-based classifier outperformed logistic regression and Word2Vec + LSTM.</p>
      <p>The models classify social media posts in test datasets into hope speech or not hope speech in
binary classification, and generalized hope, realistic hope, and unrealistic hope in case of multiclass
classification as expected. We mentioned the results in terms of macro and weighted metrics of model
evaluations: macro-precision (P), macro-recall (R), macro F1-score, and weighted average precision,
recall, F1, and accuracy. The table 4 shows the detail.</p>
      <p>Where M_Pr,M_R,M_F1,W_Pr,W_R,W_F1, and Acc represent average
macro-precision,recall,F1score,average weighted-precision,recall,F1-score and accuracy respectively.</p>
      <p>Similarly, the result of task 2 is mentioned in terms of evaluation metrics: Precision (P), Recall (R),
F1-score, and the average macro F1-scores in the table 5 below.</p>
      <p>According to the result that was published by the organizers of shared tasks, our models have
revealed a comparative performance. In task 1, our model achieved a 0.55 macro-F1 score on the
transformer-based algorithm.</p>
      <p>Similarly, in task 2 (binary classification), the better result was recorded as a 0.82 F1 score in the
English dataset and a 0.75 F1 score in the Spanish dataset whereas in the case of multiclass classification,
the better result was recorded as 0.48 F1 score in Spanish data and 0.55 F1-score in English dataset.</p>
      <p>In both tasks, task-1 and task-2 the transformer-based algorithm outperformed other baseline
algorithms and the model in the English dataset outperformed over Spanish datasets. This shows that the
selected models well experienced English data than Spanish. More importantly, the pre-processing task
has a positive impact on enhancing the model quality that utilizes English datasets than non-English
datasets. All the details are mentioned in the table 4 &amp; 5.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion and Future Work</title>
      <p>Hope is a positive frame of mind that is both present- and future-focused. It is founded on the desire
for favorable results in one’s life or the world as a whole and may also be found in motivational
speeches about those who have faced hardship. The paper addresses the development of systems to
identify and promote positive, uplifting content on social media, in contrast to the usual focus on
detecting and removing negative content. Specifically, it evaluates the HOPE_IberLEF 2024 dataset
designed to recognize positivity in social media comments with two main tasks: one focusing on hope
speech related to equality, diversity, and inclusion, and the other on future-oriented hope. It also
compares three algorithms—Logistic Regression, Word2Vec, and Transformer-based models—finding
that the Transformer-based approach outperforms the others. On the other side, we realized that the
performance of the model highly depends on the size and quality of the data sets.</p>
      <p>Finally, since the NLP task can identify hope speech on social media posts and uphold human mindsets,
the jobs ought to be transferred to other languages. Furthermore, by ofering additional algorithms
for the languages utilized here and expanding the number of dataset sizes, the performance of the
suggested model in this study should be enhanced.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements</title>
      <p>The work was done with partial support from the Mexi- can the Government through the grant
A1-S47854 of CONACYT, Mexico, grants 20241816, 20241819, and 20240951 of the Sec- retar ıa de Investigaci
on y Posgrado of the Instituto Polit ecnico Nacional, Mexico. The authors thank the CONACYT for the
computing resources brought to them through the Plataforma de Aprendizaje Profundo para Tecnolog
ıas del Lenguaje of the Laboratorio de Superc omputo of the INAOE, Mexico and ac- knowledge the
support of Microsoft through the Microsoft Latin America PhD Award.
[15] M. G. Yigezu, M. M. Woldeyohannis, A. L. Tonja, Multilingual neural machine translation for
low resourced languages: Ometo-english, in: 2021 International Conference on Information and
Communication Technology for Development for Africa (ICT4DA), IEEE, 2021, pp. 89–94.
[16] Z. Ahani, M. Shahiki Tash, Y. Ledo Mezquita, J. Angel, Utilizing deep learning models for the
identification of enhancers and super-enhancers based on genomic and epigenomic features,
Journal of Intelligent &amp; Fuzzy Systems (2024) 1–11.
[17] M. Tash, Z. Ahani, M. Zamir, O. Kolesnikova, G. Sidorov, Lidoma@ lt-edi 2024: Tamil hate speech
detection in migration discourse, in: Proceedings of the Fourth Workshop on Language Technology
for Equality, Diversity, Inclusion, 2024, pp. 184–189.
[18] M. Shahiki-Tash, J. Armenta-Segura, O. Kolesnikova, G. Sidorov, A. Gelbukh, Lidoma at
hope2023iberlef: Hope speech detection using lexical features and convolutional neural
networks, in: Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2023), co-located
with the 39th Conference of the Spanish Society for Natural Language Processing (SEPLN 2023),
CEUR-WS. org, 2023.
[19] M. G. Yigezu, M. A. Mehamed, O. Kolesnikova, T. K. Guge, A. Gelbukh, G. Sidorov, Evaluating
the efectiveness of hybrid features in fake news detection on social media, in: 2023 International
Conference on Information and Communication Technology for Development for Africa (ICT4DA),
IEEE, 2023, pp. 171–175.
[20] P. Roy, S. Bhawal, A. Kumar, B. R. Chakravarthi, Iiitsurat@ lt-edi-acl2022: Hope speech detection
using machine learning, in: Proceedings of the Second Workshop on Language Technology for
Equality, Diversity and Inclusion, 2022, pp. 120–126.
[21] K. Puranik, A. Hande, R. Priyadharshini, S. Thavareesan, B. R. Chakravarthi, Iiitt@
lt-edi-eacl2021hope speech detection: there is always hope in transformers, arXiv preprint arXiv:2104.09066
(2021).
[22] M. G. Yigezu, O. Kolesnikova, A. Gelbukh, G. Sidorov, Odio-bert: Evaluating domain task impact
in hate speech detection, Journal of Intelligent &amp; Fuzzy Systems (????) 1–12.
[23] S. Dowlagar, R. Mamidi, Edione@ lt-edi-eacl2021: Pre-trained transformers with convolutional
neural networks for hope speech detection., in: Proceedings of the First Workshop on Language
Technology for Equality, Diversity and Inclusion, 2021, pp. 86–91.
[24] M. Junaida, A. Ajees, Ku_nlp@ lt-edi-eacl2021: a multilingual hope speech detection for equality,
diversity, and inclusion using context aware embeddings, in: Proceedings of the first workshop
on language technology for equality, diversity and inclusion, 2021, pp. 79–85.
[25] D. García-Baena, M. Á. García-Cumbreras, S. M. Jiménez-Zafra, J. A. García-Díaz, R.
ValenciaGarcía, Hope speech detection in spanish: The lgbt case, Language Resources and Evaluation 57
(2023) 1487–1514.
[26] F. Balouchzahi, G. Sidorov, A. Gelbukh, Polyhope: Two-level hope speech detection from tweets,</p>
      <p>Expert Systems with Applications 225 (2023) 120078. doi:10.1016/j.eswa.2023.120078.
[27] L. Chiruzzo, S. M. Jiménez-Zafra, F. Rangel, Overview of IberLEF 2024: Natural Language
Processing Challenges for Spanish and other Iberian Languages, in: Proceedings of the Iberian Languages
Evaluation Forum (IberLEF 2024), co-located with the 40th Conference of the Spanish Society for
Natural Language Processing (SEPLN 2024), CEUR-WS.org, 2024.
[28] Z. Ahani, M. Tash, M. Zamir, I. Gelbukh, Zavira@DravidianLangTech 2024:Telugu hate speech
detection using LSTM, in: B. R. Chakravarthi, R. Priyadharshini, A. K. Madasamy, S. Thavareesan,
E. Sherly, R. Nadarajan, M. Ravikiran (Eds.), Proceedings of the Fourth Workshop on Speech, Vision,
and Language Technologies for Dravidian Languages, Association for Computational Linguistics,
St. Julian’s, Malta, 2024, pp. 107–112. URL: https://aclanthology.org/2024.dravidianlangtech-1.17.
[29] M. Zamir, M. Tash, Z. Ahani, A. Gelbukh, G. Sidorov, Tayyab@DravidianLangTech 2024:detecting
fake news in Malayalam LSTM approach and challenges, in: B. R. Chakravarthi, R. Priyadharshini,
A. K. Madasamy, S. Thavareesan, E. Sherly, R. Nadarajan, M. Ravikiran (Eds.), Proceedings of
the Fourth Workshop on Speech, Vision, and Language Technologies for Dravidian Languages,
Association for Computational Linguistics, St. Julian’s, Malta, 2024, pp. 113–118. URL: https:
//aclanthology.org/2024.dravidianlangtech-1.18.
[30] M. S. Tash, Z. Ahani, O. Kolesnikova, G. Sidorov, Analyzing emotional trends from x platform using
senticnet: A comparative analysis with cryptocurrency price, arXiv preprint arXiv:2405.03084
(2024).
[31] M. S. Tash, Z. Ahani, A. Tonja, M. Gemeda, N. Hussain, O. Kolesnikova, Word level language
identification in code-mixed kannada-english texts using traditional machine learning algorithms,
in: Proceedings of the 19th International Conference on Natural Language Processing (ICON):
Shared Task on Word Level Language Identification in Code-mixed Kannada-English Texts, 2022,
pp. 25–28.
[32] G. Y. Bade, A. A. Afaro, Object oriented software development for artificial intelligence, American</p>
      <p>Journal of Software Engineering and Applications 7 (2018) 22–24.
[33] S. Kanta, G. Sidorov, Selam@DravidianLangTech:sentiment analysis of code-mixed Dravidian
texts using SVM classification, in: B. R. Chakravarthi, R. Priyadharshini, A. K. M, S. Thavareesan,
E. Sherly (Eds.), Proceedings of the Third Workshop on Speech and Language Technologies for
Dravidian Languages, INCOMA Ltd., Shoumen, Bulgaria, Varna, Bulgaria, 2023, pp. 176–179. URL:
https://aclanthology.org/2023.dravidianlangtech-1.24.
[34] M. Zamir, M. Tash, Z. Ahani, A. Gelbukh, G. Sidorov, Lidoma@ dravidianlangtech 2024: Identifying
hate speech in telugu code-mixed: A bert multilingual, in: Proceedings of the Fourth Workshop
on Speech, Vision, and Language Technologies for Dravidian Languages, 2024, pp. 101–106.
[35] T. Achamaleh, L. Kawo, I. Batyrshini, G. Sidorov, Tewodros@ dravidianlangtech 2024: Hate speech
recognition in telugu codemixed text, in: Proceedings of the Fourth Workshop on Speech, Vision,
and Language Technologies for Dravidian Languages, 2024, pp. 96–100.
[36] G. Y. Bade, H. Seid, Development of longest-match based stemmer for texts of wolaita language,
vol 4 (2018) 79–83.
[37] G. Y. Bade, Four basic concepts of object technology: A structuring method, a reliability discipline,
an epistemological principle and a classification technique (????).
[38] C. Suneera, J. Prakash, Performance analysis of machine learning and deep learning models for
text classification, in: 2020 IEEE 17th India council international conference (INDICON), IEEE,
2020, pp. 1–6.
[39] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, I. Polosukhin,
Attention is all you need, in: I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S.
Vishwanathan, R. Garnett (Eds.), Advances in Neural Information Processing Systems, volume 30,
Curran Associates, Inc., 2017. URL: https://proceedings.neurips.cc/paper_files/paper/2017/file/
3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <article-title>Hope speech detection in youtube comments</article-title>
          ,
          <source>Social Network Analysis and Mining</source>
          <volume>12</volume>
          (
          <year>2022</year>
          )
          <fpage>75</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Chakravarthi</surname>
          </string-name>
          ,
          <article-title>HopeEDI: A multilingual hope speech detection dataset for equality, diversity, and inclusion</article-title>
          , in: M.
          <string-name>
            <surname>Nissim</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Patti</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Plank</surname>
          </string-name>
          , E. Durmus (Eds.),
          <source>Proceedings of the Third Workshop on Computational Modeling of People's Opinions</source>
          , Personality, and
          <article-title>Emotion's in Social Media, Association for Computational Linguistics</article-title>
          , Barcelona,
          <source>Spain (Online)</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>41</fpage>
          -
          <lpage>53</lpage>
          . URL: https://aclanthology.org/
          <year>2020</year>
          .peoples-
          <volume>1</volume>
          .5.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M. G.</given-names>
            <surname>Yigezu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. Y.</given-names>
            <surname>Bade</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Kolesnikova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sidorov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gelbukh</surname>
          </string-name>
          ,
          <article-title>Multilingual hope speech detection using machine learning (</article-title>
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>O.</given-names>
            <surname>Kolesnikova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. G.</given-names>
            <surname>Yigezu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gelbukh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Abitte</surname>
          </string-name>
          , G. Sidorov,
          <article-title>Detecting multilingual hate speech targeting immigrants and women on twitter</article-title>
          ,
          <source>Journal of Intelligent &amp; Fuzzy Systems</source>
          (????)
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>G. Y.</given-names>
            <surname>Bade</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Kolesnikova</surname>
          </string-name>
          , G. Sidorov,
          <string-name>
            <given-names>J.</given-names>
            <surname>Oropeza</surname>
          </string-name>
          ,
          <article-title>Social media hate and ofensive speech detection using machine learning method</article-title>
          ,
          <source>in: Proceedings of the Fourth Workshop on Speech, Vision</source>
          , and
          <article-title>Language Technologies for Dravidian Languages, Association for Computational Linguistics, St</article-title>
          .
          <source>Julian's, Malta</source>
          ,
          <year>2024</year>
          , pp.
          <fpage>240</fpage>
          -
          <lpage>244</lpage>
          . URL: https://aclanthology.org/
          <year>2024</year>
          .dravidianlangtech-
          <volume>1</volume>
          .
          <fpage>40</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M. G.</given-names>
            <surname>Yigezu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kanta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Kolesnikova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sidorov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gelbukh</surname>
          </string-name>
          , Habesha@ dravidianlangtech:
          <article-title>Abusive comment detection using deep learning approach</article-title>
          ,
          <source>in: Proceedings of the Third Workshop on Speech and Language Technologies for Dravidian Languages</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>244</fpage>
          -
          <lpage>249</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>G. Y.</given-names>
            <surname>Bade</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Kolesnikova</surname>
          </string-name>
          , G. Sidorov,
          <string-name>
            <given-names>J.</given-names>
            <surname>Oropeza</surname>
          </string-name>
          ,
          <article-title>Social media fake news classification using machine learning algorithm</article-title>
          ,
          <source>in: Proceedings of the Fourth Workshop on Speech, Vision</source>
          , and
          <article-title>Language Technologies for Dravidian Languages, Association for Computational Linguistics, St</article-title>
          .
          <source>Julian's, Malta</source>
          ,
          <year>2024</year>
          , pp.
          <fpage>24</fpage>
          -
          <lpage>29</lpage>
          . URL: https://aclanthology.org/
          <year>2024</year>
          .dravidianlangtech-
          <volume>1</volume>
          .4.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>F.</given-names>
            <surname>Balouchzahi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sidorov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gelbukh</surname>
          </string-name>
          ,
          <article-title>Polyhope: Two-level hope speech detection from tweets</article-title>
          ,
          <source>Expert Systems with Applications</source>
          <volume>225</volume>
          (
          <year>2023</year>
          )
          <fpage>120078</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>D.</given-names>
            <surname>García-Baena</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Balouchzahi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Butt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Á</surname>
          </string-name>
          .
          <string-name>
            <surname>García-Cumbreras</surname>
            ,
            <given-names>A. Lambebo</given-names>
          </string-name>
          <string-name>
            <surname>Tonja</surname>
            ,
            <given-names>J. A.</given-names>
          </string-name>
          <string-name>
            <surname>GarcíaDíaz</surname>
            , S. Bozkurt,
            <given-names>B. R.</given-names>
          </string-name>
          <string-name>
            <surname>Chakravarthi</surname>
            ,
            <given-names>H. G.</given-names>
          </string-name>
          <string-name>
            <surname>Ceballos</surname>
            , V.-
            <given-names>G.</given-names>
            Rafael, G.
          </string-name>
          <string-name>
            <surname>Sidorov</surname>
            ,
            <given-names>L. A.</given-names>
          </string-name>
          <string-name>
            <surname>Ureña-López</surname>
            ,
            <given-names>S. M.</given-names>
          </string-name>
          <string-name>
            <surname>Jiménez-Zafra</surname>
          </string-name>
          , Overview of hope at iberlef 2024:
          <article-title>Approaching hope speech detection in social media from two perspectives, for equality, diversity and inclusion and as expectations</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>73</volume>
          (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>M.</given-names>
            <surname>Dilawar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Patil</surname>
          </string-name>
          ,
          <article-title>Stress identification in social networks based on social interactions</article-title>
          ,
          <source>Open Access International Journal of Science &amp; Engineering</source>
          <volume>3</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>A.</given-names>
            <surname>Bartal</surname>
          </string-name>
          ,
          <string-name>
            <surname>K. M. Jagodnik</surname>
            ,
            <given-names>S. J.</given-names>
          </string-name>
          <string-name>
            <surname>Chan</surname>
            ,
            <given-names>M. S.</given-names>
          </string-name>
          <string-name>
            <surname>Babu</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Dekel</surname>
          </string-name>
          ,
          <article-title>Identifying women with postdelivery posttraumatic stress disorder using natural language processing of personal childbirth narratives</article-title>
          ,
          <source>American Journal of Obstetrics &amp; Gynecology MFM 5</source>
          (
          <year>2023</year>
          )
          <fpage>100834</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>G.</given-names>
            <surname>Sidorov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Balouchzahi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Butt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gelbukh</surname>
          </string-name>
          ,
          <article-title>Regret and hope on transformers: An analysis of transformers on regret and hope speech detection datasets</article-title>
          ,
          <source>Applied Sciences</source>
          <volume>13</volume>
          (
          <year>2023</year>
          )
          <fpage>3983</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Tash</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Kolesnikova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Ahani</surname>
          </string-name>
          , G. Sidorov,
          <article-title>Psycholinguistic and emotion analysis of cryptocurrency discourse on x platform</article-title>
          ,
          <source>Scientific Reports</source>
          <volume>14</volume>
          (
          <year>2024</year>
          )
          <fpage>8585</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>G. Y.</given-names>
            <surname>Bade</surname>
          </string-name>
          ,
          <article-title>Natural language processing and its challenges on omotic language group of ethiopia</article-title>
          ,
          <source>Journal of Computer Science Research</source>
          <volume>3</volume>
          (
          <year>2021</year>
          )
          <fpage>26</fpage>
          -
          <lpage>30</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>