<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Exploiting Deep Neural Networks for Tweet-based Emo ji Prediction</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andrei Catalin Coman</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giacomo Zara</string-name>
          <email>giacomo.zarag@studenti.unitn.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yaroslav Nechaev</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gianni Barlacchi</string-name>
          <email>barlacchig@fbk.eu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alessandro Moschitti</string-name>
          <email>moschitti@disi.unitn.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Fondazione Bruno Kessler</institution>
          ,
          <addr-line>Trento</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Trento</institution>
          ,
          <addr-line>Trento</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <fpage>116</fpage>
      <lpage>128</lpage>
      <abstract>
        <p>For many years now, emojis have been used in social networks and chat services in order to enrich written text with auxiliary graphical content, achieving a higher degree of empathy. In particular, given the wide use of this medium, emojis are now available in such a number and variety that every basic concept and mood is covered by at least one. For this reason, the connection between the emoji and its semantical meaning grows stronger. In this paper, we will be describing the work performed in order to develop a Machine Learning based tool that, given a tweet, predicts the most likely emoji associated with the text. The task resembles the one presented by Barbieri et al., [23], and is placed within the context of the International Workshop on Semantic Evaluation (SemEval) 2018. We designed a baseline with standard Support Vector Machines and another baseline based on fastText, which was provided as part of the Workshop. In addition, we implemented several models based on Neural Networks such as Bidirectional Long Short-Term Memory Recurrent Neural Networks and Convolutional Neural Networks. We found that the latter is the most e ective since it outperformed all our models and ranks in the 6th position out of 47 total participants. Our work aims to illustrate the potential of simpler models, which, thanks to the ne-tuning of hyper-parameters, could achieve accuracy comparable to the more complex models of the challenge.</p>
      </abstract>
      <kwd-group>
        <kwd>emoji</kwd>
        <kwd>Twitter</kwd>
        <kwd>SVM</kwd>
        <kwd>fastText</kwd>
        <kwd>Bi-LSTM</kwd>
        <kwd>Text Classi cation</kwd>
        <kwd>CNN</kwd>
        <kwd>Machine Learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        We develop a Machine Learning model, which, given the text of a tweet, predicts
the most likely associated emoji among the most common 20 shown in Table 1.
In order to do this, we have rst implemented a baseline approach based on
Support Vector Machine (SVM), and used a fastText3 classi er as an
alternative baseline. After the implementation of the baselines, we proceeded to the
development of Neural Networks based classi ers. In particular, we have
chosen to test two di erent models: a Bidirectional Long Short-Term Memory
(Bi-LSTM) Recurrent Neural Network (RNN), and a Convolutional Neural
Network (CNN). The objective was to verify whether we would manage to
outperform the results obtained by the baselines, and which of the two models
would work better. For the whole task, we used the data provided within the
International Workshop on Semantic Evaluation 20184[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        Given the wide use of emojis that we are witnessing nowadays, not just on
Twitter5 but on many other Social Media and Instant Messaging services [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ],
one of the purposes of our paper consists in studying the relation between the
text written by the user and the emojis related to it, through the implementation
of a predicting tool that attempts to mimic the reasoning according to which a
certain emoji is evoked by the written message. This turns out to be quite a
complex problem since it is highly dependent on the understanding that human
beings have of emojis, which is hardly ever the same. This ambiguity is what most
strongly limits the e ectiveness of a prediction task performed on this domain.
In fact, the results obtained by the top-ranked participants as well as ours, which
are reported in Section 6 together with a discussion, show how di cult such a
task can be. Our work, however, aims to be an alternative to the more complex
models used by the other participants, with the objective of demonstrating that
even simpler models are able to easily overcome the baselines and compete with
the more performing ones, thanks to the ne-tuning of hyper-parameters.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Related work</title>
      <p>
        Obviously, a signi cant amount of work has been performed in the scope of
the SemEval competition, by the participants of its past editions. In
particular, it is worth to mention the work of Coltekin and Rama [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], who won the
competition by means of an SVM classi er. In their feature extraction
procedure, they took into account both the character-level and word-level in order to
extract the bag-of-n-grams features. Follows the work done by Baziotis et Al.
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], who used a Bidirectional LSTM (Bi-LSTM) with attention, and pre-trained
word2vec vectors, obtaining the highest recall. The third place is occupied by the
user hgsgnlp who has not provided any publication of his work and we send to
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] for the original reference. In the next position we nd Liu [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], who instead,
4 SemEval: http://alt.qcri.org/semeval2018/
5 Twitter: http://twitter.com/
proposed a Gradient Boosting Regression Tree approach combined with a
BiLSTM on character and word n-grams. Fifth position is occupied by Lu et. Al
[
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] who adopted a Bi-LSTM with attention and used di erent embeddings such
as word, char, Part of Speech (POS) and Named Entity Recognition (NER)
embeddings together with Twitter speci c features like punctuation, all-caps and
bag-of-hashtags. After them follows the work of Beaulieu and Owusu [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] who
combined SVM with bag-of-words approach, obtaining a very high precision.
The comparison between the above mentioned top 6 teams and our work can be
seen in Table 4. Please refer to [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] for the full table of results of all participants.
      </p>
      <p>
        Among the other participants it is worth mentioning Coster et Al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] who
built an SVM classi er on characters and word n-grams, obtaining the second
best results in the Spanish task. Basile and Lino [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] applied an SVM over a
variety of features (such as tf-idf, used in this paper too, POS tags and bigrams),
obtaining a competitive system. Finally, Jin and Pedersen [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] applied a soft
voting ensemble approach combining di erent Machine Learning algorithms such
as Multinomial Naive Bayes, Logistic Regression and Random Forests.
      </p>
      <p>
        Beside sequence-to-sequence models, another fast and performing approach
to solve NLP is given by the use of CNNs. These models demonstrated to achieve
state-of-the-art performances in many NLP tasks like sentence classi cation [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]
and answer reranking [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]. For instance, Zhao and Zeng [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ] obtained better
results in emoji prediction with Twitter data by applying CNNs. A related task
has also been performed by Felbo et Al [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], who extended supervised learning
to a set of labels represented by emojis themselves, in the scope of sentiment
analysis. Similarly, Wolny [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ] has extended binary sentiment analysis to
multiclass by means of emojis. Barbieri et Al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] performed a study on emojis, with
particular respect to their usage through time.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Problem statement and baseline</title>
      <p>As previously anticipated, the task faced in this project consisted of developing
a Machine Learning based prediction framework which, given the text of a tweet,
is able to predict the most likely associated emoji among the most common 20.
In order to train our models, we have been using the dataset provided for the
SemEval competition. This dataset has been obtained by selecting only tweets
that contain exactly one emoji, removing such emoji from the text and using
it as label.</p>
      <p>The training set consists of 475.624 labeled tweets, for a total of about
330.000 unique tokens. This is the dataset that we have been using to train both
the baseline and the Neural Network models. The test set, instead, consists of
50.000 labeled tweets, for a total of about 70.000 unique tokens, resembling
the distribution of the ones in the training set.
3.1</p>
      <sec id="sec-3-1">
        <title>Support Vector Machine baseline</title>
        <p>As mentioned in the introduction, one of the baselines has been implemented
as an SVM-based classi er. The whole preprocessing procedure has been
carried out by means of the functionalities provided by Apache Spark6, an analytic
engine for data processing based on distributed computation. For the classi
cation phase, we relied on LIBLINEAR7, a library for large-scale linear classi cation.</p>
        <p>In order to be fed to the classi er, our data have rstly been processed. In
particular, the text has been set to lowercase, and then tokenized using the
NLTK8 module for Python9. Once having done this, each tweet had to be turned
into a vector, in order for it to be classi ed. For this purpose, we have decided to
apply tf-idf (Term-Frequency Inverse-Document-Frequency) analysis. The result
of the tf-idf vectorization is a m n matrix M , where m is the number of
tweets, and n is the number of distinct words in the whole dataset. The cell
M [i; j] contains the tf-idf value of the j-th term with respect to the i-th tweet.
Such value is formally de ned as follows:
tfij =
nij
jdij
idfj = log</p>
        <p>jDj
jfd : j 2 dgj
tf idfij = tfij
idfj
(1)
(2)
(3)
where nij is the number of occurrences of term j in document i, jdij is the
size, in terms of words, of document i, jDj is the size, in terms of documents, of
the whole dataset, and jfd : j 2 dgj is the number of documents in the dataset
which contain j. This is based on the idea that a word is most likely to be
meaningful for a tweet if it appears as often as possible in the tweet, and as few
times as possible in other tweets.</p>
        <p>The vectorization has been performed applying di erent ways to extract
ngrams from the text, in terms of type and strategy. In particular, we have been
iterating over di erent parameters in terms of n-ngram size, n-gram type,
presence/absence of punctuation and minimum document frequency. The n-gram
size simply indicates how many single words are to be included in each feature,
while the n-gram type indicates whether the n-gram consists of a sequence of
consecutive words or if it involves skips.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2 fastText baseline</title>
        <p>
          FastText is an open-source library that allows users to learn unsupervised text
representations [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] and perform supervised text classi cations [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ].
        </p>
        <p>
          In terms of word representations, this library follows a very similar
approach to the one proposed in [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ]. Unlike word2vec, where each word is seen as
a single unit to which one must associate a vector, fastText considers a single
word as a character-level n-gram composition, where each n-gram has its own
6 Apache Spark: https://spark.apache.org/
7 LIBLINEAR: https://www.csie.ntu.edu.tw/~cjlin/liblinear/
8 NLTK: https://www.nltk.org/
9 Python: https://www.python.org/
representation, which must be learned. This allows, for example, rare words to
still have n-grams shared with other words. The same concept is applicable to
out-of-vocabulary words, i.e., words that are present in the test set but are
missing in the train set. The representation of a missing word as a concatenation
of the vectors of the n-grams of which it is composed can solve this problem
considering that presumably the representation of each n-gram has already been
learned at an earlier stage. Concerning the classi cation part, a sof tmax
function is used to compute the probability distribution over the prede ned classes.
For a set of T tweets, this leads to minimizing the negative log-likelihood over
the classes(emojis):
        </p>
        <p>T
1 X et log(sof tmax(BAxt)) (4)</p>
        <p>T t=1
where the rst weight matrix A is a look-up table over the words, B is the
weight matrix learned during the training phase, et represents the class(emoji)
of the corresponding tweet, and xt is the normalized bag of features of the t-th
tweet. The latter vector can be seen as the average of the n-grams embeddings
that form the corresponding tweet.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Deep learning models</title>
      <p>
        In this Section we present the Neural Network models we used to perform our
emoji predictions. The rst model is an architecture based on RNNs [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], where
the recurrent cell consists of a Bi-LSTM [
        <xref ref-type="bibr" rid="ref13 ref24">13,24</xref>
        ]. As second model we decided
to implement a CNN [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], since over the years it has proved to be e ective in
the Natural Language Processing area [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Some other concepts are also exposed
including the last layer used for predictions, the loss function applied for the
train, the regularization techniques employed to contain over tting and nally
the strategy adopted to stop the train.
4.1
      </p>
      <sec id="sec-4-1">
        <title>Bidirectional Long Short-Term Memory Recurrent Neural</title>
      </sec>
      <sec id="sec-4-2">
        <title>Networks</title>
        <p>Over the years, RNNs have proven to be very e ective in tasks that involve
data in the form of sequences. With the advent of LSTMs, they have become
even more powerful, as they have enabled us to tackle the problem of long-term
dependencies. They are capable of remembering information for long periods of
time. To do this, they maintain a state that is updated each time new information
is being examined. Information update can be seen as the transition from one
state, or rather from one cell state to another. The addition or removal of
information to the cell state is regulated by structures called gates, where each
has a speci c task. These structures can be described more formally by means
of the following formulae10:
10 Understanding LSTM</p>
        <p>Understanding-LSTMs</p>
        <p>Networks:</p>
        <p>http://colah.github.io/posts/2015-08which denotes the forget gate layer. Here we decide which information we
are going to remove from the previous cell state. In order to take this decision we
look at the output ht 1 of the previous cell and the current input xt, which in
our case consists of a token. The output of the sigmoid function is a vector of
numbers between 0 and 1 for each element in the cell state Ct 1. This number
indicates the importance of that piece of information. A value close to zero
indicates an insigni cant element, while a number close to one tells us that we
de nitely want to preserve it. In the above formula Wf indicates the weight
matrix and bf the bias value. These parameters are learned by the network
during the training phase.</p>
        <p>it = (Wi [ht 1; xt] + bi)
this gate is called the input gate layer, which decides which values we are
going to update in the cell state.</p>
        <p>Cet = tanh(WC [ht 1; xt] + bC )
after deciding which values to update, we need to generate new candidate
values to add to the state.</p>
        <p>here we update the old cell state Ct 1, thus producing the new cell state Ct.
(6)
(7)
(8)
(9)
(10)
Ct = ft</p>
        <p>Ct 1 + it</p>
        <p>Cet
ot = (Wo [ht 1; xt] + bo)</p>
        <p>ht = ot tanh(Ct)
the last step is to create the actual output of the cell. Through ot we decide
which values of the state cell will be output. Then the cell state values are
pushed between -1 and 1 through tanh function and by means of the
elementwise multiplication we output only those values we decided with ot.</p>
        <p>The memory capability of the LSTM-based RNNs acquires a very signi cant
importance when it comes to written text, which is our case of study, since the
ability to correctly interpreting one word is strongly based on the knowledge of
the context, i.e. the rest of the sentence. A further extension of this feature is
represented by the choice of making the LSTM layer bidirectional. The idea is
pretty simple: the recurrent layer is duplicated, resulting in two layers put "side
by side". The sequential input is then provided both in its original form and in a
reversed one. This is based on the idea that, in addition to the knowledge of the
past, also the knowledge of the future can be exploited for correctly interpreting
the information that is currently dealt with.
4.2</p>
      </sec>
      <sec id="sec-4-3">
        <title>Convolutional Neural Networks</title>
        <p>
          As with the previous model, CNNs have been used to solve various types of
problems. One of the rst applications of this neural-architecture was in the visual
domain, where they were employed for image and video recognition. However,
their versatility has led them to be used in Recommendation Systems [
          <xref ref-type="bibr" rid="ref27">27</xref>
          ] and
even in Natural Language Processing tasks [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. In contrast to classic Multilayer
Perceptron Neural Networks, CNNs have distinguished themselves through three
essential characteristics, namely:
{ number of parameters to train: in classic neural network architectures,
i.e. those with a lot of dense hidden layers and units, the number of
parameters to be trained grows really fast. With CNNs instead, the amount of train
parameters is given by the size and quantity of convolution lters, which in
some cases are drastically less
{ parameters sharing: a feature detector that's useful in one part of the
sentence is probably useful in another part of the sentence
{ sparsity of connections: in each convolution layer, each output value is
depends only on a small number of inputs
Those type of networks are based on the idea that the tweet text input, which
is put in form of a matrix of word embeddings, can be analyzed by means of a
sliding window of a xed size, which runs step by step through the matrix,
overlapping on di erent word embeddings, thus on di erent words of the original
text. For each step, a new single feature is computed by a linear combination
of the elements covered by the window and the weights of the corresponding
lter. Usually, after the convolution, a pooling layer is applied, which in our
case consists of a Max Pooling, whose objective is to preserve only the most
important features by taking the maximum from a well-de ned set of features
contained in the previous layer. The objective of the network is to learn the
weights of the various lters that have been used.
4.3
        </p>
      </sec>
      <sec id="sec-4-4">
        <title>Softmax, Loss Function and Regularization</title>
        <p>An additional fully connected softmax layer is added to the output layer of each
of the two models presented in the previous Subsections. It is responsible for
computing the probability distribution over the classes(emojis):
p(y = jjx) =
exT j
n=1 exT n
PN
(11)
where j is the target label and n is the weight vector of the n-th class. The
vector x instead, represents the abstract representation of the input tweet text
before the last softmax layer. Both models have been trained to minimize the
cross-entropy cost function shown in the following Equation:</p>
        <p>M M
Cost = log Y p(yijsi) = X[yi log si + (1 yi) log(1 si)] (12)
i=1 i=1
where s is the output of the softmax layer and contains all the parameters
optimized by the network.</p>
        <p>To avoid having a high variance, that is, a model that would learn a decision
function so complex that it could over t the training, we decided to increase
the cost function with L2-norm regularization terms for the parameters of the
network. The new cost function is updated as follows:
Cost =</p>
        <p>M
log Y p(yijsi) + jj jj22 =
i=1</p>
        <p>M
X[yi log si + (1 yi) log(1 si)] + jj jj2
2
i=1
where contains all the parameters optimized by the network.</p>
        <p>
          As a further method of regularization we decided to use the Dropout technique,
which prevents feature co-adaptation by setting to zero (dropping out) a portion
of hidden units during the forward propagation phase [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ]. In order to be able to
decide where to stop training the models, we have adopted an Early Stopping
strategy. This technique can be seen as an additional method of regularization
that prevents over tting. We extracted a validation set from the training set
and monitored its improvements in terms of F1-score. We then set a patience
value of 2, which means that if there is no improvement for three epochs in a
row, the training ends.
(13)
5
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Experiments</title>
      <p>We performed an extensive set of experiments, intended to tune our models
and identify the combination of hyper-parameters and word embeddings which
would lead to the best results.
5.1</p>
      <sec id="sec-5-1">
        <title>Word embeddings</title>
        <p>
          The rst important aspect to take into account when applying a Neural Network
model to textual data is is the logic according to which words are embedded for
the system. For our set of experiments, we have decided to apply three di erent
strategies:
{ randomly initialized word embeddings: the embeddings for the word
are initialized randomly, and subsequently learned by the network together
with the weights
{ GloVe11 embeddings: the embeddings are those provided by the GloVe
algorithm[
          <xref ref-type="bibr" rid="ref22">22</xref>
          ], and they are maintained over the whole process
{ trainable GloVe embeddings: the embeddings are those provided by the
GloVe algorithm, but they can then be modi ed by the learning process of
the network
11 GloVe: https://nlp.stanford.edu/projects/glove/
5.2
        </p>
      </sec>
      <sec id="sec-5-2">
        <title>Neural Networks setup</title>
        <p>In this subsection, we will describe the combination of Neural Network
parameters which eventually worked best, after an iterative experimental process
performed with respect to the aspects described in Section 4. The models are t on
the data, including as parameter the weights of the classes, computed according
to their occurrences in the dataset. The training is performed over a maximum
of 50 epochs, with batches of size 2048.</p>
        <p>Layer</p>
        <p>Parameters</p>
        <p>Value
Embedding</p>
        <p>Dropout
1D convolution
1D max pooling</p>
        <p>Flatten
Dropout
Dense
As for the CNN scheme, we came up with a model that has a 1D convolution
layer at its core, and it is structured as shown in Table 2. The RNN has a
BiLSTM layer as its core, and it is structured as shown in Table 3. Both models
are then compiled using cross entropy as loss function, accuracy as metric
and Adam as optimizer, with the following parameters:
{ learning rate: 0.001
{ 1: 0.9
{ 2: 0.999
{ learning rate decay over each update: 0.001</p>
        <p>
          Accuracy Precision Recall F1
Experiment
SVM Baseline
CNN (no GloVe)
CNN ( xed GloVe)
CNN (trainable GloVe)
LSTM (no GloVe)
LSTM ( xed GloVe)
LSTM (trainable GloVe)
#1 Tubingen-Oslo [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] 0.4709
#2 NTUA-SLP [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] 0.4474
#3 hgsgnlp [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] 0.4555
#4 EmoNLP [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] 0.4746
#5 ECNU [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] 0.4630
#6 UMDuluth-CS8761 [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] 0.4573
#7 SemEval Baseline [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] 0.4256
In this section, we will be describing and discussing the results we have obtained
by applying the models described in the previous sections. For each subtask we
have performed a complete set of experiments, tuning the di erent parameters
involved in the algorithm. Here we will report the nal results for the baselines
and each embedding strategy adopted for the Neural Network approach. We also
include in the table the scores achieved by the top 5 participants of the SemEval
competition, for a more interesting comparison. The results are show in Table 4.
As we can see, the fastText implementation outperformed the SVM one,
providing a solid baseline to start our work from. The Deep Learning approach
has generally outperformed both the baselines, indicating the trainable GloVe
embeddings as the best choice. Between the two Neural Network strategies, the
CNN model turned out to work better than the Bi-LSTM one. As for the
comparison with the SemEval ranking, as we can see, our algorithm would earn the
6th position.
7
        </p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusions</title>
      <p>The task described in this paper consisted in studying the relation between
written text and emojis, with particular respect to the scope of the social network
Twitter. This was done by implementing and evaluating a Machine Learning
based tool capable of predicting, given the text of a tweet, its most likely
related emoji, according to the task proposed for the SemEval 2018 competition.
In order to achieve this target, we rst implemented a SVM-based baseline,
comparing it with the fastText algorithm. Subsequently, we have implemented
two di erent Neural Network architectures, based respectively on CNNs and
Bi-LSTM RNNs. Our baseline was outperformed by the fastText based one,
and together they provided a fair starting point. The Neural Networks themselves
outperformed the baseline, indicating the CNN scheme and the GloVe word
embeddings as the best choice. The results obtained turned out to be positively
competitive with respect to the ranking of the competition, since they would
place our work in the 6th position, with 47 total participants. This shows us
that simple deep neural networks can also reach quite good results by improving
them through ne-tuning of hyper-parameters, regularization, and optimization
of the models themselves. In terms of future work, it would be interesting to
improve the results by more thoroughly tuning the hyper-parameters of the Neural
Networks or by exploring new models, and to participate to the next edition of
the SemEval competition.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Francesco</given-names>
            <surname>Barbieri</surname>
          </string-name>
          , Jose Camacho-Collados, Francesco Ronzano, Luis Espinosa Anke, Miguel Ballesteros, Valerio Basile, Viviana Patti, and
          <string-name>
            <given-names>Horacio</given-names>
            <surname>Saggion</surname>
          </string-name>
          .
          <article-title>Semeval 2018 task 2: Multilingual emoji prediction</article-title>
          .
          <source>In Proceedings of The 12th International Workshop on Semantic Evaluation</source>
          , pages
          <volume>24</volume>
          {
          <fpage>33</fpage>
          . Association for Computational Linguistics,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Francesco</given-names>
            <surname>Barbieri</surname>
          </string-name>
          , Lu s Marujo, Pradeep Karuturi,
          <string-name>
            <given-names>William</given-names>
            <surname>Brendel</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Horacio</given-names>
            <surname>Saggion</surname>
          </string-name>
          .
          <article-title>Exploring emoji usage and prediction through a temporal variation lens</article-title>
          .
          <source>CoRR</source>
          , abs/
          <year>1805</year>
          .00731,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Angelo</given-names>
            <surname>Basile</surname>
          </string-name>
          and
          <string-name>
            <given-names>Kenny W.</given-names>
            <surname>Lino</surname>
          </string-name>
          . Tajjeb at semeval
          <article-title>-2018 task 2: Traditional approaches just do the job with emoji prediction</article-title>
          .
          <source>pages 470{476</source>
          , 01
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Christos</given-names>
            <surname>Baziotis</surname>
          </string-name>
          , Athanasiou Nikolaos, Athanasia Kolovou, Georgios Paraskevopoulos, Nikolaos Ellinas, and
          <string-name>
            <given-names>Alexandros</given-names>
            <surname>Potamianos</surname>
          </string-name>
          .
          <article-title>Ntua-slp at semeval-2018 task 2: Predicting emojis using rnns with context-aware attention</article-title>
          .
          <source>In SemEval@NAACL-HLT</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Jonathan</given-names>
            <surname>Beaulieu</surname>
          </string-name>
          and Dennis Asamoah Owusu.
          <article-title>Umduluth-cs8761 at semeval2018 task 2: Emojis: Too many choices</article-title>
          ?
          <source>In Proceedings of The 12th International Workshop on Semantic Evaluation</source>
          , pages
          <volume>400</volume>
          {
          <fpage>404</fpage>
          . Association for Computational Linguistics,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Piotr</given-names>
            <surname>Bojanowski</surname>
          </string-name>
          , Edouard Grave, Armand Joulin, and
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          .
          <article-title>Enriching word vectors with subword information</article-title>
          .
          <source>arXiv preprint arXiv:1607.04606</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Ronan</given-names>
            <surname>Collobert</surname>
          </string-name>
          and
          <string-name>
            <given-names>Jason</given-names>
            <surname>Weston</surname>
          </string-name>
          .
          <article-title>A uni ed architecture for natural language processing: Deep neural networks with multitask learning</article-title>
          .
          <source>In Proceedings of the 25th International Conference on Machine Learning, ICML '08</source>
          , pages
          <fpage>160</fpage>
          {
          <fpage>167</fpage>
          , New York, NY, USA,
          <year>2008</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>Cagr</given-names>
            <surname>Co</surname>
          </string-name>
          <article-title>ltekin and Taraka Rama. Tubingen-oslo at semeval-2018 task 2: Svms perform better than rnns in emoji prediction</article-title>
          .
          <source>In Proceedings of The 12th International Workshop on Semantic Evaluation</source>
          , pages
          <volume>34</volume>
          {
          <fpage>38</fpage>
          . Association for Computational Linguistics,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>Alexis</given-names>
            <surname>Conneau</surname>
          </string-name>
          , Holger Schwenk, Loc Barrault, and
          <string-name>
            <given-names>Yann</given-names>
            <surname>Lecun</surname>
          </string-name>
          .
          <article-title>Very deep convolutional networks for text classi cation</article-title>
          .
          <source>pages 1107{1116</source>
          , 01
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Jol</surname>
            <given-names>Coster</given-names>
          </string-name>
          , Reinder Gerard Dalen, and Nathalie Adrinne Jacqueline Stierman.
          <article-title>Hatching chick at semeval-2018 task 2: Multilingual emoji prediction</article-title>
          .
          <source>pages 445{ 448</source>
          , 01
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Bjarke</surname>
            <given-names>Felbo</given-names>
          </string-name>
          , Alan Mislove, Anders Sgaard, Iyad Rahwan, and
          <string-name>
            <given-names>Sune</given-names>
            <surname>Lehmann</surname>
          </string-name>
          .
          <article-title>Using millions of emoji occurrences to learn any-domain representations for detecting sentiment, emotion and</article-title>
          <string-name>
            <surname>sarcasm. 08</surname>
          </string-name>
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Tim</surname>
          </string-name>
          <article-title>High eld and Tama Leaver. Instagrammatics and digital methods: studying visual social media, from sel es and gifs to memes and emoji</article-title>
          .
          <source>Communication Research and Practice</source>
          ,
          <volume>2</volume>
          (
          <issue>1</issue>
          ):
          <volume>47</volume>
          {
          <fpage>62</fpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <article-title>Sepp Hochreiter and Jurgen Schmidhuber. Long short-term memory</article-title>
          .
          <source>Neural computation</source>
          ,
          <volume>9</volume>
          (
          <issue>8</issue>
          ):
          <volume>1735</volume>
          {
          <fpage>1780</fpage>
          ,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>J. J.</surname>
          </string-name>
          <article-title>Hop eld. Neural networks and physical systems with emergent collective computational abilities</article-title>
          .
          <source>Proceedings of the National Academy of Sciences of the United States of America</source>
          ,
          <volume>79</volume>
          (
          <issue>8</issue>
          ):
          <volume>2554</volume>
          {
          <fpage>2558</fpage>
          ,
          <string-name>
            <surname>April</surname>
          </string-name>
          <year>1982</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <given-names>Shuning</given-names>
            <surname>Jin</surname>
          </string-name>
          and
          <string-name>
            <given-names>Ted</given-names>
            <surname>Pedersen</surname>
          </string-name>
          .
          <article-title>Duluth urop at semeval-2018 task 2: Multilingual emoji prediction with ensemble learning and oversampling</article-title>
          .
          <source>In Proceedings of The 12th International Workshop on Semantic Evaluation</source>
          , pages
          <volume>482</volume>
          {
          <fpage>485</fpage>
          . Association for Computational Linguistics,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Armand</surname>
            <given-names>Joulin</given-names>
          </string-name>
          , Edouard Grave, Piotr Bojanowski, and
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          .
          <article-title>Bag of tricks for e cient text classi cation</article-title>
          .
          <source>arXiv preprint arXiv:1607.01759</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <given-names>Yoon</given-names>
            <surname>Kim</surname>
          </string-name>
          .
          <article-title>Convolutional neural networks for sentence classi cation</article-title>
          .
          <source>arXiv preprint arXiv:1408.5882</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Yann</surname>
            <given-names>LeCun</given-names>
          </string-name>
          and
          <string-name>
            <given-names>Yoshua</given-names>
            <surname>Bengio</surname>
          </string-name>
          .
          <article-title>The handbook of brain theory and neural networks. chapter Convolutional Networks for Images, Speech,</article-title>
          and Time Series, pages
          <volume>255</volume>
          {
          <fpage>258</fpage>
          . MIT Press, Cambridge, MA, USA,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <given-names>Man</given-names>
            <surname>Liu</surname>
          </string-name>
          . Emonlp at semeval
          <article-title>-2018 task 2: English emoji prediction with gradient boosting regression tree method and bidirectional lstm</article-title>
          .
          <source>pages 390{394</source>
          , 01
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Xingwu</surname>
            <given-names>Lu</given-names>
          </string-name>
          , Xin Mao,
          <string-name>
            <given-names>Man</given-names>
            <surname>Lan</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Yuanbin</given-names>
            <surname>Wu</surname>
          </string-name>
          . Ecnu at semeval
          <article-title>-2018 task 2: Leverage traditional nlp features and neural networks methods to address twitter emoji prediction task</article-title>
          .
          <source>In Proceedings of The 12th International Workshop on Semantic Evaluation</source>
          , pages
          <volume>433</volume>
          {
          <fpage>437</fpage>
          . Association for Computational Linguistics,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Tomas</surname>
            <given-names>Mikolov</given-names>
          </string-name>
          , Kai Chen, Greg Corrado, and
          <article-title>Je rey Dean. E cient estimation of word representations in vector space</article-title>
          .
          <source>CoRR, abs/1301.3781</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22. Je rey Pennington, Richard Socher, and
          <string-name>
            <given-names>Christopher D.</given-names>
            <surname>Manning</surname>
          </string-name>
          . Glove:
          <article-title>Global vectors for word representation</article-title>
          .
          <source>In Empirical Methods in Natural Language Processing (EMNLP)</source>
          , pages
          <fpage>1532</fpage>
          {
          <fpage>1543</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Horacio</surname>
            <given-names>Saggion</given-names>
          </string-name>
          , Miguel Ballesteros, and
          <string-name>
            <given-names>Francesco</given-names>
            <surname>Barbieri</surname>
          </string-name>
          .
          <article-title>Are emojis predictable</article-title>
          ? In EACL,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <given-names>M.</given-names>
            <surname>Schuster</surname>
          </string-name>
          and
          <string-name>
            <surname>K.K. Paliwal.</surname>
          </string-name>
          <article-title>Bidirectional recurrent neural networks</article-title>
          .
          <source>Trans. Sig. Proc.</source>
          ,
          <volume>45</volume>
          (
          <issue>11</issue>
          ):
          <volume>2673</volume>
          {
          <fpage>2681</fpage>
          ,
          <string-name>
            <surname>November</surname>
          </string-name>
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Aliaksei</surname>
            <given-names>Severyn</given-names>
          </string-name>
          , Massimo Nicosia, Gianni Barlacchi, and
          <string-name>
            <given-names>Alessandro</given-names>
            <surname>Moschitti</surname>
          </string-name>
          .
          <article-title>Distributional neural networks for automatic resolution of crossword puzzles</article-title>
          .
          <source>In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing</source>
          (Volume
          <volume>2</volume>
          :
          <string-name>
            <surname>Short</surname>
            <given-names>Papers)</given-names>
          </string-name>
          , volume
          <volume>2</volume>
          , pages
          <fpage>199</fpage>
          {
          <fpage>204</fpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Nitish</surname>
            <given-names>Srivastava</given-names>
          </string-name>
          , Geo rey Hinton,
          <source>Alex Krizhevsky</source>
          , Ilya Sutskever, and
          <string-name>
            <given-names>Ruslan</given-names>
            <surname>Salakhutdinov</surname>
          </string-name>
          .
          <article-title>Dropout: A simple way to prevent neural networks from over tting</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          ,
          <volume>15</volume>
          :
          <year>1929</year>
          {
          <year>1958</year>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27. Aaron van den Oord, Sander Dieleman, and
          <string-name>
            <given-names>Benjamin</given-names>
            <surname>Schrauwen</surname>
          </string-name>
          .
          <article-title>Deep contentbased music recommendation</article-title>
          . In C. J.
          <string-name>
            <surname>C. Burges</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Bottou</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Welling</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Ghahramani</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          K. Q. Weinberger, editors,
          <source>Advances in Neural Information Processing Systems</source>
          <volume>26</volume>
          , pages
          <fpage>2643</fpage>
          {
          <fpage>2651</fpage>
          . Curran Associates, Inc.,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <given-names>Wieslaw</given-names>
            <surname>Wolny</surname>
          </string-name>
          .
          <article-title>Emotion analysis of twitter data that use emoticons and emoji ideograms</article-title>
          .
          <source>In ISD</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <given-names>Luda</given-names>
            <surname>Zhao</surname>
          </string-name>
          and
          <string-name>
            <given-names>Connie</given-names>
            <surname>Zeng</surname>
          </string-name>
          .
          <article-title>Using neural networks to predict emoji usage from twitter data</article-title>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>