<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Novel Machine Learning-based Sentiment Analysis Method for Chinese Social Media Considering Chinese Slang Lexicon and Emoticons</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Da Li</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rafal Rzepka</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michal Ptaszynski</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kenji Araki</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, Kitami Institute of Technilogy</institution>
          ,
          <addr-line>Kitami</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Graduate School of Information Science and Technology Hokkaido University</institution>
          ,
          <addr-line>Sapporo</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Internet slang is an informal language used in everyday online communication which quickly becomes adopted or discarded by new generations. Similarly, pictograms (emoticons/emojis) have been widely used in social media as a mean for graphical expression of emotions. People can convey delicate nuances through textual information when supported with emoticons. Furthermore, we also noticed that when people use new words and pictograms, they tend to express a kind of humorous emotion which is difficult to clearly classify as positive or negative. Therefore, it is important to fully understand the influence of Internet slang and emoticons on social media. In this paper, we propose a machine learning method considering Internet slang and emoticons for sentiment analysis of Weibo, the most popular Chinese social media platform. In the first step, we collected 448 frequent Internet slang expressions as a slang lexicon, then we converted the 109 Weibo emoticons into textual features creating Chinese emoticon lexicon. To test the capability of recognizing humorous posts, we utilized both lexicons with several machine learning approaches, k-Nearest Neighbors, Decision Tree, Random Forest, Logistic Regression, Na¨ıve Bayes and Support Vector Machine for detecting humorous expressions on Chinese social media. Our experimental results show that the proposed method can significantly improve the performance for detecting expressions which are difficult to polarize into positivenegative categories.</p>
      </abstract>
      <kwd-group>
        <kwd>sentiment analysis</kwd>
        <kwd>machine learning</kwd>
        <kwd>social media</kwd>
        <kwd>Internet slang emoticons</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Nowadays, people have become increasingly accustomed to expressing their opinions
online, especially on social media such as Twitter, Facebook or Weibo - the biggest
Chinese social media network that was launched in 2009. The rapid growth of such
platforms provides rich multimedia data in large quantities for various research
opportunities as sentiment analysis which focuses on automatic sentiment prediction on given
contents. Microblog data contain a vast amount of valuable sentiment information not
only for the commercial use, but also for psychology, cognitive linguistics or political
science. Sentiment analysis has been widely used in real world applications by
analyzing the online user-generated data, such as election prediction, opinion mining and
business-related activity analysis [25]. Sentiment analysis of microblogs becomes an
important area of research in the field of Natural Language Processing. Study of
sentiment in microblogs in English language has undergone major developments in recent
years [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Chinese sentiment analysis research, on the other hand, is still at relatively
early stage [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] especially when it comes to lexicons and emoticons usage.
      </p>
      <p>
        Pictograms (emoticons/emojis) have been widely used in social media as a mean for
graphical expression of emotions. According to the study about Instagram, emojis are
present in up to 57% of online messages in many countries3. For example, “face with
tears of joy”, an emoji that means that somebody is in an extremely good mood, was
regarded as the 2015 word of the year by The Oxford Dictionary [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. In our opinion
ignoring emoticons in sentiment research is unjustifiable, because they convey a
significant emotional information and play an important role in expressing emotions and
opinions in social media [
        <xref ref-type="bibr" rid="ref13 ref5">13, 5</xref>
        ].
      </p>
      <p>
        Internet slang is ubiquitous on the Internet. The emergence of new social contexts
like micro-blogs, question-answering forums, and social networks has enabled slang
and non-standard expressions to abound on the web. Despite this, slang has been
traditionally viewed as a form of non-standard language, a form of language that is not the
focus of linguistic analysis and has largely been neglected [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        Furthermore, we also noticed that when people use new words and pictograms, they
tend to express a kind of humorous emotion which is difficult to be easily classified as
positive or negative. It seems that some emoticons are used just for fun, self-mockery
or jocosity which expresses an implicit humor which might be characteristic to Chinese
culture. Emoticons and slang seem to play an important role in expressing this kind of
emotion. There is a high possibility that this phenomenon can cause a significant
difficulty in sentiment recognition task. Figure1 shows an example of a Weibo microblog
posted with emoticons and Internet slang. In the second line of the post, (lei
jue bu ai) is a Chinese informal contraction meaning (hen
lei, gan jue zi ji bu hui zai ai le which means “too tired for romance”). Such
abbreviations are popular and usually extracted from popular phrases and shorten into four
characters in general, and become a new chengyu, a type of traditional Chinese idiomatic
expression most of which consist of four characters. Chengyu are considered as
collected wisdom of the Chinese culture. Through the insights learned from chengyu, we
can express and discover wise men’s experiences, moral concepts, or admonishments
from the older generations of Chinese. Nowadays, chengyu still plays an important
role in Chinese conversations and education. When a new chengyu is introduced it can
also convey a humorous content. Examples of such abbreviations are: lei jue bu ai, ren
jian bu chai (life is so hard that some lies are better not exposed), xi da pu ben (news
so exhilarating that everyone is celebrating and spreading it around the world) and so
on. When it comes to emoticons, new ones are introduced by social media companies,
but their meaning can change with time. For example was originally and emoji
meant for expressing “bye-bye” gesture. However, it seems that gradually Weibo users
3 https://www.quintly.com/blog/instagram-emoji-study
started using this emoji for expressing artificial smile and refuse or self-mockery4. In
the research of [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], it was shown that this emoji expresses humorous emotion rather
than negative polarity. For example, in the following post: “After jogging, I’m starving.
Someone sent me a picture of kebab. I’m too tired for romance ”.
      </p>
      <p>To address this phenomenon, in this paper we focus on the Internet slang and
emoticons used on Weibo in order to establish if both slang and emoticons improve sentiment
4 https://qz.com/944693
analysis by recognizing humorous entries which are difficult to polarize. To perform
experiments, we collected 448 frequent Chinese Internet slang expressions as a slang
lexicon, then we converted 109 Weibo emoticons into textual features creating
Chinese emoticon lexicon. Then we utilized both lexicons with several machine learning
approaches, k-Nearest Neighbors, Decision Tree, Random Forest, Logistic Regression,
Na¨ıve Bayes and Support Vector Machine for detecting humorous expressions on
Chinese social media. Our experimental results show that the proposed method can
significantly improve the performance for detecting expressions which are difficult to polarize
into positive-negative categories.</p>
      <p>Our main contributions are as follows:
– We collected 448 frequent Chinese Internet slang expressions as a Chinese slang
lexicon.
– We converted the 109 Weibo emoticons into textual features creating Chinese
emoticon lexicon.
– We empirically confirmed implicit humor characteristic to Chinese culture visible
on Weibo and utilized both lexicons with several machine learning approaches for
detecting humorous expressions on Weibo and confirmed that using both slang and
emoticons improves previously proposed method.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Research</title>
      <p>At present, the sentiment analysis technology generally can be divided into two
categories: rule-based methods relying on sentiment lexicons, and machine learning-based
methods relying on annotated data.
2.1</p>
      <sec id="sec-2-1">
        <title>Rule-based methods</title>
        <p>
          Zhang et al. [24] proposed a rule-based approach with two phases: a) the sentiment of
each sentence is first decided based on word dependency to aggregate the sentences
sentiments and then b) the sentiment of each document is calculated. Zagibalov et al.
[23] presented a method that does not require any annotated corpus training data and
only requires information on commonly occurring negations and adverbials. Li et al.
stated that polarities and strengths judgment of sentiment words comply with a
Gaussian distribution, and thus proposed a Normal distribution-based sentiment computation
method which allows quantitative analysis of semantic fuzziness of sentiment words
in Chinese language [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. Zhuo et al. presented a novel approach based on the fuzzy
semantic model by using an emotion degree lexicon and a fuzzy semantic model [26].
Their model includes text preprocessing, syntactic analysis, and emotion word
processing. However, optimal results of Zhuo’s model were achieved only when the task was
clearly defined. Wu et al. presented an approach to leverage Web resources to construct a
English Slang Sentiment Dictionary (SlangSD) that is easy to expand [21]. They
empirically showed the advantages of using SlangSD, the newly-built slang sentiment word
dictionary for sentiment classification, and provided examples demonstrating its ease of
use with a sentiment analysis system.
2.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Machine learning-based methods</title>
        <p>
          Tan and Zhang conducted an empirical study of sentiment categorization on Chinese
documents [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]. They tested four features – mutual information, information gain,
chisquare, and document frequency; and five learning algorithms: centroid classifier,
kNearest Neighbor, Winnow classifier, Na¨ıve Bayes (NB) and Support Vector Machine
(SVM). Their results showed that the information gain and SVM features provided the
best performances for sentiment classification coupled with domain or topic
dependent classifiers. There are also researchers who have combined the machine learning
approach with the lexicon-based approach. Chen et al. proposed a novel sentiment
classification method which incorporated existing Chinese sentiment lexicon and
convolutional neural network [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. The results showed that their approach outperforms the
convolutional neural network (CNN) model only with word embedding features [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
However, all these approaches did not consider emoticons.
        </p>
        <p>
          Recently, a powerful system utilizing emoji in Twitter sentiment analysis model
called DeepMoji was proposed [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. Its creators trained 1,246 million tweets containing
one of 64 common emoticons by Bi-directional Long Short-Term Memory (Bi-LSTM)
model and applied it to interpret the meaning behind the online messages. DeepMoji
is also the most advanced sarcasm-detecting model, with an accuracy rate of 82.4%
even outperforming human detectors who managed to acquire 76.1% accuracy rate.
Sarcasm reverses the emotion of the literal text, therefore sarcasm-detecting capability
can play a significant role in sentiment analysis, especially in case of social media.
Although sarcasm and irony tend to convey negative emotions in general, we found
that in Chinese social media (Weibo in our example), in addition to the expression of
positive and negative emotions, people tend to express a kind of humorous emotion that
escapes the traditional bi-polarity.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Lexicon of Chinese Online Slang</title>
      <p>Chinese Internet slang is informal language used to express ideas on the Chinese
Internet in response to events, to mass media and foreign cultures. It also expresses a
natural human desire to simplify and update language. Slang that first appears on-line is
often adopted to become widely used in everyday life. It includes content relating to
all aspects of social life, mass media, economic, political situation etc. Internet slang
is arguably the fastest-changing aspect of a language, created by a number of different
influences, technology, mass media and foreign culture amongst others.</p>
      <p>Because Internet slang is not easy to extract automatically, it can cause a significant
difficulty in sentiment detecting task. For improving the performance of Chinese social
media sentiment analysis, we created a Chinese Internet slang lexicon (examples shown
in Table 1). We manually extracted 448 frequent Internet slang terms from the Internet
New Words Ranking List, Baidu Baike5, Wikipedia6 and social media systems such as
Baidu Tieba7 and Weibo8 between 2010 and 2018, and stored them as Chinese Internet
Slang Lexicon. After analysis we observed that the entries fall under seven following
categories:
– Numbers: such as 233 (“laughter/lol”: Chinese use 233 to express “can’t stop
laughing” because 233 is an emotional sign in a Chinese BBS site9 and the sign is the
NO.233 in the list of all emoticons); 213 (“a person who is very stupid”); 520/521
(“I love you”).
– Latin alphabet abbreviations: Chinese users commonly use a QWERTY keyboard
with pinyin enabled. Upper case letters are quick to type and require no
transformation. (Lower case letters spell words). Latin alphabet abbreviations (rather
than Chinese characters) are also sometimes used to evade censorship. Such as SB
(“dumb cunt”); YY (“fantasizing/sexual thoughts”); TT (“condom”).
– Chinese contractions: e.g. ren jian bu chai (“life is so hard that some lies are
better not exposed”: This comes from the lyrics of a song entitled “Shuo Huang”
(“Lies”), by Taiwanese singer Yoga Lin. This slang reflects that some people,
especially young people in China, are disappointed by reality); lei jue bu ai (“too tired
for romance”: this slang phrase is a literal abbreviation of the Chinese phrase “too
tired to fall in love anymore”. It originated from an article on the Douban website, a
Chinese social networking service website allowing registered users to record
information and create content related to film, books, music, recent events and activities
in Chinese cities. The article was posted by a 13-year-old boy who grumbled about
his single status and expressed his weariness and frustration towards romantic love.
The article went viral on the Chinese Internet, and the phrase was subsequently
used as a sarcastic way to convey depression when encountering misfortunes or
setbacks in life); gao da shang (“high-end, impressive, and high-class”: a popular
5 https://baike.baidu.com
6 https://en.wikipedia.org
7 https://tieba.baidu.com
8 https://www.weibo.com
9 https://www.mop.com
meme used to describe objects, people, behavior, or ideas that became popular in
late 2013).
– Neologisms: diao si (“loser”: The word diao si is used to describe young males who
were born into a poor family and are unable to improve their financial status. People
usually use this phrase in an ironic and self-deprecating way); ye shi zui le (“nothing
to say”: it is a way to gently express your frustration with someone or something
that is completely unreasonable and unacceptable); dan shen gou (“single dog”: a
term which single people in China use to poke fun at themselves for being single).
– Phrases with altered or extended meanings: hao or tu hao (“vulgar tycoon”: This
word refers to irritating online game players who buy large amounts of game weapons
in order to be gloried by others. Starting from late 2013, the meaning has changed
and now is widely used to describe nouveau riche people in China who are wealthy
but less cultured.); bei tai (“spare tire”: A girlfriend or boyfriend kept as a “backup”,
“plan B” , just in case of breaking up with the current partner).
– Puns and wordplay: (“river crab”: pun on , another Chinese characters
pronounced he xie, meaning “harmony”).
– Slang derived from foreign language: (The word gong kou comes from the
Japanese katakana ero, which translated from English “erotic” into the abbreviation
of the katakana , meaning “sensual”).
4</p>
    </sec>
    <sec id="sec-4">
      <title>Lexicon of Chinese Social Media Emoticons</title>
      <p>
        In the real-life (offline) dialogue between human beings, besides tone changes, we
usually express emotions with body language. In social networks, this can partially be
achieved by using emoticons [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        There are many unknown factors in constantly changing moods of human beings,
but communication with emoticons has become a global phenomenon. On the other
hand, because of different ethnic and cultural differences, misunderstandings when
using facial emoticons is not uncommon [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. We also have noticed previously
mentioned humorous emotion in Weibo microblog entries containing emoticons which are
often difficult to interpret as positive or negative. It seems that some emoticons are used
just for fun, self-mockery or jocosity which expresses an implicit humor characteristic
in Chinese culture. Emoticons seem to play an important role in expressing this kind of
emotion. There is a high possibility that this phenomenon can cause a significant
difficulty in sentiment detecting task, therefore we decided to build a lexicon of emoticons
before adding them to our system for classifying emotions in Weibo.
      </p>
      <p>When we collected microblog data, we discovered that Weibo emoticons are
transformed by API into Chinese characters, for example, will be convert into
(“smile”). This provided us with the possibility of building Chinese emoticon lexicon.
Therefore, we selected the 109 Weibo emoticons (see Figure 2) which can be
transformed into Chinese characters, and converted them into textual features to create
Chinese emoticon lexicon. Several examples are shown in Table 2.
Inspired by above mentioned works on Internet slang and emoticons, in order to test the
influence of them, we utilized both lexicons with several machine learning approaches,
k-Nearest Neighbors (k-NN), Decision Tree (DT), Random Forest (RF), Logistic
Regression (LR), Na¨ıve Bayes (NB) and Support Vector Machine (SVM) for detecting
humorous expressions on social media. We did not tested deep learning approaches as the data
size was not sufficient.</p>
      <p>In the first step, we add the Chinese slang lexicon and Chinese emoticon lexicon
to segmentation tool for matching new words and emoticons. Then we use the updated
tool to segment the sentences of large data set. Second, we apply the segmentation
results into the word embedding tool for training word vectors. Next, we apply the word
embedding model which considered Internet slang and emoticons to train a machine
learning model with training data. Finally, we input testing data into machine learning
model, and we can obtain the sentiment probability of a Weibo post which considers
the effect of emoticons and Internet slang.
6</p>
    </sec>
    <sec id="sec-5">
      <title>Experiments</title>
      <p>In order to verify the validity of our proposed method, we performed series of
experiments described below.
6.1</p>
      <sec id="sec-5-1">
        <title>Preprocessing</title>
        <p>
          Initializing word vectors with those obtained from an unsupervised neural language
model is a popular method to improve performance in the absence of a large supervised
training set. For our experiment we collected a large dataset (7.6 million posts) from
Weibo API from May 2015 to July 2017 to be used for calculating word embeddings.
First, we deleted the images, and videos treating them as noise. Second, we applied
Chinese Internet slang lexicon and Chinese emoticon lexicon into the dictionary of
Python Chinese word segmentation module Jieba10. Next, we used Jieba to segment the
sentences of the microblogs, and applied the segmentation results into the word2vec
model [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] for training word vectors. The vectors have dimensionality of 300 and were
trained using the continuous skip-gram model.
        </p>
        <p>Next, we collected 3,000 Weibo posts containing the emoticons. To use these posts
as our training data, we asked three Chinese native speakers to annotate them into two
categories: “humorous”, and “non-humorous”. After one annotator labelled polarities of
all posts, two other native speakers confirmed correctness of his annotations. Whenever
there was a disagreement, all decided the final polarity through discussion.
6.2</p>
      </sec>
      <sec id="sec-5-2">
        <title>Applied Classifiers</title>
        <p>
          Logistic Regression Logistic regression model is confirmed to be used in many tasks
such as document classification [22]. In Logistic regression model, we generally correct
overfitting with regularization. Regularization adds a penalty term on model to reduce
the freedom of the model. Hence, the model will be less likely to fit the noise of the
training data and will improve the generalization abilities of the model. We train the
model with L2 penalty regularization called Ridge regression in our experiments.
Support Vector Machine Support vector machine [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] is a supervised learning model
with associated learning algorithms that analyzes data used for classification. An SVM
model is a representation of the examples as points in space, mapped so that the
examples of the separate categories are divided by a clear gap that is as wide as possible.
New examples are then mapped into that same space and predicted to belong to a
category based on which side of the gap they fall. In addition, it uses kernel trick, implicitly
mapping their inputs into high-dimensional feature spaces. In our experiments, we used
the radial basis function kernel.
        </p>
        <p>
          Na¨ıve Bayes Na¨ıve Bayes classifier is based on applying Bayes theorem with strong
independence assumptions between the features. Na¨ıve Bayes has been studied
extensively since the 1950s. It was introduced under a different name into the text retrieval
community in the early 1960s, and remains a baseline method for text categorization
[
          <xref ref-type="bibr" rid="ref15">15</xref>
          ], the problem of judging documents as belonging to one category or the other with
word frequencies as the features. With appropriate pre-processing, it is competitive in
text classification task with more advanced methods including support vector machines.
In our experiments, we set the parameter of alpha to 0.01.
k-Nearest Neighbors In pattern recognition, the k-Nearest Neighbors algorithm is a
non-parametric method used for classification and regression. In both cases, the input
10 https://github.com/fxsjy/jieba
consists of the k closest training examples in the feature space [20]. The output depends
on whether k-NN is used for classification or regression. The number of neighbors is
set to 5 in our experiments.
        </p>
        <p>
          Random Forest Random forests is an ensemble learning method for classification,
regression and other tasks, that operate by constructing a multitude of decision trees at
training time and outputting the class that is the mode of the classes (classification) or
mean prediction (regression) of the individual trees [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. Random decision forests correct
for decision trees habit of overfitting to their training set.
        </p>
        <p>
          Decision Tree Decision tree is a decision support tool that uses a tree-like graph or
model of decisions and their possible consequences, including chance event outcomes,
resource costs, and utility. Decision tree is commonly used in operations research,
specifically in decision analysis to help identify a strategy most likely to reach a goal,
but is also a popular tool in machine learning [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ].
6.3
        </p>
      </sec>
      <sec id="sec-5-3">
        <title>Performance Test</title>
        <p>Using trained word2vec model, we passed word vectors of training data into the machine
learning models to train the model. We collected and annotated 300 Weibo entries with
emoticons as a test set, deleted images, and videos. Then we used the above mentioned
methods to calculate scores of the precision, recall and F1-score. We compared the
results of humorous detecting by machine learning only, machine learning considering
Internet slang only, and machine learning approaches considering emoticons only. The
results are shown in Table 3, Table 4 and Table 5, respectively. The Table 6 introduces
results of the experiment where both Internet slang and emoticons were used, and Table
7 shows the results of F1-score with above methods.</p>
        <p>The results show that considering Internet slang and emoticons. Limited to small
annotated data, the precision of the humor / non-humor classification was relatively
low, but by considering Internet slang and emoticons, the F1-score of each classifier
outperformed previous method by 1.39% (LR), 2.13% (SVM), 2.90% (NB), 0.69%
(k-NN), 0.84% (RF) and 3.89% (DT). Our proposed approach has improved the
performance showing that low-cost, small-scale data labeling is able to outperform widely
used state-of-the-art when emoticon and slang information is added to the learning
process.
7</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Considerations</title>
      <p>In our proposed approach, we paid more attention to the emoticons and Internet slang
in microblogs and investigated how adding these features separately and together
influences the previously proposed method for recognizing humorous posts which are
problematic when it comes to semantic analysis. Figure 3) shown an example of a microblog
which was correctly classified by our proposed method as “humorous” while the
baseline recognized it incorrectly as non-humorous. This post contains word (yi
ke sai ting which is a homophone of English word “exciting”). The baseline does not
know this expression and the parser divides it as (yi ke / sai ting which
means “a rowing boat”). When this expression is accompanied by emoticon, they
both improve the performance of classification and predict the implicit humorous
meaning.</p>
      <p>Error analysis showed that some posts were wrongly predicted due to proper nouns
missing in the parser’s dictionary which brought clearly negative impact on the results.
In Figure 4 we show an example of such misclassification into “non-humorous”
category annotated as “humorous” by annotators. Name of a ticketing website Da mai wang
was parsed incorrectly, and one shifted character caused mis-recognition of humorous
word. Weibo microblogs contain numerous ideograms deliberately altered from their
everyday meaning, what makes them difficult to parse and match. We think that adding
new named entities into the parser’s dictionary may significantly improve the results in
the future. We observed that when emotions are expressed online, emoticons might play
a greater role than it is usually considered, therefore we will experiment with weight of
the emoticons in the future.
In this paper, we proposed adding Chinese Internet slang and emoticons for automatic
classification of humorous posts on social media platform Weibo in order to
separate them from clearly positive and negative ones. We collected 448 frequent Internet
slang expressions and created a slang lexicon, then we converted the 109 Weibo
emoticons into textual features creating Chinese emoticon lexicon. To test the influence of
slang and emoticons on sentiment analysis task, we utilized both lexicons with several
machine learning-based classifiers, namely k-Nearest Neighbors, Decision Tree,
Random Forest, Logistic Regression, Na¨ıve Bayes and Support Vector Machine for
detecting humorous expressions on Chinese social media. Our experimental results show that
the proposed additions can significantly improve the F1-score for detecting humorous
expressions which are difficult to polarize into positive-negative categories.</p>
      <p>For improving the performance of the proposed method, in near future we are going
to increase the size of both slang and emoticon lexicons to improve further
classification results. Furthermore, we plan to add image processing for classifying stickers
which also seem to convey rich emotional information. Our ultimate goal is to
investigate how much the newly introduced features are beneficial for sentiment analysis by
feeding them to a deep learning model which should allow us to construct a high-quality
sentiment recognizer for wider spectrum of sentiment in Chinese language.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgment</title>
      <p>This work was supported by JSPS KAKENHI Grant Number 17K00295.
20. Weinberger, K.Q., Saul, L.K.: Distance metric learning for large margin nearest neighbor
classification. Journal of Machine Learning Research 10(Feb), 207–244 (2009)
21. Wu, L., Morstatter, F., Liu, H.: Slangsd: Building and using a sentiment dictionary of slang
words for short-text sentiment classification. arXiv preprint arXiv:1608.05129 (2016)
22. Yu, H.F., Huang, F.L., Lin, C.J.: Dual coordinate descent methods for logistic regression and
maximum entropy models. Machine Learning 85(1-2), 41–75 (2011)
23. Zagibalov, T., Carroll, J.: Automatic seed word selection for unsupervised sentiment
classification of chinese text. In: Proceedings of the 22nd International Conference on
Computational Linguistics-Volume 1. pp. 1073–1080. Association for Computational Linguistics
(2008)
24. Zhang, C., Zeng, D., Li, J., Wang, F.Y., Zuo, W.: Sentiment analysis of chinese documents:
From sentence to document level. Journal of the American Society for Information Science
and Technology 60(12), 2474–2487 (2009)
25. Zhao, P., Jia, J., An, Y., Liang, J., Xie, L., Luo, J.: Analyzing and predicting emoji usages
in social media. In: Companion of the The Web Conference 2018 on The Web Conference
2018. pp. 327–334. International World Wide Web Conferences Steering Committee (2018)
26. Zhuo, S., Wu, X., Luo, X.: Chinese text sentiment analysis based on fuzzy semantic model.</p>
      <p>In: Cognitive Informatics &amp; Cognitive Computing (ICCI* CC), 2014 IEEE 13th International
Conference on. pp. 535–540. IEEE (2014)</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Aldunate</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gonza´</surname>
          </string-name>
          lez-Iba´n˜ez, R.:
          <article-title>An integrated review of emoticons in computer-mediated communication</article-title>
          .
          <source>Frontiers in psychology 7</source>
          ,
          <year>2061</year>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gui</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          :
          <article-title>Combining convolution neural network and word sentiment sequence features for chinese text sentiment analysis</article-title>
          .
          <source>Journal of Chinese Information Processing</source>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Cortes</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vapnik</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Support-vector networks</article-title>
          .
          <source>Machine learning 20(3)</source>
          ,
          <fpage>273</fpage>
          -
          <lpage>297</lpage>
          (
          <year>1995</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Felbo</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mislove</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Søgaard</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rahwan</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lehmann</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>Using millions of emoji occurrences to learn any-domain representations for detecting sentiment, emotion and sarcasm</article-title>
          .
          <source>arXiv preprint arXiv:1708.00524</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Guibon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ochs</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bellot</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>From emojis to sentiment analysis</article-title>
          .
          <source>In: WACAI</source>
          <year>2016</year>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Ho</surname>
          </string-name>
          , T.K.:
          <article-title>Random decision forests</article-title>
          .
          <source>In: Document analysis and recognition</source>
          ,
          <year>1995</year>
          .,
          <source>proceedings of the third international conference on. vol. 1</source>
          , pp.
          <fpage>278</fpage>
          -
          <lpage>282</lpage>
          . IEEE (
          <year>1995</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Convolutional neural networks for sentence classification</article-title>
          .
          <source>arXiv preprint arXiv:1408.5882</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Kulkarni</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
          </string-name>
          , W.Y.:
          <article-title>Tfw, damngina, juvie, and hotsie-totsie: On the linguistic and social aspects of internet slang</article-title>
          .
          <source>arXiv preprint arXiv:1712.08291</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rzepka</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ptaszynski</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Araki</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Emoticon-aware recurrent neural network model for chinese sentiment analysis</article-title>
          .
          <source>In: The Ninth IEEE International Conference on Awareness Science and Technology (iCAST</source>
          <year>2018</year>
          )
          <article-title>(</article-title>
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Su</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>A method of polarity computation of chinese sentiment words based on gaussian distribution</article-title>
          .
          <source>In: International Conference on Intelligent Text Processing and Computational Linguistics</source>
          . pp.
          <fpage>53</fpage>
          -
          <lpage>61</lpage>
          . Springer (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Efficient estimation of word representations in vector space</article-title>
          .
          <source>arXiv preprint arXiv:1301.3781</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Moschini</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>The” face with tears of joy” emoji. a socio-semiotic and multimodal insight into a japan-america mash-up</article-title>
          .
          <source>HERMES-Journal of Language and Communication in Business (55)</source>
          ,
          <fpage>11</fpage>
          -
          <lpage>25</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Novak</surname>
            ,
            <given-names>P.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smailovic</surname>
            <given-names>´</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Sluban</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Mozeticˇ</surname>
          </string-name>
          , I.:
          <article-title>Sentiment of emojis</article-title>
          .
          <source>PloS one</source>
          <volume>10</volume>
          (
          <issue>12</issue>
          ),
          <year>e0144296</year>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Peng</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cambria</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hussain</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>A review of sentiment analysis research in chinese language</article-title>
          .
          <source>Cognitive Computation</source>
          <volume>9</volume>
          (
          <issue>4</issue>
          ),
          <fpage>423</fpage>
          -
          <lpage>435</lpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Rish</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          , et al.:
          <article-title>An empirical study of the naive bayes classifier</article-title>
          .
          <source>In: IJCAI 2001 workshop on empirical methods in artificial intelligence</source>
          .
          <source>vol. 3</source>
          , pp.
          <fpage>41</fpage>
          -
          <lpage>46</lpage>
          . IBM New York (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Rzepka</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Okumura</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ptaszynski</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Worlds linking faces - meaning and possibilities of contemporary pictograms</article-title>
          .
          <source>Journal of the Japanese Society for Artificial Intelligence</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Sharma</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kaur</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>Classification in pattern recognition: A review</article-title>
          .
          <source>International Journal of Advanced Research in Computer Science and Software Engineering</source>
          <volume>3</volume>
          (
          <issue>4</issue>
          ) (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Tan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , Zhang,
          <string-name>
            <surname>J.:</surname>
          </string-name>
          <article-title>An empirical study of sentiment analysis for chinese documents</article-title>
          .
          <source>Expert Systems with applications 34(4)</source>
          ,
          <fpage>2622</fpage>
          -
          <lpage>2629</lpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ji</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bao</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          :
          <article-title>A depression detection model based on sentiment analysis in micro-blog social network</article-title>
          .
          <source>In: Pacific-Asia Conference on Knowledge Discovery and Data Mining</source>
          . pp.
          <fpage>201</fpage>
          -
          <lpage>213</lpage>
          . Springer (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>