<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Stock article title sentiment-based classification using PhoBERT⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>n Son Tung[</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>n Ngo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Long[</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>ng Tr</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>n Thu Th</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Duong T. Thu Phuong[</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>n Nguy</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>National Economics University</institution>
          ,
          <addr-line>Hanoi</addr-line>
          ,
          <country country="VN">Vietnam</country>
        </aff>
      </contrib-group>
      <fpage>225</fpage>
      <lpage>233</lpage>
      <abstract>
        <p>Text classification is a typical and important part of supervised learning, it has several applications in economics and attracted the attention of many stock market investors. For a long time, the news is frequently an unanticipated stock investment variable that instantaneously influences stock price directions. In front of an enormous volume of news, investors are always searching for models that automatically categorize news quickly and accurately. Thus, in this research, we have utilized diferent models like PhoBERT, SVM, Logistic Regression, LSTM, Random Forest, and Naive Bayes to classify news articles into three categories [negative, neutral, or positive] based on their titles. The results demonstrated that after training with a dataset of over 1000 news samples from CafeF.vn, the PhoBERT model outperformed other models with an accuracy up to 93%. The code and dataset is available at https://github.com/209sontung/Vietnamese-stock-article-classification.</p>
      </abstract>
      <kwd-group>
        <kwd>classification</kwd>
        <kwd>• PhoBERT • sentiment analysis • stock arti- cles</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Text classification is a traditional word processing problem using machine
learning. The first idea is to map a text to a known topic from a finite set of topics
based on the semantics of the text. Text documents are typically used for
classiifcation, which is done based on selected documents and features. However, the
classes are chosen before the experiment analysis, which is referred to as
supervised machine learning operations. Due to the growing number of documents,
the demand for text classification is expanding, and the tasks are getting
increasingly diverse, such as sentiment analysis of reviews and news categorization, so
on. An article in a newspaper, for example, could fall under one (or more) of
these categories (such as sports, health, information technology, etc.). The use of
spam filtering in email [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], web services [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ][
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], fake currency identification [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ],
fake news identification [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], and opinion mining techniques [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] are also prominent
and important applications in this field. Automatically categorizing text into a
certain topic makes it easier to organize, store, and query documents later.
      </p>
      <p>Forecasting is always a dificult task in the stock market because it is highly
volatile and dynamic. Many methods have been proposed to forecast the future
direction of the stock market. Financial news, for instance, has an immediate
positive or negative impact on stock prices. Before purchasing a stock, investors,
for instance, evaluate a company based on its activities on its oficial website and
ifnancial news about the company. However, investors can not fully assess such
vast amounts of financial news data on their own. As a result, investors require
a model that can assist them in quickly sorting through financial news articles.</p>
      <p>
        In this research, we collected a dataset of over 1000 financial news
articles about the stock market from the website CafeF.vn. Afterward, we used
LSTM [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], PhoBert[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], SVM[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], and other models to categorize news articles
from the above dataset into three categories [positive, negative, and neutral].
The PhoBERT model provided outstanding accuracy results of up to 93% after
training.
      </p>
      <p>The remainder of this paper is organized as follows. Section 2 introduces
related works . The proposed model is introduced in Section 3. Section 4 discusses
the results of experiments and is followed by a conclusion in Section 5.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related works</title>
      <p>Sentiment analysis, also referred to as a classification task, aims to forecast the
overall sentiment of a text, which could be a tweet or a review of a movie or
product. The main goal is to determine if the text’s conveyed impression is
positive or negative, in some cases, with a score or confidence metric.</p>
      <p>
        In English, a lot of publications on sentiment analysis have been undertaken.
For the problem of sentiment classification, Pang et al compared multiple
supervised learning algorithms [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], including Naive Bayes [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], KNN [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], Maximum
Entropy Models [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], and Support Vector Machines. They tested several types
of features and achieved the highest accuracy of 82.9% on a corpus of movie
reviews. Zhou utilized the Stanford Sentiment Treebank (SST) dataset [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] to
describe the sentiment categorization of movie reviews. In comparison to
multilayered CNN [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] and RNN [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] models, the architecture combining CNN and
LSTM models produced better performance. The two classes (positive and
negative) dataset had an accuracy of 87.7%, whereas the vfie classes (very positive,
positive, neutral, very negative, negative) dataset had a 49.2%. In another
sentiment analysis study performed on the SST dataset, Manish Munikar et al. [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]
applied BERT [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] - the latest state-of-the-art in the NLP [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] field proposed by
Google in 2018. The architecture contained a dropout regularization and
softmax classifier layers on top of the pre-trained BERT layer. Their proposed model
was presented that achieved the highest of 94.7% correct in SST-2 and 84.2%
when performed in SST-5, surpassing every aforementioned technique.
      </p>
      <p>
        In Vietnamese, Kieu and Pham [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] performed studies on a corpus of
computer product reviews by ofering a rule-based system for sentiment
classification in Vietnam utilizing the GATE framework [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. This approach reached
67.35% precision overall, but designing the rules seems to be a dificult and
time-consuming task. Quan et al. [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] presented a multi-channel LSTM - CNN
model for Vietnamese sentiment analysis that combines Long Short-Term
Memory (LSTM) and CNN. This combination had an accuracy of 87.72% on the VS
dataset and 59.61% on VLSP, which was also proposed in their research.
      </p>
      <p>In this paper, we proposed a Vietnamese stock news sentiment classification
model, which is a novel approach to sentiment analysis in Vietnamese. The
proposed model achieved 93.12% accuracy and was constructed using a
pretrained PhoBERT, a state-of-the-art language model for Vietnamese based on
BERT architecture. In addition, we built a dataset that included 1000 titles
of financial articles taken from CafeF.vn and labeled them into three groups
[negative, neutral, or positive].
3</p>
    </sec>
    <sec id="sec-3">
      <title>Proposed Model</title>
      <p>Our proposed model consists of 2 stages described in Fig. 1. The input is the
ifnancial article’s headlines, which will proceed through the first stage to
preprocess the data to convert it into a format that the PhoBERT model can
understand and improve its accuracy. Following that, in the second stage, our
PhoBERT-based model will be tasked with assessing content from the header
broadcast and categorizing it into one of three classes represented as -1, 0, or 1
(i.e., -1 as negative, 0 as neutral, and 1 as positive direction).</p>
      <sec id="sec-3-1">
        <title>Preprocessing</title>
        <p>
          The preprocessing procedure was separated into two phases. In Phase 1, first, we
applied VnCoreNLP’s Named entity recognition [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ] to extract all the proper
nouns and replace those words that signify location with the word ”loc” or
”name” for the organization name, stock code, or person’s name. To avoid any
confusion when the model predicts, the punctuation was then removed, hence
increasing the model’s accuracy. Considering the fact that white space is also
utilized to separate syllables that make up words in Vietnamese, in the last step
of Phase 1, we adopted Rdrsegmenter from VnCoreNLP to separate words for
input data. Furthermore, as an input for the PhoBERT model the title needed
to be tokenized, therefore we utilized BPE tokenizer [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ].
        </p>
        <p>In Phase 2, we had the symbol vocabulary with the character vocabulary,
and each word was represented as a sequence of characters with a unique
end-ofword symbol ” &lt; /s &gt; ” that allowed us to recover the original tokenization after
translation. In example, we counted all symbol pairs iteratively and replaced each
occurrence of the most common pair (”A”, ”B”) with the new symbol ”AB”.
Each merge process generates a new symbol that represents an n-gram of
characters. BPE does not require a shortlist because frequently occurring character
n-grams (or complete words) are finally combined into a single symbol. Thus,
the amount of the final symbol vocabulary is equal to the original vocabulary.</p>
        <p>Then we mapped each subword to its corresponding ID in the PhoBERT
vocabulary, and because each title is varied in length, we employed pad sequences
to match them all in length. i.e. sentences that shorter than 125 subwords are
padded with 0 at the end, while longer are trimmed to produce 125.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Training details</title>
        <p>First, we installed all the necessary materials included transformers library,
VnCoreNLP Python wrapper and its word segmentation component (i.e.
RDRSegmenter), then fastBPE to convert the input text into a list of subwords. After
train and test dataset are done prepared as described in Section 3.1, DataLoader
is created to load data into the model. Then we loaded P hoBERTBASE from
HuggingFace’s transformers library as the pre-trained model. The optimizer we
chose for the training stage is AdamW optimizer, which is an improved version
of Adam optimizer from the transformers library. In addition, this study used
batch size = 32 with 10 epochs divided into two stages with diferent learning
rate. The initially learning rate which we utilized was α = 5e − 6 in the first 5
epochs in order for the loss to converge faster. Once the first training phase was
completed, we reduced the learning rate to α = 5e − 7 to achieved the smallest
potential loss. The model was then trained for another 5 epochs since the loss
on the validation dataset appears to stabilize after this number of cycles.</p>
        <p>Finally, the softmax classicfiation layer (which includes three nodes
corresponding to three classes in the dataset) will output the probabilities of the
input text belonging to each of the class labels, with the total of the
probabilities equal to 1. The dense layer consists of a fully connected neural network with
the softmax activation function. The softmax function σ : RK
in (1).
→ RK is given
σ (z)i =</p>
        <p>
          ezi
PK
j=1 ezj
f or i = 1, ..., K
(1)
where z = (z1,...,zK ) ∈ RK is the softmax layer’s intermediate output (also
called logits). The predicted label for the input is then chosen from the output
node with the highest likelihood. The output of the proposed model will be
represented as -1, 0, or 1. The entire training process was deployed on PyTorch
[
          <xref ref-type="bibr" rid="ref26">26</xref>
          ] framework.
3.3
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>Method</title>
        <p>
          BERT. BERT stands for Bidirectional Encoder Representations from
Transformer. It is a transformer-based [
          <xref ref-type="bibr" rid="ref27">27</xref>
          ] machine learning technique developed by
Google for pre-training in natural language processing (NLP). In 2018, Jacob
Devlin and his colleagues created and published BERT. BERT includes two
original models in English. The first model is BERTBASE : It consists of 12 encoders
with 12 bidirectional self-attention heads. The second model is BERTLARGE :
24 encoders with 16 bidirectional self-attention heads. BERT is designed for
pretraining from unlabeled texts with 800 million words by BooksCorpus and 2,500
million words by Wikipedia.
        </p>
        <p>BERT has performed more than 10 natural language processing tasks with good
results. It has improved the GLUE benchmark to 80.5%, pushed MultiNLI
accuracy to 86.7%, absolute 5.1 point improvement in SQuAD v2.0 Test F1, etc.
With L is the number of sub-layer blocks in the transformer, H: the size of the
embedding vector, A: the number of heads in the multi-head layer, model BERT
has two architectures as follows:
– BERTBASE (L=12, H=768, A=12): Total parameters are 110 million.
– BERTLARGE (L=24, H=1024, A=16): Total parameters are 340 million.
PhoBERT. The BERT model’s release marked a watershed moment in the NLP
industry. Following the public release of the BERT model, a slew of open-source
BERT training programs have sprung up. There are also numerous unilingual
and multilingual BERT pre-train models that are commonly used. Since then,
PhoBERT has been particularly trained for Vietnamese and released by VinAI
Research in March 2020.</p>
        <p>
          PhoBERT is based on the design and approach of RoBERTa [
          <xref ref-type="bibr" rid="ref28">28</xref>
          ], which was
introduced by Facebook in 2019 and is an improvement over the original BERT.
PhoBERT was trained from about 20GB of data, including approximately 1GB
of the Vietnamese Wikipedia Corpus and 19GB remaining from the Vietnamese
News Corpus. This type of data is also ideal for training a model like BERT.
PhoBERT, similar to BERT, is available in two versions. The first version is
P hoBERTBASE with 12 transformer blocks and the second version with 24
transformer blocks is named P hoBERTLARGE .
        </p>
        <p>VnCoreNLP. VnCoreNLP (A Vietnamese Natural Language Processing
Toolkit) is a Java Natural Language Processing toolkit designed to aid NLP research
in Vietnam. Through essential NLP components such as word segmentation,
POS tagging, and NER, VnCoreNLP provides extensive linguistic annotations.
In the course of NLP research, the Vietnamese standard dataset was published.
In early 2013, the first VLSP evaluation campaign used datasets for word
segmentation and POS tagging. In 2014, a high-quality dependency treebank was
published, and a NER dataset was published for the 2016 VLSP review
campaign. The architectural system design is depicted in Fig. 2.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experimental results</title>
      <sec id="sec-4-1">
        <title>Dataset</title>
        <p>To be able to use PhoBERT to evaluate and categorize the news’ impact, we
provided a dataset that included 1000 titles of financial articles taken from CafeF.vn
and labeled them into three groups [negative, neutral, or positive] with the help
of experts. The dataset contains 187 articles having a negative impact, 248
articles with no impact, and 565 articles with a positive impact. After that, we
divided the dataset into three sets, 80% for training, 10% for validation and
10% for testing. The training set was used to train the model, validation set was
utilized to tune the hyper-parameter. Finally, the result of model was evaluated
on testing set. The examples of our dataset are shown in Table 1.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Result</title>
        <p>After experimenting with 6 diferent models on the same dataset using various
preprocessing techniques in order to achieve the best possible results, Table 2
was obtained.</p>
        <p>As seen in the table above, our model, PhoBERT, outperformed other
popular and sophisticated NLP models with 93.12% accuracy and was 10.54% higher
than the second-highest approach using Logistics Regression. Our model also
achieved the best performances in other metrics such as precision, recall and F1
score.
In this research, sentiment classification was performed on Vietnamese stock
articles collected from the site CafeF.vn. The whole dataset consists of 1000 titles
divided into positive, neutral, and negative news. Because of the fact that white
space is also utilized to separate syllables that make up words in Vietnamese,
we utilize Rdrsegmenter from VnCoreNLP (a word splitting library published
by the author of PhoBERT) to separate words for input data. Then, using the
BPE Encoder, convert text into a list of subwords, then map each subword
to its ID in the PhoBERT vocabulary. The proposed model in this study
employed a state-of-the-art language model for Vietnamese named PhoBERT, and
the entire training process has been deployed on PyTorch. As a result, our
approach achieved 93.12% accuracy, surpassing other popular and sophisticated
NLP models when performed on the same dataset.</p>
        <p>As previously stated, financial news can directly impact stock prices, so our
proposed model could be used to assist with stock price forecasting problems
using machine learning. In future work, we also want to explore the efect of
using word embeddings on sentiment classification. Furthermore, we aim to
investigate multiclass categorization of news data using various Deep Learning
models, which will comprise classes such as the economy, sports, health, and
technology. Also, we intend to extend sentiment classification to more domains
by crawling data from numerous Vietnamese websites, such as product reviews,
hotel reviews, and book reviews.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bhowmick</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hazarika</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <article-title>: Machine Learning for E-mail Spam Filtering: Review, Techniques and</article-title>
          <string-name>
            <surname>Trends. ArXiv.</surname>
          </string-name>
          (
          <year>2016</year>
          ) https://doi.org/10.1007/
          <fpage>978</fpage>
          -981-10-4765- 7 61
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2. J, R.K.,
          <string-name>
            <surname>G</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          , P, S.:
          <article-title>Email Spam Detection using Machine Learning Techniques</article-title>
          .
          <source>IARJSET. 8</source>
          ,
          <fpage>189</fpage>
          -
          <lpage>193</lpage>
          (
          <year>2020</year>
          ). https://doi.org/10.17148/iarjset.
          <year>2021</year>
          .
          <volume>8632</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Shafi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Qamar</surname>
          </string-name>
          , U.: [WiP]
          <article-title>Web Services Classification Using an Improved Text Mining Technique</article-title>
          ,
          <source>2018 IEEE 11th Conference on Service-Oriented Computing and Applications (SOCA)</source>
          .
          <volume>210</volume>
          -
          <fpage>215</fpage>
          (
          <year>2018</year>
          ) https://doi.org/10.1109/SOCA.
          <year>2018</year>
          .
          <volume>00037</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Crasso</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zunino</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Campo</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>AWSC: An approach to Web service classification based on machine learning techniques</article-title>
          .
          <source>INTELIGENCIA ARTIFICIAL. 12</source>
          , (
          <year>2008</year>
          ). https://doi.org/10.4114/ia.v12i37.
          <fpage>955</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>P</given-names>
            <surname>Gayathri</surname>
          </string-name>
          <article-title>: Texture Classification for Fake Indian Currency Detection</article-title>
          .
          <source>International Journal of Engineering Research and. V9</source>
          , (
          <year>2020</year>
          ). https://doi.org/10.17577/ijertv9is060211.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Jain</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kasbe</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          : Fake News Detection,
          <source>2018 IEEE International Students' Conference on Electrical, Electronics and Computer Science (SCEECS)</source>
          .
          <volume>1</volume>
          -
          <fpage>5</fpage>
          (
          <year>2018</year>
          ) https://doi.org/10.1109/SCEECS.
          <year>2018</year>
          .
          <volume>8546944</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Pang</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Opinion mining and sentiment analysis</article-title>
          .
          <source>Foundations and Trends in Information Retrieval. 2</source>
          ,
          <fpage>1</fpage>
          -
          <lpage>135</lpage>
          (
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Staudemeyer</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Morris</surname>
          </string-name>
          , E.:
          <article-title>-Understanding LSTM - a tutorial into Long ShortTerm Memory Recurrent Neural Networks</article-title>
          . (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Nguyen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nguyen</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>PhoBERT: Pre-trained language models for Vietnamese. Association for Computational Linguistics (</article-title>
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>M. A. Hearst</surname>
            ,
            <given-names>S. T.</given-names>
          </string-name>
          <string-name>
            <surname>Dumais</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Osuna</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Platt</surname>
            and
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Scholkopf</surname>
          </string-name>
          :
          <article-title>Support vector machines</article-title>
          ,
          <source>in IEEE Intelligent Systems and their Applications</source>
          , vol.
          <volume>13</volume>
          , no.
          <issue>4</issue>
          , pp.
          <fpage>18</fpage>
          -
          <lpage>28</lpage>
          (
          <year>1998</year>
          ), doi: 10.1109/5254.708428.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Pang</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Opinion mining and sentiment analysis</article-title>
          .
          <source>Foundations and Trends in Information Retrieval. 2</source>
          ,
          <fpage>1</fpage>
          -
          <lpage>135</lpage>
          (
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          : Nıva¨ e Bayes Text Classifier,
          <source>2007 IEEE International Conference on Granular Computing (GRC</source>
          <year>2007</year>
          ).
          <fpage>708</fpage>
          -
          <lpage>708</lpage>
          (
          <year>2007</year>
          ). https://doi.org/10.1109/GrC.
          <year>2007</year>
          .
          <volume>40</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Cunningham</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Delany</surname>
          </string-name>
          , S.J.
          <article-title>: k-Nearest Neighbour Classifiers: 2nd Edition (with Python examples)</article-title>
          . arXiv:
          <year>2004</year>
          .04523 [cs, stat]. (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Ziebart</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maas</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bagnell</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dey</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Maximum Entropy Inverse Reinforcement Learning</article-title>
          . (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Deep Learning for Sentiment Analysis : A Survey</article-title>
          . arXiv:
          <year>1801</year>
          .07883 [cs, stat]. (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Gu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ma</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shahroudy</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shuai</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cai</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Recent advances in convolutional neural networks</article-title>
          .
          <source>Pattern Recognition</source>
          .
          <volume>77</volume>
          ,
          <fpage>354</fpage>
          -
          <lpage>377</lpage>
          (
          <year>2018</year>
          ). https://doi.org/10.1016/j.patcog.
          <year>2017</year>
          .
          <volume>10</volume>
          .013.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Sherstinsky</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Fundamentals of Recurrent Neural Network (RNN) and Long Short-Term Memory (LSTM) network</article-title>
          . Physica D: Nonlinear Phenomena.
          <volume>404</volume>
          ,
          <issue>132306</issue>
          (
          <year>2020</year>
          ). https://doi.org/10.1016/j.physd.
          <year>2019</year>
          .
          <volume>132306</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Munikar</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shakya</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shrestha</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Fine-grained Sentiment Classification using BERT</article-title>
          . (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Devlin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
          </string-name>
          , M.-W.,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Google</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Language</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding</article-title>
          . (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Gelbukh</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Natural language processing</article-title>
          .
          <source>Fifth International Conference on Hybrid Intelligent Systems (HIS'05)</source>
          . (
          <year>2005</year>
          ) https://doi.org/10.1109/ICHIS.
          <year>2005</year>
          .
          <volume>79</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Kieu</surname>
            ,
            <given-names>B.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pham</surname>
            ,
            <given-names>S.B.</given-names>
          </string-name>
          :
          <article-title>Sentiment Analysis for Vietnamese</article-title>
          .
          <source>2010 Second International Conference on Knowledge and Systems Engineering</source>
          . (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Huynh</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hoang</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>GATE framework based metadata extraction from scientific papers</article-title>
          .
          <source>2010 International Conference on Education and Management Technology</source>
          .
          <article-title>(</article-title>
          <year>2010</year>
          ). https://doi.org/10.1109/ICEMT.
          <year>2010</year>
          .
          <volume>5657675</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Vo</surname>
            ,
            <given-names>Q.-H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nguyen</surname>
          </string-name>
          , H.-T.,
          <string-name>
            <surname>Le</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nguyen</surname>
          </string-name>
          , M.-L.
          <article-title>: Multi-channel LSTMCNN model for Vietnamese sentiment analysis</article-title>
          ,
          <source>2017 9th International Conference on Knowledge and Systems Engineering (KSE)</source>
          .
          <volume>24</volume>
          -
          <fpage>29</fpage>
          . (
          <year>2017</year>
          ) https://doi.org/10.1109/KSE.
          <year>2017</year>
          .
          <volume>8119429</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Vu</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Quoc Nguyen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nguyen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dras</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Johnson</surname>
          </string-name>
          , M.:
          <string-name>
            <surname>VnCoreNLP: A Vietnamese Natural Language Processing Toolkit.</surname>
          </string-name>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cho</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gu</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>Neural Machine Translation with Byte-Level Subwords</article-title>
          .
          <source>Proceedings of the AAAI Conference on Artificial Intelligence</source>
          .
          <volume>34</volume>
          ,
          <fpage>9154</fpage>
          -
          <lpage>9160</lpage>
          (
          <year>2020</year>
          ). https://doi.org/10.1609/aaai.v34i05.
          <fpage>6451</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Paszke</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gross</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Massa</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lerer</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bradbury</surname>
            <given-names>Google</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Chanan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            ,
            <surname>Killeen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            ,
            <surname>Gimelshein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Antiga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Desmaison</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          , Kop¨f,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Devito</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            ,
            <surname>Raison</surname>
          </string-name>
          <string-name>
            <surname>Nabla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Tejani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Chilamkurthy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Ai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            ,
            <surname>Steiner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Facebook</surname>
          </string-name>
          , L.:
          <article-title>PyTorch: An Imperative Style, High-Performance Deep Learning Library</article-title>
          . (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Han</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xiao</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          : Transformer in Transformer.
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ott</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goyal</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Du</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joshi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Levy</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lewis</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zettlemoyer</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stoyanov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Allen</surname>
            ,
            <given-names>P.:</given-names>
          </string-name>
          <article-title>RoBERTa: A Robustly Optimized BERT Pretraining Approach</article-title>
          . (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>