<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>December</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Transformer-based Model for Text Classification in Ukrainian</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Larysa Katerynych</string-name>
          <email>katerynych@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maksym Veres</string-name>
          <email>veres@ukr.net</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Eduard Safarov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Taras Shevchenko National University of Kyiv</institution>
          ,
          <addr-line>Academician Glushkov Avenue 4d, Kyiv, 03680</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <volume>0</volume>
      <fpage>1</fpage>
      <lpage>03</lpage>
      <abstract>
        <p>The purpose of this paper is to find a solution for printed text classification in Ukrainian, as well as to choose the means for its implementation. The paper considers the problem of identification of short texts by their scientific topic. A model, built for classification, is described. The current state of NLP and transfer learning is studied too. In practice, the effectiveness of the implemented methods is proven, which allows to obtain good results of text classification. These methods and approaches include concepts such as transfer learning, NLP, BERT. The model is built using Python programming language and some of its machine learning libraries. A multilanguage BERT Ukrainian texts from school subjects. After that, the questions from the external independent evaluation are submitted to the input, and the model classifies them. Text classification, deep learning, recurrent neural networks, long short-term memory, convolutional neural natural representations from transformers, local interpretable model-agnostic explanations.</p>
      </abstract>
      <kwd-group>
        <kwd>networks</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
    </sec>
    <sec id="sec-2">
      <title>2. Transfer learning in NLP</title>
      <sec id="sec-2-1">
        <title>Classic machine learning technique is depicted in the Figure 1. As can be seen, each separate task</title>
        <p>
          demands training of a separate NN with its own model and data. If a new problem arises for NN to
solve, it could be difficult to build an effective system for this purpose. Transfer learning is depicted in
the Figure 2. It is a technique, where a deep learning model, taught on a large dataset, is used to perform
similar tasks on another dataset [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. This model of deep learning is known as a pre-trained model [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ].
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Most tasks in NLP, such as text classification, machine translation, etc., are sequence modeling tasks.</title>
      </sec>
      <sec id="sec-2-3">
        <title>Classic machine learning models and NN cannot fixate the consecutive information, present in the text.</title>
      </sec>
      <sec id="sec-2-4">
        <title>Therefore, RNN have been used, as these architectures can model the sequential information.</title>
        <p>2022 Copyright for this paper by its authors.</p>
      </sec>
      <sec id="sec-2-5">
        <title>However, these periodic NN have their drawbacks. One of the main problems is that RNN cannot</title>
        <p>
          be parallelized (as opposed to a linear NN [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]) because it accepts one input at a time. In the case of a
text sequence, RNN or LSTM accept one token at a time as input. Therefore, training such model on a
large dataset will take a long time. As mentioned above, in 2018, the transformer was introduced by
Google, which gave a significant impulse to NLP systems [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. Soon, a wide range of models, based on
transformers, were offered for various NLP tasks. There are many advantages to using
transformerbased models, but the most important are the following: [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]
1. These models do not process input sequence token-by-token. They take the entire sequence at
once, which is a significant improvement over RNN-based models, as the model can now be
accelerated by graphics processing unit (GPU).
        </p>
      </sec>
      <sec id="sec-2-6">
        <title>2. Labeled data is not required for the preparation of these models. It is needed to provide a huge</title>
        <p>amount of unlabeled text data to prepare a model based on the transformer. This trained model can
be used for other NLP tasks. They can include text classification; named-entity recognition (NER),
for example, people, geographical, company names; text generation; etc.</p>
      </sec>
      <sec id="sec-2-7">
        <title>Bidirectional encoder representations from transformers (BERT) and the second version of</title>
        <p>generative pre-trained transformer (GPT-2) are the most popular NLP models based on transformers.</p>
      </sec>
      <sec id="sec-2-8">
        <title>For example, it is possible to use the previously trained BERT model to classify Ukrainian text.</title>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Model fine-tuning</title>
      <sec id="sec-3-1">
        <title>BERT is a large NN architecture with big number of parameters that can range from 100 to 300 mln.</title>
      </sec>
      <sec id="sec-3-2">
        <title>Therefore, training the BERT model from scratch on a small dataset will lead to overtraining.</title>
        <p>
          It is better to use a pre-trained BERT model as a starting point. These pre-trained models are usually
trained on big datasets. There are several options for BERT. All of them are listed on the official page
[
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. A universal multilingual BERT for 104 languages was chosen. It uses a dictionary of whole words
as well as the most common syllables. A part of the dictionary for multilingual version of BERT can be
seen in the Figure 3. This dictionary shows which words NN uses as input. These are whole words, for
example, кілька. But most Ukrainian words are broken down into syllables. So, the word прийшов will
be split into прий and ##шов.
        </p>
        <p>
          The training of the model can be continued on another, relatively smaller new dataset. This process
is known as fine-tuning of the model [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. Fine-tuning strategies depend on various factors, but the most
important are the size of the new dataset and its similarity to the original dataset. Given that the nature
of a typical NN for NLP is more universal in the early layers and becomes more closely related to a
specific dataset on subsequent layers, four main scenarios can be identified:
1. The new dataset is smaller and similar in content to the original dataset. If the amount of d ata
is small, then it makes no sense to fine-tune the NN due to overfitting. Since the data is similar to
the original, it can be assumed that the distinguishing features in NN will be relevant for this dataset
as well. Therefore, the optimal solution is to train the linear classifier as a distinctive feature of NN.
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>2. The new dataset is relatively large and similar in content to the original dataset. Since there is</title>
        <p>more data, the overfitting does not take place, if the entire NN is being fine-tuned.</p>
      </sec>
      <sec id="sec-3-4">
        <title>3. The new dataset is smaller and significantly different in content from the original dataset. Since</title>
        <p>the amount of data is small, only a linear classifier will be sufficient. Since the data is significantly
different, it is better to train the classifier not from the top of NN, which contains more specific data.</p>
      </sec>
      <sec id="sec-3-5">
        <title>Instead, it is better to train the classifier by activating it on earlier layers of NN.</title>
      </sec>
      <sec id="sec-3-6">
        <title>4. The new dataset is relatively large and differs significantly in content from the original dataset.</title>
      </sec>
      <sec id="sec-3-7">
        <title>Since the dataset is very large, it is possible to train the entire NN from scratch. Nevertheless, in</title>
        <p>practice, it is often still more advantageous to use it to initialize weights from a pre -trained model.</p>
      </sec>
      <sec id="sec-3-8">
        <title>In this case, there is enough data to fine-tune the entire NN.</title>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. BERT architecture</title>
      <sec id="sec-4-1">
        <title>First, BERT is based on the transformer architecture as stated above.</title>
      </sec>
      <sec id="sec-4-2">
        <title>Second, BERT is pre-trained on a large set of unlabeled text, including the entire Wikipedia (that is</title>
      </sec>
      <sec id="sec-4-3">
        <title>2.5 bln words) and the BooksCorpus (800 mln words). This process took 4 days for 16 tensor processing</title>
        <p>units (TPU). The pre-training step is half the success of BERT. The reason is when a model trains on a
large text body, it begins to gain a deeper understanding of how the given language “works”.</p>
      </sec>
      <sec id="sec-4-4">
        <title>Third, BERT is a deep bidirectional model. Bidirectionality means that BERT learns information</title>
        <p>from both left and right side of the token context at the training stage. A similar method of bidirectional
processing is also used by embeddings from language model (ELMo) system developed by Paul Allen
Institute of Artificial Intelligence. However, BERT demonstrates a more complex relationship between
the layers of language representation, so it is considered deeply bidirectional and ELMo is superficially
bidirectional. For comparison, the visualization of NN architectures of different types is displayed in
the Figure 4. They include an example of one-way processing method – OpenAI GPT.
5. ktrain library</p>
        <p>
          ktrain library [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] for Python programming language can be used for BERT fine-tuning. This is a
wrapper for Keras framework that helps to build, train, and deploy NN models with minimum amount
of code. ktrain provides means for: [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]
 learning speed regulation, which will help to find the initial level of learning for the model;
 visual graphs of learning speed to increase productivity;
 pre-trained models for text data (e.g., text classification, NER), images (e.g., image
classification), graphs (e.g., link prediction);
 methods that allow downloading and pre-processing text and images in various formats;
 verification of data that have been misclassified to improve the model;
 an application programming interface (API) for saving and deploying models and
preprocessing data.
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>6. Problem statement</title>
      <p>As an example of a practical solution to the problem of Ukrainian text classification, BERT model
can be considered and taught to classify scientific topics of given texts. These texts will cover the
following 7 school subjects:
 history of Ukraine,
 physics,
 geography,
 biology,
 mathematics (algebra and geometry),
 Ukrainian (language and literature),
 chemistry.</p>
      <sec id="sec-5-1">
        <title>The task is to create a system that determines automatically to which subject the question relates.</title>
      </sec>
      <sec id="sec-5-2">
        <title>The dataset used for this case consists of electronic textbooks on the specified subjects. To test the</title>
        <p>model, the questions from the tests, that were offered to school graduates at the external independent
evaluation (also known as “ЗНО”) of 2021, were considered. A sample can be seen in the Figure 5. It
should also be noted that only 11th (final) grade textbooks are included in the dataset, while the external
evaluation questions cover several years of study and relate to a wider range of knowledge. Google</p>
      </sec>
      <sec id="sec-5-3">
        <title>Colab was used for development.</title>
        <p>6.1.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Download and process input data</title>
      <p>ktrain library is needed to get started:
!pip install ktrain
import ktrain
from ktrain import text</p>
      <p>
        Publicly available [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] textbooks for remote studying are considered as input data. The textbooks
were converted to TXT format and divided into smaller files. ktrain automatically detects natural
language and character encoding, processes the data, and sets up the model:
(x_train, y_train), (x_test, y_test), preproc = text.texts_from_folder(
'/content/drive/MyDrive/dataset/',
maxlen=75,
max_features=10000,
preprocess_mode='bert',
train_test_names=['train', 'test'],
val_pct=0.1,
classes=['history', 'physics', 'math', 'geography', 'biology', 'Ukrainian',
'chemistry'])
      </p>
      <sec id="sec-6-1">
        <title>The first argument is the path to the dataset folder. The maxlen argument specifies the maximum</title>
        <p>number of words (512 for BERT, but it is better to use less to reduce memory usage and increase speed)
in each file, with extra words being cut off. maxlen=75 because the input text files are small.</p>
      </sec>
      <sec id="sec-6-2">
        <title>The text must be pre-processed for usage with BERT. This is achieved by setting the</title>
        <p>preprocess_mode value to 'bert'. The BERT model and vocabulary will be loaded automatically, if
necessary. val_pct=0.1 means automatically selecting 10% of the data for validation.</p>
        <p>Finally, the texts_from_folder function expects the following directory structure:
&lt;folder&gt;
train
&lt;subject_1&gt;
&lt;subject_2&gt;
&lt;subject_3&gt;
...
test
&lt;subject_1&gt;
&lt;subject_2&gt;
&lt;subject_3&gt;
...</p>
        <p>So, the dataset folder with corresponding content for 7 classes 'history', 'physics', 'math',
'geography', 'biology', 'Ukrainian', 'chemistry' is created. The result may look like this:
detected encoding: UTF-8-SIG
downloading pretrained BERT model (multi_cased_L-12_H-768_A-12.zip)...
extracting pretrained BERT model...
done.
cleanup downloaded zip...
done.
preprocessing train...
language: uk
done.</p>
        <p>Is Multi-Label? False
preprocessing test...
language: uk
done.
6.2.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Use BERT learner object for content in Ukrainian</title>
      <sec id="sec-7-1">
        <title>Creating a model and wrapping it with the learner:</title>
        <p>model = text.text_classifier('bert', (x_train, y_train), preproc=preproc)
learner = ktrain.get_learner(model,
train_data=(x_train, y_train),
val_data=(x_test, y_test),
batch_size=32)</p>
      </sec>
      <sec id="sec-7-2">
        <title>The first argument of the get_learner function uses a pre-trained BERT model with a randomly</title>
        <p>initialized end dense layer. The second and third arguments are training and verification data,
respectively. The last argument of get_learner is the packet size. A small batch_size=32 is used
based on Google's recommendations.
6.3.</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>Model training</title>
      <sec id="sec-8-1">
        <title>To train the model, the optimal level of training is found, that corresponds to the problem being</title>
        <p>solved. ktrain offers an effective method called lr_find, which trains a model with different learning
metrics and creates a graph of model loss as the learning level increases.</p>
      </sec>
      <sec id="sec-8-2">
        <title>Loss graph is displayed in the Figure 6.</title>
        <p>learner.lr_find()
learner.lr_plot()</p>
      </sec>
      <sec id="sec-8-3">
        <title>Result:</title>
      </sec>
      <sec id="sec-8-4">
        <title>Validation:</title>
      </sec>
      <sec id="sec-8-5">
        <title>Result:</title>
        <p>learner.validate(val_data=(x_test, y_test))</p>
        <p>
          The graph of the loss function shows that the classifier provides minimum losses when the training
level is 10-5...10-4. Model training:
learner.fit_onecycle(2e-5, 1)
fit_onecycle function from ktrain library is used. This function utilizes the onecycle learning
speed policy, which linearly increases the learning speed during the first half of learning and then
reduces the learning speed for the second half [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ].
        </p>
        <p>begin training using onecycle policy with max lr of 2e-05...</p>
        <p>542/542 [==============================] - 710s 1s/step - loss: 0.5202
accuracy: 0.8294 - val_loss: 0.2498 - val_accuracy: 0.9182
precision
recall f1-score</p>
        <p>support</p>
      </sec>
      <sec id="sec-8-6">
        <title>As can be seen from the results of validation, training reaches 85-98% accuracy (precision) in one epoch. Classification errors:</title>
        <p>learner.view_top_losses(n=10, preproc=preproc)</p>
      </sec>
      <sec id="sec-8-7">
        <title>Creating a predictor:</title>
        <p>p = ktrain.get_predictor(learner.model, preproc)</p>
      </sec>
      <sec id="sec-8-8">
        <title>It is better to save the model for later use:</title>
        <p>p.save('./drive/MyDrive/predictor')</p>
        <p>After downloading it and trying to offer models of questions from the external independent
evaluation of 2021:
fin_bert_model = ktrain.load_predictor('./drive/MyDrive/predictor')
p.predict("Вільгельм фон Габсбург-Лотрінген – австрійський архікнязь, полковник
армії УНР, поет. Під яким псевдонімом він відомий як полковник Українських січових
стрільців (УСС)?")
history
p.predict("Усю воду із широкої посудини перелили у високу вузьку порожню
посудину. Якими стануть сила тиску й тиск води на дно вузької посудини після цього
порівняно із силою тиску и тиском цієї води на дно широкої посудини? Уважайте, що
посудини мають циліндричну форму")
physics
p.predict("Установіть відповідність між графіком (1 — 3) функції, визначеної на
проміжку [— 4; 4], та її властивістю (А — Д)")
math
p.predict("Укажіть материк, на якому лежить Україна")
geography
p.predict("Під час експерименту декілька яєць морських їжаків помістили в морську
воду, де їх запліднили. У цю воду добавили мічений Тритієм (3Н) тимідиловий нуклеотид
(рис. 1), який поглинали клітини ембріонів")
biology
p.predict("Зображене в уривку - Христос Воскресе! – І розвіявся морок. Упали
кайдани з невольницьких рук. - Спішіть! Байдаки у відкритому морі! Поблизу нема ні
галер, ні фелюк! суголосне з подіями твору")</p>
        <p>Ukrainian
fin_bert_model.predict(Однакові кульки, підвішені на нитках, заряджені так, як
це показано на рисунках. У якому з випадків правильно зображено положення цих
кульок, зумовлене їхньою взаємодією?")
physics
fin_bert_model.predict(“Яка владна інституція звернулася із цитованою відозвою
до населення України?")
history
fin_bert_model.predict("Установіть відповідність між виразом (1
твердженням про його значення (А — Д), яке є правильним, якщо a = -2")
math
— 3) і
fin_bert_model.predict("З будь-якої точки Світового океану можна дістатися в
будь-яку іншу, не перетнувши суходіл. Це доводить, що Світовий океан —")
geography
fin_bert_model.predict("Уключення міченої сполуки в молекули клітин ембріона
відбувається під час")
biology
fin_bert_model.predict("Однаковий звук позначають букви, підкреслені в окремих
словах речення")</p>
        <p>Ukrainian
fin_bert_model.predict("Укажіть формулу вуглекислого газу")
chemistry
Taking a closer look at how the model makes conclusions for certain inputs:
fin_bert_model.explain("Мічена
молекули")
сполука
в
клітинах
ембріонів
потрапляє
в</p>
      </sec>
      <sec id="sec-8-9">
        <title>Result:</title>
        <p>y=biology (probability 0.989, score 5.285) top features
Contribution Feature
+5.804 Highlighted in text (sum)
-0.519 &lt;BIAS&gt;
мічена сполука в клітинах ембріонів потрапляє в молекули</p>
      </sec>
      <sec id="sec-8-10">
        <title>This visualization is generated using a technique called local interpretable model-agnostic</title>
        <p>
          explanations (LIME) [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. It helps to understand the relative importance of different words for the final
prediction using a linear interpreted model. “Green” words contribute to the correct classification
(biology), “red” words reduce the probability of correct prediction. Shade of color indicates the strength
or size of the coefficients in the final linear model.
        </p>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>7. Other approaches</title>
      <p>There are a number of solutions representing classification of English texts. Many of them provide
algorithms to build text classifiers. But it’s hard to find works describing algorithms to construct a
classifier of Ukrainian text. The main problem, however, is not an algorithm itself, but the lack of
resources for experiments to train the classifier. A researcher can:
 use Brownian Corps of the Ukrainian Language (BCUL),
 use specific dictionaries,
 create own dataset.</p>
      <p>
        Deep learning tends to use large and robust datasets in order to perform well. One of the attempts to
perform Ukrainian text classification is described in the paper [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], where its authors consider random
forest classifier, support vector machines (SVM), naive Bayes classifier and logistic regression
algorithms as well as BCUL dataset. The best result is shown by SVM model. Its average accuracy is
80%. Another approach to the problem proposes a solution, which can detect sentiments from Ukrainian
text (namely hotel reviews) and classify them for positive and negative ones [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Besides, it analyzes
reviews about given hotel and summarizes its most important positive and negative properties. The
dataset, which contains user reviews in Ukrainian was parsed from Booking.com and TripAdvisor.
      </p>
      <sec id="sec-9-1">
        <title>Following techniques were considered: fastText (created by Facebook Artificial Intelligence Research lab), Seq2Seq, CNN, RNN and recurrent convolution neural network (RCNN). RCNN model gives the best accuracy on available dataset: 85% for text classification and 86% for sentence classification.</title>
      </sec>
    </sec>
    <sec id="sec-10">
      <title>8. Conclusions</title>
      <sec id="sec-10-1">
        <title>BERT and ktrain were used to solve the problem of classifying documents in Ukrainian. A complex</title>
        <p>model like BERT can be applied to any problem in any language (out of 104 most popular ones in the
world, including Ukrainian, for which there is a pre-trained BERT) using ktrain.</p>
      </sec>
      <sec id="sec-10-2">
        <title>Testing accuracy of 85-98% was achieved in one learning epoch. Although the effectiveness of</title>
      </sec>
      <sec id="sec-10-3">
        <title>BERT was proven, it is relatively slow both in terms of learning and predictions for new data. Therefore,</title>
        <p>if the training lasts more than one epoch, it may be better to omit the val_data argument from
get_learner and check the accuracy only after training.</p>
      </sec>
      <sec id="sec-10-4">
        <title>This can be done in ktrain using the learner.validate method, as shown in code samples above.</title>
      </sec>
      <sec id="sec-10-5">
        <title>BERT can be quite demanding on memory. If errors, indicating that the GPU memory limits have been</title>
        <p>exceeded, are encountered, it is possible to reduce the value of either maxlen, or batch_size
parameters. If the model performs well after training, it should be saved for future classifications. When
using BERT, Keras’s built-in load_model function does not work, although model.save_weights and
model.load_weights can still be used to save weights and load them. But to load the model, the
learner.load_model function from ktrain should be used.</p>
      </sec>
      <sec id="sec-10-6">
        <title>Finally, the proposed BERT solution for Ukrainian text classification was compared to other approaches to this problem. As can be seen, the described model performs well, and its accuracy is high enough.</title>
      </sec>
    </sec>
    <sec id="sec-11">
      <title>9. References</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>I.</given-names>
            <surname>Goodfellow</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bengio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Courville</surname>
          </string-name>
          ,
          <article-title>Deep learning</article-title>
          , The MIT Press,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>E.</given-names>
            <surname>Olivas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Guerrero</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sober</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Benedito</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lopez</surname>
          </string-name>
          ,
          <source>Handbook of Research on Machine Learning Applications and Trends: Algorithms</source>
          , Methods and Techniques, IGI Publishing,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>L.</given-names>
            <surname>Katerynych</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Veres</surname>
          </string-name>
          , E. Safarov,
          <source>Neural Networks' Learning Process Acceleration, in: Proceedings of the 12th International Scientific and Practical Conference of Programming</source>
          ,
          <source>UkrPROG'</source>
          <year>2020</year>
          ,
          <article-title>Problems in Programming Scientific Journal</article-title>
          , Kyiv,
          <year>2020</year>
          , pp.
          <fpage>313</fpage>
          -
          <lpage>321</lpage>
          . doi:
          <volume>10</volume>
          .15407/pp2020.
          <fpage>02</fpage>
          -
          <lpage>03</lpage>
          .
          <fpage>313</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Petrov</surname>
          </string-name>
          .
          <source>Multilingual BERT models</source>
          ,
          <year>2019</year>
          . URL: https://github.com/googleresearch/bert/blob/master/multilingual.md.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chang</surname>
          </string-name>
          , Open Sourcing BERT:
          <article-title>State-of-the-Art Pre-training for</article-title>
          <source>Natural Language Processing</source>
          ,
          <year>2018</year>
          . URL: https://ai.googleblog.com/
          <year>2018</year>
          /11/open-sourcing
          <article-title>-bert-state-of-art-pre</article-title>
          .html.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Maiya</surname>
          </string-name>
          .
          <source>ktrain: A Low-Code Library for Augmented Machine Learning</source>
          ,
          <year>2020</year>
          . URL: https://arxiv.org/pdf/
          <year>2004</year>
          .10703.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>S.</given-names>
            <surname>Ravichandiran</surname>
          </string-name>
          ,
          <article-title>Getting Started with Google BERT: Build and Train State-of-the-</article-title>
          <string-name>
            <surname>Art Natural Language Processing Models using</surname>
            <given-names>BERT</given-names>
          </string-name>
          , Packt Publishing Ltd,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>L.</given-names>
            <surname>Smith:</surname>
          </string-name>
          <article-title>A Disciplined Approach to Neural Network Hyper-parameters: Part 1, 2018</article-title>
          . URL: https://arxiv.org/pdf/
          <year>1803</year>
          .09820.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Electronic</given-names>
            <surname>Versions</surname>
          </string-name>
          of Textbooks, Institute for Modernization of Educational Content,
          <year>2021</year>
          . URL: https://lib.imzo.gov.ua/yelektronn-vers-pdruchnikv.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>M.</given-names>
            <surname>Ribeiro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Guestrin</surname>
          </string-name>
          ,
          <article-title>"Why Should I Trust You?": Explaining the Predictions of any Classifier</article-title>
          ,
          <source>in: 22nd International Conference on Knowledge Discovery and Data Mining</source>
          , San Francisco, CA, USA,
          <year>August</year>
          13-
          <issue>17</issue>
          ,
          <year>2016</year>
          , pp.
          <fpage>1135</fpage>
          -
          <lpage>1144</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>K.</given-names>
            <surname>Bobrovnyk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Dukhnovska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Piroh</surname>
          </string-name>
          , Thematic Classification of Ukrainian Texts, Difficulties of its Introductions,
          <source>in: Control Systems and Computers</source>
          , Kyiv, Ukraine. doi:
          <volume>10</volume>
          .15407/usim.
          <year>2019</year>
          .
          <volume>01</volume>
          .041.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>D.</given-names>
            <surname>Babenko</surname>
          </string-name>
          ,
          <article-title>Determining sentiment and important properties of Ukrainian-language user reviews</article-title>
          ,
          <source>Master's thesis</source>
          , Ukrainian Catholic University, Lviv, Ukraine,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>