<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Gradient Boosted Trees for Identification of Complex Words in Context</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Raksha Agarwal</string-name>
          <email>raksha.agarwal@maths.iitd.ac.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Niladri Chatterjee</string-name>
          <email>niladri@maths.iitd.ac.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Complex Word Identification, Linguistic Features, CatBoost, Domain Adaptation</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Indian Institute of Technology Delhi</institution>
          ,
          <addr-line>Hauz Khas, Delhi-110016</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <fpage>12</fpage>
      <lpage>28</lpage>
      <abstract>
        <p>Determining whether a particular word is complex in a given context is an important task for modern NLP, as the presence of complex words may hinder smooth communication. The present work focuses on developing a binary classifier for predicting the complexity of a target word. A set of 51 features, pertaining to eight diferent classes, has been identified for the said purpose. Four diferent classifiers have been used, and their performance is compared. CatBoost registered the best performance when tested on CWI2016 dataset, and for the News and Wikinews categories for CWI2018 dataset. In fact, the CatBoost system supersedes the top performers for the 2016 and 2018 contests for the above-mentioned cases. The optimal feature subsets for the datasets are obtained using recursive feature elimination.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Presence of dificult words in a text can lower readability and comprehension for second language
learners as well as for native speakers with low literacy levels and reading dificulties [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. This
can lead to miscommunication of ideas and/or misunderstanding of contents. Automatic
identification of dificult-to-understand words in a given sentence has been considered as a
core part of Lexical Simplification (LS) systems by several works in the past [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ]. This task is
commonly referred to as Complex Word Identification (CWI). Absence of CWI from LS systems,
and adopting a ’Simplify Everything’ [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] approach may obscure the meaning of the source
sentence due to redundant substitutions of simple words.
      </p>
      <p>
        CWI systems are categorized into four types, namely Threshold-based, Lexicon-based, Implicit
CWI and Machine learning-assisted [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Threshold-based system segregate complex and simple
words by setting a threshold value on a simplicity metric, such as word frequency [
        <xref ref-type="bibr" rid="ref2 ref6">2, 6</xref>
        ].
Lexiconbased systems make use of domain-specific lexicons for CWI to replace a complex word with
a simple word/phrase with similar meaning [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Implicit CWI systems, instead of identifying
complex words, focus on determining whether or not a word can be replaced by a simpler
alternative [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. CWI systems of the above types ignore the efect of context in determining the
complexity of a word. Machine learning-assisted CWI systems are enabled to design classifiers
on an extensive feature space comprising shallow features of the target word (e.g. length of the
Proceedings of the First Workshop on Current Trends in Text Simplification (CTTS 2021), co-located with SEPLN 2021.
CEUR
word, POS tag, frequency in lexicon) as well as sentence level features.
      </p>
      <p>In the present work, Gradient Boosted Tree classifiers are trained on a feature space
comprising of hand-crafted linguistic features along with vector-based similarity features for CWI. In
particular, we have considered three diferent boosting techniques, namely, CatBoost, XGBoost,
LGBM, and compared their performance with traditional Random Forest Classifier. The
experiments were conducted on the CWI2016 and CWI2018 datasets. Our experiments establish the
superiority of the CatBoost algorithm over others.</p>
      <p>The rest of the paper is organised as follows. Section 2 presents the existing related works
on CWI. In Section 3, individual features of the proposed feature space are described. Details
about the datasets, classification algorithms and optimal feature subset selection are presented
in Section 4. Results and domain adaptation study are presented in Section 5. The paper is
concluded in Section 6.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Works</title>
      <p>CWI has attracted the attention of various researchers in the past few years with organization
of two shared tasks, namely CWI 2016 and CWI 2018 [9, 10]. While CWI 2016 focused only on
the English language, the CWI 2018 extended the scope of CWI 2016 by using German, French
and Spanish in addition to English. In the present work, CWI has been performed only for
English language.</p>
      <p>
        For CWI 2016, a Performance-Oriented Soft Voting ensemble of Threshold-based,
Lexiconbased, and Machine Learning-based classifiers trained on morphological, lexical and semantic
features of the target word achieved the highest score [11]. Majority of the other systems
used Machine Learning-assisted CWI techniques, such as SVMs [12, 13, 14], Random Forests
[15, 16, 17], Decision Trees [18], and ensemble systems [19, 20]. Additionally, threshold based
methods trained on word frequencies were also used [
        <xref ref-type="bibr" rid="ref10 ref9">21, 22, 23</xref>
        ].
      </p>
      <p>
        For CWI 2018, CAMB system [
        <xref ref-type="bibr" rid="ref11">24</xref>
        ] trained on an extensive set of hand-crafted features
consisting of lexical, psycholinguistic and lexicon based features achieved the highest macro F1
score. It used Adaboost and Random forest Classifiers. Some systems [
        <xref ref-type="bibr" rid="ref12 ref13 ref14">25, 26, 27</xref>
        ] used Word2vec
and GloVe word embeddings along with hand-crafted features. Post CWI 2018 task SEQ [
        <xref ref-type="bibr" rid="ref15">28</xref>
        ], a
BiLSTM-based sequence labelling method with GloVe word embeddings, outperformed CAMB
system.
      </p>
      <p>
        Finnimore et al.[
        <xref ref-type="bibr" rid="ref16">29</xref>
        ] trained a Logistic Regression model with a suitably chosen feature set
for CWI. Their feature set contained 25 features based on the target word/MWE, sub-word level
features, and sentence-level features. However, their system did not outperform CAMB. Sheang
[
        <xref ref-type="bibr" rid="ref17">30</xref>
        ] presented an approach to CWI based on Convolutional Neural Networks (CNN) trained
on pre-trained word embeddings with morphological and linguistic features. Ehara et al.[
        <xref ref-type="bibr" rid="ref18">31</xref>
        ]
developed a graph-based method for CWI based on similarities among corpus word frequencies.
Detailed analysis of the similarity and distance between the word-frequency distributions of
ifve corpora was conducted using four diferent measures. Zaharia et al. [
        <xref ref-type="bibr" rid="ref19">32</xref>
        ] performed CWI
using multilingual and language-specific Transformer models, multilingual word embeddings
(non-Transformer), and diferent fine-tuning techniques. Crosslingual Zero-shot, One-shot and
Few-shot transfer evaluations were also performed. Aprosio et al. [
        <xref ref-type="bibr" rid="ref20">33</xref>
        ] presented a pipeline for
personalised complex word detection adapting to the mother tongue of non-native speakers,
and based on false friend identification. Their system utilized manually curated datasets of
cognates and false friends for four language pairs.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Feature Space</title>
      <sec id="sec-3-1">
        <title>3.1. Lexical Features</title>
        <p>For the present work we have used a feature space consisting of 51 features, classified into eight
categories as described in the following subsections.</p>
        <p>Shallow features, such as Number of characters (Nchar ), vowels (Nvow), phonemes (Nphon),
syllables (Nsyl), morphemes (Nmorph), and percentage of upper case characters (UpCase) in the
target token are used to model lexical characteristics of the token. A feature (IsNE) is used to
indicate whether the input token is a Named Entity. The language of etymological1 origin (e.g.,
French, Latin) of the target word is also considered as a feature, named EtymOrig. A Boolean
feature (IsStopword) is used to indicate whether the token is a stopword. This feature has been
extracted using NLTK’s list of English stopwords. Both simple Universal POS tag (UnivTag) and
detailed Penn POS tag (PennTag)2 of the input token are also considered as features.</p>
        <p>Number of synsets (Nsyn), hyponyms (Nhypo) and hypernyms (Nhyper ) of the target word
as extracted from NLTK WordNet are considered as features as well. The number of characters
in the words immediately preceding and succeeding the target tokens, named NcharPrev and
NcharNext, respectively, are also included in the feature space. Additionally, we have also
considered two sentence level features, namely the total number of tokens in the sentence
(LenSent), and the relative position (Relpos) of the input token in the sentence.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Lexicon based features</title>
        <p>Two lexicon based Boolean features, namely InGoogle and InOgden were also considered.
a. InGoogle indicates whether the input token belongs to the list of 10,000 most common</p>
        <p>English words, determined by n-gram frequency in the Google’s Trillion Word Corpus3.
b. InOgden indicates the presence of input tokens in the list of 1000 words included in</p>
        <p>Ogden’s Basic English4.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Frequency based features</title>
        <p>Several frequency based features have been used to model the familiarity of the target word:
1https://pypi.org/project/ety/
2IsNE, UnivTag and PennTag have been extracted using spaCy
3https://github.com/first20hours/google-10000-english
4http://ogden.basic-english.org
a. Frequency of the word in Ogden’s Basic English (OgdenFreq), Exquisite Corpus (ECFreq)
and SUBTLEX (SUBTFreq). Exquisite Corpus5 compiles texts from several domains.</p>
        <p>
          SUBTLEX6 contains frequency of 51M words calculated on a corpus of Movie Subtitles.
b. Contextual Diversity (ContDiverse) reported in SUBTLEX is also used as a feature.
Contextual Diversity is computed as the percentage of movies in which the target word
appears.
c. Furthermore, frequency of the input tokens given in the L count of Thorndike and Lorge
[
          <xref ref-type="bibr" rid="ref21">34</xref>
          ], and London-Lund Corpus of English Conversation by Brown [
          <xref ref-type="bibr" rid="ref22">35</xref>
          ] are also used as
features. These are named TLFreq and BrownFreq, respectively.
        </p>
      </sec>
      <sec id="sec-3-4">
        <title>3.4. Character Language Model features</title>
        <p>
          Words with unusual letter sequence may add to the complexity of the word. In order to
incorporate this, the probability of the input token, calculated using bigram and trigram character
language models, have been considered as features (BiCharProb, TriCharProb). These
probabilities are expected to be lower for words with unusual sequence of letters. Letter counts from
Google’s Trillion Word Corpus7 are used to calculate the letter bigram and trigram probabilities
[
          <xref ref-type="bibr" rid="ref23">36</xref>
          ].
        </p>
      </sec>
      <sec id="sec-3-5">
        <title>3.5. Psycholinguistic Features</title>
        <p>
          The cognitive processes in the human brain is influenced by the psycholinguistic properties of
a word when presented with either written or spoken forms [
          <xref ref-type="bibr" rid="ref24">37</xref>
          ]. These properties include Age
of acquisition (AOA), Concreteness (Conc), Imageability (Imag) and Meaningfulness ratings,
namely MeanC and MeanP of the target word. In the present work these features are extracted
using MRC psycholinguistic database [
          <xref ref-type="bibr" rid="ref25">38</xref>
          ]. Additionally, target token’s written frequency of
occurrence (KFFreq), and the number of categories (KFNcats) and number of samples (KFNsamp)
of text in which the target word was found are also used as features [
          <xref ref-type="bibr" rid="ref26">39</xref>
          ].
        </p>
      </sec>
      <sec id="sec-3-6">
        <title>3.6. Language Model Features</title>
        <p>
          Statistical n-gram language models are used to study the collocation of words in sentences,
and determine the probability of a sequence of words. A trigram language model8 trained on
the Gigaword corpus [
          <xref ref-type="bibr" rid="ref27">40</xref>
          ] has been used to extract two features (FragScore3, FragScore5) which
measure the language model score of a word sequence containing the target token, and the
context words in the source sentence in a window of size 3 and 5, respectively [
          <xref ref-type="bibr" rid="ref23 ref28">36, 41</xref>
          ]. The
above-mentioned scores of the given target word help to determine whether or not the word is
used in an unusual context in the given source sentence.
        </p>
        <p>5https://pypi.org/project/wordfreq/
6https://github.com/Wonderlic-AI/wonderlicnlp
7http://norvig.com/ngrams/count_2l.txt, http://norvig.com/ngrams/count_3l.txt
8http://www.keithv.com/software/giga/lm_giga_64k_nvp_3gram.zip</p>
      </sec>
      <sec id="sec-3-7">
        <title>3.7. Dependency features</title>
        <p>The Dependency tree of a sentence helps to understand the relationship between diferent words
of a given sentence. In this respect, the dependency tag of the input token with its syntactical
head (DepTag) and, POS tag of the head (HeadPOS) are considered as features. Additionally, two
features are extracted from the dependency tree, namely depth of the input token in the tree
(TokDepth), and the number of children of the input token (NChild).</p>
      </sec>
      <sec id="sec-3-8">
        <title>3.8. Vector Similarity features</title>
        <p>
          In order to incorporate some additional information about the agreement between the target
word and its context the following vector similarity features are included in the feature space.
GloVe 300-dimensional word embeddings [
          <xref ref-type="bibr" rid="ref29">42</xref>
          ] have been used to calculate the similarity. Other
word embeddings, viz. Word2Vec[
          <xref ref-type="bibr" rid="ref30">43</xref>
          ] and FastText[
          <xref ref-type="bibr" rid="ref31">44</xref>
          ] gave inferior results in our preliminary
studies.
        </p>
        <p>a. The cosine similarity of the target token with the root token and syntactical head in the
dependency tree are also taken as features (RootSim, HeadSim).
b. The average similarity of the target token with its siblings and children in the dependency
tree are also considered (AvgChildSim, AvgSibSim). Maximum similarity of the target
token with its siblings is also considered (MaxSibSim).
c. To further measure the compatibility of the target token with its context, average similarity
with  words to its immediate left and right are extracted for  = 1 and  = 5 . These
features are named as LeftSim5 , RightSim5, LeftSim1 and RightSim1, respectively.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Experimental Details</title>
      <p>This section presents the details of datasets, classification algorithms, and the feature selection
approach.</p>
      <sec id="sec-4-1">
        <title>4.1. Datasets</title>
        <p>In this section the details about the datasets of CWI 2016 and CWI 2018 [9, 10] are presented.
CWI 2016 dataset contained 9200 sentences. The target words were manually annotated by
400 non-native English speakers as complex or non-complex. CWI 2018 used the sentences
from three diferent text genres, namely News (professionally written news), WikiNews (news
written by amateurs), and articles from Wikipedia [10]. Here, words and phrases of length up
to 10 words were annotated by 183 annotators comprising both native and non-native English
speakers. Along with the binary complex vs. non-complex label it also contains a probabilistic
label representing the proportion of annotators that labelled the item as complex. In this work
we have focused on the binary classification only. The data statistics is presented in Table 1. For
CWI 2016, the systems were evaluated using a new metric G-Score, which is the harmonic mean
of Accuracy and Recall. We have also reported the Accuracy, Precision, Recall and F1-Score.
For CWI 2018, macro-average F1-score has been used for evaluation. The above metrics have
been chosen as per the instructions of CWI 2016 [9] and CWI 2018 [10].</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Classification Algorithms</title>
        <p>
          In the present work, we have experimented with the following classification algorithms:
• Random Forest (RF): RF Classifiers train a multitude of decision trees on various
subsamples of the dataset, and use averaging to improve accuracy and control over-fitting
[
          <xref ref-type="bibr" rid="ref32">45</xref>
          ].
• XGBoost : It turns weakly learned decision trees into strong learners by training upon
residuals instead of aggregation [
          <xref ref-type="bibr" rid="ref33">46</xref>
          ].
• Light Gradient Boosting Machine (LGBM) : This is a histogram-based boosting algorithm.
        </p>
        <p>
          Here, boosting is performed using a specialised gradient-based one-sided sampling of
data points of large gradients [
          <xref ref-type="bibr" rid="ref34">47</xref>
          ].
• CatBoost: This method makes better utilisation of the categorical features which are
otherwise converted to numerical features in traditional gradient boosting [
          <xref ref-type="bibr" rid="ref35">48</xref>
          ]. Here,
Oblivious trees are used as base predictors. These trees use the same splitting criterion
across the entire level of the tree, making it less prone to overfitting.
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Feature Selection</title>
        <p>Recursive feature elimination is performed for maximizing the mean 5-fold cross-validation
F1-score. We observed that the same set of features did not perform equally well for the two
datasets. For CWI 2016, all features except Nhypo, KFNcats, NChild and AOA; and for CWI 2018,
all features except SUBTFreq and HeadSim are included in the optimal feature subset.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Results and Analysis</title>
      <p>
        The results for CWI 2016 dataset corresponding to the diferent classification algorithms
mentioned in Section 4.2 are mentioned in Table 2. The proposed CWI approach has been compared
with state-of-the-art and other top performing systems on the two datasets. It can be observed
that CatBoost classifiers outperform other algorithms as well as other existing baselines in terms
of Accuracy, Precision, F1-Score, and G-Score. Since many features in the proposed feature
space are categorical in nature, the superior performance of CatBoost may be attributed to its
efective Ordered Target Encoding [
        <xref ref-type="bibr" rid="ref35">48</xref>
        ] for categorical feature preprocessing. For CWI 2018,
unlike existing works [
        <xref ref-type="bibr" rid="ref11 ref14">24, 27</xref>
        ] a unified system is trained for all the three sub-datasets. Initially
      </p>
      <sec id="sec-5-1">
        <title>System</title>
      </sec>
      <sec id="sec-5-2">
        <title>Classifier</title>
      </sec>
      <sec id="sec-5-3">
        <title>CatBoost</title>
      </sec>
      <sec id="sec-5-4">
        <title>XGBoost RF</title>
      </sec>
      <sec id="sec-5-5">
        <title>LGBM</title>
      </sec>
      <sec id="sec-5-6">
        <title>SV000gg [11] Voting</title>
        <p>TALN [16] RF</p>
      </sec>
      <sec id="sec-5-7">
        <title>UWB [21] MaxEntropy</title>
      </sec>
      <sec id="sec-5-8">
        <title>PLUJAGH [22] Threshold</title>
        <p>all samples corresponding to phrases are discarded from the train and test set; and the classifiers
are trained only on single words.</p>
        <p>The results for CWI 2018 dataset for single word targets are presented in Table 3. CatBoost
classifiers achieved the highest score in terms of Accuracy and Macro-F1 for the News and
WikiNews. For Wikipedia, a low score is achieved which may be due to sub-optimality of the
feature space corresponding to this data split.</p>
        <p>To analyze the efect of individual feature subsets, the corresponding subset of features is
removed from the feature space. CatBoost classifiers are trained on the reduced feature space
and results are reported in Table 4. All the feature subsets contribute to increasing the overall
performance of the system. While removing Lexical features led to the largest decline in scores,
contextual features corresponding to Language Model, Dependency Trees and Vector Similarity
also emerged as important features.</p>
        <sec id="sec-5-8-1">
          <title>5.1. Complex Phrase Identification</title>
          <p>CWI 2018 dataset contains phrases along with single words. In order to predict the complexity
of target phrases two approaches have been used. In the first approach, all phrases are marked
as complex. In the second approach, if the mean predicted complexity of individual component
words is above a threshold then the phrase is marked as complex. In subsequent discussions
these are referred as Greedy and Threshold, respectively. Individual word complexities are
derived using the predictions obtained by CatBoost classifiers. The Accuracy and Macro-F1
for complexity of phrases corresponding to 10 equally spaced threshold values between 0 to 1
is presented in Figure 1. The threshold is chosen to be 0.5 because macro-F1 is maximized at
this value. It can be noted that the threshold value 0 corresponding to maximum accuracy is
equivalent to the Greedy approach. The results for the entire CWI 2018 dataset including both
words and phrases are presented in Table 5. Greedy approach achieves the best scores for News
and WikiNews. The scores for Threshold approach is low because a huge proportion (about
80%) of phrases for both training and test set are complex.</p>
          <p>
            McNemar’s test [
            <xref ref-type="bibr" rid="ref36">49, 50</xref>
            ] has been used to compare the performance of the CatBoost with
well-known baselines, SEQ and CAMB systems, for CWI 2018 dataset. For each system we have
constructed 2 × 2 contingency tables for both Words and Word+Phrases. Figures 2a, 2b show
the values for CatBoost v/s SEQ, while Figures 2c, 2d show the values for CatBoost v/s CAMB.
According to the test, the null hypothesis (equal performance of the systems) is rejected with
99% confidence for both Words and Words+Phrases for CAMB system; and it is rejected with
95% and 90% confidence for Words and Words+Phrases, respectively for SEQ system.
(a) For Words ( 2 = 3.935,  = 0.047 )
(b) For Words+Phrases ( 2 = 3.51,  = 0.061 )
(c) For Words ( 2 = 30.29,  &lt; 0.00001 )
(d) For Words+Phrases ( 2 = 30.3,  &lt; 0.00001 )
          </p>
        </sec>
        <sec id="sec-5-8-2">
          <title>5.2. Domain Adaptation Study</title>
          <p>In this section we study the importance of domain specific data for training CWI systems. Here,
the classifiers are trained on a Source dataset to predict the complexity for Target test dataset. A
fraction, denoted by Tgt, of training samples from the Target dataset along with Source training
data is included in the training set. Tgt=0 corresponds to the case when the training data
contains no samples from the Target dataset. The features used for training the classifier for
each Source dataset is as described in Section 4.3. Figure 3 depicts the Accuracy and Macro-F1
scores corresponding to diferent classifiers and Tgt values. As expected, the highest scores
for each of the classifiers are obtained for Tgt=100% i.e. when the the entire training data of
the Target dataset is used for training. However, it can be observed that for CWI 2018 (See</p>
          <p>(a) Source: CWI2016, Target: CWI2018
(b) Source: CWI2018, Target: CWI2016</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>The present work aims at training and testing with diferent Gradient Boosted Tree Classifiers
for determining whether a given word of a sentence is complex in its context. Four classifiers,
namely CatBoost, XGBoost, LGBM and Random Forest, are trained on a feature space comprising
lexical, semantic and frequency based features. The context of the target word is incorporated
in the feature space using features derived from Language Model, Dependency Parse Trees
and Cosine similarity of the target with its context. Experiments on two datasets, namely
CWI2016 and CWI2018 indicate that the classifier trained on the proposed feature space using
CatBoost algorithm outperforms known baseline works. Domain adaptation has also been
studied between the two datasets to determine the generalizability of the proposed CWI system.
It was observed that even with 50% reduction in domain-specific training data, the performance
is not degraded significantly. This observation is very encouraging as it extends the applicability
of the above techniques to unseen data belonging to a variety of domains.</p>
      <p>One major takeaway from the present work is that although it is based on Tree classifiers
and hand-crafted feature space, the gradient boosting (CatBoost) system outperforms other
techniques based on Deep Neural Networks. In future, we would also like to assess the
performance of the proposed feature space with other machine learning schemes, such as LSTM,
BiLSTM, CNN.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>Raksha Agarwal acknowledges Council of Scientific and Industrial Research (CSIR), India for
supporting the research under Grant no: SPM-06/086(0267)/2018-EMR-I.
[9] G. Paetzold, L. Specia, Semeval 2016 task 11: Complex word identification, in: Proceedings
of the 10th International Workshop on Semantic Evaluation (SemEval-2016), 2016, pp.
560–569.
[10] S. M. Yimam, C. Biemann, S. Malmasi, G. Paetzold, L. Specia, S. Štajner, A. Tack, M. Zampieri,
A report on the complex word identification shared task 2018, in: Proceedings of the
Thirteenth Workshop on Innovative Use of NLP for Building Educational Applications,
Association for Computational Linguistics, New Orleans, Louisiana, 2018, pp. 66–78. URL:
https://www.aclweb.org/anthology/W18-0507. doi:10.18653/v1/W18-0507.
[11] G. Paetzold, L. Specia, Sv000gg at semeval-2016 task 11: Heavy gauge complex word
identification with system voting, in: Proceedings of the 10th International Workshop on
Semantic Evaluation (SemEval-2016), 2016, pp. 969–974.
[12] S. Sanjay, A. Kumar M, K. P. Soman, AmritaCEN at SemEval-2016 task 11: Complex word
identification using word embedding, in: Proceedings of the 10th International Workshop
on Semantic Evaluation (SemEval-2016), Association for Computational Linguistics, San
Diego, California, 2016, pp. 1022–1027. URL: https://aclanthology.org/S16-1159. doi:10.
18653/v1/S16-1159.
[13] O. Kuru, Ai-ku at semeval-2016 task 11: Word embeddings and substring features for
complex word identification, in: Proceedings of the 10th International Workshop on
Semantic Evaluation (SemEval-2016), 2016, pp. 1042–1046.
[14] P. K. Choubey, S. Pateria, Garuda &amp; bhasha at semeval-2016 task 11: Complex word
identification using aggregated learning models, in: Proceedings of the 10th International
Workshop on Semantic Evaluation (SemEval-2016), 2016, pp. 1006–1010.
[15] J. Brooke, A. L. Uitdenbogerd, T. Baldwin, Melbourne at semeval 2016 task 11: Classifying
type-level word complexity using random forests with corpus and word list features, in:
Proceedings of the 10th International Workshop on Semantic Evaluation (SemEval-2016),
2016, pp. 975–981.
[16] F. Ronzano, A. Abura’ed, L. Espinosa-Anke, H. Saggion, TALN at SemEval-2016 task 11:
Modelling complex words by contextual, lexical and semantic features, in: Proceedings
of the 10th International Workshop on Semantic Evaluation (SemEval-2016), Association
for Computational Linguistics, San Diego, California, 2016, pp. 1011–1016. URL: https:
//aclanthology.org/S16-1157. doi:10.18653/v1/S16-1157.
[17] E. Davoodi, L. Kosseim, Clac at semeval-2016 task 11: Exploring linguistic and
psycholinguistic features for complex word identification, arXiv preprint arXiv:1709.02843 (2017).
[18] M. Quijada, J. Medero, Hmc at semeval-2016 task 11: Identifying complex words using
depth-limited decision trees, in: Proceedings of the 10th International Workshop on
Semantic Evaluation (SemEval-2016), 2016, pp. 1034–1037.
[19] S. Malmasi, M. Zampieri, Maza at semeval-2016 task 11: Detecting lexical complexity
using a decision stump meta-classifier, in: Proceedings of the 10th International Workshop
on Semantic Evaluation (SemEval-2016), 2016, pp. 991–995.
[20] G. Nat, Sensible at semeval-2016 task 11: Neural nonsense mangled in ensemble mess, in:
Proceedings of the 10th International Workshop on Semantic Evaluation (SemEval-2016),
2016, pp. 963–968.
[21] M. Konkol, Uwb at semeval-2016 task 11: Exploring features for complex word
identiifcation, in: Proceedings of the 10th International Workshop on Semantic Evaluation
or percentages, Psychometrika 12 (1947) 153–157.
[50] A. L. Edwards, Note on the “correction for continuity” in testing the significance of the
diference between correlated proportions, Psychometrika 13 (1948) 185–187.</p>
    </sec>
    <sec id="sec-8">
      <title>A. Feature Space</title>
      <p>A List of all the features included in the feature space is presented in Table A.1
Nchar, Nvow, Nphon, Nsyl, Nmorph, UpCase, NcharPrev, NcharNext</p>
      <p>IsNE, EtymOrig, IsStopword
Lexical UnivTag, PennTag, Nsyn, Nhypo, Nhyper</p>
      <p>LenSent, Relpos
Lexicon InGoogle, InOgden
Frequency OgdenFreq, ECFreq, SUBTFreq, ContDiverse, TLFreq, BrownFreq
Psycholinguistic AOA, Conc, Imag, Meanc, Meanp, KFfreq, KFNcats, KFNsamp
Character LM BiCharProb, TriCharProb
Language Model FragScore3, FragScore5
Dependency DepTag, HeadPOS, TokDepth, NChild
Vector Similarity LReoftSoitmSi5m,, HRiegahdtSSiimm,5A,vLgeCftSihmil1dS,imR,igAhvtgSSimib1Sim, MaxSibSim</p>
    </sec>
    <sec id="sec-9">
      <title>B. Feature Importance</title>
      <p>For CatBoost the top 5 features corresponding to feature importance based on loss function
change are mentioned in Table B.1. For each feature this value represents the diference between
the loss value of the model with this feature and without it. For both the datsets, the feature
importance is positive for all the features included in the feature space.</p>
      <p>Table B.1</p>
      <sec id="sec-9-1">
        <title>Top 5 Features CWI2016</title>
      </sec>
      <sec id="sec-9-2">
        <title>Feature Importance CWI2018</title>
      </sec>
      <sec id="sec-9-3">
        <title>Feature Importance</title>
        <p>NChar</p>
        <p>Nphon
BrownFreq
EtymOrig
ECFreq</p>
      </sec>
    </sec>
    <sec id="sec-10">
      <title>C. Domain Adaptation</title>
      <p>The Macro-F1 and Accuracy values for diferent classifiers are mentioned in Table
C.1.</p>
      <sec id="sec-10-1">
        <title>Source Tgt</title>
      </sec>
      <sec id="sec-10-2">
        <title>CatBooost</title>
      </sec>
      <sec id="sec-10-3">
        <title>XGBoost</title>
      </sec>
      <sec id="sec-10-4">
        <title>LGBM RF</title>
      </sec>
      <sec id="sec-10-5">
        <title>Accuracy Macro-F1 Accuracy Macro-F1 Accuracy Macro-F1 Accuracy Macro-F1</title>
      </sec>
      <sec id="sec-10-6">
        <title>Target: CWI2018 CWI 2016</title>
        <p>CWI 2018 2550%%</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Zampieri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Malmasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Paetzold</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Specia</surname>
          </string-name>
          ,
          <article-title>Complex word identification: Challenges in data annotation and system performance</article-title>
          ,
          <source>arXiv preprint arXiv:1710.04989</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Shardlow</surname>
          </string-name>
          ,
          <article-title>A comparison of techniques to automatically identify complex words</article-title>
          .,
          <source>in: 51st Annual Meeting of the Association for Computational Linguistics Proceedings of the Student Research Workshop</source>
          ,
          <year>2013</year>
          , pp.
          <fpage>103</fpage>
          -
          <lpage>109</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>G. H.</given-names>
            <surname>Paetzold</surname>
          </string-name>
          , L. Specia,
          <string-name>
            <surname>Plumberr:</surname>
          </string-name>
          <article-title>An automatic error identification framework for lexical simplification</article-title>
          ,
          <source>in: Proceedings of the first international workshop on Quality Assessment for Text Simplification (QATS)</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Shardlow</surname>
          </string-name>
          ,
          <article-title>A survey of automated text simplification</article-title>
          ,
          <source>International Journal of Advanced Computer Science and Applications</source>
          <volume>4</volume>
          (
          <year>2014</year>
          )
          <fpage>58</fpage>
          -
          <lpage>70</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>P.</given-names>
            <surname>Sikka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Mago</surname>
          </string-name>
          ,
          <article-title>A survey on text simplification</article-title>
          , ArXiv abs/
          <year>2008</year>
          .08612 (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>G.</given-names>
            <surname>Leroy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. E.</given-names>
            <surname>Endicott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Kauchak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Mouradi</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Just, User evaluation of the efects of a text simplification algorithm using term familiarity on perception, understanding, learning, and information retention</article-title>
          ,
          <source>Journal of medical Internet research 15</source>
          (
          <year>2013</year>
          )
          <article-title>e144</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>N.</given-names>
            <surname>Elhadad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Sutaria</surname>
          </string-name>
          ,
          <article-title>Mining a lexicon of technical terms and lay equivalents, in: Biological, translational, and clinical language processing</article-title>
          ,
          <source>Association for Computational Linguistics</source>
          , Prague, Czech Republic,
          <year>2007</year>
          , pp.
          <fpage>49</fpage>
          -
          <lpage>56</lpage>
          . URL: https://aclanthology.org/W07-1007.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>C.</given-names>
            <surname>Horn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Manduca</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Kauchak</surname>
          </string-name>
          ,
          <article-title>Learning a lexical simplifier using wikipedia, in: Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics</article-title>
          (Volume
          <volume>2</volume>
          :
          <string-name>
            <surname>Short</surname>
            <given-names>Papers)</given-names>
          </string-name>
          ,
          <year>2014</year>
          , pp.
          <fpage>458</fpage>
          -
          <lpage>463</lpage>
          . (SemEval-2016),
          <year>2016</year>
          , pp.
          <fpage>1038</fpage>
          -
          <lpage>1041</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>K.</given-names>
            <surname>Wróbel</surname>
          </string-name>
          , Plujagh at semeval
          <article-title>-2016 task 11: Simple system for complex word identification</article-title>
          ,
          <source>in: Proceedings of the 10th International Workshop on Semantic Evaluation (SemEval2016)</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>953</fpage>
          -
          <lpage>957</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>D.</given-names>
            <surname>Kauchak</surname>
          </string-name>
          , Pomona at semeval-2016 task 11:
          <article-title>Predicting word complexity based on corpus frequency</article-title>
          ,
          <source>in: Proceedings of the 10th International Workshop on Semantic Evaluation (SemEval-2016)</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>1047</fpage>
          -
          <lpage>1051</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>S.</given-names>
            <surname>Gooding</surname>
          </string-name>
          , E. Kochmar, Camb at cwi shared task 2018:
          <article-title>Complex word identification with ensemble-based voting</article-title>
          ,
          <source>in: Proceedings of the Thirteenth Workshop on Innovative Use of NLP for Building Educational Applications</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>184</fpage>
          -
          <lpage>194</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [25] A. AbuRa'ed, H. Saggion,
          <article-title>LaSTUS/TALN at complex word identification (CWI) 2018 shared task</article-title>
          ,
          <source>in: Proceedings of the Thirteenth Workshop on Innovative Use of NLP for Building Educational Applications</source>
          , Association for Computational Linguistics, New Orleans, Louisiana,
          <year>2018</year>
          , pp.
          <fpage>159</fpage>
          -
          <lpage>165</lpage>
          . URL: https://www.aclweb.org/anthology/W18-0517. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>W18</fpage>
          -0517.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [26]
          <string-name>
            <surname>D. De Hertog</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Tack</surname>
          </string-name>
          ,
          <article-title>Deep learning architecture for complexword identification</article-title>
          , in: Thirteenth Workshop of Innovative Use of
          <article-title>NLP for Building Educational Applications, Association for Computational Linguistics (ACL); New Orleans</article-title>
          , Louisiana,
          <year>2018</year>
          , pp.
          <fpage>328</fpage>
          -
          <lpage>334</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>S. T.</given-names>
            <surname>Aroyehun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Angel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. A. P.</given-names>
            <surname>Alvarez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gelbukh</surname>
          </string-name>
          ,
          <article-title>Complex word identification: Convolutional neural network vs. feature engineering</article-title>
          ,
          <source>in: Proceedings of the thirteenth workshop on innovative use of NLP for building educational applications</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>322</fpage>
          -
          <lpage>327</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>S.</given-names>
            <surname>Gooding</surname>
          </string-name>
          , E. Kochmar,
          <article-title>Complex word identification as a sequence labelling task</article-title>
          ,
          <source>in: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>1148</fpage>
          -
          <lpage>1153</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>P.</given-names>
            <surname>Finnimore</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Fritzsch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>King</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sneyd</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. U.</given-names>
            <surname>Rehman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Alva-Manchego</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Vlachos</surname>
          </string-name>
          ,
          <article-title>Strong baselines for complex word identification across multiple languages</article-title>
          , arXiv preprint arXiv:
          <year>1904</year>
          .
          <volume>05953</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [30]
          <string-name>
            <surname>K. C. Sheang</surname>
          </string-name>
          ,
          <article-title>Multilingual complex word identification: Convolutional neural networks with morphological and linguistic features</article-title>
          ,
          <source>in: Proceedings of the Student Research Workshop Associated with RANLP</source>
          <year>2019</year>
          ,
          <year>2019</year>
          , pp.
          <fpage>83</fpage>
          -
          <lpage>89</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ehara</surname>
          </string-name>
          ,
          <article-title>Graph-based analysis of similarities between word frequency distributions of various corpora for complex word identification</article-title>
          ,
          <source>in: 2019 18th IEEE International Conference On Machine Learning And Applications (ICMLA)</source>
          , IEEE,
          <year>2019</year>
          , pp.
          <fpage>1982</fpage>
          -
          <lpage>1986</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>G.-E.</given-names>
            <surname>Zaharia</surname>
          </string-name>
          , D.-C. Cercel,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dascalu</surname>
          </string-name>
          ,
          <article-title>Cross-lingual transfer learning for complex word identification</article-title>
          ,
          <source>in: 2020 IEEE 32nd International Conference on Tools with Artificial Intelligence (ICTAI)</source>
          , IEEE,
          <year>2020</year>
          , pp.
          <fpage>384</fpage>
          -
          <lpage>390</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>A.</given-names>
            <surname>Palmero Aprosio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Menini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Tonelli</surname>
          </string-name>
          ,
          <article-title>Adaptive complex word identification through false friend detection</article-title>
          ,
          <source>in: Proceedings of the 28th ACM Conference on User Modeling, Adaptation and Personalization</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>192</fpage>
          -
          <lpage>200</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>E. L.</given-names>
            <surname>Thorndike</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Lorge</surname>
          </string-name>
          ,
          <article-title>The teacher's word book of 30,000 words</article-title>
          . (
          <year>1944</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [35]
          <string-name>
            <surname>G. D. Brown</surname>
          </string-name>
          ,
          <article-title>A frequency count of 190,000 words in the london-lund corpus of english conversation</article-title>
          ,
          <source>Behavior Research Methods, Instruments, &amp; Computers</source>
          <volume>16</volume>
          (
          <year>1984</year>
          )
          <fpage>502</fpage>
          -
          <lpage>532</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>R.</given-names>
            <surname>Agarwal</surname>
          </string-name>
          , N. Chatterjee,
          <article-title>LangResearchLab_NC at CMCL2021 shared task: Predicting gaze behaviour using linguistic features and tree regressors</article-title>
          ,
          <source>in: Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics</source>
          , Association for Computational Linguistics, Online,
          <year>2021</year>
          , pp.
          <fpage>79</fpage>
          -
          <lpage>84</lpage>
          . URL: https://www.aclweb.org/anthology/2021.cmcl-
          <volume>1</volume>
          . 8. doi:
          <volume>10</volume>
          .18653/v1/
          <year>2021</year>
          .cmcl-
          <volume>1</volume>
          .8.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>G.</given-names>
            <surname>Paetzold</surname>
          </string-name>
          , L. Specia,
          <article-title>Inferring psycholinguistic properties of words, in: Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Association for Computational Linguistics</article-title>
          , San Diego, California,
          <year>2016</year>
          , pp.
          <fpage>435</fpage>
          -
          <lpage>440</lpage>
          . URL: https://www.aclweb.org/anthology/N16-1050.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>M.</given-names>
            <surname>Wilson</surname>
          </string-name>
          ,
          <article-title>Mrc psycholinguistic database: Machine-usable dictionary</article-title>
          , version
          <volume>2</volume>
          .00, Behavior research methods, instruments, &amp; computers
          <volume>20</volume>
          (
          <year>1988</year>
          )
          <fpage>6</fpage>
          -
          <lpage>10</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [39]
          <string-name>
            <given-names>H.</given-names>
            <surname>Kučera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. N.</given-names>
            <surname>Francis</surname>
          </string-name>
          ,
          <article-title>Computational analysis of present-day American English</article-title>
          , University Press of New England,
          <year>1967</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [40]
          <string-name>
            <given-names>D.</given-names>
            <surname>Graf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Maeda</surname>
          </string-name>
          , English gigaword,
          <source>Linguistic Data Consortium, Philadelphia</source>
          <volume>4</volume>
          (
          <year>2003</year>
          )
          <fpage>34</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [41]
          <string-name>
            <given-names>R.</given-names>
            <surname>Agarwal</surname>
          </string-name>
          , N. Chatterjee, LangResearchLab NC at SemEval
          <article-title>-2021 task 1: Linguistic feature based modelling for lexical complexity</article-title>
          ,
          <source>in: Proceedings of the 15th International Workshop on Semantic Evaluation (SemEval-2021)</source>
          ,
          <article-title>Association for Computational Linguistics</article-title>
          , Online,
          <year>2021</year>
          , pp.
          <fpage>120</fpage>
          -
          <lpage>125</lpage>
          . URL: https://aclanthology.org/
          <year>2021</year>
          .semeval-
          <volume>1</volume>
          .10. doi:
          <volume>10</volume>
          .18653/v1/
          <year>2021</year>
          . semeval-
          <volume>1</volume>
          .
          <fpage>10</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [42]
          <string-name>
            <given-names>J.</given-names>
            <surname>Pennington</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Socher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. D.</given-names>
            <surname>Manning</surname>
          </string-name>
          , Glove:
          <article-title>Global vectors for word representation</article-title>
          ,
          <source>in: Empirical Methods in Natural Language Processing (EMNLP)</source>
          ,
          <year>2014</year>
          , pp.
          <fpage>1532</fpage>
          -
          <lpage>1543</lpage>
          . URL: http://www.aclweb.org/anthology/D14-1162.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [43]
          <string-name>
            <given-names>T.</given-names>
            <surname>Mikolov</surname>
          </string-name>
          , E. Grave,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bojanowski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Puhrsch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Joulin</surname>
          </string-name>
          ,
          <article-title>Advances in pre-training distributed word representations</article-title>
          ,
          <source>in: Proceedings of the International Conference on Language Resources and Evaluation (LREC</source>
          <year>2018</year>
          ),
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [44]
          <string-name>
            <given-names>P.</given-names>
            <surname>Bojanowski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Grave</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Joulin</surname>
          </string-name>
          , T. Mikolov,
          <article-title>Enriching word vectors with subword information, Transactions of the Association for Computational Linguistics 5 (</article-title>
          <year>2017</year>
          )
          <fpage>135</fpage>
          -
          <lpage>146</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [45]
          <string-name>
            <surname>T. K. Ho</surname>
          </string-name>
          ,
          <article-title>The random subspace method for constructing decision forests</article-title>
          ,
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          <volume>20</volume>
          (
          <year>1998</year>
          )
          <fpage>832</fpage>
          -
          <lpage>844</lpage>
          . doi:
          <volume>10</volume>
          .1109/34.709601.
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [46]
          <string-name>
            <given-names>T.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Guestrin</surname>
          </string-name>
          ,
          <article-title>Xgboost: A scalable tree boosting system</article-title>
          ,
          <source>in: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</source>
          , KDD '16,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2016</year>
          , p.
          <fpage>785</fpage>
          -
          <lpage>794</lpage>
          . URL: https://doi.org/10.1145/2939672.2939785. doi:
          <volume>10</volume>
          .1145/2939672.2939785.
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [47]
          <string-name>
            <given-names>G.</given-names>
            <surname>Ke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Meng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Finley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Chen</surname>
          </string-name>
          , W. Ma,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Ye</surname>
          </string-name>
          , T.-Y. Liu,
          <article-title>Lightgbm: A highly eficient gradient boosting decision tree</article-title>
          , in: I. Guyon,
          <string-name>
            <given-names>U. V.</given-names>
            <surname>Luxburg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bengio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wallach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fergus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Vishwanathan</surname>
          </string-name>
          , R. Garnett (Eds.),
          <source>Advances in Neural Information Processing Systems</source>
          , volume
          <volume>30</volume>
          ,
          <string-name>
            <surname>Curran</surname>
            <given-names>Associates</given-names>
          </string-name>
          , Inc.,
          <year>2017</year>
          . URL: https://proceedings.neurips.cc/ paper/2017/file/6449f44a102fde848669bdd9eb6b76fa-Paper.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [48]
          <string-name>
            <given-names>L.</given-names>
            <surname>Prokhorenkova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Gusev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Vorobev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. V.</given-names>
            <surname>Dorogush</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gulin</surname>
          </string-name>
          ,
          <article-title>Catboost: unbiased boosting with categorical features</article-title>
          , in: S. Bengio,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wallach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Larochelle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Grauman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Cesa-Bianchi</surname>
          </string-name>
          , R. Garnett (Eds.),
          <source>Advances in Neural Information Processing Systems</source>
          , volume
          <volume>31</volume>
          ,
          <string-name>
            <surname>Curran</surname>
            <given-names>Associates</given-names>
          </string-name>
          , Inc.,
          <year>2018</year>
          . URL: https://proceedings.neurips.cc/paper/2018/ file/14491b756b3a51daac41c24863285549-Paper.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [49]
          <string-name>
            <given-names>Q.</given-names>
            <surname>McNemar</surname>
          </string-name>
          ,
          <article-title>Note on the sampling error of the diference between correlated proportions</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>