Hate Speech Detection with Machine-Translated Data: The Role of Annotation Scheme, Class Imbalance and Undersampling Camilla Casula Sara Tonelli Fondazione Bruno Kessler Fondazione Bruno Kessler Trento, Italy Trento, Italy ccasula@fbk.eu satonelli@fbk.eu Abstract One possible solution to alleviate data sparseness is the use of machine translated data from English While using machine-translated data for to less resourced languages for training classifiers, supervised training can alleviate data exploiting the large amount of data available for sparseness problems when dealing with English. This has already been used in the con- less-resourced languages, it is important text of hate speech detection (Sohn and Lee, 2019; that the source data are not only correctly Casula et al., 2020) but results have not been con- translated, but also follow the same anno- sistent across languages. tation scheme and possibly class balance An additional issue is the fact that there is no as the smaller dataset in the target lan- shared fixed definition within the NLP community guage. We therefore present an evaluation of what type of language constitutes hate speech. of hate speech detection in Italian using Indeed, there are typically large differences among machine-translated data from English and hate speech and abusive language datasets in terms comparing three settings, in order to un- of annotation frameworks and their applications derstand the impact of training size, class in practice (Caselli et al., 2020). In addition to distribution and annotation scheme.1 this, there can be large variations between datasets 1 Introduction in terms of size and class balance. Possible is- sues affecting the behaviour of classifiers trained The task of detecting hate speech on social me- on machine-translated data, such as different class dia has been attracting increasing attention due to distribution in source and target language, or dif- the negative effects this phenomenon can have on ferent annotation scheme, have not been analysed. online communities and society as a whole. The In order to fill this gap, we explore the impact of development of systems which can effectively de- these differences between datasets when perform- tect hate speech has therefore become increasingly ing hate speech detection in Italian using machine- important for academics and tech companies alike. translated data from English. Our goal is to ad- One of the difficulties of producing accurate dress the three following questions: hate speech detection systems is the need for large, high-quality datasets, the creation of which is time • What performance can we expect by us- and resource-consuming. English can count on the ing only machine translated data, given that highest number of hate speech detection datasets, translation quality for social media language as well as the ones with the largest sizes, with may be problematic? up to 150k posts for a single dataset (Gomez et al., 2020). Other languages such as Italian, on • Is it better to use a larger translated set for the other hand, can count on fewer datasets which training, even by merging slightly different tend to be smaller (Vidgen and Derczynski, 2020). classes, or a smaller, more precise one? Given that machine learning methods are typically used for this task, the use of small datasets can • What is the impact of class imbalance, and to lead to overfitting problems due to the lack of lin- what extent can undersampling be effective? guistic variation (Vidgen and Derczynski, 2020). 1 The above questions are addressed by compar- Copyright c 2020 for this paper by its authors. Use per- mitted under Creative Commons License Attribution 4.0 In- ing three experimental settings that are described ternational (CC BY 4.0). in Section 4 and evaluated in Section 5. 2 Related Work schemes across corpora, as well as similar annota- tion schemes being applied in different ways. In recent years, the number of research works fo- cused on the detection of hate speech on social me- 3 Data dia has remarkably increased, mostly due to the growing awareness regarding the societal impact Since tweets containing hate speech or abusive these platforms can have. language constitute a very small subset (between Computational methods for detecting the pres- 0.1% and 3% depending on the label used) of all ence of hate speech on the web have become nec- tweets being posted (Founta et al., 2018), ran- essary due to the extremely large amounts of user- dom samples are generally not used for annota- generated content being posted each day. These tion, because the final datasets would contain an methods typically rely on supervised learning, in extremely low number of positive class examples, the form of both traditional machine learning (e.g. which would make classification difficult. The support vector classifiers) and deep learning ap- typical solution to this is to preselect posts that proaches (Schmidt and Wiegand, 2017). Given are likely to contain hateful language by search- the increased attention towards this topic, more ing for specific hate-related keywords. While this and more shared tasks regarding hate speech and method is effective for gathering more instances of abusive language detection have emerged, such hate speech, it can make datasets biased, which is as the HaSpeeDe task at Evalita 2018 (Bosco et a main issue in hate speech datasets (Wiegand et al., 2018), OffensEval (Zampieri et al., 2019) and al., 2019). HatEval (Basile et al., 2019) at SemEval 2019, The dataset we chose for training our system is and the multilingual OffensEval at SemEval 2020 described in Founta et al. (2018). This dataset was (Zampieri et al., 2020). not created starting from a set of predefined of- Systems based on Transformers architectures fensive terms or hashtags in order to reduce bias, such as BERT (Devlin et al., 2019) have proven which was an important factor in our choice. The effective for hate speech detection and classifica- method used by Founta et al. (2018) to increase tion in both English (Zampieri et al., 2019) and the percentage of hateful/abusive tweets is boosted Italian (Polignano et al., 2019a). These systems random sampling, in which a portion of the dataset are generally pre-trained on large unlabeled cor- is “boosted” with tweets that are more likely to be- pora through two self-supervised tasks (next sen- long in the minority classes. The boosted set of tence prediction and masked language modeling) tweets is created using text analysis and machine to create language models which can then be fine- learning (Founta et al., 2018). tuned to a variety of downstream tasks using la- The dataset was annotated through crowdsourc- beled data. ing using the labels hateful, abusive, spam, and AlBERTo (Polignano et al., 2019b) is a BERT- normal. The definition of hate speech given by based system which was pre-trained on Italian Founta et al. (2018) to the annotators, based on Twitter data, and it currently defines the state of existing literature on the topic, is: the art for hate speech detection in Italian (Polig- Hate Speech: Language used to express nano et al., 2019a). hatred towards a targeted individual or Recently, more attention has been directed to- group, or is intended to be derogatory, wards the quality of hate and abuse detection sys- to humiliate, or to insult the members tems. Vidgen et al. (2019) investigate the flaws of the group, on the basis of attributes presented by most abusive language detection such as race, religion, ethnic origin, sex- datasets in circulation: they can contain systematic ual orientation, disability, or gender. biases towards certain types and targets of abuse, they are subject to degradation over time, they typ- The abusive label, on the other hand, is the re- ically present very low inter-annotator agreement, sult of three separate labels (abusive, offensive, and they can vary greatly with respect to quality, and aggressive) being combined. In preliminary size, and class balance. Vidgen and Derczynski annotation rounds, Founta et al. (2018) found that (2020) further analyse the role of datasets in the these three labels were significantly correlated, so detection of abuse, addressing issues such as the they grouped them together. The definition of abu- use of different task descriptions and annotation sive language given to the annotators is: Abusive Language: Any strongly im- In order to investigate this, we compare three polite, rude or hurtful language using different experimental settings. In the first one, profanity, that can show a debasement of we fine-tune AlBERTo on the translated tweets someone or something, or show intense in Founta et al. (2018) after merging the hate- emotion. ful and abusive classes together, mapping them to a single hateful class as required by the bi- While the Founta et al. (2018) dataset was orig- nary classification task at Evalita 2018. In a sec- inally comprised of 80k tweets, Twitter datasets ond setting, AlBERTo is fine-tuned on the hate- can often be subject to degradation due to tweets ful class alone, discarding all tweets annotated as being removed over time and not accessible any- abusive in Founta et al. (2018). We hypothesize more through tweet IDs (Vidgen et al., 2019). Af- this setting may perform better when tested on the ter retrieving all available tweets and after remov- HaSpeeDe data, given the higher similarity in an- ing tweets annotated as spam, the total number notation framework. of tweets we use for training is 12,379, of which Simply removing tweets annotated as abusive, 727 are annotated as hateful and 1,792 as abusive. however, can throw off the balance between Before translating the data into Italian, we pre- classes. More specifically, when training the sys- process it using the Ekphrasis tool 2 to tokenise the tem on both abusive and hateful tweets the hate- text and normalise user mentions, URLs (replaced ful+abusive class constitutes about 20% of our by and respectively), as well as data, while when we only use tweets annotated numbers, which are substituted with a number as hateful this percentage drops to 7%, potentially tag. We then use the Google Translate API to affecting classification results. In particular, the translate the data into Italian, in order to use it as data we use for testing has a different class bal- training data for our classifier. ance, with 30% of tweets marked as hateful. In For testing, we use the test portion of the Twit- order to assess the impact of class imbalance on ter dataset used in the Hate Speech Detection our results, we further evaluate each setting using (HaSpeeDe) task at Evalita 2018 (Bosco et al., undersampling (Kubat, 2000; Sun et al., 2009), a 2018), consisting of 1,000 Italian tweets manu- technique typically used for imbalanced classifi- ally annotated for hate speech against immigrants. cation, in which we reduce the number of tweets This dataset is a simplified version of the dataset belonging to the majority class, so that the overall described in (Sanguinetti et al., 2018), in which percentage of tweets containing hate increases. more fine-grained labels are used. Given that undersampling our data reduces the total size of tweets available for training, the re- 4 Experimental Setup sulting datasets for each annotation scheme con- We experiment with the fine-tuning of AlBERTo siderably differ in size. We therefore consider a (Polignano et al., 2019b), a BERT-based language third setting, in which we use further random un- model pre-trained on Italian Twitter data, using dersampling (Kubat, 2000; Sun et al., 2009) to data that was automatically translated from En- match the larger dataset (hateful+abusive) with the glish. This model has achieved state-of-the-art re- smaller one (hateful only), so that the two annota- sults when fine-tuned on the training data from the tions can be effectively compared in a setting with HaSpeeDe task at Evalita 2018 (Polignano et al., equal class balance and sample size. 2019a). In summary, the three data settings we train our Our goal is that of exploring the impact of dif- system on are: ferent annotation schemes and class balance when using machine-translated data for hate speech de- 1. Hateful and abusive tweets, using undersam- tection. Indeed, merging fine-grained classes into pling to progressively lower class imbalance; coarser ones has been a common and accepted practice when creating larger training sets from a 2. Hateful only tweets, again using undersam- smaller one (e.g. Founta et al. (2019)). This step pling to progressively lower class imbalance; has been performed also to compare classification 3. Hateful and abusive tweets, both using un- in different languages (Corazza et al., 2020). dersampling to progressively lower class im- 2 https://github.com/cbaziotis/ekphrasis balance as in the previous settings, and using further random undersampling to match the classes are therefore extremely imbalanced before low sample sizes of setting 2. undersampling. Predictably, with the classes be- ing this imbalanced, the system identifies all test Our AlBERTo fine-tuning architecture consists instances as belonging to the majority class. This of a pooling layer for extracting the AlBERTo hid- again happens with the minority class comprising den representation for each sequence, followed 20% of the training data. by a dropout layer (dropout rate 0.2), two dense layers of size 768 and 128 and, finally, a soft- Setting 2: Hateful only max layer. We use L2 regularization (λ=0.01), % hate Size (tweets) Macro-F1 Hate class F1 7% 10,587 0.40 0 Adam optimizer (2e-5 learning rate), and categor- 20% 3,635 0.40 0 ical cross-entropy loss. We train the system for 5 30% 2,423 0.65 0.54 epochs with batch size 32. 40% 1,818 0.52 0.56 5 Results and Discussion Table 2: Scores obtained when fine-tuning Al- BERTo on tweets labeled as hateful only. We measure the classification results using both macro-F1 score and minority class F1 score. We Similarly to Setting 1, the best classification repeat each run five times in order to compensate performance in this case is achieved with 30% of for random initialization, and we report the aver- minority class tweets. Interestingly, the best per- age scores of these runs. formance is comparable to the one obtained in Set- 5.1 Setting 1: Hateful + Abusive Tweets ting 1, even though in this case the number of training samples available is much lower, suggest- The classification results obtained when fine- ing that more task-specific training instances can tuning AlBERTo on both abusive and hateful impact performance. We can note a difference tweets combined can be observed in Table 1. with the minority class at 40% of total data, in Setting 1: Hateful + abusive which the performance drops in terms of macro-F1 % hate Size (tweets) Macro-F1 Hate class F1 score, likely due to the very small number of sam- 20% 12,379 0.40 0 ples available for training and the consequent lack 30% 8,397 0.64 0.52 of linguistic variation. The hate class F1 score, 40% 6,298 0.63 0.57 however, remains stable. Table 1: Scores obtained when fine-tuning Al- State-of-the-art results obtained by fine-tuning BERTo on both hateful and abusive tweets. AlBERTo on the same Evalita dataset as reported in Polignano et al. (2019a) reach 0.80 macro-F1 The class balance of the dataset prior to un- and 0.73 F1 on the hate class, which we can con- dersampling is 20% hateful + abusive tweets and sider an upper-bound for our task, obtained in 80% non-hateful, which amounts to 12,379 tweets a fully-supervised monolingual setting. On the total. With this class balance, the system per- other hand, the most frequent label baseline is forms the worst, classifying every tweet as be- 0.40 macro-F1, which is clearly outperformed us- longing to the majority non-hateful class. On the ing only machine-translated data. other hand, with a higher percentage of minor- ity class instances, the classification results im- 5.3 Setting 3: Hateful + Abusive Tweets prove, in spite of the considerably smaller amount (Random Undersampling) of training data available. These results suggest Since there are large differences in size between that consistency in class balance can play a bigger the hateful+abusive annotation and the hateful- role than training data size in classification results only annotation, we randomly undersample the in this context. hateful+abusive training data so that it matches the size of the hateful-only training data, in order to 5.2 Setting 2: Hateful Only Tweets allow us to effectively compare the impact of each The performance of the system when fine-tuned on annotation framework on our results. The classifi- tweets labeled as hateful only is reported in Table cation performance is reported in Table 3. 2. As previously mentioned, only 7% of tweets If we compare the results of Setting 3 with in the dataset we use are labeled as hateful. The those of Setting 2, it is clear that using more task- Setting 3: Hateful + abusive (random undersampling) (Pamungkas et al., 2020). In the case of abusive % hate Size (tweets) Macro-F1 Hate class F1 30% 2,423 0.58 0.38 tweets, we observe that the offenses are less direct 40% 1,818 0.59 0.51 and therefore slurs tend to be translated poorly. See for example the following sentence, which Table 3: Scores obtained when fine-tuning Al- is labeled as abusive in the Founta et al. (2018) BERTo on tweets labeled as hateful and abusive, dataset: after random undersampling. (1) use that ugly ass design [...] specific data, in this case hateful-only tweets, can utilizzare quel disegno asino brutto [...] lead to a larger improvement in performance when use that design donkey ugly [...] the amount of training data is the same. This sug- Here, “ass” is translated with “asino” (“don- gests that consistency in annotation between train- key”), effectively removing the profanity in the ing and test data can have a positive impact on translated tweet and changing completely the classification, although it is not fundamental to meaning of the message. help classification of hate speech detection with On the other hand, when profanities are used machine translated data. In fact, other aspects such in a more direct way, or when they are expressed as class balance can also play an important role. through unambiguous words such as “idiot” and 5.4 Qualitative Analysis “stupid”, they tend to be translated correctly, con- Another aspect affecting classification, which we tributing to a correct classification. Example 2 have not considered so far, is the quality of ma- shows a hateful tweet which was translated almost chine translation, a particularly challenging task correctly, retaining its offensiveness in the target on social media data (Michel and Neubig, 2018). language. In order to assess the impact of translation qual- (2) what happens when you put idiots in charge ity on our results, two annotators with linguistic cosa succede quando si mette idioti in carica background manually analysed 500 samples from the training data, consisting of 300 tweets anno- 6 Conclusions tated as normal, 100 as hateful, and 100 as abu- sive. Each annotator checked manually 250 ran- In this paper we analysed the impact of machine- dom tweets from this sample. Translation qual- translated data on Italian hate speech detection in ity was evaluated using the semantic adequacy an- a zero-shot setting. Our experiments show that notation scheme proposed in Dorr et al. (2011, when using machine-translated data for training p. 807). Annotations are judged on a scale be- it is possible to learn a classification model that tween -3 and 3, with scores below 0 for inadequate clearly outperforms the most-frequent baseline, translations and above 0 for adequate ones. The even if translation quality is affected by the jar- averaged annotations for each class are reported in gon used in social media data. We found that Table 4. using more task-specific data can have a positive impact on classification performance even with Normal Hateful Abusive Overall lower sample sizes compared to larger, less tar- Average 0.438 0.527 -0.043 0.368 geted datasets. Table 4: Average translation quality scores. Consistency in class distribution of training and test data can have a bigger impact than the size of Overall, translations tend towards adequacy, but the training set, or the annotation scheme. Indeed, the average scores are below 1 for all classes. using only the original training set translated into Interestingly, tweets annotated as abusive show Italian, without undersampling, classification per- poorer translation quality than other classes. This formance would be poor. could help explain the small differences in classi- In the future, we plan to extend this kind of eval- fication performance between our experiments. uation to new language pairs and new datasets, to A major role is played in this context by profan- check whether the findings obtained on the En- ities, which are often used to offend a target but glish – Italian pair are confirmed also with other can also appear in non derogatory messages ex- languages. changed among members of the same community References abusive behavior. In 12th International AAAI Con- ference on Web and Social Media. Valerio Basile, Cristina Bosco, Elisabetta Fersini, Debora Nozza, Viviana Patti, Francisco Manuel Antigoni Maria Founta, Despoina Chatzakou, Nicolas Rangel Pardo, Paolo Rosso, and Manuela San- Kourtellis, Jeremy Blackburn, Athena Vakali, and Il- guinetti. 2019. SemEval-2019 task 5: Multilin- ias Leontiadis. 2019. A unified deep learning archi- gual detection of hate speech against immigrants and tecture for abuse detection. In Proceedings of the women in twitter. In Proceedings of the 13th Inter- 10th ACM Conference on Web Science, WebSci ’19, national Workshop on Semantic Evaluation, pages page 105–114, New York, NY, USA. Association for 54–63, Minneapolis, Minnesota, USA, June. Asso- Computing Machinery. ciation for Computational Linguistics. Raul Gomez, Jaume Gibert, Lluis Gomez, and Dimos- Cristina Bosco, Dell’Orletta Felice, Fabio Poletto, thenis Karatzas. 2020. Exploring hate speech de- Manuela Sanguinetti, and Tesconi Maurizio. 2018. tection in multimodal publications. In 2020 IEEE Overview of the evalita 2018 hate speech detection Winter Conference on Applications of Computer Vi- task. In EVALITA 2018-Sixth Evaluation Campaign sion (WACV), pages 1459–1467, 03. of Natural Language Processing and Speech Tools for Italian, volume 2263, pages 1–9, Turin, Italy. M. Kubat. 2000. Addressing the curse of imbalanced CEUR. training sets: One-sided selection. Fourteenth Inter- national Conference on Machine Learning, 06. Tommaso Caselli, Valerio Basile, Jelena Mitrovic, Inga Kartoziya, and Michael Granitzer. 2020. I feel of- Paul Michel and Graham Neubig. 2018. MTNT: A fended, don’t be abusive! implicit/explicit messages testbed for machine translation of noisy text. In Pro- in offensive and abusive language. In Nicoletta ceedings of the 2018 Conference on Empirical Meth- Calzolari, Frédéric Béchet, Philippe Blache, Khalid ods in Natural Language Processing, pages 543– Choukri, Christopher Cieri, Thierry Declerck, Sara 553, Brussels, Belgium, October-November. Asso- Goggi, Hitoshi Isahara, Bente Maegaard, Joseph ciation for Computational Linguistics. Mariani, Hélène Mazo, Asunción Moreno, Jan Odijk, and Stelios Piperidis, editors, Proceedings of Endang Wahyu Pamungkas, Valerio Basile, and Vi- The 12th Language Resources and Evaluation Con- viana Patti. 2020. Do you really want to hurt ference, LREC 2020, Marseille, France, May 11-16, me? predicting abusive swearing in social media. 2020, pages 6193–6202. European Language Re- In Nicoletta Calzolari, Frédéric Béchet, Philippe sources Association. Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hitoshi Isahara, Bente Mae- Camilla Casula, Alessio Palmero Aprosio, Stefano gaard, Joseph Mariani, Hélène Mazo, Asunción Menini, and Sara Tonelli. 2020. Fbk-dh at semeval- Moreno, Jan Odijk, and Stelios Piperidis, edi- 2020 task 12: Using multi-channel bert for multilin- tors, Proceedings of The 12th Language Resources gual offensive language detection. In Proceedings and Evaluation Conference, LREC 2020, Marseille, of Offenseval. France, May 11-16, 2020, pages 6237–6246. Euro- pean Language Resources Association. Michele Corazza, Stefano Menini, Elena Cabrio, Sara Tonelli, and Serena Villata. 2020. A multilingual Marco Polignano, Pierpaolo Basile, Marco de Gemmis, evaluation for online hate speech detection. ACM and Giovanni Semeraro. 2019a. Hate speech detec- Trans. Internet Techn., 20(2):10:1–10:22. tion through alberto italian language understanding model. In NL4AI@ AI* IA. Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Marco Polignano, Pierpaolo Basile, Marco de Gem- deep bidirectional transformers for language under- mis, Giovanni Semeraro, and Valerio Basile. 2019b. standing. In Proceedings of the 2019 Conference of AlBERTo: Italian BERT Language Understanding the North American Chapter of the Association for Model for NLP Challenging Tasks Based on Tweets. Computational Linguistics: Human Language Tech- In Proceedings of the Sixth Italian Conference on nologies, Volume 1 (Long and Short Papers), pages Computational Linguistics (CLiC-it 2019), volume 4171–4186, Minneapolis, Minnesota, June. Associ- 2481. CEUR. ation for Computational Linguistics. Manuela Sanguinetti, Fabio Poletto, Cristina Bosco, Bonnie J Dorr, Joseph Olive, John McCary, and Caitlin Viviana Patti, and Marco Stranisci. 2018. An Ital- Christianson, 2011. Machine Translation Evalua- ian twitter corpus of hate speech against immigrants. tion and Optimization, pages 745 – 843. Springer In Proceedings of the Eleventh International Confer- New York. ence on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan, May. European Language Antigoni Maria Founta, Constantinos Djouvas, De- Resources Association (ELRA). spoina Chatzakou, Ilias Leontiadis, Jeremy Black- burn, Gianluca Stringhini, Athena Vakali, Michael Anna Schmidt and Michael Wiegand. 2017. A survey Sirivianos, and Nicolas Kourtellis. 2018. Large on hate speech detection using natural language pro- scale crowdsourcing and characterization of twitter cessing. In Proceedings of the Fifth International Workshop on Natural Language Processing for So- cial Media, pages 1–10, Valencia, Spain, April. As- sociation for Computational Linguistics. Hajung Sohn and Hyunju Lee. 2019. Mc-bert4hate: Hate speech detection using multi-channel bert for different languages and translations. In 2019 In- ternational Conference on Data Mining Workshops (ICDMW), pages 551–559. IEEE. Y. Sun, A. Wong, and M. Kamel. 2009. Classification of imbalanced data: a review. Int. J. Pattern Recog- nit. Artif. Intell., 23:687–719. Bertie Vidgen and Leon Derczynski. 2020. Direc- tions in abusive language training data: Garbage in, garbage out. ArXiv, abs/2004.01670. Bertie Vidgen, Alex Harris, Dong Nguyen, Rebekah Tromble, Scott Hale, and Helen Margetts. 2019. Challenges and frontiers in abusive content detec- tion. In Proceedings of the Third Workshop on Abu- sive Language Online, pages 80–93, Florence, Italy, August. Association for Computational Linguistics. Michael Wiegand, Josef Ruppenhofer, and Thomas Kleinbauer. 2019. Detection of Abusive Lan- guage: the Problem of Biased Datasets. In Proceed- ings of the 2019 Conference of the North American Chapter of the Association for Computational Lin- guistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 602–608, Min- neapolis, Minnesota, June. Association for Compu- tational Linguistics. Marcos Zampieri, Shervin Malmasi, Preslav Nakov, Sara Rosenthal, Noura Farra, and Ritesh Kumar. 2019. Semeval-2019 task 6: Identifying and catego- rizing offensive language in social media (offense- val). In Proceedings of the 13th International Work- shop on Semantic Evaluation, pages 75–86, Min- neapolis, Minnesota, USA. Association for Compu- tational Linguistics. Marcos Zampieri, Preslav Nakov, Sara Rosenthal, Pepa Atanasova, Georgi Karadzhov, Hamdy Mubarak, Leon Derczynski, Zeses Pitenis, and Çağrı Çöltekin. 2020. SemEval-2020 Task 12: Multilingual Offen- sive Language Identification in Social Media (Of- fensEval 2020). In Proceedings of the 14th Inter- national Workshop on Semantic Evaluation. Associ- ation for Computational Linguistics.