<?xml version="1.0" encoding="UTF-8"?>
<TEI xml:space="preserve" xmlns="http://www.tei-c.org/ns/1.0" 
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" 
xsi:schemaLocation="http://www.tei-c.org/ns/1.0 https://raw.githubusercontent.com/kermitt2/grobid/master/grobid-home/schemas/xsd/Grobid.xsd"
 xmlns:xlink="http://www.w3.org/1999/xlink">
	<teiHeader xml:lang="en">
		<fileDesc>
			<titleStmt>
				<title level="a" type="main">The LaboCIC at HOMO-MEX 2024: Using BERT to Classify Hate-LGTB Speech</title>
			</titleStmt>
			<publicationStmt>
				<publisher/>
				<availability status="unknown"><licence/></availability>
			</publicationStmt>
			<sourceDesc>
				<biblStruct>
					<analytic>
						<author>
							<persName><forename type="first">Daniel</forename><forename type="middle">Yacob</forename><surname>Espinosa</surname></persName>
							<email>espinosagonzalezdaniel@gmail.com</email>
							<affiliation key="aff0">
								<orgName type="department" key="dep1">Instituto Politécnico Nacional (IPN)</orgName>
								<orgName type="department" key="dep2">Centro de Investigación en Computación (CIC)</orgName>
								<address>
									<settlement>Mexico City</settlement>
									<country key="MX">Mexico</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Grigori</forename><surname>Sidorov</surname></persName>
							<email>sidorov@cic.ipn.mx</email>
							<affiliation key="aff0">
								<orgName type="department" key="dep1">Instituto Politécnico Nacional (IPN)</orgName>
								<orgName type="department" key="dep2">Centro de Investigación en Computación (CIC)</orgName>
								<address>
									<settlement>Mexico City</settlement>
									<country key="MX">Mexico</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Eusebio</forename><surname>Ricárdez Vázquez</surname></persName>
							<affiliation key="aff0">
								<orgName type="department" key="dep1">Instituto Politécnico Nacional (IPN)</orgName>
								<orgName type="department" key="dep2">Centro de Investigación en Computación (CIC)</orgName>
								<address>
									<settlement>Mexico City</settlement>
									<country key="MX">Mexico</country>
								</address>
							</affiliation>
						</author>
						<title level="a" type="main">The LaboCIC at HOMO-MEX 2024: Using BERT to Classify Hate-LGTB Speech</title>
					</analytic>
					<monogr>
						<idno type="ISSN">1613-0073</idno>
					</monogr>
					<idno type="MD5">EDACF3821138735C6ACF13F77B226103</idno>
				</biblStruct>
			</sourceDesc>
		</fileDesc>
		<encodingDesc>
			<appInfo>
				<application version="0.7.2" ident="GROBID" when="2025-04-23T19:41+0000">
					<desc>GROBID - A machine learning software for extracting information from scholarly documents</desc>
					<ref target="https://github.com/kermitt2/grobid"/>
				</application>
			</appInfo>
		</encodingDesc>
		<profileDesc>
			<textClass>
				<keywords>
					<term>BERT</term>
					<term>tweets</term>
					<term>LGBT+</term>
					<term>hate speech</term>
					<term>hate LGBT+</term>
				</keywords>
			</textClass>
			<abstract>
<div xmlns="http://www.tei-c.org/ns/1.0"><p>Social media has emerged as a crucial space for communication, especially with the increase in its use during the pandemic. These platforms enable the exchange of information and connection among users, being particularly significant for communities such as the LGBT+ movement. However, cyberbullying towards the LGBT+ community on social media has severe consequences, including psychological harm, isolation, low self-esteem, and in extreme cases, physical violence and even deaths. Despite the policies and tools implemented by platforms to combat hate, their effectiveness varies, and the LGBT+ community remains highly exposed to these behaviors.</p><p>One of the current challenges in content moderation on social media is identifying satire and irony, complicating the classification of messages as hate content. In this context, we participated in the tasks proposed by Homo-Mex [1], focusing on Task 1 and Task 3. Task 1 centers on the classification of tweets with hate content directed at the LGBT+ community, while Task 3 involves binary classification of songs in Spanish. To solve these problems, we used BERT, achieving results of 92.37% F1-score for Task 1 and 89.14% F1-score for Task 3. This research work aims to improve artificial intelligence systems for the categorization of hate speech, contributing to the creation of safer digital spaces for the LGBT+ community.</p></div>
			</abstract>
		</profileDesc>
	</teiHeader>
	<text xml:lang="en">
		<body>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1.">Introduction</head><p>Social media has become a fundamental space for communication, and since the pandemic in recent times, the use of these digital spaces has been increasing. This is due to the exchange of information among users and their interactions. Many of these users find social media to be a safe place to connect with communities such as the LGBT+ movement. The term LGBT+ is a way to recognize and group people who identify with sexual orientations different from heterosexuality and cisgender. However, on many digital platforms, social biases still exist, which divide the population with this type of content. Cyberbullying towards the LGBT+ community on social media has serious consequences, both for the individuals who are directly attacked and for the community as a whole. Such acts can lead to psychological harm, feelings of isolation, low self-esteem, and in extreme cases, situations of physical violence and intimidation, with even some cases resulting in fatalities <ref type="bibr" target="#b1">[2]</ref>. Social media platforms have implemented various policies and tools to combat the spread of hate, such as content moderation, reporting systems, and the promotion of safe spaces, although their effectiveness and consistency can vary. These measures are implemented for all social media users, but due to the impact and large community of the LGBT+ movement, these behaviors are more exposed <ref type="bibr" target="#b2">[3]</ref>.</p><p>Research by Andrew Flores indicates that the LGBT+ community is more likely to be victims of hate crimes and discrimination. These crimes are primarily motivated by social prejudices, and when members of these communities seek help, very few countries and institutions offer specialized support for this type of aggression <ref type="bibr" target="#b3">[4]</ref>. It is noted that the highest incidence of these crimes occurs in close circles: schools, workplaces, or the homes of the affected individuals, and that the harm caused by these crimes is more severe and violent than other incidents.</p><p>There is an ongoing problem that is still under investigation: satire and irony. Many of the messages displayed on social media have this aspect, making it more difficult to classify them as hate content towards certain individuals or communities <ref type="bibr" target="#b4">[5]</ref>.</p><p>For this occasion, we decided to participate in the task created by Homo-Mex <ref type="bibr" target="#b0">[1]</ref>. Homo-Mex, for this year, decided to launch three research tasks for this year is IberLEF <ref type="bibr" target="#b5">[6]</ref>. For this research work, we chose to work on Task 1 and Task 3. All the tasks are related to hate speech towards the LGBT+ community; given the issues that such behaviors can cause, Homo-Mex is tasks aim to improve artificial intelligence systems for categorizing these topics.</p><p>In Task 1, the objective is to classify tweets with hate content directed at the LGBT+ community. For Task 2, within a set of tweets, a marker is placed to identify the specific type of community the hate is directed towards: lesbophobia, gayphobia, biphobia, transphobia, LGBT+phobia, or not LGBT+ related. For the final task, Task 3, the goal is to perform binary classification for a set of songs in Spanish. In this research work, we will focus only on Task 1 and Task 3.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.">Dataset</head><p>In Task 1, the objective is to classify tweets with hate content directed at the LGBT+ community, using a dataset composed of 7,000 tweets in Spanish for training and 220 tweets for testing. These tweets exhibit typical characteristics of social media texts, such as mentions of other users, hashtags, links, and emojis. To address this task, various preprocessing techniques were implemented to clean and organize the text. Subsequently, natural language processing (NLP) models, particularly BERT and its variations, were applied to identify patterns and perform the classification.</p><p>In Task 3, the training dataset consisted of 600 Spanish songs, and the testing dataset comprised 246 songs. Most of these songs are segmented according to their musical structure into parts such as intro, chorus, verses, bridge, refrains, and outro. We considered it important to implement a preprocessing layer, as these elements are not deemed significant enough to be used in the research. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.">Metodology</head><p>We conducted several experiments with tweets and found that in most cases, we recommend adding a preprocessing layer to the data. Since Tasks 1 and 3 have different characteristics, we implemented different preprocessing approaches for each dataset. For Task 1, we have tweets, which, as we know, often include informal language, abbreviations, and emojis. In contrast, Task 3 involves songs, typically filled with informal language and slang, which can be mixed with metaphors, ambiguities, or even vulgar expressions disguised with other texts. We will start with preprocessing layers before feeding the data into the models, ensuring that our entire research is related to BERT and some of its variations.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.1.">Pre-processing steps</head><p>As with any text-related task, we always recommend using a preprocessing layer before directly applying models. Preprocessing the data helps transform, clean, and prepare it for analysis or modeling. This improves data quality and makes it easier to use in algorithms. Additionally, preprocessing helps enhance model performance and facilitates the interpretation and analysis of the results. Without a preprocessing layer, the outcomes of an analysis or machine learning model can be inaccurate or misleading.</p><p>The following configuration was used exclusively for the tweets in Task 1:</p><p>Lowercase All tweets were converted to lowercase to standardize the texts. Links Links were replaced with the tag 'enlace'.</p><p>Hashtags Hashtags were modified and replaced with the tag 'hashtag'.</p><p>User Mentions User mentions were replaced with the tag 'mención de usuario'.</p><p>Emojis No changes were made to emojis as they integrate well with the model.</p><p>Other Symbols All symbols not recognized within the ASCII reference standard were removed.</p><p>Due to Task 3 involving songs throughout its dataset, the following configurations were made:</p><p>Lowercase All texts were converted to lowercase.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Musical Structures Marks of musical structures were removed.</head><p>Other Symbols All symbols not recognized within the ASCII reference standard were removed from the songs.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.">Experiments</head><p>In previous works, we have been involved in tweet classification, primarily focusing on bots <ref type="bibr" target="#b6">[7]</ref>. Therefore, we wanted to test our methodology to observe its behavior with a different task. In this case, we used a structure of word and character N-grams. The tests conducted with this structure were for both tasks.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Table 2</head><p>Results Task 1 of F1-score with N-grams Structure N-grams char word F1-Score 5-7-8-9 3-4-5 75.21</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Table 3</head><p>Results Task 3 of F1-score with N-grams Structure N-grams char word F1-Score 5-7-8-9 3-4-5 66.30</p><p>In Task 1, focused on classifying tweets with hate content, we achieved expected results, demonstrating the effectiveness of N-grams in this context, although there was still significant room for improvement. In Task 3, which involved the binary classification of songs in Spanish, N-grams also proved to be useful; however, the challenge was greater due to the different nature of the texts and the more creative and varied use of language in song lyrics. Thanks to the previous results, which were far from accurate, we decided to try another methodology that we have been using in recent years.</p><p>For PAN 2023, we conducted research on Crypto-influencers classification using BERT <ref type="bibr" target="#b7">[8]</ref>. In our experiments for PAN 2023, the standout models were BERT, RoBERTa <ref type="bibr" target="#b8">[9]</ref>, and BERTweet <ref type="bibr" target="#b9">[10]</ref>. Therefore, we decided to test which of these models we could use for this research. It is important to mention that for all models used in the experiments, we applied the following configuration: a batch size of 16, 9 training epochs, and a GPU usage limit of 15GB. We used PyTorch with pre-trained models on the GPU. These configurations were applied to all three BERT variants and evaluated using F1-score. For the next Task, in our case Task 3, we decided to conduct the same experiment with all three BERT variants. These were the results obtained for this task. For these experiments, we consider utilizing more GPU power if necessary when using RoBERTa, as this model resulted in significantly longer training times. Surely, with additional computational resources, we could potentially leverage its robustness more effectively.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.">Conclusions</head><p>We believe that these types of problems are currently underexplored in research, despite their great importance. Additionally, something we greatly appreciated about this research was the datasets, as they were all in Spanish, whereas most models are trained and built for the English language. Often, the problem lies in the lack of Spanish data to train these models.</p><p>For Task 1, we were surprised by the results of BERTweet. Although this model typically performs well, given its primary use with tweets, an important aspect is that it was trained on English tweets. This led us to reflect on the importance of having robust models available and trained in multiple languages. RoBERTa also showed promising results, but its performance in Spanish was noticeably inferior, further emphasizing the need for models specifically trained for different languages.</p><p>For Task 3, the most notable aspect is that we used the same models as for Task 1. Despite being different data sets, they showed similar results. We suppose that this is why BERtweet performance does not show more outstanding results, mainly due to its training method and the fact that the data were in English, lacking the necessary approach to achieve good classification results. This data offers us valuable lessons on the importance of training data in the effectiveness of artificial intelligence models. The variability in the data and the adaptation of the model to different languages and types of content are crucial factors that affect performance.</p><p>These experiments provided a unique opportunity to evaluate the flexibility and adaptability of our models in different linguistic and thematic contexts. It also allows us to better understand the limitations of current models and highlights the need to develop more robust training techniques that consider the particularities of language, the context of the data, and the cultural context.</p><p>These findings led us to pose several additional research questions. How can we improve language models to perform equally well in Spanish as they do in English? What strategies can be adopted to generate and curate more Spanish datasets to enable more effective training of these models?</p><p>We would like to contribute to the creation of a model similar to BERTweet, trained exclusively with large-scale Spanish tweets. This would not only improve the classification of bots and the detection of</p></div><figure xmlns="http://www.tei-c.org/ns/1.0" type="table" xml:id="tab_0"><head>Table 1</head><label>1</label><figDesc>Dataset of HomoMex 2024<ref type="bibr" target="#b0">[1]</ref> </figDesc><table><row><cell>Dataset</cell><cell>Train</cell><cell>Test</cell></row><row><cell cols="3">Task 1 7000 tweets 220 tweets</cell></row><row><cell>Task 3</cell><cell>600 songs</cell><cell>246 songs</cell></row></table></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" type="table" xml:id="tab_1"><head>Table 4</head><label>4</label><figDesc></figDesc><table><row><cell>Results of Task 1 F1-Score and BERT Variations</cell><cell></cell></row><row><cell>Model</cell><cell>F1-Score</cell></row><row><cell>RoBerta</cell><cell>87.27</cell></row><row><cell>BERTweet</cell><cell>88.41</cell></row><row><cell>BERT</cell><cell>92.37</cell></row></table></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" type="table" xml:id="tab_2"><head>Table 5</head><label>5</label><figDesc></figDesc><table><row><cell>Results of Task 3 F1-Score and BERT Variations</cell><cell></cell></row><row><cell>Model</cell><cell>F1-Score</cell></row><row><cell>RoBerta</cell><cell>83.99</cell></row><row><cell>BERTweet</cell><cell>74.02</cell></row><row><cell>BERT</cell><cell>89.14</cell></row></table></figure>
		</body>
		<back>

			<div type="acknowledgement">
<div xmlns="http://www.tei-c.org/ns/1.0"><p>fake news on Spanish-speaking social networks but could also be applied to other natural language processing tasks, such as hate speech in different internet communities.</p></div>
			</div>

			<div type="references">

				<listBibl>

<biblStruct xml:id="b0">
	<analytic>
		<title level="a" type="main">Overview of homo-mex at iberlef 2024: Hate speech detection towards the mexican spanish speaking lgbt+ population</title>
		<author>
			<persName><forename type="first">H</forename><surname>Gómez-Adorno</surname></persName>
		</author>
		<author>
			<persName><forename type="first">G</forename><surname>Bel-Enguix</surname></persName>
		</author>
		<author>
			<persName><forename type="first">H</forename><surname>Calvo</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Vásquez</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><forename type="middle">T</forename><surname>Andersen</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Ojeda-Trueba</surname></persName>
		</author>
		<author>
			<persName><forename type="first">T</forename><surname>Alcáhntara</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Soto</surname></persName>
		</author>
		<author>
			<persName><forename type="first">C</forename><surname>Macias</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Natural Language Processing</title>
		<imprint>
			<biblScope unit="volume">73</biblScope>
			<date type="published" when="2024">2024</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b1">
	<monogr>
		<author>
			<persName><forename type="first">Z</forename><surname>Akmeşe</surname></persName>
		</author>
		<author>
			<persName><forename type="first">K</forename><surname>Deniz</surname></persName>
		</author>
		<title level="m">Hate speech in social media: Lgbti persons</title>
				<imprint>
			<date type="published" when="2017">2017</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b2">
	<analytic>
		<title level="a" type="main">Hate crimes against trans people: Assessing emotions, behaviors, and attitudes toward criminal justice agencies</title>
		<author>
			<persName><forename type="first">M</forename><forename type="middle">A</forename><surname>Walters</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Paterson</surname></persName>
		</author>
		<author>
			<persName><forename type="first">R</forename><surname>Brown</surname></persName>
		</author>
		<author>
			<persName><forename type="first">L</forename><surname>Mcdonnell</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">J Interpers Violence</title>
		<imprint>
			<biblScope unit="volume">35</biblScope>
			<biblScope unit="page" from="4583" to="4613" />
			<date type="published" when="2017">2017</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b3">
	<analytic>
		<title level="a" type="main">Hate crimes against lgbt people: National crime victimization survey, 2017-2019</title>
		<author>
			<persName><forename type="first">A</forename><surname>Flores</surname></persName>
		</author>
		<author>
			<persName><forename type="first">R</forename><surname>Stotzer</surname></persName>
		</author>
		<author>
			<persName><forename type="first">I</forename><surname>Meyer</surname></persName>
		</author>
		<author>
			<persName><forename type="first">L</forename><surname>Langton</surname></persName>
		</author>
		<idno type="DOI">10.1371/journal.pone.0279363</idno>
	</analytic>
	<monogr>
		<title level="j">PLOS ONE</title>
		<imprint>
			<biblScope unit="volume">17</biblScope>
			<biblScope unit="page">e0279363</biblScope>
			<date type="published" when="2022">2022</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b4">
	<analytic>
		<title level="a" type="main">Bert-based ironic authors profiling</title>
		<author>
			<persName><forename type="first">W</forename><surname>Yu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">B</forename><forename type="middle">T</forename><surname>Boenninghoff</surname></persName>
		</author>
		<author>
			<persName><forename type="first">D</forename><surname>Kolossa</surname></persName>
		</author>
		<ptr target="https://api.semanticscholar.org/CorpusID:251471104" />
	</analytic>
	<monogr>
		<title level="m">Conference and Labs of the Evaluation Forum</title>
				<imprint>
			<date type="published" when="2022">2022</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b5">
	<analytic>
		<title level="a" type="main">Overview of IberLEF 2024: Natural Language Processing Challenges for Spanish and other Iberian Languages</title>
		<author>
			<persName><forename type="first">L</forename><surname>Chiruzzo</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><forename type="middle">M</forename><surname>Jiménez-Zafra</surname></persName>
		</author>
		<author>
			<persName><forename type="first">F</forename><surname>Rangel</surname></persName>
		</author>
		<ptr target=".org" />
	</analytic>
	<monogr>
		<title level="m">Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2024), co-located with the 40th Conference of the Spanish Society for Natural Language Processing</title>
				<meeting>the Iberian Languages Evaluation Forum (IberLEF 2024), co-located with the 40th Conference of the Spanish Society for Natural Language Processing<address><addrLine>SEPLN</addrLine></address></meeting>
		<imprint>
			<publisher>CEUR-WS</publisher>
			<date type="published" when="2024">2024. 2024</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b6">
	<analytic>
		<title level="a" type="main">Bots and Gender Profiling using Character Bigrams</title>
		<author>
			<persName><forename type="first">D</forename><surname>Espinosa</surname></persName>
		</author>
		<author>
			<persName><forename type="first">H</forename><surname>Gómez-Adorno</surname></persName>
		</author>
		<author>
			<persName><forename type="first">G</forename><surname>Sidorov</surname></persName>
		</author>
		<ptr target="http://ceur-ws.org/Vol-2380/" />
	</analytic>
	<monogr>
		<title level="m">CLEF 2019 Labs and Workshops</title>
		<title level="s">Notebook Papers</title>
		<editor>
			<persName><forename type="first">L</forename><surname>Cappellato</surname></persName>
		</editor>
		<editor>
			<persName><forename type="first">N</forename><surname>Ferro</surname></persName>
		</editor>
		<editor>
			<persName><forename type="first">D</forename><surname>Losada</surname></persName>
		</editor>
		<editor>
			<persName><forename type="first">H</forename><surname>Müller</surname></persName>
		</editor>
		<imprint>
			<date type="published" when="2019">2019</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b7">
	<analytic>
		<title level="a" type="main">Using BERT to profiling cryptocurrency influencers</title>
		<author>
			<persName><forename type="first">D</forename><forename type="middle">Y</forename><surname>Espinosa</surname></persName>
		</author>
		<author>
			<persName><forename type="first">G</forename><surname>Sidorov</surname></persName>
		</author>
		<idno>WS.org</idno>
		<ptr target="https://ceur-ws.org/Vol-3497/paper-207.pdf" />
	</analytic>
	<monogr>
		<title level="m">Working Notes of the Conference and Labs of the Evaluation Forum (CLEF 2023)</title>
		<title level="s">CEUR Workshop Proceedings</title>
		<editor>
			<persName><forename type="first">M</forename><surname>Aliannejadi</surname></persName>
		</editor>
		<editor>
			<persName><forename type="first">G</forename><surname>Faggioli</surname></persName>
		</editor>
		<editor>
			<persName><forename type="first">N</forename><surname>Ferro</surname></persName>
		</editor>
		<editor>
			<persName><forename type="first">M</forename><surname>Vlachos</surname></persName>
		</editor>
		<meeting><address><addrLine>Thessaloniki, Greece</addrLine></address></meeting>
		<imprint>
			<date type="published" when="2023">September 18th to 21st, 2023. 2023</date>
			<biblScope unit="volume">3497</biblScope>
			<biblScope unit="page" from="2568" to="2573" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b8">
	<monogr>
		<title level="m" type="main">Roberta: A robustly optimized bert pretraining approach</title>
		<author>
			<persName><forename type="first">Y</forename><surname>Liu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Ott</surname></persName>
		</author>
		<author>
			<persName><forename type="first">N</forename><surname>Goyal</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Du</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Joshi</surname></persName>
		</author>
		<author>
			<persName><forename type="first">D</forename><surname>Chen</surname></persName>
		</author>
		<author>
			<persName><forename type="first">O</forename><surname>Levy</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Lewis</surname></persName>
		</author>
		<author>
			<persName><forename type="first">L</forename><surname>Zettlemoyer</surname></persName>
		</author>
		<author>
			<persName><forename type="first">V</forename><surname>Stoyanov</surname></persName>
		</author>
		<idno type="arXiv">arXiv:1907.11692</idno>
		<imprint>
			<date type="published" when="2019">2019</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b9">
	<analytic>
		<title level="a" type="main">BERTweet: A pre-trained language model for English tweets</title>
		<author>
			<persName><forename type="first">D</forename><forename type="middle">Q</forename><surname>Nguyen</surname></persName>
		</author>
		<author>
			<persName><forename type="first">T</forename><surname>Vu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><forename type="middle">Tuan</forename><surname>Nguyen</surname></persName>
		</author>
		<idno type="DOI">10.18653/v1/2020.emnlp-demos.2</idno>
		<ptr target="https://aclanthology.org/2020.emnlp-demos.2.doi:10.18653/v1/2020.emnlp-demos.2" />
	</analytic>
	<monogr>
		<title level="m">Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, Association for Computational Linguistics</title>
				<meeting>the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, Association for Computational Linguistics</meeting>
		<imprint>
			<date type="published" when="2020">2020</date>
			<biblScope unit="page" from="9" to="14" />
		</imprint>
	</monogr>
</biblStruct>

				</listBibl>
			</div>
		</back>
	</text>
</TEI>
