<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>October</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Analysing gender-based violence against Colombian public figures on Twitter</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Juan Sebastian Chaparro-Saenz</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ixent Galpin</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Facultad de Ciencias Naturales e Ingeniería, Universidad de Bogotá Jorge Tadeo Lozano</institution>
          ,
          <addr-line>Bogotá</addr-line>
          ,
          <country country="CO">Colombia</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <volume>2</volume>
      <fpage>8</fpage>
      <lpage>30</lpage>
      <abstract>
        <p>This work proposes the development of a methodology that standardises the extraction, processing and analysis of natural language data for the study of gender-based violence evidenced on the Twitter social network. We develop a tool that may be exploited by diferent organisations, foundations, corporations, associations or state institutions that promote, exercise and disseminate human rights in Colombia and elsewhere. In this work, we take as a case study ten prominent female public figures in Colombia in the artistic, political and journalistic spheres. We extract a total of 39,629 tweet responses during a turbulent national strike amid the COVID-19 pandemic, and carry out topic identification and sentiment analysis. While we observe diferences between the diferent roles based on natural language processing with diferent libraries, the are notable negative terms in the topics identified which are of concern as they may incite gender-based violence. It is expected that this proposed tool will benefit the decision-making of these institutions to issue early warnings, together with the exercise of the protection, prevention and defence of women's rights.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Sentiment Analysis</kwd>
        <kwd>Topic Identification</kwd>
        <kwd>Text Mining</kwd>
        <kwd>Gender-based violence</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Violence against women has been highlighted as a problem that generates an impact of
significant importance to society [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Such violence may take diferent forms, including physical,
sexual, psychological, economic, verbal or written. These diferent types of violence can be
exercised by diferent actors, including partners, colleagues, fellow students and even
adversaries in the political field [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. In Colombia, according to the Defensoría del Pueblo (Human
Rights Ombudsman), there has been a increase in gender-based violence in the country since the
start of the COVID-19 pandemic [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. This includes diferent types of violence such as physical
violence (18%), sexual violence (6%), psychological violence (42%), violence against assets (6%)
and economic violence within a home (27%) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        Currently, society is constantly growing and evolving as part of the the digital era [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. This
in its wake allows diferent types of digital content to be created, giving way to a world in
which it is possible to constantly interact, becoming routine for human beings. According to
ifgures approximately 4.5 billion people currently use the Internet and 3.8 billion users are
registered in a social network such as Facebook, Twitter, YouTube, WhatsApp, among others [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
As such, this work focuses on the exploration of misogynistic postings on the Twitter social
network. This is deemed to be of high importance as such negative tweets in turn trigger
demonstrated psychological violence in Colombia, through intimidating comments, harassment,
threats, contempt, mockery, humiliation, among others [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        The microblogging Twitter platform is often considered to be a barometer of society [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ],
given that postings (or tweets) can be publically posted in real time. It enables people to express
their opinions through publications maintaining freedoom of expression on any topic that is
being debated at a moment in time. Twitter has become the de facto medium where opinions
are expressed on diferent issues, especially in the political sphere, which usually generate
contrasting and often conflicting points of view. On many occasions, these points of view reflect
in gender violence [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. In Colombia, Twitter currently has 3.2 million users, which accounts for
approximately 7.8% of the population. While there is a significant digital divide in Colombia
between rural and urban areas [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], there is also a gender-based one [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. In the case of Twitter,
this is also visible: 62.9% of users are male, compared to 37.1% female [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <p>
        In this work, we select ten Colombian public figures in the political, artistic, or journalistic
spheres. We collect and analyse tweets that correspond to responses to the ten public figures
under study, based on the controversies that may arise from their opinions. We identify the
predominant topics in these responses, and apply automated techniques to gauge sentiment
(negative or positive) and the degree of subjectivity or objectivity. We apply the well-established
CRISP-DM methodology [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] to ensure a robust analysis. This paper is structured as follows.
Section 2 presents related work. Subsequently, we broadly follow the first five steps of the
CRISP-DM methodology: Section 3 presents steps 1 and 2, business and data understanding. In
Section 4, we describe the data processing carried out prior to analysis (step 3 of CRISP-DM).
The next section describes the models applied. In Section 6, we present the results obtained
through the evaluation of the models. Finally, we present a discussion in Section 7, and Section 8
concludes.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>At the time of writing, there is considerable research based on text mining applied to social
problems such as violence, health, poverty, among others, that in turn supports decision-making.</p>
      <p>
        For example, Cremades et al. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] proposes the development of a system based on artificial
intelligence capable of predicting a potential suicide. Such a system would be based on the
application of a methodology that involves the collection, compilation and selection of text
for processing. Saura et al. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] analyse text from the #BlackFriday hashtag on Twitter. The
authors conclude that companies do not use the social network as a marketing strategy since
the study of this publication is based on the selection of opinions with a criterion related to
exclusive ofers. In summary, the study identifies the neutral, positive and negative sentiment
on the consolidation of writings on a selection criterion [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
      <p>
        The proposed methodology is integrated by stages for the final result. For this reason, it
is necessary to employ data engineering approaches for the extraction and consolidation of
the sample under study. Barriga et al. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] employ technological elements for the extraction,
consolidation and analysis of texts from the social network Twitter, such as the platform’s own
API that allows it to be integrated into the development of software under the Java language.
Thus, it enables the use of methods for the extraction of tweets. Once the previous process occurs,
there is a step involving warehousing of the information in two data models; relational and
non-relational. A non-relational database engine, MongoDb, is used to contain the information
of the extracted account profile and its corresponding timelines, treated as JSON files. Finally,
the author aims to develop a web tool for the extraction and storage of data from the social
network Twitter, which allow interoperability with external tools for data analysis [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
      </p>
      <p>
        Silva et al. [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] present an investigation directly related to social problems based on gender
violence. In this case, technology in the location of some type of violence is applied, as it
could be evidenced in verbal or written expression. Thus, a model for detecting aggressiveness
towards women in opinions published on social networks was created through the application
of machine learning techniques. It carries out the choice of a Twitter corpus, by means of
extraction with the platform’s API, performs a data cleaning process and makes use of Microsoft
Excel for punctuation, grammar and spelling cleaning. Characteristics of the tweet are selected
under a process developed in Java for further evaluation. Classification and evaluation of
texts is carried out using a collection of machine learning algorithms with the Weka library.
Finally, the result is the classification of a tweet as “aggressive” or “slightly aggressive” given
the performance of each model evaluated.
      </p>
      <p>
        Various studies analyse the behaviour of sentiments using Twitter hashtags. Evovli et al. [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]
investigate, through the analysis of tweets, hatred which was directed towards Muslims
following the United Kingdom Brexit referendum in 2016. A qualitative analysis is carried out
of tweets with the hashtags #IslamIsTheProblem and #Muslimterrorists. In this study,
variables that afect the study such as trolls and bots are taken into account. Evovli proposes
that Islamophobia may be mitigated by considering the connections in the social network
graphs [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Business and Data Understanding</title>
      <p>
        To understand the problem, it is appropriate to understand the type of abusive and specific
language towards women that will lead to the development of the tool in this work. This type
of violent language directed at women is known as misogyny, defined as the hatred or prejudice
towards women as a result of a belief that they are the weaker sex. Pamungkas et al. [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]
detected the presence of misogyny through a series of cross-lingual classification experiments.
      </p>
      <p>The empowerment of women as protagonists in the diferent roles of society has been
noteworthy in recent years. The use of social networks has been a means of promoting ideals,
relating and acquiring visibility. For this reason, we see the need to standardise a methodology
for the analysis of gender violence for a specific group of women. We select a group who not
only has political participation in the country, but are also an interdisciplinary group in order
to address and characterise through the study the diferent types of language expressed in the
opinions expressed about the opinions of these women.</p>
      <p>For the proposed objetive, it is necessary to consolidate a representative sample of data
according to the criteria that have been proposed. Thus, as a first step, a data engineering
Name
Claudia Lopez
Martha Lucia Ramirez
Margarita Rosa de Francisco
Vicky Davila
Maria Fernanda Carrascal
Angelica Lozano
Paloma Valencia
Claudia Gurissati
Adriana Lucia
Angela Robledo</p>
    </sec>
    <sec id="sec-4">
      <title>4. Data Preparation</title>
      <p>Taking into account the data collected from the profiles listed in Table 1, the development
of a tool is carried out using Python1, a scripting language commonly used for data science
projects. Python as a language has become a well-established tool for data analysis and the main
advantage is that it can be used without licensing costs. In other words, it is an open and free
technology compared to proprietary technologies. As shown in Figure 1, the tool comprises four
phases that govern the design and reflect critical aspects in accordance with the methodology
adopted, viz.</p>
      <p>1. Tweet extraction
2. Data cleaning
3. NoSQL database storage on Firestore Cloud
4. Text preprocessing.</p>
      <sec id="sec-4-1">
        <title>4.1. Tweet extraction</title>
        <p>In this phase, version 2 of the Twitter application platform programming interface is adopted2
to search for tweets with specific criteria. It was configured to find all that tweets that reply to
tweets created by the profiles listed in Table 1.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Data cleaning</title>
        <p>Once the information was extracted, a treatment was carried out to eliminate repeated tweets,
tweets written by the account owner and false tweets, in order to clean and build a valid data
set for further study.</p>
        <p>2https://developer.twitter.com/en/docs/twitter-api/early-access</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. NoSQL database storage on Firestore Cloud</title>
        <p>
          Based on the information processed and collected, it is necessary for this tool to make use of
catalog services that allow the set of semistructured data to be managed in an agile and fast way.
In such a way, it is sought that this technological service complies with some characteristics
that interoperate within the proposed model that makes use of a development based on Python
language and consolidates a unified and solid base for consulting a large number of data. For
this reason, the Google Cloud Firestore3 database was chosen, which is a document-oriented
NoSQL engine. Since our data set is treated as a JSON file in this way, Firestore allowed us to
store each data set in diferent collections that relate a set of tweets per profile. Each document
contains a set of key-value pairs that uses few resources and contains fields with assigned
values [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ].
        </p>
      </sec>
      <sec id="sec-4-4">
        <title>4.4. Text preprocessing</title>
        <p>
          By virtue of the compression on the objective of this study, it is necessary to become familiar with
the initial data collection in storage in the NoSQL service, verifying the quality and quantity of
these. In this sense, and in accordance with the flow of the CRISP-DM methodology, it is essential
to understand the problem that needs to be solved, which is why the diferent techniques for
data pre-processing are executed. After the discovery and preparation stage, resulting elements
derived from threads are obtained, such as data cleaning, tokenization, stopword removal,
stemming and lemmatization, closely related for the topic and classification model [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]. For the
application of these techniques the NLTK4 and Spacy5 libraries were employed, which provide
the possibility of treating the process in Spanish. Thus, the text preprocessing proposed for the
present study consists of:
• Data cleaning: In this step the elimination of characters, numbers, punctuations was
carried out.
• Tokenization: In this step we convert sentences into a word list.
• Stopword Removal: A list of words that do not provide correct information to the model
is constructed and therefore the removal of these words is carried out.
• Lemmatisation: This technique is used to reduce the dimensionality of a word, that is,
to take the verb to its infinitive form. In addition, sufixes that derive in quantities of
something are removed.
• Stemming: As with the lemmatisation step that seeks to reduce words, this technique is
done by applying structure rules on a set of letters joined in a word.
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Modeling</title>
      <p>Within the framework of the process proposed for the scope of this modeling phase, we explored
the topics and sentiments according to the corpus obtained from each Twitter profile in this
study.</p>
      <p>3https://firebase .google.com/docs
4https://www.nltk.org/
5https://spacy.io/</p>
      <sec id="sec-5-1">
        <title>5.1. Topic Analysis using Latent Dirichlet Allocation</title>
        <p>
          To understand the structure of the adopted model that determines the topics, Latent
Dirichlet Allocation (LDA) was used, which consists of a three-level hierarchical Bayesian model.
According to Blei et al. [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ], “Each topic is, in turn, modeled as an infinite mixture over an
underlying set of topic probabilities. In the context of text modeling, the topic probabilities
provide an explicit representation of a document”. We use the LDA Gensin Models library6
for implementation. The corpus is partitioned for each profile, allowing a visualisation with
the most relevant topics to be generated. Once we have treated the data in which we seek to
reduce the dimension of the vocabulary for the optimisation of the model, subsequently it can
be identified that some words do not have an adequate meaning in isolation. In comparison, it
is more meaningful if they are considered with adjacent words. Thus, it is necessary to consider
n-grams [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ]. As such, the data set is explored and the frequency with which the words appear
together is identified, classifying them in bi-grams or tri-grams. Once we obtain the processed
and cleaned tweets, we can build a dictionary to continue the training of the LDA model. In this
step, the corpus vocabulary is built in which all the unique words in the data set are assigned a
unique ID.
        </p>
        <p>The subsequent model training step includes the use of the dictionary corresponding to the
scenario, to refer a class to each of the topics. During development, it will be necessary to iterate
multiple times to return the topics resulting from the most probable words. In order to find
the optimal parameters for the LDA model defined in the development of the tool, they were
initially defined in Table 2. With these parameters it is possible to contrast the most relevant
topics and the slightly more frequent words.</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Sentiment analysis</title>
        <p>We subsequently carry out a sentiment analysis over each tweet. We use the Vader7 and
Textblob8 libraries, which enable the measurement of polarity (i.e., how negative or positive
a tweet is), and subjectivity (i.e., whether a tweet corresponds to fact or opinion) associated
with a tweet. In accordance with the capabilities of these libraries, they do not support the
Spanish language for sentiment analysis. For this reason, a process for the translation of text
6https://radimrehurek.com/gensim/models/ldamodel.html
7https://pypi.org/project/vaderSentiment/
8https://textblob.readthedocs.io/en/dev/
was developed that consists of the installation of the “Translate Text” extension9, which allows,
without any limitation, to translate documents stored on Google Firestore Cloud into the English
language.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Results</title>
      <p>For the results chain, an orientation derived from mechanisms such as the word cloud, LDA
gensim model, Vader sentiment library and Textblob library were considered to understand
and identify the data that would allow discussions and insights on the violence evidenced
on Twitter. The results obtained correspond to an estimate that would be obtained in a real
scenario. The error of this estimate can be given by the concentration of elements that do not
have an adequate meaning, the precision of the libraries, and the quality of the processed data.
However, the method adopted constitutes an approximation of a real and a valid scenario to
discuss violence against women.</p>
      <sec id="sec-6-1">
        <title>6.1. Topic Analysis using Latent Dirichlet Allocation</title>
        <p>(a) Claudia Lopez</p>
        <p>(b) Martha Lucia Ramirez
(c) Paloma Valencia
(d) Angela Robledo</p>
        <p>Name</p>
        <p>Claudia Lopez</p>
        <p>Martha Lucia Ramirez
Margarita Rosa de Francisco</p>
        <p>Vicky Davila
Maria Fernanda Carrascal</p>
        <p>Angelica Lozano
Paloma Valencia
Claudia Gurissati</p>
        <p>Adriana Lucia
Angela Robledo</p>
        <p>Most salient negative terms</p>
        <p>vandalo,ineptar,mierda
narco, muñeca_mafia, narcotráfico, narca_lucir, narca, muñeca
izquierda_destruir, definitivamente_tocar_izquierda_destruir
rata, bobo, uribista, doble_moral, asesinar, titere</p>
        <p>robar
angélico, muñeca jugadita, mierda
criminal, victima, matarife, terrorista, hijueputa, delincuente</p>
        <p>cobarde
vandalo, maldad
terrorista, viejo, delincuente, vandalo, criminal</p>
        <p>Figure 2 presents a graphical representation of the vocabulary as a visual resource of the most
significant words within the data set obtained from the opinions whose content is integrated by
the single user of the platform. Notably we see these cases with words in common between
them such as “narco” and “vieja [a derogatory term used in Colombia to refer to a woman]”, as
well as outstanding words such as “mierda [shit]”, “bruja [witch]”, “inepta [inept]”, all of them
with a negative and violent connotation. These terms clearly suggest the presence of tweets
with derogatory, discriminatory and stigmatising content. Furthermore, this finding reflects
a social and cultural problem derived from history in which men have subjected women in
multiple ways, developing dominance and stigmatising them as the weaker sex.</p>
        <p>Figure 3 presents a visualisation with the identification of topics observed, with an adjustment
of the  parameter equal to 1. Following trial and error, four topics are found to yield the most
meaningful characterisation of the data for each profile. This parameter will determine the
weight given to the probability of a word on the topic. Now, as long as this adjustment is closer
to 1, it will return a set of terms characterised by their probability in the topic. Each circle in
the visualisation represents a topic, and the larger it is, the more dominant it will be compared
to other topics. For the specific analysis of Martha Lucia Ramirez and Paloma Valencia, the
most negative words are notably evident in comparison with the others, within the set of most
outstanding terms for each topic, shown on the right hand side of each visualisation. In the case
of the vice-president Martha Lucia Ramirez, we see that the term “narco” is the most prevalent
within the major topic, and in turn is related to the second topic. Within the topics words are
observed that suggest that the tweets contain grotesque, stigmatising and derogatory messages.</p>
        <p>Table 3 presents the most salient negative terms for the topics identified in each profile,
including n-grams with a political and slightly negative context. This finding allows us to show
the presence of violent, derogatory and stigmatizing comments in this case in two profiles with
a political roles. When observing the results for the singer Adriana Lucia, the absence of violent
terms is observed, in comparison with the political roles where we observe a large index of
elements that make up misogyny, that is, it is considered that in this area there is a violation of
rights.
(a) Overall Term Frequency
(b) Topic 1
(c) Topic 4
Figure 3: Exploring Topics Using LDA Gensim for Martha Lucia Ramirez</p>
        <p>93</p>
      </sec>
      <sec id="sec-6-2">
        <title>6.2. Sentiment analysis</title>
        <p>(a) Claudia Lopez
or personal judgement. In all cases, tweets tend to contain messages with factual information
rather than subjective opinions.</p>
        <p>Figure 6 presents shows the correlation between polarity and subjectivity. We observe that
for Martha Lucia Ramirez, there is a no significant correlation between polarity and subjectivity.
However, in the case of Paloma Valencia, there is a slight negative correlation between polarity
and subjectivity, i.e., the more positive a tweet is, the more factual it tends to be.
(a) Martha Lucia Ramirez
(b) Paloma Valencia</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>7. Discussion</title>
      <p>
        With the findings obtained, the presence of violent elements directed at female public figures in
Colombia was notable. Taking into account the profiles of the public figures, it was evident that
in the political sphere there is evidence of controversy by opinion. In summary, the texts aimed
at the political profile tend to present a negative sentiment compared to the artistic profile. As
mentioned by the Colombian Human Rights Ombudsman’s Ofice in its 2021 annual bulletin,
recently situations of psychological violence have worsened and concentrated in the Colombian
population [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. For this reason, it is considered that in the digital scene there is the presence
of shocks of political thought leading to the formation of comments with a negative meaning
directed at women of a public nature. This situation is a growing social problem, which gives
way to the participation of technology as a basis for analysing the information generated on
social networks such as Twitter, Facebook or Instagram.
      </p>
      <p>It is evident that the coherence of a dialogue with good values must be satisfied with the
absence of violence and the presence of respect. However, in this study the presence of negative
elements towards women was observed, which is interpreted as a violation of human rights,
autonomy and freedom of expression. With the existence of the Human Rights Ombudsman’s
Ofice as a control body in Colombia, an entity that currently has the function of issuing early
warnings of violence in the field of the armed conflict and also in charge of preventing the
violation of human rights. For this reason, the proposal of this article as a methodology for
the monitoring of violence in the digital scenario, a potentially valuable source of information
can be consolidated as support for the emission of early warnings of gender violence in social
networks, with established criteria such as the frequency, intensity and quantity of negative
opinions. After the deployment of this tool, it is hoped that it will be possible to contribute to
the mitigation of the the impact of acts of psychological violence for social reconstruction in
the Colombian population.</p>
    </sec>
    <sec id="sec-8">
      <title>8. Conclusions</title>
      <p>In this work, we identify elements of psychological violence against women in Colombia
based on Twitter responses to public figures. The result of this study allows progress in social
construction by providing tools to control institutions such as the Human Rights Ombudsman
in Colombia for the prevention of potential human rights violations. We present a methodology
to standardise the tweet extraction process, consolidate and analyse the information from the
social network. This constitutes an instrument for the real dimensioning of psychological
violence. The analysis carried out in said methodology includes the subjective component of a
text, used for interpretation in relation to the positive or negative polarity found. Thus, with the
adoption of the proposed methodology, the state control bodies and human rights organisations
in Colombia and elsewhere can agree on a unified criterion of comparison, a fact that will
contribute to the homogenisation of the protection, promulgation and prevention of human
rights.</p>
      <p>It is hoped that after applying the proposed methods for the analysis and having studied the
data processed by the architecture proposed, it will be possible to work and involve experts in
the field of women’s rights and gender, experts in linguistics to optimise the tool and generate
new strategies that promote human rights. Further work includes incorporating data from other
other data sources such as Instagram, Facebook or WhatsApp, bearing in mind that to optimise
data analysis it is appropriate to integrate data from diverse sources, since the problem studied
is evident on the other social networks as well. We also consider that it would be fruitful to
include a male control group, so as to eliminate references of non-validated sexist discrimination
based on gender-neutral derogatory comments. Furthermore, we envisage that it would be
appropriate to incorporate linguistic techniques that contribute to the more precise detection of
misogyny in these digital settings, in addition to correlating other properties resulting from the
integration with specialised areas in human rights and linguistics.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>U.</given-names>
            <surname>Women</surname>
          </string-name>
          ,
          <source>The world for women and girls annual report 2019-2020</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M. I. L.</given-names>
            <surname>Vélez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. M. E.</given-names>
            <surname>Jaramillo</surname>
          </string-name>
          ,
          <article-title>Derechos laborales y de la seguridad social para las mujeres en colombia en cumplimiento</article-title>
          de la ley 1257 de 2008, Revista de Derecho (
          <year>2015</year>
          )
          <fpage>269</fpage>
          -
          <lpage>296</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D. D.</given-names>
            <surname>Pueblo</surname>
          </string-name>
          ,
          <article-title>Situación de las mujeres y personas con orientación sexual e identidad de género diversas, refugiadas y migrantes en colombia, Women's rights (</article-title>
          <year>2021</year>
          )
          <article-title>10</article-title>
          . URL: https://www.defensoria.gov.co/public/pdf/Boletin_Situacion_Mujer_
          <year>2020</year>
          .pdf.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Oussous</surname>
          </string-name>
          , F.-
          <string-name>
            <given-names>Z.</given-names>
            <surname>Benjelloun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Lahcen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Belfkih</surname>
          </string-name>
          ,
          <article-title>Big data technologies: A survey</article-title>
          ,
          <source>Journal of King Saud University-Computer and Information Sciences</source>
          <volume>30</volume>
          (
          <year>2018</year>
          )
          <fpage>431</fpage>
          -
          <lpage>448</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Kemp</surname>
          </string-name>
          ,
          <year>2020</year>
          ,
          <year>Digital 2020</year>
          :
          <article-title>3.8 billion people use social media</article-title>
          , URL: https:// wearesocial.com/blog/2020/01/digital-2020-3
          <article-title>-8-billion-people-use-social-media.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>E.</given-names>
            <surname>Van der Klashorst</surname>
          </string-name>
          , S. Safarikova,
          <article-title>Twitter as barometer of public opinion on the female athlete: The case of caster semenya</article-title>
          ,
          <source>African Journal for Physical Activity and Health Sciences (AJPHES) 24</source>
          (
          <year>2018</year>
          )
          <fpage>649</fpage>
          -
          <lpage>658</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Khatua</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Cambria</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Khatua</surname>
          </string-name>
          ,
          <article-title>Sounds of silence breakers: Exploring sexual violence on twitter</article-title>
          ,
          <source>in: 2018 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM)</source>
          , IEEE,
          <year>2018</year>
          , pp.
          <fpage>397</fpage>
          -
          <lpage>400</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S. D. M.</given-names>
            <surname>Dussan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Leon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Garcia-Bedoya</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Galpin</surname>
          </string-name>
          ,
          <article-title>Exploring the colombian digital divide using moodle logs through supervised learning</article-title>
          ,
          <source>Interactive Technology and Smart Education</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>D. F.</given-names>
            <surname>Martinez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. N.</given-names>
            <surname>Pacheco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. F.</given-names>
            <surname>Payan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. C.</given-names>
            <surname>Cepeda</surname>
          </string-name>
          ,
          <article-title>Exploring the digital gender divide: Insights from the colombian case</article-title>
          ,
          <source>IDIA2020</source>
          (
          <year>2020</year>
          )
          <fpage>69</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Y. M.</given-names>
            <surname>Shum</surname>
          </string-name>
          ,
          <year>2020</year>
          , Situación digital,
          <source>internet y redes sociales colombia</source>
          <year>2020</year>
          , URL: https: //yiminshum.com/social-media-colombia-2020/.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>R.</given-names>
            <surname>Wirth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hipp</surname>
          </string-name>
          , Crisp-dm:
          <article-title>Towards a standard process model for data mining, in: Proceedings of the 4th international conference on the practical applications of knowledge discovery and data mining</article-title>
          , volume
          <volume>1</volume>
          , Springer-Verlag London, UK,
          <year>2000</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>11</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S. Z.</given-names>
            <surname>Cremades</surname>
          </string-name>
          ,
          <article-title>Redes sociales para la prevención del suicidio juvenil, 3C TIC</article-title>
          .
          <article-title>Cuadernos de desarrollo aplicados a las TIC (</article-title>
          <year>2019</year>
          )
          <fpage>54</fpage>
          -
          <lpage>69</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Saura</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Reyes-Menéndez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Palos-Sanchez</surname>
          </string-name>
          ,
          <article-title>Un análisis de sentimiento en twitter con machine learning: Identificando el sentimiento sobre las ofertas de# blackfriday</article-title>
          ,
          <source>Revista Espacios</source>
          <volume>39</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>J. C.</given-names>
            <surname>Barriga Mariño</surname>
          </string-name>
          , et al., Desarrollo y aplicación de una herramienta de extracción y almacenamiento de datos de
          <article-title>twitter a un contexto social de violencia política, technology (</article-title>
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>R.</given-names>
            <surname>Silva</surname>
          </string-name>
          , et al.,
          <article-title>Detección de violencia verbal hacia las mujeres en redes sociales mediante técnicas de aprendizaje automático, technology (</article-title>
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>G.</given-names>
            <surname>Evolvi</surname>
          </string-name>
          ,
          <article-title>Hate in a tweet: Exploring internet-based islamophobic discourses</article-title>
          ,
          <source>Religions</source>
          <volume>9</volume>
          (
          <year>2018</year>
          )
          <fpage>307</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>E. W.</given-names>
            <surname>Pamungkas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Basile</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Patti</surname>
          </string-name>
          ,
          <article-title>Misogyny detection in twitter: a multilingual and cross-domain study</article-title>
          ,
          <source>Information Processing &amp; Management</source>
          <volume>57</volume>
          (
          <year>2020</year>
          )
          <fpage>102360</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Google</surname>
          </string-name>
          ,
          <year>2021</year>
          ,
          <article-title>Cloud firestore data model</article-title>
          , URL: https://firebase .google.com/docs/firestore/ data-model.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>G. A.</given-names>
            <surname>García</surname>
          </string-name>
          <string-name>
            <surname>Vélez</surname>
          </string-name>
          ,
          <article-title>Aplicación de la metodología crisp-dm a la recolección y análisis de datos georreferenciados desde twitter, technology (</article-title>
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>D. M. Blei</surname>
            ,
            <given-names>A. Y.</given-names>
          </string-name>
          <string-name>
            <surname>Ng</surname>
            ,
            <given-names>M. I. Jordan</given-names>
          </string-name>
          , Latent dirichlet allocation,
          <source>the Journal of machine Learning research 3</source>
          (
          <year>2003</year>
          )
          <fpage>993</fpage>
          -
          <lpage>1022</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>F. I. Nicolai</given-names>
            <surname>Manaut</surname>
          </string-name>
          , Sistema de análisis de
          <article-title>tópicos para interacciones cliente-call center, technology (</article-title>
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>