<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Rhodes, Greece
damian.a.furman@gmail.com (D. A. Furman)
https://damifur.github.io/ (D. A. Furman)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>An Initial Exploration of How Argumentative Information Impacts Automatic Generation of Counter-Narratives Against Hate Speech</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Damián Ariel Furman</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pablo Torres</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>José A. Rodríguez</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Diego Letzen</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vanina Martínez</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Laura Alonso Alemany</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Arti cial Intelligence Research Institute (IIIA-CSIC)</institution>
          ,
          <addr-line>Barcelona</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>CONICET</institution>
          ,
          <addr-line>Godoy Cruz 2290, Buenos Aires</addr-line>
          ,
          <country country="AR">Argentina</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Universidad Nacional de Córdoba</institution>
          ,
          <country country="AR">Argentina</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of Buenos Aires (UBA))</institution>
          ,
          <addr-line>Intendente Güiraldes 2160 - Ciudad Universitaria, Buenos Aires</addr-line>
          ,
          <country country="AR">Argentina</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>Fighting hate speech through automatic counter-narrative generation is gaining interest because of the increasing capabilities of Large Language Models. However, counter-narrative generation is a challenging task that can bene t from insightful analyses of text. In this work, we present an approach to improve the generation of counter-narratives by providing Large Language Models with high-quality examples. In addition, we show that enhancing the original hate speech with an argumentative analysis, identifying justi cations and conclusions, together with collectives and the properties associated to them, seems to produce some improvements, specially with with smaller training datasets, helping to orient the generation towards a particular response strategy. The dataset of counter-narratives with argumentative information is made publicly available. Warning: This work contains o ensive and hateful text that may be distressing. It does not represent the views of the authors.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Counter-narrative generation</kwd>
        <kwd>Hate speech</kwd>
        <kwd>Argument mining</kwd>
        <kwd>Large Language Models</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>In social media platforms, hate speech is ampli ed beyond human scale, spreading faster and
increasing their reach, with negative impacts in societies, like polarization or an increase in
violent episodes against targeted communities or individuals. It is because of these known
consequences that many legal systems typify it as a crime, at least in some of its forms.</p>
      <p>
        The predominant strategy adopted so far to counter hate speech in social media is to recognize,
block and delete these messages and/or the users that generated it. This strategy has two main
disadvantages. The rst one is that blocking and deleting may prevent a hate message from
spreading, but does not counter its consequences on those who were already reached by it.
The second one is that there is no place for subtleties or shades while de ning hate speech: it
must be done as a binary classi cation because the consequence of that classi cation is binary.
This can generate accusations of overblocking or censorship, and not just because of errors
in automated systems, which have been shown to be highly biased [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], but because blocking
seems to be an overly simplistic approach to deal with the inherent complexity of hate speech.
      </p>
      <p>
        An alternative to blocking that has been gaining attention in the last years, is to "oppose
hate content with counter-narratives (i.e. informed textual responses)" [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ]1. This way, the
consequences of errors in the hate classi cation are minimized, overblocking is avoided, and it
helps to spread a message against hate that can reach people that are not necessarily convinced,
or even not involved in the conversation.
      </p>
      <p>However, the huge volume of online hate messages makes the manual generation of
counternarratives an impossible task. In this scenario, automating the generation of counter-narratives
is an appealing avenue, but the task poses a great challenge due to the complex linguistic and
communicative patterns involved in argumentation.</p>
      <p>
        Traditional machine learning approaches have typically produced less than satisfactory
results for argumentation mining and generation. However, the recent availability of Large
Language Models (LLMs) provides a promising approach to address the task of counter-narrative
generation. Indeed, LLMs seem capable of generating satisfactory text for many tasks. Thorburn
and Kruger [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] showed that a version of ChatGPT can tackle 6 argumentative reasoning tasks
with some degree of success. They also nd that netuning the LLM parameters outperforms
prompt-only based approaches.
      </p>
      <p>
        However, as Hinton and Wagemans [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] show in their in-depth analysis of the argumentative
capabilities of GPT-3, the argumentative text generated by LLMs tends to show some weaknesses.
Although the language they use is clearly argumentative, as is the structure of arguments they
create, most of them are not considered acceptable by humans, falling in fallacies like ’begging
the question’ and providing mostly irrelevant information.
      </p>
      <p>In this paper we present an initial exploration of the impact of argumentative information
in improving the quality of arguments generated by LLMs, more concretely, in improving the
quality of automatically generated counter-narratives against hate speech. We compare di erent
scenarios: LLMs without any speci c adaptation to the task or domain, with ne-tuning using
a dataset of counter-narratives, in a few-shot approach, and providing additional information
about some of the argumentative aspects of the hate speech.</p>
      <p>
        To assess the quality of the counter-narratives generated in the di erent scenarios, we carry
out a preliminary evaluation with human judges, who achieved moderate agreement between
each other. Based on those judgements, we can say that argumentative information by itself
does not produce an improvement in the counter-narratives, but high-quality, speci cally
targeted ne-tuning seems to have a positive impact. Argumentative information does produce
improvements in scenarios with very small training data and very speci c ne-tuning, which
seems promising to produce highly tailored counter-narratives, as in Gupta et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>The rest of the paper is organized as follows. In the next section, we review relevant work
related to automated counter-narrative generation and argumentative analysis of hate speech.
Then in Section 3 we describe our dataset of counter-narratives, with which we carry out the
1No Hate Speech Movement Campaign: http://www. nohatespeechmovement.org/
comparison of scenarios described in Section 4, where we also describe extensively our approach
to the evaluation of generated counter-narratives, based on human judgements, and the prompts
used to obtain the counter-narratives. Results analyzed in Section 5 show how ne-tuned LLMs
and argumentative information provide better results, which we illustrate with some examples.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related work</title>
      <p>Automated counter-narrative generation has been recently tackled by leveraging the rapid
advances in neural natural language generation. As with most natural language generation
tasks in recent years, the basic machine learning approach has been to train or ne-tune a
generative neural network with examples speci c to the target task.</p>
      <p>
        The CONAN dataset [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] is, to our knowledge, the rst dataset with counter-narratives. It
has 4078 Hate Speech – Counter Narrative original pairs manually written by NGO operators,
translated to three languages: English, French and Italian. Data was augmented using automatic
paraphrasing and translations between languages to obtain 15024 nal pairs of hate speech –
counter-narrative. Unfortunately, this dataset is not representative of the language in social
media.
      </p>
      <p>
        Similar approaches were carried out by Qian et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] and Ziems et al. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Qian et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]’s
dataset consists of reddit and Gab conversations where Mechanical Turkers identi ed hate
speech and wrote responses.Ziems et al. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] did not produce new text, but labeled COVID-19
related tweets as hate, counter-speech or neutral based on their hatefulness towards Asians.
      </p>
      <p>
        In follow-up work to the seminal CONAN work, Tekiro lu et al.[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] applied LLMs to assist
experts in creating the corpus, with GPT-2 generating a set of counter-narratives for a given
hate speech and experts editing and ltering them. Fanton et al.[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] iteratively re ned a LLM
where the automatically generated counter-narratives were ltered and post-edited by experts
and then fed them to the LLM as further training examples to ne-tune it, in a number of
iterations. Bonaldi et al. [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] apply this same approach to obtain a machine-generated dataset of
dialogues between people producing hate speech and experts in hate countering. As a further
enhancement in the LLM-based methodology, Chung et al. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] enhanced the LLM assistance
with a knowledge-based retrieval architecture to enrich counter-narrative generation.
      </p>
      <p>
        Ashida and Komachi [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] use LLMs for generation with a prompting approach, instead of
ne-tuning them with manually created or curated examples. They also propose a methodology
to evaluate the generated output, based on human evaluation of some samples. This same
approach is applied by Vallecillo-Rodríguez et al. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] to create a dataset of counter-narratives
for Spanish. Both these approaches are targeted to user-generated text, closely related to social
media.
      </p>
      <p>However, none of the aforementioned datasets or approaches to counter-narrative generation
includes or integrates any additional annotated information apart from the hate message,
possibly its context, and its response. That is why we consider an alternative approach that
aims to reach generalization not by the sheer number of examples, but by providing a richer
analysis of such examples that guides the model in nding adequate generalizations. We believe
that information about the argumentative structure of hate speech, may be used as constraints
for automatic counter-narrative generation.</p>
      <p>
        Chung et al. [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] address an argumentative aspect of hate speech countering. They classify
counter-narratives by type, using a LLM, and showing that knowledge about the type of
counternarratives can be successfully transferred across languages, but they do not use this information
to generate counter-narratives.
      </p>
      <p>
        To our knowledge, ours is the only corpus where tweets of hate speech have been annotated
with argumentative information: ASOHMO [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], based on the Hateval corpus [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. This dataset
enriches the argumentative tweets of Hateval [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] with a manual analysis of their argumentative
aspects, following an adaptation of the proposal of Wagemans [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], an analytic approach to
represent the semantics of the core schemes proposed by Walton et al. [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], with fewer categories
based on a limited set of general argument features. The following argumentative aspects are
manually identi ed in tweets:
• Justi cations and Conclusions.
• Type of Justi cation and Conclusion: Fact, Policy or Value.
• A Pivot signalling the argumentative relation between Justi cation and Premise.
• Two domain-speci c components: theCollective which is the target of hate, and the
      </p>
      <p>Property that is assigned to such Collective.</p>
      <p>In this work, we present counter-narratives manually associated to the hate tweets in
ASOHMO and present an initial exploration of the impact of di erent kinds of information
(counter-narratives, counter-narratives by subtype and information about argumentative
components) in improving automatic generation of counter-narratives.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Creating counter-narratives associated to argumentative aspects of hate speech</title>
      <p>
        Here we present CONEAS (Counter-Narratives Exploiting Argumentative Structure), a dataset
of counter-narratives de ned according to the argumentative information labeled on tweets
from ASOHMO [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. Each argumentative tweet is paired with counter-narratives of three
di erent types de ned by applying systematic transformations over argumentative components
of the tweet, and a fourth type consisting of any counter-narrative that does not fall under any
of the other three.
      </p>
      <p>All counter-narratives, regardless of their type, also follow the guidelines of the Get The
Trolls Out project2: don’t be aggressive or abusive, don’t spread hate yourself, try to de-escalate the
conversation, respond thinking on a wider audience than the person posting the original tweet and
try to build a narrative. Annotators were suggested to try to write at least one counter-narrative
of each type but only if they came naturally, otherwise they could leave it blank.</p>
      <p>The instructions to generate each type of counter-narrative are as follows:</p>
      <sec id="sec-3-1">
        <title>Negate Relation Between Justification And Conclusion (Type A) Negate the implied re</title>
        <p>lation between the justi cation and the conclusion.
HATE TWEET:
user must deport all illegal migrants india already reeling under constant
threat of muslim radicals curb population
Justification: india already reeling under constant threat of muslim radicals curb
population (fact)
Conclusion: must deport all illegal migrants (policy)
Collective: illegal migrants
Property: muslim radicals
COUNTER NARRATIVE A (Negate relation between justification and conclusion)</p>
        <p>Deporting illegal migrants will not mitigate the problems with muslim radicals.</p>
        <p>COUNTER NARRATIVE B (Negate relation between collective and property)</p>
        <p>Illegal migrants are not necessarily muslim radicals.</p>
        <p>COUNTER NARRATIVE C (Negate justification based on type)</p>
        <p>It is not true that India is reeling under threat of muslim radicals.</p>
      </sec>
      <sec id="sec-3-2">
        <title>FREE COUNTER NARRATIVE (Free)</title>
        <p>Deporting illegal migrants without consideration to their circumstances is an inhumane move.</p>
      </sec>
      <sec id="sec-3-3">
        <title>Negate association between Collective and Property (type B) Attack the relation between</title>
        <p>the property, action or consequence that is being assigned to the targeted group and the
targeted group itself.</p>
      </sec>
      <sec id="sec-3-4">
        <title>Attack Justification based on it is type (Type C) If the justi cation is a fact, then the fact</title>
        <p>must be put into question or sources must be asked to prove that fact. If it is of type
“value”, it must be highlighted that the premise is actually an opinion, possibly relativizing
it as a xenophobous opinion. If it is a “policy”, a counter policy must be provided.
Free Counter-Narrative (type D) All counter-narratives that the annotator comes up with
and do not fall within any of the other three types.</p>
        <p>An example of each type of counter-narrative can be seen in Figure 1. Our dataset3 consists
of a total of 1722 counter-narratives for 725 argumentative tweets in English and 355
counternarratives for 144 tweets in Spanish (an average of 2.38 and 2.47 per tweet respectively). Table 1
shows the percentage of tweets that has a counter-narrative of each type.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Experiments</title>
      <p>We designed a series of experiments to assess the impact of high-quality examples and
argumentative information in the automatic generation of counter-narratives via prompting LLMs.
We want to explore the following approaches:
Fine-tuned vs Few-shot Use a LLM that has been trained for general purposes to generate
counter-narratives by prompting the LLM with some examples of the desired input-output,
as shown in the left column of Figure 3, or take a general LLM and ne-tune it with the
examples of hate tweets associated to manually generated counter-narratives.</p>
      <sec id="sec-4-1">
        <title>With or without argumentative information We want to assess the impact of di erent</title>
        <p>combinations of argumentative information provided within the input of the model:
Collective and Property; Justi cation, Conclusion and Pivot; and all types.</p>
        <p>With specific kinds of counter-narratives We pretrained two models for each type of
counternarrative using only that type: one without extra information and another adding
argumentative information relevant for the correspondent type (Justi cation and Conclusion
for type A, Collective and Property for type B and Justi cation for type C).</p>
        <p>
          Small or Big size of the same kind of LLM We want to compare performance of a larger
model with higher hardware requirements against a smaller one, ne-tuned, cheaper to
run but requiring a speci c annotated dataset. After testing behavior of similar
alternatives (Bloom, GPT-J and GPT2), we chose Flan-T5 [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ], an open model with base (250M
parameters) and XL (3B parameters) versions that is instrution- ne-tuned.
        </p>
        <p>Few-shot experiments were conducted for Flan-T5 Base (small) and XL (larger) models.
ne-tuning was only conducted on Flan-T5 Base due to computational resource constraints.</p>
        <p>We conducted some manual evaluation of prospective to nd optimal parameters for
generation, and we found that using Beam Search with 5 beams yielded the best results, so this is the
con guration we used throughout the paper.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.1. Fine-tuning of the LLM with counter-narratives</title>
        <p>To ne-tune FLAN-T5 with our dataset of counter-narratives, we randomly split our dataset in
training, development and test partitions, assuring that all counter-narratives for the same hate
tweet are contained into the same partition. Details can be seen on Table 1.</p>
        <p>Train
Dev
Test</p>
        <p>English Spanish
#Tweets #CNs % corpus A B C #Tweets #CNs % corpus
509 1201 69.8% 496 238 467 105 257 72.4%
71 173 10.0% 67 38 68 12 27 7.6%
145 348 20.2% 138 74 136 27 71 20%</p>
        <p>Proportion of tweets with counter-narrative
96% 47% 90%</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.2. Experiments based on few-shot</title>
        <p>For the few-shot experiments, the prompt has an instruction followed by two random
examples taken from the test partition of the dataset. For each example, the hate tweet and its
corresponding counter-narrative are enclosed in special tokens de ning the start and end.</p>
      </sec>
      <sec id="sec-4-4">
        <title>4.3. Evaluation method for generated counter-narratives</title>
        <p>
          Evaluation of counter-narratives is not straightforward. So far, no automatic technique has been
found satisfactory for this speci c purpose. Automatic metrics proposed for other NLP tasks,
like BLEU [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] for automatic translation or ROUGE [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] for summarization, are not adequate
for this task because they rely strongly on word or n-gram overlap with manually generated
examples. These measures are disputed in the NLP community because, among other factors,
they can’t be adapted to cases where there can be many possible good outputs of the model,
with signi cant di erences between themselves, such as our case. We discarded these measures
after comparing di erent counter-narratives of a same tweet from our dataset and noting that
many of them scored 0 on both.
        </p>
        <p>
          Faced with the lack of appropriate automatic metrics adequate for the task, many authors
have conducted manual evaluations for automatically generated counter-narratives. Manual
evaluations typically distinguish di erent aspects of the adequacy of a given text as a
counternarrative for another. Chung et al. [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] evaluate three aspect of the adequacy of
counternarratives: Suitableness (if the counter-narrative was suited as a response to the original hate
message), Informativeness (how speci c or generic the response is) andIntra-coherence (internal
coherence of the counter-narrative regardless of the message it is responding to). Ashida and
Komachi [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ], on the other hand, assess these three other aspects: O ensiveness, Stance (towards
the original tweet) and Informativeness (same as Chung et al. [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]).
        </p>
        <p>Based on these previous works, we have put together a rst version of criteria to manually
evaluate4 the adequacy of counter-narratives, considering four di erent aspects:
• O ensiveness: if the tweet is o ensive to either the target group, the author of the tweet
or any other group or person. Possible values are: O ensive; Possibly O ensive/Not clear;
Not o ensive.
• Stance: if the tweet supports or counters the speci c message of the hate tweet. Possible
values are: Supports the original message; Not clear/Changes subject wrt original tweet;
Counters the original message. Stance incorporates a certain notion of suitableness,
since it assigns value "Changes the subject" if the counter-narrative is not responding
speci cally to the standpoint of the original tweet.
• Informativeness: Evaluates the complexity and speci city of the generated text. Only
counter-narratives with a "Counters" Stance are evaluated. Possible values are:
1. Generic statement: replies that don’t incorporate any information mentioned on
the tweet and could counter many di erent hate messages (e.g "I don’t think so" or
"That is not true").
2. Speci c but not argumentative: the reply is a simple statement, possibly
composed of a single sentence without providing justi cation for the stance but referring
to some speci c aspect of the original tweet. Usually they comply with a formula
composed of a pre x (like "I don’t think that" or "Do you have proof that") and a
verbatim copy of some part of the hate tweet.
4Results of the evaluation can be found on https://shorturl.at/aetFZ
3. Speci c and Argumentative: counter-narratives with some degree of elaboration
of the information contained on the hate message. We identi ed three common
patterns that we associate with this value:</p>
        <p>A - replies that take more than one element from the original message and stablish
some relation between them (e.g. "I don’t see the relation between {element
from the original message} and {other element from the original message}").</p>
        <p>B - A simple statement declaring stance over a single element from the original
tweet but adding a second coordinated statement with personal appreciations
about it (e.g. "I don’t think we should {some policy mentioned on the tweet}. It is
a bad idea").</p>
        <p>
          C - An argumentative reply based on information not mentioned explicitly on
the original tweet, but necessarily inferred, showing a comprehensive
understanding of the meaning of the hate message (e.g. a reply to a tweet concluding
with #BuildTheWall saying "Building a wall would cost the taxpayers more" or
"Building a wall won’t give you more control over illegal tra cking").
• Felicity: This category is related to Chung et al. [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]’s Intra-Coherence, but also
considering additional dimensions like syntactical and semantic correctness. It evaluates
independently of the original tweet, if the generated text sounds, by itself, uent and
correct. There are three possible values: The text is incoherent or semantic or syntactically
incorrect; The text is coherent with small errors like incoordination of genre/tense/etc.
or repeating parts of the original text without adapting them to the text being generated;
The text is uent and sounds correct.
        </p>
        <p>Aggregating the results for these four categories, we de ne two extra concepts: Good and
Excellent counter-narratives. Good counter-narratives will be those with optimal values on
O ensiveness, Stance and Felicity. Excellent counter-narratives will be those that also have the
optimal value for Informativeness. We believe Informativeness is the most valuable of the four
categories, that is why it is determinant in characterizing Excellent counter-narratives. The
Good indicator shows that productions are not harmful or totally random.</p>
        <p>We are planning to improve the kind of information that is currently captured in the
Informativeness category in a second version of the evaluation criteria.</p>
      </sec>
      <sec id="sec-4-5">
        <title>4.4. Annotation environment and agreement</title>
        <p>To properly evaluate the quality of the generated counter-narratives with the presented method,
we conducted a preliminary manual evaluation. We evaluated three random subsets of 20
hate tweets in English and 10 in Spanish. One contains only tweets associated with
counternarratives of both types A and C on our dataset, and was used to evaluate models ne-tuned
only with these kinds of counter-narratives. Another contains only tweets associated with
counter-narratives of type B and was also used to evaluate models ne-tuned only with this
type of counter-narratives. The last subset contains tweets with counter-narrative pairs of all
types, and was used for all the rest of the experiments.</p>
        <p>We generated one counter-narrative for each tweet in the corresponding evaluation subset
for each combination of features to be assessed: few-shot, ne-tuned, with di erent kinds of
O ensiveness
Stance
Informativeness
Felicity
argumentative information, with di erent sizes of LLM. For the larger version of FLAN-T5 we
only applied the few-shot approach, and, after assessing no improvement on the smaller version,
we aborted the rest of experiments with this version of the LLM to reduce the carbon footprint
of our experiments. The results for the 18 experiments can be seen in Table 3.</p>
        <p>Then, three annotators labeled each tweet according to the four categories described above.
The nal value for each category was obtained by calculating the value with more votes (at least
two annotators agreed on the value). In total, each annotator labeled 540 hate
tweet/counternarrative pairs. Of all these, there were 10 cases where each of the three annotators labeled a
di erent value. In these cases, we adopted a conservative criterion and assigned the worst of
the three possible values.</p>
        <p>
          Table 2 shows the agreement scores between the three annotators, calculated using Cohen’s
Kappa [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ]. In most cases, agreement ranges from Moderate (0.41 &lt; &lt; 0.60) to Substantial
(0.61 &lt; &lt; 0.80), except for the agreement achieved by annotator 3 against the other two
on the category of Felicity which is just Fair (0.21 &lt; &lt; 0.40)5. As can be expected for such
an interpretative task, agreement between annotators can be improved. However, this initial
assessment served as a starting approach to assess the impact of di erent factors in the quality
of generated counter-arguments.
        </p>
        <p>We are currently working on a second version of the evaluation criteria, with more insightful
categories, expanding on Informativeness and trying to capture argument acceptability,
relevance and persuasiveness. We will check whether this improved criteria improve inter-annotator
agreement. If so, we will engage a higher number of judges and aim to obtain a more reliable
assessment of the quality of automatically generated counter-narratives.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Analysis of results</title>
      <p>Results of the manual evaluation of di erent strategies for counter-narrative generation for
English can be seen in Table 3. A summary of this table can be seen in Figure 2, which displays
the aggregated proportion of Good and Excellent counter-narratives for each strategy.</p>
      <p>
        We can clearly see that the larger versions of the model (XL) produce counter-narratives that
are less satisfactory in general, and that argumentative information only decreases the quality
of the generated text. Fine-tuned models produce better counter-narratives in general, even
if smaller. A very valuable conclusion that can be obtained from these results is that a small
number of high quality examples produce a much bigger improvement in performance than
5The interpretation of the ranges of values of the kappa coe cient is according to Landis and Koch [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ].
Approaches
Base
Base All
Base Collective
Base Premises
XL
XL All
XL Collective
XL Premises
Base
Base All
Base Collective
Base Premises
Base CNs A
Base CNs A Premises
Base CNs B
Base CNs B Collective
Base CNs C
Base CNs C Justification
40%
10%
15%
20%
10%
0%
10%
0%
25%
0%
0%
0%
5%
0%
5%
5%
20%
5%
0%
5%
5%
5%
0%
10%
5%
0%
35%
10%
35%
30%
25%
40%
0%
10%
5%
20%
15%
10%
5%
5%
10%
5%
20%
30%
15%
15%
5%
10%
35%
25%
80%
85%
65%
25%
70%
45%
90%
70%
45%
15%
0%
25%
80%
80%
85%
80%
60%
65%
10%
15%
25%
55%
using larger models, which are also more taxing.
      </p>
      <p>If we focus on Informativeness (third dimension of evaluation in Table 3, we can see that the
approaches that produce most informative counter-narratives are ne-tuned (lower half of the
Table), without a detriment in any of the other dimensions of evaluation. Interestingly, when
ne-tuned only with counter-narratives of a single type, providing argumentative information
consistently improves the informativeness of the counter-narratives, even if only slightly. We
have to take into account that such approaches use a much smaller number of counter-narratives,
as can be seen in Table 1. Even in the case of type B counter-narratives, with extremely few
examples to ne-tune, argumentative information produces an improvement in informativeness.</p>
      <p>When we make a qualitative analysis of the generated counter-narratives, we can see that
providing argumentative information about the hate tweet does yield counter-narratives that
are more speci c and informative, as can be seen in Figure3. Models counting with this
information frequently use it by negating the relation between Collective and Property or
between Justi cation and Conclusion.</p>
      <p>Results obtained for counter-narratives for Spanish hate tweets were much worse, as could
be expected given the much smaller number of examples for ne-tuning and that base LLMs</p>
      <p>Tweet with argumentative information:
street interview whit italians "send all migrants
back to where they came from they block
streets to pray " - free speech time - https://t
co/d5dqr8pg3r @user | Justification: street
interview whit italians "send all migrants back
to where they came from they block streets
to pray " (fact) | Conclusion: "send all
migrants back to where they came from they
block streets to pray " (policy) | Pivot: migrants
- they - they</p>
      <p>Counter-narrative:
I don’t think it’s a good idea to send all
migrants back to where they came from.</p>
      <p>Tweet without argumentative information:
street interview whit italians "send all migrants
back to where they came from they block
streets to pray " - free speech time - https://t
co/d5dqr8pg3r @user</p>
      <p>Counter-narrative:
I don’t think it’s the right thing to do.
perform worse for tasks in Spanish in general. Indeed, values for Informativeness and Felicity
almost never reach more than 10% positive, and Stance and O ensiveness are almost never
beyond 30% positiveness. However, the same tendency as for English could be observed:
netuned models perform better than non- ne-tuned models, even if the latter are bigger. Moreover,
argumentative information seems to make a bigger impact in improving the generated
counternarratives than in the case of English, with increases in the range of 30%-50% in the reduction
of negative scores for O ensiveness and Stance, although a decrease in Felicity. Given these
encouraging results with such few examples, we will be increasing the number of examples
with argumentative information in future work.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusions and future work</title>
      <p>We have presented an approach to generate counter-narratives against hate speech in social
media by prompting large language models with information about some argumentative aspects
of the original hate speech. We have carried out a small manual evaluation of the quality of
generated counter-narratives. This evaluation is preliminary, with a small number of judgements
and moderate to substantial inter-annotator agreement, but we have found promising tendencies.</p>
      <p>We have shown that argumentative information by itself does not improve the quality of
counter-narratives generated by LLMs, on the contrary, it may even be detrimental, specially
in the case of bigger models. However, ne-tuning a smaller model with a small corpus of
high-quality examples of pairs hate speech – counter-narrative yields some improvement in
performance. This nding has a signi cant impact both because smaller language models are
more accessible to low-budget scenarios, and because of their smaller carbon footprint.</p>
      <p>We have also shown that some kinds of argumentative information do have some positive
impact in generating more speci c, more informative counter-narratives. In particular, we have
found that the types of counter-narrative that negate the relation between the Justi cation
and the Conclusion and that negate the Justi cation have an improvement in performance if
argumentative information about the Justi cation and the Conclusion is provided.</p>
      <p>Moreover, we have also found that argumentative information makes a positive impact
in scenarios with very few tweets, as shown by our experiments for Spanish. Although the
quality of the counter-narratives generated for Spanish is much lower than for English, the fact
that argumentative information has a positive impact is encouraging, and we will continue to
annotate examples for Spanish to improve the generation of counter-narratives.</p>
      <p>We will also explore other aspects of the quality of counter-narratives, with a more insightful,
more extensive human evaluation. We will also explore the interaction between argumentative
information and other aspects, like vocabulary, level of formality, and culture.</p>
      <p>Finally, the evaluation of counter-narratives is still far from being solved. We are currently
considering di erent avenues to improve it, as it is a crucial step to advance the eld. We are
working on obtaining a higher number of judgements, but also on more insightful guidelines
that re ect more valuable aspects of counter-narratives, more related to argument acceptability.</p>
    </sec>
    <sec id="sec-7">
      <title>7. Acknowledgments</title>
      <p>This work was funded in part by Secretaría de Investigación Cientí ca y Tecnológica FCEN–UBA
(RESCS-2020-345-E-UBA-REC), CONICET under the PIP (grant 11220200101408CO), Agencia
Nacional de Promoción Cientí ca y Tecnológica, Argentina under grants PICT-2018-0475
(PRH2014-0007), PICT-2020- SERIEA-01481, and the NAACL Regional Americas Fund (2022). This
work used computational resources from CCAD – UNC (https://ccad.unc.edu.ar/), which are
part of SNCAD – MinCyT, Argentina. We specially want to thank two anonymous reviewers
that contributed to improve this work with their thoughtful and constructive comments.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>T.</given-names>
            <surname>Davidson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Bhattacharya</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Weber</surname>
          </string-name>
          ,
          <article-title>Racial bias in hate speech and abusive language detection datasets</article-title>
          ,
          <source>in: Proceedings of Third Workshop on Abusive Language Online</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Benesch</surname>
          </string-name>
          ,
          <article-title>Countering dangerous speech: New ideas for genocide prevention</article-title>
          ,
          <source>United States Holocaust Memorial Museum</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Y.-L.</given-names>
            <surname>Chung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Kuzmenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. S.</given-names>
            <surname>Tekiroglu</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Guerini, CONAN - COunter NArratives through nichesourcing: a multilingual dataset of responses to ght online hate speech</article-title>
          ,
          <source>in: ACL</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>L.</given-names>
            <surname>Thorburn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kruger</surname>
          </string-name>
          ,
          <article-title>Optimizing language models for argumentative reasoning</article-title>
          ,
          <source>in: Proceedings of the 1st Workshop on Argumentation &amp; Machine Learning co-located with 9th International Conference on Computational Models of Argument (COMMA</source>
          <year>2022</year>
          ),
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hinton</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. H. M. Wagemans</surname>
          </string-name>
          ,
          <article-title>How persuasive is ai-generated argumentation? an analysis of the quality of an argumentative text produced by the GPT-3 AI text generator</article-title>
          ,
          <source>Argument Comput</source>
          .
          <volume>14</volume>
          (
          <year>2023</year>
          )
          <fpage>59</fpage>
          -
          <lpage>74</lpage>
          . URL: https://doi.org/10.3233/AAC-210026. doi:
          <volume>10</volume>
          . 3233/AAC-210026.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>R.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Desai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Goel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bandhakavi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Chakraborty</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Akhtar</surname>
          </string-name>
          ,
          <article-title>Counterspeeches up my sleeve! intent distribution learning and persistent fusion for intent-conditioned counterspeech generation, in: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics</article-title>
          (Volume
          <volume>1</volume>
          :
          <string-name>
            <surname>Long</surname>
            <given-names>Papers)</given-names>
          </string-name>
          ,
          <source>Association for Computational Linguistics</source>
          , Toronto, Canada,
          <year>2023</year>
          , pp.
          <fpage>5792</fpage>
          -
          <lpage>5809</lpage>
          . URL: https://aclanthology.org/
          <year>2023</year>
          .
          <article-title>acl-long</article-title>
          .
          <volume>318</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Qian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bethke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. M.</given-names>
            <surname>Belding</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>A benchmark dataset for learning to intervene in online hate speech</article-title>
          , CoRR abs/
          <year>1909</year>
          .04251 (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>C.</given-names>
            <surname>Ziems</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Soni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <article-title>Racism is a virus: Anti-asian hate and counterhate in social media during the covid-</article-title>
          19 crisis,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S. S.</given-names>
            <surname>Tekiro lu</surname>
          </string-name>
          , Y.-L. Chung,
          <string-name>
            <given-names>M.</given-names>
            <surname>Guerini</surname>
          </string-name>
          ,
          <article-title>Generating counter narratives against online hate speech: Data and strategies</article-title>
          , in: ACL,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>M.</given-names>
            <surname>Fanton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Bonaldi</surname>
          </string-name>
          ,
          <string-name>
            <surname>S. S.</surname>
          </string-name>
          <article-title>Tekiro lu, M. Guerini, Human-in-the-loop for data collection: a multi-target counter narrative dataset to ght online hate speech</article-title>
          ,
          <source>in: ACK</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>H.</given-names>
            <surname>Bonaldi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dellantonio</surname>
          </string-name>
          ,
          <string-name>
            <surname>S. S.</surname>
          </string-name>
          <article-title>Tekiro lu, M. Guerini, Human-machine collaboration approaches to build a dialogue dataset for hate speech countering</article-title>
          ,
          <source>in: EMNLP</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Y.-L. Chung</surname>
            ,
            <given-names>S. S.</given-names>
          </string-name>
          <article-title>Tekiro lu, M. Guerini, Towards knowledge-grounded counter narrative generation for hate speech</article-title>
          ,
          <source>in: Findings of the ACL-IJCNLP</source>
          <year>2021</year>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>M.</given-names>
            <surname>Ashida</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Komachi</surname>
          </string-name>
          ,
          <article-title>Towards automatic generation of messages countering online hate speech and microaggressions</article-title>
          ,
          <source>in: Proceedings of Sixth Workshop on Online Abuse and Harms (WOAH)</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>M.</given-names>
            <surname>Vallecillo-Rodríguez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. Montejo</given-names>
            <surname>Ráez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Martín-Valdivia</surname>
          </string-name>
          ,
          <article-title>Automatic counter-narrative generation for hate speech in spanish</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>71</volume>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Y.-L. Chung</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Guerini</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Agerri</surname>
          </string-name>
          ,
          <article-title>Multilingual counter narrative type classi cation</article-title>
          ,
          <source>in: Proceedings of the 8th Workshop on Argument Mining</source>
          , Association for Computational Linguistics, Punta Cana, Dominican Republic,
          <year>2021</year>
          , pp.
          <fpage>125</fpage>
          -
          <lpage>132</lpage>
          . URL: https://aclanthology. org/
          <year>2021</year>
          .argmining-
          <volume>1</volume>
          .12. doi:
          <volume>10</volume>
          .18653/v1/
          <year>2021</year>
          .argmining-
          <volume>1</volume>
          .
          <fpage>12</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>D.</given-names>
            <surname>Furman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Torres</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Rodríguez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Letzen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Martínez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. Alonso</given-names>
            <surname>Alemany</surname>
          </string-name>
          ,
          <article-title>Which argumentative aspects of hate speech in social media can be reliably identi ed?</article-title>
          ,
          <source>in: Proceedings of Fourth International Workshop on Designing Meaning Representations, co-located with IWCS</source>
          <year>2023</year>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>V.</given-names>
            <surname>Basile</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bosco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Fersini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Nozza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Patti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. M.</given-names>
            <surname>Rangel Pardo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          , M. Sanguinetti, SemEval
          <article-title>-2019 task 5: Multilingual detection of hate speech against immigrants and women in Twitter</article-title>
          ,
          <source>in: Proceedings of 13th International Workshop on Semantic Evaluation</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>J. H. M. Wagemans</surname>
          </string-name>
          ,
          <article-title>Constructing a periodic table of arguments</article-title>
          ,
          <source>in: Proceedings of 11th International Conference of the Ontario Society for the Study of Argumentation</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>D.</given-names>
            <surname>Walton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Reed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Macagno</surname>
          </string-name>
          , Argumentation Schemes,
          <string-name>
            <surname>CUP</surname>
          </string-name>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>K.</given-names>
            <surname>Papineni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Roukos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Ward</surname>
          </string-name>
          , W.-J. Zhu,
          <article-title>Bleu: a method for automatic evaluation of machine translation</article-title>
          ,
          <source>in: ACL</source>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>C.-Y. Lin</surname>
            ,
            <given-names>ROUGE:</given-names>
          </string-name>
          <article-title>A package for automatic evaluation of summaries</article-title>
          , in: Text Summarization Branches Out,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>J.</given-names>
            <surname>Cohen</surname>
          </string-name>
          ,
          <article-title>A Coe cient of Agreement for Nominal Scales</article-title>
          ,
          <source>Educational and Psychological Measurement</source>
          <volume>20</volume>
          (
          <year>1960</year>
          )
          <fpage>37</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Landis</surname>
          </string-name>
          ,
          <string-name>
            <surname>G. G. Koch,</surname>
          </string-name>
          <article-title>The measurement of observer agreement for categorical data</article-title>
          ,
          <source>Biometrics</source>
          <volume>33</volume>
          (
          <year>1977</year>
          )
          <fpage>159</fpage>
          -
          <lpage>174</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>