<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Beauty Oluokun, Guilherme Paulino-Passos∗, Antonio Rago∗ and Francesca Toni</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="editor">
          <string-name>Hagen, Germany</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computing, Imperial College London</institution>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Debates and disagreements are fundamental to human progression as they challenge assumptions and foster personal growth. We have witnessed public discourse migrating online rapidly, with platforms like Reddit's r/AmITheAsshole subreddit ofering a space for users to seek judgements on their actions and engage in moral debates. However, automated methods for modelling such debates and making predictions regarding human judgement therein are lacking in the literature. In this paper, we investigate how to model and predict within these online debates using computational argumentation, a set of formalisms known to excel in representing and reasoning with knowledge in a human-like manner. Concretely, we introduce a pipeline for modelling and predicting human judgement within Reddit threads using argument mining and quantitative bipolar argumentation frameworks under gradual semantics to imitate an outside observer or arbitrator. We demonstrate that our approach achieves a reasonable degree of accuracy in this domain and, interestingly, that our model's behaviour when diferent gradual semantics are applied correlates fairly well with their theoretical properties.</p>
      </abstract>
      <kwd-group>
        <kwd>bipolar argumentation</kwd>
        <kwd>gradual semantics</kwd>
        <kwd>online debate</kwd>
        <kwd>human judgement</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        Debates and disagreements are common in everyday life, from the seemingly insignificant,
e.g. arguments between young siblings over toys, to those which concern billions of dollars,
e.g. corporate legal battles. Though they can often be
unpleasant at that moment, disputes,
i.e. processes of argumentation, are pivotal to human progression as they underpin all human
reasoning [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], enabling us to present new ideas to each other and foster personal growth. When
people with difering perspectives converse, they can challenge each other’s assumptions which
can lead to better informed decisions, especially in the political and legal realms but also on a
smaller-scale, e.g. domestic issues such as disputes between family and friends.
      </p>
      <p>
        Nowadays, public discourse is commonly held online, particularly on social media, where
participants are almost invisible to one another. This means that, when taking part in such
discourse, one cannot always be sure how their contribution to a debate afected other
participants or the outcome of the dispute. Thus, in order to predict such efects, it would be beneficial
to have models for human debate that easily assist in investigating the efect of the debate
structures and individual arguments in the conclusion of disputes. One such set of formalisms
for modelling debates is computational argumentation (see [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ] for overviews). For example,
argumentative techniques have previously been used to gather insights on online debates in
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], focusing on what views were being expressed and why.
      </p>
      <p>
        To understand why argumentation has been so efective in this task, we need to consider its
core components, as well as those of the application domain. Online debates normally consist
of chains of comments and replies which naturally create a tree-like structure. Argumentation
Frameworks (AFs) enable us to model a set of arguments that attack or support one another,
most commonly as graphs and often as trees, and have been beneficial in areas of explainable
AI, e.g. in modelling disagreements between an AI model’s output and human observer by
evaluating the dialectical acceptability or strength of the arguments [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Baroni et. al. in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]
outline the various extensions of a basic AF and unify them into a more generalised framework:
Quantitative Bipolar Argumentation Frameworks (QBAFs). Following this, Cocarascu et al.
built on Baroni’s work to deploy QBAFs in an Argumentative Dialogical Agent (ADA) model to
improve review aggregation explanations in sites such as Metacritic and Rotten Tomatoes [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
ADA utilised sentiment analysis, argument mining, and gradual evaluation for QBAFs, showing
how QBAFs can be successful in modelling a network of reviews that interact with one another.
      </p>
      <p>In this work, we propose an argumentative approach to modelling an outside observer (an
arbitrator, a judge, or simply the audience) of an online debate in order to create an automatic
model of human judgement. We leverage QBAFs to model human debates and provide tools
to further investigate the efect of the debate structure and individual arguments in the final
decision or conclusion of the dispute. We deploy our approach in the particularly intriguing
corner of the digital space that is the Reddit r/AmITheAsshole (AITA) forum. This online
community provides a space for users to explain a situation in their lives where they may have
acted like an “asshole” and receive a judgement from other users on their morality. These online
debates involve thousands of participants, from across the globe, ofering their opinions in
response to the original poster (OP) or other comments. We model AITA debates as QBAFs,
showing the framework closely mimics the flow of a thread of comments under a Reddit post
that would be read by a user of the app. As our argumentative approach is able to predict
the verdict of an AITA thread with a high success rate, we tentatively posit that it could be a
plausible model of human reasoning during the debate, which is worthy of extensive future
investigations.</p>
      <p>Our contributions are as follows:
• We present a comprehensive pipeline for modelling Reddit’s AITA threads using QBAFs to
predict their verdicts. This pipeline facilitates the mining of arguments and their relations
and the analysis of argumentative structures from online discussions.
• We evaluate our pipeline’s prediction performance wrt accuracy and discuss the
advantages and limitations of employing QBAF and diferent gradual evaluation methods in
existing literature for modelling online AITA debates.
• We provide a public dataset of 823 example threads from the AITA subreddit that have
the verdicts labelled1.
1Dataset and code available at github.com/BOluokun/reddit-argumentation.</p>
    </sec>
    <sec id="sec-3">
      <title>2. Background and Related Work</title>
      <p>
        AITA In this online forum, the OP’s situation can be judged as one of four tags: You are The
Asshole (YTA), indicating that the OP is the only one in the wrong; Not The Asshole (NTA),
meaning the OP has done no wrong and the other party in the conflict is an asshole ; No Assholes
Here (NAH), stating that no one in the situation can be rightly labelled an asshole; and Everyone
Sucks Here (ESH), expressing that both (or all) sides in the situation have acted as assholes.
They capture the four possible combinations between the OP being in the wrong or not, and
the other party (or parties) of the conflict being in the wrong or not. In order to simplify the
problem, we focus only on whether the OP is in the wrong or not. That is, we map both YTA
and ESH into YTA, and both NTA and NAH into NTA. The overall verdict comes from the
judgement indicated in the top-level comment with the highest upvote score. Nonetheless, there
may be disagreements between comments supporting the NTA and NAH verdicts, likewise
between YTA and ESH verdicts, which we leave to future work. This is similar to the approach
in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], where Efstathiadis et. al. utilised a BERT model [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] that achieved an accuracy of 62%
when classifying posts and an accuracy of 86% when classifying comments. In that work, issues
may have been caused by isolating the comments from the original posts. Thus, our approach
of modelling the interactions between comments and the post, through AFs, could yield better
results.
      </p>
      <p>(a) AITA thread 1 with the NTA verdict.</p>
      <p>(b) AITA Reddit thread 2 with the YTA verdict.</p>
      <p>
        Quantitative Bipolar Argumentation Frameworks QBAFs [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] are quadruples ⟨ , A, S,  ⟩
where  is a finite set (of arguments), A ⊆  × and S ⊆  × are the attack and support relations,
respectively, and  ∶  →  assigns base scores to arguments, representing the arguments’
intrinsic strengths from within a given evaluation range  (we use  = [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ] throughout).
In our application, the arguments represent the comments in the AITA Reddit thread. For
any  ∈  , the attackers of  are A( ) = { ∈  ∣ (,  ) ∈ A} and the supporters of  are
S( ) = { ∈  ∣ (,  ) ∈ S}. QBAFs can be visualised as graphs, with arguments as nodes
labelled by their base scores and relations as edges labelled by + (for support) or - (for attack).
      </p>
      <p>Then, we say that F = ⟨ , A, S,  ⟩ is a QBAF for  ∈  if ∄(,  ) ∈ A ∪ S for any  ∈  , for all
 ∈  ⧵ { } there is a path in F from  to  , and ∄ ∈  with a path from  to  . Argument 
is called the explanandum. In our application,  is the specific statement in which the initial
stances of the participants are polar, and always refer to the statement “OP is NTA”. In a QBAF
for  , all other arguments are ‘related to’  and there are no circular paths between arguments.
We can visualise an agent’s QBAF as a tree rooted at  .</p>
      <p>
        Gradual semantics QBAFs can be paired with a gradual semantics  [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] which credits
arguments with a dialectical strength within  . Dialectical strengths encapsulate a participant’s
views on the quality of or belief in arguments within a QBAF based on a combination of the base
scores and perceived strengths of the arguments’ attackers and supporters. We focus on four
gradual semantics: the Quantitative Argumentation Debate (QuAD) [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] and Discontinuity-Free
Quantitative Argumentation Debate (DF-QuAD) [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] algorithms, Quadratic Energy Model (QEM)
[
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] and Exponent-based restricted semantics (Ebs) [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. We omit most of the formal definitions
for lack of space, but, for illustration, we give the definition of Ebs. This formula shows how
this semantics calculates the dialectical strength of an argument,  , using its base score,  ( ),
and an aggregation of the dialectical strengths of its attackers and supporters - in this case
given by  ( ).
      </p>
      <p>Definition 1 (Ebs Gradual Semantics). For F = ⟨ , A, S,  ⟩ and any  ∈  ,
 (F ,  ) = 1 −</p>
      <p>1 −  ( )2
1 +  ( ) ⋅ 2 ( )
where  ( ) =</p>
      <p>∑  (F ,  ) −
∈ S( )</p>
      <p>∑  (F ,  ).
∈ A( )</p>
      <p>
        Diferent gradual semantics may satisfy diferent properties, characterising their functionality,
including attainability (A) [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], informally enforcing that all strength scores in  are possible for
an argument, given any base score and a suitable choice of attackers and supporters; (strict)
bivariate monotony ((S)M) [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], informally describing that attacks cannot benefit their targets
and equally, supports cannot harm their targets; and (strict) Franklin ((S)F) [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], informally that
a supporter is never more important than an attacker of equal strength. We decided to focus on
the four chosen semantics, as they provided a good variation in our selected properties, all used
the same evaluation range, and they allowed for variable base score functions. More gradual
semantics could be tested in future work, e.g. the Restricted Euler-based semantics [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
Relation-based Argument Mining To obtain QBAFs from r/AmITheAsshole threads, we
use relation-based argument mining, amounting to classifying pairs of texts (a child and a
parent) as ‘Attack’, ‘Support’ or ‘No’ (depending on whether the child disagrees, agrees, or is
unrelated to the parent, respectively). Specifically, in our experiments, we use Large Language
Models (LLMs) [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] to perform relation-based argument mining using the few-shot prompting
method outlined by Gorur et al. [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
      </p>
    </sec>
    <sec id="sec-4">
      <title>3. Online Debates as Argumentation Frameworks</title>
      <p>In this section we detail our method for modelling an “arbitrator” for the AITA debates using
QBAFs under gradual semantics, as outlined in Figure 2.</p>
      <p>First, we obtain a QBAF for “OP is NTA” (the explanandum) as follows. The explanadum has
a single supporter representing the original post. The nested comments under the post are then
added to the QBAF recursively as attackers or supporters of their parent (post or comment)
depending on the classification by the relation-based argument mining component. We discard
a comment and its replies when it has a ‘No’ relation to its parent. This algorithm for converting
AITA threads into QBAFs is guaranteed to produce well-formed QBAFs for the explanandum
“OP is NTA”, as each edge between arguments is exclusively either an attack or a support, all
comments are connected to the explanandum through a reply chain and no argument attacks
or supports itself. These QBAFs are also guaranteed to be acyclic (in that there is no path from
any argument to itself).</p>
      <p>In the QBAF we need to assign appropriate base scores ( ) for all arguments. We propose two
elementary methods for this. A basic method is to use a ‘fixed’  which assigns each argument
(including the explanandum) the same value:
 fixed ( ) = 
(1)
for a chosen value  ∈  . This can be interpreted as every comment in the thread being equally
influential and trustworthy. This is quite a naive approach but a useful baseline to evaluate the
importance of the structure of the debate - chains of supports and attacks - in determining the
verdict rather than features of the comments themselves. Figure 3a shows an example when
using  fixed for an AITA thread with DF-QuAD gradual semantics ( 1) and diferent values of  :
 1 = 0.1,  2 = 0.25 and  3 = 0.4.</p>
      <p>The second method is setting  to vary with the number of upvotes a comment has. A
comment having a greater upvote score indicates more participants who are reading the thread
agree with the content of that comment. This suggests that a comment with a high upvote
score should have a high base score, and a comment with a low upvote score should have a low
base score. The range of upvotes varies between AITA threads and is unbounded so it is not
appropriate to select an absolute scale with specific ‘high’ and ‘low’ values. Instead, for a thread
 , We find the maximum and minimum upvotes (  max and  min respectively) and define a base
score function such that for an arbitrary comment  with  max ≠  min:
 upvote( ) =  +  ×</p>
      <p>−  min
 max −  min
so we set, for  representing one of these arguments:
where  and  are parameters to be chosen and   is the upvote score of comment  . Using eq. (2)
to determine base scores gives the comment with the lowest number of upvotes a base score
of  and the comment with the highest number of upvotes a base score of  +  . For any two
comments ,  , if score( ) &lt; score( ), then  upvote( ) &lt;  upvote( ) and if score( ) = score( ),
then  upvote( ) =
 upvote( ). The explanandum and the original post do not have upvote scores
 upvote( ) =  + 2

(2)
(3)
Equation (3) is also used if  max =</p>
      <p>min. Figure 3b shows an example of evaluating the stance
of the explanandum in a QBAF with QuAD ( 2) and Ebs ( 3) with  upvote where  = 0.05 and
 = 0.35. For all four gradual semantics, if an argument has no supporters or attackers its
strength is equal to its base score ( ), we see argument 7 has fewer upvotes than argument 6 and
thus is assigned a lower base score with  upvote. Note that, if ,  ∈ 
(defined in eq. ( 1)) and  upvote (defined in eq. ( 2) and .
(3)) are well-defined.</p>
      <p>and  +  ∈  then  fixed
(a) Thread 1 arguments’ strengths with  fixed</p>
      <p>(b) Thread 2 arguments’ strengths with  upvote
diferent base scores and gradual semantics.</p>
      <p>Once the relations between comments and replies in a thread are assigned and the base scores
decided, the QBAF classifier , i.e. the AITA “arbitrator”, or audience, can be built specific to that
thread. The QBAF classifier predicts the verdict of the AITA thread using a chosen gradual
semantics to evaluate the explanandum, giving a score in the evaluation range. If the score is
strictly above the neutral value in  (0.5), we predict the stance NTA (the positive class) as the
QBAF has calculated a strong belief in the argument “OP is NTA”. Contrastingly, if the score
is below or equal to the neutral value we predict the stance YTA (negative class) because the
QBAF has calculated a lack of belief in the explanandum, suggesting the OP has likely been an
asshole. Normally, an argument can have a neutral stance (strength) but we have restricted our
use of strength scores so the QBAFs act as binary classifiers.</p>
      <p>
        Implementation details We build QBAFs by applying a depth-first algorithm on a JSON
representation of AITA threads which produces QBAF objects encapsulating NetworkX [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]
directed graphs (representing the arguments in QBAFs as nodes, with the edges having either an
‘Attack’ or ‘Support’ label). Each QBAF object has tau, semantics and eval_range attributes
which are functions corresponding to the base score function ( ), gradual semantics ( ) and
evaluation range ( ), respectively.
      </p>
      <p>
        For the relation-based argument mining, we used the method of [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] (of-the-shelf) with
a 4-bit quantisation of the Mistral-7B-Instruct-v0.2, due to hardware limitations. However,
memory and storage limitations did not reduce the efectiveness of the LLM as [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] reported
similar accuracy and macro F1 as Llama70B-4bit. Moreover, the message sent to the LLM
consists of a 7-shot prompt primer followed by the parent comment (Arg1), the child comment
(Arg2) and the line “Relation:” in an identical form to the examples.
      </p>
    </sec>
    <sec id="sec-5">
      <title>4. Evaluation</title>
      <p>Experiment Set-Up Numerous databanks of posts from r/AmITheAsshole exist publicly2.
However, many of these separate the original post content from the comments underneath and
thus lose the structure of the discussion about the post that is integral to our research. Therefore,
we decided to use Reddit’s API, through PRAW 3, to scrape snippets of AITA threads. Our
dataset is stored in a TSV file and has a total of 823 entries. Reddit’s public API was employed
to scrape threads from the AITA subreddit and from an initial sample of 1000 threads; we only
retained those which were correctly labelled with a verdict leaving an arbitrary number of 823
entries. For each, we recorded the title, the verdict, the filename of the JSON containing the
thread contents, and the total number of comments (including the original post). Overall, the
dataset has 390 NTA entries and 433 YTA entries, which is not a significant class imbalance.</p>
      <p>Of the 823 example threads in the data set, 165 examples were set aside for testing. 5-fold
cross-validation, performed on the remaining 658 examples, was used to determine the optimal
values of  and  for eq. (1) and eq. (2) for each gradual semantics. All eight combinations of
gradual semantics ( ) and base score functions ( ) for QBAF classifiers were tuned and tested.</p>
      <p>
        To tune  for the QBAF classifiers using  fixed , we tested each classifier with 100 values for 
ranging from 0.01 to 0.5. This range was chosen because preliminary testing suggested that
the performance of  &gt; 0.5 would be too low. The optimal  chosen for a classifier was the
one which produced the highest mean F1 score across the 5 folds. Likewise, to tune  and 
2E.g. dvc.ai/blog/a-public-reddit-dataset.
3praw.readthedocs.io/en/latest/
for the  upvote classifiers, we tested 1600 combinations of (,  ) with  ranging from 0.01 to
0.2 and  ranging from 0.01 to 0.8. These ranges were chosen to explore performance over
the full evaluation range  = [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ] as 0.2 + 0.8 = 1 is the maximum possible base score and
0.01 + 0.01 = 0.02 is very close to the minimum.
      </p>
      <p>We did not utilise an LLM as a baseline in our experiments as we were aiming for explainability:
an LLM could be tuned to perform well on our dataset, however its black-box nature would likely
make it dificult to interpret how its prediction was formulated. Exploring the performance of a
purely LLM-based system is left for future work.</p>
      <p>Experiment Results The optimal parameters for  and  are as in Table 1. The performance
of the eight QBAF classifiers with respect to their optimal values for  and  are shown in
Table 2.</p>
      <p>Note that DF-QuAD, QuAD and Ebs have an optimal  below 0.25 whereas QEM prefers
an  above 0.4. With  upvote, all four gradual semantics prefer a low value for  (below 0.05)
paired with a higher value for  . However, like with  fixed , QEM finds higher base scores more
optimal with  &gt; 0.7 while the other three semantics prefer smaller ranges. Figure 3a highlights
how the choice of  can greatly influence the prediction of the QBAF arbitrator. All arguments
support their parent, which should result in a prediction that agrees with the true verdict of
NTA, however, DF-QuAD with  1 = 0.1 is not capable of building strength in the explanandum
which is greater than 0.5. Furthermore,  3 = 0.4 provides the correct prediction but with an
extremely high strength, placing too much trust in individual arguments. Thus the aggregation
of attackers and supporters may cause too extreme changes from the base score of an argument.</p>
      <p>
        Overall, Table 2 highlights that QBAF classifiers can be successful at determining the verdicts
of AITA threads, with the lowest F1 score being 0.7374 for Ebs with  fixed and the best performing
QBAF classifier using Ebs and  upvote with an F1 score of 0.8261, closely followed by DF-QuAD
with the highest ROC-AUC score. The variation in performance across the gradual semantics
seems to be correlated with the properties each satisfy (see Table 3)4, as discussed below.
4The proofs for the semantics’ satisfaction of these properties are shown in [
        <xref ref-type="bibr" rid="ref13 ref14 ref6">6, 13, 14</xref>
        ] except for Attainability (A) for
QEM, which is satisfied but we omit the proof for lack of space.
      </p>
      <p>DF-QuAD</p>
      <p>QEM
QuAD</p>
      <p>Ebs</p>
      <p>Discussion Predictably, all  fixed versions of the QBAF classifiers performed worse than their
 upvote counterparts. However, the diferences for DF-QuAD and QEM were less significant,
suggesting that with the appropriate gradual semantics, the structure of the debate can be
reasonably suficient in determining a arbitrator’s verdict. When using  fixed , each argument is
equally weighted and all four gradual semantics satisfy monotony, though not necessarily the
strict version, therefore an argument’s strength is afected solely by the number of attackers
and supporters it has, not the quality of its attacking and supporting arguments.</p>
      <p>The performance of the Ebs classifiers were the most interesting since, depending on the
choice of base score, results were disparate. That is, using  upvote produced the best results
among all semantics for  1 score, while using  fixed gave the worst. As shown in Table 2, recall
was high using both  fixed and  upvote, but precision increased with  upvote. QuAD’s performance
was also significantly improved by changing from  fixed to  upvote. Like Ebs,  upvote did not really
afect its recall but improved its precision; however, QuAD’s average recall is quite low so it still
performs the worst overall out of the four gradual semantics. DF-QuAD improves on the QuAD
gradual semantics by removing the discontinuities and introducing the strict Franklin property
while maintaining attainability and monotony. For example, in Figure 3b, Ebs is able to correctly
predict the YTA verdict as the attack from argument 2 (the top comment) overcomes the support
from argument 5 and decreases the strength of the original post -  upvote(1) = 0.2250 by eq. (3).
Contrastingly, QuAD narrowly fails and predicts NTA as the strength of supporting argument
5 is higher than the strength of the attacking argument 2, despite argument 2 having a greater
base score due to a higher upvote score. This imbalance in the afects of supports and attacks is
likely due to QuAD not satisfying the Franklin property.</p>
      <p>Across all QBAF classifiers tested, recall is higher than precision, suggesting that
argumentation frameworks are particularly useful for identifying the positive NTA verdict but it struggles
to align with the negative YTA verdict. DF-QuAD and QEM have the most stable performance
with the two versions of  presented, exhibiting less significant changes in recall and precision .
This is probably due to their satisfaction of attainability, bivariate montony and strict Franklin.</p>
      <p>Note that the lack of an of-the-shelf method (based on LLMs) for the relation-based argument
mining, not requiring training, could lead to some unpredictable results and errors in the
prediction or relations between comments. For example, if in our first example (Figure 3a), the
LLM determined that the top comment (argument 2) attacked the original post instead, all the
calculation in the QBAF would have the wrong efect. Even if all other relations are correctly
determined, all the gradual semantics we have explored would decrease the strength of the
original post (argument 1) based on the strength of the top comment, causing the QBAF to
calculate  (0) &lt; 0.5 and to incorrectly identify the OP as an asshole.</p>
    </sec>
    <sec id="sec-6">
      <title>5. Conclusions and Future Work</title>
      <p>
        In this paper, we have presented a novel approach to using symbolic techniques to model
and explore human behaviour during debate and moral judgement. We outline a method for
modelling an online debate, such as a thread on Reddit’s r/AmITheAsshole subreddit, and
explore how various choices of base scores ( ) and gradual semantics ( ) afect the performance
of the QBAF as a classifier due to argumentation properties. Our work can be seen as steps
towards building explainable and accurate models of online debate and human judgement which
are lacking in the literature. Our QBAF classifier for Reddit threads was fairly successful and
appears to rival other NLP-focused methods such as those in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Moreover, it maintains a
beneficial level of human-interpretability, as a user may select any argument and calculate its
strength to determine how important it may have been to the verdict, in addition to eficiently
examining the structure of the debate.
      </p>
      <p>We suggest future work should first focus on refining the relation-based argument mining
specifically for the purpose of mining social media threads. A fine-tuned transformer model
could be developed to decrease the error in QBAF models of the Reddit threads. The experiments
could then be redone to determine if an improvement in performance is possible. Moreover,
the methodology we have outlined in this project is transferable and can be easily tweaked
to be applied to other written debates. Firstly, it would be simple to apply the pipeline to
other subreddits or other social media sites such as X (Twitter) and Facebook. Other suggested
areas, away from online settings, include legal debates and cases, in which one could use
argumentation frameworks to predict a judge’s verdict on a case after the prosecution and
defence have laid out their arguments. For example, a model could be developed to assign base
scores to pieces of evidence and legal arguments in a case based on how trustworthy they are or
how influential they could be. Then, a QBAF (with an appropriate gradual semantics) could be
applied to predict if a jury or judge would likely pass a guilty verdict or not. This may be useful
in the legal field in aiding decisions on which cases to take to trial or determining which need
more evidence to be gathered. When applying QBAFs and gradual semantics in other contexts,
it would thus be ideal to explore more informative algorithms, e.g. based on NLP, to determine
the base score function  in specific situations. Additionally, more attention needs to be placed
on improving the relation-based argument mining for the chosen debate context.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>Paulino-Passos, Rago and Toni were partially funded by the European Research Council (ERC)
under the European Union’s Horizon 2020 research and innovation programme (grant agreement
No. 101020934). Rago and Toni were partially funded by J.P. Morgan and by the Royal Academy
of Engineering under the Research Chairs and Senior Research Fellowships scheme. Any views
or opinions expressed herein are solely those of the authors.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>H.</given-names>
            <surname>Mercier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Sperber</surname>
          </string-name>
          ,
          <article-title>Why do humans reason? arguments for an argumentative theory</article-title>
          ,
          <source>Behavioral and brain sciences 34</source>
          (
          <year>2011</year>
          )
          <fpage>57</fpage>
          -
          <lpage>74</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>P.</given-names>
            <surname>Baroni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Gabbay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Giacomin</surname>
          </string-name>
          , L. van der Torre (Eds.), Handbook of Formal Argumentation, College Publications,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>K.</given-names>
            <surname>Atkinson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Baroni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Giacomin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hunter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Prakken</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Reed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. R.</given-names>
            <surname>Simari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Thimm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Villata</surname>
          </string-name>
          , Towards artificial argumentation,
          <source>AI</source>
          Magazine
          <volume>38</volume>
          (
          <year>2017</year>
          )
          <fpage>25</fpage>
          -
          <lpage>36</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Lawrence</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Reed</surname>
          </string-name>
          ,
          <article-title>Using complex argumentative interactions to reconstruct the argumentative structure of large-scale debates</article-title>
          , in: ArgMining@EMNLP,
          <year>2017</year>
          , pp.
          <fpage>108</fpage>
          -
          <lpage>117</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>K.</given-names>
            <surname>Cyras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rago</surname>
          </string-name>
          , E. Albini,
          <string-name>
            <given-names>P.</given-names>
            <surname>Baroni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Toni</surname>
          </string-name>
          ,
          <string-name>
            <surname>Argumentative</surname>
            <given-names>XAI</given-names>
          </string-name>
          :
          <article-title>A survey</article-title>
          ,
          <source>in: IJCAI</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>4392</fpage>
          -
          <lpage>4399</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>P.</given-names>
            <surname>Baroni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rago</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Toni</surname>
          </string-name>
          ,
          <article-title>From fine-grained properties to broad principles for gradual argumentation: A principled spectrum</article-title>
          ,
          <source>Int. J. Approx. Reason</source>
          .
          <volume>105</volume>
          (
          <year>2019</year>
          )
          <fpage>252</fpage>
          -
          <lpage>286</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>O.</given-names>
            <surname>Cocarascu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rago</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Toni</surname>
          </string-name>
          ,
          <article-title>Extracting Dialogical Explanations for Review Aggregations with Argumentative Dialogical Agents</article-title>
          , in: AAMAS,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>I. S.</given-names>
            <surname>Efstathiadis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Paulino-Passos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Toni</surname>
          </string-name>
          ,
          <article-title>Explainable patterns for distinction and prediction of moral judgement on reddit</article-title>
          ,
          <source>CoRR abs/2201</source>
          .11155 (
          <year>2022</year>
          ). arXiv:
          <volume>2201</volume>
          .
          <fpage>11155</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          ,
          <article-title>BERT: pre-training of deep bidirectional transformers for language understanding</article-title>
          ,
          <source>in: NAACL-HLT</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>P.</given-names>
            <surname>Baroni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rago</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Toni</surname>
          </string-name>
          ,
          <article-title>How many properties do we need for gradual argumentation?</article-title>
          , in: AAAI,
          <year>2018</year>
          , pp.
          <fpage>1736</fpage>
          -
          <lpage>1743</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>P.</given-names>
            <surname>Baroni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Romano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Toni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Aurisicchio</surname>
          </string-name>
          , G. Bertanza,
          <article-title>Automatic evaluation of design alternatives with quantitative argumentation</article-title>
          ,
          <source>Argument Comput. 6</source>
          (
          <year>2015</year>
          )
          <fpage>24</fpage>
          -
          <lpage>49</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>A.</given-names>
            <surname>Rago</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Toni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Aurisicchio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Baroni</surname>
          </string-name>
          ,
          <article-title>Discontinuity-free decision support with quantitative argumentation debates</article-title>
          ,
          <source>in: KR</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>63</fpage>
          -
          <lpage>73</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>N.</given-names>
            <surname>Potyka</surname>
          </string-name>
          ,
          <article-title>Continuous dynamical systems for weighted bipolar argumentation</article-title>
          ,
          <source>in: KR</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>148</fpage>
          -
          <lpage>157</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>L.</given-names>
            <surname>Amgoud</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ben-Naim</surname>
          </string-name>
          ,
          <article-title>Evaluation of arguments in weighted bipolar graphs</article-title>
          ,
          <source>Int. J. Approx. Reason</source>
          .
          <volume>99</volume>
          (
          <year>2018</year>
          )
          <fpage>39</fpage>
          -
          <lpage>55</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>L.</given-names>
            <surname>Amgoud</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ben-Naim</surname>
          </string-name>
          ,
          <article-title>Evaluation of arguments in weighted bipolar graphs</article-title>
          ,
          <source>in: ECSQARU</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>25</fpage>
          -
          <lpage>35</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>B.</given-names>
            <surname>Min</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Ross</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Sulem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. P. B.</given-names>
            <surname>Veyseh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. H.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Sainz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Agirre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Heintz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Roth</surname>
          </string-name>
          ,
          <article-title>Recent advances in natural language processing via large pre-trained language models: A survey</article-title>
          ,
          <source>ACM Comput. Surv</source>
          .
          <volume>56</volume>
          (
          <year>2024</year>
          )
          <volume>30</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>30</lpage>
          :
          <fpage>40</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>D.</given-names>
            <surname>Gorur</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rago</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Toni</surname>
          </string-name>
          ,
          <article-title>Can large language models perform relation-based argument mining?</article-title>
          ,
          <source>CoRR abs/2402</source>
          .11243 (
          <year>2024</year>
          ). arXiv:
          <volume>2402</volume>
          .
          <fpage>11243</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>A.</given-names>
            <surname>Hagberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Schult</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Swart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Hagberg</surname>
          </string-name>
          , Exploring Network Structure, Dynamics, and
          <article-title>Function using NetworkX</article-title>
          , in: SciPy,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>