<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Exploring the potential of defeasible argumentation for quantitative inferences in real-world contexts: An assessment of computational trust</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Technological University Dublin</institution>
          ,
          <addr-line>Dublin</addr-line>
          ,
          <country country="IE">Ireland</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Argumentation has recently shown appealing properties for inference under uncertainty and conflicting knowledge. However, there is a lack of studies focused on the examination of its capacity of exploiting real-world knowledge bases for performing quantitative, case-by-case inferences. This study performs an analysis of the inferential capacity of a set of argument-based models, designed by a human reasoner, for the problem of trust assessment. Precisely, these models are exploited using data from Wikipedia, and are aimed at inferring the trustworthiness of its editors. A comparison against non-deductive approaches revealed that these models were superior according to values inferred to recognised trustworthy editors. This research contributes to the field of argumentation by employing a replicable modular design which is suitable for modelling reasoning under uncertainty applied to distinct real-world domains.</p>
      </abstract>
      <kwd-group>
        <kwd>Defeasible Argumentation</kwd>
        <kwd>Argumentation Theory</kwd>
        <kwd>Explainable Artificial Intelligence</kwd>
        <kwd>Non-monotonic Reasoning</kwd>
        <kwd>Computational Trust</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Trust is a crucial human construct investigated within several disciplines, such as
psychology, sociology and philosophy, with many applications. It is an ill-defined
construct, whose formalisation lies, among others, in the domain of knowledge
representation and reasoning. It is a complex phenomenon, essential to support decision-making
processes and delegation in uncertain domains. Many definitions of trust can be found
in the literature [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. Briefly, it can be described as a prediction that a trusted entity will
bring to completion the expectations of a trustor in some specific context. A
computational model of trust is one that brings this prediction to fruition when software agents
are involved. Such models have emerged, aimed at making use of the notion of
human trust in open digital worlds [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. They help an agent to collect, aggregate, quantify
and classify evidence to inform its decision about how/whether to interact with another
agent. Reasoning applied for the definition of computational models of trust is likely
suitable to be modelled by defeasible argumentation [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Within Artificial Intelligence,
defeasible argumentation is aimed at developing computational models of arguments
[
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. These models are typically built upon layers specialised for the definition of
internal structure of arguments, the resolution of conflicts between arguments and the
possible resolution strategies for reaching a justifiable conclusion. Here, the modelling
of reasoning via defeasible argumentation and, in turn, applied to the inference of
computational trust, is proposed in the context of Wikipedia editors. The goal is to design
knowledge-driven, argument-based models capable of assigning a trust value in the
range [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ] R to editors on a case-by-case basis. One means complete trust should
be assigned to an editor, while 0 means an absence of trust assigned to the editor. These
models are built upon domain knowledge and instantiated by quantitative data, thus can
provide numerical inferences. As with [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], assigning trust is assumed to be a reasoning
process or a rational decision grounded on evidence made by rational agents.
Moreover, it is assumed to be a defeasible reasoning process, whose underlying beliefs can
be negated by new information. For example, an initial analysis might conclude that a
Wikipedia editor should be assigned a high trustworthiness value, due to a large amount
of previous interactions performed by him/her. However, if the reputation achieved by
this agent after performing these interactions is not positive, then a new low
trustworthiness value might be inferred instead, retracting the previous conclusion. The fact that
these pieces of evidence and arguments can be withdrawn in light of new information
allows this process to be seen as a form of defeasible reasoning activity. If successful,
this reasoning activity might reinforce the generalisability of defeasible argumentation
for carrying out quantitative, case-by-case inferences with uncertain and conflicting
evidence, such as performed in other domains [
        <xref ref-type="bibr" rid="ref23 ref25 ref26">23, 25, 26</xref>
        ]. Thus, the research question
under investigation is: “Can the consideration of conflicts and their resolution through
defeasible argumentation lead to a better inference of trust of Wikipedia editors than a
non-deductive aggregation of evidence?”
      </p>
      <p>The remainder of this paper continues with Section 2 providing the related work
on computational trust. Section 3 defines the concept of better inference of trust in the
context of this study, and presents the design of an empirical experiment for tackling
the research question. The results, the analysis and the discussion of this experiment are
provided in Section 4. Lastly, Section 5 concludes the study and suggests future work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        The first computational model of trust was proposed in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Its goal was to enable
artificial agents to make trust-based decisions in the domain of Distributed Artificial
Intelligence. In general, trust evidence includes recommendation, reputation, past
interactions, credentials and many other factors that might lead to contradicting assessments
of trust. In this paper, the context under evaluation comes from the Wikipedia project.
This project is under constant change from different types of contributors, ranging from
domain experts and to casual contributors, to vandals and committed editors. Several
works have attempted to compute the trust of Wikipedia editors and Wikipedia articles.
For instance, [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] presents a content-driven reputation system for Wikipedia editors,
assuming that the reputation of editors can be used as a rough guide to the trust assigned
to articles edited by them. In turn, reputation is assigned according to the longevity
of the text inserted and the longevity of the text edited by each editor. In a subsequent
work, [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] computes the trust of a word in a Wikipedia article according to the reputation
of the original editor of the word, as well as the reputation of editors who edited
content in the vicinity of the word. The study demonstrates that text labelled as high trust
has a significantly lower chance of being edited in the future. Similarly, [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ] explores
the revision history of an article to assess the trustworthiness of the article through a
dynamic Bayesian network. A trust value is defined in the range [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ] R, where 0
means complete untrustworthiness and 1 means complete trustworthiness. A set of 200
articles was evaluated and correctly classified in approximately 83% of cases
according to a trust value threshold. The classes considered were featured articles (assumed
to be highly trustworthy for being thoroughly reviewed) and clean-up articles (marked
for major revision by editors). In short, other works evaluate the trust of Wikipedia’s
contributors through a multi-agent trust model [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] and the Wikipedia editor reputation
through the stability of content inserted [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Several works have examined the relation
between defeasible reasoning and computational trust [
        <xref ref-type="bibr" rid="ref16 ref18">16, 18</xref>
        ], or proposed
argumentbased approaches for reasoning about trust [
        <xref ref-type="bibr" rid="ref28 ref3">3, 28</xref>
        ]. However, to the best of the authors’
knowledge, the use of defeasible argumentation, instantiated by quantitative
information, for the inference of trust of Wikipedia editors as a numerical scalar has not been
attempted so far. Hence, it is expected that inferential models built with defeasible
argumentation might provide a useful approach to produce knowledge-driven, case-by-case
inferences of trust. This investigation also extends previous works [
        <xref ref-type="bibr" rid="ref14 ref23 ref25 ref26">23, 14, 25, 26</xref>
        ] which
have adopted a similar approach, but in different domains of application. Thus, an
additional goal comes from enhancing the generalisability of defeasible argumentation as
an effective approach to reason with quantitative, uncertain and conflicting information
in real-world contexts.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Design and Methodology</title>
      <p>
        A primary research study was designed, which included a comparison between the
inferences produced by defeasible argumentation models and two baseline inferences
constructed for comparison purposes. The baselines were computed by measures of
central tendency (average and weighted average) of the features employed for the
inference of computational trust, also resulting in a value in the range [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ] R. Two
knowledge bases in the form of logical expressions that can be adapted as
computational arguments were produced by the first author of this paper. These were employed
for the development of argument-based models. These models follow the five-layer
modelling approach proposed in [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] and employed in other studies [
        <xref ref-type="bibr" rid="ref23 ref25 ref26">23, 25, 26</xref>
        ]: 1)
definition of the structure of arguments, 2) definition of their conflicts, 3) their
evaluation 4) the computation of the acceptance status of each argument and 5) their final
accrual. A comparison of the inferences produced by defeasible argumentation models
and baseline measures was done by assessing the values assigned to Barnstar editors.
A Barnstar1 represents an award used by Wikipedia to recognise valuable editors. It is
a non-automatic award bestowed from a Wikipedia editor to another Wikipedia editor.
1 https://en.wikipedia.org/wiki/Wikipedia:Barnstars
      </p>
      <p>Therefore, it is not a ground truth for trust. Instead, it is used as a proxy measure in
order to evaluate the produced inferences. Two metrics are employed for comparison
of inferential models: rank of Barnstars and spread of values assigned to Barnstars.
When sorting editors in descending order by their assigned trust values, it is assumed
that the ranking of the best models will result in Barnstar editors being placed at the
highest positions. Non-Barnstar editors may also be highly trustworthy. Nonetheless,
Barnstar editors still should, presumably, be ranked at the highest positions. Moreover,
since trust is not a binary concept, it is expected that the distribution of the trust values
assigned by these same models to Barnstar editors should have a positive, continuous
spread. Spread is measured by the standard deviation of the values assigned to Barnstar
editors. Figure 1 summarises the design of the research.</p>
      <p>
        Knowledge-base 1 [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]
      </p>
      <p>
        Knowledge-base 2 [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]
Argument-based models
1. Structure of arguments
2. Conflicts of arguments
3. Evaluation of conflicts
      </p>
      <p>4. Acceptance status
5. Accrual of arguments</p>
      <p>Argument-based
models’ inferences</p>
      <p>Design of argument- based
models from
domain knowledge
Instantiation
of models</p>
      <p>
        Dataset
Inferences through
average and weighted
average of features
Comparison of
spread and rank
of Barnstar editors
An XML dump of the Portuguese-language Wikipedia was selected for examination2. It
contained 1; 076; 396 articles, 1; 798; 363 editors and 67 Barnstar editors up to
December 2018. The rationale behind this decision was merely the suitability of the dump for
the available computational resources. No natural language information contained in
each article was analysed, but only quantitative data related to editors. Each Wikipedia
page is identified by its title and it has a number of associated revisions containing:
i) its own ID; ii) a time stamp; iii) a contributor (editor) identified by a user name or
IP address if anonymous; iv) an optional commentary left by the editor; v) the
current number of bytes of the page on current revision; vi) and an optional tag indicating
whether the revision is minor or major and should be reviewed by other editors. From
the data contained in each revision, the author applied its knowledge and intuition in
this domain to design a set of quantitative features believed by him to be useful for the
inference of trust3. Table 1 list this set of features associated to each editor (including
2 File ptwiki-20190201-stub-meta-history.xml, downloaded on 2 January 2019 from
https://dumps.wikimedia.org/.
3 Another human reasoner might have produced a different set of features, which could lead to
different assignments of trust.
anonymous ones identified by their IP). Some of these features such as presence,
regularity and frequency were first proposed in [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. A time window of 30 days was selected
for evaluation of the frequency and regularity factors, in line with the statistical
examination performed by the Wikimedia Foundation’s Analytics that also selects this time
window for some of its analysis of Wikipedia dumps. Designed features were in turn
employed for constructing two knowledge bases, as exemplified in the next sections.
Due to space limitations, these can be found in a public repository [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ].
      </p>
      <sec id="sec-3-1">
        <title>Layer 1 - Definition of the structure of arguments The first step of this argumentation</title>
        <p>
          process focuses on the construction of forecast arguments [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. Here, these are extended
in order to allow the manipulation of numerical inputs, similarly to the structure adopted
in fuzzy inference rules [
          <xref ref-type="bibr" rid="ref27">27</xref>
          ].
        </p>
        <p>Definition 1 (Forecast argument). A generic forecast argument arg is defined, without
loss of generalisability for AND and OR operators, as:
(i3</p>
        <p>arg: (i1 [l1; u1] AND i2 [l2; u2] ) OR
[l3; u3] AND i4 [l4; u4]) ! conclusion
[lc; uc]
Where in 2 R is the input value of the feature n with numerical range [ln 2 R, un 2 R];
the range [lc 2 R, uc 2 R] is the numerical range of the conclusion level being inferred,
or in this case the trust level; and AND and OR are boolean logical operators. This
structure includes a set of premises (believed to influence the conclusion being inferred)
and a conclusion derivable by applying an inference rule !. It is an uncertain
implication which is used to represent a defeasible argument. Premises and conclusions are
strictly bounded in numerical ranges. In order to facilitate the reasoning process, natural
language terms (for instance low and high) are also mapped to these numerical ranges.</p>
        <p>
          Both linguistic terms and numerical ranges are usually provided by the knowledge base
designer. In this paper, some of these values were defined based on the statistical
analysis of Wikipedia dumps provided by the Wikimedia Foundation’s Analytics, while
others were defined intuitively based on the author’s experience with digital
collaborative environments. Examples of forecast arguments using natural language terms and
their respective numerical ranges employed in this study are:
- arg1: low activity factor [
          <xref ref-type="bibr" rid="ref5">0, 5</xref>
          ] ! low trust [0, 0.25)
- arg2: high regularity factor [0.75, 1] ! high trust [0.75, 1]
- arg3: medium low pres. factor [0.25, 0.50) ! medium low trust[0.25, 0.50)
Layer 2 - Definition of the conflicts of arguments In order to evaluate
inconsistencies, the notion of mitigating argument [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] is introduced. These are arguments that
attack other forecast arguments or other mitigating arguments. Both forecast and
mitigating arguments are special defeasible rules, as defined in [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. Informally, if their
premises hold then presumably (defeasibly) their conclusions also hold. Different types
of attacks, and consequently, mitigating arguments, exist in the literature [
          <xref ref-type="bibr" rid="ref20 ref4">20, 4</xref>
          ]. In the
present study, three types are employed: undermining, undercutting and rebuttal attack.
Table 2 lists their definitions and examples. Note that the coexistence of arguments
inferring different conclusions might be possible according to some expert’s reasoning,
hence not all arguments with different conclusions lead to rebuttal attacks. The
computation of the acceptability status of arguments and final numerical scalar being produced
by such models is performed in the next layers. This computation is made via abstract
argumentation theory as proposed by [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. In this case, all attacks are seen as a binary
relation. All the designed arguments and attacks can now be seen as an argumentation
framework (AF) depicted in Fig. 2.
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Layer 3 - Evaluation of the conflicts of arguments At this stage an AF can be elicited</title>
        <p>with data. Forecast and mitigating arguments can be activated or discarded, based on
whether their premises evaluate true or false. Attacks between activated arguments will
be evaluated before being activated as well. As mentioned in the previous layer, attacks
Undercutting
Rebuttal</p>
        <sec id="sec-3-2-1">
          <title>A set of premises and an inference ) to</title>
          <p>an argument B (forecast or mitigating):
premises ) :B</p>
        </sec>
        <sec id="sec-3-2-2">
          <title>A bi-direction inference , between</title>
          <p>
            forecast arguments that support mutually
exclusive conclusions:
f orecast arg: , f orecast arg:
low frequency factor
AND low regular.
factor AND low activ.
fac. ) : arg3
arg2 , arg3
usually have a form of a binary relation. In a binary relation, a successful (activated)
attack occurs whenever both of its source (argument attacking) and its target
(argument being attacked) are activated. However, this study also makes use of the notion of
strength of arguments as presented in [
            <xref ref-type="bibr" rid="ref20">20</xref>
            ]. In this case, an attack is considered
successful only if the strength of its source is equal to, or greater than, the strength of its target.
To define the strength of an argument, feature weights are defined based on a pairwise
comparison between the 9 employed features (Table 1) performed by the knowledge
base designer. Hence, they will be numbers in the range [
            <xref ref-type="bibr" rid="ref8">0, 8</xref>
            ] N, being 0 if a feature
is considered less important than any other feature for the inference of computational
trust, and 8 if it is considered more important than any other feature. The weight of a
feature will also represent the strength of the argument employing this feature. These
weights can be seen in the full knowledge bases [
            <xref ref-type="bibr" rid="ref24">24</xref>
            ].
          </p>
          <p>
            Layer 4 - Definition of the acceptance status of arguments Given a set of activated
attacks and arguments, acceptability semantics [
            <xref ref-type="bibr" rid="ref6 ref9">9, 6</xref>
            ] are applied to compute the
acceptance status of each argument, that is, its acceptability. Extension-based and
rankingbased semantics are used to evaluate the overall interaction of arguments across the set,
in order to select the arguments that should ultimately be accepted. In this study, two
extension-based semantics (grounded and preferred [
            <xref ref-type="bibr" rid="ref9">9</xref>
            ]) and one ranking-based
semantics (categoriser [
            <xref ref-type="bibr" rid="ref6">6</xref>
            ]) are employed.
          </p>
          <p>
            Layer 5 - Accrual of acceptable arguments In the last step of the reasoning
process, a final inference must be produced. In the case of extension-based semantics, if
multiple extensions are computed, the cardinality of an extension (number of accepted
arguments) is used as a mechanism for the quantification of its credibility. Intuitively, a
larger extension of arguments might be seen as more credible than smaller extensions. If
the computed extensions have all the same cardinality, these are all brought forward in
the reasoning process. After the selection of the larger extension/s or best-ranked
argument/s, a single scalar is produced through the accrual of the values inferred by forecast
arguments. Mitigating arguments have already completed their role by contributing to
the resolution of conflicting information and thus are not considered in this layer. In
order to infer a crisp value at the end of the reasoning process, it is also necessary to
infer a crisp value by each accepted forecast argument. Following Definition 1, this is
done as proposed in [
            <xref ref-type="bibr" rid="ref22">22</xref>
            ]:
          </p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>Definition 2 (Crisp conclusion of a forecast argument). The crisp value of a conclu</title>
        <p>sion mapped to a numerical range [lc; uc] in a generic forecast argument arg (Definition
1) is given by the function:
f (arg) = Rmjuaxc Rlcmjin (v
v = min[max(i1; i2); max(i3; i4)]
Rmax) + uc, where Rmax = min[max(u1; u2); max(u3; u4)]</p>
        <p>Rmin = min[max(l1; l2); max(l3; l4)]</p>
        <sec id="sec-3-3-1">
          <title>Three cases are possible depending on the values of lc and uc:</title>
          <p>1. lc &lt; uc: the higher the value of the premises of arg, the higher the value of f (arg).
2. lc &gt; uc: the higher the value of the premises of arg, the lower the value of f (arg).
3. lc = uc: arg already infers a crisp value and Definition 2 is not necessary.</p>
          <p>Finally, the accrual of the crisp values inferred by forecast arguments will result in
the trust value inferred by this reasoning process. This accrual can be made in different
ways, for instance considering measures of central tendency. The average is accounted
in this study for models that use a binary relation of attacks, while the weighted average
is accounted for models that use the notion of strengths of arguments. Note that in the
case of two preferred extensions with the same number of accepted forecast arguments,
the outcome of the preferred semantics is the mean of its two extensions.</p>
          <p>
            Table 3 summarises the design of the argument-based models with different
parameters for each of their layers. Let us point out that the literature of defeasible
argumentation is vast [
            <xref ref-type="bibr" rid="ref17 ref7">7, 17</xref>
            ], and allows for many other configurations. Hence, we do not propose
an optimal set of models. Instead, we have borrowed well known parameters that are
believed to be enough for an initial account of the proposed assessment of trust and
adequate for the knowledge bases in hand.
The data extracted from a Portuguese Wikipedia dump was used to elicit the designed
argument-based models (Table 3) and baseline instruments. The inferences produced by
them were employed for the evaluation of the rank and spread of trust values assigned to
Barnstar editors. Table 4 lists the procedure of calculation of each metric, while Figure
3 depicts the respective results.
          </p>
          <p>s
tr
a
s
ran40
B
fo20
k
an 0</p>
          <p>R</p>
          <p>KB1</p>
          <p>KB2</p>
          <p>Baselines</p>
          <p>Fig. 3a depicts the resulting normalised sum of Barnstar ranks, indicating if
Barnstar editors were ranked at the highest positions or not. It is possible to observe that
the computed ranks by argument-based models were effective, ranging from 0.87 to
12.95. This suggests that defeasible argumentation was capable of capturing, to some
degree, the notions of the ill-defined construct of trust. In contrast, the baseline
instruments (average and weighted average), presented poor performance (ranks equal to
21.1 and 40.4 respectively). It implied that the reasoning performed by argument-based
models was able to greatly improve the use of the selected features for the ranking of
Barnstars. Among the argument-based models, the inferences produced by those built
with KB1 (Af1-6g) did not result in a rank of Barnstar significantly different. A
possible reason might be the simplified topology of KB1 (Figure 2a). In contrast, note that
the inferences produced by models built with KB2 resulted in a higher variance of the
rank of Barnstar (1.102 - 12.955), with model A9 not being reported due to the higher
number of cases with no inference (52.33%). It implies that, as expected, acceptability
semantics are more significant when employed over AFs of greater topological
complexity. The other metric evaluated, the spread of the trust values assigned to Barnstar
editors, was measured through the standard deviation (s ) of these values. Figure 3b
depicts the results for this metric. The inferences of models built with KB1 (Af1-6g) had
low variance and robust results. The inferences of models built with KB2 (Af7-12g)
had higher variance, with the best results achieved by the categoriser and preferred
semantics with no strength of arguments (A7 and A8). In comparison to the baseline
instruments, argument-based models achieved better results except when built with KB2
and strength of arguments (Af10-12g). It might be argued that a single set of strengths
was selected for all the exploited data, thus it is not adequate for case-by-case reasoning.</p>
          <p>Figure 3c reports the sum of the ranks achieved by each model for each metric of
evaluation. While relative differences are lost when models are ranked, this sum still
provides a general account on their performance. Argument-based models seem to
confirm the likely superior inferential capacity of defeasible argumentation compared to
the selected baselines. In particular, models A4 and A6 built with KB1 presented the
best general solutions. The exception comes from model A9 (grounded semantics, no
strength of arguments, and KB2). This configuration of parameters led to a high number
of cases with no inference (52.33%). The grounded semantics, with no strength of
arguments, is a sceptical approach, and, as expected, likely unable to solve a high number
of rebuttals. In summary, the use of defeasible argumentation for the inference of trust
of the Wikipedia editors could be seen as more appealing than the compared baseline
instruments. Such instruments do not take into account possible conflicts among the
selected pieces of evidence. Thus, the results of this study indicate that the assumption of
assigning a numerical trust value to Wikipedia editors as a form of defeasible reasoning
process is likely valid. Hence, it is a promising reasoning technique because it offers a
flexible approach for translating different knowledge bases and beliefs of human
reasoners into computational rules. Moreover, it allows the creation of models that can be
extended, falsified, and replicated, supporting the enhancement of the understanding
of computational trust itself. These advantages are observed also against data-driven
techniques, even the ones able to produce interpretable solutions such as decision-trees.
5</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusions and Future Work</title>
      <p>
        This study presented an empirical evaluation of defeasible argumentation for the
inference of computational trust in the context of the Wikipedia project. It employed two
knowledge bases formed by computational rules and grounded on the domain
knowledge of a human reasoner. A primary research has been conducted including the
construction of inferential models using defeasible argumentation. These were employed to
represent the reasoning applied to assess and infer the trust of Wikipedia editors as a
numerical metric. Moreover, they were elicited with real-world, quantitative data provided
by publicly available Wikipedia dumps. The output of these models were scalars
representing a trust value assigned to each editor in the range [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ] R. The selected metrics
for the evaluation of their inferential capacity were the spread and rank of trust values
assigned to editors recognised as trustworthy by the Wikipedia community. Findings
indicated that models built with defeasible argumentation outperformed non-deductive
calculations in both metrics. Therefore, the assessment of computational trust as a form
of defeasible reasoning process is presumably plausible. Thus, this research contributes
to the field of defeasible argumentation by exemplifying a practical use of this reasoning
approach seldom reported in the literature. This use is done via a modular design which
is suitable for modelling reasoning applied to distinct real-world domains. For instance,
previous works have employed this design for the inference of other phenomena, such
as human mental workload [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] and risk of mortality in elderly individuals [
        <xref ref-type="bibr" rid="ref25 ref26">25, 26</xref>
        ].
Therefore, the results presented here reinforce the generalisability of defeasible
argumentation for knowledge representation and production of quantitative inferences in
distinct domains characterized by uncertain and conflicting evidence. Future work will
concentrate on replicating this experiment by considering other reasoning approaches,
such as fuzzy reasoning and expert systems, and taking into account knowledge bases
built by multiple reasoners and/or including human-in-the-loop alternatives [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] for the
automation of the creation of arguments and attacks.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Adler</surname>
          </string-name>
          , B.T.,
          <string-name>
            <surname>de Alfaro</surname>
          </string-name>
          , L.:
          <article-title>A content-driven reputation system for the wikipedia</article-title>
          .
          <source>In: Proceedings of the 16th Int. Conf on World Wide Web</source>
          . pp.
          <fpage>261</fpage>
          -
          <lpage>270</lpage>
          . WWW '07,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Adler</surname>
            ,
            <given-names>B.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chatterjee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>de Alfaro</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Faella</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pye</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Raman</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Assigning trust to wikipedia content</article-title>
          .
          <source>In: Proc. of the 4th Int. Symposium on Wikis</source>
          . pp.
          <volume>26</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>26</lpage>
          :
          <fpage>12</fpage>
          . WikiSym '08,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Amgoud</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Demolombe</surname>
            ,
            <given-names>R.:</given-names>
          </string-name>
          <article-title>An argumentation-based approach for reasoning about trust in information sources</article-title>
          .
          <source>Argument &amp; Computation</source>
          <volume>5</volume>
          (
          <issue>2-3</issue>
          ),
          <fpage>191</fpage>
          -
          <lpage>215</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Amgoud</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vesic</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Rich preference-based argumentation frameworks</article-title>
          .
          <source>International Journal of Approximate Reasoning</source>
          <volume>55</volume>
          (
          <issue>2</issue>
          ),
          <fpage>585</fpage>
          -
          <lpage>606</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Barakat</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bradley</surname>
            ,
            <given-names>A.P.</given-names>
          </string-name>
          :
          <article-title>Rule extraction from support vector machines: A review</article-title>
          .
          <source>Neurocomputing</source>
          <volume>74</volume>
          (
          <issue>1</issue>
          ),
          <fpage>178</fpage>
          -
          <lpage>190</lpage>
          (
          <year>2010</year>
          ), artificial Brains
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Besnard</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hunter</surname>
            ,
            <given-names>A.:</given-names>
          </string-name>
          <article-title>A logic-based theory of deductive arguments</article-title>
          .
          <source>Artificial Intelligence</source>
          <volume>128</volume>
          (
          <issue>1-2</issue>
          ),
          <fpage>203</fpage>
          -
          <lpage>235</lpage>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Bryant</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krause</surname>
            ,
            <given-names>P.:</given-names>
          </string-name>
          <article-title>A review of current defeasible reasoning implementations</article-title>
          .
          <source>The Knowledge Engineering Review</source>
          <volume>23</volume>
          (
          <issue>3</issue>
          ),
          <fpage>227</fpage>
          -
          <lpage>260</lpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Dondio</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Longo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Computing trust as a form of presumptive reasoning</article-title>
          . In:
          <article-title>Web Intelligence (WI) and Intelligent Agent Technologies (IAT)</article-title>
          ,
          <source>IEEE/WIC/ACM Int. Joint Conf. on. vol. 2</source>
          , pp.
          <fpage>274</fpage>
          -
          <lpage>281</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Dung</surname>
            ,
            <given-names>P.M.</given-names>
          </string-name>
          :
          <article-title>On the acceptability of arguments and its fundamental role in nonmonotonic reasoning, logic programming and n-person games</article-title>
          .
          <source>Artificial intelligence</source>
          <volume>77</volume>
          (
          <issue>2</issue>
          ),
          <fpage>321</fpage>
          -
          <lpage>358</lpage>
          (
          <year>1995</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Javanmardi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lopes</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baldi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Modeling user reputation in wikis</article-title>
          .
          <source>Statistical Analysis and Data Mining: The ASA Data Science Journal</source>
          <volume>3</volume>
          (
          <issue>2</issue>
          ),
          <fpage>126</fpage>
          -
          <lpage>139</lpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Krupa</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vercouter</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , Hu¨bner,
          <string-name>
            <given-names>J.F.</given-names>
            ,
            <surname>Herzig</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          :
          <article-title>Trust based evaluation of wikipedia's contributors</article-title>
          . In: Aldewereld,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Dignum</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            ,
            <surname>Picard</surname>
          </string-name>
          ,
          <string-name>
            <surname>G</surname>
          </string-name>
          . (eds.) Engineering Societies in the Agents World X. pp.
          <fpage>148</fpage>
          -
          <lpage>161</lpage>
          . Springer Berlin Heidelberg (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Longo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Argumentation for knowledge representation, conflict resolution, defeasible inference and its integration with machine learning</article-title>
          .
          <source>In: Machine Learning for Health Informatics</source>
          . pp.
          <fpage>183</fpage>
          -
          <lpage>208</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Longo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dondio</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barrett</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Temporal factors to evaluate trustworthiness of virtual identities</article-title>
          .
          <source>In: 2007 Third International Conference on Security and Privacy in Communications Networks and the Workshops-SecureComm</source>
          <year>2007</year>
          . pp.
          <fpage>11</fpage>
          -
          <lpage>19</lpage>
          . IEEE (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Longo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rizzo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dondio</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Examining the modelling capabilities of defeasible argumentation and non-monotonic fuzzy reasoning</article-title>
          .
          <source>Knowledge-Based Systems</source>
          p. (in Press) (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Marsh</surname>
            ,
            <given-names>S.P.</given-names>
          </string-name>
          :
          <article-title>Formalizing Trust as a Computational Concept</article-title>
          .
          <source>Ph.d. thesis</source>
          , University of Stirling, Department of Computer Science and Mathematics (
          <year>1994</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Matt</surname>
            ,
            <given-names>P.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Morgem</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toni</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Combining statistics and arguments to compute trust</article-title>
          .
          <source>In: 9th International Conference on Autonomous Agents and Multiagent Systems</source>
          , Toronto. vol.
          <volume>1</volume>
          , pp.
          <fpage>209</fpage>
          -
          <lpage>216</lpage>
          . ACM (May
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Modgil</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prakken</surname>
          </string-name>
          , H.:
          <article-title>A general account of argumentation with preferences</article-title>
          .
          <source>Artificial Intelligence</source>
          <volume>195</volume>
          ,
          <fpage>361</fpage>
          -
          <lpage>397</lpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Parsons</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Atkinson</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McBurney</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sklar</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Singh</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haigh</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Levitt</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rowe</surname>
          </string-name>
          , J.:
          <article-title>Argument schemes for reasoning about trust</article-title>
          .
          <source>Argument &amp; Computation</source>
          <volume>5</volume>
          (
          <issue>2-3</issue>
          ),
          <fpage>160</fpage>
          -
          <lpage>190</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Parsons</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McBurney</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sklar</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          :
          <article-title>Reasoning about trust using argumentation: A position paper</article-title>
          .
          <source>In: Workshop on Argumentation in Multi-Agent Systems</source>
          . pp.
          <fpage>159</fpage>
          -
          <lpage>170</lpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Pollock</surname>
            ,
            <given-names>J.L.</given-names>
          </string-name>
          :
          <article-title>Cognitive carpentry: A blueprint for how to build a person</article-title>
          . Mit Press (
          <year>1995</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Prakken</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sartor</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>The role of logic in computational models of legal argument: A critical survey</article-title>
          . In: Kakas,
          <string-name>
            <given-names>A.C.</given-names>
            ,
            <surname>Sadri</surname>
          </string-name>
          ,
          <string-name>
            <surname>F</surname>
          </string-name>
          . (eds.)
          <article-title>Computational Logic: Logic Programming and Beyond: Essays in Honour of Robert A</article-title>
          .
          <string-name>
            <surname>Kowalski Part</surname>
            <given-names>II</given-names>
          </string-name>
          , pp.
          <fpage>342</fpage>
          -
          <lpage>381</lpage>
          . Springer, Berlin, Heidelberg (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Rizzo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Evaluating the Impact of Defeasible Argumentation as a Modelling Technique for Reasoning under Uncertainty</article-title>
          .
          <source>Ph.D. thesis</source>
          , Technological University Dublin (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Rizzo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Longo</surname>
            ,
            <given-names>L.:</given-names>
          </string-name>
          <article-title>An empirical evaluation of the inferential capacity of defeasible argumentation, non-monotonic fuzzy reasoning and expert systems</article-title>
          .
          <source>Expert Systems with Applications</source>
          <volume>147</volume>
          , (in press) (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Rizzo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Longo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Structured knowledge bases for the inference of computational trust of Wikipedia editors (</article-title>
          <year>2020</year>
          , accessed
          <issue>May 5</issue>
          ,
          <year>2020</year>
          ), doi.org/10.6084/m9.figshare. 12249770
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Rizzo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Majnaric</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dondio</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Longo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>An investigation of argumentation theory for the prediction of survival in elderly using biomarkers</article-title>
          . In: Iliadis,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Maglogiannis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            ,
            <surname>Plagianakos</surname>
          </string-name>
          , V. (eds.)
          <source>Artificial Intelligence Applications and Innovations</source>
          . pp.
          <fpage>385</fpage>
          -
          <lpage>397</lpage>
          . Springer International Publishing,
          <string-name>
            <surname>Cham</surname>
          </string-name>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Rizzo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Majnaric</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Longo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>A comparative study of defeasible argumentation and non-monotonic fuzzy reasoning for elderly survival prediction using biomarkers</article-title>
          . In: Ghidini,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Magnini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Passerini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Traverso</surname>
          </string-name>
          , P. (eds.)
          <source>AI*IA 2018 - Advances in Artificial Intelligence</source>
          . pp.
          <fpage>197</fpage>
          -
          <lpage>209</lpage>
          . Springer International Publishing,
          <string-name>
            <surname>Cham</surname>
          </string-name>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Ross</surname>
          </string-name>
          , T.J.:
          <article-title>Fuzzy Logic with Engineering Applications</article-title>
          . New York:
          <string-name>
            <surname>McGraw-Hill</surname>
          </string-name>
          (
          <year>1995</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Tang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cai</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McBurney</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sklar</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parsons</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>Using argumentation to reason about trust and belief</article-title>
          .
          <source>Journal of Logic and Computation</source>
          <volume>22</volume>
          (
          <issue>5</issue>
          ),
          <fpage>979</fpage>
          -
          <lpage>1018</lpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Zeng</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alhossaini</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ding</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fikes</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McGuinness</surname>
            ,
            <given-names>D.L.</given-names>
          </string-name>
          :
          <article-title>Computing trust from revision history</article-title>
          .
          <source>In: Intl. Conf. on Privacy, Security and Trust</source>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>