<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Logic-Based Explainable Framework for Relation Classification of Human Rights Violations</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Bimal Bhattarai</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rupsa Saha</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ole-Christofer Granmo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vladimir Zadorozhny</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jiawei Xu</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Agder</institution>
          ,
          <addr-line>Grimstad</addr-line>
          ,
          <country country="NO">Norway</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Pittsburgh</institution>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <fpage>14</fpage>
      <lpage>21</lpage>
      <abstract>
        <p>Using a Relational Tsetlin Machine (RTM) for analysis of semi-structured data allows the use of inherent relational structures present in natural language text to get an explainable classification of data. A finite Herbrand model derives Horn Clauses from the model, which are simple yet powerful logical tools that can build an abstract view of the world. We use the same to analyze human rights violation data. We show concretely how natural language can be transformed into a relational structure, and further use the Relational Tsetlin Machine to not only classify incidents as serious and non-serious violations but also explore the patterns learned by the RTM in order to arrive that those decisions. Furthermore, the distilled Horn Clauses show a precise understanding of the concepts involved without the drawback of textual ambiguity.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Tsetlin Machine (TM)</kwd>
        <kwd>Explainable</kwd>
        <kwd>Natural Language Processing</kwd>
        <kwd>Relational TM</kwd>
        <kwd>Logic-based Reasoning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        tropy models, and similar statistical approaches, trained
on large amounts of data; (c) Identifying and matching
Training Artificial Intelligence ( AI) to answer natural surface-level patterns with templates for response
generlanguage questions is a crucial part of the quest for en- ation. Often hybrid approaches are preferred for higher
coding a human-equivalent understanding of the world performance. But most QA systems sufer from a lack of
in machines. Massive structured Knowledge Bases (KBs), ability to generalize, and either have restrictive use cases
such as Freebase [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] have been the cornerstone of such or require massive amounts of knowledge. The absence
eforts. A common challenge lies in the appropriate in- of explanations of decisions taken by models also makes
terpretation of language by AI agents, both for building it dificult to identify problem areas and ofer resolutions
the KBs themselves (from existing natural language re- or improvements. [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ].
sources), as well as to identify the information required Tsetlin machines (TMs) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] use propositional logic
and provided from questions. The need for abstraction structures to build human-readable reasoning patterns
from specific (and limited) examples to build concepts from data. TMs’s pattern recognition capabilities have
and rules about the world, in general, is another chal- been successfully demonstrated in natural language
unlenge. What is a standard inductive reasoning problem if derstanding [
        <xref ref-type="bibr" rid="ref10 ref5 ref6 ref7 ref8 ref9">5, 6, 7, 8, 9, 10</xref>
        ], though none have explored
the KB is completely consistent and error-free, becomes logical decision making. The propositional clauses
conextremely diferent when allowing for all the uncertainty, structed by a TM have high discriminative power and
indeterminacy, errors, exceptions, and conflicts that are constitute a global description of the task learnt [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
present in real-world data, more so when the data has Apart from maintaining accuracy comparable to
statebeen extracted and structured by AI even partially. of-the-art machine learning techniques, the method also
      </p>
      <p>
        Various methods currently in use for Question An- has provided a smaller memory footprint and faster
inferswering (QA) are broad: (a) Tokenization, POS tagging, ence than more traditional neural network-based
modparsing, and other linguistic approaches to derive a pre- els [
        <xref ref-type="bibr" rid="ref12 ref13">12, 13, 14, 15, 16</xref>
        ]. Furthermore, [17] shows that
cise query from natural language questions, which can TMs can be fault-tolerant, able to mask stuck-at faults.
further be deployed on a structured database; (b) Sup- However, although TMs can express any propositional
port Vector Machines, Bayesian Classifiers, Maximum En- formula by using a disjunctive normal form, first-order
logic is required to obtain the computing power
equiva21st International Workshop on Nonmonotonic Reasoning, lent to a universal Turing machine. The more recently
September 2–4, 2023, Rhodes, Greece proposed Relational Tsetlin Machine introduces a first
*$Cobrirmeaspl.obhnadtitnagraaiu@thuoiar..no (B. Bhattarai); rupsa.saha@uia.no order TM framework with Herbrand semantics, with an
(R. Saha); ole.granmo@uia.no (O. Granmo); vladimir@sis.pitt.edu eye towards QA applications [18].
(V. Zadorozhny); jix20@pitt.edu (J. Xu) In this paper, we aim to use the RTM model to approach
© 2023 Copyright for this paper by its authors. Use permitted under Creative Commons License categorization and QA on a Human Rights Violation
CPWrEooUrckReshdoinpgs IhStpN:/c1e6u1r3-w-0s.o7r3g ACttEribUutRion W4.0oInrtekrnsahtioonpal (PCCroBYce4.0e).dings (CEUR-WS.org)
dataset, showcasing explainability with non-recursive sum. For  = 1, the voting target is  , whereas for  = 0,
ifrst-order Horn clauses built from specific examples. the voting target is −  . Observe that when the vote
aggregate approaches the user-specified threshold, the
likelihood of reinforcing a sentence rapidly decreases to
2. Background zero. This guarantees that clauses are distributed evenly
throughout the frequently occurring patterns, rather than
2.1. Propositional Tsetlin Machine omitting some and focusing only on others. The TM
The TM, first proposed in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], is a revolutionary tech- makes use of both Type I and Type II feedback. Type I
nique to pattern classification, regression, and novelty feedback is intended to generate frequent patterns and
detection [
        <xref ref-type="bibr" rid="ref6 ref7">19, 20, 6, 7</xref>
        ]. The base unit of a TM is called Type II feedback is intended to strengthen the
discrima Tsetlin automaton (TA), which learns the best action inating capacity of the patterns (for details, see [19]).
from a set of available actions in its environment. A “reg- When  = 1, Type I feedback is supplied stochastically
ular” TM can also be referred to as a Propositional TM, to clauses with positive polarity; when  = 0, Type I
due to the nature of its input and output operations. feedback is given to clauses with negative polarity. Each
      </p>
      <p>In a “Propositional” TM, a team of base TAs collec- clause strengthens its TA, in turn, depending on the
foltively generates propositional formulas using conjunctive lowing criteria: (1) its output  (); (2) the TA’s action
clauses. Its input is of the vector form  = (1, . . . , ), — Include or Exclude; and (3) the value of the literal 
and the TM learns clauses that decide if an input is to be allocated to the TA. Type I feedback is governed by two
in class  = 0 or  = 1. The learnt clauses are conjunc- rules:
tive combinations of elements from a subset of a literal
set made of input features and their respective negations
¯ = ¬ = 1 − . Each clause can be represented as:
 () = ⋀︀∈  = ∏︀∈ .</p>
      <p>(1)
E.g., the clause  () = 1 ∧ 2 = 12 consists of the
literals  = {1, 2} and outputs 1 if 1 = 2 = 1.</p>
      <p>The number of clauses used is a user-defined parameter
. Each new configuration of clauses created is subjected
to feedback, which controls the distribution of frequently
occurring patterns, as well as increasing the
discriminating power of individual patterns. Of the total number of
clauses, half vote in favor of  = 1 i.e., positive
polarity clauses (+), whereas the other half vote in favor of
 = 0 i.e., negative polarity clauses (− ). Classification
is performed based on a majority vote using equation 2
and the unit step function: ˆ = () = 1 if  ≥ 0 else 0
(for details, see [19]).</p>
      <p>= ∑︀=/12 +() −
∑︀=/12 − ().</p>
      <p>(2)
• When  () = 1 and  = 1 the Include is
rewarded and Exclude is penalized with probability
− 1
 . This reinforcement is powerful (triggered
with high probability) and causes the clause to
recall and refine the pattern it detects in .1
• When  () = 0 or  = 0 the Include is
penalized and Exclude is rewarded with probability
1 . This reinforcement is weak (activated with
low probability) and coarsens infrequent patterns,
hence increasing their frequency.</p>
      <sec id="sec-1-1">
        <title>As mentioned before, the user-configurable parameter</title>
        <p>determines pattern frequency; a larger  results in fewer
patterns.</p>
        <p>When  = 0, Type II feedback is supplied
stochastically to sentences with positive polarity; when  = 1,
Type II feedback is given to clauses with negative
polarity. Whenever  () = 1 and  = 0, it penalizes
Exclude. Thus, this feedback generates literals for
diferentiating between ! = 0 and  = 1 by evaluating the
clause to 0 when confronted with its rival class.</p>
        <p>While this vanilla TM setup operates on propositional
input variables  = (1 . . . , ), to generate
propositional conjunctive clauses, the Relational Tsetlin
Machine (RTM), described next, processes relations to
generate Horn clauses.</p>
      </sec>
      <sec id="sec-1-2">
        <title>E.g. the XOR-relation can be encoded as the classifier</title>
        <p>^ =  (1¯2 + ¯12 − 12 − ¯1¯2). The TM leverages
a team of TA for learning, one TA per literals in . Each
TA performs one of two actions - Include or Exclude- and
determines whether to include the literal  assigned in its
clause. TM encompasses an online learning system that
processes one training example ,  at a time. The TA 3. Relational Tsetlin Machine
generates a new configuration of clauses 1+, . . . , −/2,
before computing a voting total . Following that, feed- The work in [18] introduced the RTM as an extension
back is distributed statistically to each TA team. The dif- to the vanilla TM, encoding relations found in natural
ference  between the clipped voting total  and language using a logic-based representation for the TM.
a user-defined voting threshold  determines the likeli- The notion of RTM is based on a logic program using a
hood that each TA team will get feedback. Take note that
the voting amount is normalized by clipping the voting
1Tisarkeepnlaocteedthbayt w1.hen true positives are boosted, the probability −  1</p>
        <p>Input Sentence
The 1992 Anti-terrorist Law suspended the
requirement that the police obtain warrants
in order to make arrest. Who is Perpetrator?</p>
        <p>Who is Victim?</p>
        <p>Dataset
Relation
Extraction</p>
        <p>Entity
Extraction</p>
        <p>Entity</p>
        <p>Generalization
Generate Feature</p>
        <p>Vectors
Train TM</p>
        <p>Extract Clauses
ifnite Herbrand model [21, 22]. The ability to characterize more compact representation than is possible with
learning using Horn clauses is particularly beneficial merely propositional clauses. To illustrate this, assume
since Horn clauses are both simple and strong enough to  = {1, 2, . . . , } be  variables representing the
describe any logical formula [22]. constants occurring in an observation (˜ , ˜). Here,  is</p>
        <p>The RTM is a three-step process based on mapping the the maximum number of distinct constants required for
learning problem to a pattern recognition problem using each observation (˜ , ˜), each needing its own variable.
vanilla TM: We map the atoms to propositional inputs to create a
- By mapping relations to propositional inputs, a method vanilla TM learning problem. That is, each propositional
for dealing with relations and constants is devised. We input  represents a unique atom with a specific
varibegin with Horn clauses without variables. Let a set able configuration:  ≡ ( 1 ,  2 , . . . ,   ), with
of constants  = {1, 2, . . . , } be finite and a set  being the arity of . As a result, the number of
conof relations  = {1, 2, . . . , } of arity  ≥ 1,  ∈ stants in  has no efect on the number of propositional
{1, 2, . . . , } form the finite Herbrand base HB = inputs  required to describe the problem. Rather than
{1(1, 2, . . . , 1 ), 1(2, 1, . . . , 1 ), that, this is determined by the number of variables in 
. . . , (1, 2, . . . ,  ), (2, 1, . . . ,  ), . . .} con- as well as the number of relations in . For a particular
sisting of all  ground atoms that can be expressed observation (˜ , ˜), we first replace the constants in ˜
using  and . Additionally, we have a logic program with variables, from left to right. Accordingly, the
corre with program rules defined as non-recursive Horn sponding constants in ˜ are also replaced with the same
clauses. Each Horn clause has the following form: variables. Remaining constants in ˜ are arbitrarily
replaced with additional variables. The propositional input
0 ← 1, 2, · · · , . (3) vector is regenerated.</p>
        <p>Here, ,  ∈ {0, . . . , }, is an atom (1, 2, . . . ,  ) - A convolution approach with a standard TM as
dewith variables 1, 2, . . . ,  , or its negation scribed in Section 2.1 handles a large number of
alterna¬(1, 2, . . . ,  ). The arity of  is denoted by . tive constant-to-variable mappings as a standard pattern
We map every atom in HB to a propositional input , recognition problem.
getting the propositional input vector  = (1, . . . , ) The RTM is summarized in Algorithm 3.
(cf. Section 2.1). Decoupling the constants, i.e. step 2 above, serves
- The horn clauses with variables are introduced to to generalize the clause-based relations, and to allow
decouple the TM from the constants, resulting in a quicker learning with less input. By seeking Horn clauses
Algorithm 1 Relational TM</p>
        <p>input Convolutional Tsetlin Machine TM, Example pool , Number of training rounds 
1: procedure Train(TM, , )
2: for  ← 1, . . . ,  do
3: (˜ , ˜) ← ObtainTrainingExample()
4: ′ ← ObtainConstants(˜)
5: (˜ ′, ˜′) ← VariablesReplaceConstants(˜ , ˜, ′)
6: ′′ ← ObtainConstants(˜ ′)
7:  ← GenerateVariablePermutations(˜ ′, ′′)
8: UpdateConvolutionalTM(TM, , ˜′)
9: end for
10: end procedure
that employ variables rather than constants, we can pri- words or phrases to specific parts of judgments and
asoritize atoms above variable configurations. If  is the pects using the following ontology:
largest number of unique constants involved in any
particular observation, the number of atoms is bound by
(), where  is the largest arity of relations in .</p>
        <p>Given the possibility of diferent methods to assign
variables to constants, the preceding technique may result
in duplicate rules. One may wind up with identical rules
with just a syntactic variation, i.e., the same rules
represented using diferent variable symbols. To prevent
the creation of unnecessary rules, the Relational TM
generates all feasible permutations of variable assignments.</p>
        <p>Finally, we conduct a convolution over the permutations
to process them.
• Perpetrator - what entity is being judged?
• JPr - judgment of presence or absence
• JQu - judgment of quantity or intensity
• JGT - judgment of giving or taking
• SAVP - what is the specific aspect being judged</p>
        <p>(protection/violation)
• GAHR - general aspect of human rights
• Victim - who is the victim of the action?
• Negation - words that negate the meaning of the
sentence (usually not)</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>4. Using RTM to Analyse Data on</title>
    </sec>
    <sec id="sec-3">
      <title>Human Rights Violations</title>
      <sec id="sec-3-1">
        <title>In this section, we present a case study for analyzing human rights violation data using the previously described Relational Tsetlin Machine. The data is derived from the PULSAR system [23].</title>
        <p>PULSAR (Parsing Unstructured Language into
Sentiment-Aspect Representations) aims at parsing
human rights reports into sentence-level judgments and
linking judgments to specific aspects of human rights.
It handles a large corpus of yearly reports from human
rights non-governmental organizations (HRNGOs), as
well as the State Department and Amnesty International
Annual Reports.</p>
        <p>The parser uses aspect-based sentiment analysis
(ABSA) to separate judgments (sentiments) from things
being judged (aspects). As an example, consider the
following sentence:
• I like (sentiment, judgment) status of human rights
in this country (aspect).</p>
      </sec>
      <sec id="sec-3-2">
        <title>PULSAR refines this approach by mapping specific</title>
      </sec>
      <sec id="sec-3-3">
        <title>Using the above ontology PULSAR produces a series of judgments about human rights violations/protections, e.g.:</title>
        <p>• Security forces (Perpetrator) are (JPr) regularly
(JQu) participating (JGT) in the abducting (SAVP,
GAHR) of minorities (Victim).</p>
        <p>We use PULSAR output to produce either a ‘suppress’
or a ‘support’ relation between two entities, followed
by a query. The entities are the subject and the object
of the relation present in the sentence. In the case of a
‘suppress’ relation, Entity A and B are ‘Perpetrator’ and
‘Victim’. In a ‘support’ relation, they are ‘Supporter’ and
‘Beneficiary’. Based on this relation, one needs to identify
whether a violation of rights has occurred.</p>
        <p>To assist us in constructing the task, we make the
following assumptions:
1. All statements only include information about
relation “Support” or “Suppress”.
2. All questions are limited to information on the
relation of subject and object.
3. Both “Support” and “Suppress” each entail two
entities, such that
 / (,  ) :  ∈
{},  ∈ {}</p>
        <sec id="sec-3-3-1">
          <title>Sentence</title>
        </sec>
        <sec id="sec-3-3-2">
          <title>Sentence</title>
          <p>The 1992 Anti-terrorist Law suspended
the requirement that the police obtain Support
warrants in order to make arrest.</p>
          <p>Support (Supporter-Beneficiary)
The Government contends disciplinary
action against police who are guilty of Suppress
violating human rights.</p>
          <p>Suppress (Perpetrator-Victim)</p>
        </sec>
        <sec id="sec-3-3-3">
          <title>Relation</title>
        </sec>
        <sec id="sec-3-3-4">
          <title>Relation</title>
        </sec>
        <sec id="sec-3-3-5">
          <title>EntityA (Subj.)</title>
          <p>Law</p>
        </sec>
        <sec id="sec-3-3-6">
          <title>EntityA (Subj.)</title>
          <p>Govt.</p>
        </sec>
        <sec id="sec-3-3-7">
          <title>EntityB (Obj.)</title>
          <p>Police</p>
        </sec>
        <sec id="sec-3-3-8">
          <title>EntityB (Obj.)</title>
          <p>Police</p>
          <p>The first step is to convert the text into a
machineunderstandable relational representation. A relation in The relational factor is given by the diference in variance
this context refers to a relationship between two (or more)  =  − , resulting in:
elements of a text. Once identified, the entities may be
tgheenereralalitzioednsfo(wrfhuertthheerr rreedduuccetdiovniaofgseenaerrcahliszpaaticoen. Foirnnalolyt), RelA = ⎧⎪⎨ iiff  &gt;&lt; 00,, (5)
are used as input features for a standard TM setup for ⎪⎩  if  = 0.
the categorization of the text.</p>
          <p>Relation Extraction: Our text is composed of simple Entity Extraction: After identifying the relations, we
sentences, each of which has a sentence containing only must determine the textual components that contribute
one relation. The relations found in the query are deter- to the formation of those relationships. This enables us
mined using other features of a dataset and are linguis- to enhance the representation with restrictions,
allowtically related. Table 1 illustrates examples of Relation ing the RTM to learn rules that most accurately reflect
Extraction on our dataset. Each statement is associated action and consequences in a logical manner. The
rewith the relation of either “Support" or “Suppress”, while trieved entities can be combined with the information
the query is associated with either the Subject or the about the external world knowledge to create a richer
Object. Using the query, we can extract the relation as representation. For example, the concept that the subject
well as identify the entity. and object in a ‘Suppress’ relation can also be termed
The relation between the entities are calculated based as ‘Perpetrator’ and ‘Victim’, is an example of external
knowledge. Notably, as per Fig. 1, it is not feasible to
on positive valence count () and negative valence begin answering the query until both Relation Extraction
cdoounneta(sshow).nTihneEaqsusiagtinomne4n: t of the relation () is and Entity Extraction have been completed. Additionally,
knowledge of the relation enables us to filter down the produced during learning. At the end of the training, the
potential entities that will successfully respond to the relations captured by the TM constitute a global picture
query. of the learning, i.e. what the model has learned in general.</p>
          <p>Entity Generalization: One disadvantage of the Additionally, the global picture can be seen as a
descriprelational representation is that as more sentences are tion of the task itself, as understood by the machine. We
analyzed, the number of potential relations grows ex- also have access to a local snapshot that is unique to each
ponentially. One strategy to limit the spread is to rel- input instance. This contains just the clauses that
deegate particular entities to a more generic identifica- scribe the relations associated with a particular instance.
tion. Consider the following two instances from Table We use instances from a human rights dataset to
1: “Text: The 1992 Anti-terrorist Law suspended the showcase our work. The dataset contains human
requirement that the police obtain warrants in order rights-related sentences. Each sentence constitutes a
to make an arrest. Q1: Who is Supporter? Q2: Who perpetrator, a victim, and a query related to them. The
is Beneficiary?" and “Text: The Government contends relation that exists between them is either support or
disciplinary action against police who are guilty of vi- suppress. Based on this relation we determine whether
olating human rights. Q1: Who is Perpetrator? Q2: the rights of the victim are violated and we put them in
Who is Victim?". Processing the texts as per the pre- the respective label of “severity” or “non-severity”. We
vious section, we end up with six distinct relations: Sup- show the clauses obtained for one such example chosen
port(Law, Police), Supporter(Law), Beneficiary(Police), from the dataset:
Suppress(Government, Police), Perpetrator(Government), Input: The Government maintained that this waiting
Victim(Police). However, in order to answer any of the period was necessary to determine whether a woman may
queries, we simply need the relations associated with still be carrying the child of her former spouse.
that query. As a result, we can reduce both sentences Output: no violation
to four relations: Support(subj1, obj1), Suppress(subj2,
obj2), Subj(subj1 / subj2), and Obj(obj1 / obj2). Prioriti- To assist us in constructing the task, we make the
zation requires that the entities included in the query following assumptions:
relation be generalized first. Any instances of those
entities in the relations before the query are substituted.</p>
          <p>The entities present in other relations are subsequently
substituted with variables irrelevant for answering the
query (Table 2).
4.1. Classification
Having reduced the text into a relational feature set, we
now proceed to the classification task. We train RTM
with 650 clauses, threshold = 600, and specificity = 25
for 200 epochs. The LSTM with 100 memory units, 100
embedding sizes, spatial dropout, and softmax activation
is used. The CNN with 32-dimensional embedding,
convolution layer with max pooling, and sigmoid activation
is used. Table 3 shows that the RTM outperforms LSTM,
CNN, and vanilla TM considerably over 10 independent
runs. The vanilla TM, due to its disregard for the
relationship between variables, exhibits poor performance as
it lacks the ability to generalize these relationships. Both
LSTM and CNN demonstrate comparable performance.</p>
          <p>However, the utilization of the Herbrand model and horn
clause in RTM enhances its robustness and generalization
capability, surpassing the performance of other methods
by approximately 20%.
4.2. Explainability
One of the major reasons for choosing TMs for this task
is the inherent explainable structure built into the clauses
1. All statement sentences include just information</p>
          <p>about the relation “Support” or “Suppress”.
2. All questions are limited to information on the</p>
          <p>relation between perpetrator and victim.
3. Relation “Support” entails two entities, such
that  / (,  ) :  ∈
{},  ∈ {}
4. Relation “Suppress” entails two entities, such
that  / (,  ) :  ∈
{},  ∈ {}
5. “Support” or “Suppress” is a time-bound relation,
its impact is overtaken by a subsequent
comparable action.</p>
        </sec>
      </sec>
      <sec id="sec-3-4">
        <title>After Relation and Entity Extraction:</title>
        <p>Input =&gt; Support(Govt., woman), Support(Govt.,
child), NotSuppress(Govt., woman), Not Suppress(Govt.,
child), Query (Subject), Query (Object). 2
Clauses Without Entity Generalization: The
complexity of tasks depends upon the collection of
sentences from which the model has to identify the
relations. For our experiment, we number each input
for two potential relations (i.e., support or suppress).
Based on this relation the classification is done into one</p>
      </sec>
      <sec id="sec-3-5">
        <title>2Due to the fact that the TM requires binary features, each input is</title>
        <p>transformed to a vector with each element representing the
existence (or lack) of the relationship instances.
of two labels (i.e., severe or non-severe). The initial stage (, ) ←
in preparing data for the TM is to reduce the input to (),  (),
relation-entity bindings. These bindings comprise the (, ).
feature set against which the TM is trained.</p>
        <p>Clauses are of the form: Support(Govt, woman) AND Characterizing learning using Horn clauses is
interSupport(Govt, child) AND NotSuppress(Govt, woman) esting because they are simple but powerful enough to
AND Query(Govt) AND Query(Child). describe any logical formula [22]. Multiple learned Horn
Clauses can form a deductive framework on a dataset.</p>
        <p>Clauses With Entity Generalization: We follow the
same procedure for making relation-entity binding for
relational input features as explained above. Following
that, we group the features by entity type to generalize
the information. After all, entities are replaced with
general placeholders, we train TM with binary features
having general entities as features.</p>
      </sec>
      <sec id="sec-3-6">
        <title>Performing Entity Generalization:</title>
        <p>Input =&gt; Support(Subj1, Obj1), Support(Subj2, Obj2),
NotSuppress(Subj1, Obj1), NotSuppress(Subj1, Obj2),
Q(Subj1), Q(Obj1 / Obj2).</p>
        <p>Clauses are of the form: Query(Subj1) AND
Query(Obj2) AND Support(Subj1, Obj1) AND
Support(Subj1, Obj2) AND NotSuppress(Subj1, Obj2) AND
NotSuppress(Subj1, Obj1).</p>
        <p>These clauses ofer a more compact view of the task
without distractions from unimportant constants.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5. Horn Clause Representation</title>
      <sec id="sec-4-1">
        <title>Horn clause representation of the above example is given in Table 4. Using generalization clauses 6 and 7 can be further replaced by the following single clause :</title>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>6. Conclusion</title>
      <sec id="sec-5-1">
        <title>A Relational Tsetlin Machine is used to reduce the text</title>
        <p>to relational input and classify human rights violations.
In many real-world datasets, especially those with
polarizing information such as human rights, it is imperative
to gain a deeper understanding of the classification logic
behind the actual classes. While the information can
hide behind ambiguous language, the framework must
have mechanisms to reduce the ambiguity as much as
possible, so that the resultant classification is easy and
straightforward to interpret. The use of RTM allows us
to obtain precise explanations in terms of Horn Clauses
that can form the basis of extended logical frameworks,
without compromising on accuracy.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>K.</given-names>
            <surname>Bollacker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Evans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Paritosh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Sturge</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Taylor</surname>
          </string-name>
          , Freebase:
          <article-title>a collaboratively created graph database for structuring human knowledge</article-title>
          ,
          <source>in: Proceedings of the 2008 ACM SIGMOD international conference on Management of data</source>
          ,
          <year>2008</year>
          , pp.
          <fpage>1247</fpage>
          -
          <lpage>1250</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M. A. C.</given-names>
            <surname>Soares</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. S.</given-names>
            <surname>Parreiras</surname>
          </string-name>
          ,
          <article-title>A literature review C. Granmo, From Arithmetic to Logic Based on question answering techniques, paradigms and AI: A Comparative Analysis of Neural Networks systems</article-title>
          ,
          <source>Journal of King</source>
          Saud University-Computer and Tsetlin Machine,
          <source>in: 27th IEEE International and Information Sciences</source>
          <volume>32</volume>
          (
          <year>2020</year>
          )
          <fpage>635</fpage>
          -
          <lpage>646</lpage>
          . Conference on Electronics Circuits and Systems
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Pundge</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Khillare</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. N.</given-names>
            <surname>Mahender</surname>
          </string-name>
          ,
          <source>Question (ICECS2020)</source>
          , IEEE,
          <year>2020</year>
          .
          <article-title>answering system, approaches</article-title>
          and techniques: a [14]
          <string-name>
            <surname>K. D. Abeyrathna</surname>
            ,
            <given-names>O.-C.</given-names>
          </string-name>
          <string-name>
            <surname>Granmo</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Goodwin</surname>
          </string-name>
          , Exreview,
          <source>International Journal of Computer</source>
          Applica- tending
          <source>the Tsetlin Machine With Integer-Weighted tions 141</source>
          (
          <year>2016</year>
          )
          <fpage>0975</fpage>
          -
          <lpage>8887</lpage>
          .
          <article-title>Clauses for Increased Interpretability</article-title>
          , IEEE Access
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>O</given-names>
            <surname>.-C. Granmo</surname>
          </string-name>
          ,
          <source>The tsetlin machine - a game 9</source>
          (
          <year>2021</year>
          ).
          <article-title>theoretic bandit driven approach to optimal pat-</article-title>
          [15]
          <string-name>
            <given-names>J.</given-names>
            <surname>Lei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Rahman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Shafik</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <article-title>Wheeldon, tern recognition with propositional logic, ArXiv A</article-title>
          .
          <string-name>
            <surname>Yakovlev</surname>
            ,
            <given-names>O.-C.</given-names>
          </string-name>
          <string-name>
            <surname>Granmo</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Kawsar</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          . Mathur, abs/
          <year>1804</year>
          .01508 (
          <year>2018</year>
          ).
          <article-title>Low-Power Audio Keyword Spotting using Tsetlin</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>R. K.</given-names>
            <surname>Yadav</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jiao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.-C.</given-names>
            <surname>Granmo</surname>
          </string-name>
          , M. Good- Machines, arXiv preprint arXiv:
          <volume>2101</volume>
          .11336 (
          <year>2021</year>
          ).
          <article-title>win, Human-level interpretable learning for aspect-</article-title>
          [16]
          <string-name>
            <surname>K. D. Abeyrathna</surname>
            ,
            <given-names>A. A. O.</given-names>
          </string-name>
          <string-name>
            <surname>Abouzeid</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <article-title>Bhattarai, based sentiment analysis</article-title>
          , in: The
          <string-name>
            <surname>Thirty-Fifth C. Giri</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Glimsdal</surname>
            ,
            <given-names>O.-C.</given-names>
          </string-name>
          <string-name>
            <surname>Granmo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Jiao</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Saha</surname>
          </string-name>
          , AAAI Conference on Artificial
          <string-name>
            <surname>Intelligence (AAAI- J. Sharma</surname>
            ,
            <given-names>S. A.</given-names>
          </string-name>
          <string-name>
            <surname>Tunheim</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Zhang</surname>
          </string-name>
          ,
          <source>Building con21)</source>
          ,
          <year>2021</year>
          .
          <article-title>cise logical patterns by constraining tsetlin machine</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>B.</given-names>
            <surname>Bhattarai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.-C.</given-names>
            <surname>Granmo</surname>
          </string-name>
          , L. Jiao,
          <article-title>Measuring clause size, in: INTERNATIONAL JOINT CONFERthe novelty of natural language text using the con- ENCE ON ARTIFICIAL INTELLIGENCE (IJCAI), junctive clauses of a tsetlin machine text classifier</article-title>
          ,
          <year>2023</year>
          . in: Proceedings of the 13th International Confer- [17]
          <string-name>
            <given-names>R.</given-names>
            <surname>Shafik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Wheeldon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Yakovlev</surname>
          </string-name>
          ,
          <source>Explainability ence on Agents and Artificial Intelligence -</source>
          Vol- and
          <source>Dependability Analysis of Learning Automata ume 2: ICAART„</source>
          <year>2021</year>
          , pp.
          <fpage>410</fpage>
          -
          <lpage>417</lpage>
          . doi:
          <volume>10</volume>
          .5220/ based AI Hardware,
          <source>in: IEEE 26th International 0010382204100417. Symposium on On-Line Testing and Robust System</source>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>B.</given-names>
            <surname>Bhattarai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.-C.</given-names>
            <surname>Granmo</surname>
          </string-name>
          , L. Jiao,
          <article-title>Word-level hu- Design (IOLTS)</article-title>
          , IEEE,
          <year>2020</year>
          .
          <article-title>man interpretable scoring mechanism for novel text</article-title>
          [18]
          <string-name>
            <given-names>R.</given-names>
            <surname>Saha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.-C.</given-names>
            <surname>Granmo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. I.</given-names>
            <surname>Zadorozhny</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Gooddetection using tsetlin machines, Applied Intelli- win, A relational tsetlin machine with applications gence (2022). to natural language understanding</article-title>
          , Journal of In-
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>B.</given-names>
            <surname>Bhattarai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.-C.</given-names>
            <surname>Granmo</surname>
          </string-name>
          , L. Jiao,
          <source>Explainable telligent Information Systems</source>
          (
          <year>2022</year>
          )
          <fpage>1</fpage>
          -
          <lpage>28</lpage>
          .
          <article-title>tsetlin machine framework for fake news detection [19]</article-title>
          <string-name>
            <surname>O.-C. Granmo</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Glimsdal</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Jiao</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>Goodwin, with credibility score assessment</article-title>
          ,
          <source>in: LREC</source>
          ,
          <year>2022</year>
          . C. W. Omlin,
          <string-name>
            <given-names>G. T.</given-names>
            <surname>Berge</surname>
          </string-name>
          , The convolutional tsetlin
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>R.</given-names>
            <surname>Saha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.-C.</given-names>
            <surname>Granmo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Goodwin</surname>
          </string-name>
          , Mining inter- machine, arXiv preprint arXiv:
          <year>1905</year>
          .
          <volume>09688</volume>
          (
          <year>2019</year>
          ).
          <article-title>pretable rules for sentiment and</article-title>
          semantic relation [20]
          <string-name>
            <surname>K. D. Abeyrathna</surname>
            ,
            <given-names>O.-C.</given-names>
          </string-name>
          <string-name>
            <surname>Granmo</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Zhang</surname>
          </string-name>
          , L. Jiao,
          <article-title>analysis using tsetlin machines</article-title>
          , in: International M.
          <article-title>Goodwin, The regression tsetlin machine: a Conference on Innovative Techniques and Appli- novel approach to interpretable nonlinear regrescations of</article-title>
          <source>Artificial Intelligence</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>67</fpage>
          -
          <lpage>78</lpage>
          . sion,
          <source>Philosophical Transactions of the Royal Sodoi:10.1007/978-3-030-63799-6\_5. ciety A</source>
          <volume>378</volume>
          (
          <year>2019</year>
          )
          <article-title>20190165</article-title>
          . doi:
          <volume>10</volume>
          .1098/rsta.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>B.</given-names>
            <surname>Bhattarai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.-C.</given-names>
            <surname>Granmo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jiao</surname>
          </string-name>
          , An interpretable
          <year>2019</year>
          .
          <volume>0165</volume>
          .
          <article-title>knowledge representation framework for natural</article-title>
          [21]
          <string-name>
            <given-names>J.</given-names>
            <surname>Lloyd</surname>
          </string-name>
          ,
          <article-title>Foundations of Logic Programming, language processing with cross-domain application</article-title>
          , Springer-Verlag, New York,
          <year>1984</year>
          . in: Advances in Information Retrieval: 45th Euro- [22]
          <string-name>
            <given-names>R.</given-names>
            <surname>Kowalski</surname>
          </string-name>
          ,
          <article-title>Logic programming</article-title>
          ,
          <source>in: Computapean Conference on Information Retrieval, ECIR tional Logic</source>
          , volume
          <volume>9</volume>
          ,
          <year>2014</year>
          , pp.
          <fpage>523</fpage>
          -
          <lpage>569</lpage>
          . doi:10.
          <year>2023</year>
          , Dublin, Ireland, April 2-
          <issue>6</issue>
          ,
          <year>2023</year>
          , Proceedings, 1016/B978-0
          <source>-444-51624-4</source>
          .
          <fpage>50012</fpage>
          -
          <lpage>5</lpage>
          .
          <string-name>
            <surname>Part</surname>
            <given-names>I</given-names>
          </string-name>
          ,
          <year>2023</year>
          , pp.
          <fpage>167</fpage>
          -
          <lpage>181</lpage>
          . [23]
          <string-name>
            <given-names>B.</given-names>
            <surname>Park</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Greene</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Colaresi</surname>
          </string-name>
          , How to teach ma-
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>C. D. Blakely</surname>
            ,
            <given-names>O.-C.</given-names>
          </string-name>
          <string-name>
            <surname>Granmo</surname>
          </string-name>
          ,
          <string-name>
            <surname>Closed-Form</surname>
          </string-name>
          Expres
          <article-title>- chines to read human rights reports and identify sions for Global and Local Interpretation of Tsetlin judgments at scale, Journal of Human Rights 19 Machines with Applications to Explaining High-</article-title>
          (
          <year>2020</year>
          )
          <fpage>99</fpage>
          -
          <lpage>116</lpage>
          . Dimensional Data, arXiv preprint arXiv:
          <year>2007</year>
          .
          <volume>13885</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>A.</given-names>
            <surname>Wheeldon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Shafik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Yakovlev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Edwards</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Haddadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.-C.</given-names>
            <surname>Granmo</surname>
          </string-name>
          , Tsetlin Machine:
          <article-title>A New Paradigm for Pervasive AI</article-title>
          , in: SCONA Workshop at Design,
          <article-title>Automation and Test in Europe (DATE</article-title>
          <year>2020</year>
          ),
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>J.</given-names>
            <surname>Lei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Wheeldon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Shafik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Yakovlev</surname>
          </string-name>
          , O.-
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>