<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Segmenting and Clustering Noisy Arguments</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Lorik Dumani</string-name>
          <email>dumani@uni-trier.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christin Katharina Kreutz</string-name>
          <email>kreutzch@uni-trier.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Manuel Biertz</string-name>
          <email>biertz@uni-trier.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alex Witry</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ralf Schenkel</string-name>
          <email>schenkel@uni-trier.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Trier University</institution>
          ,
          <addr-line>54286 Trier</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Automated argument retrieval for queries is desirable, e.g., as it helps in decision making or convincing others of certain actions. An argument consists of a claim supported or attacked by at least one premise. The claim describes a controversial viewpoint that should not be accepted without evidence given by premises. Premises are composed of Elementary Discourse Units (EDUs) which are their smallest contextual components. Oftentimes argument search engines find similar claims to a query first before returning their premises. Due to heterogeneous data sources, premises often appear repeatedly in diferent syntactic forms. From an information retrieval perspective, it is essential to rank premises relevant for a query claim highly in a duplicate-free manner. The main challenge in clustering them is to avoid redundancies as premises frequently address various aspects, i.e., consist of multiple EDUs. So, two tasks can be defined: segmentation of premises in EDUs and clustering of similar EDUs. In this paper we make two contributions: Our first contribution is the introduction of a noisy dataset with 480 premises for 30 queries crawled from debate portals which serves as a gold standard for the segmentation of premises into EDUs and the clustering of EDUs. Our second contribution consists of first baselines for the two mentioned tasks, for which we evaluated various methods. Our results show that an uncurated dataset is a major challenge and that clustering EDUs is only reasonable with premises as context information.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Computational argumentation is an important building block in decision
making applications. Retrieving supporting and opposing premises for controversial
claims can help to make informed decisions on the topic or, when seen from a
diferent viewpoint, to persuade others to take particular standpoints or even
actions. In line with existing work in this field, we consider arguments that consist
of a claim that is supported or attacked by at least one premise [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. The claim
is the central component of an argument, and it is usually controversial [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ].
The premises increase or decrease the claim’s acceptance [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. The stance of a
premise indicates if it supports (pro) or attacks (con) the claim. Table 1 shows an
example for an argument consisting of a claim supported or opposed by premises.
      </p>
      <p>
        In the NLP community researchers either address argument mining, i.e., the
analysis of the structure of arguments in natural language texts (see the work of
Cabrio and Villata [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] for an overview of recent contributions), or an
informationseeking perspective, i.e., the identification of relevant premises associated with
a predefined claim [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. Due to the rapidly increasing need for argumentative
queries, established search engines that only retrieve relevant documents will
no longer be suficient. Instead, argument search engines are required that can
provide the best pro and con premises for a query claim. In fact, various
argument search engines [
        <xref ref-type="bibr" rid="ref22 ref27">27,22</xref>
        ] have recently been developed. These systems usually
work on claims and premises that were either mined from texts beforehand or
extracted from dedicated argument websites such as idebate.org. Their workflow
usually starts with finding result claims similar to the query claim. Then they
locate the result premises belonging to these claims to present them as output.
      </p>
      <p>However, these systems face a number of challenges since claims and premises
are formulated in natural language. First, premises that are semantically (mostly)
equivalent occur repeatedly in diferent textual representations since they appear
in diferent sources, but should be retrieved only once to avoid duplicates. This
requires the clustering of similar premises for result presentation. Second,
discussions on debate portals, but also in natural language arguments are often
not well-structured, such that a single supporting or attacking piece of text can
address several aspects and thus should be represented as multiple premises. For
example, a sentence supporting the viewpoint that aviation fuel should be taxed
could address two aspects, the potential danger for the environment and the
current low tax rate on aviation fuel. Directly using such sentences as formal
premises, as seen in premise p3 in Table 1, would make it impossible to retrieve
a duplicate-free and complete list of premises.</p>
      <p>
        This issue can be avoided by dividing the premises into their core aspects and
clustering them instead of whole premises. In the literature, the smallest
contextual components of a text are called Elementary Discourse Units (EDUs) [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ].
Obtaining high quality EDUs [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] from text (discourse segmentation) is a crucial
task preceding all eforts in parsing or representing discourses [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. Thereby, it
takes a pragmatic perspective, i.e., links between discourse segments are
established not on semantic grounds but on the author’s (assumed) intention [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. For
the explorative purposes outlined here, only the concept of EDUs as smallest,
non-overlapping units of intra-text-discourse – mostly clauses – is picked up [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
      </p>
      <p>
        In this paper we address the aforementioned limitations and deal with the
segmentation of textual premises into EDUs and the clustering of EDUs based
on their semantic similarity. Contrasting previous research on both of these
tasks that worked with manually curated and thus high-quality argument
colEDU1(p1) = “Less CO2 emissions lead to a clean environment”
EDU1(p2) = “Higher taxes would not change anything”
EDU1(p3) = “It does not matter that the costs for aviation are already high”
EDU2(p3) = “as the environment can be protected by less CO2 emissions”
lections, we use a dataset that was crawled from debate portals [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Unlike
other datasets, the premises in this dataset contain a considerably higher
number of sentences and often cover multiple aspects (which is at odds with our
generally micro-structural approach to arguments). In addition, as an uncurated
real-world dataset, it contains many ill-formulated sentences and other defects.
Our contribution is two-fold: First we provide a real-life dataset consisting of
480 premises retrieved for 30 query claims that are segmented into 4,752 EDUs.
Then, for each query claim the belonging EDUs have been manually clustered
by semantic equivalence. Second, we report our first results for the two tasks of
EDU identification and EDU clustering on this dataset.
      </p>
      <p>Our proposed method works as follows: for a given set of textual premises
returned by an argument search engine for a query claim, we first identify the
EDUs for each result. In the second step, we focus on the clustering of EDUs. To
accomplish this, we first generate embeddings and then we cluster those with an
agglomerative clustering algorithm. As an example, consider Table 1 again. Here,
premise p3 is composed of two EDUs EDU1(p3) and EDU2(p3) (see Figure 1). In
addition to that, EDU2(p3) and EDU1(p1) (where EDU1(p1) is the only EDU of
p1) have the same meaning and therefore should be assigned to the same cluster.</p>
      <p>The remainder of this paper is structured as follows: Section 2 provides an
overview of related work addressing the segmentation of argumentative texts
into EDUs and clustering algorithms. In Section 3 the dataset and its manual
annotation is described in more detail. Then, in Section 4 we present and evaluate
our methods for extraction and clustering of EDUs. Section 5 concludes our work
and provides future research directions.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        There is a plethora of research on discourse segmentation of text but to the best
of our knowledge, existing approaches are designed for curated datasets. A
rulebased approach including a post processing step for identification of starts and
ends of EDUs was proposed by Carreras et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Among other features they
utilize chunks tags and sentence patterns. Soricut and Marcu [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] introduced a
probabilistic approach based on syntactic parse trees. Tofiloski et al. [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]
perform EDU segmentation based on syntactic and lexical features with the goal
of capturing only interesting, not all EDUs. Here, every EDU is conditioned to
contain a verb. Others suggest a classifier able to decide whether a word is the
beginning, middle or end of a nested EDU using features derived from Part of
Speech (POS) tags, chunk tags or dependencies to the root element [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. In a
recent paper, Trautmann et al. [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] also argue that “spans of tokens” rather than
whole sentences should be annotated and define this task as Argument Unit
Recognition and Classification. We omit preprocessing of text and utilization of
preconditions which is applicable to a supervised scenario as it might flaw an
approach based on uncurated data as no guarantees can be made for a real-world,
possibly defective, crawled dataset from debate portals.
      </p>
      <p>
        The clustering of similar arguments is still a recent field of research. Boltuzic
and Snajder [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] applied Word2Vec [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] with hierarchical clustering for debate
portals. Reimers et al. [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] experiment with contextualized word embedding
methods such as ELMo [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] and BERT [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and show that these can be used to
classify and cluster topic-dependent arguments. They use hierarchical clustering
with a stopping threshold which is determined on the training set to obtain
clusters of premises. However, they do not specify a concrete value. Further,
Reimers et al. note that premises sometimes cover diferent aspects. Hence, we
divide premises into their EDUs and cluster these instead. Like them, we also
use uncurated data and make use of ELMo and BERT. We additionally utilize
the embedding methods InferSent [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], and Flair [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Contrasting Reimers et
al., we only consider relevant premises for the clustering as we intend to start
with a step-by-step approach.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Dataset and Labeling</title>
      <p>
        We make use of the argumentation dataset introduced in our prior work [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]
where we crawled four debate portals and extracted claims with their associated
textual premises. In a follow-up work [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], we built a benchmark collection for
argument retrieval based on that dataset. In this former work, we picked 232
randomly chosen claims on the topic energy and used them as query claims to
pool the most similar result claims retrieved by standard IR methods. In the
latter [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], for 30 of these query claims, we collected the premises of all pooled
result claims and manually assessed their relevance with respect to the query
claim, using a three-fold scale (“very relevant”, “relevant”, “not relevant”). This
resulted in 1,195 tuples of the form (query claim, result claim, result premise,
assessment). Following the practice at TREC (Text REtrieval Conference), a
premise is relevant if it has at least one relevant EDU, and very relevant if it
contains no aspect not relevant to the initial query claim.
      </p>
      <p>
        In this paper, we only included result premises that were assessed with “very
relevant” or “relevant” to keep the efort for manual assessment reasonable. This
means we consider 480 tuples for our new dataset. For each of these 480 result
premises, the EDUs were identified by one annotator who is a research assistant
from political science and has a deep understanding of argumentation theory. For
this segmentation, the annotator followed the manual by Carlson and Marcu [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
This resulted in a total of 4,752 EDUs for the 480 premises (on average 9.9 EDUs
per premise), indicating that premises in debate portals usually cover plenty of
aspects and segmentation is indispensable for argument retrieval and clustering.
      </p>
      <p>
        In a next step, the EDUs were manually clustered by identifying
semantically equivalent EDUs and putting them in the same cluster. This was done with
support of a modified variant of the OVA tool [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] ( http://ova.uni-trier.de/) for
modeling complex argumentations, which was enhanced to be capable to store
text positions. Since EDUs cannot be further divided by definition, clusters were
formed manually that include all EDUs with the same meaning. For each of the
30 query claims, an OVA view was created where all EDUs identified in result
premises for this query were represented as nodes. A human annotator then
clustered these nodes by creating an artificial node for each cluster identified and
then connecting all semantically identical EDUs to the cluster node by
dragging edges. Additionally, to make the clustering more readable, the annotator
created three artificial clusters “PRO”, “CON”, and “CLAIMS” and referenced
the previously formed artificial clusters to them depending on their stance with
respect to the query. In this paper we will not consider stances. However, since
we are making the dataset available (on request), they can be important for
further work, for example, for those who also want to use additional distinctions
according to the stance.
      </p>
      <p>Figure 2 illustrates a screenshot of the clustering annotation tool. Not all
EDUs could plausibly be treated as a single premise (e.g., EDUs that are
postmodifiers to noun phrases), thus we also allowed to mark EDUs as context
information for other EDUs. For the clustering task, we clustered 1,044 EDUs for
11 queries, distributed to 622 clusters. Because of time constraints, we did not
manage to cluster all EDUs of all 30 queries here, and instead only analyzed 11,
which are after all more than 1,000 clustered EDUs. The annotators’ feedback
was that the visualization helped to keep an overview as there were almost 100
EDUs per query to cluster.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Methodology and Evaluation</title>
      <p>
        This section describes our approaches for segmenting premises into EDUs and
clustering them, as well as an evaluation of the performance of these methods
with respect to the ground truth. Figure 3 provides a schematic overview of
the diferent steps. In general, our approach will retrieve clusters of EDUs for
input query claims. Given a query claim qi as well as similar result claims ci;j
with associated premises pi;j;k. These relevant result claims are retrieved by
application of our prior work [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. In the first step of this approach, premises
are divided into EDUs, in the second step all EDUs of premises linked to result
claims for our query claim will be clustered.
4.1
      </p>
      <sec id="sec-4-1">
        <title>Step 1: Segmentation of Premises into EDUs</title>
        <p>
          We first compare diferent approaches for segmentation of premises into EDUs
to the ground truth segmentation from the 30 claims. We focus on basic
segmentation methods generating sequential, i.e., non-overlapping EDUs in order
to obtain insight into their performance on a real-world dataset as they are often
used as a preprocessing step in more sophisticated segmenters [
          <xref ref-type="bibr" rid="ref1 ref21 ref25 ref6">6,21,25,1</xref>
          ].
        </p>
        <p>
          As an initial baseline (sentence baseline), we split premises into sentences
with CoreNLP (stanfordnlp.github.io/CoreNLP) and considered each sentence
as an EDU. As CoreNLP also allows to extract a text’s PennTree, which
contains the POS tag for each term and displays the closeness of the terms in a
hierarchical structure, we also identified EDUs by cutting the PennTree of premises
(tree cut) at height cutofs from 1 to 10, denoted by tci;1 i 10 in the following.
Additionally, we obtained subclauses from sentences which we also regarded as
EDUs by applying Tregex (nlp.stanford.edu/software/tregex) (subclauses). We
also implemented a rule-based splitter (splitter ) which does consider the
peculiarities of our dataset but difers from the ground truth [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. This splitter is kind
of an extension of the sentence baseline, thus sentence boundaries and all kinds
of punctuation marks are seen as discourse boundaries [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] so those are used
to split premises into EDUs. Further, before conjunctions and terms or phrases
indicating subclauses, boundaries are included.
        </p>
        <p>
          Table 2 shows the performance of the diferent approaches which we compare
in terms of their precision, recall, F1 score, specificity, and accuracy. Out of the
tree cut approaches, tc3, i.e., the tree cut with cutof at height 3, obtained the
highest F1 score. The rule-based splitter achieved the highest F1 score of all
tested methods. This method results in a high precision combined with a lower
recall which is a property of conservative approaches [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ]. The comparably poor
results for the other approaches may occur since classical preprocessing steps are
unfit for approximating human annotations on uncurated real-world datasets. A
Kruskal-Wallis test1 on F1 scores for boundaries of EDUs computed for each
premise of the sentence baseline, splitter and tc3 holds for p = 0:05. Thus, the
splitter method is significantly better than the other methods.
        </p>
        <p>Evaluation of EDUs In order to evaluate the quality of EDUs obtained by
the annotators as well as our best approaches we constructed triples of the
form (EDUground_truth, EDUtc3 , EDUsplitter) for 50 randomly chosen premises.
Within each triple, the EDUs were ranked by their subjective perceived quality
by a reviewer who is an expert in computational argumentation and familiar
with argumentation theory. Note that it was not shown to the assessor how each
EDU was determined and the ordering within triples was shufled. The expert
assessor assigned ranks from 1 to 3 with 1 being the best, ties were permitted.</p>
        <p>The ground truth achieved an average rank of 1.66 (#1: 22 times, #2: 23
times, #3: 5 times), tc3 did perform equally well (#1: 23 times, #2: 21 times, #3:
6 times). The splitter method performed considerably worse with an average rank
of 2.64 (#1: 6 times, #2: 6 times, #3: 38 times). As the ground truth would be
expected to outperform other approaches clearly, this outcome indicates firstly
the dificulty in the annotation process, secondly the subjective perception of
what is better and what is less good, and thirdly the dificulty in correctly
capturing language with computers. Figure 4 shows an example of both the
manually created EDUs and those created by the splitter method.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Step 2: Clustering of EDUs</title>
        <p>In order to build clusters of EDUs automatically for each of the eleven claims,
ifrst we obtained the embedding vectors of EDUs using ELMo, BERT, Flair,
and InferSent2. For this task, we consider the segmentation of premises into
1 A Kruskal-Wallis test was used as data in the three groups is not normally
distributed; this was tested with a Shapiro-Wilk test.
2 We used the implementations provided by https://github.com/facebookresearch/</p>
        <p>InferSent and https://github.com/flairNLP/flair.</p>
        <p>EDUs by ground truth:
[From what I understand,] [the cheap oil is something] [that will not only effect the economy
in the long run,] [but it will also hurt those] [who want to receive retirement or disability
benefits at the federal level.] [It's great to finally have cheaper gas than that] [which was
nearly $3 in the past.] [I do think] [it might have an adverse effect on our economy.]
EDUs by splitter:
[From what I understand,] [the cheap oil is something] [that will not only effect the
economy in the long run,] [but it will also hurt those] [who want] [to receive retirement] [or
disability benefits at the federal level.] [It's great] [to finally have cheaper
gas] [than] [that] [which was nearly $3 in the past.] [I do think it might have an adverse
effect on our economy.]</p>
        <p>
          EDUs given by the ground truth of Section 4.1. Otherwise, an automatic
external evaluation would be infeasible. We derived eight vectors per EDU and
embedding technique by extending EDUs with context information, i.e., we
obtained tuples (EDU; ctx) with context ctx from all combinations of the power
set P(fpremise; result claim; query claimg). After that, we performed an
agglomerative (hierarchical) clustering of the EDUs of all claims related to the
query for each of the eleven queries as it is the state-of-the-art for clustering
arguments [
          <xref ref-type="bibr" rid="ref19 ref3">3,19</xref>
          ]. Then, since we do not know the number of clusters a priori,
we performed a dynamic tree cut [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. The advantage of this approach over other
approaches such as k-means is that there is no need to specify a final number
of result clusters, which is not known in our case. The benefit of agglomerative
clustering over divisible clustering is certainly the lower runtime. As a
straightforward baseline, all EDUs from the same premise are assigned to the same
cluster (BLpremiseAsCluster). Two additional baselines consist of one big cluster
containing all EDUs (BLoneCluster), as well as many clusters, each containing one
EDU (BLownClusters). The quality of the clustering was measured with external
and internal evaluation measures. While external evaluation measures base on
previous knowledge, in our case the ground truth clustering formed by the
assessor, the internal evaluation measures base on information that only involves
the vectors of the datasets themselves [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ].
        </p>
        <p>With regard to the external cluster evaluation metrics, we measured the
following three: the purity, the adjusted mutual information (AMI), and the
adjusted Rand index (Rand). For the internal cluster evaluation, we measured the
Calinski-Harabasz index (CHI) and the Davies-Bouldin index (DBI).3 Concise
descriptions of the metrics can be found in Table 3. The results of the external
and internal evaluations can be found in Table 4.</p>
        <p>We can observe that BLpremiseAsCluster outperforms all methods for the
external evaluation measures except for the perfect purity of BLownClusters. In
general BLoneCluster and BLownClusters do not produce surprising results for the
external evaluation. CHI and DBI are undefined for their number of clusters.
3 We used the implementations provided by https://scikit-learn.org/stable/modules/
classes.html#module-sklearn.metrics.cluster.</p>
        <p>It is remarkable that the methods that perform best have the corresponding
premises as additional context information, while the worst performing methods
do not utilize them. In fact, for each evaluation measure, all 16 methods that
include the premise as context information achieve better values than those 16
that do not include it. The inclusion of a query claim and result claim as context
information seems to have no influence on the ranking, because the methods with
and without usage of this context information in the ranking are sometimes
better and sometimes worse. Thus, clustering EDUs always requires the context
information in the premise. Kruskal-Wallis tests with p = 0:05 were conducted
on the three external and two internal measures of the eleven query claims of the
three best performing methods as well as the baseline.4 For AMI, CHI and DBI
significant diferences were found. For purity and Rand, no significant diferences
could be found between the four groups.</p>
        <p>For the internal cluster evaluations all methods that include the premise
as context information produce better outcomes than those computed for the
baseline BLpremiseAsCluster clusters. The best values were achieved when using
EDUs computed with ELMo or BERT embeddings. This observation clearly
shows challenges in automatic clustering of arguments in dificult datasets. We
conducted Mann-Whitney U tests on the five measures from the eleven
clusterings for each of the three best methods and their counterpart without utilization
of the premise as context-information (e.g. ELMoe;p;q and ELMoe;q were
observed as a pair) with p = 0:05.5 We found significant diferences in values for
purity, AMI, Rand, CHI and DBI for ELMo as well as InferSent; for the two
experiments with BERT embeddings, significant diferences were found for all
4 Kruskal-Wallis tests were used as except for purity, data is not normally distributed
in the four groups; this was tested with Shapiro-Wilk tests.
5 Mann-Whitney U tests were used as for all pairs, some of the measures are not
normally distributed; this was tested with Shapiro-Wilk tests.
external measures and CHI. From this observation we derive the usefulness of
premises as context information for the overall clustering quality.
Error Analysis of the Clustering We performed an additional manual
evaluation of the clustering by including the three best performing methods shown in
Table 4, as well as the initial manual clustering. For this evaluation we randomly
picked 30 clusters which contain at least three EDUs per cluster (120 clusters
in total) and added a new EDU to each of them, which two human
annotators (diferent from the one who constructed the ground truth in Section 3) had
to spot to determine the perceived soundness of the clustering. This new EDU
originated from the same premise or, if no EDU was available there, from the
same query. For each cluster at most five EDUs were shown. They were shufled
and the new EDU was placed at a random position. Additionally, we include
a random baseline. Here, for each of the 120 evaluation clusters, the intruding
EDU was picked at random.</p>
        <p>
          Only the query and the EDUs were presented to the annotators. For the
manually labeled clusters, both annotators managed to identify 16 out of 30
false EDUs. For InferSente;p, BERTe;p;r, and ELMoe;p;q, it was 11.5, 8.5,
and 7 out of 30 on average, respectively.6 The inter-annotator agreement,
calculated with Krippendorf’s [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ], was 0.463 on a nominal scale, implying that the
agreement is moderate. The random baseline picked a total of 9, 4, and 6 wrong
EDUs. We found no significant diferences with Kruskal-Wallis tests for
clustering based on BERT, Elmo and InferSent embeddings for the number of correctly
identified intruding EDUs by the two annotators and the random baseline. Yet,
for the ground truth, significant diferences were found. The results show that
6 The diferences in the annotations were two times 1, once 0, and once 4.
the automatic clustering of EDUs by semantics still lags behind manual
annotation. However, they also reveal that even the manually produced clustering is
ambiguous, as one would have expected to find (almost) all the wrong EDUs.
Overall, the annotators’ impression was that it was a very dificult task to spot
the intruding EDU because except for the query no context information was
given. In most cases, the query did not really help in identifying the out-of-place
EDU. In contrast, when creating the ground truth, the (other) annotator first
read the whole texts associated with result claims and then decided which EDUs
should be clustered. This is an important diference.
5
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion and Future Work</title>
      <p>Segmenting complex premises and clustering of semantically similar premises
are important tasks in the retrieval of arguments, as argument retrieval systems
need to deal with complex natural-language statements and should not show
duplicate results. This is even a problem for arguments extracted from debate
portals since single textual premises often address a variety of aspects. In this
paper we discussed the segmentation of premises into EDUs, as well as clustering
these from an uncurated dataset. Our results show that segmenting premises into
their EDUs in such a dataset with rule-based procedures that are suitable for
curated datasets is feasible, in particular by following either a precision or a
recall-oriented approach. Furthermore, we have seen that clustering EDUs only
performs comparably well with the associated premises as context information
at least. The segmentation of EDUs from noisy texts remains a dificult task for
now. We provide the labeled data of EDUs and clusters of EDUs so that future
argument mining methods can use it for evaluation of their performance.</p>
      <p>Future work will include extracting unique EDUs using context information
and further analyzing properties of real-world datasets which impede manual
EDU extraction and clustering. With these insights, an annotation support
system could be constructed to help manually identifying and clustering EDUs.
Acknowledgments We would like to thank Anna-Katharina Ludwig for her
invaluable help in clustering the EDUs and Patrick J. Neumann for his help in
the implementation.</p>
      <p>This work has been funded by the Deutsche Forschungsgemeinschaft (DFG)
within the project ReCAP, Grant Number 375342983 - 2018-2020, as part of the
Priority Program ”Robust Argumentation Machines (RATIO)” (SPP-1999).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Afantenos</surname>
            ,
            <given-names>S.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Denis</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Muller</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Danlos</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Learning recursive segments for discourse parsing</article-title>
          .
          <source>In: LREC</source>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Akbik</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blythe</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vollgraf</surname>
          </string-name>
          , R.:
          <article-title>Contextual string embeddings for sequence labeling</article-title>
          .
          <source>In: COLING</source>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Boltuzic</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Snajder</surname>
          </string-name>
          , J.:
          <article-title>Identifying prominent arguments in online debates using semantic textual similarity</article-title>
          . In: ArgMining@
          <string-name>
            <surname>HLT-NAACL</surname>
          </string-name>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Cabrio</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Villata</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Five years of argument mining: a data-driven analysis</article-title>
          .
          <source>In: IJCAI</source>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Carlson</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marcu</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          : Discourse Tagging Reference Manual, https://www.isi. edu/~marcu/discourse/tagging-ref-manual.pdf
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Carreras</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Màrquez</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Boosting trees for clause splitting</article-title>
          .
          <source>In: ACL</source>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Conneau</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kiela</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schwenk</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barrault</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bordes</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Supervised learning of universal sentence representations from natural language inference data</article-title>
          .
          <source>In: EMNLP</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Devlin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toutanova</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>BERT: pre-training of deep bidirectional transformers for language understanding</article-title>
          .
          <source>In: NAACL-HLT</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Dumani</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neumann</surname>
            ,
            <given-names>P.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schenkel</surname>
            ,
            <given-names>R.:</given-names>
          </string-name>
          <article-title>A framework for argument retrieval - ranking argument clusters by frequency and specificity</article-title>
          .
          <source>In: ECIR. LNCS</source>
          , vol.
          <volume>12035</volume>
          . Springer (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Dumani</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schenkel</surname>
            ,
            <given-names>R.:</given-names>
          </string-name>
          <article-title>A systematic comparison of methods for finding good premises for claims</article-title>
          .
          <source>In: SIGIR</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>van Eemeren</surname>
            ,
            <given-names>F.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garssen</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krabbe</surname>
            ,
            <given-names>E.C.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Henkemans</surname>
            ,
            <given-names>A.F.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Verheij</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wagemans</surname>
            ,
            <given-names>J.H.M</given-names>
          </string-name>
          . (eds.):
          <source>Handbook of Argumentation Theory</source>
          . Springer (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Janier</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lawrence</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reed</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          : OVA+
          <article-title>: an argument analysis interface</article-title>
          .
          <source>In: COMMA</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Krippendorf</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Estimating the reliability, systematic error and random error of interval data</article-title>
          .
          <source>Educational and Psychological Measurement</source>
          <volume>30</volume>
          (
          <year>1970</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Langfelder</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Horvath</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Defining clusters from a hierarchical cluster tree: the Dynamic Tree Cut package for R</article-title>
          .
          <source>Bioinformatics</source>
          <volume>24</volume>
          (
          <issue>5</issue>
          ) (11
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Mann</surname>
            ,
            <given-names>W.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thompson</surname>
            ,
            <given-names>S.A.</given-names>
          </string-name>
          :
          <article-title>Rhetorical structure theory: Toward a functional theory of text organization</article-title>
          .
          <source>Text &amp; Talk</source>
          <volume>8</volume>
          (
          <issue>3</issue>
          ) (
          <year>1988</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          .
          <source>In: NeurIPS</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Peldszus</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stede</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          : From Argument Diagrams to Argumentation
          <source>Mining in Texts. IJCINI 7</source>
          (
          <issue>1</issue>
          ) (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Peters</surname>
            ,
            <given-names>M.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neumann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Iyyer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gardner</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zettlemoyer</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Deep contextualized word representations</article-title>
          .
          <source>In: NAACL-HLT</source>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Reimers</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schiller</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Beck</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Daxenberger</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stab</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gurevych</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Classification and clustering of arguments with contextualized word embeddings</article-title>
          .
          <source>In: ACL</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Rendón</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Abundez</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arizmendi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Quiroz</surname>
            ,
            <given-names>E.M.:</given-names>
          </string-name>
          <article-title>Internal versus external cluster validation indexes</article-title>
          .
          <source>Int. J. Comput. Commun</source>
          .
          <volume>5</volume>
          (
          <issue>1</issue>
          ) (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Soricut</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marcu</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Sentence level discourse parsing using syntactic and lexical information</article-title>
          .
          <source>In: HLT-NAACL</source>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Stab</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Daxenberger</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stahlhut</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Miller</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schiller</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tauchmann</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eger</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gurevych</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Argumentext: Searching for arguments in heterogeneous sources</article-title>
          .
          <source>In: NAACL-HTL</source>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Stab</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gurevych</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Identifying argumentative discourse structures in persuasive essays</article-title>
          .
          <source>In: EMNLP</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Stede</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Afantenos</surname>
            ,
            <given-names>S.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peldszus</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Asher</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perret</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>Parallel discourse annotations on a corpus of short texts</article-title>
          . In: LREC (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Tofiloski</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brooke</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Taboada</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>A syntactic and lexical-based discourse segmenter</article-title>
          .
          <source>In: ACL and AFNLP</source>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Trautmann</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Daxenberger</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stab</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schütze</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gurevych</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Fine-grained argument unit recognition and classification</article-title>
          . In: AAAI (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Wachsmuth</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khatib</surname>
            ,
            <given-names>K.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ajjour</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Puschmann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Qu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dorsch</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Morari</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bevendorf</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Building an argument search engine for the web</article-title>
          . In: ArgMining@EMNLP (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>