<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>April</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Argumentation mining in scientific literature: From computational linguistics to biomedicine</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Pablo Accuosto</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mariana Neves</string-name>
          <email>mariana.lara-neves@bfr.bund.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Horacio Saggion</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>German Federal Institute for Risk Assessment (BfR)</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>LaSTUS/TALN Group, Universitat Pompeu Fabra</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <volume>1</volume>
      <issue>2021</issue>
      <fpage>20</fpage>
      <lpage>36</lpage>
      <abstract>
        <p>In this work we propose to tackle the limitations posed by the lack of annotated data for argument mining in scientific texts by annotating argumentative units and relations in research abstracts in two scientific domains. We evaluate our annotations by computing inter-annotator agreements, which range from moderate to substantial according to the dificulty level of the tasks and domains. We use our newly annotated corpus to fine-tune BERT-based models for argument mining in single and multi-task settings, ifnally exploring the adaptation of models trained in one scientific discipline (computational linguistics) to predict the argumentative structure of abstracts in a diferent one (biomedicine).</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;argument mining</kwd>
        <kwd>scientific corpora</kwd>
        <kwd>domain adaptation</kwd>
        <kwd>transformer models</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>The accelerated pace at which scientific knowledge is produced makes its discovery and
assessment a challenging task. Natural language processing (NLP) technologies, in general, and
text-mining tools, in particular, have become increasingly essential to identify and characterize
the most relevant information produced in a given scientific discipline.</p>
      <p>
        In order to assess a research article it is necessary to consider its logic, rhetoric and dialectic
quality dimensions [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. It is therefore not enough to identify the claims made by its authors but
also the evidence that they provide to support them. NLP tools that help to identify the main
argumentative elements of a given text and how they are connected to each other can support
the assessment of a given article. The automatic identification of arguments, its components
and relations in texts is known as argument mining or argumentation mining [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The tasks
involved in the automatic extraction of arguments from texts (claim/premise identification,
prediction of argumentative structure) are not substantially diferent to other text mining tasks
for which neural-based supervised learning methods produce state-of-the-art results (e.g.: text
segmentation, sequence labelling and entity linking) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. These approaches, however, rely on
large volumes of annotated data which are dificult to obtain for complex tasks such as argument
mining. Scarcity of annotated corpora, therefore, limits the possibilities of using supervised
machine learning algorithms for the identification of argumentative units and relations in texts
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. This obstacle is greater when dealing with scientific discourse: the inherent complexity of
scientific texts makes it very dificult to carry out large annotation eforts with lay annotators.
      </p>
      <sec id="sec-1-1">
        <title>1.1. Contributions</title>
        <p>
          In previous works [
          <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
          ] we proposed an annotation scheme for argumentative units and relations
which considered the specificities of the scientific discourse and conducted a pilot annotation
experiment with 60 abstracts from papers in the computational linguistics domain. These
pilot annotations were done by one person as they were intended: i) to analyze the possibility
of leveraging information contained in discourse-level annotations in order to improve the
performance of argument mining models trained with a small number of abstracts and, ii) to
explore the potential value of the trained models to predict the acceptance/rejection of the
manuscripts in computational linguistics conferences, which was considered as a proxy for
argumentative quality aspects of the abstracts. In those pilot experiments we trained BiLSTM
models with CRF classifiers on top and used contextualized word embeddings obtained by
means of pre-trained ELMo encoders [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. In this work we:
1. Refine our previous annotation scheme to better account for the argumentative structure of
scientific abstracts and to simplify the annotation process.
2. Make available SciARG, a corpus obtained by applying our new scheme to the annotation of
510 scientific abstracts in two domains: computational linguistics (CL) and biomedicine. Three
annotators participated in the annotation process in the CL domain, while two annotators
were involved in the annotation of biomedical abstracts.
3. Assess the consistency of SciARG annotations by analyzing inter-annotator agreement.
4. Use the SciARG corpus to fine-tune and evaluate BERT-based argument mining models, both
in single and multi-task settings.
5. Analyze the potential of adapting models trained with CL abstracts -the original discipline
for which the annotation schema was developed- to the biomedical domain.
        </p>
        <p>The SciARG corpus and the code used in the experiments described in this work are made
publicly available as a contribution to the research community.1</p>
        <p>The rest of the paper is organized as follows: in Section 2 we describe previous work aimed at
identifying arguments in scientific texts. In Section 3 we describe the data used to generate the
corpus, our proposed annotation scheme and the annotation process. In Section 4 we describe
the experiments conducted with the generated corpus and, in Section 5, we analyze the results
obtained. In Section 6, we present our conclusions and suggest potential follow-ups.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        The inherent complexity and ambiguity of the scientific language makes the identification
of arguments in scientific texts a particularly challenging task [
        <xref ref-type="bibr" rid="ref10 ref8 ref9">8, 9, 10</xref>
        ]. The Argumentative
Zoning (AZ) model [
        <xref ref-type="bibr" rid="ref11 ref12">11, 12</xref>
        ] and the CoreSC scheme [
        <xref ref-type="bibr" rid="ref13 ref14">13, 14</xref>
        ] provide relevant antecedents in
this area. AZ includes categories used to annotate knowledge claims made by the authors of the
      </p>
      <sec id="sec-2-1">
        <title>1SciARG is available at https://github.com/LaSTUS-TALN-UPF/SciARG</title>
        <p>
          papers and to establish connections with previous works. CoreSC, in turn, provides a readable
representation of the research process described by the paper. Diferences and similarities
between the two schemes are studied in [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]. AZ was originally applied to the annotation of
computational linguistics texts and CoreSC in physical chemistry and bio-chemistry articles.
Dernoncourt and Lee [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] released in 2017 the PubMed 200k RCT dataset as a resource to
train sentence classifiers for unstructured abstracts. The dataset was constructed by retrieving
195,654 structured abstracts of randomized controlled trials from the 2016 MEDLINE/PubMed
Baseline Database2 and labelling each sentence with the name of the section it belongs to. It is
relevant to note that the aforementioned corpora and datasets are aimed at the classification
of the rhetorical role of sentences but not the discourse relations between them. In this work
we intend to establish a bridge between these two annotation levels. Lawrence and Reed [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]
and Lippi and Torroni [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] provide thorough analyses of argument mining initiatives in various
types of texts and domains, including legal documents [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ], online discussions [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ], Wikipedia
articles [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ], newspapers [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ], student essays [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] and television debates [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ], while Habernal
and Gurevych [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ] and Schulz et al. [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ] explore argument mining in collections of texts from
multiple, diverse sources. Few annotation eforts have focused on the analysis of arguments in
scientific articles when compared to the number of works aimed at identifying argumentative
components and relations in other textual genres. The annotation of 24 German scientific
articles in the educational domain by Kirschner et al. [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] is one of the first works intended
for the analysis of the whole argumentative structure of scientific texts, considering not only
argumentative components but also how they are linked to each other. Lauscher et al. [
          <xref ref-type="bibr" rid="ref25 ref26">25, 26</xref>
          ]
carried out experiments in which they enriched, with an argumentation layer, 40 papers in the
area of computer graphics included in the DrInventor Scientific Corpus [
          <xref ref-type="bibr" rid="ref27">27</xref>
          ]. As mentioned in
Section 1, we have previously conducted experiments with 60 computational linguistic abstracts
aimed at analyzing the potential benefits obtained by enriching argument mining models with
discourse-level knowledge [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. SciARG Corpus</title>
      <p>
        In this section we describe the source data used as a basis of the SciARG corpus as well as the
annotation schema that we propose. We describe the annotation process and assess the quality
of the produced annotations by considering inter-annotator agreement measures.
3.1. Data
The SciARG corpus covers two knowledge areas: computational linguistics and biomedicine.
We refer to these sub-corpora as CL and BIO, respectively.
• CL corpus. Includes 225 computational linguistics abstracts from the ACL Anthology [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ].3
These abstracts are a subset of the 798 abstracts annotated with discourse relations in the
2The MEDLINE database of life sciences and biomedical information (www.nlm.nih.gov/bsd/medline.html) is
maintained by the U.S. National Library of Medicine and available through the PubMed (pubmed.ncbi.nlm.nih.gov)
search engine.
      </p>
      <p>3In particular, from the Proceedings of the 2014 Conference on Empirical Methods in NLP (EMNLP).</p>
      <p>
        Discourse Dependency TreeBank for Scientific Abstracts (SciDTB) [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ]. 4
• BIO corpus. Includes 285 biomedical abstracts of articles from MEDLINE/PubMed. These
abstracts are a sample of those used by Neves et al. [
        <xref ref-type="bibr" rid="ref32">32</xref>
        ] for the evaluation of argumentation
in the biomedical domain. The sample was selected in a stratified way in order to include all
annotations types considered in the referred work.
      </p>
      <sec id="sec-3-1">
        <title>3.2. Annotation scheme</title>
        <p>
          In this work we focus on the analysis of the way in which authors logically structure information
in abstracts to persuade potential readers about the relevance and validity of their proposals. Our
annotation scheme is aimed at capturing the underlying argumentative structure departing from
its linguistic realization. It is therefore relevant to consider previous works that characterize
the diferent constituent elements of scientific abstracts. Several works have been dedicated to
the study, from a genre analysis perspective, of the rhetorical structure of scientific articles and
its parts [
          <xref ref-type="bibr" rid="ref33 ref34">33, 34</xref>
          ]. Based on these works, a broad categorization of the most frequent rhetorical
moves in scientific abstracts can be considered: i) contextualization of the research topic; ii)
limitations in existing solutions; iii) main purpose of the current work; iv) description of the
methodology; v) summary of the main results; vi) conclusions.5 Based from this general
structure of scientific abstracts we propose a fine-grained scheme that considers a sentence as
the annotation unit and contains 11 types of units (Table 1) and six types of directed relations
(Table 2).6 Each of the unit types can, in turn, be mapped to a coarse-grained category. The use
of fine or coarse-grained types depend on specific usages of the corpus. 7
        </p>
        <p>Description
high level description of the proposed approach/solution
processes/tools/methods that are part of the proposal
data obtained from experiments
direct interpretation of observed data
results and the means by which they were obtained
high-level interpretation/generalization of results
secondary methods/processes not part of the proposal
known problem/limitation addressed by the proposal
new ideas/paths for known problems/limitations
known information to support the proposed approach
additional information (definitions/examples)
Coarse
proposal
proposal
outcomes
outcomes
outcomes
outcomes
methods
motivation
motivation
motivation
other</p>
        <p>
          An annotated abstract can be seen as a directed graph with the sentences as its nodes and
the relations between them as the edges. In order to gain in uniformity of the annotations,
4This allows us to continue exploring the interaction between argumentative, rhetoric and discourse annotation
levels in scientific abstracts, as originally proposed by Peldszus and Stede [
          <xref ref-type="bibr" rid="ref30 ref31">30, 31</xref>
          ] for other textual genres.
5Minor variations to this general structure depend on the knowledge area.
6We omit the attack relation as there were no attacks identified in any of the abstracts analyzed.
7In the context of this work we use the fine-grained types.
        </p>
        <p>
          Description of the child node function
provides new supporting information/evidence for the parent
provides additional information relevant to assess/contextualize the parent
describe methods through which supporting evidence is obtained
provides information essential to understand/contextualize the parent
describes a step that comes after the step described by the parent in a process
provides non-essential information
reduce the level of ambiguity and simplify the annotation process, we only consider trees as
valid annotations graphs. In addition to types and relations, annotators were asked to identify
the unit that describe the most significant contribution of the work as the main unit. Fig. 1
shows an example of the tree resulting from annotating the abstract from [
          <xref ref-type="bibr" rid="ref35">35</xref>
          ].
        </p>
        <p>
          In previous work we considered argumentative units at sub-sentence level and explored the
relation between discourse and argumentative levels [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. Having observed that discourse-level
annotations can be leveraged to identify argumentative relations within sentences we decided,
in this work, to focus at the sentence level and leave the prediction of intra-sentence relations
as a second step in an argument mining pipeline. This allows us to facilitate the annotation
process and it also contributes to bridge the gap between annotations aimed at identifying
the rhetorical role of sentences (such as AZ and CoreSC) and those aimed at finding discourse
relations between -and within- them (such as SciDTB).
        </p>
        <p>Most sentences in computational linguistics abstracts contain one type of argumentative unit.
A relatively frequent exception to this are sentences in which mentions to methods are included
in the results. For this specific case we introduce in our schema the unit type results-means.8 For
other cases in which more than one type of unit can be identified, our scheme allows annotators
to register this information by assigning a second type to the sentence. In the annotation process
annotators were asked to weight the relevance of the diferent types of information contained
in the sentence to make a decision with respect to the main and secondary types.</p>
        <p>It is frequent to find, in abstracts, that authors build up supporting evidence or justifications
for implicit or explicit claims in more than one sentence. Consider the example in Fig. 1.
Nodes (4) (motivation-problem) and (7) (motivation-background) provide partial information
that, when considered together, contribute to justify the proposed work described in node (2).
From a discourse analysis perspective, this would be represented by a multi-nuclear relation
which could be annotated by introducing a diferent type of node in the argumentative tree.
This, on one hand, introduces some practical dificulties in the automatic processing of the
annotations, as will become evident when we describe the experiments in Section 4 and, on the
other hand, does not allow to capture the hierarchical relation between nodes (4) and (7). We
opt, instead, to introduce the relation info-required to account for these cases. In this example,
we indicate that there is an info-required relation that goes from node (7) to node (4) and a
support relation that goes from node (4) to node (2). When looking for supporting evidence
for the sentence in node (2), therefore, we would consider not only their direct children but</p>
        <sec id="sec-3-1-1">
          <title>8Examples for all types of units are included in the supplemental material.</title>
          <p>
            also the chains of sentences below them linked by info-required relations. Depending on the
specific dimensions of the argumentation quality to analyze, therefore, diferent subsets of
units and relations can be considered. Most argument mining works focus on logic aspects of
argumentation and, in particular, in the arguments’ cogency [
            <xref ref-type="bibr" rid="ref1">1</xref>
            ]. This argumentative dimension
is conveyed by relations of type support (or attack). It has been noted, nevertheless, that the
diferent argumentative dimensions correlate with the perceived overall argumentative quality
of the text [
            <xref ref-type="bibr" rid="ref1 ref36">1, 36</xref>
            ]. Should we consider only support relations, our annotations would not capture
the link between a proposal and its implementation details, which improves the text’s clarity
and persuades the reader about the validity of the proposal and, therefore, the perceived overall
argumentative strength of the text.
          </p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>3.3. Annotation process</title>
        <p>
          The annotation guidelines used in this work are made available online.9 The first step of the
annotation process is to have the texts of the abstracts splitted into sentences. In the case of the
CL corpus, the source files used are already segmented into sentences and elementary discourse
units in the SciDTB corpus. For the biomedical abstracts the sentence segmentation is done by
means of the syntok tool.10 The annotation was done by means of a modified version of GraPAT
(Graph-based Potsdam Annotation Tool) [
          <xref ref-type="bibr" rid="ref37">37</xref>
          ] according to the specific needs of the task.
        </p>
        <p>We first developed and adjusted the annotation scheme with the CL corpus and then used it
to annotate the BIO corpus. One of the goals of this work is to assess the applicability of the
proposed scheme to diferent domains. In Section 3.5 we analyze the main diferences between
both corpora and the resulting annotations. The annotation of the CL corpus was done by
9https://github.com/LaSTUS-TALN-UPF/SciARG/blob/main/Annotation_Guidelines_Arguments_SciDTB.pdf
10github.com/fnl/syntok
three expert annotators, 1, 2, 3 (two NLP researchers and one computational linguist) in
three rounds. The first two rounds were aimed at training the annotators, clarifying doubts and
making the necessary adjustments to the annotation scheme and tool. As a result of the whole
process 225 CL abstracts were annotated, having 30 abstracts annotated by the three annotators
to compute inter-annotator agreement. For the BIO corpus only two of the CL annotators (1,
2) could participate in the annotation process. In this case no training phase was needed and
there were no substantial modifications to the annotation tool or scheme. As a result of this
process 285 abstracts were annotated, of which 50 were annotated by both annotators. For our
experiments we split both sub-corpora into training and test sets, as described in Section 5.1.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.4. Agreement</title>
        <p>
          In this section we assess the reliability of our annotations by considering inter-annotators’
agreement. Table 3 shows the agreements obtained for CL and BIO sub-corpora. In the case of
CL, where we have three annotators, we report the average of the pairwise agreements with
their corresponding standard deviations. In order to compute the agreements we consider exact
matches between pairs of labels assigned by two annotators. For the parent attachment task
the label is the absolute position of the parent sentence in the document. In addition to each
sub-task-specific agreement, we report the agreement observed when considering simultaneous
exact matches for all the tasks. It is relevant to note that annotations within one document
cannot be consider as completely independent from each other, which presents a limitation
when interpreting the significance of Cohen’s  coeficient. 11 In Table 3, therefore, in addition
to Cohen’s  s, we directly report the accuracy obtained for the diferent tasks without any
presupposition with respect to the independence of the annotations. When considering a pair
of annotated documents as labeled trees, the accuracy indicates the number of changes (with
respect to the number of nodes) that would be necessary to do in one tree to obtain the other one.
It can, therefore, be interpreted as an edit-distance measure that allows to estimate the degree of
agreement between two annotators with respect to the argumentative roles of the nodes when
considering the document as a whole. In all cases annotation agreements fall between moderate
and substantial levels: substantial agreements are obtained in general in the CL corpus (and
almost perfect agreement when coarse-grained types are considered), while agreements in the
identification of unit types and relations are lower in the BIO corpus. In addition to the fact
that the annotation scheme was designed and adjusted specifically for the CL domain, lower
agreements are expected in BIO as abstracts have a higher level of complexity than CL ones
in terms of their structure, the number of units that they contain and their lengths, as shown
in Section 3.5. It is also relevant to note that annotators have a high level of familiarity with
CL texts while they are not experts in the BIO domain. When analyzing discrepancies in the
annotation of the BIO corpus we observe that units of types observation and result give origin
to systematic disagreements between annotators 1 and 2. In fact, annotator 2 annotated
as observation 64% of the units that annotator 1 annotated as result, which makes us believe
that a clear distinction between these two types is dificult to establish without specific domain
11As a decision made at one node of the argumentative structure afects decisions made in other nodes. This
problem has already been observed by Marcu et al. [
          <xref ref-type="bibr" rid="ref38">38</xref>
          ] when evaluating inter-annotator agreement of discourse
annotations.
knowledge. When coarse-grained types are considered, in fact, these two types of units are
not distinguished and the level of agreement reaches 0.93 Cohen’s  . This also afects the
attachment of these units to their parents as they are also considered diferently in terms of the
argumentative role that they play.
        </p>
      </sec>
      <sec id="sec-3-4">
        <title>3.5. Corpus statistics and analysis</title>
        <p>Substantial diferences can be observed between the CL and BIO sub-corpora. Abstracts in
BIO are, in general, longer and argumentatively more complex than those in CL (Table 4). It
is frequent in BIO to find abstracts that describe a series of experiments, each one with their
results. In some cases, results from one experiment are used to motivate and/or justify new
ones. This level of detail is not present in CL abstracts. In general, the description of research
outcomes and their interpretation is much more complex in BIO abstracts, which leads to a
significant diference in the number of units of type observation, result and conclusion when
compared to CL abstracts.12</p>
        <p>The distinction between the plain report of observed data, the interpretation of results and
the extraction of conclusions from them is more ambiguous in BIO than in CL and, therefore,
diferentiating these types of units is more dificult. The distances between units and their
parents are greater in BIO. In fact, nearly 19% of the times a unit is 5 or more units away from
its parent. In CL this occurs only in 2% of the cases. In 69% of the cases CL units are only one or
two units away from its parent when considering the CL corpus. In BIO this occurs only 58% of
12While in CL 3% of the units are of type observation, 19% of type result or result-means and 4% of type conclusion,
in BIO there are, respectively, 18%, 26% and 11% units of these types. More details are provided in the supplemental
material.
the times. In both domains backward relations are more frequent than forward relations: the
parent occurs before the child 68% and 66% of the times in CL and BIO, respectively.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Experiments</title>
      <p>4.1. Tasks
In this section we describe the experiments carried out in order to, given a scientific abstract,
predict the nodes and relations needed to represent its argumentative structure.
• Unit type: Given the text of a sentence, predict its type. The class to predict in this case is
one of the 11 fine-grained types described in Table 1.
• Relation direction: Given two sentences, predict whether a forward or backward relation
exists between them (e.g: whether the first unit is a child of the second one in the argumentative
tree or vice versa). We model this task as a three-class classification task where, given two
sentences, the possible classes to predict are forw, back or none, indicating, respectively, that
there is a directed relation from the first to the second sentence, from the second to the first
sentence, or that the two sentences are not related.
• Relation type: Given a sentence, predict the label of the relation with its parent (its
argumentative and/or discourse function) or none for the root node. The class to predict in this
case is one of the 6 relations described in Table 2.
• Main unit: Given the text of a sentence, predict whether it is the main unit.13</p>
      <p>There are clear links between the four tasks. For instance, the main unit is, in most cases,
the root of the argumentative tree. Associations can also be established between a unit’s type
and its function: in the most frequent case, units of type result are used to support units of type
proposal or conclusion. It is therefore natural to explore the possibility of training the tasks
jointly, in a multi-task setting, which we compare to the results obtained when training the
results independently, in single-task settings.</p>
      <sec id="sec-4-1">
        <title>4.2. Experimental setup</title>
        <p>
          Transformer-based encoders [
          <xref ref-type="bibr" rid="ref39">39</xref>
          ] such as BERT [
          <xref ref-type="bibr" rid="ref40">40</xref>
          ] currently provide state-of-the-art
performance for semantic text classification tasks. For our experiments we make use of the BERT
implementations available as part of HuggingFace’s Transformers library [
          <xref ref-type="bibr" rid="ref41">41</xref>
          ]. We use the cased
version of SciBERT [
          <xref ref-type="bibr" rid="ref42">42</xref>
          ] as base model, as it is trained on texts in the same domains as the ones
covered by our corpus. We apply the standard method of considering the representation of
the [CLS] token and feed it into linear classifiers. A softmax function is then applied to the
classifier’s output in order to obtain the distribution of probabilities for the predicted labels. In
the multi-task setting the BERT layers are shared among all the tasks only training
independently the task-specific heads. We follow the common practice of modeling the identification of
relations between pairs of sentences as a classification task using as input the sequence obtained
13The main unit is considered to be the unit where the main proposed approach/solution is described.
by concatenating the tokens occurring in each sequence separated by the [SEP] special token.
In order to predict the relations present in a given abstract we consider all the pairs formed by a
sentence and the sentences that occur after it in the text. As a result of the prediction we should
obtain the label forw if the first sentence is a child of the second one in the argumentative tree,
back if the relation is established in the opposite direction, and none if the two sentences are not
in a direct relation. The most frequent case, given any two sentences, is that they are not related.
In order to train the model with more positive examples, when a relation exists between two
sentences we sample it twice in the training set: once for each direction, with the corresponding
forw / back labels.14 For evaluation we consider each pair only once, in the order in which they
appear in the text.
        </p>
        <p>
          When fine-tuning our models in each domain, we consider the median number of tokens in
the input sequences and set the maximum sequence length to its double. We consider
crossentropy as the loss function to optimize. We use the Adam optimizer with a learning rate of
2e-5 and a warm-up period of 10% of the learning steps. We set a dropout probability of 0.1
for multi-task settings and 0.2 for single-task ones. The batch size used is of 16 instances with
gradient accumulation of 2 batches. These hyperparameters were set based on five-fold
crossvalidation evaluations in the training set. While the general recommendation is to fine-tune
BERT for 2 to 4 epochs [
          <xref ref-type="bibr" rid="ref40">40</xref>
          ], we observed that more epochs were required to train our tasks,
considering the large number of classes and the relatively small number of training instances.15
In Section 5 we report the results obtained for each task for 5, 10 and 15 epochs, so it is possible
to observe how each combination of task and domain impact on the training time required both
in single and multi-task settings, which would be more dificult to observe if we considered
either the best cross-validation epoch or a fixed number of epochs, as frequently done when
using BERT. We also observed that the models’ overall performances improved when including,
as additional tokens, information about the sentences positions in the abstracts as well as their
relative distance and order. We add special tokens to the standard BERT tokenizer to represent
this information.16
        </p>
        <p>As mentioned in Sections 1.1 and 3, our annotation scheme was specifically developed to
account for argumentative types and relations in computational linguistic abstracts. One of the
goals of this work is to explore i) the applicability of this scheme to other scientific disciplines and
ii) whether models trained in the CL domain can be easily adapted to predict the argumentative
structure of abstracts in other scientific areas. In particular, in the BIO domain. We use the
newly annotated set of biomedical abstracts in order to respond to both research questions.</p>
        <p>We are also interested in exploring to what extent models trained with annotations in CL
contain task-specific information that can be exploited to predict argumentative types and
relations in scientific abstracts with a more complex structure and in another discipline. We
therefore analyze the results obtained by keeping the weights of a model fine-tuned with the
CL abstracts fixed and only training a linear classifier on top of it with the BIO abstracts.</p>
        <p>14I.e.: if sentence 2 is a child of sentence 1 in the argumentative tree, we include the instances (2, 1, forw)
and (1, 2, back) in the training set.</p>
        <p>15While there are 3 classes and 13,874 training instances for the BIO/relation type task, we only have 1,049
training instances and 11 classes for CL/unit type.</p>
        <p>16I.e.: "[CLS] [AFTER] [DISTANCE-1] [POS-1] This paper presents ... [SEP] We observe ..."</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Results</title>
      <p>We present in this section the performance of the models trained in each domain (BIO, CL), as
well as the adaptation of CL models to BIO by training a small number of additional parameters.</p>
      <sec id="sec-5-1">
        <title>5.1. Evaluation</title>
        <p>In the CL domain 30 abstracts were annotated in common by three annotators (1, 2, 3).
This set is used for evaluation, while the rest of the 195 abstracts is used to train the models.
We generate a set of consensus annotations by assigning, to each instance, the majority label
considering the annotations by 1, 2 and 3. In the few cases in which there is total discrepancy
among the three annotators we keep the label assigned by the annotator with the highest average
of pairwise agreement with the other two annotators. For BIO it is not possible to do this since
we only have two annotators (1, 2) that annotated 50 abstracts, while the rest of the 235
abstracts were annotated only by one annotator (1). Therefore, in these experiments we only
use the annotations produced by 1. As test set, in this case, we consider the subset of 35
abstracts17 annotated by 1 with the highest levels of agreement with the annotations made by
2, keeping the remaining 250 abstracts annotated by 1 as training set.</p>
        <p>In Table 5 we report the results obtained by the set of experiments described in Section 4
when evaluating against the consensus annotations. For the types of units and relations we
use weighted-averaged F1-scores, as we want to consider the contribution of each label to the
results in proportion to their frequency. For the evaluation of the direction of the relations,
instead, we use macro-averaged scores, which are more sensitive to the minority classes. If
we were to use micro-averaged or weighted-averaged scores in these cases we would obtain
misleading high numbers for F1, given the large proportion of none labels which are correctly
classified. This is also the case for the prediction of the main unit. As expected, considering
the greater argumentative complexity of the BIO abstracts, which is also reflected in the lower
levels of inter-annotator agreements, the performance of the models trained and evaluated
with the BIO annotations is lower than the one obtained with the CL annotations (Table 5).
In the BIO domain the models trained jointly in a multi-task settings tend to perform better
than those in which these tasks are trained independently. In the case of CL the diference
between both settings is less evident: while there is a clear advantage of the multi-task setting
in the prediction of the types of units, better results are obtained for the prediction of the parent
relations in a single-task setting. We also observe that the BERT models fine-tuned with the CL
annotations (CL-BERT) without in-domain fine-tuning perform competitively when compared
to the models in which BERT is fine-tuned with the BIO annotations.</p>
        <p>It is relevant to note that the CL-BERT model with frozen weights performs significantly
better in the prediction of the BIO annotations than the frozen SciBERT encoder that we consider
as baseline. This confirms that the model fine-tuned with CL annotations is able to capture
information about the argumentative structure of scientific abstracts independent of the specific
discipline in which it was trained.</p>
        <p>17The number of 35 is chosen in order to keep the training-test sets percentages similar in both domains.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusions</title>
      <p>In this work we propose a new sentence-level annotation scheme for the identification of
argumentative units and relations in scientific abstracts, which we apply to the annotation of
510 documents in two highly specialized domains: computational linguistics and biomedicine.
The resulting corpus, as well as the code used to train and evaluate models trained with it,
is made publicly available. The results obtained in our experiments encourage us to think
that, in spite of the fact that the annotation scheme was originally developed and refined for
the CL domain, it can be successfully applied to other scientific disciplines. This work also
opens up new research paths, including further exploration of domain adaptation techniques
for argument mining models in challenging domains as is the case of scientific articles.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>This work was (partly) supported by the Spanish Government under the María de Maeztu Units
of Excellence Programme (MDM-2015-0502) and by the Research and Innovation Agency of
Uruguay (ANII). We also acknowledge support from the project Context-aware Multilingual
Text Simplification (ConMuTeS) PID2019-109066GB-I00/AEI/10.13039/501100011033 awarded
by Ministerio de Ciencia, Innovación y Universidades (MCIU) and by Agencia Estatal de
Investigación (AEI) of Spain.</p>
    </sec>
    <sec id="sec-8">
      <title>A. Supplemental material</title>
      <sec id="sec-8-1">
        <title>A.1. Types of units</title>
        <p>proposal
We present a novel approach to improve word alignment for statistical machine translation ( SMT ) .
proposal-implementation
We observe , identify , and detect naturally occurring signals of interestingness in click transitions on the
Web between source and target documents , which we collect from commercial Web browser logs .
observation
Our method produces a gain of +1.68 BLEU on NIST OpenMT04 for the phrase-based system , and a gain
of +1.28 BLEU on NIST OpenMT06 for the hierarchical phrase-based system .
result
Experimental results show statistically significant improvements of BLEU score in both cases over the
baseline systems .
means
We conducted experiments on two standard benchmarks : Chinese PropBank and English PropBank .
result-means
Results on the Switchboard disfluency tagged corpus show utterance-final accuracy on a par with
stateof-the-art incremental repair detection methods , but with better incremental accuracy , faster
time-todetection and less computational overhead .
conclusion
This transfer learning approach brings a clear performance gain over features based on the traditional
bag-of-visual-word approach .
motivation-problem
However , fundamental problems on efectively incorporating the word embedding features within the
framework of linear models remain .
motivation-hypothesis
Combining the two tasks can potentially improve the eficiency of the overall pipeline system and reduce
error propagation .
motivation-background
Recent work has shown success in using continuous word embeddings learned from unlabeled data as
features to improve supervised NLP systems , which is regarded as a simple semi-supervised learning
mechanism .
information-additional
The structure of argumentation consists of several components ( i.e. claims and premises ) that are connected
with argumentative relations .</p>
      </sec>
      <sec id="sec-8-2">
        <title>A.2. Distribution of types and relations in CL and BIO sub-corpora</title>
        <p>39
157
54
27
69
103
157
20
24</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>H.</given-names>
            <surname>Wachsmuth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Naderi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Hou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bilu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Prabhakaran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. A.</given-names>
            <surname>Thijm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Hirst</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <article-title>Computational argumentation quality assessment in natural language</article-title>
          ,
          <source>in: Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics (ACL</source>
          <year>2017</year>
          )
          <article-title>(Volume 1: Long Papers</article-title>
          ),
          <year>2017</year>
          , pp.
          <fpage>176</fpage>
          -
          <lpage>187</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J.</given-names>
            <surname>Lawrence</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Reed</surname>
          </string-name>
          ,
          <article-title>Argument mining: A survey, Computational Linguistics (</article-title>
          <year>2019</year>
          )
          <fpage>1</fpage>
          -
          <lpage>54</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Lippi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Torroni</surname>
          </string-name>
          ,
          <article-title>Argumentation mining: State of the art and emerging trends</article-title>
          ,
          <source>ACM Trans. Internet Technol</source>
          .
          <volume>16</volume>
          (
          <year>2016</year>
          )
          <volume>10</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          :
          <fpage>25</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>C.</given-names>
            <surname>Stab</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Gurevych</surname>
          </string-name>
          ,
          <article-title>Parsing argumentation structures in persuasive essays</article-title>
          ,
          <source>Computational Linguistics</source>
          <volume>43</volume>
          (
          <year>2017</year>
          )
          <fpage>619</fpage>
          -
          <lpage>659</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>P.</given-names>
            <surname>Accuosto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Saggion</surname>
          </string-name>
          ,
          <article-title>Mining arguments in scientific abstracts with discourse-level embeddings</article-title>
          , Data &amp; Knowledge
          <string-name>
            <surname>Engineering</surname>
          </string-name>
          (
          <year>2020</year>
          )
          <fpage>101840</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>P.</given-names>
            <surname>Accuosto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Saggion</surname>
          </string-name>
          ,
          <article-title>Transferring knowledge from discourse to arguments: A case study with scientific abstracts</article-title>
          ,
          <source>in: Proceedings of the 6th Workshop on Argument Mining (ArgMining</source>
          <year>2019</year>
          ),
          <article-title>Association for Computational Linguistics</article-title>
          , Florence, Italy,
          <year>2019</year>
          , pp.
          <fpage>41</fpage>
          -
          <lpage>51</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>W19</fpage>
          -4505.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Peters</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Neumann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Iyyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gardner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Clark</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          ,
          <article-title>Deep contextualized word representations</article-title>
          ,
          <source>in: Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>1</volume>
          (
          <string-name>
            <surname>Long</surname>
            <given-names>Papers)</given-names>
          </string-name>
          ,
          <source>Association for Computational Linguistics</source>
          , New Orleans, Louisiana,
          <year>2018</year>
          , pp.
          <fpage>2227</fpage>
          -
          <lpage>2237</lpage>
          . URL: https://www.aclweb.org/anthology/N18-1202. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>N18</fpage>
          -1202.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>C.</given-names>
            <surname>Stab</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Kirschner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Eckle-Kohler</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Gurevych</surname>
          </string-name>
          ,
          <article-title>Argumentation mining in persuasive essays and scientific articles from the discourse structure perspective</article-title>
          ,
          <source>in: Proceedings of the Workshop on Frontiers and Connections between Argumentation Theory and Natural Language Processing, Forlì-Cesena, Italy, July 21-25</source>
          ,
          <year>2014</year>
          ,
          <year>2014</year>
          , pp.
          <fpage>21</fpage>
          -
          <lpage>25</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>C.</given-names>
            <surname>Kirschner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Eckle-Kohler</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Gurevych</surname>
          </string-name>
          ,
          <article-title>Linking the thoughts: Analysis of argumentation structures in scientific publications</article-title>
          ,
          <source>in: Proceedings of the 2nd Workshop on Argumentation Mining</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>11</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>N.</given-names>
            <surname>Green</surname>
          </string-name>
          ,
          <article-title>Identifying argumentation schemes in genetics research articles</article-title>
          ,
          <source>in: Proceedings of the 2nd Workshop on Argumentation Mining</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>12</fpage>
          -
          <lpage>21</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>S.</given-names>
            <surname>Teufel</surname>
          </string-name>
          , et al.,
          <article-title>Argumentative zoning: Information extraction from scientific text</article-title>
          ,
          <source>Ph.D. thesis</source>
          , University of Edinburgh,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S.</given-names>
            <surname>Teufel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Siddharthan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Batchelor</surname>
          </string-name>
          ,
          <article-title>Towards discipline-independent argumentative zoning: Evidence from chemistry and computational linguistics</article-title>
          ,
          <source>in: Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing (EMNLP</source>
          <year>2009</year>
          )
          <article-title>(Volume 3), Association for Computational Linguistics</article-title>
          ,
          <year>2009</year>
          , pp.
          <fpage>1493</fpage>
          -
          <lpage>1502</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>M.</given-names>
            <surname>Liakata</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Saha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dobnik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Batchelor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Rebholz-Schuhmann</surname>
          </string-name>
          ,
          <article-title>Automatic recognition of conceptualization zones in scientific articles and two life science applications</article-title>
          ,
          <source>Bioinformatics</source>
          <volume>28</volume>
          (
          <year>2012</year>
          )
          <fpage>991</fpage>
          -
          <lpage>1000</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>M.</given-names>
            <surname>Liakata</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. N.</given-names>
            <surname>Soldatova</surname>
          </string-name>
          , et al.,
          <article-title>Semantic annotation of papers: Interface &amp; enrichment tool (SAPIENT)</article-title>
          ,
          <source>in: Proceedings of the Workshop on Current Trends in Biomedical Natural Language Processing, Association for Computational Linguistics</source>
          ,
          <year>2009</year>
          , pp.
          <fpage>193</fpage>
          -
          <lpage>200</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M.</given-names>
            <surname>Liakata</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Teufel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Siddharthan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Batchelor</surname>
          </string-name>
          ,
          <article-title>Corpora for the conceptualisation and zoning of scientific papers</article-title>
          ,
          <source>in: Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC'10)</source>
          ,
          <source>European Language Resources Association (ELRA)</source>
          , Valletta, Malta,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>F.</given-names>
            <surname>Dernoncourt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. Y.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <article-title>Pubmed 200k rct: A dataset for sequential sentence classification in medical abstracts</article-title>
          ,
          <source>arXiv preprint arXiv:1710.06071</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>R.</given-names>
            <surname>Mochales-Palau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.-F.</given-names>
            <surname>Moens</surname>
          </string-name>
          ,
          <article-title>Argumentation mining: The detection, classification and structure of arguments in text</article-title>
          ,
          <source>in: Proceedings of the 12th International Conference on Artificial Intelligence and Law (ICAIL</source>
          <year>2009</year>
          ), ACM,
          <year>2009</year>
          , pp.
          <fpage>98</fpage>
          -
          <lpage>107</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>T.</given-names>
            <surname>Goudas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Louizos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Petasis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Karkaletsis</surname>
          </string-name>
          ,
          <article-title>Argument extraction from news, blogs, and social media</article-title>
          ,
          <source>in: Hellenic Conference on Artificial Intelligence</source>
          , Springer,
          <year>2014</year>
          , pp.
          <fpage>287</fpage>
          -
          <lpage>299</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>E.</given-names>
            <surname>Aharoni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Dankin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Gutfreund</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lavee</surname>
          </string-name>
          ,
          <string-name>
            <surname>R. Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Rinott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Slonim</surname>
          </string-name>
          , Contextdependent evidence detection,
          <year>2018</year>
          . US Patent App.
          <volume>14</volume>
          /720,
          <fpage>847</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>E.</given-names>
            <surname>Florou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Konstantopoulos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Koukourikos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Karampiperis</surname>
          </string-name>
          ,
          <article-title>Argument extraction for supporting public policy formulation</article-title>
          ,
          <source>in: Proceedings of the 7th Workshop on Language Technology for Cultural Heritage</source>
          ,
          <source>Social Sciences, and Humanities</source>
          , Association for Computational Linguistics, Sofia, Bulgaria,
          <year>2013</year>
          , pp.
          <fpage>49</fpage>
          -
          <lpage>54</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>C.</given-names>
            <surname>Stab</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Gurevych</surname>
          </string-name>
          ,
          <article-title>Annotating argument components and relations in persuasive essays</article-title>
          ,
          <source>in: Proceedings of COLING</source>
          <year>2014</year>
          ,
          <source>the 25th International Conference on Computational Linguistics: Technical Papers</source>
          , Dublin City University and Association for Computational Linguistics, Dublin, Ireland,
          <year>2014</year>
          , pp.
          <fpage>1501</fpage>
          -
          <lpage>1510</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>J.</given-names>
            <surname>Visser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Konat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Duthie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Koszowy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Budzynska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Reed</surname>
          </string-name>
          ,
          <article-title>Argumentation in the 2016 us presidential elections: Annotated corpora of television debates and social media reaction, Language Resources and Evaluation (</article-title>
          <year>2019</year>
          )
          <fpage>1</fpage>
          -
          <lpage>32</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>I.</given-names>
            <surname>Habernal</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Gurevych</surname>
          </string-name>
          ,
          <article-title>Argumentation mining in user-generated web discourse</article-title>
          ,
          <source>Computational Linguistics</source>
          <volume>43</volume>
          (
          <year>2017</year>
          )
          <fpage>125</fpage>
          -
          <lpage>179</lpage>
          . doi:
          <volume>10</volume>
          .1162/COLI\_a\_
          <volume>00276</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>C.</given-names>
            <surname>Schulz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Eger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Daxenberger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Kahse</surname>
          </string-name>
          ,
          <string-name>
            <surname>I.</surname>
          </string-name>
          <article-title>Gurevych, Multi-task learning for argumentation mining in low-resource settings</article-title>
          ,
          <source>in: Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>2</volume>
          (
          <string-name>
            <surname>Short</surname>
            <given-names>Papers)</given-names>
          </string-name>
          ,
          <source>Association for Computational Linguistics</source>
          , New Orleans, Louisiana,
          <year>2018</year>
          , pp.
          <fpage>35</fpage>
          -
          <lpage>41</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>N18</fpage>
          -2006.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>A.</given-names>
            <surname>Lauscher</surname>
          </string-name>
          , G. Glavaš,
          <string-name>
            <given-names>S. P.</given-names>
            <surname>Ponzetto</surname>
          </string-name>
          ,
          <article-title>An argument-annotated corpus of scientific publications</article-title>
          ,
          <source>in: Proceedings of the 5th Workshop on Argument Mining (ArgMining</source>
          <year>2018</year>
          ),
          <year>2018</year>
          , pp.
          <fpage>40</fpage>
          -
          <lpage>46</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>A.</given-names>
            <surname>Lauscher</surname>
          </string-name>
          , G. Glavaš, K. Eckert,
          <article-title>ArguminSci: A tool for analyzing argumentation and rhetorical aspects in scientific writing</article-title>
          ,
          <source>in: Proceedings of the 5th Workshop on Argument Mining (ArgMining</source>
          <year>2018</year>
          ),
          <year>2018</year>
          , pp.
          <fpage>22</fpage>
          -
          <lpage>28</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>B.</given-names>
            <surname>Fisas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ronzano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Saggion</surname>
          </string-name>
          ,
          <article-title>A multi-layered annotated corpus of scientific papers</article-title>
          .,
          <source>in: Proceedings of the 2016 The International Conference on Language Resources and Evaluation</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>D. R.</given-names>
            <surname>Radev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Muthukrishnan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Qazvinian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Abu-Jbara</surname>
          </string-name>
          ,
          <article-title>The ACL Anthology network corpus</article-title>
          ,
          <source>Language Resources and Evaluation</source>
          <volume>47</volume>
          (
          <year>2013</year>
          )
          <fpage>919</fpage>
          -
          <lpage>944</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>A.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>SciDTB: Discourse dependency TreeBank for scientific abstracts</article-title>
          ,
          <source>in: Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (ACL</source>
          <year>2018</year>
          )
          <article-title>(Volume 2: Short Papers), Association for Computational Linguistics</article-title>
          , Melbourne, Australia,
          <year>2018</year>
          , pp.
          <fpage>444</fpage>
          -
          <lpage>449</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>A.</given-names>
            <surname>Peldszus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Stede</surname>
          </string-name>
          ,
          <article-title>Rhetorical structure and argumentation structure in monologue text</article-title>
          ,
          <source>in: Proceedings of the Third Workshop on Argument Mining (ArgMining</source>
          <year>2016</year>
          ),
          <year>2016</year>
          , pp.
          <fpage>103</fpage>
          -
          <lpage>112</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>A.</given-names>
            <surname>Peldszus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Stede</surname>
          </string-name>
          ,
          <article-title>An annotated corpus of argumentative microtexts</article-title>
          ,
          <source>in: Proceedings of the First Conference on Argumentation</source>
          , Lisbon, Portugal,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>M.</given-names>
            <surname>Neves</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Butzke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Grune</surname>
          </string-name>
          ,
          <article-title>Evaluation of scientific elements for text similarity in biomedical publications</article-title>
          ,
          <source>in: Proceedings of the 6th Workshop on Argument Mining</source>
          , Association for Computational Linguistics, Florence, Italy,
          <year>2019</year>
          , pp.
          <fpage>124</fpage>
          -
          <lpage>135</lpage>
          . URL: https: //www.aclweb.org/anthology/W19-4515. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>W19</fpage>
          -4515.
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>J.</given-names>
            <surname>Swales</surname>
          </string-name>
          ,
          <source>Genre analysis: English in academic and research settings</source>
          , Cambridge University Press,
          <year>1990</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <surname>M. B. Dos</surname>
            <given-names>Santos</given-names>
          </string-name>
          ,
          <article-title>The textual organization of research paper abstracts in applied linguistics</article-title>
          ,
          <source>Text-Interdisciplinary Journal for the Study of Discourse</source>
          <volume>16</volume>
          (
          <year>1996</year>
          )
          <fpage>481</fpage>
          -
          <lpage>500</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>Z.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wang</surname>
          </string-name>
          , T. Liu,
          <article-title>Transformation from discontinuous to continuous word alignment improves translation quality</article-title>
          ,
          <source>in: Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP)</source>
          ,
          <year>2014</year>
          , pp.
          <fpage>147</fpage>
          -
          <lpage>152</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>L.</given-names>
            <surname>Ng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lauscher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tetreault</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Napoles</surname>
          </string-name>
          ,
          <article-title>Creating a domain-diverse corpus for theorybased argument quality assessment</article-title>
          , arXiv preprint arXiv:
          <year>2011</year>
          .
          <volume>01589</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>J.</given-names>
            <surname>Sonntag</surname>
          </string-name>
          , M. Stede,
          <article-title>GraPAT: A tool for graph annotations</article-title>
          ,
          <source>in: Proceedings of the 2014 The International Conference on Language Resources and Evaluation</source>
          ,
          <year>2014</year>
          , pp.
          <fpage>4147</fpage>
          -
          <lpage>4151</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>D.</given-names>
            <surname>Marcu</surname>
          </string-name>
          , E. Amorrortu,
          <string-name>
            <given-names>M.</given-names>
            <surname>Romera</surname>
          </string-name>
          ,
          <article-title>Experiments in constructing a corpus of discourse trees</article-title>
          ,
          <source>in: Towards Standards and Tools for Discourse Tagging</source>
          ,
          <year>1999</year>
          . URL: https://www. aclweb.org/anthology/W99-0307.
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [39]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaswani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Parmar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Gomez</surname>
          </string-name>
          , Ł. Kaiser,
          <string-name>
            <surname>I. Polosukhin</surname>
          </string-name>
          ,
          <article-title>Attention is all you need</article-title>
          ,
          <source>in: Advances in neural information processing systems</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>5998</fpage>
          -
          <lpage>6008</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          [40]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , BERT:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          ,
          <source>in: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>1</volume>
          (Long and Short Papers),
          <year>2019</year>
          , pp.
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          [41]
          <string-name>
            <given-names>T.</given-names>
            <surname>Wolf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Debut</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Sanh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chaumond</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Delangue</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Moi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Cistac</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Rault</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Louf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Funtowicz</surname>
          </string-name>
          , et al.,
          <article-title>HuggingFace's Transformers: State-of-the-art natural language processing</article-title>
          ,
          <source>ArXiv</source>
          (
          <year>2019</year>
          ) arXiv-
          <fpage>1910</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          [42]
          <string-name>
            <given-names>I.</given-names>
            <surname>Beltagy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lo</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Cohan,
          <article-title>SciBERT: A pretrained language model for scientific text</article-title>
          ,
          <source>in: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLPIJCNLP)</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>3606</fpage>
          -
          <lpage>3611</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>