<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards Data Augmentation for DRS-to-Text Generation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Muhammad Saad Amin</string-name>
          <email>muhammadsaad.amin@unito.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alessandro Mazzei</string-name>
          <email>alessandro.mazzei@unito.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luca Anselma</string-name>
          <email>luca.anselma@unito.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Turin</institution>
          ,
          <addr-line>Corso Svizzera 185, Turin, 10149</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The data augmentation approach is becoming very popular in Natural Language Generation (NLG). Different approaches have been utilized in NLP and NLG to augment data and increase training examples for the neural model. Yet no studies have performed augmentation on logical input i.e., Discourse Representation Structures (DRS). We present data augmentation in DRS i.e., DRS taken from the PMB corpus, for the DRS-to-Text generation task. We conducted our experiments on a standard bi-LSTM-based sequence-to-sequence model thus creating an endto-end neural approach for generating English sentences from DRS. We evaluated the output generated from word-level and character-level decoders with the help of reference-based evaluation metrics like BLEU, ROUGE, METEOR, NIST, and CIDEr. The practical implementation of augmented DRS succeeded in achieving better results compared to DRS without augmentation. To prove the significance of our model, we conducted statistical significance tests i.e., the Shapiro-Wilk Test (to check data normality) and the Wilcoxon Test (to test model significance). Wilcoxon results states that our model is significantly better with the p-value = 2.37e-05 for Char-level model and p-value = 7.78e-07 for Word-level model.</p>
      </abstract>
      <kwd-group>
        <kwd>1 Bi-LSTM</kwd>
        <kwd>Data Augmentation</kwd>
        <kwd>DRS-to-Text Generation</kwd>
        <kwd>Neural Network</kwd>
        <kwd>Parallel Meaning Bank (PMB)</kwd>
        <kwd>Statistical Significance Test</kwd>
        <kwd>Shapiro-Wilk Test</kwd>
        <kwd>Wilcoxon Test</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Data augmentation is an approach utilized to increase the number of examples for training a neural
model without explicitly adding new data examples [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. This approach is becoming very trendy in many
NLP and NLG applications nowadays. This is due to the complex nature of tasks being addressed.
Previously, most of the researchers working in the Computer Vision (CV) domain use different
augmentation techniques i.e., cropping, flipping, color jittering, rotating, etc. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. This CV augmentation
approach is very applicable to increase the number of examples as rotated, flipped or cropped versions
of an image are also an image. But augmentation approach for NLP and NLG is not so easy to implement
due to the discrete nature of sentences [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. That means, if our sentence augmentation is not good, it will
result in ungrammatical sentences and thus result in the bad performance of the model.
      </p>
      <p>
        Discourse Representation Structure (DRS) is derived from Discourse Representation Theory (DRT)
that is the formal representation of data as first order logic. Initial works in formal meaning
representation focused on the generation of DRS from text, an approach referred to as parsing [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. This
work was directed toward mapping of words with their relevant logical representation and formulation.
But very few works have been implemented in translation i.e., generating sentences from Discourse
Representation Structures (DRS). Recently, different authors have implemented a bi-LSTM-based
neural sequence-to-sequence model to generate sentences from DRS [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. But till now to our knowledge,
no work has been done to augment DRS i.e., formal logical representation and translation of the logical
representation. Keeping in mind this research gap, we worked on DRS augmentation to check whether
this approach will help in improving model performance as increased metrics scores.
      </p>
      <p>The research questions that we addressed in these experiments are listed as follows:
1. Is it possible to augment Formal Meaning Representation based on logical inputs i.e., DRS?
2. How augmentation can be performed in DRS and the translation of DRS as both belong to two
different directions?
3. Does augmentation in DRS result in increased model performance?
4. How to statistically justify the results with the help of Significance Tests?</p>
      <p>So, in a nutshell, we can say that our main contribution is twofold. First, we have developed a way
of augmenting logical inputs (DRS) and their respective translations. The initial format of DRS is the
Box Format, and this version of DRS cannot be embedded into the neural network directly. To make
DRS an input for the neural network we must flatten the Box format of DRS into Clausal format and
then Clausal format is preprocessed into Absolute DRS format to be fed into a Neural Network (NN).
Getting corpus data from PMB, we performed an augmentation approach on the Clausal format of DRS
so that it can be preprocessed and passed to the neural model. A graphical depiction of the Box and
Clausal format of DRS along with the translation is shown in Figure 1 below.</p>
      <p>Both formats of DRS have the same meaning but to augment and embed DRS into NN, we must
transform from Box format into Clausal format. So, we argued that the NN trained with augmented data
produces better results. Secondly, we have applied statistical significance tests on the DRS-to-Text
generation task to verify that better results are not achieved accidentally. For the implementation of
statistical significance tests, the choice of the right test is another problem. Among a series of parametric
and non-parametric tests, the choice of the right significance test is a tricky move. A detailed description
of both contributions will be discussed in the latter sections.</p>
      <p>The remaining paper is structured as follows: literature insights are described in Section 2. Section
3 describes the data and the approach used to augment logical input and respective translation of DRS.
The methodology implemented to conduct the experiment is discussed in Section 4. Results are
discussed in Section 5, and the conclusion and future work are described in Section 6.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Literature Insights</title>
      <p>
        Literature insights into data augmentation in Natural Language Processing (NLP) and Generation
(NLG) clearly state that this domain is still underexplored [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Many researchers in NLP have used
different approaches to augment the data examples. Based on the text processing challenges, different
Rule-based and Model-based approaches have been proposed by researchers in this domain [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
Comparing the approaches, there exist some pros and cons of augmentation. Rule-based techniques are
easily implementable but sometimes create more diverse data which is not required for data
augmentation [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. The data which is neither too similar nor too different from the original examples are
considered good augmented data. Because similar or too different data moves towards overfitting of the
model. Similarly, model-based approaches are considered good for augmentation, but it is very difficult
to develop and utilize model-based augmentation approaches for increasing data every time [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
      <p>
        Considering Rule-based techniques, different researchers proposed different approaches based on
the nature of the task being executed. Feature Space Data Augmentation [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], Easy Data Augmentation
based on random insertion, deletion, and swap [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], Paraphrase Identification [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], and Dependency
Tree Morphing [13] are some of the rule-based approaches implemented in the literature. Similarly,
MixUp (also referred as Mixed Sample Data Augmentation Technique, MSDA) [14], CutMix [15],
CutOut [16], Copy-Paste [17], and Seq2MixUp [18] approaches are derived from
In-interpolationbased techniques. Different Model-based techniques include BackTranslation [19], SCPN [20],
Semantic Text Exchange (STE) [21], ContextualAux [22], Lambada [23], XLDA [24], SeqMix [25],
SlotSub-LM [26], UBT &amp; TBT [27], Soft Con-textual DA [28], Data Diversification [29], DiPS [30], and
Augmented SBERT [31].
      </p>
      <p>In our implementation, we have used a Rule-based approach to augment the data. We defined a rule
of verb change with the help of SpaCy NLP pipeline to transform the data in present, past, and future
tenses. Basically, in the DRS-to-Text generation system we have two formats as input to the Neural
Network i.e., DRS and its respective translation as shown in fig. 1. Keeping in mind the aspect and
nature of data used in our experimental implementation, we have to augment DRS and also the
translation of the DRS. The nature of both types of data is totally different i.e., one is a logical input
(DRS) and the other on is a linear text i.e., translation of DRS. By using a Rule-based approach, we
successfully augment the DRS and the translation of DRS to increase the number of relevant examples,
thus achieving higher results.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Data and Augmentation Approach</title>
      <p>Originally, DRS is presented in Box format as it is easy to understand and analyze the structure. Box
representation has unique labels i.e., b1, b2, b3… Each box has 2 layers stated as top-layer and the
bottom layer. The top layer of DRS contains Discourse Referents i.e., x1, t1, and the bottom layer of
DRS contains conditions over these Discourse Referents. Each referent or condition belongs to a unique
box label. For example, b2 person.n.01 x1 contains three types of information i.e., b2 as box label, x1 as
discourse referent, and person.n.01 as a predicate that is disambiguated with senses (senses are provided
in wordnet, synsets) e.g., person.n.01, time.n.08.</p>
      <p>The box format of DRS is not convenient for modeling purposes; therefore, we convert the Box
format into the clausal format. The clausal format or the absolute format is easily readable by the neural
network. In clausal format, the variables and the conditions of the box format are converted into clauses.
For example, top box layer variables are converted into clauses by a special condition called “REF” i.e.,
b2 REF x1 which states that discourse variable x1 is bound in box b2.</p>
      <p>DRS is also referred as the logical representation of components like semantic relations (Agent,
Patient, Theme), operators (REF, NOT), the concepts (touch.v.01), variable indices (b1, x1), and deictic
constants (now, speaker, hearer). By altering the values of these components, one can augment the
DRS. There are multiple ways of augmenting a DRS based on tense-change, polarity-change,
namechange, quantity-change, and by changing numbers. Among all these possible formats of DRS
augmentation, we worked on tense-change approach. In tense-change, the tense of original DRS is
converted into the present, past, and future tense as shown in Figure 2.</p>
      <p>Tense-change augmentation is also referred to as a verb-based (word that describes the action in the
sentence) augmentation approach because we are transforming verbs i.e., present à past and future,
past à present and future, and future à present and past. By default, the tense change variants are
taken as a present, past, and future indefinite tenses.</p>
      <p>3.1.</p>
    </sec>
    <sec id="sec-4">
      <title>Left side: DRS Augmentation</title>
      <p>DRS is a logical combination of events, and entities, and the relationships between these entities.
Certain semantic phenomena are also covered in DRS including pronouns, presuppositions,
quantification, negation, discourse relations, etc. Among different variants of DRS available on The
Parallel Meaning Bank (PMB) corpus, we have used fully interpretable version of DRS. The reason
behind this choice is the representation of information in DRS. In this version of DRS, we have WordNet
synset-based verbs, adverbs, nouns, and adjectives. And Verbnet based semantic relations.</p>
      <p>For augmenting DRS, we worked on a verb-based augmentation approach. To change the relation
between entities of DRS, we adopted a simple string-replacement approach to replace one string with
another string as shown in Fig.2. While iterating through each DRS, we first identified the time in which
a verb is presented e.g., EQU t1 “now”, TPR t1 “now”, and TPR “now” t1. These three formats
represent verbs in any format of the present, past, or future tense. After, the identification of DRS in
one format, we performed string replacement to convert a verb happening only in one type of tense into
multiple types of different tenses e.g., have not à does not, did not, will not etc. This is how to augment
the DRS which is the logical section of our input data. But during the augmentation of DRS, we kept
track of the relevant translations of respective DRS as well. But just like DRS, augmentation of its
translation is not just a string replacement approach. For the augmentation of linear text into different
sentences, we used a Rule-based approach to convert sentences discussed in section 3.2 below.
3.2.</p>
    </sec>
    <sec id="sec-5">
      <title>Right side: Text Augmentation</title>
      <p>Text augmentation as tense change is a very challenging task in NLP. For our implementation, we
have used SpaCy pipeline to transform English sentences from one type of tense into another type based
on the transformation performed in DRS. For implementation, we used SQLite database to keep track
of the sentences with a max length of 1000 characters. We applied this pipeline to process the initial
sentence and worked on sentence patterns to learn the structure of the sentence (conjugates, singular,
plural, past, present, and future).</p>
      <p>In tense transformation e.g., tense change, there are also other factors that must be kept in mind
while reconstructing the sentence. Some major points of consideration include active and passive,
imperative, negation, singular and plural, subject and object, nouns, progressive and perfect, infinitive,
first person, ambiguous, POS, and perfect participles sentences. We have not worked only on simple
and positive sentences but based on the translation of DRS, we have to deal with all types of tenses
mentioned above. Table 1 elaborates on the examples associated with each type of tense form to identify
the complexity of the task addressed.</p>
      <p>If a sentence is presented as present perfect, present perfect continuous, or present continuous than
it is converted into present indefinite as the default mode of tense change is the indefinite mode. The
same strategy is also applied to other types of continuous, perfect and perfect continuous forms of past
and future sentences.</p>
      <sec id="sec-5-1">
        <title>Present to Past &amp; Future</title>
      </sec>
      <sec id="sec-5-2">
        <title>Past to Present &amp; Future</title>
      </sec>
      <sec id="sec-5-3">
        <title>Future to Present &amp; Past</title>
      </sec>
      <sec id="sec-5-4">
        <title>First person</title>
      </sec>
      <sec id="sec-5-5">
        <title>Infinitive</title>
      </sec>
      <sec id="sec-5-6">
        <title>Ambiguous-POS</title>
      </sec>
      <sec id="sec-5-7">
        <title>Plural</title>
      </sec>
      <sec id="sec-5-8">
        <title>Third person singular</title>
      </sec>
      <sec id="sec-5-9">
        <title>Taking will as noun</title>
      </sec>
      <sec id="sec-5-10">
        <title>Perfect tense</title>
      </sec>
      <sec id="sec-5-11">
        <title>Continuous tense</title>
      </sec>
      <sec id="sec-5-12">
        <title>Double tense change</title>
      </sec>
      <sec id="sec-5-13">
        <title>Negation</title>
      </sec>
      <sec id="sec-5-14">
        <title>Future perfect</title>
      </sec>
      <sec id="sec-5-15">
        <title>Passive tenses</title>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>4. Experimental Implementation</title>
      <p>
        For the implementation of the experiment, a series of experimental steps are executed to perform the
task under observation. For implementing augmentation in DRS-to-Text generation, we performed
Rulebased and string replacement based on operations on DRS data. After performing data augmentation,
we must put the augmented data into a bi-LSTM-based neural network to analyze the performance of
our approach. For Neural Machine Translation (NMT) tasks, LSTM has been considered as the best
model due to its ability to remember the connection between long-term input sequences [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Depending
on literature-based suggestions, we also used bi-LSTM-based sequence-to-sequence model to translate
DRS into English sentences.
      </p>
      <p>
        DRS-to-Text is a particular logic to language generation task where input is the first-order logic and
output is the corresponding linear text. This is not a generalized text generation task from graphs, tables,
or images. Therefore, we must use a sequence-to-sequence model capable of remembering long
sequences, and bi-LSTM is proven successful in remembering long logical input sequences [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
Different pre-trained language models like BERT, ELMo, and ROBERTa have been used previously
for parsing e.g., Text-to-AMR and Text-to-DRS. Still, for translation and generation, most of the
researchers have focused only on bi-LSTM-based architectures [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Dealing with a very specific task,
we have not tried other Transformer-based i.e., BERT, GPT, and BART architectures for
logic-tolanguage implementation. But this can be a very interesting future direction to explore further
architectures that can beat bi-LSTM for logic to language-based text generation task.
      </p>
      <p>Neural Architecture. For the implementation of the experiment, we have used the encoder-decoder
architecture of the NMT module. Bi-directional LSTM operates input sequences in both directions. The
encoder part of the model encodes DRS representation, and the decoder module decodes DRS into its
respective English sentences. To conduct this experiment, we have used GPUs with CUDA based
parallel computing platform to speed up the experimental performance. The hyperparameter setting for
our experiment is shown in Table 2 mentioning the parameters and their corresponding values.</p>
      <p>Dataset. We have used the English version of the Parallel Meaning Bank (PMB) 3.0.0 dataset for
our experiment, having gold standard (fully annotated corpus) 6620, 885, and 898 training, validation,
and testing examples. Based on the nature of our implementation, we have used Gold-PMB dataset in
both formats i.e., with augmentation and without augmentation, to check the increase in the evaluation
scores. Then we expanded the training examples by adding Silver-PMB (partially manually annotated
data) 97,598 training examples with Gold-PMB training examples. Collectively, to train our model
without data augmentation, we have 104,218 training, 885 validation, and 898 testing examples. In the
second experiment i.e., DRS-to-Text generation with augmentation, we only performed data
augmentation on training examples. We did not augment, validation, or testing examples of the dataset.
After train augmentation, we were having 26,480 training examples in the case of augmentation in
Gold-PMB, and 4,16,872 training examples in the case of augmentation in Gold-Silver-PMB. Validation
and testing files of PMB data are not augmented in our experiment. We also added only training
examples of Silver-PMB with Gold-PMB to increase the number of training examples for our neural
model. All dataset examples with and without augmentation are mentioned in Table 3 below.</p>
      <p>Implementation Pipeline. The implementation pipeline includes all the steps involved in English
text generation from DRS. Our main focus of this experiment is to perform data augmentation in DRS
and analyze the accuracy improvement. So, we choose the clausal format of augmented DRS and
preprocess it to make meaningful entities as atomic entities. This representation of DRS is meaningful
for a neural network to understand the input pattern and perform well. The complete implementation
pipeline is shown in Figure 3 below.</p>
      <p>
        The encoder part of bi-LSTM encodes the DRS and converts it into vector form. This vector form is
then embedded into the decoder part to be converted into respective English sentences. The neural
model-generated English sentences are then compared with the reference English sentences to calculate
the evaluation scores. For the evaluation of generated sentences, we are using 5 different automatic
evaluation metrics like BLEU, ROUGE, NIST, METEOR, and CIDEr to check the syntax, semantics,
relevance, and grammatical structure of the generated text. We have compared our results with
stateof-the-art DRS-to-Text results of authors in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and proved that augmentation is helpful in getting better
results as compared to results generated without augmentation.
      </p>
    </sec>
    <sec id="sec-7">
      <title>5. Results</title>
      <p>Results are the outcomes received after the implementation of the proposed methodology. Here we
discuss our findings and try to prove the research questions addressed previously. In the implementation
of DRS-to-Text generation, we conducted two experiments based on the types of PMB datasets. Our
first experiment is also referred to as the baseline experiment conducted on the Gold-PMB dataset. We
performed two different experiments on the gold dataset i.e., an experiment without augmentation on
the PMB-Gold dataset, and an experiment with augmentation on the PMB-Gold dataset. We analyzed
character-level and word-level results of the model and achieved high evaluation scores in all formats
of evaluation metrics. Baseline results are mentioned in Table 4 with all descriptions of the dataset and
evaluation metrics.</p>
      <p>Our second experiment is based on certain findings: first, if we add training examples of Silver-PMB
data (not fully manually annotated corpus) with Gold-PMB data (fully annotated corpus), will it also
go for an increase in evaluation scores? Secondly, can we achieve higher evaluation scores as compared
to the Gold-PMB augmentation? Finally, we also must compare our augmentation-based results with
literature models. So, to prove our hypothesis, we augmented the Gold and Silver PMB training
examples and conducted the experiment. We succeeded in achieving high evaluation scores of all
metrics but this time the score was not as high as we achieved in the Gold-PMB experiment. This is
possibly due to the addition of certain DRS examples which were not fully manually annotated by the
experts. A noise in SILVER-PMB data propagated through all the variants of dataset with and without
augmentation. This causes into less increase in evaluation scores. Just like the augmentation results of
Gold-PMB, we also analyzed character-level and word-level results of the neural model. We also
compared the results with the literature and our implementation of the model with and without
augmentation. All results are mentioned in Table 4.</p>
      <p>The table reflects the successful implementation of our proposed hypothesis. In the literature, to the
best of our knowledge, there is no implementation of augmentation in DRS but there are other
implementations of DRS for language translations. To strengthen our hypothesis, we conducted a
baseline experiment on a fully manually annotated gold corpus. Our baseline experiment strengthens
our claim and then we further embedded Silver data into Gold and performed augmentation tasks. The
first 2 experimental findings are of baseline experiments with and without augmentation. It is clearly
shown in a bold format that we achieved efficient results for the augmented version of the DRS-to-Text
implementation. The remaining 3 experiments are listed as the literature-based implementation of the
author in 3rd row of Table 4. The 4th and 5th rows are our implementations on the gold and silver
datasets with and without augmentation. And the 5th row (in bold) also highlights our
augmentationbased results as the high scorer in its regard.</p>
      <p>Statistical Significance Tests. To prove our model’s achievement statistically, we conducted certain
statistical significance tests as well [32]. Significance tests are becoming a new trend in the NLG domain
nowadays. Significance tests are applied when two different models are applied to the same data, or the
same model is applied to two different datasets. In our case, we applied the same bi-LSTM-based
sequence-to-sequence model on two different data samples i.e., dataset without augmentation and
dataset with augmentation. The purpose of doing these tests is to verify that the good results of one
model are not achieved accidentally. Therefore, among a series of parametric and non-parametric tests,
we choose the right test for our experiment based on two findings. First, we determined whether our
data is normally distributed or not.</p>
      <p>To check the normality of the data, we conducted Shapiro-Wilk Test. We choose this test because it
is highly effective as compared to other tests used to check data normality. In our case, our data were
not normally distributed and therefore we have to move towards non-parametric tests. If our data was
normally distributed, then only a t-test would be enough to check model significance [32]. Among a list
of non-parametric tests, we choose Wilcoxon Test due to two reasons. First, we choose the Wilcoxon
test because it is highly suitable for the data which is coming from automatic evaluation metrics e.g.,
BLEU, ROUGE, METEOR, etc. Secondly, we choose this because it has the highest statistical
significance as compared to other non-parametric tests working on scores coming from automatic
evaluation metrics.</p>
      <p>For the implementation of significance tests, we calculated the sentence-wise score of BLEU for
model-generated test data and Gold reference data having approximately 1K examples. We conducted
character level and word level significance tests and found that our augmentation models are
significantly better with p-value = 2.37e-05 for the Char-level model and p-value = 7.78e-07 for the
Word-level model.</p>
    </sec>
    <sec id="sec-8">
      <title>6. Conclusion and Future Work</title>
      <p>Data augmentation is a very challenging task in NLP and NLG. The main goal of augmentation is
to increase training examples for the neural model without explicitly adding new data for training. In
this contrast, we have implemented a data augmentation approach in DRS for text generation tasks. We
conducted two experiments on PMB gold and gold-silver datasets. We achieved high evaluation scores
of BLEU, ROUGE, METEOR, NIST, and CIDEr in the case of a model trained on augmented data.
Furthermore, we conducted statistical significance tests to prove model performance on both
characterlevel and word-level translations. We found that our augmentation models are significantly better with
p-value = 2.37e-05 for Char-level model and p-value = 7.78e-07 for Word-level model.</p>
      <p>In future, we will extend this experiment by applying other data augmentation approaches on logical
forms (DRS) with respect to polarity change, number change, quantity change, and name change in the
same DRS. We are also focusing on applying augmentation on low-resource languages like ITALIAN,
FRENCH, and DUTCH.</p>
    </sec>
    <sec id="sec-9">
      <title>7. References</title>
      <p>[13] Gözde Gül ¸Sahin and Mark Steedman. 2018. Data augmentation via dependency tree morphing
for lowresource languages. In Proceedings of the 2018 Conference on Empirical Methods in
Natural Language Processing, pages 5004–5009, Brussels, Belgium. Association for
Computational Linguistics.
[14] Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz. 2017. mixup: Beyond
empirical risk minimization. Proceedings of ICLR.
[15] Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon
Yoo. 2019. Cutmix: Regularization strategy to train strong classifiers with localizable features. In
Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6023–6032.
[16] Terrance DeVries and Graham W Taylor. 2017. Improved regularization of convolutional neural
networks with cutout. arXiv preprint.
[17] Golnaz Ghiasi, Yin Cui, Aravind Srinivas, Rui Qian, Tsung-Yi Lin, Ekin D. Cubuk, Quoc V. Le,
and Barret Zoph. 2020. Simple copy-paste is a strong data augmentation method for instance
segmentation. arXiv preprint.
[18] Demi Guo, Yoon Kim, and Alexander Rush. 2020. Sequence-level mixed sample data
augmentation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language
Processing (EMNLP), pages 5547–5552, Online. Association for Computational Linguistics.
[19] Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016. Improving Neural Machine Translation
Models with Monolingual Data. In Proceedings of the 54th Annual Meeting of the Association for
Computational Linguistics (Volume 1: Long Papers), pages 86–96, Berlin, Germany. Association
for Computational Linguistics.
[20] John Wieting and Kevin Gimpel. 2017. Revisiting Recurrent Networks for Paraphrastic Sentence
Embeddings. In Proceedings of the 55th Annual Meeting of the Association for Computational
Linguistics (Volume 1: Long Papers), pages 2078–2088, Vancouver, Canada. Association for
Computational Linguistics.
[21] Steven Y. Feng, Aaron W. Li, and Jesse Hoey. 2019. Keep calm and switch on! Preserving
sentiment and fluency in semantic text exchange. In Proceedings of the 2019 Conference on
Empirical Methods in Natural Language Processing and the 9th International Joint Conference on
Natural Language Processing (EMNLP-IJCNLP), pages 2701–2711, Hong Kong, China.</p>
      <p>Association for Computational Linguistics.
[22] Sosuke Kobayashi. 2018. Contextual augmentation: Data augmentation by words with
paradigmatic relations. In Proceedings of the 2018 Conference of the North American Chapter of
the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short
Papers), pages 452–457, New Orleans, Louisiana. Association for Computational Linguistics.
[23] Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor, George Kour, Segev
Shlomov, Naama Tepper, and Naama Zwerdling. 2020. Do not have enough data? Deep learning
to the rescue! In Proceedings of AAAI, pages 7383–7390.
[24] Jasdeep Singh, Bryan McCann, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher. 2019.</p>
      <p>Xlda: Cross-lingual data augmentation for natural language inference and question answering.
arXiv preprint arXiv:1905.11471.
[25] Rongzhi Zhang, Yue Yu, and Chao Zhang. 2020. SeqMix: Augmenting Active Sequence Labeling
via Sequence Mixup. In Proceedings of the 2020 Conference on Empirical Methods in Natural
Language Processing (EMNLP), pages 8566–8579, Online. Association for Computational
Linguistics.
[26] Samuel Louvan and Bernardo Magnini. 2020. Simple is better! lightweight data augmentation for
low resource slot filling and intent classification. In Proceedings of the 34th Pacific Asia
Conference on Language, Information and Computation, pages 167– 177, Hanoi, Vietnam.</p>
      <p>Association for Computational Linguistics.
[27] Vaibhav Vaibhav, Sumeet Singh, Craig Stewart, and Graham Neubig. 2019. Improving Robustness
of Machine Translation with Synthetic Noise. In Proceedings of the 2019 Conference of the North
American Chapter of the Association for Computational Linguistics: Human Language
Technologies, Volume 1 (Long and Short Papers), pages 1916–1920, Minneapolis, Minnesota.</p>
      <p>Association for Computational Linguistics.
[28] Fei Gao, Jinhua Zhu, Lijun Wu, Yingce Xia, Tao Qin, Xueqi Cheng, Wengang Zhou, and Tie-Yan
Liu. 2019. Soft contextual data augmentation for neural machine translation. In Proceedings of the
57th Annual Meeting of the Association for Computational Linguistics, pages 5539–5544,
Florence, Italy. Association for Computational Linguistics.
[29] Xuan-Phi Nguyen, Shafiq Joty, Kui Wu, and Ai Ti Aw. 2020. Data diversification: A simple
strategy for neural machine translation. In Advances in Neural Information Processing Systems,
volume 33, pages 10018–10029. Curran Associates, Inc.
[30] Ashutosh Kumar, Satwik Bhattamishra, Manik Bhandari, and Partha Talukdar. 2019a. Submodular
optimization-based diverse paraphrasing and its effectiveness in data augmentation. In Proceedings
of the 2019 Conference of the North American Chapter of the Association for Computational
Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 3609–
3619, Minneapolis, Minnesota. Association for Computational Linguistics.
[31] Nandan Thakur, Nils Reimers, Johannes Daxenberger, and Iryna Gurevych. 2021. Augmented
sbert: Data augmentation method for improving bi-encoders for pairwise sentence scoring tasks.</p>
      <p>Proceedings of NAACL.
[32] Dror, R., Baumer, G., Shlomov, S., &amp; Reichart, R. (2018, July). The hitchhiker’s guide to testing
statistical significance in natural language processing. In Proceedings of the 56th Annual Meeting
of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 1383-1392).
[33] D. Nozza, L. Passaro, M. Polignano, Preface to the Sixth Workshop on Natural Language for
Artificial Intelligence (NL4AI), in: D. Nozza, L. C. Passaro, M. Polignano (Eds.), Proceedings of
the Sixth Workshop on Natural Language for Artificial Intelligence (NL4AI 2022) co-located with
21th International Conference of the Italian Association for Artificial Intelligence (AI*IA 2022),
November 30, 2022, CEUR-WS.org, 2022.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Yutai</given-names>
            <surname>Hou</surname>
          </string-name>
          , Yijia Liu, Wanxiang Che, and Ting Liu.
          <year>2018</year>
          .
          <article-title>Sequence-to-sequence data augmentation for dialogue language understanding</article-title>
          .
          <source>In Proceedings of the 27th International Conference on Computational Linguistics</source>
          , pages
          <fpage>1234</fpage>
          -
          <lpage>1245</lpage>
          ,
          <string-name>
            <given-names>Santa</given-names>
            <surname>Fe</surname>
          </string-name>
          , New Mexico, USA. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Connor</given-names>
            <surname>Shorten and Taghi M Khoshgoftaar</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>A survey on Image Data Augmentation for Deep Learning</article-title>
          .
          <source>Journal of Big Data</source>
          ,
          <volume>6</volume>
          (
          <issue>1</issue>
          ):
          <fpage>60</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Ruibo</given-names>
            <surname>Liu</surname>
          </string-name>
          , Guangxuan Xu, Chenyan Jia, Weicheng Ma, Lili Wang, and
          <string-name>
            <given-names>Soroush</given-names>
            <surname>Vosoughi</surname>
          </string-name>
          .
          <year>2020b</year>
          .
          <article-title>Data boost: Text data augmentation through reinforcement learning guided conditional generation</article-title>
          .
          <source>In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)</source>
          , pages
          <fpage>9031</fpage>
          -
          <lpage>9041</lpage>
          , Online. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Rik</surname>
            <given-names>van Noord</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lasha</surname>
            <given-names>Abzianidze</given-names>
          </string-name>
          , Antonio Toral, and
          <string-name>
            <given-names>Johan</given-names>
            <surname>Bos</surname>
          </string-name>
          . 2018b.
          <article-title>Exploring neural methods for parsing discourse representation structures</article-title>
          .
          <source>Transactions of the Association for Computational Linguistics</source>
          ,
          <volume>6</volume>
          :
          <fpage>619</fpage>
          -
          <lpage>633</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>van</surname>
            <given-names>Noord</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Bisazza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            , &amp;
            <surname>Bos</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          (
          <year>2021</year>
          ).
          <article-title>Evaluating Text Generation from Discourse Representation Structures</article-title>
          . In A.
          <string-name>
            <surname>Bosselut</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Durmus</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Prashant Gangal</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Gehrmann</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Jernite</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Perez-Beltrachini</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Shaikh</surname>
          </string-name>
          , &amp; W. Xu (Eds.),
          <source>Proceedings of the 1st Workshop on Natural Language Generation</source>
          , Evaluation, and
          <string-name>
            <surname>Metrics</surname>
          </string-name>
          (GEM
          <year>2021</year>
          )
          <article-title>(pp</article-title>
          .
          <fpage>73</fpage>
          -
          <lpage>83</lpage>
          ).
          <article-title>Association for Computational Linguistics (ACL)</article-title>
          . https://doi.org/10.18653/v1/
          <year>2021</year>
          .gem-
          <volume>1</volume>
          .8.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Jason</given-names>
            <surname>Wei</surname>
          </string-name>
          and
          <string-name>
            <given-names>Kai</given-names>
            <surname>Zou</surname>
          </string-name>
          .
          <year>2019</year>
          . EDA:
          <article-title>Easy data augmentation techniques for boosting performance on text classification tasks</article-title>
          .
          <source>In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)</source>
          , pages
          <fpage>6382</fpage>
          -
          <lpage>6388</lpage>
          ,
          <string-name>
            <surname>Hong</surname>
            <given-names>Kong</given-names>
          </string-name>
          , China. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Feng</surname>
            ,
            <given-names>S. Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gangal</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wei</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chandar</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vosoughi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mitamura</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Hovy</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          (
          <year>2021</year>
          ).
          <article-title>A survey of data augmentation approaches for NLP</article-title>
          .
          <source>arXiv preprint arXiv:2105</source>
          .
          <fpage>03075</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Xiang</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <source>Junbo Zhao, and Yann LeCun</source>
          .
          <year>2015</year>
          .
          <article-title>Character-Level Convolutional Networks for Text Classification</article-title>
          .
          <source>In Proceedings of the 28th International Conference on Neural Information Processing Systems - Volume 1, NIPS'15, page 649-657</source>
          , Cambridge, MA, USA. MIT Press.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Ashutosh</given-names>
            <surname>Kumar</surname>
          </string-name>
          , Satwik Bhattamishra, Manik Bhandari, and
          <string-name>
            <given-names>Partha</given-names>
            <surname>Talukdar</surname>
          </string-name>
          . 2019a.
          <article-title>Submodular optimization-based diverse paraphrasing and its effectiveness in data augmentation</article-title>
          .
          <source>In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>1</volume>
          (Long and Short Papers), pages
          <fpage>3609</fpage>
          -
          <lpage>3619</lpage>
          , Minneapolis, Minnesota. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Eli</surname>
            <given-names>Schwartz</given-names>
          </string-name>
          , Leonid Karlinsky, Joseph Shtok, Sivan Harary, Mattias Marder, Abhishek Kumar, Rogerio Feris, Raja Giryes, and
          <string-name>
            <surname>Alex M Bronstein</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>δencoder: an effective sample synthesis method for few-shot object recognition</article-title>
          .
          <source>In Proceedings of the 32nd International Conference on Neural Information Processing Systems</source>
          , pages
          <fpage>2850</fpage>
          -
          <lpage>2860</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Jason</given-names>
            <surname>Wei</surname>
          </string-name>
          and
          <string-name>
            <given-names>Kai</given-names>
            <surname>Zou</surname>
          </string-name>
          .
          <year>2019</year>
          . EDA:
          <article-title>Easy data augmentation techniques for boosting performance on text classification tasks</article-title>
          .
          <source>In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)</source>
          , pages
          <fpage>6382</fpage>
          -
          <lpage>6388</lpage>
          ,
          <string-name>
            <surname>Hong</surname>
            <given-names>Kong</given-names>
          </string-name>
          , China. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Hannah</surname>
            <given-names>Chen</given-names>
          </string-name>
          , Yangfeng Ji, and
          <string-name>
            <given-names>David</given-names>
            <surname>Evans</surname>
          </string-name>
          .
          <year>2020b</year>
          .
          <article-title>Finding friends and flipping frenemies: Automatic paraphrase dataset augmentation using graph theory</article-title>
          .
          <source>In Findings of the Association for Computational Linguistics: EMNLP</source>
          <year>2020</year>
          , pages
          <fpage>4741</fpage>
          -
          <lpage>4751</lpage>
          , Online. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>