<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Thessaloniki, Greece
$ stefan.reicho@student.uibk.ac.at (S. Reicho); adam.jatowt@uibk.ac.at (A. Jatowt)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Innsbruck @ JOKER2023 Task 1: Data Augmentation Techniques for Humor Recognition in Text</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Stefan Reicho</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Adam Jatowt</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Innsbruck</institution>
          ,
          <addr-line>6020 Innsbruck</addr-line>
          ,
          <country country="AT">Austria</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>We explore the use of data augmentation techniques and their impact on pun detection in English, leveraging the oficial JOKER Task 1 - English Pun Detection dataset. The paper reports on various data augmentation strategies such as synonym replacement, back-translation, paraphrasing, shortening, and extension, and their respective impact on pun detection performance. An important aspect of the study is ensuring that the augmentation techniques retain the humorous efect present in the original text. The results from the experiments suggest however that the augmentation techniques did not substantially enhance pun detection capabilities.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;data augmentation</kwd>
        <kwd>humor detection</kwd>
        <kwd>JOKER</kwd>
        <kwd>wordplay</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Humor is a crucial aspect of human communication that lends nuance to interactions and
enhances our understanding of complex situations. One of the primary forms of humor in
English is punning or wordplay, characterized by the use of words in such a way that multiple
meanings or efects occur. The ability to detect puns in text is an interesting albeit challenging
task in Natural Language Processing (NLP).</p>
      <p>
        This paper presents our empirical study of the impact of various data augmentation techniques
on the performance of pun detection, based on the oficial JOKER Task 1 - English Pun Detection
dataset [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. In particular, we explore techniques such as synonym replacement, back-translation,
paraphrasing, shortening, and extension. Our study utilizes the transformer-based BERT[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]
model as the underlying classifier for the pun detection task.
      </p>
      <p>We present a detailed discussion on the various data augmentation techniques used, as
well as the measures taken to ensure the quality of augmented data. Despite the multiple
techniques employed, our study suggests that the data augmentation methods we apply have
not significantly enhanced pun detection capabilities. However, this paper ofers some insights
into pun detection and the potential influence of data augmentation techniques, providing a
foundation for future research in this area.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Methodology</title>
      <sec id="sec-2-1">
        <title>2.1. Dataset</title>
        <p>
          The primary source of data for this research was the oficial JOKER Task 1 - English Pun
Detection dataset [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Pun Detection Classifier</title>
        <p>The pun detection task was performed using a classifier model built on the transformer
architecture. The choice for the underlying pre-trained model was the BERT model, specifically the
’bert-base-uncased’ variant. This model was selected for its high efectiveness in handling a
wide range of NLP tasks, including text classification.</p>
        <p>The BERT model was fine-tuned using the Adam optimizer [ 3] with a learning rate of 1e-5
and a weight decay of 0.01. Our training procedure used a batch size of 32 and was conducted
over two epochs.</p>
        <p>To measure the performance of the model, the CrossEntropyLoss function was used.</p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Data Augmentation</title>
        <p>To further improve the ability to successfully classify sentences containing puns, without tuning
the classifier, data augmentation could be applied.</p>
        <p>Although applying diferent data augmentation techniques to humorous data containing
wordplay can be challenging without losing the humorous efect, we decided to take this path
and perform experiments in order to understand to what extent diferent data augmentation
technologies would be applicable for wordplay detection.</p>
        <p>In the following, we describe a few simple techniques that were implemented and
experimented with.</p>
        <sec id="sec-2-3-1">
          <title>2.3.1. Synonym replacement</title>
          <p>Data augmentation by substitution involves replacing parts of the original data to create slightly
diferent versions, yet keeping the underlying meaning the same. One popular substitution
technique is synonym replacement, where certain words are replaced by their synonyms. For
example the sentence "The quick brown fox jumps over the fence" could become "The fast
brown fox jumps over the fence", by replacing "quick" with one of its synonyms "fast".</p>
          <p>Synonym replacement can be implemented in a variety of ways. Two popular approaches are
Wordnet-based [4] and Word2vec-based [5] synonym augmentation. To implement synonym
replacement, this project uses the TextAugment [6] library, which provides easy-to-use methods
for data augmentation.</p>
        </sec>
        <sec id="sec-2-3-2">
          <title>2.3.2. Wordnet-based synonym augmentation</title>
          <p>TextAugment’s Wordnet-based augmentation algorithm chooses ’r’ words from a sentence and
replaces them with one of their synonyms ’s’, where ’r’ and ’s’ are randomly selected. [6] Note
that the synonyms selected by Wordnet are context insensitive. This might lead to less accurate
replacements if a word strongly depends on its context.</p>
        </sec>
        <sec id="sec-2-3-3">
          <title>2.3.3. Word2vec-based synonym augmentation</title>
          <p>Word2vec-based augmentation uses a diferent approach by utilizing a word embedding model.
Words that are close to each other in the model’s vector space are likely to have a similar
meaning, but must not strictly be synonyms. In contrast to Wordnet-based augmentation, it
can handle words based on their context, making it possible to suggest more suitable synonyms
in some situations.</p>
          <p>As for the embedding model, we are using the pre-trained Google News Word2vec model,
which was trained upon roughly 100 billion words from diferent Google News articles. However,
we encountered an incompatibility issue between the TextAugment library and the latest
gensim version (v4.3.1). Starting with gensim version 4.0.0, the ’wv’ attribute was removed
from the ’KeyedVectors’ class. To resolve this breaking change, we had to slightly modify the
TextAugment library as a workaround.</p>
        </sec>
        <sec id="sec-2-3-4">
          <title>2.3.4. Paraphrasing by back-translation</title>
          <p>Another family of data augmentation techniques is augmentation by paraphrasing. In this case,
instead of only substituting certain parts of the text, the entire text is modified by rewording
or restructuring it (by paraphrasing). By paraphrasing, the example from before: "The quick
brown fox jumps over the fence" could become "The fence is jumped over by the quick brown
fox".</p>
          <p>Back-translation is a type of paraphrasing, in which text is translated from its original
language to some other language and then back-translated to its original language. Often, this
results in a reworded or restructured sentence while retaining its original meaning.</p>
          <p>For the translation of sentences, we used the MarianMTModel [7], which was trained on
a large dataset of texts. First, each sentence is translated from English to German using the
pre-trained model ’Helsinki-NLP/opus-mt-en-de’, and then back-translated from German to
English using the ’Helsinki-NLP/opus-mt-de-en’ model. After, the same was done for English
and French using the ’Helsinki-NLP/opus-mt-en-fr’ and ’Helsinki-NLP/opus-mt-fr-en’ models.
This results in 2 new augmented samples for every input sentence.</p>
        </sec>
        <sec id="sec-2-3-5">
          <title>2.3.5. Paraphrasing using PEGASUS</title>
          <p>PEGASUS [8] stands for Pre-training with Extracted Gap-sentences for Abstractive
Summarization and is used for tasks like text summarization. However, it is also capable of rephrasing text
while retaining its original meaning. This makes PEGASUS not only suitable for summarization
but also for data augmentation.</p>
          <p>For our experiments, we use the ’tuner007/pegasus_paraphrase’ model, a variant of PEGASUS,
which was fine-tuned specifically for the task of paraphrasing.</p>
        </sec>
        <sec id="sec-2-3-6">
          <title>2.3.6. Shortening</title>
          <p>In this approach, a text or sentence is reduced in length. One way to achieve this is by removing
unnecessary details, i.e., the least important parts of the text. However, removing a word from a
sentence, in our case, a sentence that might contain pun wordplay, without altering its meaning
or humorous efect, turns out to be quite challenging. Every word in a sentence containing
wordplay might contribute to its meaning.</p>
          <p>To implement augmentation by shortening we applied the Rapid Automatic Keyword
Extraction (RAKE) algorithm[9]. RAKE selects the key words and phrases of a sentence by analyzing
the frequency of word appearance and its co-occurrence with other words in the text. By
removing some of the words that were not identified as key words, a new shortened sentence
can be formed.</p>
        </sec>
        <sec id="sec-2-3-7">
          <title>2.3.7. Extension</title>
          <p>Augmentation via extension is another widely used technique. Since there are only new words
added, but none removed, it could be considered one of the safer techniques in regards to not
altering the text’s meaning.</p>
          <p>In our implementation, we use GPT-2 [10] to generate a new additional sentence based on
the original text. This sentence is then randomly appended either at the beginning or the end
of the original text.</p>
        </sec>
      </sec>
      <sec id="sec-2-4">
        <title>2.4. Improving augmented data quality</title>
        <p>When applying any of those techniques to humorous texts containing wordplay, special care
must be taken to not lose the humorous efect. In the experiments, it was inevitable that there
were samples generated in which the humorous efect was lost. Therefore, after augmentation,
we filter out / reject some generated samples that do not meet certain criteria.</p>
        <p>The following filters were used:
• Duplicate Filter: If a generated text sample replicates the original word for word,
it is discarded. Having duplicate data in the dataset can lead to overfitting. Thus, it is
important to prevent augmentation techniques from generating duplicate samples.
• Last Words Filter: Most of the time the punning word is one of the last two words in
the sample [11]. By requiring every augmented sample to end with the same two words,
it’s possible to filter out generated samples that are likely to have lost their punning word
in the process of augmentation.
• Similarity Filter: Filtering out generated samples that are semantically "too far"
from the original sample can also ensure a higher quality of the entire augmented dataset.
Here it is done by using Sentence-BERT [12] (SBERT), more specifically using the
’allMiniLM-L6-v2’ [13] model, which is able to generate sentence embeddings that capture
semantic information. These embeddings are then used to determine the similarity
between the original and augmented samples. The similarity is calculated using the cosine
distance. If this similarity measure falls below a certain threshold, then the generated
sample will not be included in the augmented set.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Evaluation</title>
      <sec id="sec-3-1">
        <title>3.1. Test set evaluation</title>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Further experiments on train set</title>
        <p>In addition to the work we submitted to JOKER, which involved making predictions on the test
set, we conducted more experiments using only the training set, which we further divided into
testing and training subsets via the application of K-Fold Cross Validation with k = 4.</p>
        <p>Table 2 provides evaluations of diferent datasets generated using various DA techniques.
The technique "None" implies an unaltered dataset and can be used as a standard for other
techniques to be compared against. Note that due to the randomness in machine learning
processes, we noticed fluctuations of around 1 percent in the results. Therefore, it is challenging
to determine if any of those techniques provide a significant advantage for data augmentation
in detecting wordplay.</p>
        <p>According to the evaluation table, back-translation exhibits a notable increase in size, because
every sample is once Back-translated from German and French. Here, adding filters helps to
significantly reduce the data size, but it seems to have almost no efect on performance.</p>
        <p>Synonym replacement via WordNet and Word2Vec yields similar performance. The use of
WordNet with filters seems to produce the highest accuracy of 74%.</p>
        <p>Pegasus paraphrase technique’s unfiltered values are missing, because it generated so many
samples (including bad quality ones) which were too many to train the classifier on our setup.
But after applying filters, and reducing the size this technique also does not show an increased
performance.</p>
        <p>The application of filters in the extension technique results in the highest recall value (92.61%),
contributing to the best F1 score in the table (80.07%), but with only a minor improvement in
accuracy.</p>
        <p>Overall, synonym replacement (WordNet) with filters seems to be the most balanced technique
in terms of performance and accuracy. Yet, if recall and F1 score are of higher importance,
extension with filters could be a preferable choice.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusion</title>
      <p>In conclusion, we found that the accuracy of detecting puns on the JOKER Task 1 test data was
much lower compared to our experiments where we divided the training data into separate
training and testing groups. One explanation could be that training and test data have diferent
origins. Our attempts to enhance the performance via data augmentation approaches specific
to humorous text and pun wordplay did not yield substantial improvements in pun detection
capabilities. It is worth mentioning that larger language models, such as GPT-4 might have a
superior accuracy on pun detection. Therefore, it may be an area worth exploring for further
enhancement of pun detection techniques.
[3] D. P. Kingma, J. Ba, Adam: A method for stochastic optimization, 2017. arXiv:1412.6980.
[4] G. A. Miller, R. Beckwith, C. Fellbaum, D. Gross, K. J. Miller, Introduction to wordnet: An
on-line lexical database, International journal of lexicography 3 (1990) 235–244.
[5] T. Mikolov, K. Chen, G. Corrado, J. Dean, Eficient estimation of word representations in
vector space, arXiv preprint arXiv:1301.3781 (2013).
[6] V. Marivate, T. Sefara, Improving short text classification through global augmentation
methods, in: International Cross-Domain Conference for Machine Learning and
Knowledge Extraction, Springer, 2020, pp. 385–399.
[7] M. Junczys-Dowmunt, R. Grundkiewicz, T. Dwojak, H. Hoang, K. Heafield, T. Neckermann,
F. Seide, U. Germann, A. F. Aji, N. Bogoychev, A. F. T. Martins, A. Birch, Marian: Fast
neural machine translation in C++, in: Proceedings of ACL 2018, System Demonstrations,
Association for Computational Linguistics, Melbourne, Australia, 2018, pp. 116–121. URL:
https://aclanthology.org/P18-4020. doi:10.18653/v1/P18-4020.
[8] J. Zhang, Y. Zhao, M. Saleh, P. Liu, PEGASUS: Pre-training with extracted gap-sentences
for abstractive summarization, in: H. D. III, A. Singh (Eds.), Proceedings of the 37th
International Conference on Machine Learning, volume 119 of Proceedings of Machine
Learning Research, PMLR, 2020, pp. 11328–11339. URL: https://proceedings.mlr.press/v119/
zhang20ae.html.
[9] S. Rose, D. Engel, N. Cramer, W. Cowley, Automatic keyword extraction from individual
documents, Text mining: applications and theory (2010) 1–20.
[10] A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al., Language models are
unsupervised multitask learners, OpenAI blog 1 (2019) 9.
[11] L. Ermakova, A.-G. Bosser, A. Jatowt, T. Miller, The joker corpus: English–french parallel
data for multilingual wordplay recognition, in: Proceedings of the 46th International ACM
SIGIR Conference on Research and Development in Information Retrieval, SIGIR, ACM
Press, 2023.
[12] N. Reimers, I. Gurevych, Sentence-bert: Sentence embeddings using siamese bert-networks,
2019. arXiv:1908.10084.
[13] W. Wang, F. Wei, L. Dong, H. Bao, N. Yang, M. Zhou, Minilm: Deep self-attention distillation
for task-agnostic compression of pre-trained transformers, 2020. arXiv:2002.10957.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>L.</given-names>
            <surname>Ermakova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.-G.</given-names>
            <surname>Bosser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. M. P.</given-names>
            <surname>Preciado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sidorov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jatowt</surname>
          </string-name>
          ,
          <article-title>Science for fun: The CLEF 2023 JOKER track on automatic wordplay analysis</article-title>
          , in: J.
          <string-name>
            <surname>Kamps</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Crestani</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Maistro</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Joho</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Davis</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Gurrin</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          <string-name>
            <surname>Kruschwitz</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          . Caputo (Eds.),
          <source>Advances in Information Retrieval: 45th European Conference on Information Retrieval</source>
          ,
          <string-name>
            <surname>ECIR</surname>
          </string-name>
          <year>2023</year>
          , Dublin, Ireland, April 2-
          <issue>6</issue>
          , Proceedings,
          <string-name>
            <surname>Part</surname>
            <given-names>III</given-names>
          </string-name>
          , volume
          <volume>13982</volume>
          of Lecture Notes in Computer Science, Springer, Berlin, Heidelberg,
          <year>2023</year>
          , pp.
          <fpage>546</fpage>
          -
          <lpage>556</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>031</fpage>
          -28241-6_
          <fpage>63</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , Bert:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          ,
          <year>2019</year>
          . arXiv:
          <year>1810</year>
          .04805.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>