<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Deep Learning meets Post-modern Poetry</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Timo Baumann</string-name>
          <email>baumann@informatik.uni-hamburg.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Burkhard Meyer-Sickendiek</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Informatics, Universitat Hamburg</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Literary Studies, Freie Universitat Berlin</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <fpage>30</fpage>
      <lpage>36</lpage>
      <abstract>
        <p>We summarize our project Rhythmicalizer in which we analyze a corpus of post-modern poetry in a combination of qualitative hermeneutical and computational methods, as we have run the project over the course of the past three years (and preparing it for some time before that). Interdisciplinary work is always challenging and we here focus on some of the highlights of our collaboration.</p>
      </abstract>
      <kwd-group>
        <kwd>Literary Studies</kwd>
        <kwd>Machine Learning</kwd>
        <kwd>Meta-Research</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>http://www.rhythmicalizer.net</p>
    </sec>
    <sec id="sec-2">
      <title>Introduction</title>
      <p>At least 80 % of modern and post-modern poems exhibit neither rhyme nor
metrical schemes like iamb or trochee. However, does this mean that they are
free of any rhythmical features? Of course not and the US American research
on free verse prosody claims the opposite: Modern poets like Whitman, the
Imagists, the Beat poets as well as contemporary Slam poets have developed a
post-metrical idea of prosody, using rhythmical features of everyday language,
prose, and musical styles like Jazz or Hip Hop, yielding a large and complex
variety in their poetic prosodies which, however, appear to be much harder to
quantify and regularize than traditional patterns.</p>
      <p>
        In our joint project, we examine the largest portal for spoken poetry
Lyrikline1 and analyze and classify such rhythmical patterns in a human-in-the-loop
approach in which we interleave manual annotation with computational modelling
and data-based analysis [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>The remainder of this paper is structured as follows: in the next section,
we describe the research questions that we set out to address in our project;
Copyright 2020 for this paper by its authors. Use permitted under Creative Commons
License Attribution 4.0 International (CC BY 4.0).</p>
      <p>This work is funded by the Volkswagen Foundation in the programme `Mixed Methods
in the Humanities? Funding possibilities for the combination and the interaction of
qualitative hermeneutical and digital methods' (funding codes 91926 and 93255). We
wish to thank Hussein Hussein for his valuable contributions to the project.
1 http://lyrikline.org
in Sections 3 and 4, we describe the methods used and the results obtained,
respectively, and in Section 5 we describe { as much as is adequate in a public
setting { the lessons taken during the course of the project. (It is still unclear, if
all of these lessons also were lessons learned.)
2</p>
    </sec>
    <sec id="sec-3">
      <title>Research Questions</title>
      <p>
        Automating the analysis and di erentiating the prosodies of free verse
postmodern poetry comes with a multitude of challenges, including the manual
analysis and classi cation of su ciently many poems as training material for
an automated method. Yet, the amount of data that can be acquired (even in
the order of hundreds of poems) is still very little for modern machine learning
methods. However, the by far largest challenge of this endeavour lies in the nature
of free verse prosody, which is an almost purely spoken aspect of a poem and
is not readily observable from the pure textual form as in traditional metric
schemes (see, e.g., [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]). We thus base our analyses on the audible form of the
poem, spoken aloud by the original author, as a gold standard for the intended
prosodic realization of the poem (instead of, e.g., attempting to derive such
features in a generic way from the written form). The use of speech data greatly
improves our modelling ability, yet it also increases the complexity of the data
and the need for automated analyses.
      </p>
      <p>Our primary research questions are hence: (a) can we relate theories on the
classi cation of the prosodies found in post-modern poetry to poems that we
nd our collection of spoken poems, (b) can we pre-process the poems in our
collections in such a way that they can be automatically analyzed, (c) can we
build an automatic classi cation system that uses quantitative features derived
from the poems to yield the classi cation despite the signi cant data sparsity,
and (d) can we gather additional, new insight from the automatic classi cation
in the humanities.
3</p>
    </sec>
    <sec id="sec-4">
      <title>Method</title>
      <p>We here describe the setup of our project, in particular how we went about
de ning and implementing the philological classi cation, as well as the processes
used to prepare and automatically classify our data. Over the course of the
project, we have extended and sometimes changed our methods, as is re ected in
our publications, and we here only describe a subset of the methods used.
3.1</p>
      <p>
        Theoretical Foundation: Grammetrical Ranking and Rhymic
Phrasing
Our classi cation of poetry is based on two theoretical approaches: 1) the Idea
of grammetrical ranking and 2) the Idea of rhythmic phrasing. 1) The term
grammetrics, coined by Donald Wesling, is a hybridization of grammar and
31/143
metrics: the key hypothesis is that the interplay of sentence-structure and
linestructure can be accounted for more economically by simultaneous than by
successive analysis. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] In poetry as a kind of versi ed language, the singular
sentence interacts with verse periods (syllable, foot, part-line, line, rhymed pair
or stanza, whole poem), a process for which Wesling nds `scissoring' an apt
metaphor: \Grammetrics assumes that meter and grammarcan be scissored by
each other, that the cutting places can be graphed with some precision. One
blade of the shears is meter, the other grammar. When they work against each
other, they divide the poem. It is their purpose and necessity to work against
each other." [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]
      </p>
      <p>
        2) The concept of rhythmic phrasing was developed by Richard Cureton. For
Cureton, rhythm embraces what has traditionally been regarded as very di erent
kinds of perceptual phenomena. Cureton divided rhythm - a global term covering
all relations of strength and weakness - into three distinct components, which he
terms meter, grouping, and prolongation. Meter involves the perception of beats in
regular patterns; grouping involves the apprehension of linguistic units organized
around a single peak of prominence; and prolongation involves the experience of
anticipation and arrival. Cureton claimed a hierarchical, multi-dimensional, and
preferential treatment of poetic rhythm, going back to [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] treatment of rhythm in
Tonal Music. Lerdahl and Jackendo claimed four di erent rhythmic dimensions:
a) The grouping structure in terms of a hierarchical segmentation, b) the metrical
structure in terms of a regular alternation of strong and weak beats at a number
of hierarchical levels, c) the time span-reduction in terms of an organizations
uniting time-spans at all temporal levels of a work, and d) the prologational
reduction in terms of a `psychological' awareness of tensing and relaxing patterns
in a given piece.
      </p>
      <p>As a result, the prosodic hierarchy to be considered for free verse is
considerably more complex than for metric poetry and our working hypothesis of this
hierarchy is depicted in Figure 1 (left). As can be seen, all levels of the linguistic
hierarchy can carry poetic prosodic meaning, from the segment up to the periodic
sentence. Based on this prosodic hierarchy, Figure 1 (right) depicts a
categorization of some poets' works along two axes, the governing prosodic unit (x-axis)
and the degree of iso/heterochronicity, or regularity of temporal arrangement
(y-axis). The respective rhythm derives from the combined prosodic units and the
degree of its isochronic (or heterochronic) succession. The green lines are meant
as a localization of poets according to the interplay of grammetrical ranking and
prosodic succession. The blue brackets mark the time span reduction (Rubato)
as well as the prolongational reduction (Phrasing). According to the idea of
time-span-reduction, the rubato is caused by a deviation from the isochronous
rhythm. And according to the idea of prolongational reduction, the phrasing is
divided into three di erent articulation techniques (legato, portato, staccato).
3.2</p>
      <p>Manual Analysis and Annotation
Based on our theoretical foundation, we devised a number of prosodic classes of
poems and built a manually curated collection of poems for each class. Over the
32/143
course of our project, we extended the kinds of poems to di erentiate, starting
with only various kinds of sound poetry (letristic vs. syllabic decompositions) via
kinds of poems that di er by the way that enjambments are realized (variable foot
poems vs. unemphasized enjambments vs. gestic rhythms) to a large collection
of 18 classes. Presently, we work on identifying and specifying the relation of the
classes to each other which will help inform the automatic classi cation methods
described below.
3.3</p>
      <p>Data Extraction and Preparation
We work with the website Lyrikline which hosts a large collection of modern and
post-modern readout poetry. Lyrikline was created by the Literaturwerkstatt
Berlin and hosts contemporary international poetry as audio les (read by the
authors) and texts (original versions &amp; translations), so it o ers the melodies,
sounds, and rhythms of international poetry, recited by the authors themselves.
It covers more than 10,000 poems from about 1000 international poets from more
than 60 di erent countries. Nearly 80 % of these are postmetrical poems. We
focus our analysis on the roughly 2400 German-language poems on the page.2</p>
      <p>
        To enable our analyses, we perform forced alignment [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] (which we re ne
manually as necessary) in order to be able to relate speech and text line-by-line.
2 An extraction program for extracting poems (text, audio and meta-data) is available
by request from the rst author.
33/143
concatenated representation
r
e
d
o
c
n
e
r
e
d
o
c
n
e
line
representation
encoder
      </p>
      <p>attention
RNN RNN ...</p>
      <p>RNN RNN ...</p>
      <p>RNN</p>
      <p>RNN
...
characters of poem line
acoustic features of
features of following
poem line pause
line representation
.
.</p>
      <p>.
line representation
class
softmax
decision
layer</p>
      <p>In addition, we use various tools to extract syntactic information from the text
as required.
3.4</p>
      <p>Classi cation Methods
We devised our project with primary concerns on data sparsity (i.e., the limitation
that there is too little data to fully train advanced machine learning algorithms)
as well as interpretability (i.e., the ability of a method to explain what aspects of
the data determine its behaviour).</p>
      <p>Thus, we rst focused our e orts on building interpretable classi ers (e.g.,
decision trees) on speci c interpretable features (e.g., the presence or absence of a
verb in a line). However, based on the (relatively obvious) fact that the prosodic
structure of a poem is re ected in most of its lines, and that our interpretable
features need to be aggregated across the poem's lines anyway, we also build a
hierarchical model for our poems, as depicted in Figure 2.</p>
      <p>
        Our deep-learning based hierarchical attention model [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] uses multiple
recurrent neural-network-based encoders for each line, focusing on the text, the
speech, and the quality of the pause following the line, respectively, and using
inner-attention. We aggregate across the lines with another recurrent layer. The
architecture is more thoroughly described in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Given the fact that most lines
exhibit the structural properties that yield a poem's prosodic classi cation, we
use a two-stage training procedure, in which we rst train the line-by-line
classication in isolation and only afterwards train the full network including the nal
decision layer.
34/143
      </p>
    </sec>
    <sec id="sec-5">
      <title>Results</title>
      <p>Over the course of our project, we have performed a number of experiments; we
here only give a very rough overview of the results and refer the reader to the
corresponding publications for further detail.</p>
      <p>
        In [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], we described our rst experiments on classifying poetry using a
hierarchical attention network. We found that free verse prosodies can be classi ed
along a uency continuum and that the classi er's mis-classi cations cluster along
this uency continuum. Via ablation studies (i.e., leaving out certain features to
nd their relative importance for the nal results), we found that, depending on
the classes analyzed, not only is the acoustic realization of the speech itself highly
relevant for classi cation but also the realization of the pause following each
line of the poem. In fact, our classi er is highly reliable in determining di erent
kinds of enjambments [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and that an additional annotation of enjambments in
the poems is unnecessary (unlike what we had originally hypothesized). We also
tried traditional classi cation approaches using engineered features [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] and in a
comparison with our deep learning-based method found these to be inferior [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ],
although they have direct access on theoretically grounded features. Recently, we
work on integrating such symbolic features into our deep learning-based approach.
      </p>
      <p>We nd that pre-training and data augmentation are crucial for the success
of neural learning techniques on the (relatively) small data sets typically found
in digital humanities. Another way to success on di cult-to-describe material
like sound poetry is the granularity of modelling: instead of using words as the
textual unit of processing, our recurrent networks use the sequence of individual
characters found in the poem. This has (at least) the following two advantages:
there is a xed set of characters (as compared to the endless possibilities for
words) which means that our models has far fewer parameters to train and
cannot be confronted with `out-of-vocabulary words' when applied to a new
poem. Secondly, the concept of `word' does not necessarily re ect what we nd
in sound poetry and the prosodic features are actually very well re ected by the
stream of characters (e.g., the recurrence of consonant-vowel pairs) as compared
to words.</p>
      <p>
        Finally, and coming back to our original ideas [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], we have implemented an
interface for humanist-in-the-loop classi cation and analysis of our corpus [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and
have gradually applied our method to eventually cover the full corpus.
5
      </p>
    </sec>
    <sec id="sec-6">
      <title>Lessons (Learned?)</title>
      <p>Over the course of our interdisciplinary project, the common research question
has been a strong uniting factor given the sometimes di ering research interests
involved (creating classi cation systems that work despite data sparsity vs. the
interest in nding evidence in data for humanistic theories). A strong enabling
factor in the joint work was the ability to openly ask about the other eld's
theory, applications, limitations, and to not be shy to ask again if the given
answer was unclear. It often helped to agree to disagree and to move on despite
of this, or to compromise on what should be done.
35/143</p>
      <p>The project sometimes was hindered by technical limitations, such as forced
alignment of text to speech (or syntactic analysis) breaking down for sound poetry,
and often being too inaccurate (or with too low coverage) for large amounts of the
material. This meant that a great deal of manual annotation of the base material
had to be performed, despite the later stages being successfully automated. This
means that the classi ers produced in the project cannot always be applied
to new data in a fully automatic way and the work required to prepare data
(although relatively straightforward) can be more than would be required for a
manual prosodic analysis itself.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Baumann</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hussein</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <article-title>Burkhard: Analysis of rhythmic phrasing: Feature engineering vs. representation learning for classifying readout poetry</article-title>
          .
          <source>In: Proceedings of the Joint LaTeCH&amp;CLfL Workshop</source>
          . pp.
          <volume>44</volume>
          {
          <fpage>49</fpage>
          . Association for Computational Linguistics, Santa Fe, USA (Sep
          <year>2018</year>
          ), https://aclanthology.info/ papers/W18-4505/w18-
          <fpage>4505</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Baumann</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hussein</surname>
            , H., Meyer-Sickendiek,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Analysing the focus of a hierarchical attention network: The importance of enjambments when classifying post-modern poetry</article-title>
          .
          <source>In: Proceedings of Interspeech</source>
          . pp.
          <volume>2162</volume>
          {
          <fpage>2166</fpage>
          .
          <string-name>
            <surname>Hyderabad</surname>
          </string-name>
          ,
          <string-name>
            <surname>India</surname>
          </string-name>
          (Sep
          <year>2018</year>
          ). https://doi.org/10.21437/Interspeech.2018-
          <fpage>2533</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Baumann</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hussein</surname>
            , H., Meyer-Sickendiek,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Style detection for free verse poetry from text and speech</article-title>
          .
          <source>In: Proceedings of the 27th International Conference on Computational Linguistics (COLING</source>
          <year>2018</year>
          ). pp.
          <year>1929</year>
          {
          <year>1940</year>
          . Santa Fe, USA (Aug
          <year>2018</year>
          ), https://aclanthology.info/papers/C18-1164/c18-
          <fpage>1164</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Baumann</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hussein</surname>
            , H., Meyer-Sickendiek,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elbeshausen</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>A tool for human-in-the-loop analysis and exploration of (not only) prosodic classi cations for post-modern poetry</article-title>
          .
          <source>In: Proceedings of INF-DH</source>
          . pp.
          <volume>151</volume>
          {
          <fpage>156</fpage>
          . Gesellschaft fur Informatik, Kassel, Germany (Sep
          <year>2019</year>
          ). https://doi.org/10.18420/inf2019 ws15
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Baumann</surname>
            , T., Meyer-Sickendiek,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Large-scale analysis of spoken free-verse poetry</article-title>
          .
          <source>In: Proceedings of Language Technology Resources and Tools for Digital Humanities (LT4DH)</source>
          . Osaka,
          <string-name>
            <surname>Japan</surname>
          </string-name>
          (Dec
          <year>2016</year>
          ), https://www.aclweb.org/anthology/ W16-4017
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Bobenhausen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>The metricalizer2{automated metrical markup of german poetry</article-title>
          .
          <source>Current Trends in Metrical Analysis</source>
          , Bern: Peter Lang pp.
          <volume>119</volume>
          {
          <issue>131</issue>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Hussein</surname>
            , H., Meyer-Sickendiek,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baumann</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Automatic detection of enjambment in german readout poetry</article-title>
          .
          <source>In: Proceedings of Speech Prosody. Poznan, Poland (Jun</source>
          <year>2018</year>
          ). https://doi.org/10.21437/SpeechProsody.2018-
          <fpage>67</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Katsamanis</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Black</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Georgiou</surname>
            ,
            <given-names>P.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goldstein</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Narayanan</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>SailAlign: Robust long speech-text alignment</article-title>
          .
          <source>In: Proc. of Workshop on New Tools and Methods for Very-Large Scale Phonetics Research</source>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Lerdahl</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jackendo</surname>
          </string-name>
          , R.:
          <article-title>A Generative Theory of Tonal Music</article-title>
          . MIT Press series
          <article-title>on cognitive theory and mental representation</article-title>
          , MIT Press (
          <year>1983</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Wesling</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          : The Scissors of Meter: Grammetrics and Reading. University of Michigan Press (
          <year>1996</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dyer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>He</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smola</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hovy</surname>
          </string-name>
          , E.:
          <article-title>Hierarchical attention networks for document classi cation</article-title>
          .
          <source>In: Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          . pp.
          <volume>1480</volume>
          {
          <issue>1489</issue>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>