<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Comparing Joint Speech and Synchronous Speech: What Happens When We Add More Speakers?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Fred Cummins</string-name>
          <email>fred.cummins@ucd.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>School of Computer Science University College Dublin</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Eight subjects were recorded speaking texts in synchrony in groups of 1, 2, 4, 6 and 8 speakers at a time. The question asked was whether either speech timing or inter-speaker synchronisation would be manifestly di erent as more speakers speak together. Results showed no appreciable changes to overall speech timing as number of speakers varied. Examination of asynchrony scores likewise showed no e ect of number of speakers, even as texts from very di erent genres were used. The results suggest that synchronous speaking in a laboratory is importantly di erent from joint speech found in ritual, protest, and elsewhere.</p>
      </abstract>
      <kwd-group>
        <kwd>Joint Speech</kwd>
        <kwd>Synchronisation</kwd>
        <kwd>Synchronous Speech</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Background</title>
      <p>
        Joint speech is found wherever multiple people utter the same words at the
same time [
        <xref ref-type="bibr" rid="ref3 ref7">7, 3</xref>
        ]. This empirically grounded de nition serves to pick out many
important domains of human activity, including ritual and prayer, protest, the
activities of many sports fans, and a variety of practices in early education.
It is thus a fundamental part of human social activity and is found in every
known culture and throughout history. Its thematisation as an object of scienti c
inquiry is, however relatively recent [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. When faced with an underexplored topic
that extends into many forms of culturally saturated collective activities, it is
something of a creative challenge to nd experimental and modelling avenues
that can illuminate aspects of the activity that make human behaviour and
experience within these domains intelligible.
      </p>
      <p>
        One experimental avenue that has produced some limited insights is
provided by the vehicle of synchronous speech [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. This is a joint speaking task
conducted in laboratories, in which volunteers are asked to read texts selected
by the experimenter with the express goal of speaking in synchrony with one
another. Many of the more interesting characteristics of joint speech do not
survive translation to the constrained setting of the laboratory. The passion that
inspires the protesters, the piety of the faithful, or the enthusiasm of the soccer
supporter cannot be reproduced at will by reading unmotivated texts. However,
the ability to remain in time with other speakers is interesting in its own right,
and this, more mechanical, aspect of joint speech lends itself to study using the
synchronous speech approach.
      </p>
      <p>
        Past work on synchronous speaking has used dyads (pairs of speakers)
exclusively. The body of ndings includes the observation that speakers can
synchronise e ectively and without practice [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. This is interesting in its own right
as speaking is an intrinsically plastic activity that adapts to context and to
cospeakers e ortlessly, yet in the laboratory, such expressive and context-sensitive
variability is discarded without any obvious e ort. In the constrained situation
of a synchronous speech experiment, an average asynchrony of about 40 ms. is
typical. Asynchrony after a long pause is slightly greater, but collapses back to
the 40 ms. level within a syllable or two. If speakers are back to back and thus
cannot see each other, asynchrony at phrase starts is slightly elevated by about
10 ms. but again, it quickly relaxes to the same values. The synchrony elicited
by this experimental methodology is not strictly comparable to that found in
the wild, e.g. during group prayers or in the street. One way in which it di ers
is in the production of a speci c kind of speech error unique to the synchronous
speaking situation. It has been repeatedly observed that a small hesitation or
error on the part of one speaker can frequently lead to an abrupt and simultaneous
cessation of speaking by both speakers, which is a kind of production error only
found under these conditions. The presence of this error suggests that the two
speakers are tightly coupled, and thus non-independent, when performing the
task [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. This is evidence for the existence of an interpersonal synergy [
        <xref ref-type="bibr" rid="ref10 ref12">12, 10</xref>
        ].
Coupling between speakers has also been documented in an fMRI study in which
cortical activity during live synchronisation with an experimenter was
interestingly di erentiated from activity when speaking along with a recording of the
experimenter, even when the subjects were unaware that there was a di erence
in conditions, or that recordings were being used at all [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
      <p>When joint speech is encountered, there are typically very many more than
two speakers involved. Synchronisation is much more lax than found in these
dyadic synchronous speech experiments, and, of course, the motivation and
context are radically di erent. In what follows, a small experiment is reported that
employs the synchronous speech methodology, but extends it to more speakers,
allowing a comparison of the temporal alignment of 2, 4, 6 and even 8
speakers at a time. We pose two substantial research questions. Will the addition of
more speakers lead to systematic changes in the overall temporal patterning of
speech? And will the addition of more speakers lead to less overall synchrony
within the group? This will allow us to better understand the di erence in group
performance between the experimental situation of synchronous speaking and
the collective behaviour of joint speaking in the wild. We know that prosody
in chanting is often substantially altered, but this may be a function of
repetition, which is frequently present, or of associated body movements, such as st
pumping, or of many other potential causes. We do not know if it is a necessary
e ect of group synchronisation, though observation of joint speech done
instrumentally, such as in public swearing of an oath of allegiance, typically exhibits
less prosodic stylisation than well practiced pledges, chants, and prayers. We
also know that synchrony among speakers in ritual and protest is typically not
as tight as that found under laboratory conditions of synchronous speaking.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Methods</title>
      <p>Eight speakers took part (age range: 21{46, 2 Female, 6 Male). All were student
volunteers and native speakers of English. No reward was given for
participation and all experimental procedures were approved by the UCD research ethics
committee. Recording was done in a single class room (approx 3 m x 4 m)
and subjects each wore head mounted microphones to minimise cross-speaker
contamination. Each subject was recorded on a spearate channel. An ambient
recording of all speakers was also captured with a separate microphone.</p>
      <p>Five texts were prepared that spanned genres. They included one regular and
one irregular wordlist, the Hail Mary prayer, a short prose piece with irregular
stressed syllable onsets, and a metrically complex poem. All texts had been used
in previous synchronous speech experiments. There were 40 trials. On each trial,
1, 2, 4, or 6 subjects were randomly selected, or all 8 subjects took part, and the
5 texts were read in randomized order. This provided eight complete recordings
of all texts for each xed number of speakers. Speaking began on a signal from
the experimenter. No tempo instructions or other guidance was given.</p>
      <p>
        There are several ways to estimate synchrony across any given pair of speakers
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Perhaps the simplest is to identify the point at which the intensity rise in a
strong stressed syllable occurs (this is an estimate of a P-centre [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]), and to use
those as landmarks whose time of occurrance can be compared. Stressed syllable
onsets were computed for all recordings and texts using the algorithm reported
in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <p>Initial results will be presented here for the rst time. A rst question to be
asked is whether overall speech timing is a ected by the constraint of speaking
in groups of di erent size. To pursue this issue, we rst examine the intervals
between stressed syllable onsets in the two wordlists. The regular word list
consisted of eight trochees (Borrow, Dancer, Butter, Dagger, Boiler, Doggie, Body,
Deeper ) while the irregular wordlist had a similar strong-weak alternation, but
with an irregular patterning of word edges (Debug, Debasing, Ball, Degree,
Bandana, Beach, Beginning, Duck ).</p>
      <p>
        Fig. 1 shows interval durations for the two wordlists as a function of
position in list (x-axis) and speaker numbers. Each list displays idiosyncratic timing
properties that are highly preserved across changes to speaker number, and,
crucially, there is no evidence here for any systematic change as speaker numbers
vary from 1 to 8. The alternating pattern of relatively shorter and longer
intervals is of a kind with previously presented results from both speech and typing
[
        <xref ref-type="bibr" rid="ref11 ref2">2, 11</xref>
        ].
      </p>
      <p>We repeat the analysis now with a much more rhythmically complex text.
The poem \Kill a cat" displays a binary rhythm in lines 1, 2, and 5, but switches
to a triple metre in lines 3 and 41. The text is:</p>
      <p>Kill a cat, kill a cat
Bash its brains in with a bat.</p>
      <p>Their nine lives expire
When tossed in a re, so
Kill a cat today.
1 Apologies to cat lovers everywhere.
N=1</p>
      <p>N=8
)
c
e
s
i(no .06
tr
a
u
D
)
c
e
s
i(no .06
tr
a
u
D
are computed on a pairwise basis, there are many more pairwise comparisons
when n = 8 than when n = 2 (224 vs. 8).</p>
      <p>From the gure it is clear that asynchrony remains more or less constant
as the number of speakers increases, irrespective of text type. There is some
evidence that the simpler texts (word lists) result in slightly tighter synchrony
than the complex poem or prose passages, but the principal result here is a lack
of e ect of the number of speakers.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Discussion</title>
      <p>
        Joint speech and synchronous speech are not the same thing. The rst is found in
a wide variety of situations in which collective purposes, passions, and aspirations
are made manifest, and in which di erent kinds of identies are brought forth or
enacted, whether that be as Buddhist monks, concerned citizens, or Arsenal
supporters. The latter is found in the laboratory, and although there too people
speak in synchrony with one another, their purposes are starkly di erent, and
the speech produced has its own speci c characteristics. It usually exhibits a very
strong synchrony that is alien to joint speech in the wild. But the circumstances
under which joint speech is produced vary greatly and they frequently include
contextual factors and historical particulars that leave their stamp on the speech.
Thus prayers and ritual incantations are typically repeated many times over.
Protest chants are repeated and fequently augmented with st pumping, clapping
or noise making of various kinds. Football chants too typically involve repetition
and musical elements. (In studying joint speech, a determinate border between
speech and song or music can no longer be sustained [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].)
trochees
irregular
poem
prayer
prose
      </p>
      <p>A signi cant di erence between the context in which joint speech and
synchronous speech occur is the restriction of all previous synchronous speech
experiments to dyads, whereas joint speech typically involves many more speakers.
The present work thus served to see whether this factor, number of speakers, can
account for any of the observed di erences between the two kinds of speech. The
result is in the negative, both when overall temporal patterning is examined,
and when pairwise asynchronies are computed. Although a negative result, this
is a nding of some signi cance, as it allows us now to more clearly distinguish
between the two kinds of speaking, and to recognise that synchronous speaking
is not free of its own contextual imprematur that cannot be attributed simply
to the smaller number of participants.</p>
      <p>One important lesson to be learned from this small experiment is that the
study of joint speech demands a great deal of sensitivity to the context in which
it occurs. In keeping with a general trend in phonetic and linguistic research, it
is necessary to recognise that the laboratory is not a neutral space in which a
pristine object of study, speech, can be found. It is, rather, itself a rich context
imbued with its own characteristics that indelibly mark the speech produced
therein. The enhanced synchrony found in synchronous speech, and the
associated characteristic phenomenon of joint cessation when an error occurs, these
are not a result of a lack of numbers, but a product of an altered situation of
speaking.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Cummins</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>On synchronous speech</article-title>
          .
          <source>Acoustic Research Letters Online</source>
          <volume>3</volume>
          (
          <issue>1</issue>
          ),
          <volume>7</volume>
          {
          <fpage>11</fpage>
          (
          <year>2002</year>
          ). https://doi.org/doi = 10.1121/1.1416672
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Cummins</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Rhythmic organization of read word lists</article-title>
          .
          <source>Journal of the Acoustical Society of America</source>
          <volume>112</volume>
          (
          <issue>5</issue>
          :2),
          <volume>2443</volume>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Cummins</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Practice and performance in speech produced synchronously</article-title>
          .
          <source>Journal of Phonetics</source>
          <volume>31</volume>
          (
          <issue>2</issue>
          ),
          <volume>139</volume>
          {
          <fpage>148</fpage>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Cummins</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Measuring synchronization among speakers reading together</article-title>
          .
          <source>In: Proc. ISCA Workshop on Experimental Linguistics</source>
          . pp.
          <volume>105</volume>
          {
          <fpage>108</fpage>
          . Athens, GR (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Cummins</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Joint speech: The missing link between speech and music</article-title>
          ? Percepta|Revista de Cognic~
          <source>ao Musical</source>
          <volume>1</volume>
          (
          <issue>1</issue>
          ),
          <volume>17</volume>
          {
          <fpage>32</fpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Cummins</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>The remarkable unremarkableness of joint speech</article-title>
          .
          <source>In: Proceedings of the 10th International Seminar on Speech Production</source>
          . pp.
          <volume>73</volume>
          {
          <fpage>77</fpage>
          . Cologne, DE (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Cummins</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>The Ground From Which We Speak: Joint Speech and the Collective Subject</article-title>
          . Cambridge Scholars (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Cummins</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Port</surname>
            ,
            <given-names>R.F.</given-names>
          </string-name>
          :
          <article-title>Rhythmic constraints on stress timing in English</article-title>
          .
          <source>Journal of Phonetics</source>
          <volume>26</volume>
          (
          <issue>2</issue>
          ),
          <volume>145</volume>
          {
          <fpage>171</fpage>
          (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Jasmin</surname>
            ,
            <given-names>K.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McGettigan</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Agnew</surname>
            ,
            <given-names>Z.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lavan</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Josephs</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cummins</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Scott</surname>
            ,
            <given-names>S.K.</given-names>
          </string-name>
          :
          <article-title>Cohesion and joint speech: Right hemisphere contributions to synchronized vocal production</article-title>
          .
          <source>The Journal of Neuroscience</source>
          <volume>36</volume>
          (
          <issue>17</issue>
          ),
          <volume>4669</volume>
          {
          <fpage>4680</fpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Riley</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Richardson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shockley</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ramenzoni</surname>
            ,
            <given-names>V.C.</given-names>
          </string-name>
          :
          <article-title>Interpersonal synergies</article-title>
          .
          <source>Frontiers in Psychology 2</source>
          ,
          <issue>38</issue>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Rosenbaum</surname>
            ,
            <given-names>D.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kenny</surname>
            ,
            <given-names>S.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Derr</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          :
          <article-title>Hierarchical control of rapid movement sequences</article-title>
          .
          <source>Journal of Experimental Psychology: Human Perception and Performance</source>
          <volume>9</volume>
          (
          <issue>1</issue>
          ),
          <volume>86</volume>
          {
          <fpage>102</fpage>
          (
          <year>1983</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Schmidt</surname>
            ,
            <given-names>R.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carello</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Turvey</surname>
          </string-name>
          , M.T.:
          <article-title>Phase transitions and critical uctuations in the visual coordination of rhythmic movements between people</article-title>
          .
          <source>Journal of Experimental Psychology: Human Perception and Performance</source>
          <volume>16</volume>
          (
          <issue>2</issue>
          ),
          <volume>227</volume>
          (
          <year>1990</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Scott</surname>
          </string-name>
          , S.K.:
          <article-title>P-centers in Speech: An Acoustic Analysis</article-title>
          .
          <source>Ph.D. thesis</source>
          , University College London (
          <year>1993</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>