<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Visualization: The Missing Factor in Simultaneous Speech Translation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sara Papi</string-name>
          <email>spapi@fbk.eu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Matteo Negri</string-name>
          <email>negri@fbk.eu</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marco Turchi</string-name>
          <email>turchi@fbk.eu</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>. Fondazione Bruno Kessler</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Italy</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>. University of Trento</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Simultaneous speech translation (SimulST) is the task in which output generation has to be performed on partial, incremental speech input. In recent years, SimulST has become popular due to the spread of multilingual application scenarios, like international live conferences and streaming lectures, in which on-the-fly speech translation can facilitate users' access to audio-visual content. In this paper, we analyze the characteristics of the SimulST systems developed so far, discussing their strengths and weaknesses. We then concentrate on the evaluation framework required to properly assess systems' effectiveness. To this end, we raise the need for a broader performance analysis, also including the user experience standpoint. We argue that SimulST systems, indeed, should be evaluated not only in terms of quality/latency measures, but also via task-oriented metrics accounting, for instance, for the visualization strategy adopted. In light of this, we highlight which are the goals achieved by the community and what is still missing.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>Simultaneous speech translation (SimulST) is the
task in which the translation of a source language
speech has to be performed on partial,
incremental input. This is a key feature to achieve low
latency in scenarios like streaming conferences and
lectures, where the text has to be displayed
following as much as possible the pace of the speech.</p>
      <p>Copyright © 2021 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).</p>
      <p>SimulST is indeed a complex task in which the
difficulties of performing speech recognition from
partial inputs are exacerbated by the problem to
project meaning across languages. Despite the
increasing demand for such a system, the problem is
still far from being solved.</p>
      <p>So far, research efforts mainly focused on the
quality/latency trade-off, i.e. producing high
quality outputs in the shortest possible time, balancing
the need for a good translation with the necessity
of a rapid text generation. Previous studies,
however, disregard how the translation is displayed
and, consequently, how it is actually perceived by
the end users. After a concise survey of the state
of the art in the field, in this paper we posit that,
from the users’ experience standpoint, output
visualization is at least as important as having a good
translation in a short time. This raises the need
for a broader, task-oriented and human-centered
analysis of SimulST systems’ performance, also
accounting for this third crucial factor.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Background</title>
      <p>
        As in the case of offline speech translation, the
adoption of cascade architectures
        <xref ref-type="bibr" rid="ref35 ref37">(Stentiford and
Steer, 1988; Waibel et al., 1991)</xref>
        was the first
attempt made by the SimulST community to tackle
the problem of generating text from partial,
incremental input. Cascade systems
        <xref ref-type="bibr" rid="ref15 ref16 ref20 ref27 ref3 ref4 ref40">(Fu¨gen, 2009;
Fujita et al., 2013; Niehues et al., 2018; Xiong
et al., 2019; Arivazhagan et al., 2020b)</xref>
        involve
a pipeline of two components. First, a
streaming automatic speech recognition (ASR) module
transcribes the input speech into the
corresponding text
        <xref ref-type="bibr" rid="ref21 ref23 ref38">(Wang et al., 2020; Moritz et al., 2020)</xref>
        .
Then, a simultaneous text-to-text translation
module translates the partial transcription into
targetlanguage text
        <xref ref-type="bibr" rid="ref1 ref10 ref18 ref20 ref42 ref43 ref44">(Gu et al., 2017; Dalvi et al., 2018;
Ma et al., 2019; Arivazhagan et al., 2019)</xref>
        . This
approach suffers from error propagation, a
wellknown problem even in the offline scenario, where
the transcription errors made by the ASR module
are propagated to the MT module, which cannot
recover from them as it does not have direct
access to the audio. Another strong limitation of
cascaded systems is the extra latency added by
the two-step pipeline, since the MT module has to
wait until the streaming ASR output is produced.
      </p>
      <p>
        To overcome these issues, the
direct models initially proposed in
B[Pleaseinsertintopreamble]rard et al.
(2016; Weiss et al. (2017) represent a valid
alternative that is gaining increasing traction
        <xref ref-type="bibr" rid="ref5">(Bentivogli et al., 2021)</xref>
        . Direct ST models are
composed of an encoder, usually bidirectional,
and a decoder. The encoder starts from the
audio features extracted from the input signal and
computes a hidden representation; the decoder
transforms this representation into target language
text. Direct modeling becomes crucial in the
simultaneous scenario, as it reduces the overall
system’s latency due to the absence of
intermediate symbolic representation steps. Despite
the data scarcity issue caused by the limited
availability of speech-to-translation corpora, the
adoption of direct architectures showed to be
promising
        <xref ref-type="bibr" rid="ref32 ref39 ref41">(Weiss et al., 2017; Ren et al., 2020;
Zeng et al., 2021)</xref>
        , driving recent efforts towards
the development of increasingly powerful and
efficient models.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Architectural Challenges</title>
      <p>
        This section surveys the direct SimulST models
developed so far, highlighting strengths and
weaknesses of the current architectures and decision
policies – i.e. the strategies used by the system to
decide whether to output a partial translation or to
wait for more audio information. We discuss
ongoing research on architectural improvements of
encoder-decoder models, as well as popular
approaches like offline training and re-translation.
All these works concentrate on reducing systems
latency, targeting a better quality/latency trade-off.
Encoding Strategy. Few studies
        <xref ref-type="bibr" rid="ref12 ref13 ref24 ref25">(Elbayad et al.,
2020a; Nguyen et al., 2021b)</xref>
        tried to improve the
encoder part of simultaneous systems. Elbayad et
al. (2020a) and Nguyen et al. (2021b) introduced
the use of unidirectional encoders instead of
standard bidirectional encoders (i.e. the encoder states
are not updated after each read action) to speed up
the decoding phase. Nguyen et al. (2021b) also
proposed an encoding strategy called
Overlapand-Compensate, where the encoder exploits extra
frames provided from the past that were discarded
during the previous encoding step. The
segmentation problem is a crucial aspect in SimulST, where
the system needs to split a long audio input into
smaller chunks (speech frames) in order to process
them. Different segmentation techniques can be
adopted to extract this information, starting from
the easiest one based on fixed time windows
        <xref ref-type="bibr" rid="ref21 ref22">(Ma
et al., 2020b)</xref>
        to the dynamic ones based on
automatically detected word boundaries
        <xref ref-type="bibr" rid="ref41 ref7">(Zeng et al.,
2021; Chen et al., 2021)</xref>
        . Ma et al. (2020b) also
studied the dynamic segmentation based on oracle
boundaries but they discovered that, in their
scenario, it had worse performance compared to that
of the fixed segmentation.
      </p>
      <p>
        Decoding Strategy. Some efforts have been
made to improve the decoding strategy as it
strongly correlates to the decision policy of
simultaneous systems. Speculative beam search, or
SBS,
        <xref ref-type="bibr" rid="ref42 ref43 ref44">(Zheng et al., 2019c)</xref>
        represents the first
successful attempt to use beam search in SimulST.
This technique consists in hallucinating several
prediction steps in the future in order to make
more accurate decisions based on the best
“speculative” prediction obtained. Also Zeng et al.
(2021) integrate the beam search in the decoding
strategy, developing the wait-k-stride-N strategy.
In particular, the authors bypass output
speculation by directly applying beam search, after
waiting for k words, on a word stride of size N (i.e., on
N words at a time) instead of one single word as
prescribed by the standard wait-k. Nguyen et al.
(2021a) analyzed several decoding strategies
relying on different output token granularities, such as
characters and Byte Pair Encoding (BPE),
showing that the latter yields lower latency.
      </p>
      <p>Offline or Online training? An alternative
approach to simultaneous training is the offline (or
full-sentence) training of the system and its
subsequent use as a simultaneous one. Nguyen et
al. (2021a) explored this solution with an
LSTMbased direct ST system, analyzing the
effectiveness of different decoding strategies. Interestingly,
the offline approach does not only preserve overall
performance despite the switch of modality, it also
improves system’s ability to generate well-formed
sentences. These results are confirmed by Chen
et al. (2021), who successfully exploit a direct ST
system jointly trained in an offline fashion with an
ASR one.</p>
      <p>
        Another point of view: re-translation.
Retranslation
        <xref ref-type="bibr" rid="ref13 ref26 ref27 ref3 ref3 ref4 ref4">(Niehues et al., 2016; Niehues et al.,
2018; Arivazhagan et al., 2020a; Arivazhagan et
al., 2020b)</xref>
        consists in re-generating the output
from scratch (e.g. after a fixed amount of time)
for as long as new information is received. This
approach ensures high quality (the final output is
produced with all the available context) and low
latency (partial translations can be generated with
ifxed, controllable delay). This, however, comes at
the cost of strong output instability (the so-called
lfickering , due to continuous updates of the
displayed translations) which is not optimal from the
user experience standpoint. To this end, some
metrics have been developed to measure the instability
phenomenon, such as the Erasure
        <xref ref-type="bibr" rid="ref3 ref4">(Arivazhagan et
al., 2020b)</xref>
        , which measures the number of tokens
that were deleted from the emitted translation to
produce the next translation.
      </p>
      <p>Decision Policy. In simultaneous settings, the
model has to decide, at each time step, if the
available information is enough to produce a partial
translation – i.e. to perform a write action
using the information received until that step (audio
chunk/s in case of SimulST or token/s in case of
simultaneous MT) – or if it has to wait and perform
a read action to receive new information from the
input. Possible decision policies result in
different ways to balance the quality/latency trade-off.
On one side, more read actions provide the
system with larger context useful to generate
translations of higher quality. On the other side, this
counterbalances the increased, sometimes
unacceptable latency. To address this problem, two
types of policy have been proposed so far: fixed
and adaptive. While fixed decision policies look at
the number of ingested tokens (or speech chunks,
in the speech scenario), in the adaptive ones the
decision is taken by also looking at the contextual
information extracted from the input.</p>
      <p>
        While little research focused on adaptive
policies
        <xref ref-type="bibr" rid="ref18 ref42 ref43 ref44 ref45">(Gu et al., 2017; Zheng et al., 2019a; Zheng
et al., 2020)</xref>
        due to the hard and time-consuming
training
        <xref ref-type="bibr" rid="ref1 ref42 ref43 ref44">(Zheng et al., 2019b; Arivazhagan et al.,
2019)</xref>
        , the adoption of very easy-to-train fixed
policies is the typical choice. Indeed, the most
widely used policy is a fixed one, called wait-k
        <xref ref-type="bibr" rid="ref20 ref42 ref43 ref44">(Ma et al., 2019)</xref>
        . Simple yet effective, it is based
on waiting for k source words before starting to
generate the target sentence, as shown in Table 1.
source It was a
wait-3 - -
wait-5 - -
way
Es
that
ging
parents
um
Es
...
eine
ging
      </p>
      <p>
        As the original wait-k implementation is based
on textual source data, Ma et al. (2020b) adapted
it to the audio domain by waiting for k fixed
time frames (audio chunks or speech frames)
rather than k words. However, this simplistic
approach does not consider various aspects of
human speech, such as different speech rates,
duration, pauses, and silences. In
        <xref ref-type="bibr" rid="ref32">(Ren et al., 2020)</xref>
        ,
the adaptation was done differently, by including
a Connectionist Temporal Classification
(CTC)based
        <xref ref-type="bibr" rid="ref17">(Graves et al., 2006)</xref>
        segmentation
module that is able to determine word boundaries. In
this case, the wait-k strategy is applied by
waiting for k pauses between words that are
automatically detected by the segmenter. Similarly, Zeng et
al. (2021) employed the CTC-based segmentation
method but applying a wait-k-stride-N policy to
allow re-ranking during the decoding phase. The
wait-k-stride-N model emits more than one word
at a time, slightly increasing the latency, since the
output is prompted after the stride is processed.
This small increase in latency, however, allows the
model to perform beam search on the stride, which
has been shown to be effective in improving
translation quality
        <xref ref-type="bibr" rid="ref36">(Sutskever et al., 2014)</xref>
        . Decoding
more than one word at a time is the approach also
employed by Nguyen et al. (2021a), who showed
that emitting two words increases the quality of the
translation without any relevant impact on latency.
Another way of applying the wait-k strategy was
proposed by Chen et al. (2021), where a streaming
ASR system is used to guide the direct ST
decoding. They look at the ASR beam to decide how
many tokens have been emitted within the partial
audio segment, hence having the information to
apply the original wait-k policy in a
straightforward way. An interesting solution is also the one
by Elbayad et al. (2020a), who jointly train a
direct model across multiple wait-k paths. Once the
sentence has been encoded, they optimize the
system by uniformly sampling the k value for the
decoding step. Even though they reach good
performance by using a single-path training with k=7
and a different k value for testing, the multi-path
approach proved to be effective. One of its
advantages is that no k value has to be specified for the
training, which allows to avoid the training from
scratch of several models for different values of k.
Retrospective. All the aspects analyzed in this
section highlight several research directions
already taken by the simultaneous community,
which have to be studied more in depth. Among
all, the audio or text segmentation strategy clearly
emerges as a fundamental factor of simultaneous
systems, and the ambivalent results obtained in
several studies point out that this aspect has to be
better clarified. Moreover, the presence of
extensive literature on the wait-k policy shows that it
represents one of the topics of greatest interest to
the community, which continues to work on it to
further improve its effectiveness as it directly
impacts on the systems’ performance, especially
latency. Unfortunately, all these studies focus on the
architecture enhancements and decision policies
despite the absence of a unique and clear
evaluation framework to perform a correct and complete
analysis of the system.
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>Evaluation Challenges</title>
      <p>
        A good simultaneous model should produce a high
quality translation with reasonable timing, as
waiting too long will negatively affect the user
experience. Offline MT and ST communities commonly
use the well-established BLEU metric
        <xref ref-type="bibr" rid="ref28 ref30">(Papineni
et al., 2002; Post, 2018)</xref>
        to measure the quality of
the output translation, but a simultaneous system
also needs a metric that accounts for the time spent
by the system to output the partial translation.
Simultaneous MT (SimulMT) is the task in which
a real-time translation is produced having a
partial source text at disposal. Since SimulMT was
the first yet easiest simultaneous scenario studied
by the community, a set of metrics was previously
introduced for the textual input-output translation
part.
      </p>
    </sec>
    <sec id="sec-5">
      <title>Latency Metrics for SimulMT. The first met</title>
      <p>
        ric, the Average Proportion (AP), was proposed by
Cho and Esipova (2016) and measures the average
proportion of source input read when generating a
target prediction, that is the sum of the tokens read
when generating the partial target. However, AP
is not length-invariant, i.e. the value of the metric
depends on the input and output lengths and is not
evenly distributed on the [0, 1] interval
        <xref ref-type="bibr" rid="ref20 ref42 ref43 ref44">(Ma et al.,
2019)</xref>
        , making this metric strongly unreliable.
      </p>
      <p>To overcome all these problems, Ma et al.
(2019) introduced the Average Lagging (AL) that
directly describes the lagging behind the ideal
policy, i.e. a policy that produces the output
exactly at the same time as the speech source. As a
downside, Average Lagging is not differentiable,
which is, instead, a useful property, especially if
the metric is likely to be added in the system’s loss
computation. For this reason, Cherry and Foster
(2019) proposed the Differential Average Lagging
(DAL), introducing a minimum delay after each
operation.</p>
      <p>
        Another way of measuring the lagging is to
compute the alignment difficulty of a source-target
pair. Hence, Elbayad et al. (2020b) proposed
the Lagging Difficulty (LD) metric that exploits
the use of the fast-align
        <xref ref-type="bibr" rid="ref11">(Dyer et al., 2013)</xref>
        tool to estimate the source and target alignments.
Then, they infer the reference decoding path and
compute the AL metric. The authors claimed the
LD to be a realistic measure of the simultaneous
translation as it also evaluates how a translation
is easy to align considering the context available
when decoding.
      </p>
      <p>
        Latency Metrics for SimulST. The most
popular AP, AL and DAL metrics were successively
adapted by the SimulST community to the speech
scenario by converting, for instance, the number of
words to the sum of the speech segment durations,
as per
        <xref ref-type="bibr" rid="ref13 ref21 ref22">(Ma et al., 2020a)</xref>
        . Later, Ma et al. (2020b)
raised the issue of using computational unaware
metrics and proposed computational aware
metrics accounting for the time spent by the model
to generate the output. Unfortunately,
computing such metrics is not easy at all in absence of
a unique and reproducible environment that can
be used to evaluate the model’s performance. To
this end, Ma et al. (2020a) proposed SimulEvala
tool which computes the metrics by simulating
a real-time scenario with a server-client scheme.
This toolkit automatically evaluates simultaneous
translations (both text and speech) given a
customizable agent that can be defined by the user and
that will depend on the adopted policy. Despite the
progress in the metrics for evaluating quality and
latency, no studies have been conducted on the
effective correlation with user experience. This
represents a missing key point in the current
evaluation framework landscape, giving rise to the need
for a tool that combines quality and latency
metrics with application-oriented metrics (e.g.,
reading speed), which are strongly correlated to the
visualization and, as an ultimate goal, to the user
experience.
5
      </p>
    </sec>
    <sec id="sec-6">
      <title>The Missing Factor: Visualization</title>
      <p>
        In the previous section, we introduced the most
popular metrics used to evaluate the
simultaneous systems’ performance. These metrics account
for the quality and the latency of the system
without capturing the user needs. Although many
researchers acknowledge the importance of human
evaluation, this current partial view can push the
community in the wrong direction, in which all the
efforts are focused on the quality/latency factors
while the problem experienced by the user is of
another kind. Indeed, the third factor that matters
and strongly influences the human understanding
of a – even very good – translation is the
visualization strategy adopted. The visualization
problem and the need to present the text in a readable
fashion for the user was only faced in our
previous work
        <xref ref-type="bibr" rid="ref19">(Karakanta et al., 2021)</xref>
        . In the paper,
we raised the need for a clearer and less
distracting visualization of the SimulST system’s
generated texts by presenting them as subtitles (text
segmented in lines preserving coherent information).
We proposed different visualization strategies to
better assess the online display problem,
attempting to simulate a setting where human
understanding is at the core of our analysis.
      </p>
    </sec>
    <sec id="sec-7">
      <title>Visualization modalities. The standard word</title>
      <p>
        for-word visualization method
        <xref ref-type="bibr" rid="ref20 ref42 ref43 ref44">(Ma et al., 2019)</xref>
        , in
which the words appear sequentially on the screen
as they are generated, could be strongly
suboptimal for the human understanding
(RomeroFresco, 2011). Infact, the word-for-word approach
has two main problems: i) the emission rate of
words (some go too fast, some too slow) is
irregular and the users waste more time reading the
text because their eyes have to make more
movements, and ii) emission of pieces of text that do
not correspond to linguistic units/chunks,
requiring more cognitive effort. Moreover, when the
maximum length of the subtitle (that depends on
the dimensions of the screen) is reached, the
subtitle disappears without giving the user enough time
to read the last words emitted. As this will
negatively impact the user experience, we propose in
        <xref ref-type="bibr" rid="ref19">(Karakanta et al., 2021)</xref>
        to adopt different
visualization modes that better accommodate the
human reading requirements. We first introduced
the block visualization mode, for which an
entire subtitle is displayed at once (usually one or
two lines maximum) as soon as the system has
ifnished generating it. This display mode is the
easiest to read for the user because it prevents
re-reading phenomena
        <xref ref-type="bibr" rid="ref31">(Rajendran et al., 2013)</xref>
        and unnecessary/excessive eye fixations
(RomeroFresco, 2010), reducing the human effort.
However, we discovered that the latency introduced by
waiting for an entire subtitle is too high to let this
visualization mode be used in many simultaneous
scenarios. As a consequence, we proposed the
scrolling lines visualization mode that displays the
subtitles line by line. Every time a new line
becomes available, it appears at the bottom of the
screen, while the previous (older) line is scrolled
to the upper line. In this way, there are always
two lines displayed on the screen. To evaluate the
performance of the system in the different
visualization modes, we also proposed an ad-hoc
calculation of the reading speed (characters per second
or CPS) that correlates with the human judgment
of the subtitles
        <xref ref-type="bibr" rid="ref29">(Perego et al., 2010)</xref>
        . The reading
speed shows how fast a user needs to read in
order not to miss any part of the subtitle. The lower
the reading speed, the better is the model’s
output since a fast reading speed increases the
cognitive load and leaves less time to look at the image.
The scrolling line method offers the best balance
between latency and a comfortable reading speed
resulting to be the best choice for the
simultaneous scenario. On the other hand, this approach
requires segmented text (i.e. a text that is divided
into subtitles), thus the system needs to be able to
simultaneously generate transcripts or translations
together with proper subtitle delimiters. However,
building a simultaneous subtitling system
combines the difficulties of the simultaneous setting
with the constraint of having a text formatted in
proper subtitles. Since both these research
directions are still evolving, a lot of work is required to
achieve good results.
      </p>
      <p>
        The lack of studies on this aspects highlights the
shortcomings of the actual SimulST systems,
individuating possible improvements that will allow
the systems to evolve in a more organic and
complete way according to the user needs. Moreover,
to completely assess the subtitling scenario, a
system has to be able to jointly produce timestamps
metadata linked to the word emitted, a task that has
not been addressed so far. The need for this kind
of system represents an interesting direction to
follow for the simultaneous community. In the light
of this, the researcher should also take into account
the three quality-latency-visualization factors in
their analyses. We are convinced that these are
the most promising aspects to work on to build the
best SimulST system for the audience and that
human evaluation has to have a crucial role in future
studies. We also believe that interdisciplinary
dialogue with other fields such as cognitive studies,
media accessibility and human-computer
interaction would be very insightful to evaluate SimulST
outputs from communicative perspectives
        <xref ref-type="bibr" rid="ref14 ref19 ref24 ref25 ref41 ref7">(Fantinuoli and Prandi, 2021)</xref>
        .
6
      </p>
    </sec>
    <sec id="sec-8">
      <title>Conclusions and Future Directions</title>
      <p>SimulST systems have become increasingly
popular in recent years and many efforts have been
made to build robust and efficient models.
Despite the difficulties introduced by the online
framework, these models have rapidly improved,
achieving comparable results to the offline
systems. However, many research directions have not
been explored enough (e.g., the adoption of
dynamic or fixed segmentation, the offline or the
online training). First among all, the visualization
strategy that is adopted to display the output of the
simultaneous systems is an important and largely
under-analyzed aspect of the simultaneous
experience. We posit that the presence of
applicationoriented metrics (e.g., reading speed), which are
strongly related to the visualization and, as an
ultimate goal, to the user experience, is the factor
that misses in the actual evaluation environment.
Indeed, this paper points out that BLEU and
Average Lagging are not the only metrics that matter to
effectively evaluate a SimulST model, even if they
are fundamental to judge a correct and real-timed
translation. We hope that this will inspire the
community to work on this critical aspect in the future.</p>
    </sec>
    <sec id="sec-9">
      <title>Acknowledgement</title>
      <p>This work has been carried out as part of the
project Smarter Interpreting (https://kunven
o.digital/) financed by CDTI Neotec funds.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Naveen</given-names>
            <surname>Arivazhagan</surname>
          </string-name>
          , Colin Cherry, Wolfgang Macherey, Chung-Cheng Chiu, Semih Yavuz, Ruoming Pang,
          <string-name>
            <given-names>Wei</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Colin</given-names>
            <surname>Raffel</surname>
          </string-name>
          .
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <article-title>Monotonic infinite lookback attention for simultaneous machine translation</article-title>
          .
          <source>In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics</source>
          , pages
          <fpage>1313</fpage>
          -
          <lpage>1323</lpage>
          , Florence, Italy, July. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Naveen</given-names>
            <surname>Arivazhagan</surname>
          </string-name>
          , Colin Cherry, Wolfgang Macherey, and
          <string-name>
            <given-names>George</given-names>
            <surname>Foster</surname>
          </string-name>
          . 2020a.
          <article-title>Retranslation versus streaming for simultaneous translation</article-title>
          .
          <source>In Proceedings of the 17th International Conference on Spoken Language Translation</source>
          , pages
          <fpage>220</fpage>
          -
          <lpage>227</lpage>
          , Online, July. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Naveen</given-names>
            <surname>Arivazhagan</surname>
          </string-name>
          , Colin Cherry, Isabelle Te, Wolfgang Macherey, Pallavi Baljekar, and
          <string-name>
            <given-names>George</given-names>
            <surname>Foster</surname>
          </string-name>
          . 2020b.
          <article-title>Re-translation strategies for long form, simultaneous, spoken language translation</article-title>
          .
          <source>In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source>
          , pages
          <fpage>7919</fpage>
          -
          <lpage>7923</lpage>
          . IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Luisa</given-names>
            <surname>Bentivogli</surname>
          </string-name>
          , Mauro Cettolo, Marco Gaido, Alina Karakanta, Alberto Martinelli, Matteo Negri, and
          <string-name>
            <given-names>Marco</given-names>
            <surname>Turchi</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Cascade versus direct speech translation: Do the differences still make a difference? In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th</article-title>
          <source>International Joint Conference on Natural Language Processing (Volume 1: Long Papers)</source>
          , pages
          <fpage>2873</fpage>
          -
          <lpage>2887</lpage>
          , Online, August. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          Alexandre Be´rard, Olivier Pietquin, Christophe Servan, and
          <string-name>
            <given-names>Laurent</given-names>
            <surname>Besacier</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Listen and Translate: A Proof of Concept for End-to-End Speech-to-Text Translation</article-title>
          . In NIPS Workshop on end
          <article-title>-to-end learning for speech and audio processing</article-title>
          , Barcelona, Spain, December.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Junkun</given-names>
            <surname>Chen</surname>
          </string-name>
          , Mingbo Ma, Renjie Zheng, and
          <string-name>
            <given-names>Liang</given-names>
            <surname>Huang</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Direct simultaneous speech-to-text translation assisted by synchronized streaming ASR. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021</article-title>
          , pages
          <fpage>4618</fpage>
          -
          <lpage>4624</lpage>
          , Online, August. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Colin</given-names>
            <surname>Cherry</surname>
          </string-name>
          and
          <string-name>
            <given-names>George</given-names>
            <surname>Foster</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Thinking slow about latency evaluation for simultaneous machine translation</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Kyunghyun</given-names>
            <surname>Cho</surname>
          </string-name>
          and
          <string-name>
            <given-names>Masha</given-names>
            <surname>Esipova</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Can neural machine translation do simultaneous translation?</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Fahim</given-names>
            <surname>Dalvi</surname>
          </string-name>
          , Nadir Durrani, Hassan Sajjad, and
          <string-name>
            <given-names>Stephan</given-names>
            <surname>Vogel</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Incremental decoding and training methods for simultaneous translation in neural machine translation</article-title>
          .
          <source>In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>2</volume>
          (
          <issue>Short Papers)</issue>
          , pages
          <fpage>493</fpage>
          -
          <lpage>499</lpage>
          , New Orleans, Louisiana, June. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Chris</given-names>
            <surname>Dyer</surname>
          </string-name>
          , Victor Chahuneau, and
          <string-name>
            <surname>Noah</surname>
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Smith</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>A simple, fast, and effective reparameterization of IBM model 2</article-title>
          .
          <source>In Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , pages
          <fpage>644</fpage>
          -
          <lpage>648</lpage>
          , Atlanta, Georgia, June. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Maha</given-names>
            <surname>Elbayad</surname>
          </string-name>
          , Laurent Besacier, and
          <string-name>
            <given-names>Jakob</given-names>
            <surname>Verbeek</surname>
          </string-name>
          . 2020a.
          <article-title>Efficient Wait-k Models for Simultaneous Machine Translation</article-title>
          .
          <source>In Proc. Interspeech</source>
          <year>2020</year>
          , pages
          <fpage>1461</fpage>
          -
          <lpage>1465</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Maha</given-names>
            <surname>Elbayad</surname>
          </string-name>
          , Michael Ustaszewski, Emmanuelle Esperanc¸
          <article-title>a-Rodier, Francis Brunet-Manquat, Jakob Verbeek, and Laurent Besacier</article-title>
          . 2020b.
          <article-title>Online versus offline NMT quality: An in-depth analysis on English-German and German-English</article-title>
          .
          <source>In Proceedings of the 28th International Conference on Computational Linguistics</source>
          , pages
          <fpage>5047</fpage>
          -
          <lpage>5058</lpage>
          , Barcelona, Spain (Online), December.
          <source>International Committee on Computational Linguistics.</source>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Claudio</given-names>
            <surname>Fantinuoli</surname>
          </string-name>
          and
          <string-name>
            <given-names>Bianca</given-names>
            <surname>Prandi</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Towards the evaluation of automatic simultaneous speech translation from a communicative perspective</article-title>
          .
          <source>In Proceedings of the 18th International Conference on Spoken Language Translation (IWSLT</source>
          <year>2021</year>
          ), pages
          <fpage>245</fpage>
          -
          <lpage>254</lpage>
          , Bangkok, Thailand (online),
          <source>August. Association for Computational Linguistics.</source>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>C.</given-names>
            <surname>Fu</surname>
          </string-name>
          ¨gen.
          <year>2009</year>
          .
          <article-title>A system for simultaneous translation of lectures and speeches</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Tomoki</given-names>
            <surname>Fujita</surname>
          </string-name>
          , Graham Neubig,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sakti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Toda</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Satoshi</given-names>
            <surname>Nakamura</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Simple, lexicalized choice of translation timing for simultaneous speech translation</article-title>
          .
          <source>In INTERSPEECH.</source>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Alex</given-names>
            <surname>Graves</surname>
          </string-name>
          , Santiago Ferna´ndez,
          <source>Faustino Gomez, and Ju¨rgen Schmidhuber</source>
          .
          <year>2006</year>
          .
          <article-title>Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks</article-title>
          .
          <source>In Proceedings of the 23rd International Conference on Machine Learning, ICML '06, page 369-376</source>
          , New York, NY, USA. Association for Computing Machinery.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Jiatao</given-names>
            <surname>Gu</surname>
          </string-name>
          , Graham Neubig, Kyunghyun Cho, and
          <string-name>
            <surname>Victor</surname>
            <given-names>O.K.</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Learning to translate in real-time with neural machine translation</article-title>
          .
          <source>In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume</source>
          <volume>1</volume>
          ,
          <string-name>
            <surname>Long</surname>
            <given-names>Papers</given-names>
          </string-name>
          , pages
          <fpage>1053</fpage>
          -
          <lpage>1062</lpage>
          , Valencia, Spain, April. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Alina</given-names>
            <surname>Karakanta</surname>
          </string-name>
          , Sara Papi, Matteo Negri, and
          <string-name>
            <given-names>Marco</given-names>
            <surname>Turchi</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Simultaneous speech translation for live subtitling: from delay to display</article-title>
          .
          <source>In Proceedings of the 1st Workshop on Automatic Spoken Language Translation in Real-World Settings (ASLTRW)</source>
          , pages
          <fpage>35</fpage>
          -
          <lpage>48</lpage>
          , Virtual,
          <string-name>
            <surname>August.</surname>
          </string-name>
          <article-title>Association for Machine Translation in the Americas</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Mingbo</given-names>
            <surname>Ma</surname>
          </string-name>
          , Liang Huang, Hao Xiong, Renjie Zheng, Kaibo Liu, Baigong Zheng, Chuanqiang Zhang, Zhongjun He, Hairong Liu,
          <string-name>
            <given-names>Xing</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <surname>Hua Wu</surname>
            , and
            <given-names>Haifeng</given-names>
          </string-name>
          <string-name>
            <surname>Wang</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>STACL: Simultaneous translation with implicit anticipation and controllable latency using prefix-to-prefix framework</article-title>
          .
          <source>In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics</source>
          , pages
          <fpage>3025</fpage>
          -
          <lpage>3036</lpage>
          , Florence, Italy, July. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Xutai</given-names>
            <surname>Ma</surname>
          </string-name>
          , Mohammad Javad Dousti, Changhan Wang,
          <string-name>
            <surname>Jiatao Gu</surname>
            , and
            <given-names>Juan</given-names>
          </string-name>
          <string-name>
            <surname>Pino</surname>
          </string-name>
          .
          <year>2020a</year>
          .
          <article-title>SIMULEVAL: An evaluation toolkit for simultaneous translation</article-title>
          .
          <source>In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations</source>
          , pages
          <fpage>144</fpage>
          -
          <lpage>150</lpage>
          , Online, October. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>Xutai</given-names>
            <surname>Ma</surname>
          </string-name>
          , Juan Pino, and
          <string-name>
            <given-names>Philipp</given-names>
            <surname>Koehn</surname>
          </string-name>
          .
          <year>2020b</year>
          .
          <article-title>SimulMT to SimulST: Adapting simultaneous text translation to end-to-end simultaneous speech translation</article-title>
          .
          <source>In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing</source>
          , pages
          <fpage>582</fpage>
          -
          <lpage>587</lpage>
          , Suzhou, China, December. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <given-names>Niko</given-names>
            <surname>Moritz</surname>
          </string-name>
          , Takaaki Hori, and
          <string-name>
            <given-names>Jonathan</given-names>
            <surname>Le</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Streaming automatic speech recognition with the transformer model</article-title>
          .
          <source>In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source>
          , pages
          <fpage>6074</fpage>
          -
          <lpage>6078</lpage>
          . IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>Ha</given-names>
            <surname>Nguyen</surname>
          </string-name>
          , Yannick Este`ve, and Laurent Besacier.
          <year>2021a</year>
          .
          <article-title>An empirical study of end-to-end simultaneous speech translation decoding strategies</article-title>
          .
          <source>In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source>
          , pages
          <fpage>7528</fpage>
          -
          <lpage>7532</lpage>
          . IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <given-names>Ha</given-names>
            <surname>Nguyen</surname>
          </string-name>
          , Yannick Este`ve, and Laurent Besacier. 2021b.
          <article-title>Impact of Encoding and Segmentation Strategies on End-to-End Simultaneous Speech Translation</article-title>
          .
          <source>In Proc. Interspeech</source>
          <year>2021</year>
          , pages
          <fpage>2371</fpage>
          -
          <lpage>2375</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <given-names>J.</given-names>
            <surname>Niehues</surname>
          </string-name>
          , T. Nguyen, Eunah Cho,
          <string-name>
            <surname>Thanh-Le</surname>
            <given-names>Ha</given-names>
          </string-name>
          , Kevin Kilgour, Markus Mu¨ller, Matthias Sperber,
          <string-name>
            <surname>S.</surname>
          </string-name>
          <article-title>Stu¨ker, and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Waibel</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Dynamic transcription for low-latency speech translation</article-title>
          .
          <source>In INTERSPEECH.</source>
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <given-names>J.</given-names>
            <surname>Niehues</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ngoc-Quan</surname>
            <given-names>Pham</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thanh-Le</surname>
            <given-names>Ha</given-names>
          </string-name>
          ,
          <article-title>Matthias Sperber, and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Waibel</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Low-latency neural speech translation</article-title>
          .
          <source>In INTERSPEECH.</source>
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <string-name>
            <given-names>Kishore</given-names>
            <surname>Papineni</surname>
          </string-name>
          , Salim Roukos, Todd Ward, and
          <string-name>
            <given-names>WeiJing</given-names>
            <surname>Zhu</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>Bleu: a method for automatic evaluation of machine translation</article-title>
          .
          <source>In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics</source>
          , pages
          <fpage>311</fpage>
          -
          <lpage>318</lpage>
          , Philadelphia, Pennsylvania, USA, July. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <given-names>Elisa</given-names>
            <surname>Perego</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Del Missier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Porta</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Mosconi</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>The cognitive effectiveness of subtitle processing</article-title>
          .
          <source>Media Psychology</source>
          ,
          <volume>13</volume>
          :
          <fpage>243</fpage>
          - -
          <lpage>272</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <given-names>Matt</given-names>
            <surname>Post</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>A call for clarity in reporting BLEU scores</article-title>
          .
          <source>In Proceedings of the Third Conference on Machine Translation: Research Papers</source>
          , pages
          <fpage>186</fpage>
          -
          <lpage>191</lpage>
          , Brussels, Belgium, October. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          <string-name>
            <surname>Dhevi J. Rajendran</surname>
          </string-name>
          , Andrew T. Duchowski, Pilar Orero, Juan Mart´ınez, and
          <string-name>
            <surname>Pablo</surname>
          </string-name>
          Romero-Fresco.
          <year>2013</year>
          .
          <article-title>Effects of text chunking on subtitling: A quantitative and qualitative examination</article-title>
          .
          <source>Perspectives</source>
          ,
          <volume>21</volume>
          (
          <issue>1</issue>
          ):
          <fpage>5</fpage>
          -
          <lpage>21</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          <string-name>
            <given-names>Yi</given-names>
            <surname>Ren</surname>
          </string-name>
          , Jinglin Liu, Xu Tan, Chen Zhang, Tao Qin,
          <string-name>
            <given-names>Zhou</given-names>
            <surname>Zhao</surname>
          </string-name>
          , and
          <string-name>
            <surname>Tie-Yan Liu</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>SimulSpeech: End-to-end simultaneous speech to text translation</article-title>
          .
          <source>In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics</source>
          , pages
          <fpage>3787</fpage>
          -
          <lpage>3796</lpage>
          , Online, July. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          <string-name>
            <given-names>Pablo</given-names>
            <surname>Romero-Fresco</surname>
          </string-name>
          ,
          <year>2010</year>
          .
          <article-title>Standing on quicksand: hearing viewers' comprehension and reading patterns of respoken subtitles for the news</article-title>
          , pages
          <fpage>175</fpage>
          -
          <lpage>194</lpage>
          . Brill, Leiden, The Netherlands.
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          <string-name>
            <given-names>Pablo</given-names>
            <surname>Romero-Fresco</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Subtitling through speech recognition: Respeaking</article-title>
          . Manchester: St. Jerome.
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          <string-name>
            <surname>Frederick W. M. Stentiford</surname>
            and
            <given-names>Martin G.</given-names>
          </string-name>
          <string-name>
            <surname>Steer</surname>
          </string-name>
          .
          <year>1988</year>
          .
          <source>Machine Translation of Speech. British Telecom Technology Journal</source>
          ,
          <volume>6</volume>
          (
          <issue>2</issue>
          ):
          <fpage>116</fpage>
          -
          <lpage>122</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          <string-name>
            <given-names>Ilya</given-names>
            <surname>Sutskever</surname>
          </string-name>
          , Oriol Vinyals, and
          <string-name>
            <surname>Quoc</surname>
            <given-names>V</given-names>
          </string-name>
          <string-name>
            <surname>Le</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Sequence to sequence learning with neural networks</article-title>
          . In Z. Ghahramani,
          <string-name>
            <given-names>M.</given-names>
            <surname>Welling</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Cortes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Lawrence</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K. Q.</given-names>
            <surname>Weinberger</surname>
          </string-name>
          , editors,
          <source>Advances in Neural Information Processing Systems</source>
          , volume
          <volume>27</volume>
          . Curran Associates, Inc.
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          <string-name>
            <given-names>Alex</given-names>
            <surname>Waibel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Ajay N.</given-names>
            <surname>Jain</surname>
          </string-name>
          , Arthur E.
          <string-name>
            <surname>McNair</surname>
          </string-name>
          ,
          <string-name>
            <surname>Hiroaki Saito</surname>
            , Alexander G. Hauptmann, and
            <given-names>Joe</given-names>
          </string-name>
          <string-name>
            <surname>Tebelskis</surname>
          </string-name>
          .
          <year>1991</year>
          .
          <article-title>JANUS: A Speech-to-Speech Translation System Using Connectionist and Symbolic Processing Strategies</article-title>
          .
          <source>In Proceedings of the International Conference on Acoustics, Speech and Signal Processing, ICASSP 1991</source>
          , pages
          <fpage>793</fpage>
          -
          <lpage>796</lpage>
          , Toronto, Canada, May
          <volume>14</volume>
          -17.
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          <string-name>
            <surname>Chengyi</surname>
            <given-names>Wang</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu Wu</surname>
            , Shujie Liu,
            <given-names>Jinyu</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>Liang</given-names>
          </string-name>
          <string-name>
            <surname>Lu</surname>
            , Guoli Ye, and
            <given-names>Ming</given-names>
          </string-name>
          <string-name>
            <surname>Zhou</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Low latency endto-end streaming speech recognition with a scout network</article-title>
          . arXiv preprint arXiv:
          <year>2003</year>
          .10369.
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          <string-name>
            <given-names>Ron J.</given-names>
            <surname>Weiss</surname>
          </string-name>
          , Jan Chorowski, Navdeep Jaitly,
          <string-name>
            <surname>Yonghui Wu</surname>
            , and
            <given-names>Zhifeng</given-names>
          </string-name>
          <string-name>
            <surname>Chen</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Sequence-toSequence Models Can Directly Translate Foreign Speech</article-title>
          .
          <source>In Proceedings of Interspeech</source>
          <year>2017</year>
          , pages
          <fpage>2625</fpage>
          -
          <lpage>2629</lpage>
          , Stockholm, Sweden,
          <year>August</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          <string-name>
            <given-names>Hao</given-names>
            <surname>Xiong</surname>
          </string-name>
          , Ruiqing Zhang, Chuanqiang Zhang, Zhongjun He,
          <string-name>
            <surname>Hua Wu</surname>
            , and
            <given-names>Haifeng</given-names>
          </string-name>
          <string-name>
            <surname>Wang</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Dutongchuan: Context-aware translation model for simultaneous interpreting</article-title>
          . arXiv preprint arXiv:
          <year>1907</year>
          .12984.
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          <string-name>
            <given-names>Xingshan</given-names>
            <surname>Zeng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Liangyou</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Qun</given-names>
            <surname>Liu</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>RealTranS: End-to-end simultaneous speech translation with convolutional weighted-shrinking transformer</article-title>
          .
          <source>In Findings of the Association for Computational Linguistics: ACL-IJCNLP</source>
          <year>2021</year>
          , pages
          <fpage>2461</fpage>
          -
          <lpage>2474</lpage>
          , Online, August. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          <string-name>
            <given-names>Baigong</given-names>
            <surname>Zheng</surname>
          </string-name>
          , Renjie Zheng, Mingbo Ma, and
          <string-name>
            <given-names>Liang</given-names>
            <surname>Huang</surname>
          </string-name>
          . 2019a.
          <article-title>Simpler and faster learning of adaptive policies for simultaneous translation</article-title>
          .
          <source>In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)</source>
          , pages
          <fpage>1349</fpage>
          -
          <lpage>1354</lpage>
          ,
          <string-name>
            <surname>Hong</surname>
            <given-names>Kong</given-names>
          </string-name>
          , China, November. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          <string-name>
            <given-names>Baigong</given-names>
            <surname>Zheng</surname>
          </string-name>
          , Renjie Zheng, Mingbo Ma, and
          <string-name>
            <given-names>Liang</given-names>
            <surname>Huang</surname>
          </string-name>
          . 2019b.
          <article-title>Simultaneous translation with flexible policy via restricted imitation learning</article-title>
          .
          <source>In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics</source>
          , pages
          <fpage>5816</fpage>
          -
          <lpage>5822</lpage>
          , Florence, Italy, July. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          <string-name>
            <given-names>Renjie</given-names>
            <surname>Zheng</surname>
          </string-name>
          , Mingbo Ma, Baigong Zheng, and
          <string-name>
            <given-names>Liang</given-names>
            <surname>Huang</surname>
          </string-name>
          . 2019c.
          <article-title>Speculative beam search for simultaneous translation</article-title>
          .
          <source>In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)</source>
          , pages
          <fpage>1395</fpage>
          -
          <lpage>1402</lpage>
          ,
          <string-name>
            <surname>Hong</surname>
            <given-names>Kong</given-names>
          </string-name>
          , China, November. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref45">
        <mixed-citation>
          <string-name>
            <given-names>Baigong</given-names>
            <surname>Zheng</surname>
          </string-name>
          , Kaibo Liu, Renjie Zheng, M. Ma, Hairong Liu, and
          <string-name>
            <given-names>L.</given-names>
            <surname>Huang</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Simultaneous translation policies: From fixed to adaptive</article-title>
          . ArXiv, abs/
          <year>2004</year>
          .13169.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>