<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>STREAMS: Live Streaming and Micro-batching with Automated Multilingual Services</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Aitor Álvarez</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thierry Etchegoyhen</string-name>
          <email>tetchegoyhen@vicomtech.org</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Joaquín Arellano</string-name>
          <email>jarellano@vicomtech.org</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Víctor Ruíz</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Haritz Arzelus</string-name>
          <email>harzelus@vicomtech.org</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>David Ponce</string-name>
          <email>adponce@vicomtech.org</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ander González-Docasal</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Harritxu Gete Ugarte</string-name>
          <email>hgete@vicomtech.org</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Aragon Institute for Engineering Research, University of Zaragoza</institution>
          ,
          <addr-line>50009 Zaragoza</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Fundación Vicomtech, Basque Research and Technology Alliance (BRTA)</institution>
          ,
          <addr-line>Donostia-San Sebastián, 20009</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of the Basque Country UPV/EHU</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>We present STREAMS, a batch and streaming platform which integrates and manages AI services for rich transcription, translation, voice synthesis and audio description. Its main goal is to provide a unified platform to enhance the processes and products of companies working in multiple sectors. STREAMS has been successfully deployed in real-life batch and multilingual streaming scenarios, ofering a rich array of multilingual services with innovative methods, in particular to enhance portability, responsiveness and output readability.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Streaming</kwd>
        <kwd>Micro-batching</kwd>
        <kwd>Rich Transcription</kwd>
        <kwd>Machine Translation</kwd>
        <kwd>Speech Synthesis</kwd>
        <kwd>Subtitling</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>1https://haiko.es/
2https://jarkatza.com/
3https://www.mondragonlingua.com/en
4https://goiena.eus/
5https://mixer.eus/
6https://www.noticiasdegipuzkoa.eus/
7https://www.ondavasca.com/
8https://www.vicomtech.org/en</p>
    </sec>
    <sec id="sec-2">
      <title>2. STREAMS platform</title>
      <sec id="sec-2-1">
        <title>The core architecture of STREAMS is shown in Figure 1.</title>
        <p>The solution can operate in both batch and streaming
modes, via independent deployments.</p>
        <p>STREAMS consists of the following main components:
client requests management, pipelines definition
and execution, scalability and error handling. The
batch and streaming services are deployed
independently, although they are both served within
the same solution.
• Two web-based graphical user interface (GUI)
frontends, one for transcription and translation
purposes, and the other for speech synthesis. 2.1. Frontends
Both GUIs operate with the same backend in The STREAMS GUIs aim to facilitate the communication
batch mode. between the user and the backend within a batch scenario.
• A main REST API, which exposes the functional- They were designed from a usability and user experience
ities of the backend to the GUIs and/or the user perspective, prioritising simplicity.
for direct integration. The principal internal com- The first GUI, shown in Figure 2, provides users with
munications between the diferent modules are diferent input options, such as text and audio file, and
alperformed through REST APIs as well. lows them to select transcription and translation models
• A backend, which orchestrates all the function- to perform the desired tasks. Additionally, it integrates
alities of the solution, including the AI services, two main text-boxes to present the transcription and</p>
      </sec>
      <sec id="sec-2-2">
        <title>We describe each of these components in turn below.</title>
        <p>translation results. It is worth noting that the transcrip- The batch operational mode operates on previously
tion and translation results can be downloaded in difer- generated contents in video, audio and texts formats,
genent formats (txt, json and srt), to be used for diferent erating transcriptions and translation in diferent output
applications. formats. In addition to providing results in text plain</p>
        <p>The second GUI for speech synthesis, shown in Fig- (txt), the STREAMS solution also generates outputs in
ure 3, includes a main text-box for the input text and a JSON format, which includes confidence scores and time
drop-down menu to select the desired voice. stamps at the word level for both transcription and
trans</p>
        <p>The GUIs were developed using the Angular frame- lation, and SRT subtitling format.
work9 and deployed via a Nginx web server10. Figure 2 The streaming operational mode supports the
genershows the graphical interface for transcription and trans- ation of multilingual subtitles from a video transmitted
lation purposes. over the common RTMP and HLS streaming protocols.
The streaming backend was built to manage both the
2.2. REST API input video and audio channels, generate a new subtitle
channel in the original or translated languages, and
conThe principal REST API connected to the Commander struct a new output stream with the 3 channels of video,
serves as the main interface between the GUI and the audio and subtitles perfectly synchronised in time.
backend for batch processing. Additionally, it provides Managing the input video and audio channels during
an alternative way for the user to directly access all the the subtitle generation process enables the generation
features of the solution via http requests, allowing third of perfectly aligned subtitles within a controlled timing
party systems to be built on top of STREAMS and thus of the video broadcast, which should not exceed a
conextend their functionality. The principal communication ifgurable maximum delay. Depending on the content
between all modules of the solution are performed via and speech rates, a maximum delay of between 15 and
REST APIs as well. 30 seconds with respect to the original video was
suficient in our tests to guarantee good quality and stable
2.3. Backends results over time, without information loss. Based on
the experience of the companies in the consortium that
The STREAMS solution ofers two types of backends specialise in streaming, these delays were considered
perdepending on the operational mode under which contents fectly acceptable in streaming applications designed for
are provided as input, supporting batch and streaming web broadcasts. Finally, the output video streams can be
processing modes. transmitted over the same RTMP and HLS streaming
pro9https://angular.io/ tocols, being thus compatible with the most important
10https://www.nginx.com/ commercial video streaming platforms.</p>
        <p>Both batch and streaming backends are composed of engine, the live transcription engine implements an
onseveral modules which support the features of the solu- line decoding process, in which the recognition models
tion. Whereas the backend for batch processing includes are previously loaded in memory and the features are
all the modules, the streaming backend is focused on processed in real time, without having to wait until the
generating good quality and well-timed subtitles, thus audio has finished to start decoding.
integrating three main modules dedicated to live
transcription, machine translation and subtitles generation. 2.3.2. Machine Translation
All the main modules are described in turn in the
following subsections.</p>
      </sec>
      <sec id="sec-2-3">
        <title>The Machine Translation (MT) module is based on Vi</title>
        <p>
          comtech’s proprietary Itzuli Translator engine, a robust
2.3.1. Ofline and Live transcription and scalable text translation system, which can be
deployed under Kubernetes orchestration12 or as a
stanThe Ofline and Live transcription modules are based on dalone platform in a dedicated server. The module
inteVicomtech’s propietary Transkit software library, which grates MarianNMT [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] by default in its own backend to
implements a diferent logic depending on the opera- perform eficient NMT inference.
tional mode. The library ofers easy access to speech tran- On top of the Itzuli MT engine, a dedicated layer was
scription functionalities through a REST API, supports developed to support diferent processing scenarios. The
concurrent processing, can be deployed as a standalone module can thus handle input delivered in raw text
forapplication or in scalable mode with automatic request mat, or as arrays of transcribed words with their
cortrafic balancing, and includes dynamic management of responding timestamps for streaming scenarios. Input
decentralised transcription instances. management, including preprocessing and bufering in
        </p>
        <p>
          The library is composed of 7 technological modules streaming scenarios, is handled by the MT module, which
connected via configurable pipelines. The modules corre- can return either a fully postprocessed text or arrays of
spond to an audio transcoder which integrates the FFm- translated word with an approximated time distribution.
peg11 tool, an acoustic segmenter based on the Voice The STREAMS MT module also provides support for
Activity Detector module proposed by [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], an Auto- end-to-end batch translation of subtitle files in SRT
formatic Speech Recognition (ASR) module built on top mat. In this scenario, source text is automatically
exof Kaldi [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], a module for automatic punctuation and tracted from the original file, preserving the time codes,
capitalisation [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], a rule-based text normaliser, and a fi- sentences are reconstructed from the original subtitles to
nal postprocessing module in charge of generating the provide better sentence-level translation, and the
transdiferent output formats. lated text is reinserted into the original time code
struc
        </p>
        <p>Although all these components can be used for batch ture via an information redistribution algorithm. The
processing, 3 modules are mainly used for streaming, translated text is thus distributed into each target
subwhich correspond to live transcription, punctuation and title according to the original character distribution in
capitalisation, as well as normalisation. Unlike the ofline the corresponding source subtitle, with additional
mini11https://fmpeg.org/
mal heuristics to enforce key subtitling constraints in the
target language, e.g., enforcing subtitle breaks on final
punctuation markers or preventing subtitle lines with
single words.</p>
        <sec id="sec-2-3-1">
          <title>2.3.3. Subtitle generator</title>
          <p>The main purpose of the subtitle generator component
is to provide quality subtitles that adhere to standards of
readability and usability, featuring appropriate
segmentation and persistence, among other criteria. This module
can operate in batch or streaming scenarios, returning
formatted subtitles from either arrays of words with their
timestamps or plain text.</p>
          <p>The module includes a portable and eficient
segmentation component, to provide quality subtitles that balance
grammatical segmentation and adequate persistence. It
includes a character-counting method, which ensures
that subtitles respect the maximum number of characters
per line. The second method for subtitle segmentation
is a simple rule-based approach, which uses character
counting as well as enforcing segmentation on
punctuation marks. These methods are included as fail-safe
in scenarios where minimal latency is critical, but are
typically discarded in favour of more advanced methods
within the platform in most scenarios.</p>
          <p>
            Two additional classes of methods are provided for
quality subtitle generation within STREAMS. The first
one relies on supervised models, based on Conditional
Random Fields (CRF) as proposed by [
            <xref ref-type="bibr" rid="ref7">7</xref>
            ], and DNNs.
These models provide quality segmentation but require
an ofline training phase on annotated data. The second
type of method is fully unsupervised and is based on
the masking prediction of a Masked Language Model
(MLM)13, as proposed by [
            <xref ref-type="bibr" rid="ref9">9</xref>
            ]. This approach provides
coverage for all languages supported by the MLM model and
delivers quality segmentation at a minimal cost. Within
STREAMS, this approach was further extended to
integrate other subtitling constraints, such as the optimal
number of characters per second, subtitle persistence,
or the length balance between lines in a subtitle. The
importance of each constraint can be configured to
generate subtitles that adhere to the standards of specific
broadcasters.
          </p>
        </sec>
      </sec>
      <sec id="sec-2-4">
        <title>The module consists of 3 main components, each one</title>
        <p>in charge of a specific part of the process: pre-processing
the input text to normalise numbers and special
characters, transforming the input text into a waveform via
Text-To-Speech (TTS) technology , and finally,
formatting the output into the requested audio format.</p>
        <p>
          The core TTS technology is based on the Tacotron-2
architecture [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. This model consists of a
sequence-tosequence model, which includes an encoder, a decoder
with attention and a final post-processing convolutional
neural network (CNN). For the STREAMS solution, one
synthetic voice was developed and integrated for each of
the 4 languages emphasised within the project, enabling
the generation of audio content based on input text. It
is worth noting that combining this TTS module with
the ASR and MT modules, a pipeline for speech
translation purposes could be easily created using the proposed
architecture, among other interesting applications.
        </p>
        <sec id="sec-2-4-1">
          <title>2.3.5. Audio-description</title>
          <p>Audio-description (AD) is a growing media access
service which complements subtitling by providing audio
information for key visual elements in media content, a
particularly useful feature for blind and/or visually
impaired people. This information is commonly generated
manually by dubbers in recording studios, using
narrative and neutral voices to describe the scenes and their
visual elements.</p>
          <p>The STREAMS solution incorporates an AD module
which enables the automatic generation of this
information using the TTS module, given that the synthetic
voices were built on content with a narrative and
journalistic style. Taking as input a script with the time-marked
descriptions of the scenes, the AD module generates a
master audio of the same length of the script and/or
content, including each audio segment at the corresponding
timestamp. The DocParser sub-module developed for
the platform (see Figure 1) aids in processing scripts in
various formats and extracting information as needed.
2.3.6. Aligner</p>
        </sec>
      </sec>
      <sec id="sec-2-5">
        <title>The Aligner is a complementary module that enables</title>
        <p>
          generating subtitles or time-marked word transcriptions
2.3.4. Speech synthesis from an audio input and its corresponding script.
This module was developed and integrated in response
The speech synthesis module is based on Vicomtech’s to the need expressed by several entities, such as dubbing
proprietary VicomTTS software library for neural speech companies, to generate subtitles from pre-existing script
synthesis, which supports concurrent processing, deploy- and audio. Using the acoustic models and the
forcedment under local or scalable modes with automatic re- recognition mode of the Ofline Transcription module,
quest trafic balancing, as well as dynamic management the Aligner module generates well-timed subtitles by
of decentralised synthesis instances, via REST API access. performing an alignment between the text and the audio,
thus recovering the timestamps at the word level. The
13We use multilingual BERT by default [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ].
        </p>
      </sec>
      <sec id="sec-2-6">
        <title>DocParser sub-module is used to process the input scripts and extract the texts properly as well.</title>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Conclusions</title>
      <p>We described STREAMS, a batch and streaming platform
which integrates and manages AI services for rich
transcription, translation, voice synthesis and audio
description. STREAMS is meant to enhance the processes and
products of companies working in multiple sectors and
has been successfully deployed in real-life batch and
multilingual streaming scenarios during the project, ofering
a rich array of multilingual services with innovative
methods, in particular to enhance portability, responsiveness
and output readability.</p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgments</title>
      <sec id="sec-4-1">
        <title>STREAMS was partially funded by the Basque Business</title>
        <p>Development Agency, SPRI, under grant agreement
ZL2023/00211.</p>
        <p>The authors would also like to thank the companies in
the consortium that have contributed with their
knowledge and experience to the project and the STREAMS
solution.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>R.</given-names>
            <surname>Prabhavalkar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Hori</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. N.</given-names>
            <surname>Sainath</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Schlüter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Watanabe</surname>
          </string-name>
          ,
          <article-title>End-to-end speech recognition: A survey</article-title>
          ,
          <source>IEEE/ACM Transactions on Audio, Speech, and Language Processing</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaswani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Parmar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Gomez</surname>
          </string-name>
          , Ł. Kaiser,
          <string-name>
            <surname>I. Polosukhin</surname>
          </string-name>
          ,
          <article-title>Attention is all you need</article-title>
          ,
          <source>in: Advances in Neural Information Processing Systems</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>6000</fpage>
          -
          <lpage>6010</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>H.</given-names>
            <surname>Dinkel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <article-title>Voice activity detection in the wild: A data-driven approach using teacher-student training</article-title>
          ,
          <source>IEEE/ACM Transactions on Audio, Speech, and Language Processing</source>
          <volume>29</volume>
          (
          <year>2021</year>
          )
          <fpage>1542</fpage>
          -
          <lpage>1555</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D.</given-names>
            <surname>Povey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ghoshal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Boulianne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Burget</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Glembek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hannemann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Motlicek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Qian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Schwarz</surname>
          </string-name>
          , et al.,
          <article-title>The kaldi speech recognition toolkit</article-title>
          , in: IEEE 2011 workshop
          <article-title>on automatic speech recognition and understanding</article-title>
          ,
          <source>CONF, IEEE Signal Processing Society</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>González-Docasal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>García-Pablos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Arzelus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Álvarez</surname>
          </string-name>
          ,
          <article-title>Autopunct: A bert-based automatic punctuation and capitalisation system for spanish and basque</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>67</volume>
          (
          <year>2021</year>
          )
          <fpage>59</fpage>
          -
          <lpage>68</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Junczys-Dowmunt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Grundkiewicz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Dwojak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Hoang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Heafield</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Neckermann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Seide</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.</given-names>
            <surname>Germann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. Fikri</given-names>
            <surname>Aji</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Bogoychev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. F. T.</given-names>
            <surname>Martins</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Birch</surname>
          </string-name>
          , Marian:
          <article-title>Fast neural machine translation in C++</article-title>
          ,
          <source>in: Proc. of ACL</source>
          <year>2018</year>
          ,
          <year>2018</year>
          , pp.
          <fpage>116</fpage>
          -
          <lpage>121</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Alvarez</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.-D.</surname>
            Martínez-Hinarejos,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Arzelus</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Balenciaga</surname>
          </string-name>
          ,
          <source>A. del Pozo</source>
          ,
          <article-title>Improving the automatic segmentation of subtitles through conditional random field</article-title>
          ,
          <source>Speech Communication</source>
          <volume>88</volume>
          (
          <year>2017</year>
          )
          <fpage>83</fpage>
          -
          <lpage>95</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , BERT:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          , in: J.
          <string-name>
            <surname>Burstein</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Doran</surname>
          </string-name>
          , T. Solorio (Eds.),
          <source>Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>1</volume>
          (Long and Short Papers),
          <source>Association for Computational Linguistics</source>
          , Minneapolis, Minnesota,
          <year>2019</year>
          , pp.
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>D.</given-names>
            <surname>Ponce</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Etchegoyhen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Ruiz</surname>
          </string-name>
          ,
          <article-title>Unsupervised subtitle segmentation with masked language models</article-title>
          , in: A.
          <string-name>
            <surname>Rogers</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Boyd-Graber</surname>
          </string-name>
          , N. Okazaki (Eds.),
          <source>Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume</source>
          <volume>2</volume>
          :
          <string-name>
            <surname>Short</surname>
            <given-names>Papers)</given-names>
          </string-name>
          ,
          <source>Association for Computational Linguistics</source>
          , Toronto, Canada,
          <year>2023</year>
          , pp.
          <fpage>771</fpage>
          -
          <lpage>781</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Skerry-Ryan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Stanton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Weiss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Jaitly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xiao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bengio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Le</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Agiomyrgiannakis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Clark</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. A.</given-names>
            <surname>Saurous</surname>
          </string-name>
          , Tacotron: Towards End-to-
          <source>End Speech Synthesis</source>
          (
          <year>2017</year>
          )
          <fpage>4006</fpage>
          -
          <lpage>4010</lpage>
          . doi:
          <volume>10</volume>
          .21437/Interspeech. 2017-
          <volume>1452</volume>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>