<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>EMPHASIS: Empowering Decision Making with Higher Productivity by Means of HyperAutomation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Montse Cuadros</string-name>
          <email>mcuadros@vicomtech.org</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aitor Álvarez</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Naiara Pérez</string-name>
          <email>naiara.perez@ehu.eus</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Juan Manuel Martín</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pablo Turón</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Elena Zotova</string-name>
          <email>ezotova@vicomtech.org</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Haritz Arzelus</string-name>
          <email>harzelus@vicomtech.org</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Joaquin Arellano</string-name>
          <email>jarellano@vicomtech.org</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Arantza del Pozo</string-name>
          <email>adelpozo@vicomtech.org</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>HiTZ Center - Ixa, University of the Basque Country UPV/EHU, Manuel Lardizabal Pasealekua 1, Donostia/San-Sebastian</institution>
          ,
          <addr-line>20018</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Vicomtech Foundation, Basque Research and Technology Alliance (BRTA)</institution>
          ,
          <addr-line>Mikeletegi Pasealekua 57, Donostia/San-Sebastian, 20009</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The EMPHASIS project aims to support companies in the hyper-automation of tasks by researching and developing tools adaptable to diferent use cases using speech and natural language processing technologies. The toolset is deployed on the EMPHASIS platform, allowing the building of fully customisable and flexible processing pipelines with a stack of diferent text and speech components adaptable to specific domains and languages. All the adaptation of components is available as a feature in the main EMPHASIS tools, allowing end-users to adjust the technology to the domains and languages required by each application scenario.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;NLP</kwd>
        <kwd>NER</kwd>
        <kwd>Text classification</kwd>
        <kwd>Document classification</kwd>
        <kwd>Question Answering</kwd>
        <kwd>Hyperautomatization</kwd>
        <kwd>ASR</kwd>
        <kwd>Diarization</kwd>
        <kwd>Emotions</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
    </sec>
    <sec id="sec-2">
      <title>2. Consortium and funding body</title>
      <sec id="sec-2-1">
        <title>EMPHASIS has been partially funded by the Basque</title>
        <p>government through the Hazitek Estratégico 2021
programme of the SPRI Group under grant agreement ZE- 4. Challenges
2021/00039. It has run from 05/2021 to 12/2023. The
consortium is led by Grupo Teknei1, Ibermatica2, Natural The main technological challenges of EMPHASIS are
Vox3, Eutik4, Gureak5, Segula6, Zuchetti7 and Merkatu8. related to the application of the following technologies
The research centres involved are Vicomtech, Ibermat- to diferent languages, domains, audio and text formats,
ica’s R&amp;D business unit (I3B) and the Speech Interactive together with the development of tools to allow their
Group from the UPV/EHU9. easy adaption to the real use case of each client:
tomised and adapted to diferent use cases, following the
principles of low-code.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Goals and expected results</title>
      <sec id="sec-3-1">
        <title>The aim of this project is to take the back-ofice towards what has been called Hyperautomation or Cognitive Automation, which proposes to integrate AI with RPA solutions.</title>
        <p>The integration of speech and natural language
processing technologies makes it possible to structure
content that has not been structured until now, such as
documents (invoices, delivery notes, deeds, building permits,
municipal licences, etc.) without having to have an
associated template, or calls (in call centres, for example).
It is a matter of converting all this content (documents
and calls) into data, so that it can be managed, analysed
and the associated processes can be optimised, in order
to make progress in a real digital transformation, which
otherwise seems impossible. It is not a matter of doing
digitally the same as we do now, but rather, based on
data, defining the optimal flows and processes that
guarantee total digitisation, scalable, robust, flexible in the
face of changes, adaptable to new scenarios, and always
with reasonable costs. To this end, another key factor is
to follow the principles of no code or, failing that, low
code, i.e., to develop solutions whose integration and
implementation are as transparent as possible, avoiding
the costly processes involved in BPM (Business Process
Management) solutions.</p>
        <p>The aim is therefore to develop technology that can
be integrated as APIs in third-party solutions, so that
their configuration can be carried out by non experts. In
the same way, the aim is for these APIs to encapsulate
functions that can be customised, i.e., one of the main
objectives of the project is that the functions developed
within the framework of the project can later be
cus1https://www.teknei.com/
2https://ibermatica.com/
3https://naturalspeech.es/
4https://www.eutik.com/
5https://www.gureakmarketing.com/es/
6https://www.segulatechnologies.com/es/
7https://www.zucchetti.es/
8https://www.merkatu.com/
9https://www.ehu.eus/en/web/speech-interactive
• Cognitive Document Automation: its
objective is to automate the extraction of relevant
content from textual sources of diferent formats
(invoices, delivery notes, e-mails, text documents,
FAQ systems, conversations extracted by
automatic transcription [3]), to classify content into
diferent categories [ 4], domains, feelings, and
also extract core-entities (NER) and their
relationships that allow the understanding of its content.
The technology to be used is cutting-edge Deep
Learning technology where a paradigm shift has
been seen in the last two years thanks to the
proliferation of advanced neural architectures that
exploit language models with knowledge of the
world. Documents containing images with text
inside are also converted, and their reading is
improved by applying advanced methods on
stateof-the-art OCR tools such as LayoutLM [5].
Additionally, technology based on Question
Answering techniques can also be used in order to have
tools to search into documents and find relevant
information.
• Speech Analytics: its objective is to automate
the analysis of data with acoustic content by
transformation to text using advanced tools for
automatic enriched speech transcription and
emotion analysis [6]. The main technology to be used
is Deep Learning technology based on the latest
contributions from the scientific community to
the state of the art and which has shown great
improvements with respect to previous Machine
Learning technology.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5. Use cases</title>
      <sec id="sec-4-1">
        <title>EMPHASIS proposes a set of application scenarios that</title>
        <p>form a common denominator in terms of technological
needs and will allow us to join forces towards the
development of a global solution for the cognitive automation
of processes related to documents and audios in Spanish.
The main application scenarios can be divided at a more
technological level into two subgroups where:</p>
      </sec>
      <sec id="sec-4-2">
        <title>All in all, these technologies have been divided into different verticals or domains related to use-cases that have been defined in EMPHASIS by the diferent companies participating in the project:</title>
        <p>• Text processing technology is linked to Cognitive Contact Centres: In relation to the use of Speech
An</p>
        <p>Document Automation (CDA). alytics, the main challenges are the analysis of customer
• Speech processing technology is fundamental to responses to VoiceBots to improve the processes of
aucarry out cognitive automation related to audios, tomatic interpretation of responses and to analyse their
phone calls, etc. also called Speech Analytics. emotions. In addition, if the interaction is agent-based,
early detection of satisfaction levels or performance
compliance. Other channels to analyse apart from audio are
textual channels, including chats, e-mails or social
networks, which are sources of data that will be used to
analyse user satisfaction and decision-making to improve the
solution’s knowledge flows in this use case or to interact
with the customer automatically in search of information.</p>
        <p>In this sense, texts will be analysed in order to classify
them by subject matter. Finally, the last channel to be
used in this use case is the physical document, which
must be converted into a textual document using
computer vision and OCR techniques.</p>
        <p>Banking: Speech Analytics is used to work on
solutions that allow the acoustic analysis of audios through
the enriched transcription of their content for subsequent
classification of customers in quadrants. In this use case,
the transcription of the content will automate the
management processes of customer requests and their correct
attention. In the case of Text Analyics, the detailed
analysis of textual documents to make a correct segmentation
of their content, classify them by language and typology 6. Approach
and extract the metadata. In addition, there is a need
to go a step further and manage documents in order to Therefore, we are talking about developing basic
soluautomate the search for documents related to questions, tions for document and audio processing, but at the same
in environments such as FAQs. time, we are talking about developing tools and systems
that allow the customisation of these solutions. To give
Retail: The main objective of the project has been to an example, if one of the base solutions to be generated
provide technology for the automatic processing of doc- allows the classification of documents and whether it is
uments such as invoices, delivery notes to facilitate the an invoice or a delivery note, tools will also be generated
integration of the content of these documents into an to help in the process of annotating other documents
ERP automatically or with minimum supervision. and generating the corresponding AI models, so that
in the future the customised solution can classify
docJustice: Using Speech Analytics, the project aims to uments by identifying whether it is a building permit,
analyse audios in the legal field, through enriched tran- a deed, or a municipal licence. All modules developed
scription, identification of speakers, through biometrics in the project have been deployed as REST APIs within
[7], detection of emotions in the audios. In addition a Docker container, so each module could be used as a
to creating transcription routines with time stamps to single tool. However, the project has developed a core
quickly locate the parts corresponding to a transcribed platform where diferent sets of tools can be
concatetext. In terms of text analysis, there is a huge need of nated as fully configurable pipelines. This allows
buildusing Anonymization [8]. ing pipelines per domain and per language based on each
partner or each use-case needs. Figure 1 shows a diagram
of the EMPHASIS platform. Last but not least, the
platform allows resource, queue, and pipeline management,
and also load-balance configurations.</p>
        <p>Administration: The main objectives are the detailed
analysis of textual documents to make a correct
segmentation of their content, classify them by language and
typology and extract the metadata. In addition, there is
a need to go a step further and manage documents in
order to automate the search for documents related to
questions, in environments such as FAQs.</p>
        <p>Industry: In relation to the use of Speech Analytics,
the main challenge is to classify incidents by voice and
in real time. In terms of Textual Analytics, the main
challenges to be worked on in the project are the
classification of textual documents, analysis of e-mails and
attachments to manage content and classify it according
to the specifications of each domain.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>7. Main technological modules</title>
      <sec id="sec-5-1">
        <title>In this section, we present the main technological mod</title>
        <p>ules developed in the framework of the EMPHASIS
project based on Text and Speech Technologies, which
are integrated, as explained in section 6, into the
EMPHASIS platform.
7.1. Text technologies
training new models to classify types of documents, and
using trained models to classify documents.</p>
        <p>Text Classification This module focuses on
finetuning a foundational model (LLM), in this case, BERT [9],
to perform text classifications for a set of categories. The
training process enables the addition of diferent sources
of texts, and of symbolic features, exploited by the
calculation of an embedding layer which is concatenated
to the output hidden states of the LLMs. This module
can be used for diferent applications such as sentiment
analysis and news classification. The module has two
main functionalities: the training of new models and the
use of already trained models through a REST API.</p>
        <p>NERC and Anonymization This module contains
two main tools. The first one is a sequence-labeller trainer
that could be used to train a NERC disambiguator with
an annotated dataset at entity level. In this case, we have
implemented a trainer that accepts an arbitrary number
of entities and training a LLM to recognise and classify
entities. Again, this module could be used to train new
models or to use it with existing ones. The second tool,
named anonimisator, performs the replacement of the
detected entities of the NERC tool into a set of
candidates based on a taxonomy of entities. This taxonomy is
Document Classification This module contains the defined in a YML file and performs a set of techniques
entire document classification pipeline which is devel- based on the kind of entity, for instance replacement or
oped in two steps. The first one contains the extraction obfuscation. Moreover, in the case of entities such as
Perof the text from the documents via document parsing sons or Locations, a replacement could be done using a
(for text-based documents) or Tesseract OCR (for image- predefined dictionary or in the case of a category such as
based documents). The second one contains a document telephone numbers, an algorithm performing a random
trainer used to train a document classifier with diferent production of numbers could be used. All these possible
kinds of LLMs such as BERT-based [9] or LayoutLM [10]. settings are easily configurable from the YML file.
The goal of this trainer is to classify documents into a set
of labels. It could be done at the page level or the
document level. The module has two main functionalities:
Question Answering This module allows the user to identity or the total number of speakers in the audio is
get answers to a given question/query automatically. The required to perform the diarisation task.
system is designed to find answers in a database, which The diarisation module for the EMPHASIS solution
consists of a large corpus of texts written in natural lan- was developed on top of the Kaldi toolkit, through a
techguage and relevant to each use case. There are three nological approach based on the combination of
idenmodes of searching: (i) semantic similarity, (ii) extrac- tity vectors and a posterior classification task, similar
tive QA and (iii) lexical search. The semantic similarity to the work presented in [16]. The identity vectors (i.e.,
method is implemented with LLMs trained to distinguish x-vectors) are extracted from a time-delay neural
netsemantics [11], so they retrieve paragraphs or phrases work (TDNN) model [17], while the Probabilistic Linear
most relevant to a question. The extractive QA algorithm Discriminant Analysis (PLDA) is used to score the
simuses a language model pre-trained to extract a short span ilarity between embeddings. In the last step, the VBx
of text with the answer from a given document. Finally, resegmentation algorithm [18] is employed to obtain the
Lexical search includes the algorithms to perform an ifnal sequence of speakers given the speaker-specific
disexact string matching and a statistical ranking based tribution scored from the previous PLDA model.
on words. This module has, on the one hand, the
functionality of indexing and vectorising a database (from a
use-case) in Elasticsearch10 and, on the other hand, the
functionality of fast document retrieval using the modes
explained before.</p>
        <p>Emotion Detection The Speech Emotion Recognition
(SER) module was built with the aim of analysing the
emotions of agents and customers in telephone calls from a
call centre. To this end, and considering the lack of
available data for the Spanish language to train AI models
7.2. Speech Technologies within this complex domain, a new acoustic corpora was
collected and annotated for the task within the project
Online Transcription The Online Transcription mod- from a collection of telephone calls gathered from 2
comule aims to convert the input audio into text and it is based panies of the consortium, each with its particular domain.
on Vicomtech’s proprietary Transkit software library. This audio data was manually labelled considering the
The library ofers easy access to speech transcription particularities of each domain at categorical and the
threefunctionalities through a REST API, supports concurrent dimensional (arousal, valence and dominance) levels.
processing, can be deployed as a standalone application Regarding the technological approach, and following
or in scalable mode with automatic request trafic balanc- the current trends, an architecture composed of a feature
ing, and includes dynamic management of decentralised extractor (i.e., audio embeddings) module and a
classifitranscription instances. cation layer was implemented. As the feature extractor,</p>
        <p>The library is composed of 7 technological modules the Wav2Vec2.0 XLS-R foundation model [19] was
inteconnected via configurable pipelines. The modules corre- grated, whilst as classification layer both Support Vector
spond to an audio transcoder which integrates the FFm- Machines (SVM) and DNN based downstream models
peg11 tool, an acoustic segmenter based on the Voice were evaluated with a very similar performance.
Activity Detector (VAD) module proposed by [12], an
Automatic Speech Recognition (ASR) module built on
top of Kaldi [13], a module for automatic punctuation 8. Conclusions
and capitalisation [14], a rule-based text normaliser, and
a final postprocessing module in charge of generating The main goal of the EMPHASIS project is to support
diferent output formats. companies in the hyperautomation of tasks by
research</p>
        <p>With the aim of providing domain adaptation function- ing the most convenient tools adaptable to diferent
doalities, the TrainLM sub-module was developed, which mains and using the novel state-of-the-art technologies
enables the user to adapt the Language Model and Vo- based on deep learning technology. As a result of the
cabulary of the ASR at text level. project, companies have a platform with plug-and-play
core Text and Speech Analysis modules adaptable to
different domains and languages.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <sec id="sec-6-1">
        <title>EMPHASIS was partially funded by the Basque Business Development Agency, SPRI, under grant agreement ZE2021/00039. The authors would also like to thank the companies in the consortium that have contributed with</title>
        <p>Diarisation Speaker diarisation aims to solve the
problem of "who spoke when", which implies segmenting a
given audio over the active speaker and clustering the
segments belonging to each speaker by assigning the
same label to each one [15]. Unlike speaker recognition
technologies, no prior knowledge about the speakers’
10https://www.elastic.co/es/elasticsearch
11https://fmpeg.org/
their knowledge and experience to the project and the [11] N. Reimers, I. Gurevych, Making Monolingual
SenEMPHASIS solution. tence Embeddings Multilingual using Knowledge
Distillation, in: Proceedings of the 2020
Conference on Empirical Methods in Natural Language
References Processing, Association for Computational
Linguistics, 2020.
[1] A. Baidya, Document analysis and classification: [12] H. Dinkel, S. Wang, X. Xu, M. Wu, K. Yu, Voice
acA robotic process automation (rpa) and machine tivity detection in the wild: A data-driven approach
learning approach, in: 2021 4th International Con- using teacher-student training, IEEE/ACM
Transference on Information and Computer Technologies actions on Audio, Speech, and Language Processing
(ICICT), IEEE, 2021, pp. 33–37. 29 (2021) 1542–1555.
[2] A. Haleem, M. Javaid, R. P. Singh, S. Rab, R. Suman, [13] D. Povey, A. Ghoshal, G. Boulianne, L. Burget,
Hyperautomation for the enhancement of automa- O. Glembek, N. Goel, M. Hannemann, P. Motlicek,
tion in industries, Sensors International 2 (2021) Y. Qian, P. Schwarz, et al., The kaldi speech
recogni100124. tion toolkit, in: IEEE 2011 workshop on automatic
[3] A. Álvarez, H. Arzelus, I. G. Torre, A. González- speech recognition and understanding, CONF, IEEE
Docasal, Evaluating novel speech transcription Signal Processing Society, 2011.
architectures on the spanish rtve2020 database, Ap- [14] A. González-Docasal, A. García-Pablos, H. Arzelus,
plied Sciences 12 (2022) 1889. A. Álvarez, Autopunct: A bert-based automatic
[4] K. Kowsari, K. Jafari Meimandi, M. Heidarysafa, punctuation and capitalisation system for spanish
S. Mendu, L. Barnes, D. Brown, Text classification and basque, Procesamiento del Lenguaje Natural
algorithms: A survey, Information 10 (2019) 150. 67 (2021) 59–68.
[5] Y. Xu, M. Li, L. Cui, S. Huang, F. Wei, M. Zhou, [15] T. Park, N. Kanda, D. Dimitriadis, K. Han, S.
WatanLayoutlm: Pre-training of text and layout for doc- abe, S. Narayanan, A review of speaker diarization:
ument image understanding, in: Proceedings of Recent advances with deep learning, Computer
the 26th ACM SIGKDD International Conference Speech and Language 72 (2022) 101317.
on Knowledge Discovery &amp; Data Mining, 2020, pp. [16] G. Sell, D. Garcia-Romero, Speaker diarization with
1192–1200. plda i-vector scoring and unsupervised calibration,
[6] A. González-Docasal, N. Pérez, A. Alvarez, M. Ser- in: 2014 IEEE Spoken Language Technology
Workras, L. García-Sardiña, H. Arzelus, A. García Pablos, shop (SLT), IEEE, 2014, pp. 413–417.
M. Cuadros Oller, P. Delgado, A. Lazpiur, B. Romero, [17] V. Peddinti, D. Povey, S. Khudanpur, A time delay
Nalytics: Natural speech and text analytics, 2020- neural network architecture for eficient modeling
09. of long temporal contexts, in: Sixteenth annual
con[7] J. M. Martín-Doñas, I. G. Torre, A. Álvarez, J. Arel- ference of the international speech communication
lano, The vicomtech spoofing-aware biometric association, 2015.
system for the sasv challenge, arXiv preprint [18] F. Landini, J. Profant, M. Diez, L. Burget, Bayesian
arXiv:2204.01399 (2022). hmm clustering of x-vector sequences (vbx) in
[8] O. de Gibert Bonet, A. García Pablos, M. Cuadros, speaker diarization: theory, implementation and
M. Melero, Spanish datasets for sensitive entity analysis on standard tasks, Computer Speech &amp;
detection in the legal domain, in: Proceedings of Language 71 (2022) 101254.
the Thirteenth Language Resources and Evaluation [19] A. Baevski, Y. Zhou, A. Mohamed, M. Auli, wav2vec
Conference, European Language Resources Associ- 2.0: A framework for self-supervised learning of
ation, Marseille, France, 2022, pp. 3751–3760. URL: speech representations, Advances in neural
inforhttps://aclanthology.org/2022.lrec-1.400. mation processing systems 33 (2020) 12449–12460.
[9] J. Devlin, M. Chang, K. Lee, K. Toutanova,</p>
        <p>BERT: Pre-training of Deep Bidirectional
Transformers for Language understanding, CoRR
abs/1810.04805 (2018). URL: http://arxiv.org/abs/
1810.04805. arXiv:1810.04805.
[10] A. R. GV, Q. You, D. Dickinson, E. Bunch, G. Fung,</p>
        <p>Document Classification and Information
Extraction framework for Insurance Applications, in:
2021 Third International Conference on
Transdisciplinary AI (TransAI), 2021, pp. 8–16. doi:10.1109/
TransAI51903.2021.00010.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>