<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Unsupervised Anomaly Detection in Predictive Maintenance using Sound Data⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Antonino Ferraro</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Antonio Galli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Valerio La Gatta</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vincenzo Moscato</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marco Postiglione</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giancarlo Sperlì</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francesco Moscato</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Electrical Engineering and Information Technology, University of Naples Federico II</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Information Engineering, Electrical Engineering and Applied Mathematics, University of Salerno</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper represent an extended abstract of a recent proposal ([1]), in which the authors present a new methodology for unsupervised anomaly detection in predictive maintenance using sound data. In particular, the methodology leverages LSTM and CNN-based autoencoders to process continuous audio streams from diferent audio sources in real-world factories based on a customized sliding window. The novelties of the proposed approach include a general methodology for unsupervised anomaly detection, machine ID encoding using one-hot encoding, and conditioning an autoencoder by jointly analyzing the relationships between the mel-spectrogram and the machine ID to compute an anomaly score. The methodology achieves good performances in terms of efectiveness and in addition low inference time and memory requirements.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Deep Learning</kwd>
        <kwd>Anomaly Detection</kwd>
        <kwd>Industrial AI</kwd>
        <kwd>Predictive maintenance</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Over the past two decades, Anomalous Sound Detection (ASD) has become an increasingly
challenging task in various applications. ASD is used to identify whether the sound emitted
from an object is normal or anomalous, and early detection can prevent critical problems ([
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ]).
Traditional approaches use supervised machine learning models ([
        <xref ref-type="bibr" rid="ref2 ref4">2, 4</xref>
        ]), while unsupervised
models have also been used to distinguish between normal and abnormal situations ([
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ]).
Recently, deep learning approaches have been successfully exploited in various contexts ([
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ]).
In industrial applications, existing anomaly detection systems rely on monitoring sensors, with
most systems utilizing visual detection methods. However, these methods have limitations
such as illumination or occlusion, which can hinder their performance, particularly in real-time
applications. Additionally, sensor-based analysis systems can be vulnerable to deception, as
demonstrated by the Triton malware and Stuxnet virus. As a result, there is growing interest in
the use of additional features, such as audio streams, that can be analyzed externally by the
system to improve performance and reliability.
      </p>
      <p>
        Various Anomalous Sound Detection (ASD) systems have been developed to address the
limitations of traditional techniques ([
        <xref ref-type="bibr" rid="ref10 ref2 ref9">2, 9, 10</xref>
        ]). These systems use acoustic data features,
such as Mel-scale1 spectrograms or air pressure values, to train neural networks, typically
using autoencoder architectures ([
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]). Studies have compared the performance of diferent
autoencoder architectures for ASD, including Convolutional LSTM, sequential Convolutional
Autoencoder, LSTM-based autoencoder, dense and convolutional architectures,
Transformerbased and Conformer-based autoencoder ([
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]). A few-shot approach, called SNIPER, has
been designed by ([
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]) to overcome the problem of insuficient observed anomalies. However,
additional information related to the equipment, such as equipment ID, can potentially improve
the patterns learned from the model.
      </p>
      <p>
        This paper represent an extended abstract of a recent proposal [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], in which the authors
present a new methodology for unsupervised anomaly detection in predictive maintenance using
sound data. The methodology is flexible and eficient for real-world scenarios, allowing it to be
applied to multiple instances of the same or diferent equipment, and can instantiate diferent
types of autoencoders. The authors evaluate the methodology using LSTM and CNN-based
autoencoders to process continuous audio streams from diferent audio sources in real-world
factories based on a customized sliding window. The novelties of the proposed approach include
a general methodology for unsupervised anomaly detection, machine ID encoding using one-hot
encoding, and conditioning an autoencoder by jointly analyzing the relationships between the
mel-spectrogram and the machine ID to compute an anomaly score. The methodology achieves
low inference time and memory requirements and has been evaluated on audio streams from
multiple machine types. Finally, the proposed methodology enables application in real scenario
by analyzing continuous audio stream coming from diferent audio sources on the basis of a
customized sliding windows.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Methodology</title>
      <p>The Anomaly Detection task is a challenging aspect of predictive maintenance, which involves
identifying anomalous warning or failure states in industrial machines to improve
maintenance scheduling. Various sensor-based techniques have been proposed for addressing this
task. However, cyber-physical attacks like Triton or Stuxnet have raised new challenges that
require the investigation of novel features, including equipment sounds. This article proposes a
methodology for addressing the Anomaly Detection task, consisting of an Ofline Training and
Online Operation phase. These phases are discussed in Sections 2.1.1 and 2.1.2.
1The Mel-Scale is a frequency scale that segments the entire frequency spectrum into  uniformly spaced frequency
bins. The term ’evenly spaced’ is used to denote that the separation of frequency bins along the frequency dimension
more closely approximates the sensitivity of the human auditory system compared to the linearly spaced frequency
bands that are typically employed in spectrograms.</p>
      <sec id="sec-2-1">
        <title>2.1. Methodology Description</title>
        <p>The proposed ASD methodology is composed by two main phases: an ofline phase (Figure 1),
aiming to train an autoencoder model on the basis of extracted features from a pre-collected
normal audio clips (Figure 2), and an online operation phase (Figure 3), that supports analysis
and detection in real scenario.</p>
        <sec id="sec-2-1-1">
          <title>2.1.1. Ofline Training Phase</title>
          <p>The ofline training phase includes three modules: Audio Pre-processing, IDs Pre-processing,
and ID Conditioned autoencoder. The Audio Pre-processing consists of two components,
MelSpectrogram extractor and Normalization and Frames generator, and extracts features from
audio signals. The IDs Pre-processing encodes the ID string code of each machine version. The
ID Conditioned autoencoder jointly analyzes the outputs of the first two modules and computes
an anomaly score using an encoder-decoder architecture. Mel-Spectrogram extractor produces
log-mel-scale images, and Normalization and Frames generator segments them into overlapping
frames.</p>
          <p>Fig. 2 outlines the pertinent parameters for feature extraction from an audio signal using
Short-Time Fourier Transform (STFT), including the window length (_  ), window overlap
(ℎ_ℎ), and number of Mel scale bins () used for spectrogram transformation.</p>
          <p>The ID pre-processing step employs One-Hot Encoding to convert machine IDs to binary
sequences of equal length, allowing the encoder-decoder architecture to diferentiate between
sound signals from diferent machine versions. This ensures that frames of the same spectrogram
are associated with the same binary sequence. The ID binary sequences are incorporated into
the training process via the ID Conditioned Autoencoder module.</p>
          <p>This module consists of an autoencoder and an ID Conditioning Neural Network. The
autoencoder uses an encoder-decoder architecture to reconstruct input spectrogram frames.
The encoder encodes the input into a latent representation, which is then reconstructed by
the decoder. The ID Conditioning Neural Network maps one-hot-encoded ID arrays into
conditioning functions that are combined with the output of the decoder. The goal of ID
conditioning is to inform the model about the presence of diferent machines for the recognition
of their diferent normal behaviors.</p>
          <p>We introduce concatenation to reduce the number of false negatives because normal sound
of a machine  with ID  could be diferent from normal sound of a machine  with ID 
this could generate some false negatives.</p>
          <p>The key concept is that the autoencoder must be trained to reconstruct normal audio
spectrograms in input only if the provided ID is correct. With this assumptions, after the training, if a
normal test sample is placed in input, a low reconstruction error (in terms of mean absolute error
or mean squared error) is expected, while if there is an anomalous one, an high reconstruction
error is generated, even if this anomalous behavior is similar to a normal behaviour of another
machine. The similarity problem is so resolved by the presence of the ID.</p>
          <p>Nevertheless, the training process needs to be revised for supporting the autoencoder in
recognizing the relationships between machine identifiers (IDs) and audio signal, because the
ID conditioning in latent space is not enough. For this reason, the Label Generation module
randomly changes with a probability 1 −  the correct  binary sequence associated to an
audio signal with another one available. In particular, it adds the string match or not-match
(corresponding to the output of the Random Match - Non Match association module) for each
frame associated to the same audio clip on the basis of decisions.</p>
          <p>Furthermore, a new loss must be used and tuned in the training process because the
classical diference between the encoder input and the decoder output is not enough because the
association between ID and audio signal may not be correct.</p>
        </sec>
        <sec id="sec-2-1-2">
          <title>2.1.2. Online Operation Phase</title>
          <p>During the operational phase, our approach should be capable of processing a continuous
audio stream from machines using sliding windows. Figure 3 illustrates the architecture for
online operation, where the left side analyzes the raw audio signal stream using sliding windows.
The sliding window component samples the last  seconds from the stream every ℎ seconds.
The extracted audio signals are processed through the mel-spectrogram extractor and the frame
generator, whose details are explained in the ofline phase. The goal of the Anomalous Sound
Detection is to classify a sound signal as normal or anomalous. This subsystem consists of the
Pre-trained Autoencoder, the Reconstruction Error Calculator, and the Thresholding.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Experimental Evaluation</title>
      <p>Our aim is to investigate the sound analysis to deal with the more recent cyber-physical attacks,
whose aim is to deceive monitoring platform afecting the performance of classical predictive
maintenance tools.</p>
      <p>In this section we described the experimental evaluation of the proposed approach in terms of
eficacy and eficiency on the DCASE dataset, whose characterization is shown in Section 3.1. We
further discussed pre-processing phase for generating images from audio signals (see Section 3.2)
and autoencoder structure, also optimized the related hyperparameter (see Section 3.4). Finally,
performance metrics are described in Section 3.5.</p>
      <p>Two types of autoencoders are considered in order to investigate the ID conditioning efects,
also analyzing its compatibility with diferent encoding and decoding processes. According
to their autoencoder’s models, the overall architectures are identified as ID Conditioned LSTM
Autoencoder (IDC-LSTM-AE) and ID Conditioned Convolutional Autoencoder (IDCCAE), that are
implemented and trained on four machines available in DCASE dataset.</p>
      <sec id="sec-3-1">
        <title>3.1. Dataset and Recording Procedure</title>
        <p>
          We evaluated our methodology on the Unsupervised Detection of Anomalous Sounds for Machine
Condition Monitoring dataset, provided by DCASE 2020 TASK 2 belonging to MIMII dataset ([
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]).
In particular, it contains audio clips recorded from four diferent machine types (pumps, valves,
slide rails and fans), each one composed by four diferent versions.
        </p>
        <p>In conclusion, four models have been trained, one for each machine type, using training and
test sets of all available IDs.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Pre-Processing Phase</title>
        <p>This section describes the pre-processing operations on the dataset for performing the
experimental analysis, also discussing parameters selections regarding mel-spectrograms extraction,
normalization, frames generation and IDs pre-processing. In particular, the same parameters
are used for all machines in the mel-spectrograms extraction task, that has been performed by
using the Librosa library2: the number of bins (_) is 128, the STFT window (_  ) is
1024 and the ℎ_ℎ is 512.</p>
        <p>Due to the duration of each clip (10 seconds) and the above mentioned parameters, each
mel-spectrogram has the dimension of 128 × 313.</p>
        <p>Finally, four models for each architecture type must be trained to detect eventual anomalies.
Moreover, for match and not-match transformations an  = 0.75 is chosen, while the vector C
is chosen equal to 5, after optimization.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Autoencoder Structure</title>
        <p>This section outlines the structural design of the IDCCAE and IDC-LSTM-AE . The conditioning
phase is a sequence of mathematical operations involving the encoder output and the ID
conditioning network output (as detailed in Section 2). Specifically, the Encoded ID is analyzed
through a dense and activation layer to produce the ID conditioning network’s first output,
which is then multiplied with the encoder output. The second output is computed via a dense
layer with the same input provided to the Encoded ID, and the final representation (decoder
input) is obtained by adding the multiplication output to the second output.</p>
        <p>The encoder network comprises of five hidden layers with convolutional filters ranging from
32 to 512. Each encoder block includes a convolutional layer followed by batch normalization
and ReLU activation. The bottleneck encompasses a layer with 40 convolutional filters, which
reduces the encoder feature maps to a 40-dimensional encoded representation of the input.
The decoder network involves a fully-connected layer that reshapes its input to the encoder’s
last layer’s shape. Furthermore, five ConvolutionTransposeBlocks mirror the encoder, where
each block contains Conv2DTranspose layers, batch normalization layers, and ReLu activation
functions. Conditioning operations are as explained previously.</p>
        <p>The LSTM based autoencoder consists of an encoder with three LSTM layers (64,32 and 16
units) and a decoder which is the reversed version of the encoder with a RepeatVector layer. The
input is seen as time-series of 32 timesteps, characterized by 128 frequency amplitude features.</p>
      </sec>
      <sec id="sec-3-4">
        <title>3.4. Hyperparameters Tuning</title>
        <p>In this section we discuss about the hyperparameter tuning strategy of the proposed model. In
particular, we explore the following model parameters: the constant vector C, which must be
reconstructed by the autoencoder when the provided ID is wrong,  , being the percentage of
correct frame-ID couples in training set. We performed a grid search setting  and C to {0.9, 0.75,
0.5} and {0, 2.5, 5, 10}, respectively. During the training phase, other parameters are optimized for
improving the performances of the proposed methodology. In addition to those seen for IDCCAE
and IDC-LSTM-AE, we also optimized batch size ({64, 128, 256, 512}), number of epochs (in
the [50, 200] range) and learning rate ({10− 2, 10− 3, 10− 4, 10− 5}) as autoencoder-independent
hyperparameters using ADAM as optimizer.</p>
        <p>Finally, Mean Squared Error (MSE) has been chosen to evaluate reconstruction errors for all
models and all machine types.</p>
      </sec>
      <sec id="sec-3-5">
        <title>3.5. Evaluation and Performance Metrics</title>
        <p>In this section are described the metrics used to evaluate the performances of trained models.
The metrics used for models evaluation are the area under the receiver operating characteristic</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Model TrainingPumpInference</title>
      <p>IDCCAE 1h 26min 51s 3,19s
IDC-LSTM-AE 3min 49s 15,4s
(ROC) curve (AUC) and the partial-AUC (pAUC). The ROC curve shows the trend of the true
positive rate (TPR) in function of the false positive rate (FPR) at the variation of a parameter,
the pAUC is calculated as the AUC over a low FPR range [0, ], with  = 0.1.</p>
      <p>The anomaly score associated to a test sample is calculated taking the reconstruction errors
average over all frames extracted from it and after the application of normalization. The pAUC
is defined because it is especially important to increase the TPR under low FPR conditions, in
that if an ASD system gives false alerts frequently we cannot trust it.</p>
    </sec>
    <sec id="sec-5">
      <title>4. Results</title>
      <p>This section discusses about the results obtained by the IDCCAE and IDC-LSTM-AE models
with respect to diferent competitors in terms of eficiency and eficacy analysis.</p>
      <p>We evaluated the eficiency of the proposed models by investigating their training and
inference time, also considering their used memory. Specifically, we compare the two proposed
models based on LSTM and CNN respectively.</p>
      <p>Table 1 shows that the IDC-LSTM-AE achieves best results in terms of training although the
inference time is highest, also requires a large amount of memory (64 MiB). On the other hand,
IDCCAE requires a very high training time whilst achieving very good results in terms of both
inference time (on average 2.98s) and used memory (8 MiB).</p>
      <p>In turn, the eficacy analysis has been evaluated comparing the proposed models with respect
to diferent competitors in terms of AUC and pAUC.</p>
      <p>Table 2 shows results obtained by all models for each type of machinery. For pump machinery,
IDC-LSTM-AE achieves the best result in terms of AUC (78.29%) while IDCAE [15] shows an
increase of 0.66% in terms of pAUC w.r.t our approach. In turn, for fan machinery IDCAE
and IDCCAE are the best models in terms of AUC and pAUC respectively (77.45% and 70.33%).
In turn, for slider and valve IDCCAE achieves the highest results for pAUC metric 84.14%,
while CAE results the best model for AUC metric reporting an increase of 0.78% and 4.1% w.r.t.
IDCCAE.</p>
      <p>Finally, the results show that the proposed methodology, which leverages ID Conditioning,
Mel-Spectogram and novel loss function for improving the model’s performance, achieves the
best value in terms of AUC and pAUC for all machines considered.</p>
      <p>Model
Baseline
CAE [16]
LSTM [17]
IDCAE [15]
IDCCAE
IDC-LSTM-AE
Model
Baseline
CAE [16]
LSTM [17]
IDCAE [15]
IDCCAE
IDC-LSTM-AE
Std.Dev
0.70%
1.87%
2.21%
Std.Dev
0.53%
0.72%
2.29%</p>
      <p>Mean
52.45%
52.63%
52.05%
70.32%
70.33%
65.83%</p>
      <p>Valve
Std.Dev
0.49%
5.00%
2.99%</p>
    </sec>
    <sec id="sec-6">
      <title>5. Discussion and Conclusions</title>
      <p>Anomaly detection is a data-driven approach that employs predictive maintenance to minimize
downtime, reduce costs, and optimize maintenance procedures. Numerous anomaly detection
systems have been developed and studied in recent years. In this study, we propose a
methodology for anomaly detection in predictive maintenance that incorporates machine identifiers
into the autoencoder learning process. This method enables models to be trained on sounds
recorded in the proximity of diferent versions of the same machine type, resulting in improved
detection capabilities. We also employ algebraic operations to enhance the learning phase and
leverage the latent representation of the input generated by the encoder. Our experiments,
conducted on the DCASE 2020 Task 2 Challenge dataset, demonstrate that our approach improves
performance by up to 17.61% compared to non-conditioned autoencoder versions and baseline
models, particularly for pAUC.</p>
      <p>Our future work will focus on analyzing various types of conditioning networks, such as
Variational AutoEncoders (VAEs) or Generative Adversarial Networks (GANs), and applying
diferent pre-processing strategies to achieve better training performance. These strategies
include noise reduction to eliminate background noise commonly found in factory environments
and audio data augmentation techniques such as pitching and time-shifting. Additionally, we
plan to investigate two emerging research areas: Context Prediction and Context Histories,
which utilize time series of Contexts to record machine data in context histories for diferent
types of data analysis ([18, 19, 20]).
dataset: Sound dataset for malfunctioning industrial machine investigation and
inspection, arXiv preprint arXiv:1909.09347 (2019). doi:https://doi.org/10.48550/arXiv.
1909.09347.
[15] S. Kapka, Id-conditioned auto-encoder for unsupervised anomaly detection, arXiv preprint
arXiv:2007.05314 (2020). doi:https://doi.org/10.5281/zenodo.4061782.
[16] A. Ribeiro, L. M. Matos, P. J. Pereira, E. C. Nunes, A. L. Ferreira, P. Cortez, A. Pilastri, Deep
dense and convolutional autoencoders for unsupervised anomaly detection in machine
condition sounds, arXiv preprint arXiv:2006.10417 (2020). doi:https://doi.org/10.
48550/arXiv.2006.10417.
[17] A. Jalali, A. Schindler, B. Haslhofer, Dcase challenge 2020: Unsupervised anomalous sound
detection of machinery with deep autoencoders (2020).
[18] J. H. da Rosa, J. L. Barbosa, G. D. Ribeiro, Oracon: An adaptive model for context prediction,
Expert Systems with Applications 45 (2016) 56–70. doi:https://doi.org/10.1016/j.
eswa.2015.09.016.
[19] V. La Gatta, V. Moscato, M. Postiglione, G. Sperlì, Pastle: Pivot-aided space
transformation for local explanations, Pattern Recognition Letters 149 (2021) 67–74. URL:
https://www.sciencedirect.com/science/article/pii/S0167865521002014. doi:https://doi.
org/10.1016/j.patrec.2021.05.018.
[20] A. S. Filippetto, R. Lima, J. L. V. Barbosa, A risk prediction model for software project
management based on similarity analysis of context histories, Information and Software
Technology 131 (2021) 106497. doi:https://doi.org/10.1016/j.infsof.2020.106497.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>E.</given-names>
            <surname>Di Fiore</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ferraro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Galli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Moscato</surname>
          </string-name>
          ,
          <string-name>
            <surname>G. Sperlì,</surname>
          </string-name>
          <article-title>An anomalous sound detection methodology for predictive maintenance</article-title>
          ,
          <source>Expert Systems with Applications</source>
          <volume>209</volume>
          (
          <year>2022</year>
          )
          <fpage>118324</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>E. C.</given-names>
            <surname>Nunes</surname>
          </string-name>
          ,
          <article-title>Anomalous sound detection with machine learning: A systematic review</article-title>
          ,
          <source>arXiv preprint arXiv:2102.07820</source>
          (
          <year>2021</year>
          ). doi:https://doi.org/10.48550/arXiv. 2102.07820.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>P.</given-names>
            <surname>Calvo-Bascones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Sanz-Bobi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. M.</given-names>
            <surname>Welte</surname>
          </string-name>
          ,
          <article-title>Anomaly detection method based on the deep knowledge behind behavior patterns in industrial components. application to a hydropower plant</article-title>
          ,
          <source>Computers in Industry</source>
          <volume>125</volume>
          (
          <year>2021</year>
          )
          <article-title>103376</article-title>
          . doi:https://doi.org/10. 1016/j.compind.
          <year>2020</year>
          .
          <volume>103376</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>C.</given-names>
            <surname>Yin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. N.</given-names>
            <surname>Xiong</surname>
          </string-name>
          ,
          <article-title>Anomaly detection based on convolutional recurrent autoencoder for iot time series</article-title>
          ,
          <source>IEEE Transactions on Systems, Man, and Cybernetics: Systems</source>
          (
          <year>2020</year>
          ). doi:
          <volume>10</volume>
          .1109/TSMC.
          <year>2020</year>
          .
          <volume>2968516</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Thudumu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Branch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Jin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. J.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <article-title>A comprehensive survey of anomaly detection techniques for high dimensional big data</article-title>
          ,
          <source>Journal of Big Data</source>
          <volume>7</volume>
          (
          <year>2020</year>
          )
          <fpage>1</fpage>
          -
          <lpage>30</lpage>
          . doi:https: //doi.org/10.1186/s40537-020-00320-x.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Cinque</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. Della</given-names>
            <surname>Corte</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Moscato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sperlí</surname>
          </string-name>
          ,
          <article-title>A graph-based approach to detect unexplained sequences in a log</article-title>
          ,
          <source>Expert Systems with Applications</source>
          <volume>171</volume>
          (
          <year>2021</year>
          )
          <article-title>114556</article-title>
          . doi:https://doi.org/10.1016/j.eswa.
          <year>2020</year>
          .
          <volume>114556</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>L.</given-names>
            <surname>De Simone</surname>
          </string-name>
          , E. Caputo,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cinque</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Galli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Moscato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Russo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Cesaro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Criscuolo</surname>
          </string-name>
          , G. Giannini,
          <article-title>Lstm-based failure prediction for railway rolling stock equipment</article-title>
          ,
          <source>Expert Systems with Applications</source>
          <volume>222</volume>
          (
          <year>2023</year>
          )
          <fpage>119767</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>A. De Santo</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Ferraro</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Galli</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Moscato</surname>
          </string-name>
          , G. Sperlì,
          <article-title>Evaluating time series encoding techniques for predictive maintenance</article-title>
          ,
          <source>Expert Systems with Applications</source>
          <volume>210</volume>
          (
          <year>2022</year>
          )
          <fpage>118435</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vafeiadis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Votis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Giakoumis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Tzovaras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Hamzaoui</surname>
          </string-name>
          ,
          <article-title>Audio content analysis for unobtrusive event detection in smart homes</article-title>
          ,
          <source>Engineering Applications of Artificial Intelligence</source>
          <volume>89</volume>
          (
          <year>2020</year>
          )
          <article-title>103226</article-title>
          . doi: https://doi.org/10.1016/j.engappai.
          <year>2019</year>
          .
          <volume>08</volume>
          .020.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ferraro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Moscato</surname>
          </string-name>
          , G. Sperlì,
          <article-title>Deep learning-based community detection approach on multimedia social networks</article-title>
          ,
          <source>Applied Sciences</source>
          <volume>11</volume>
          (
          <year>2021</year>
          )
          <fpage>11447</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>M.</given-names>
            <surname>Meire</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Karsmakers</surname>
          </string-name>
          ,
          <article-title>Comparison of deep autoencoder architectures for real-time acoustic based anomaly detection in assets</article-title>
          ,
          <source>in: 2019 10th IEEE International Conference on Intelligent Data Acquisition and Advanced Computing Systems: Technology and Applications (IDAACS)</source>
          , volume
          <volume>2</volume>
          ,
          <year>2019</year>
          , pp.
          <fpage>786</fpage>
          -
          <lpage>790</lpage>
          . doi:
          <volume>10</volume>
          .1109/IDAACS.
          <year>2019</year>
          .
          <volume>8924301</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>B.</given-names>
            <surname>Bayram</surname>
          </string-name>
          , T. B.
          <string-name>
            <surname>Duman</surname>
          </string-name>
          , G. Ince,
          <article-title>Real time detection of acoustic anomalies in industrial processes using sequential autoencoders</article-title>
          ,
          <source>Expert Systems</source>
          <volume>38</volume>
          (
          <year>2021</year>
          ). doi:https://doi. org/10.1111/exsy.12564.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Koizumi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Murata</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Harada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Saito</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Uematsu</surname>
          </string-name>
          , Sniper:
          <article-title>Few-shot learning for anomaly detection to minimize false-negative rate with ensured true-positive rate (</article-title>
          <year>2019</year>
          )
          <fpage>915</fpage>
          -
          <lpage>919</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICASSP.
          <year>2019</year>
          .
          <volume>8683667</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>H.</given-names>
            <surname>Purohit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Tanabe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ichige</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Endo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Nikaido</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Suefusa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kawaguchi</surname>
          </string-name>
          , Mimii
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>