<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>AI trustworthy in multimodal and healthcare scenarios</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ermanno Cordelli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Valerio Guarrasi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giulio Iannello</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Filippo Rufini</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paolo Soda</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lorenzo Tronchin</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Unit of Computer Systems and Bioinformatics, Department of Engineering</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University Campus Bio-Medico of Rome</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <fpage>29</fpage>
      <lpage>31</lpage>
      <abstract>
        <p>The pervasiveness of artificial intelligence in our daily lives has raised the need to understand and trust the outputs of the learning models, especially when involved in decision processes. As a result, eXplainable Artificial Intelligence has captured more and more interest in the scientific community, providing insights into the behaviour of these systems, ensuring algorithms fairness, transparency and trustworthiness. In this contribution we overview our work on the explainability of deep learning models applied to time series, multimodal data and towards extracting meaningful medical concepts. eXplainable Artificial Intelligence, Multimodal Learning explanations, medical concepts, multivariate time series ∗Corresponding author.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>We take into account the challenge of explaining a real</title>
      <sec id="sec-1-1">
        <title>1. Introduction</title>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Artificial Intelligence (AI) has proven to efectively sup</title>
      <p>
        port the decision process [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and in particular deep
learning techniques have achieved state-of-the-art
performance [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Despite the impressive prediction accuracy
attained in several applications, there is still the need to
As a result, eXplainable Artificial Intelligence (XAI) has
captured more and more interest in the scientific
community [
        <xref ref-type="bibr" rid="ref3 ref4 ref5">3, 4, 5</xref>
        ] since the complex nature of the models,
such as deep neural networks, makes it impossible for
of explaining multivariate TS (MTS) and showing how
to adapt diferent methodologies, originally designed for
images, to this domain.
      </p>
    </sec>
    <sec id="sec-3">
      <title>The second concerns the explainability of Multimodal</title>
      <p>Deep Learning (MDL) models, a topic at its infancy in
the current literature. With the recent availability of a
larger data repository, we expect to have the possibility
how to learn shared representation and how to combine
the unimodal networks. Nevertheless, more complex
models exacerbate the problem of understanding what
the predictions rely on, and also which modalities and
explain the decisions of the learning models proposed. to explore more complex deep architectures, studying</p>
    </sec>
    <sec id="sec-4">
      <title>XAI aims to provide an insight into the behaviour and</title>
      <p>
        the user to understand and validate the decision process. features hold an important role [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
    </sec>
    <sec id="sec-5">
      <title>Finally, the third direction we take leads toward en</title>
      <p>
        processes of these systems, ensuring algorithms fairness, suring trustworthiness and reliability specifically in the
identifying any potential bias in the training data and
allowing complex AI models to be more transparent and
understandable to humans [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>Among a growing body of literature about XAI, in our
laboratory we are directing our eforts toward three
issues. The first concerns the explainability of Deep
Learndeed, the rising capabilities of storing and registering
ing (DL) models working on Time Series (TS) data. In- clinical context. However, in the medical field,
identifydata have increased the number of temporal datasets, defined as relevant on an abstract scale is much more
chalboosting the attention on TS classification models and
raising the need to explain their decision. In this context,
we present the application and evaluation of three XAI
tection on telematics data. We dealt with the challenge
methods in a real-world multimodal task of anomaly de- explanations that users can trust and rely on.
medical domain. Indeed, this is a field where XAI has
a major impact, allowing both AI designers and
medical experts to rely on meaningful explanations of the
black box inner workings and reasoning, so that the
decision of AI-based systems can be properly understood and
adequately considered when applied in the real-world
ing anatomical structures or tissue features that can be
lenging and these elements may not be unambiguously
defined. Therefore, it is essential to develop methods
that can bridge this gap and provide more human-like</p>
      <sec id="sec-5-1">
        <title>2. Translating XAI to Multivariate</title>
      </sec>
      <sec id="sec-5-2">
        <title>Time Series</title>
        <p>
          atics data from vehicles’ black-box, where the available 2.1. Methods
modalities are acceleration MTS and velocity univariate
TS (UTS). Moreover, in this application there is also a Figure 1 shows an overview of the proposed approach
supervised classifier trained to recognise if a crash event that is able to explain a multimodal anomaly detection
occurred or not. The peculiarity of MTS is that they are architecture identifying car crash events from
telematcharacterised by complex non-linear temporal dependen- ics data of vehicles. We first train the black-box
multicies between their attributes, i.e. the points of each UTS modal model for classification, which consists of a
Conare connected with the other sequences via the time di- volutional Neural Network (CNN) that can learn local
mension. This key issue makes the development of an temporal-spatial patterns that are intuitive to visualise
XAI approach for MTS-based anomaly detection rather for the end-user. Then, the trained CNN, the test samples
challenging. The analysis of the literature shows that and the performance on the test set are used as inputs
the study on XAI for MTS is limited: indeed, more ef- of the XAI framework to generate visual explanations of
forts have been directed towards data diverse from TS, the decision provided by the model and to evaluate the
i.e. images and tabular data, not taking into account the quality of such explanation.
complex relationship retained in multivariate time series. To consider both temporal and spatial relationships
Thus as first contribution, we studied how to employ between each dimension of a MTS, we represent each
three XAI methods suited to explain models working on sample as a 2D image where each pixel does not retain
images and how to extend their application to a multi- only visual features, such as the shape, the intensity and
modal architecture working on telematics MTS acquired the texture, but also temporal features across each UTS
by car’s black-boxes [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. included in the MTS. In this way we gain the
advan
        </p>
        <p>
          A further issue that we tackled in this work is evalu- tage of analysing the MTS as a whole, leveraging the
ating the provided explanation for the MTS. Indeed, as rich literature about the explainability on images, and of
highlighted in [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], diferent kinds of explanations pro- maintaining the flexibility to search for the best
architecduced after interpreting machine learning models may ture and performance without the constrain of designing
not be equally explainable, so a pressing need is emerging a model for explicitly achieving explainability.
for quantifying the quality of the explanations produced, Therefore, given the multimodal CNN initially
dea topic that is only in its infancy in the current litera- signed, we investigated its explainability by employing
ture [
          <xref ref-type="bibr" rid="ref8">8, 9</xref>
          ]. three well-known XAI algorithms originally designed
for images, providing a saliency map, which is an
efi
        </p>
        <p>Tr + Vl Te cient way of pointing out what causes a certain outcome.
MOD 1, MOD 2 MOD 1, MOD 2 The three approaches are: (i) an agnostic solution that is
generalisable by definition to any model and returns a
comprehensible local predictor, namely LIME [10]; (ii) a
CNN TrCaNinNed XXXAAAIII 231 IITTAVE ITANO emxopdleicl-istplyecfioficrsoclountivoonlu,tnioanmaellayrGchriatde-cCtuArMes[t1h1a]t, dcoesnisgindeedr
TAUNQ LAVEU the CNN activation on the specific input sample; (iii) a
model-inspection approach, namely IG [12], a method that
derives explanation by examining the internal model
beCRCARSAHS/HNO PERFORMANCES EXPVLAISNUAATLION haviour when modifying the input sample. Indeed, as</p>
        <p>
          LIME aims to approximate the decision surface of a
comTraining Testing EExxptlraancatitoionn EAxspsleasnmateionnt fprloemxmthoedesluurrsoinggataenminotdeerlpsrectaanbnleotobnee,ptehrefeecxtpl ylafnaaittihofnusl
with respect to the original model [13]. For this reason we
Figure 1: An overview of our approach. The classifier (CNN) also customized an XAI method suited for CNNs and
anis first trained using the training (Tr) and validation (Vl) other more general suited for deep architectures, namely
sets of the dataset encompassing acceleration (MOD 1) and Grad-CAM and IG, respectively. We therefore present
speed (MOD 2) signals of the vehicle black box. Then the how to customise and employ them to deal with a
multrained model (Trained CNN) and its predictions on the test timodal architecture working on multivariate telematic
set (Te) are employed for extracting the explanations with the data [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
three diferent XAI methods. Regarding the evaluation procedure, there is no
consensus in the literature on methods to assess
explainability [14]. In this respect, we customised the two strategies
presented in [9] for UTS, by adapting them to the MTS
domain. We exploited a perturbation-based XAI
evaluation, measuring the performance drop Δ of the anomaly
detection system respectively disrupting the time points
identified by the XAI method as relevant ( Δ  ), and
random regions of the signal (Δ ). The assessment
is based on the assumption that if relevant/random
features (time points) get changed, the model’s performance
should decrease/stagnate. Finally, as we aimed to conduct
an exhaustive evaluation, we performed the assessment
procedure considering all test set samples in an hold-out
setup.
This work represents a first attempt to tailor XAI
algorithms to the multimodal nature of the data, suggesting
further research in this field. As a first direction, we will
investigate XAI methods able to provide a more
humaninterpretable representation, since the saliency maps
provided by all the XAI methods are hard to be interpreted
by users. The second direction will focus on developing
a multimodal XAI method able to explain both signals
available in the telematics data at hand (i.e. acceleration
and speed).
        </p>
        <sec id="sec-5-2-1">
          <title>2.4. Papers and available resources</title>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>This work was published in [7].</title>
      <sec id="sec-6-1">
        <title>3. Multimodal XAI</title>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Multimodal Deep Learning (MDL) studies how deep neu</title>
      <p>ral networks can learn shared representations between
2.2. Main results diferent modalities that, used together, may ofer deeper
insights into the data. It investigates when to fuse the
difOverall, the explanations obtained reasonably find the ferent modalities and how to obtain more powerful data
crucial features used from the model to perform the representations. Linking information coming from
varianomaly detection task, casting light on the black-box’s ous modalities can leverage the understanding of a field
decision process. IG and Grad-CAM are able to ex- of study. These considerations are of particular relevance
ploit the cross-correlation in UTS learned from the CNN, for the field of bio-medicine, where MDL has proven to
namely defining which UTS is most valuable for the pre- be useful [15].
diction. Instead, LIME explanation does not present this Among the diferent challenges of MDL, we focus here
evidence since it uses a surrogate model to approximate on supervised multimodal fusion applied to early identify
the CNN, being an agnostic method, so it does not directly patients at risk of the severe outcome, like intensive care
inspect the convolutional architecture’s inner workings. or death, among those afected by SARS-CoV-2, and using
To provide an exhaustive comparison between the three chest X-ray (CXR) scans and clinical data. Instead of
usmethods, Table 1 shows the results emerging from XAI ing manually designed or handcrafted modality-specific
quantitative evaluation. From the results, IG-based per- features, via DL we can automatically learn and extract
turbations account for the most significant drop in perfor- an embedded representation for each modality, which
mance and it is the only method that always exceeds the is then fused with the others in an end-to-end training
random drop. Hence, the quantitative analysis suggests process that exploits the loss backpropagation.
that IG is the more informative explainability method It is well-known that the major disadvantage of neural
for the anomaly detection task as it is able to detect the networks is their lack of interpretability. In spite of the
time points valuable for the CNN to perform the predic- importance of MDL, to the best of our knowledge, no
tion. In contrast, Grad-CAM results in the least reliable work has studied the XAI methods for fused modalities,
algorithm from the quantitative evaluation as two out of particularly in the medical field. Hence, we developed a
three times it is overtaken by random drop. Moreover, deep architecture, explainable by design, which jointly
we find the temporal trend to have a minor influence for learns modality reconstructions and sample
classificathe CNN to perform the classification task as we observe tions of the aforementioned biomedical multimodal data,
from the drops in the values of swap and mean metrics i.e., imaging and tabular data.
compared to the zero one.</p>
      <p>In general, this study provides insight into the quality
of explanation and sheds light on the most significant 3.1. Methods
features that are exploited by the CNN when it performs
the crash detection task.</p>
    </sec>
    <sec id="sec-8">
      <title>The proposed architecture is shown in Figure 2. The explanation of the decision taken is computed by applying a latent shift that simulates a counterfactual prediction revealing the features of each modality that contribute the</title>
      <sec id="sec-8-1">
        <title>3.3. Challenges and perspectives</title>
        <p>
          3.4. Papers and available resources
most to the decision and a quantitative score indicating
the modality’s importance. The latent shift is applied on This work is available in [16].
an embedded representation of the data, ℎ in Figure 2,
created by exploiting a Convolutional Autoencoder (CAE)
for the imaging modality and an Autoencoder (AE) for 4. Towards eXplainable Medical
the tabular modality, which is connected in an end-to-end Concepts
manner to a multi-layer perceptron which performs the
classification task for the prognosis of the severity of the Despite the fact that XAI has revolutionized the way we
COVID-19 virus. Exploiting the nature of the model’s ar- see machine learning models, there are still several
limitachitecture the variations on the embedded feature vector tions with respect to providing explanations closer to
huhelp us understand how the classifications are performed man perceptiveness, resulting in users not being able to
by the MLP, via the echoed perturbations on the embed- fully trust the model working [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. In this context, concept
ded vector and on the CAE’s and AE’s reconstructions. attribution methods have emerged as a new paradigm,
The explanation of the decision taken is computed by providing interpretations of the inner mechanisms of DL
applying a latent shift that, simulates a counterfactual methods by measuring the relevance of human-friendly
prediction revealing the features of each modality that concepts directly defined by the user [ 17]. In Computer
contribute the most to the decision and a quantitative Vision (CV) field the definition of semantic concepts is
score indicating the modality’s importance. simple and intuitive: it is based on key aspects contained
in images that users can easily relate to real-life contexts
3.2. Explanation assessment (e.g. ear shapes that are much more similar to a dog than
a cat). However, in the medical field, identifying
anatomTo study the validity of the proposed method, we con- ical structures or tissue features that can be defined as
ducted a reader study with four radiologists assessing relevant on an abstract scale is much more challenging
the prognosis of a subset of patients. Each radiologist and it may not be easy to define them unambiguously.
observed both data modalities simultaneously for each Therefore we are tackling this open issue focusing on
patient and performed the prognosis task. Afterward, the exploring unsupervised approaches to automatically
exradiologists attributed an importance score, on a scale tract meaningful medical semantic concepts. Specifically,
from 1 to 5, indicating how significant each modality was we are investigating the efectiveness of deep clustering
for the prognosis task. Then, to understand the most methods that perform representation learning through
important features for each modality, we asked the ra- Convolutional Autoencoders and clustering
simultanediologist to select the clinical variables and to segment ously, operating on the latent space learned from the
the areas of interest in the X-ray image, most useful to raw data distribution [18]. This approach has the
adstratify the patient. The sanity check, although very vantage of producing features that are highly correlated
time-consuming, was very useful since it showed a high with the image structures, providing semantically
meanintersection between the explanations provided by the ingful groups of images in the dataset. The key idea is
method and those of the radiologists, both for the modal- to gain new insights about the data, studying a novel
ity and the feature importance. By reducing the model’s approach able to extract from the deep latent space of the
opacity which makes it dificult for doctors and regula- Autoencoder the set of features with the highest possible
tors to trust the MDL models, we are able to improve conceptual expressiveness.
trust and transparency.
        </p>
      </sec>
      <sec id="sec-8-2">
        <title>4.1. Methods</title>
        <p>In the first stage, concept extraction is performed using
the Deep Clustering algorithm. This algorithm constructs
a low-level representation of the dataset (called latent
space H) and clusters the samples into  clusters in the
same H-space. To obtain the best spatial representation of
the hidden concepts within the dataset, we implemented
an iterative process, involving training the algorithm for
several initialization of the number of clusters into which
to separate the samples of the dataset.</p>
        <p>The overall goal is to obtain a concept extraction
process based on images structural features (patterns of
pixels) common to the largest possible number of patients
available in the dataset, so achieving a clustering
configuration that most generalizes over the patients distribution.</p>
        <p>Hence, to determine the optimal value of  we employed Acknowledgments
two sets of metrics: (i) cluster-based metrics, which
validate the clustering conditions (Silhouette-score, Davies- The work detailed in section 2 was realized in cooperation
Bouldin, Calinski-Harabasz); (ii) patient-based metrics, with Consorzio ELIS (Consel) and Generali Italia S.p.A.
which are custom metrics designed to capture how the
data splits up according to each patient’s distribution.</p>
        <p>Patient-based metrics take into account the global varia- References
tions between each diferent clustering configuration and
are used to identify the ideal conditions for semantic
clustering of patients. The optimal value of  is, then, selected
based on the best results obtained for these metrics.</p>
        <p>DL models. The ultimate goal is to develop an automated
system for the unbiased evaluation of images, using XAI
tools to explain why images were grouped together. This
approach represents a new perspective for the analysis
of biomedical data sets and it has the potential to
significantly advance the field. Using unsupervised DL models,
it is possible to discover patterns within the data that
may not be evident to human analysts. Furthermore, by
using XAI tools to explain clustering results, the system
can provide insights into underlying biology and
pathology that are not readily available through traditional
image analysis methods. Overall, this work represents
an exciting development that has the potential to have a
significant impact on basic and clinical research.</p>
      </sec>
      <sec id="sec-8-3">
        <title>4.2. Validation strategy</title>
        <p>In order to validate the efectiveness of the proposed
concept extraction model we will adopt a comparative
strategy applied on a binary task for overall survival
prediction on a cohort of 191 patients with Non-small cell
lung cancer, including both CT volumetric images (total
of 22384 slices) and clinical data. First, we will build a
baseline for comparison: we will train a set of classical
machine learning classifiers on the clinical data
computing their performance for the desired task. Second, we
will train the proposed deep clustering model on the set
of CT scans, extracting a latent vector for each patient.
Third, this vector will be used in early fusion with the
clinical data to train the same set of classifiers used as
baseline. Comparing the performance with and without
the deep clustering concept vectors will allow to check
the efectiveness of the approach, testing its ability to
extract meaningful information.</p>
      </sec>
      <sec id="sec-8-4">
        <title>4.3. Challenges and perspectives</title>
        <p>In the analysis of biomedical datasets, this work shifts the
focus from dataset construction towards maximizing data
information content. This is accomplished by identifying
patterns of pixel structures within the images that define
similarity at the semantic level, employing unsupervised
[9] U. Schlegel, H. Arnout, M. El-Assady, D. Oelke, D. A. ing models for high stakes decisions and use
interKeim, Towards a rigorous evaluation of xai meth- pretable models instead, Nature Machine
Intelliods on time series, in: 2019 IEEE/CVF Interna- gence 1 (2019) 206–215.
tional Conference on Computer Vision Workshop [14] C. Molnar, Interpretable machine learning, Lulu.
(ICCVW), IEEE, 2019, pp. 4197–4201. com, 2020.
[10] M. T. Ribeiro, S. Singh, C. Guestrin, ”Why Should I [15] D. Ramachandram, G. W. Taylor, Deep multimodal
Trust You?” Explaining the Predictions of Any Clas- learning: A survey on recent advances and trends,
sifier, in: Proceedings of the 22nd ACM SIGKDD IEEE signal processing magazine 34 (2017) 96–108.
international conference on knowledge discovery [16] V. Guarrasi, L. Tronchin, D. Albano, E. Faiella,
and data mining, 2016, pp. 1135–1144. D. Fazzini, D. Santucci, P. Soda, Multimodal
ex[11] R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, plainability via latent shift applied to covid-19
stratD. Parikh, D. Batra, Grad-CAM: Visual Explana- ification, arXiv preprint arXiv:2212.14084 (2022).
tions from Deep Networks via Gradient-based Lo- [17] M. Graziani, V. Andrearczyk, S. Marchand-Maillet,
calization, in: Proceedings of the IEEE international H. Müller, Concept attribution: Explaining cnn
conference on computer vision, 2017, pp. 618–626. decisions to physicians, Computers in biology and
[12] M. Sundararajan, A. Taly, Q. Yan, Axiomatic at- medicine 123 (2020) 103865.
tribution for deep networks, in: International [18] E. Aljalbout, V. Golkov, Y. Siddiqui, M. Strobel,
Conference on Machine Learning, PMLR, 2017, pp. D. Cremers, Clustering with deep learning:
Taxon3319–3328. omy and new methods, arXiv:1801.07648 (2018).
[13] C. Rudin, Stop explaining black box machine
learn</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>T</surname>
          </string-name>
          .-c. Fu,
          <article-title>A review on time series data mining</article-title>
          ,
          <source>Engineering Applications of Artificial Intelligence</source>
          <volume>24</volume>
          (
          <year>2011</year>
          )
          <fpage>164</fpage>
          -
          <lpage>181</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>H. I.</given-names>
            <surname>Fawaz</surname>
          </string-name>
          , G. Forestier,
          <string-name>
            <given-names>J.</given-names>
            <surname>Weber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Idoumghar</surname>
          </string-name>
          , P.-A. Muller,
          <article-title>Deep learning for time series classification: a review, Data mining and knowledge discovery 33 (</article-title>
          <year>2019</year>
          )
          <fpage>917</fpage>
          -
          <lpage>963</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>R.</given-names>
            <surname>Guidotti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Monreale</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ruggieri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Turini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Giannotti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Pedreschi</surname>
          </string-name>
          ,
          <article-title>A survey of Methods for Explaining Black Box Models, ACM computing surveys (CSUR) 51 (</article-title>
          <year>2018</year>
          )
          <fpage>1</fpage>
          -
          <lpage>42</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <article-title>Techniques for Interpretable Machine Learning</article-title>
          ,
          <source>Communications of the ACM</source>
          <volume>63</volume>
          (
          <year>2019</year>
          )
          <fpage>68</fpage>
          -
          <lpage>77</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>L. H.</given-names>
            <surname>Gilpin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Bau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. Z.</given-names>
            <surname>Yuan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bajwa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Specter</surname>
          </string-name>
          , L. Kagal, Explaining Explanations:
          <article-title>An Overview of Interpretability of Machine Learning</article-title>
          ,
          <source>in: 2018 IEEE 5th International Conference on data science and advanced analytics (DSAA)</source>
          , IEEE,
          <year>2018</year>
          , pp.
          <fpage>80</fpage>
          -
          <lpage>89</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>G.</given-names>
            <surname>Joshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Walambe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kotecha</surname>
          </string-name>
          ,
          <article-title>A review on explainability in multimodal deep neural nets</article-title>
          , IEEE Access (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>L.</given-names>
            <surname>Tronchin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Sicilia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Cordelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. R.</given-names>
            <surname>Celsi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Maccagnola</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Natale</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Soda</surname>
          </string-name>
          ,
          <article-title>Explainable ai for car crash detection using multivariate time series</article-title>
          ,
          <source>in: 2021 IEEE 20th International Conference on Cognitive Informatics</source>
          &amp;
          <article-title>Cognitive Computing (ICCI* CC)</article-title>
          , IEEE,
          <year>2021</year>
          , pp.
          <fpage>30</fpage>
          -
          <lpage>38</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>F.</given-names>
            <surname>Doshi-Velez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <article-title>Towards a rigorous science of interpretable machine learning</article-title>
          ,
          <source>arXiv preprint arXiv:1702.08608</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>