<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Journal of Hydrology 495 (2013) 38-51. doi:10.1016/j.jhydrol.2013.04.041.
[9] A. Wunsch</journal-title>
      </journal-title-group>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/hess-26-2405-2022</article-id>
      <title-group>
        <article-title>Application for Spatial, Temporal, and Spatio-Temporal Explanations</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Matteo Salis</string-name>
          <email>matteo.salis@unito.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gabriele Sartor</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marco Pellegrino</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stefano Ferraris</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Abdourrahmane M. Atto</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rosa Meo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pier Andrea Mattioli</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Torino Italy</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Workshop</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Computer Science Department - University of Turin</institution>
          ,
          <addr-line>Corso Svizzera 185, Torino</addr-line>
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Interuniversity Department of Regional and Urban Studies and Planning - Politecnico di Torino and University of Turin</institution>
          ,
          <addr-line>Viale</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>LISTIC Laboratory - Université Savoie Mont Blanc</institution>
          ,
          <addr-line>5 chemin de bellevue, 74 940 Annecy-le-vieux</addr-line>
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <volume>11700</volume>
      <fpage>38</fpage>
      <lpage>51</lpage>
      <abstract>
        <p>Over the past few years, the efects of climate change have significantly increased, directing our attention to our environmental resources. One of the most critical resources is fresh water, whose availability has been endangered by even more frequent extreme events of precipitations and droughts. Deep Learning (DL) has revealed a useful tool for accurate hydrological forecasting, but its main drawback is its black-box nature. To face this issue eXplainable Artificial Intelligence (XAI) has come out. In this work, Randomized Input Sampling for Explanation (RISE) is applied to a regression spatio-temporal model trained to predict the Water Table Depth (WTD) collected by a sensor in Vottignasco, in the northwest of Italy. An investigation of the model behaviour over spatial, temporal and spatio-temporal dimensions has been conducted formalizing S-RISE, T-RISE, and ST-RISE respectively. Results suggest the usefulness of a spatio-temporal approach (ST-RISE) and give interesting intuitions on the events that occurred over the area.</p>
      </abstract>
      <kwd-group>
        <kwd>Explainable AI</kwd>
        <kwd>Model agnostic algorithms</kwd>
        <kwd>RISE</kwd>
        <kwd>Spatio-temporal explanations</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Water is essential to all species. Groundwater resources represent one of the most relevant sources
of freshwater for human beings. Changes in rainfall patterns and temperature due to climate change
have made groundwater even more pivotal in providing fresh and clean water to communities [
        <xref ref-type="bibr" rid="ref2 ref3 ref4">2, 3, 4</xref>
        ].
Groundwater resources could be quantified by measuring the water table depth (WTD), i.e. the distance
between the ground surface and the higher surface of the groundwater body. Numerical models have
been successfully proposed to model the WTD, however, they usually depend heavily on the geophysical
properties of the area to be modelled. This requires extensive measurements and computationally
intensive simulations. Deep Learning (DL) models have been adopted to overcome this issue and
develop models which depend only on open and easy-to-measure weather data (e.g. rainfall and
temperature). More specifically, spatio-temporal DL models based on computer vision architectures
have been proposed to handle dynamical relations between weather and the hydrogeological output
[5, 6, 7]. For this model type the input is often a weather video, in which each frame  contains a weather
map. Each weather map comprises multiple channels  , in other words, the observed weather variables
(e.g. rainfall, temperature etc.). The weather video is thus fed into a DL model to forecast a specific
hydrological variable (e.g. the WTD, see Figure 1).
https://sites.google.com/view/matteo-salis/home (M. Salis)
      </p>
      <p>CEUR</p>
      <p>ceur-ws.org</p>
      <p>DL models achieved overwhelming performances [8, 9, 10, 7]; however, their black box nature is a
matter of fierce debate in the scientific community and prevents their adoption as a reliable tool to
support decision and policymaking. An entire research field named explainable Artificial Intelligence
(XAI) has come out with the aim of explaining DL models behaviour. There exist several approaches
which depend on the generalization with respect to DL architecture (model specific or model agnostic)
and to data (global explanation or local explanation)[11, 12, 13]. More specifically, an algorithm is said to
be model agnostic if it could be applied regardless of the specific type of DL architecture. Concerning data
generalization, an explanation is defined local if it is specific to a dataset instance, reversely it is named
general when it is related to the overall behaviour of the model over the whole dataset. Many strategies
have been proposed to produce explanations, some of them, like gradient-based and propagation-based
techniques, require access to some or all hidden layers. Reversely, perturbation-based techniques,
among which the local method RISE (Randomized Input Sampling for Explanation) [14], need only a
trained model by which to perform predictions. RISE was proposed for image classification explanations,
and it consists of masking  times randomly the input features of the instance to be explained. Then,
a saliency map is created highlighting pixels that contribute the most to the class probability score
maximization. This approach is rather simple and suitable when details about the black box model are
not reported and one cannot access hidden activations.</p>
      <p>The main objective of this study is to produce local explanations for a specific CNN-LSTM black
box model proposed in [7] which forecasts the weekly WTD given a weekly weather video of the
previous two years of total precipitation, minimum temperature, and maximum temperature. Given the
multidimensionality (time and space) of the input features applying a XAI algorithm is not a trivial
task, especially because many of them have been originally developed for standard classification tasks
on well-known public dataset [15, 14, 16, 17]. More specifically, in the presence of spatio-temporal data
(e.g. weather video), it is relevant to produce not only a spatial saliency map, but also a temporal, and
even more importantly a spatio-temporal saliency object for each input feature (i.e. weather variable
or video’s channel). To this aim, we have adopted the perturbation-based and model-agnostic RISE
algorithm, and we have proposed and formalized its application to produce a) a spatial explanation in
the form of a saliency 2D map (very similar to the RISE original version) which shows the most relevant
pixels; b) a temporal explanation in the form of a saliency 1D vector which highlights the most relevant
frames in the input video; and c) a spatio-temporal explanation in the form of a saliency 3D video
which highlights the most relevant pixels for each frame. Given the preliminary nature of the present
study, we have just focused on total precipitations, which is also the most interesting weather variable
for domain experts. The choice of RISE instead of other XAI methods is because of its simplicity and
because the extension to multidimensional data (e.g. video) has been already reported as a future work
by authors in the original paper [14]. Furthermore, RISE has been already applied in hydrological DL
studies [18, 5] but without making any extension and formalization to temporal and spatio-temporal
explanations.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related works</title>
      <p>Some XAI algorithms have already been developed for video-related tasks. Authors in [19] proposed
saliency Tubes for highlighting the most relevant areas in each frame of a video for a classification task.
However, this model-specific method has been developed for 3D CNN and requires access to the last
convolutional layer activations. [20] proposed a new, but still perturbation-based, method named STEP
(Spatio-Temporal Extremal Perturbation) which finds the salient elements using 3D masks generated
via a 2-step optimization procedure. In more detail, STEP finds the smallest subset of input elements
for retaining the highest prediction accuracy on a specified target label. The downside of this more
sophisticated approach is that it introduces new hyperparameters, another loss function to be chosen,
and a new heavy optimization problem to be solved. Authors in [21] criticized the adoption of a 3D
mask in case of complex action recognition tasks. Thus, they proposed to generate masks based on
optical flows which estimate apparent object motions inside a video. However, these optical flow need
to be in turn estimated with other algorithms.</p>
      <p>It is relevant to stress that our explanation task should be simpler than video classification tasks in
which high-resolution video contains complex actions to be recognized. Indeed, in the case study of [7]
the images are about 12.5km in resolution and the task is to forecast a scalar, not recognize and classify
a complex motion. For these reasons, even if STEP and Adaptive Occlusion could be very suited for
some more complex case studies, we decided to adopt the simpler and model-agnostic RISE algorithm.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Methods</title>
      <p>As already stated, RISE has been proposed to retrieve pixel relevance in a classification task, detecting
the pixels in the input images that maximize the class probability scores. The result of this method is a
spatial explanation, i.e. a saliency map over the input’s pixels, in which the higher the value, the higher
the pixel’s influence is on the model’s prediction. In the original paper [ 14] input perturbations were
performed by applying black blurred patches over the unperturbed input. Our case study has required
some modifications. The first major modification is because we are dealing with a regression task. Thus,
inspired by [18, 12, 5], we have defined a saliency object for the  -th instance after  perturbations as in
Equation 1:
 () =</p>
      <p>1
 () ⋅  =1

∑ Δ

() ⋅   () ,
prediction  ̂ 
perturbed with the  -th perturbation mask.
where  () is the masks’ mean. In other words, it is a weighted sum of random masks, where the weights
Δ (Equation 2) are the absolute diference between the unperturbed prediction  ̂  and the perturbed

 obtained respectively by feeding the model with the original input  () and the input
Δ

() = | ̂ () −  ̂ () |,</p>
      <p>The second major modification is related to the physical meaning of data contained in the input
weather video. Indeed, whilst in the case of classical RGB image blurring with zeros means inserting
black smooth patches (i.e. no information), in the case of physical raster data substitute values with
zeros has a strong physical meaning. To this aim, inspired by [22, 23, 18, 5], we decide to perturb input
(1)
(2)
images with additive Gaussian noise. In this way, when a positive noise is generated the precipitations
are increased, reversely if the negative noise is generated the precipitations are decreased.</p>
      <p>Given the multidimensional nature of the input, perturbation masks with diferent dimensions could
be created. This will bring saliency objects with diferent dimensions and meanings. Following this
idea and to accomplish our aim to define a spatial, a temporal, and a spatio-temporal explanation, we
have formalized three diferent RISE applications varying the dimensionality of the Gaussian noise that
defines the perturbations:
1. S-RISE: In this case, noise is in the form of 2D perturbation mask. For every iteration  a 2D mask
  is applied to each frame of the input video. The result is a 2D saliency map, similar to the
original RISE, which highlights the areas that are more influential on the model’s  -th prediction.
2. T-RISE: In this case, noise is in the form of 1D perturbation mask. In every iteration  a
perturbation vector   is a univariate Gaussian noise, and it is applied to each pixel time series (i.e. the
series of values of a pixel over the whole video). The resulting saliency object is a vector which
determines the most relevant frame in the input video for the  -th prediction.
3. ST-RISE: In this last formalization, noise is in the form of a multivariate 3D Gaussian perturbation
mask, i.e. every mask   is a video of the same spatial and temporal extent as the input one.
Then, the noise is applied element-wise to the unperturbed input. Thus, the result is a saliency
video which highlights the most relevant areas and frames for the  -th prediction.</p>
      <p>In the following subsections, each approach is formalized.
3.1. S-RISE
In our work, to predict a groundwater level  () ∈ ℝ we consider a set of  videos  = { () | ∈ [1;  ]} .
A video is a tensor  () with element  ℎ(,),, ∈ ℝ representing the pixel at height ℎ ∈ [1, ...,  ] , width
 ∈ [1, ...,  ] , channel  ∈ [1, ..., ] and number of frame  ∈ [1, ...,  ] . Then, the model used in this
work is defined as  ∶  →  , which returns the  -th weekly WTD defined as  ̂ () , taking as input a
video  () , representing the weather video of the  previous weeks before the target dates.</p>
      <p>As previously said, S-RISE consists in calculating Δ() for each input sample  () , for  iterations. In this
case,  ̂ () is the prediction  ( () +   ) produced by processing  () perturbed by the two-dimensional
Gaussian noise   . Specifically, in the spatial case, the two-dimensional noise of the iteration  for the
instance  is defined as:
  =  ⋅ 
− 12 [ (ℎ− ℎ2 )2 + (−  2 )2 ]
where  ∈ {−1, 1} is a random parameter to make positive or negative the perturbations, (ℎ ,   ) the
coordinates of the Gaussian noise centre which are randomly sampled from a uniform distribution,
and  2 defines the extension of the Gaussian noise in the space. Formally, in the spatial case, the
perturbation of the input over a specific channel  can be defined as</p>
      <p>() +   =  ,() +   ,
∀ ∈ [1, ...,  ] and a specific channel  ∈ [1, ..., ] , conceptually meaning adding the 2D noise map   to
each frame  on the channel  .</p>
      <p>Finally, the 2D saliency map for the  -th instance is obtained from Equation 1 after J iterations (see
pipeline on the top of Figure 2).
In the T-RISE an input video  () is perturbed over the temporal dimension, i.e. along the frames. For
this purpose, a one-dimensional perturbation mask   is generated at each iteration  as a univariate
Gaussian noise defined as follows:
12  − 2 ],
  =  ⋅  − [
(4)
(5)
with the centre of the disturbance in the frame   that is randomly sampled from a uniform distribution.
In the temporal case, the perturbation of the input over a specific channel  can be defined as:
 () +   =  ℎ(,), +   ,
(6)
∀ℎ ∈ [1, ...,  ] , ∀ ∈ [1, ...,  ] and a specific channel  ∈ [1, ..., ] , producing a new video applying the
noise vector over all the pixel time series  ℎ(,), in the same way. In this case,  () captures the influence
of each frame (i.e. time step) on the prediction and, consequently, it is represented as a saliency vector
of dimension  (see pipeline at the bottom of Figure 3).
In the last case, to assess the joint influence of spatial and temporal dimensions, RISE is applied using a
three-dimensional Gaussian noise defined as:
where  are the coordinates in space and time of the pixel which is the centre of the Gaussian noise at
iteration  and are randomly sampled from uniform distributions. Finally, the perturbation of the input
in the spatio-temporal approach can be seen as
  =  ⋅ exp{− 1 (x − ) TΣ−1(x − )},</p>
      <p>2
 () +   =   (′) +   ,
(7)
(8)
representing the addition of the noise video   to the complete weather video for a specific channel
 ∈ [1, ..., ] . Consequently, the saliency video  () identifies the most influential pixels’ value considering
jointly time and space (see pipeline in Figure 4).</p>
    </sec>
    <sec id="sec-4">
      <title>4. Results and discussions</title>
      <p>4.1. Data
In [7] authors defined their Region of Interest (ROI) as the Grana-Maira catchment in Piemonte, an
Italian administrative region (Figure 5). Weather videos were set up using total precipitations (both
snow and rain), minimum temperature and maximum temperature, i.e.  = 3 channels. In the original
paper, WTD time series data were retrieved for diferent sensors located in the same catchment, and for
each sensor, a local model was trained. However, because of our interest in method formalization, at
this stage of the research, we have decided to focus on just one sensor in the municipality of Vottignasco.
Authors made available their data already pre-processed, standardized by computing z-scores and split
into training, validation, and test. We decided to focus only on the standardized test instances, looking
at the model behaviour in a scenario mostly independent from training data.</p>
      <sec id="sec-4-1">
        <title>4.2. Implementation details</title>
        <p>We have experimented with the three proposed approaches S-RISE, T-RISE, and ST-RISE for each one of
the 104 instances present in the test set of the Vottignasco WTD series. For each approach,  has been
sampled from a discrete uniform distribution with values 1 and -1. For S-RISE  has been set to 0.75
that, given the resolution of weather data, corresponds approximately to 10km. Instead for T-RISE  has
been set to 8, which corresponds to two months, thus roughly 95% of the noise is concentrated in four
months. For both S-RISE and T-RISE the total number of iterations  to produce a saliency object has
been set to 1000. This was found, inspired by [5], by looking for the minimum number of iterations that
makes the saliency object vary less than the 2% from further perturbations. For ST-RISE the covariance
matrix Σ has been constructed by setting covariances to 0 and using the individual standard deviations
of the previous cases (0.75 for spatial dimensions and 8 for the temporal one). The total number of
iterations  in the case of ST-RISE has been increased to 5000. All experimentation has been conducted
on Google Colab using the free available resources1.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.3. Evaluation</title>
        <p>The application of RISE in its three variations is assessed through insertion/deletion [14] metrics and
visual inspection of the saliency representations validated by domain experts. In the first case, using
some metrics we have a numerical and more objective evaluation of the diferent approaches while,
in the latter, saliency representations can give a more intuitive explanation of the model’s behaviour
regarding the dimension on which RISE focuses on. More in detail, the insertion consists of iteratively
calculating for each instance  the error between the original model’s prediction  ̂ () and the one
“enabling” only a percentage  of most important pixels of the input, according to the obtained saliency</p>
        <sec id="sec-4-2-1">
          <title>1The implementation of the experiments is available at this link.</title>
          <p>representations, leaving all other pixels to 0. Reversely, the deletion starts from the original input and
iteratively deletes information zeroing out the  percent of the most relevant input pixels. Therefore,
Insertion and deletion metrics present complementary assessments.
4.3.1. Insertion &amp; Deletion
Given our regression task, both insertion and deletion are computed by looking at the mean squared
error between the original unperturbed prediction and the one obtained by feeding the inserted/deleted
perturbed input to the model.</p>
          <p>It is relevant to emphasize that S-RISE, T-RISE, and ST-RISE produce saliency objects of diferent
dimensions. The saliency video of ST-RISE is the most precise, and it allows sorting each element (tuple
of pixel and corresponding frames) of the input video by its relevance to the prediction. Diferently,
S-RISE and T-RISE could produce rankings only in space and time respectively. This means that for
each deletion or insertion step for S-RISE  elements (i.e., a pixel time series) are zero-out or enabled,
while for T-RISE  ∗  elements are set to 0 or enabled. The saliency video of ST-RISE is the only one
that allows the computation of insertion and deletion metrics for each single element of the input video.</p>
          <p>As already discussed in Section 3, another specificity of our case study is that setting pixels to 0
has a physical meaning. Given that our dataset is standardized and feature means’ are then   = 0
∀ ∈ [1, ..., ] where  = 3 , in the case of deletion, we are eliminating relevant information by replacing
the mean; while insertion starts with a video made of a constant value (i.e. 0, the feature’s mean) and at
each iteration information is restored.</p>
          <p>On the left of Figure 6 it is depicted the mean of the insertion curves computed on all the 104 instances
of the test set. On the x-axis is reported the fraction of input elements (i.e., tuple of pixel and frame)
enabled. The idea behind the insertion metric is to evaluate how much the most influential pixels
influence the prediction without nearby information. The sharpest is the drop in the metric, the higher
is the relevance of that portion of input [14]. Consequently, a lower AUC (Area Under the Curve) means
a better explanation which isolates in a more precise way the most relevant pixels. ST-RISE has revealed
the best achieving an AUC of 0.136. This means that ST-RISE has detected more precisely the most
relevant portions of the input on average on all the test instances. Of course, this result is thanks to the
higher definition of the saliency video with respect to the saliency map and saliency vector. Furthermore,
ST-RISE is in principle able to capture jointly spatial and temporal (i.e. spatio-temporal) relations, while
S-RISE and T-RISE are only able to look for explanations in spatial and temporal dimensions separately.</p>
          <p>Figure 6 depicts on the right the mean deletion curve computed averaging the single deletion curve
over the 104 test instances. It represents the increase in the error, deleting a percentage  of the most
influential pixels from the original output. In contrast to the insertion, the deletion measures the
relevance of a limited number of pixels keeping the others unchanged (i.e. maintaining contextual
information). The higher the increase in the curve, the better the explanation is. In this case, ST-RISE
confirms its higher efectiveness and S-RISE performs slightly worse than T-RISE with an AUC of,
respectively, 0.503, 0.44 and 0.476. All the curves after a first peak display a consistent drop in the error
and then all of them recover. This could be due to hallucination efects created by replacing true input
values with a constant (the mean) value [14, 24, 25].
4.3.2. Graphical Explanation
In order to deduce a more conceptual explanation, S-RISE, T-RISE and ST-RISE are analyzed through
visual inspection of their saliency representations. Indeed, Figure 7 shows the explanations generated
by RISE on two instances of the test set. The first instance presents the saliencies for the prediction at
the beginning of the spring 2022 (27-03-2022), while the latter one of the winter 20232 (25-12-2022).</p>
          <p>From the spatial perspective, terrain’s structure is pivotal in groundwater phenomena because water
under the terrain flows from mountain to valley and then plain. This means that the WTD level is
extremely related to the mountains and valleys’ weather nearby the sensor location. Thus, given that
in our ROI the highest mountains are in the northwest, we expect the model to be very sensible to
precipitation in that area. The model seems to fulfil our expectations. Indeed, the most relevant pixels
in S-RISE for both instances are focused on the west and northwest with respect to the sensor’s location.
Even if most relevant pixels are outside the catchment area, this explanation is plausible. Indeed,
hydrological catchments are mainly defined focusing on river courses and thus groundwater bodies
could be related to multiple catchments. Focusing on the Vottignasco WTD sensor, even if it is included</p>
        </sec>
        <sec id="sec-4-2-2">
          <title>2We consider as winter 2023 the period of time between 21-12-2022 and 20-03-2023.</title>
          <p>in the Grana-Maira catchment, it is right in front of the Varaita valley3, and from the domain knowledge
it is sensible to have influences from even northern valleys. This is exactly what the saliency map from
S-RISE highlights.</p>
          <p>Concerning T-RISE, the saliency vectors produced for the two instances present substantial diferences.
The saliency vector of instance 12, inferring the WTD in spring, suggests that the prediction is mostly
influenced by precipitation in winter 2022 and autumn 2021. A smaller impact on the model’s output is
also provided by data in winter 2021. Given that snow is most responsible for long-term dependencies
and rain exerts mainly mid-short-term efects on groundwater bodies, it is sensible to interpret the
saliency of winter 2021 as mainly driven by snow.</p>
          <p>The saliency vector of instance 51 shows that the most influential precipitation period for the
prediction in winter 2023 is the previous season, i.e. autumn 2022, but the former winter (2021) is
almost equivalently relevant, highlighting a very long-term dependency. This could be because of the
precipitation scarcity in 2022 and because at the beginning of winter 2023 the snow had been very rare
and thus the previous year’s snow-pack gained relevance.</p>
          <p>The major benefit of using ST-RISE is that it is possible to look for spatio-temporal relevance patterns
for each element of the input (i.e. pixel and frame tuples). For the instance 12, it appears that the
most influential pixels of the input lie almost in the same locations over time. Consequently, no
spatiotemporal events emerge from this representation which are not deducible from S-RISE and T-RISE
disjointly. Instead, for the instance 51, a slight but visible shift of the saliency over the spatial and
3The Varaita catchment is just above in the north with respect to the Grana-Maira catchment.
temporal dimensions is present. More in detail, in winter 2022 the saliency is concentrated on the
central northern area, while in autumn 2022 it has moved to the south, focusing on the central and the
lower part of the ROI. This is in line with the possible interpretation of the temporal explanation given
by the T-RISE saliency vector. Indeed, given the low precipitation in 2022 and low snow in early winter
2023, the previous year’s snow-pack (winter 2022) could have been very relevant. It could be worth
considering that in the northwest there is the Monviso, which is the highest mountain in the Cottian
Alps, and thus, that area is a considerable snow reserve. Southern areas could have more relevance as
time approaches the prediction date because of the rainfall collected by the Grana-Maira and Varaita
catchment. Indeed, as already stated, the Vottignasco sensor is just in front of the Varaita Valley and it
is very near the Maira River (which originates in the Grana-Maira catchment). This spatio-temporal
saliency pattern could not have been detected by S-RISE and T-RISE separately, it is exactly for this
reason that the usefulness of ST-RISE emerges.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions</title>
      <p>In this paper, we have proposed a formalization of RISE application for spatial (S-RISE), temporal
(T-RISE) and spatio-temporal (ST-RISE) explanations in a regression task. These implementations of
RISE have been applied to the CNN-LSTM model described in [7], predicting the WTD in Vottignasco
(Piemonte, IT) taking in input a stream of weather image data (i.e. a video). The saliency representations
produced in this setting resulted explicative of the events occurring over time and space. In particular,
ST-RISE has proved to be significantly useful in detecting jointly areas and times most relevant for a
particular channel (i.e. variable) for predicting a specific instance. In other words, its local explanations
are more fine-grained than the S-RISE and T-RISE, allowing recognition of saliency patterns that with
S-RISE and T-RISE solely would not have been possible. Nonetheless, ST-RISE is compensated for its
precision by incurring a higher computational cost, needing 5000 iterations instead of 1000 of the other
approaches.</p>
      <p>This work can be extended in multiple directions. First, in this study, we have just focused on one
single variable. It could be worthwhile to investigate also the other available variables (i.e. channels)
either individually or jointly. Moreover, a comparison with other XAI methods is needed to highlight
the merits and limits of the presented approach, especially for spatio-temporal regression tasks.</p>
      <p>RISE has proved to be useful in detecting the most relevant areas and times for WTD forecasting.
In the end, this could be extremely helpful for environmental monitoring and policy implementation,
fostering regional resilience and sustainability.
Networks With Perturbation, in: Proceedings of the IEEE/CVF Winter Conference on Applications
of Computer Vision, 2021, pp. 1120–1129.
[21] T. Uchiyama, N. Sogi, K. Niinuma, K. Fukui, Visually explaining 3D-CNN predictions for video
classification with an adaptive occlusion sensitivity analysis, in: 2023 IEEE/CVF Winter Conference
on Applications of Computer Vision (WACV), IEEE, Waikoloa, HI, USA, 2023, pp. 1513–1522.
doi:10.1109/WACV56688.2023.00156.
[22] Q. Pan, W. Hu, J. Zhu, Series Saliency: Temporal Interpretation for Multivariate Time Series</p>
      <p>Forecasting, 2020. doi:10.48550/arXiv.2012.09324. arXiv:2012.09324.
[23] U. Schlegel, D. Oelke, D. A. Keim, M. El-Assady, An Empirical Study of Explainable AI
Techniques on Deep Learning Models For Time Series Tasks, 2020. doi:10.48550/arXiv.2012.04344.
arXiv:2012.04344.
[24] U. Schlegel, H. Arnout, M. El-Assady, D. Oelke, D. A. Keim, Towards A Rigorous Evaluation Of
XAI Methods On Time Series, in: 2019 IEEE/CVF International Conference on Computer Vision
Workshop (ICCVW), 2019, pp. 4197–4201. doi:10.1109/ICCVW.2019.00516.
[25] A. Theissler, F. Spinnato, U. Schlegel, R. Guidotti, Explainable AI for Time Series Classification: A
Review, Taxonomy and Research Directions, IEEE Access 10 (2022) 100700–100724. doi:10.1109/
ACCESS.2022.3207765.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>D.</given-names>
            <surname>Aineto</surname>
          </string-name>
          , R. De Benedictis,
          <string-name>
            <given-names>M.</given-names>
            <surname>Maratea</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mittelmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Monaco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Scala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Serafini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Serina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Spegni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Tosello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Umbrico</surname>
          </string-name>
          , M. Vallati (Eds.),
          <source>Proceedings of the International Workshop on Artificial Intelligence for Climate Change, the Italian workshop on Planning and Scheduling</source>
          , the RCRA Workshop on
          <article-title>Experimental evaluation of algorithms for solving problems with combinatorial explosion, and</article-title>
          the Workshop on Strategies, Prediction, Interaction, and
          <article-title>Reasoning in Italy (AI4CC-IPS-RCRA-SPIRIT 2024), co-located with 23rd International Conference of the Italian Association for Artificial Intelligence</article-title>
          (AIxIA
          <year>2024</year>
          ), CEUR Workshop Proceedings, CEUR-WS.org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Intergovernmental</given-names>
            <surname>Panel on Climate</surname>
          </string-name>
          <string-name>
            <given-names>Change</given-names>
            ,
            <surname>Climate Change</surname>
          </string-name>
          2022 - Impacts, Adaptation and Vulnerability: Working Group II Contribution to the
          <source>Sixth Assessment Report of the Intergovernmental Panel on Climate Change</source>
          ,
          <source>Technical Report</source>
          , IPCC Cambridge University Press,
          <year>2023</year>
          . doi:
          <volume>10</volume>
          .1017/9781009325844.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>United</given-names>
            <surname>Nations</surname>
          </string-name>
          (UN),
          <source>The United Nations World Water Development Report</source>
          <year>2023</year>
          :
          <article-title>Partnerships and cooperation for water</article-title>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>World</given-names>
            <surname>Meteorological</surname>
          </string-name>
          <article-title>Organization (WMO)</article-title>
          ,
          <source>State of Global Water Resources</source>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>