<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Exploring waste-collection fleet data: challenges in a real-world use case from multiple data providers</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Simone Monaco</string-name>
          <email>simone.monaco@polito.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paolo Bethaz</string-name>
          <email>paolo.bethaz@polito.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Daniele Apiletti</string-name>
          <email>daniele.apiletti@polito.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fabrizio Pio Baldini</string-name>
          <email>fabrizio.baldini@seaeco.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Carlo Caso</string-name>
          <email>carlo.caso@t2d-digital.com</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tania Cerquitelli</string-name>
          <email>tania.cerquitelli@polito.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of control and computer engineering, Politecnico di Torino</institution>
          ,
          <addr-line>Torino</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>SEA Soluzioni Eco Ambientali S.r.l.</institution>
          ,
          <addr-line>Villanova Canavese (TO)</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>T2D Transfer to Digital Salerno</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In the age of connected vehicles, large amounts of data can be collected while driving through a variety of on-board sensors. The information collected can be used for various types of data-driven analytics that can be of great benefit to both vehicle owners, e.g., to reduce costs by means of predictive maintenance, and to society as a whole, e.g., to optimize mobility behavior. Prior to any real-world data analysis, an investigation and characterization of the available data is of utmost importance in order to evaluate the quality and quantity of the data and to set the right expectations. In this paper, we focus on the data exploration and characterization step, which is necessary to avoid inconsistencies in the collected parameters and to enable valid, data-driven modeling. The proposed data exploration considers both the frequency of samples and their values for all monitored parameters. A specific cross-provider data comparison is performed to compare values collected for the same vehicle at the same time from diferent fleet monitoring data providers. The study is applied to a real-world use case with months of data from dozens of vehicles deployed in the waste collection service managed by SEA, Soluzioni Eco Ambientali, in Italy. The analyzes uncover unexpected behaviors in the measurements and lead to their early identification, bringing great benefits to the company operating the fleet by improving data collection and enabling a safe modeling phase.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Real-world data characterization</kwd>
        <kwd>waste-collection</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <sec id="sec-1-1">
        <title>The recent rise of monitoring systems on modern ve</title>
        <p>
          hicles is opening a broad range of applications for
machine learning (ML) techniques to process this data [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].
Among the most popular use cases, analyses that can be
implemented to leverage the collected data may include,
for example, estimating the best route for the vehicle,
predicting fuel consumption, or predicting a failure on
a particular component. However, ML algorithms are
efective as long as the available data are suficiently
accurate. Then, it is fundamental to assert the
reliability of this data before going into any kind of analysis.
This evidence becomes even more important when
dealing with data coming from multiple sources, for which
collecting strategies, precision, and availability can
significantly vary. In this work, we present a real-world
use case where waste-collecting vehicles are monitored
by multiple data providers, and interventions to fix
failures on such vehicles are regularly recorded. The main
contribution of this paper are:
        </p>
      </sec>
      <sec id="sec-1-2">
        <title>1. An exploratory analysis on real data from a waste</title>
        <p>collection company, which is tracking its fleet by
means of diferent data providers.
2. An exploratory analysis of the recorded
maintenance interventions to enable a future predictive
maintenance solution.</p>
      </sec>
      <sec id="sec-1-3">
        <title>The rest of the paper is organized as follows. Section 2</title>
        <p>presents the related works, Section 3 describes the
realworld use case under analysis, and Section 4 focuses
on vehicle-tracking data exploration. Finally, Section 5
draws conclusions and future works.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Related work</title>
      <p>
        In recent years, the advent of the Internet of Things and
its rapid difusion has facilitated real-time
communication and data exchange between technological devices
that can be used in a wide variety of scenarios. When
the technologies involved are located on a vehicle (e.g.
car, truck), the term Internet of Vehicles (IoV) is used
[
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ]. This term refers to the situation in which data are
collected locally on the vehicle and then sent to a remote
storage location (cloud) where it can be deeply analyzed.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Use case description</title>
      <p>
        Many infrastructures and diferent technologies have
been implemented to support large-scale, real-time, and
reliable information services. Basically, for this purpose, We analyze real-world fleet data of garbage collection
a three-tier architecture [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] is used: the first tier with all operated by SEA, Soluzioni Eco Ambientali. The company
the sensors within the vehicle, a second tier representing acts in cooperation with the local administration in many
the communication layer, and the third tier including Italian municipalities, to handle the collection, transport,
statistics tools, support for storage, and processing in- and selection of municipal waste. To this aim, the
comfrastructure. This last tier leverages collected data from pany deploys a fleet of trucks. They are equipped with
sensors to extract information that can bring value in sensors able to track the vehicle during the trips. Every
diferent contexts. day, more than 300 vehicles operate in 14 diferent
work
      </p>
      <p>
        The GPS data which reflects the movement of vehi- sites, located in diferent cities in the north and center
cles can be traced to simulate the mobility model in IoV. of Italy. An unexpected failure of one of these vehicles
Authors in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] have presented diferent machine would need a quick reorganization of the workload of all
learning techniques for short-term congestion prediction the others, to optimally perform the service. Hence, an
using vehicle trajectory available for connected vehicles. early detection of the signals leading to possible failure
In another set of works, the goal is to correctly predict is of primary importance, both from an economic and an
fuel consumption under various driving conditions. This organizational point of view.
is done for example in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] and [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], where prediction re- In this context, we present a preliminary analysis and
sults show the good value of mean absolute percentage characterizations of such data. Specifically, we focused
errors. Finally, a great deal of research has been done to on a subset of 40 trucks, which are the most
representry to exploit the data collected to be able to do predictive tative and well monitored, i.e., those for which the data
maintenance. This implies the building of a data-driven was extensively present for most of the time. From the
model that can predict in advance a failure on a given technological point of view, all the data is stored in a data
component, thus being able to intervene promptly, avoid- lake, managed by T2D Transfer to Digital, which collects
ing breakdowns during vehicle travel. In this context, in the results of regular calls to the REST APIs of the data
[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] authors try for example to estimate the probability providers. The diferent service providers, being either
of fuel pump failure, while in [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] the goal is to predict the truck manufacturer or an external supplier, added
the presence of a Diagnostic Trouble Code ignition on sensors on the vehicle chassis and the top installations.
commercial trucks. Hence, data sources are extremely heterogeneous, and
      </p>
      <p>In many works, data is collected and managed with a also irregular in their timings, which leads to the
presclear purpose in mind. However, in other cases, compa- ence of a variety of anomalies. Among all the available
nies would like to gain general insights of the information data sources, we can identify the following types of data
they collected, and then drive the decisions on how to in the data lake.
efectively reach cost reductions and higher operational
eficiency. In such cases, an exhaustive exploration of the
available data is required, allowing the company to
possibly introduce early changes in the collection process.</p>
      <p>In this way, our work difers from the above-mentioned
ones because our focus is on the data exploration phase,
to highlight noteworthy behaviors or anomalies of
interest for the data owner. Furthermore, all the
abovementioned works have been conducted on data collected
by a single provider. Instead, in our use case, we analyze
and compare three diferent data providers, each with its
own specific sensors. A single parameter can be either
specific to a single provider, or it can be monitored by
several providers (each with its peculiarities). For this
reason, we focus on a data comparison step, not covered
in previous related works, to highlight the benefits and
weaknesses of each data provider.
• Tracking. This data describes the trips each
truck performs. We have 3 providers (named P1,
P2, and P3 in the following), each tracking
different features. The only common measures are
GPS position, odometer value, total time of
engine activity, and average speed. Some providers
also include other measures such as acceleration
and estimated fuel consumption. In addition to
the variability on the tracked features, diferent
providers also have diferent policies on the
sampling of the data, hence leading to very diverse
daily measure rates. The raw datasets collected by
the providers are time series of a total of 117880
samples of 38 tracked features for 35 vehicles for
P1, 68611 samples of 68 features for 8 vehicles
for P2, and 91072 samples of 17 features for 15
vehicles for P3.
• Maintenance. Among the goals of the current
exploratory analysis, we would like to assess the
feasibility of a future predictive model aiming to
predict failures based on the historical
maintenance interventions. To this aim, in the data lake
Invoices 7 15 8 7 4 1 1 1 6 4 2 2 2 2 2 2 1 7 6 4 4 4 4 3 3 2 2 2 2 1 1 1 1 0 0 3 2 1 0 0
P1
P2
P3
4. Exploratory analysis
we analyzed the list of all the repairs of the dif- P3 on the rows) supply information across the 40
vehiferent components of the trucks, for which the cles (on the columns). The heatmap shows the average
corresponding invoices are collected. We have a daily number of samples, for each vehicle (column) and
total of 120 invoices associated with 36 vehicles. provider (row). The first row reports the number of
invoices for each truck, each corresponding to the
maintenance interventions.</p>
      <p>Among the 40 vehicles, the first 17 share tracking data
from multiple providers at the same time, and are
described by more than 30 samples per day per vehicles
in some cases, with the average being less than 5
samples per day per vehicle. The number of maintenance
interventions (invoices) ranges from zero to 15 in the 9
months under analysis.</p>
      <p>In the following, we discuss the results of the
exploratory analysis on such data.</p>
      <sec id="sec-3-1">
        <title>In this section, we describe the data analysis process</title>
        <p>and results. The dataset of interest in the current work
includes information of 40 vehicles from May 2021 to
January 2022. Even though the data providers’ REST
APIs have been directly used to collect the data at regular
and synchronous intervals during the whole period of
time, the timestamps of some samples refer to periods
sensibly far from the range (May 2021 - January 2022).</p>
        <p>Furthermore, significant portions of the original data are 4.1. Multi-provider data comparison
duplicated, i.e., diferent calls to the provider APIs return
the very same data, also with identical timestamps. Some
of these anomalous repeated data are associated with a P1
subsequent period with no collected data. Nevertheless,
the overall trend of the vehicles in these situations sug- Location_references(4)
gests they where still moving without being tracked. For Alarm_id
this reason, we dropped all duplicated values down, since
we assumed they are only the result of irregularities of
the providers’ update systems. In the same way, we then Engine_working_hour
also removed all samples outside of the considered time cuOrrednotm_septeered
range. Performance_stats(5) GPS_coordinates</p>
        <p>Since at each API call, all variables are returned by the Guidemode_distance(4)
data provider, but not all variables efectively recorded a Fuel_Acoltnitusudmeption P3
change in their value, a feature selection step has been P2
used to remove irrelevant features. After removing
duplicate records, we also removed all those variables that we
felt would not be helpful for an upcoming data-driven
analysis. These variables were identified according to the
following criteria: (i) a missing value ratio greater than Figure 2: Venn diagram of the high-level semantic
aggrega50%, (ii) limited variance (e.g., variables whose values are tion of the features from 3 tracking data providers.
always constant).</p>
        <p>As a first indicator of the remaining information car- Since each provider has diferent variables, we
exedinality, we defined the average number of samples per cuted a semantic association to identify related features
vehicle per day as a synthetic and high-level measure. from diferent providers. The resulting features with
Figure 1 shows how the diferent providers (P1, P2, and the corresponding associations are presented in Figure 2,
Sep
2021</p>
        <p>Oct</p>
        <p>Nov</p>
        <p>Dec
where the Venn diagram shows how the semantic groups From Figure 1 we note that 10 vehicles share the
P1are shared between the 3 providers. and P3-provided odometer and engine working time. In</p>
        <p>The features in common to all the providers are the Figure 3 we report the comparison of such variables for
odometer, the speed, and the geographical position of 3 overlapping trucks (namely v13, v16, and v17) in the
the vehicles. The working time of the engine is available period between August 2021 and November 2021, after a
in 2 out of 3 providers. The odometer and the engine simple conversion of units of measure (P1 uses kilometers
working time are of particular interest in our analysis and hours, while P3 meters and minutes).
since they give a measure of the mechanical "age" of the P1 presents a temporal gap in the collected samples, of
truck, and despite some possible diferences on the ini- about one week from September 1st, for all the reported
tial ofset among diferent providers, they are supposed trucks. Since the last measure for each vehicle before
to be sampled coherently among them all. Other vari- the gap and the first one after drastically difer, we can
ables, such as the current speed, fuel consumption, and assume that within this period the trucks were still
acperformance statistics, are also potentially significant to tive. This assumption is proven by the measures from
trace the degradation of the vehicles. However, even if P3, confirming that it was associated with a lack of
inforthey share the same high-level semantics, we noticed mation from the provider itself. The issue is common to
that their values were so diferent that each provider ac- all the data of this source, and it was not reported to the
tually tracks a variable on its own, hence they cannot be knowledge of the company before the current analysis.
merged together. Moreover, this period, as highlighted by the lower plot,
is also associated with a drastic lowering in the value
4.1.1. Comparative analysis on monotonic features of the engine time. This can correspond to a sort of
internal reset of the sensors, which may have made the
The odometer and the engine working time are two mono- registered values start from a default position, slightly
tonic variables shared among diferent data providers (P1 diferent from any truck. In this situation, a fix for this
and P3). Both variables are constant when the vehicle error can be obtained by adding the first value from P3 of
is turned of, whereas both must be strictly increasing the correspondent vehicle to all measures of P1. With this
while driving. correction, each couple of curves after the gap becomes
comparable.</p>
        <p>Finally, the upper plot also shows another anomaly
concerning the odometer tracking of v17. The traces from
the two providers have significantly diferent trends that
continue diverging as time increases. Vehicle v17 traveled
with an average speed of 83 / and 23 /,
for P1 and P3, respectively. All providers’ average speed
distributions are peaked at around 60 /, as shown
in Fig. 4, but the calculated average values are on the tails
of both distributions.</p>
        <p>In this situation, as for other trucks monitored by more
than one data provider, we experienced that information
from diferent providers can reveal significant diferences
in the reported values.</p>
        <sec id="sec-3-1-1">
          <title>4.2. Analysis over time</title>
          <p>An important analysis to perform in order to understand
efectively the data collected is an exploration of the time
period in which the vehicles are monitored. How long
and in which manner the data is collected is in fact a
useful information for the company that can highlight
important diferences between the various providers.
Below, for each provider we have reported a graph
representing the time on the x axis, and the diferent trucks
monitored by the provider on the y axis. In this way, for
each truck we can easily see the start and end date of
collection and whether there were any anomalies in the
monitored period (e.g. missing data or changes in the
daily data rate).
4.2.1. P1 provider
The graph shown in Fig 5 shows all trucks monitored by
provider P1. The diferent color shades of the points
represent a diferent sampling frequency on a given day (the
darker the point, the more measurements were collected
on the same day for that truck). This graph shows that
the number of daily measurements collected is almost
the same and that all trucks were monitored during the
period from May 2021 to January 2022. The anomalies
are highlighted by the red rectangles in the figure. The
ifrst anomaly that is detected is common to all trucks, and
is a lack of data at the beginning of 2021/09 (September).</p>
          <p>This behavior can be explained by a period of company
closure, or by a general failure in the data collection
process, afecting all trucks. Other anomalies are instead
related to specific trucks, and show rather long periods
(even several months) of lack of data. In general, to see if
in all of these situations we don’t have data because the
truck was actually idled or because there was a problem
in the collection, we need to look at the cumulative
values collected from variables such as the odometer. If the
value of this variable in the last sample collected before
the period without data is equal to the value of the first
sample collected after the period without data, it means
that the truck actually did not travel in that period, so no
information was lost. Otherwise, if the two values are
different, it means that the truck moved but we don’t have
data that monitored its operations during that period. So
there was an error in the data collection phase.</p>
          <p>Sampl963es
12
15
v01
v02
v03
v04
v05
v06
v07
v08
v09
v10
v11
v12
v13
v14
v15
lepvvvv11219706
tav18
v21
v22
v23
v24
v25
v26
v27
v28
v29
v30
v31
v32
v33
v34
v35
2021-05
2021-06
2021-07
2021-08
2021-11
2021-12</p>
          <p>2022-01
2021-09 date 2021-10
A similar analysis was performed for the measurements
collected by provider P2 and shown in Figure 6. In this
case, the number of trucks monitored is much lower than
those monitored by P1. Moreover, all the trucks
considered by P2 are also monitored by P1. This is useful in
order to make a quick cross-check on anomalies between
the two providers. In particular, we can easily notice that
in the same periods in which data for P1 was missing
(both the period common to all trucks and those for
individual trucks), we now have measurements collected by
P2. This leads us to assume that the previous lack of data
were really due to P1-related failures. The only period
in which we have no measurements in either providers
is between 2021/12 and 2022/01 for vehicle v03. In this
case, the only check we can do to verify if the truck was
really stopped is to monitor the odometer value before
and after the lack of data.
we manually excluded all the elements associated with
periodical maintenance operations to be done as part of</p>
          <p>Regarding the measurements collected by P2, a partic- the security enforcement by law, then all the resulting
ular behavior can be noticed in the sampling frequency. elements have been grouped into meaningful categories
Until 2021/12 there is a very low sampling frequency, based on the failed components. Looking at the
remainequal to around one daily measurement per truck. How- ing invoices, we then observed a significant number of
ever, this behavior changes suddenly in the last months operations from the same category carried out on the
of collection, where the sampling frequency increases same vehicle within a short time (e.g., the day after). We
significantly. This seems to be due to a change in the assumed all these cases describe a situation when the
collection pattern from December onward, probably with repair was insuficient to solve the problem, hence not
the aim of collecting information with a higher level of de- relevant for a predictive-maintenance analysis. For this
tail than the single daily information previously collected. reason, we excluded such data from the analysis. The
reThe use of the trucks has not changed since December, sulting, clean number of failures, grouped by categories,
what has changed is the amount of information describ- is shown in Figure 8. For each row, besides the total
numing the behavior of a truck, with up to 200 samples per ber of failures, also the number of second occurrences,
day. i.e., those which are not the first failures for the vehicle
and the category, is reported. The failures coming from
4.2.3. P3 provider this second group are particularly interesting when
investigating the possibility of building a predictive model on
The same temporal analysis has also been performed for top of this data since they can give a preliminary measure
the trucks monitored by provider P3. Five of the trucks of the frequencies of the failures.
in Figure 7 (from v36 to v40) are monitored only by P3,
therefore it is impossible to cross-check them with the Electrical system
data of the other providers to try to better understand the Braking system
anomalies. Regarding instead the remaining trucks, they Mechanics
are all monitored by at least one between P1 and P2 and Engine lubrification system
this allows us to cross-check the collected measurements.</p>
          <p>In particular, the most evident anomalies in the data Hydraulics
collected by P3 are for those five trucks where there are Sensors
nHoowmeevaesur,rfeomretnhtes sianmtheetrpuecrkiosdafnrdomin2t0h2e1/s0a5mteo m20o2n1t/h10s,. ExhaPuoswtesrypstlaenmt TSoetcaolnd-occurrence
the other providers have collected measurements, so it is 0 5 10 Count of lab1el5ed failures 20 25
fair to assume that the lack of data for P3 is a problem
related only to that provider. Figure 8: Clean count of the failures, separated by category.</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>4.3. Maintenance analysis</title>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>Concerning the maintenance data, we analyzed the col</title>
        <p>lection of invoices indicating failures of diferent
components on the trucks. With the help of domain experts,
The entire process of cleaning and correctly
grouping the available invoices is fundamental whether the
collected data will be exploited to address a predictive
maintenance analysis, building a predictive model
capable of estimating the residual life of a truck before a
maintenance intervention is needed. The possibility of
identifying malfunctions or component issues in advance
is in fact a key aspect for automotive companies, which
can then intervene properly before a failure occurs. This
leads to significant cost savings for the company, since
a properly implemented predictive model would allow
to avoid breakdowns during vehicle travel and to
constantly monitor the condition of the various components,
scheduling maintenance interventions only when they
are really necessary. From Figure 8, we note that for a
few categories we do not have enough data to be able
to build a robust predictive model for estimating future
failures. This is due to the limited time period in which
trucks are monitored, which is not long enough to collect
enough maintenance interventions for all categories.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5. Conclusions and future works</title>
      <sec id="sec-4-1">
        <title>In this work, we presented an exploratory analysis and</title>
        <p>preliminary characterization of real-world fleet data
coming from diferent tracking providers. We analyzed and
compared the data collected by 3 providers, on 40
vehicles, over 9 months, highlighting issues and challenges.
The presence of multi-provider information for the same
vehicle, although more expensive, can help recover the
frequent issue of missing data. As future work, we plan
to extend the current work to integrate data from
different sources with the goal of designing a predictive
maintenance algorithm, by exploiting the maintenance
intervention data. Anomalous information from the
vehicle tracing, like the ones we reported in this paper, can
lead to unreliable input features and consequently
uncertain prediction. Accurate data exploration can instead
improve its quality and then overcome the initial cost
of implementing redundant monitoring strategies with
all the benefits, both in terms of cost and
environmental impact reduction, an accurate failure detection could
provide.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Theissler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pérez-Velázquez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kettelgerdes</surname>
          </string-name>
          , G. Elger,
          <article-title>Predictive maintenance enabled by machine learning: Use cases and challenges in the automotive industry</article-title>
          ,
          <source>Reliability engineering &amp; system safety 215</source>
          (
          <year>2021</year>
          )
          <fpage>107864</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>O.</given-names>
            <surname>Kaiwartya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. H.</given-names>
            <surname>Abdullah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Cao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Altameem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Prasad</surname>
          </string-name>
          , C.-
          <string-name>
            <surname>T. Lin</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Liu</surname>
          </string-name>
          ,
          <article-title>Internet of vehicles: Motivation, layered architecture, network model, challenges, and future aspects</article-title>
          ,
          <source>IEEE Access 4</source>
          (
          <year>2016</year>
          )
          <fpage>5356</fpage>
          -
          <lpage>5373</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>K. M. Alam</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Saini</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>El Saddik</surname>
          </string-name>
          ,
          <article-title>Toward social internet of vehicles: Concept, architecture, and applications</article-title>
          ,
          <source>IEEE access 3</source>
          (
          <year>2015</year>
          )
          <fpage>343</fpage>
          -
          <lpage>357</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>K.</given-names>
            <surname>Golestan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Soua</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Karray</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Kamel</surname>
          </string-name>
          ,
          <article-title>Situation awareness within the context of connected cars: A comprehensive review and recent trends</article-title>
          ,
          <source>Information Fusion</source>
          <volume>29</volume>
          (
          <year>2016</year>
          )
          <fpage>68</fpage>
          -
          <lpage>83</lpage>
          . URL: https://www.sciencedirect.com/science/ article/pii/S1566253515000743. doi:https://doi. org/10.1016/j.inffus.
          <year>2015</year>
          .
          <volume>08</volume>
          .001.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Elfar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Talebpour</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. S.</given-names>
            <surname>Mahmassani</surname>
          </string-name>
          ,
          <article-title>Machine learning approach to short-term trafic congestion prediction in a connected environment</article-title>
          ,
          <source>Transportation Research Record</source>
          <volume>2672</volume>
          (
          <year>2018</year>
          )
          <fpage>185</fpage>
          -
          <lpage>195</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S. J.</given-names>
            <surname>Kamble</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. R.</given-names>
            <surname>Kounte</surname>
          </string-name>
          ,
          <article-title>Machine learning approach on trafic congestion monitoring system in internet of vehicles</article-title>
          ,
          <source>Procedia Computer Science</source>
          <volume>171</volume>
          (
          <year>2020</year>
          )
          <fpage>2235</fpage>
          -
          <lpage>2241</lpage>
          . URL: https://www.sciencedirect.com/science/ article/pii/S1877050920312321. doi:https: //doi.org/10.1016/j.procs.
          <year>2020</year>
          .
          <volume>04</volume>
          .241, third International Conference on Computing and
          <article-title>Network Communications (CoCoNet'19).</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>R.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dubey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Pugliese</surname>
          </string-name>
          ,
          <article-title>Hybrid electric buses fuel consumption prediction based on real-world driving data</article-title>
          ,
          <source>Transportation Research Part D: Transport and Environment</source>
          <volume>91</volume>
          (
          <year>2021</year>
          )
          <article-title>102637</article-title>
          . URL: https://www.sciencedirect.com/ science/article/pii/S1361920920308221. doi:https: //doi.org/10.1016/j.trd.
          <year>2020</year>
          .
          <volume>102637</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Easa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Qu</surname>
          </string-name>
          ,
          <article-title>Modeling relationship between truck fuel consumption and driving behavior using data from internet of vehicles</article-title>
          ,
          <source>Computer-Aided Civil and Infrastructure Engineering</source>
          <volume>33</volume>
          (
          <year>2018</year>
          ). doi:
          <volume>10</volume>
          .1111/mice.12344.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>J.-D. Wu</surname>
          </string-name>
          , J.-C. Liu,
          <article-title>Development of a predictive system for car fuel consumption using an artificial neural network</article-title>
          ,
          <source>Expert Systems with Applications</source>
          <volume>38</volume>
          (
          <year>2011</year>
          )
          <fpage>4967</fpage>
          -
          <lpage>4971</lpage>
          . URL: http://www.sciencedirect.com/science/article/ pii/S0957417410011127. doi:https://doi.org/ 10.1016/j.eswa.
          <year>2010</year>
          .
          <volume>09</volume>
          .155.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>P.</given-names>
            <surname>Bethaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Cavaglion</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Cricelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Liore</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Manfredi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Salio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Regalia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Conicella</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Greco</surname>
          </string-name>
          , T. Cerquitelli,
          <article-title>Empowering commercial vehicles through data-driven methodologies</article-title>
          ,
          <source>Electronics</source>
          <volume>10</volume>
          (
          <year>2021</year>
          )
          <fpage>2381</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>