<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A p pli c a ti o n o f m a c hi n e le a r ni n g f o r im p r o vi n g c o n t e m p o r a r y E T A fo r e c a s ti n g</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>G a b ri el ė K a s p ut yt ė</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>A r n a s M at u s e vi či u s</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>a n d T o m a s K ril a vi či u s</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>1 V yt a ut as M a g n us U ni versit y, F ac ult y of I nf or m atics, De p art me nt of M at he m atics a n d St atistics, Vilei k os street 8, L T- 4 4 4 0 4 K a u n as, Lit h u a ni a 2 V yt a ut as M a g n us U ni versit y, F ac ult y of I nf or m atics, De p art me nt of A p plie d I nf or m atics, Vilei k os street 8, L T- 4 4 4 0 4 K a u n as, Lit h u a ni a 3 Ce ntre f or A p plie d Rese arc h a n d De vel o p me nt, Lit h u a ni a A b str a ct T h e p o p ul arit y of r o a d tr a ns p ort is gr o wi n g e v er hi g h er i n t h e m o d er n d a y of a g e a n d is t h e m ost p o p ul ar m o d e f or tr a ns p orti n g g o o ds [1 ]. Us u all y, tr a ns p ort ati o n i n cr e as es t h e pri m e c ost of t h e g o o ds or s er vi c es, t h us, t o i n cr e as e a c o m p a n y's pr o fit, e ff e cti v e v e hi cl e r o uti n g b e c o m es e ss e nti al. I n t his r es e ar c h, w e h a v e a n al ys e d t h e p ossi bilit y of i m pr o vi n g t h e f or e c ast of esti m at e d ti m e of arri v al b y r a n ki n g t h e dri v ers b as e d o n t h eir b e h a vi o ur d at a a n d esti m ati n g d e vi ati o ns fr o m pl a n n e d arri v al ti m e usi n g di ff er e nt m a c hi n e l e ar ni n g m et h o ds. T h e r a n ki n g w as p erf or m e d wit h T O P SI S a n d VI K O R m et h o ds, w hil e t h e f or e c asti n g w as p erf or m e d usi n g fi v e m a c hi n e l e ar ni n g al g orit h ms: d e cisi o n tr e e, r a n d o m f or est, X G B o ost, S u p p ort Ve ct or M a c hi n e a n d k - N e ar est N ei g h b o urs. T h e p erf or m a n c e of t h e f or e c asti n g m o d els w as e v al u at e d usi n g t h e a dj ust e d c o e ffi ci e nt of d et er mi n ati o n, r o ot s q u ar e m e a n err or a n d m e a n a bs ol ut e err or m etri cs. It w as c o n cl u d e d t h at t h e VI K O R m et h o d s h o ul d b e us e d t o r a n k t h e dri v ers. M or e o v er, r es e ar c h r es ults r e v e al e d t h at t h e b est f or e c asti n g p erf or m a n c e w as a c hi e v e d usi n g a n e ns e m bl e m o d el b as e d o n r a n d o m f or est a n d S u p p ort Ve ct or M a c hi n e m o d els. K e y w o r d s Esti m at e d ti m e of arri v al ( E T A), fr ei g ht tr a ns p ort, d ri v er's s c or e, ra n ki n g, m a c hi n e l e ar ni n g, T O P SI S, VI K O R, D e cisi o n Tr e e, R a n d o m F or est, X G B o ost, S u p p ort Ve ct or M a c hi n e, k - N e ar est N ei g h b o urs 1. I nt r o d u cti o n T o i m pr o v e t h e tr a n s p ar e n c y of a s u p pl y c h ai n, it s p arti ci p a nt s u s e tr a n s p ort m a n a g e m e nt s y st e m s a s w ell a s tr a c ki n g a n d m o nit ori n g s y st e m s. Ve hi cl e s ar e g etti n g e q ui p p e d wit h str o n g er mi cr o pr o c e s s or s, l ar g er m e m or y c a p a cit y a n d r e al-ti m e o p er ati n g s y st e m s. T h e n e wl y i n st all e d t e c h n ol o gi c al pl atf or m s c a n u s e m or e a d v a n c e d 2. R el at e d w o r k a p pli c ati o n s of t h e o p er ati n g s y st e m, i n cl u di n g m o d elb a s e d pr o c e s s c o ntr ol f u n cti o n s, arti fi ci al i nt elli g e n c e, T o b ett er u n d erst a n d w h at attri b ut es a n d m et h o ds c a n b e a n d c o m pr e h e n si v e c o m p ut ati o n. u s e d f or r a n ki n g dri v er s, r e s e ar c h [2 ] w a s st u di e d. T h e T h er ef or e, a n i n cr e a si n g tr e n d of i m pl e m e nti n g t h e ai m of t hi s r e s e ar c h w a s t o c at e g ori z e dri v er s a c c or dl at e st d e v el o p m e nt s i n i nf or m ati o n a n d i nt er n et t e c h- i n g t o t h eir ri s k- pr o n e n e s s b y a n al y si n g a G P S- b a s e d n ol o gi e s i n t h e tr a n s p ort s e ct or a s w ell a s a t e n d e n c y of d e vi c e' s ur b a n tr a ffi c d at a. Hi er ar c hi c al Cl u st eri n g Al g od e v el o pi n g n e w r e s e ar c h- b a s e d s y st e m s t h at c a n q ui c kl y rit h m ( H C A) a n d Pri n ci p al C o m p o n e nt A n al y si s ( P C A) a d a pt t o a n e v er- c h a n gi n g e n vir o n m e nt i s o b s er v e d. T h e w er e us e d f or t h e st atisti c al a n al ysis of t h e f oll o wi n g dri vg o al of t hi s st u d y i s t o i d e ntif y m et h o d s, b e st s uit e d f or i n g p ar a m et er s: S p e e d o v er 6 0 k m/ h , S p e e d , Ac c el er ati o n , s ol vi n g t h e e sti m at e d ti m e of arri v al ( E T A) f or e c a sti n g P ositi v e a c c el er ati o n , Br a ki n g a n d M e c h a ni c al w or k . T h e pr o bl e m. a ut h ors of t his r es e ar c h c o n cl u d e t h at w hil e it is p ossi bl e T h e r e st of t h e p a p er i s or g a ni z e d a s f oll o w s. R el at e d t o cl a s sif y t h e dri v er s a c c or di n g t o t h e s e p ar a m et er s, a w or k i n t hi s ar e a i s pr e s e nt e d i n S e cti o n 2 . S e cti o n 3 i n- l ot of e xt er n al f a ct or s, s u c h a s t h e dri vi n g e n vir o n m e nt or t h e c o n diti o n of t h e ass ess e d dri v er, ar e n ot t a k e n i nt o I V U S 2 0 2 2: 2 7t h I nter n ati o n al C o nfere nce o n I nf or m ati o n Tec h n ol o g y, a c c o u nt. M a y 1 2, 2 0 2 2, K a u n as, Lit h u a ni a A c o m p ar ati v e a n al y si s of t w o m ulti- crit eri a d e ci si o n* C orr es p o n di n g a ut h or. m a ki n g ( M C D M) m et h o d s T O P SI S a n d VI K O R i s pr e† T h es e a ut h ors c o ntri b ut e d e q u all y. s e nt e d i n [3 ]. It w a s f o u n d t h at t h e s e t w o m et h o d s u s e g a bri el e. k as p ut yt e @ v d u.lt ( G. K as p ut ytė); di ff er e nt ki n ds of n or m ali z ati o n ( T O P SI S u s es v e ct or n orar n as. m at us e vi ci us @ v d u.lt ( A. M at us e vi či us); m ali z ati o n, VI K O R - li n e ar) t o eli mi n at e t h e u nit s of t o 0m 0as0.0k-r0il0a0v1i-c8i 5u0s 9-@4v2d0u.Xlt((T.T. KKrriillaavviiččiiuuss)) crit eri o n f u n cti o n s a n d u s e di ff er e nt a g gr e g ati n g f u n c© 2022 C o p yri g ht f or t his pa per b y its a ut h ors. Use per mitte d u n der Creati ve C o m m o ns Lice nse ti o ns f or r a n ki n g. N o n et h el ess, b ot h m et h o ds w er e f o u n d CPWroEorUcksRehodipngs hItSp:/Nc1e6ur1-3w-s0.or7g3 ACttEri bUutiRo n 4.W0oI rntekrsnahti oonpal (PCrC oBcY e4.e0).di n gs ( C E U R- W S. or g ) s uit a bl e f or s ol vi n g r a n ki n g pr o bl e m s.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>tr o d u c es t h e d at as ets us e d i n t h e c urr e nt st u d y. S e cti o n 4
pr es e nts t h e s el e ct e d r a n ki n g a n d f or e c asti n g t e c h ni q u es.</p>
      <p>E x p eri m e nt al r e s ult s ar e pr o vi d e d i n S e cti o n 5 . Fi n all y,
c o n cl u di n g r e m ar k s a n d f ut ur e pl a n s ar e di s c u s s e d i n
S e cti o n 6 .</p>
      <p>3. Research data</p>
      <p>The use of machine learning (ML) techniques for the
ETA problem is discussed in [4] The authors applied
artificial neural networks (ANNs) and support vector re- 3.1. Data for ranking drivers
gression (SVR) to predict the time of arrival of container
ships. Distance to the destination, the timestamp, geo- The available readings from a vehicle monitoring system
coordinates, and weather information have been chosen that can be used to evaluate a driver’s behaviour were
as features. It was shown that SVR had performed better extracted. The readings in this research cover the period
than ANNs and that the weather data did not have a sig- from August 21, 2020, to January 1, 2022, and up to 398
nificant impact on estimating the time of vessel arrivals. observations representing diferent vehicles.</p>
      <p>In the [5] study, the -Nearest neighbours (KNN), SVR, A dataset was constructed containing values for 7
atand the random forest algorithms were evaluated as meth- tributes, namely Free-rolling distance, Engine overloaded
ods for predicting the arrival time of open-pit trucks. A distance, Highest gear distance, Excess idling, Overspeeding
site-based approach was used as the position was only time, Extreme braking events and Harsh braking events.
measured at a few discrete nodes of the route network. It
was concluded that the random forest algorithm provides 3.2. Data for forecasting model
the best prediction results. In this research, logistic transportation data was reviewed</p>
      <p>Ma et al. [6] proposed a tree-based integration method and a dataset was created for the ETA forecasting models.
to predict trafic accidents by using diferent data vari- The initial dataset includes 1758 observations and 13
ables. Predictions of the gradient boosting decision variables. The obtained information is from August 21,
tree algorithm outperformed back propagation neural 2020 to January 24, 2022. A set of explanatory variables X,
networks, support vector machines, and random forest. with vectors 1, 2, . . . , 13, is obtained. The description
However, in this study, the nonlinear relationship be- of these variables is presented in Table 1.
tween the influence characteristics and the predicted
value was not analysed.</p>
      <p>To improve travel time predictions, the author of [7] Table 1
study applied the combination of random forest and gra- Set of explanatory variables
dient boosting regression tree (GBRT) models. The aim Variable Description Variable type
was to study how reducing a large volume of raw GPS 1 Driver’s score Ordinal
data into a set of feature data afects high-quality travel 2 Tour beginning country Categorical
time predictions. Only travel time observations from 3 Tour ending country Categorical
the previous departure time intervals were found to be 4 Number of intermediate stops Discrete
tboenbeeficiuaslefdeaatusrienspauntds wwehreenrencoomotmheerndtyedpebsyothfereaault-htiomre 567 FTToouuurtrrhbbeeesggtiicnnonnuiinnnggtrydmaoynth CCCaaattteeegggooorrriiicccaaalll
information (e.g. trafic flow, speed) is available. Also, 8 Vehicle height Continuous
it is noted, that trees in GBRT models were found to be 9 Vehicle width Continuous
consistently much shorter than those of random forest 10 Vehicle length Continuous
models, leading to shorter computation times. 11 Vehicle weight Continuous</p>
      <p>To sum up, characterizing attributes of drivers in re- 12 Hours of service breaks Continuous
search are usually derivative – data obtained from vehi- 13 Planned distance Continuous
cle monitoring devices represent their driving behaviour.</p>
      <p>Since MCDM methods were found popular for conduct- Let  be the factual time when the th cargo will be
ing ranking procedures, two easily comparable MCDM delivered, and  the planned time of delivery for the th
methods will be used: TOPSIS and VIKOR. What con- cargo. Then, the deviation from the planned time of
delivcerns the ETA problem, the accuracy of the ML models ery for the th cargo ∆ will be denoted as the diference
tested in reviewed research was inconsistent. Therefore, between the planned and factual time of delivery:
a wide variety of ML methods suitable for the problem
and available data will be evaluated. Namely, decision ∆ =  − , (1)
tree, random forest, XGBoost, Support Vector Machine
(SVM) and KNN methods, as well as an ensemble of mod- where  = 1, 2, . . . , . In that case, the explanatory
els, will be tested. variable is denoted as:
 = ∆.</p>
      <p>(2)
This variable is the goal of the forecasting problem.</p>
    </sec>
    <sec id="sec-2">
      <title>4. Methodology</title>
    </sec>
    <sec id="sec-3">
      <title>4.1. VIKOR method</title>
      <p>The VIKOR method was introduced as one applicable
technique to implement within MCDM. It focuses on
ranking and selecting from a set of alternatives in the
presence of conflicting criteria, and on proposing a
compromise solution (one or more). The compromise ranking
algorithm VIKOR has the following steps [8]:
1. Determine the best * and the worst − values
of all criterion functions,  = 1, 2, . . . , . If the
th function represents a benefit then:
* = max ,

− = min .</p>
      <p>2. Compute the values  and  ,  = 1, 2, . . . ,  ,
by the relations</p>
    </sec>
    <sec id="sec-4">
      <title>4.2. TOPSIS method</title>
      <p>The basic principle of the TOPSIS method is that the
chosen alternative should have the shortest distance from the
ideal solution and the farthest distance from the
negativeideal solution [9]. The TOPSIS procedure consists of the
following steps:
1. Calculate the normalized decision matrix. The
normalized value  is calculated as</p>
      <p>= √︁∑︀
=1 2
,
[︂
 = max 

* −  ]︂ ,
* − −
where  are the weights of criteria, expressing
their relative importance.
3. Compute the values  ,  = 1, 2, . . . ,  , by the
relation</p>
      <p>= 
where
 − *
− − *
+ (1 − )
 − *
− − *</p>
      <p>,

 = ∑︁  * −  ,</p>
      <p>=1 * − −
* = min ,</p>
      <p>* = min ,

− = max ,</p>
      <p>− = max ,

and  is introduced as weight of the strategy of
"the majority of criteria" (or "the maximum group
utility"), here  = 0.5.
4. Rank the alternatives, sorting by the values , 
and , in decreasing order. The results are three
ranking lists.
5. Propose as a compromise solution the alternative
(′) which is ranked the best by the measure 
(minimum) if the following two conditions are
satisfied:</p>
      <p>C1. "Acceptable advantage":
(′′) − (′) ⩾ 
(8)
where ′′ is the alternative with second
position in the ranking list by ;  =
1/( − 1);  is the number of alternatives.</p>
      <p>C2. "Acceptable stability in decision-making":
Alternative ′ must also be the best ranked
by  or/and . Here,  is the weight of the
decision-making strategy "the majority of
criteria".</p>
      <p>where  = 1, . . . ,  ;  = 1, . . . , .
2. Calculate the weighted normalized decision
matrix. The weighted normalized value  is
calculated as</p>
      <p>=  ·  ,
where  = 1, . . . ,  ;  = 1, . . . , ,  is
the weight of the th attribute or criterion, and
∑︀</p>
      <p>=1  = 1.
3. Determine the ideal and negative-ideal solution.
* = {(max  | ∈ ′), (min  | ∈ ′′)},
 
(11)
− = {(min  | ∈ ′), (max  | ∈ ′′)},
 
(12)
where ′ is associated with benefit criteria, and
′′ is associated with cost criteria.
4. Calculate the separation measures, using the
dimensional Euclidean distance. The separation
of each alternative from the ideal solution is given
as</p>
      <p>⎯⎸ 
* = ⎷⎸∑︁( − *)2,</p>
      <p>=1
where  = 1, . . . , .</p>
      <p>Similarly, the separation from the negative-ideal
solution is given as</p>
      <p>⎯⎸ 
− = ⎷⎸∑︁( − −)2,
=1
(3)
(4)
(5)
(6)
(7)</p>
      <p>where  = 1, . . . , .
5. Calculate the relative closeness to the ideal
solution. The relative closeness of the alternative 
with respect to * is defined as
(9)
(10)
(13)
(14)
(15)
* =</p>
      <p>−
* + −
surpasses other ML algorithms by solving many data
science problems faster and more accurately than its
counterparts. Also, this algorithm has additional protection
from overfitting.</p>
      <p>The objective function to be optimized is given by
(16)
where  is the number of iterations, (, ˆ) is the
training loss function, ˆ = ∑︀</p>
      <p>=1 is the number of trees, Ω
is the regularization term,  ∈ ℱ , and ℱ is the set of
possible classification and regression trees.</p>
      <p>Writing the prediction value at step  as ˆ(), gives

ˆ() = ∑︁ () = ˆ(−1) + ().
(17)
=1
Next, a tree which optimizes our objective is chosen.</p>
      <p>() = ∑︁ (, ˆ) + ∑︁ Ω() =
=1 =1

= ∑︁ (, ˆ(−1) + ()) + Ω() + ,</p>
      <p>=1
where  is a constant.</p>
      <p>To minimize the probability of overfitting, the
complexity of the tree Ω( ) is defined as
(18)
(19)</p>
      <sec id="sec-4-1">
        <title>4.3. Decision Tree</title>
        <p>Decision tree is a type of supervised learning algorithm
that can be used in both regression and classification
problems [10]. Each node is related to an attribute, whereas
the leaves of the tree represent the final solution as the
result of combining values of the attributes.</p>
        <p>The splitting process is stopped after a particular
stopping criterion is met. For example, a given threshold for
the minimum number of observations left in a node being
reached or a given threshold for the minimum change in
the impurity measure not succeeding any more by any
variable can be a stopping criterion [11].</p>
        <p>Let  be the initial dataset made out of training
samples with known dependent variable values. At first, the
tree will be made of only a root node 1 which represents
the full set of variables. The objective is to split the nodes
into two decision nodes until a terminal node is reached,
for example splitting  into  and , then splitting 
and  into further sub-nodes until a stopping criterion
is met [12].</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.4. Random Forest</title>
        <p>Random forest is a ML algorithm that constructs a
multitude of decision trees at training time. The main principle
of constructing a random forest is that it is formed by
combining solutions from binary decision trees made
using diverse subsets of the original dataset and subsets
containing randomly selected features from the feature
set.</p>
        <p>Constructing small decision trees that only have a few
features takes up only a little of the processing time,
hence the majority of such trees’ solutions can be
combined into a single strong classifier.</p>
        <p>Steps for constructing a random forest as
presented in [10] are as follows:
1. First, assume that the number of cases in the
training set is . Then, take a random sample of these
 cases, and use this sample as the training set
for constructing the tree.
2. If there are  input variables, specify a number
 &lt;  such that at each node,  random
variables out of the  can be selected. The best split
on these  is used to split the node.
3. Each tree is subsequently grown to the largest
extent possible and no pruning is needed.
4. Aggregate the predictions of the target trees to
predict new data.</p>
        <p>5. Finally, a decision is made by the majority rule.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.5. XGBoost</title>
        <p>XGBoost is a ML algorithm that implements frameworks
based on Gradient Boosted Decision Trees [13]. XGBoost
where W is a weight vector, namely, W =
{1, 2, . . . , }; X is a set of training data made of 
number of objects,  number of attributes and an
associated class label ; and  is a scalar constant.</p>
        <p>Ω( ) =  +</p>
      </sec>
      <sec id="sec-4-4">
        <title>4.6. Support Vector Machine</title>
        <p>
          The goal of SVM is to find the maximum separating
hyperplane that would have the maximum distance between
the nearest training data objects [
          <xref ref-type="bibr" rid="ref1">14</xref>
          ]. A separating
hyperplane can be written as:
        </p>
        <p>WX +  = 0,
(21)</p>
        <p>The distance between hyperplanes, denoted as
2/||W||, has to be maximal. Consequently, this means
that ||W|| (the Euclidean norm of the vector W) has to
be minimized. To simplify calculations, the Euclidean
norm ||W|| can be swapped for ||W||2/2. Thus, the
objective function for this optimization problem is defined
as:
min 1 2</p>
        <p>W, 2 ||W|| ,
(WX + ) ≥ 1,
 = 1, . . . , ,
where the constraint (23) ensures that all objects from
the training dataset will be positioned on the correct side
of the appropriate marginal hyperplane.</p>
        <p>The Lagrange multiplier strategy allows combining
these two conditions into one:
min max
W, ≥0
{︃ 1</p>
        <p>}︃
2 ||W||2 − ∑︁ [(WX − ) − 1] .</p>
        <p>=1
(24)</p>
        <p>Kernel functions are used when the training dataset
needs to be transformed into a higher-dimensional space
due to the data being linearly inseparable.</p>
        <p>(X, X ) = (X)(X ) , ∀X, X ∈ X. (25)
In this study, the Gaussian radial basis function kernel
was used:
(X, X ) = e−||X−X ||2 ,
(22)
(23)
where the  value is derived from the following equation:
1
 ≈ MED,=1,...,(||X − X ||).
(26)
Here MED is the median. Usually parameter  is found
through trial and error.</p>
      </sec>
      <sec id="sec-4-5">
        <title>4.7. -Nearest Neighbours</title>
        <p>
          The KNN algorithm is a method based on objects likeness
[
          <xref ref-type="bibr" rid="ref2">15</xref>
          ]. In other words, the principle is to find the
predeifned number ( ) of training samples closest to the new
point. In the case of regression, the relationship between
the explanatory variables and the continuous dependent
variable is approximated by estimating the average of the
observations, which together form the so-called
neighbourhood. Its size is determined using cross-validation
while minimizing the root mean square error.
        </p>
        <p>
          The Euclidean distance was used to calculate the
distance between objects.
The significance of regression in a model is usually
calculated using the coeficient of determination [
          <xref ref-type="bibr" rid="ref3">16</xref>
          ]:
12... =
√︃
        </p>
        <p>2 ,
1 − 2
where 2 is a residual dispersion from a forecast
model, 2 – dispersion of .</p>
        <p>
          However, the adjusted coeficient of determination
2 is better suited for comparing regression models
as it avoids the inaccuracy, caused by numerous factors
in the coeficient of determination [
          <xref ref-type="bibr" rid="ref3">16</xref>
          ]:
        </p>
        <p>2 = 1 − (1 − 2) ·  − − −1 1 ,
where  is the number of observations available for
analysis,  is the number of variables.</p>
        <p>
          Moreover, statisticians are used to measuring accuracy
by computing mean square error (MSE), or its square
root conventionally abbreviated by RMSE (for root mean
square error). The latter is in the same units as the
measured variable and so is a better descriptive statistic.
Moreover, it is the most popular evaluation metric used
in regression problems. RMSE follows an assumption
that errors are unbiased and follow a normal distribution.
RMSE metric is given by [
          <xref ref-type="bibr" rid="ref4">17</xref>
          ]:
⎯
        </p>
        <p>= ⎷⎸⎸ 1 ∑︁( − )2,
=1
(29)
where  are the observations,  are the predicted
values of a variable.</p>
        <p>Moreover, the average magnitude of the forecast errors
can be measured by the mean absolute error:</p>
        <p>= 1 ∑︁ | − |. (30)</p>
        <p>
          =1
In this case, the direction of errors is not being considered
[
          <xref ref-type="bibr" rid="ref4">17</xref>
          ].
        </p>
        <p>There are various ways to improve models depending
on the technique involved. The most popular way is
to construct ensemble models. Once there are multiple
models that produce a score for a particular outcome,
they can be combined to produce ensemble scores. For
example, a new score can be calculated as the average
of two classifiers and then assess it as a further model.
Usually, the area under the curve improves for these
ensemble models.
(27)
(28)</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Results</title>
    </sec>
    <sec id="sec-6">
      <title>5.1. Driver ranking</title>
      <sec id="sec-6-1">
        <title>Selection of attribute weights. To begin the ranking</title>
        <p>procedure, first attribute weights had to be established.</p>
        <p>Since the drivers must be ranked in compliance with at- Table 3
tribute priorities dictated by a company, their importance Score distribution with diferent weight sets
was evaluated by an expert on a scale from 1 to 10. This is
presented in Table 2. Thus, we get the first set of weights:
where 1 = 0.09 is the weight of Free rolling distance,
2 = 0.19 is the weight of Engine overloaded distance,
Excess idling and Overspeeding time, 3 = 0.13 is the
weight of Highest gear distance, 4 = 0.15 is the weight
of Extreme braking events, 5 = 0.07 is the weight of
Harsh braking events.</p>
        <p>However, since the importance of attributes can be
biased, a baseline weight model was also tested.</p>
        <p>Results of ranking methods. Ranking of the drivers
was performed using the generated sets of weights.
Criterion values, computed by TOPSIS () and VIKOR ()
methods, were used to rank the drivers. However, in
many instances the diference between two values of the
same criteria had been minute, hence the values were
grouped. Values were grouped using a ten-point system.</p>
        <p>Considering the results of diferent ranking methods
presented in Table 3, the method with the most
logical ranking of the drivers was confirmed to be with the
VIKOR method. In addition, the first weight set should be
used when creating a dataset for the forecasting models,
since it would comply with the attribute priorities from
the company and no significant diference was observed
between the two tested sets of weights.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>5.2. Forecasting models</title>
      <p>For improving the forecast of ETA, it was enough to
forecast deviation from planned duration, because this
variable had already been computed by routing service.</p>
      <p>In that case, the goal was to forecast the deviation from
planned tour duration. Overall five ML methods were
tested: decision tree, random forest, XGBoost, support
vector machine (SVM) and -Nearest neighbours (KNN).</p>
      <p>Quantitative variables were normalized using min-max
normalization, while the categorical variables were
transformed and added to the models by replacing them with
binary dummy variables.</p>
      <p>Therefore, when applying the random selection and
assignment of the set indices to the test and training sets,
75% of the dataset was assigned to the training sample and
25% to the test sample. Cross-validation was used for the
selection of optimal parameters in all five models. During
this procedure in the regression models, the sample data
was divided into 10 groups.</p>
      <p>Further, the optimal parameters of all models are
determined:
load of data. Such partial selection of a sample ˆ – of the SVM model. The adjusted coeficient of
occurs once per each iteration. The maximum determination of this ensemble model resulted in 0.7672.
number of repetitions was determined to be 150. Another way to form an ensemble of models is by using
4. The -Nearest neighbour model. The  param- the weighted sum method. In this type of ensemble, the
eter for KNN model was established to be equal to prediction value of the better model (in this case the
4. This value was selected by changing the param- random forest model) is determined to have a weighting
eter value from 1 to 10 and determining which coeficient 1, that is less than 1, but not less than 0.5.
parameter has the smallest RMSE value. The Eu- However, the total amount of weights must be equal
clidean distance measure was used to estimate to one, therefore, the weighting coeficient of the other
the distance between the points. model (SVM model) 2 shall be greater than zero, but less
5. The SVM model. The number of support vectors than 0.5. Then, the new predicted value could then be
was established to be 1008 and the Gaussian radial obtained as follows:
basis function kernel was used.
(32)
(33)
(34)
The equation
must be met, thus (32) can be written as:
ˆ = 1 · ˆ + 2 · ˆ .</p>
      <p>1 = 1 − 2
ˆ = (1 − 2) · ˆ + 2 · ˆ .</p>
      <p>In order to find with which weight the adjusted
coefifcient of determination of the ensemble model obtains
the highest value, the value of 2 was being changed
from 0.01, 0.02, 0.03, and so on to 0.49. The experiment
resulted in a maximum value of 2 (0.7795) when 2
was equal to 0.2. The second way of constructing an
ensemble model resulted in a higher 2 than the first
method, hence, the second was more suitable.
Therefore, the expression of the final ensemble model was as
follows:
ˆ = 0.8 · ˆ + 0.2 · ˆ .</p>
      <p>(35)
The forecast graph of the created ensemble model is
presented in Fig. 1. Some outliers remained poorly predicted,
but the overall prediction is accurate. Metrics evaluating</p>
      <p>The accuracy of all five constructed models was
evaluated by predicting values of deviation from planned tour
time for unseen test set data. The adjusted coeficient of
determination 2 , RMSE, and MEA were calculated to
determine and compare the suitability of the prediction
models. The results are presented in Table 4. It can be
seen that for all three accuracy measures, the best results
for predicting test data were obtained using a random
forest model, where the mean absolute diference between
the predicted and actual values had been almost 684, the
square root of the average squared diferences between
the predicted and actual values had been 1 120.15, and the
adjusted coeficient of determination had been 77.57%.</p>
      <p>XGBoost model yielded quite similar results, where 2
had been lower by only 3.3%, the MAE error had been
higher by 93.24 units and the RMSE had been higher by
90.83. The worst prediction results were obtained using
KNN method, for which 2 had been less than 25%.</p>
      <p>A possibility to improve the models by forming an
ensemble model was observed, hence a decision was made
to try a combination of predictions from two models:
the random forest and SVM. Several combinations were
made. The first method estimated the average of the
forecasts of both models:
ˆ =
ˆ + ˆ ,
2</p>
      <p>(31)
where ˆ is the predicted value of the th observation
for the model ensemble, ˆ are the predicted values
of the th observation of the random forest model and
the created ensemble model are presented in Table 5. In
comparison to the results of individual models (Table 4),
higher accuracy could be observed in all three metrics of
the ensemble model. Nonetheless, the improvement in
accuracy was not significant: the adjusted coeficient of</p>
    </sec>
    <sec id="sec-8">
      <title>6. Conclusions</title>
    </sec>
    <sec id="sec-9">
      <title>Acknowledgments</title>
      <p>We thank Tadas Motieju¯nas, Darius Ežerskis, Viktorija
Naudužaitė, Donatas Šumyla and UAB Ruptela (https:
//www.ruptela.com/) for cooperation and useful insights.
determination was higher than the best individual model
by only 0.38%, the RMSE was lower by 1.87, and MAE
was lower by 0.47.
In this research, in order to improve the forecast of
contemporary ETA, the possibility to rank the drivers
based on their behaviour data and predict deviations from
planned arrival time using diferent ML methods were
analysed. For this purpose, a dataset consisting of vehicle
monitoring data was used for ranking the drivers with
TOPSIS and VIKOR methods. It was found that the results
of the VIKOR method with the company’s attribute
importance weight set produced the most suitable drivers’
scores. Then, these scores were used to supplement a
new dataset constructed for ML methods. Moreover, five
methods: decision tree, random forest, XGBoost, SVM,
KNN, were used to create the deviation from the planned
tour duration forecasting model. Finally, the ensemble
model based on the random forest and SVM resulted in
the most accurate results (2 = 77.95%).</p>
      <p>In the future, it is planned to continue the
construction of the improved ETA prediction model by including
real-world parameters that a vehicle takes into account
while driving a certain route. For example, the need to
stop for mandatory driving breaks or filling up would be
considered.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>J.</given-names>
            <surname>Han</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kamber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pei</surname>
          </string-name>
          , Data Mining:
          <article-title>Concepts and Techniques</article-title>
          ,
          <string-name>
            <surname>Second Edition</surname>
          </string-name>
          (The Morgan Kaufmann Series in
          <source>Data Management Systems)</source>
          , 2 ed., Morgan Kaufmann,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Murni</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Kosasih</surname>
            ,
            <given-names>A. F.</given-names>
          </string-name>
          <string-name>
            <surname>Oeoen</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Handhika</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <string-name>
            <surname>Sari</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Lestari</surname>
          </string-name>
          ,
          <article-title>Travel time estimation for destination in bali using knn-regression method with tensorlfow</article-title>
          ,
          <source>IOP Conference Series: Materials Science and Engineering</source>
          <volume>854</volume>
          (
          <year>2020</year>
          )
          <article-title>012061</article-title>
          . doi:
          <volume>10</volume>
          .1088/
          <fpage>1757</fpage>
          -899X/854/1/012061.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>H.</given-names>
            <surname>Altland</surname>
          </string-name>
          ,
          <article-title>Regression analysis: Statistical modeling of a response variable</article-title>
          ,
          <source>Technometrics</source>
          <volume>41</volume>
          (
          <year>2012</year>
          )
          <fpage>367</fpage>
          -
          <lpage>368</lpage>
          . doi:
          <volume>10</volume>
          .1080/00401706.
          <year>1999</year>
          .
          <volume>10485936</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>C.</given-names>
            <surname>Chatfield</surname>
          </string-name>
          , Time-series forecasting,
          <source>Significance</source>
          <volume>2</volume>
          (
          <year>2005</year>
          )
          <fpage>131</fpage>
          -
          <lpage>133</lpage>
          . URL: https://rss.onlinelibrary.wiley.com/doi/abs/10. 1111/j.1740-
          <fpage>9713</fpage>
          .
          <year>2005</year>
          .
          <volume>00117</volume>
          .x. doi:https://doi. org/10.1111/j.1740-
          <fpage>9713</fpage>
          .
          <year>2005</year>
          .
          <volume>00117</volume>
          .x.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>