<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Maximum Likelihood Estimation with Deep Learning for Multiple Sclerosis Progression Prediction</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Tsvetan Asamov</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Petar Ivanov</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anna Aksenova</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dimitar Taskov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Svetla Boytcheva</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Medical University - Sofia</institution>
          ,
          <country country="BG">Bulgaria</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Multiprofile Hospital for Active Treatment in Neurology and Psychiatry "St. Naum" - Sofia</institution>
          ,
          <country country="BG">Bulgaria</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Ontotext</institution>
          ,
          <country country="BG">Bulgaria</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We develop a maximum likelihood estimation approach for intelligent disease progression prediction. We use patients' covariates and employ a multi-layer perceptron to approximate the optimal distribution parameters for a given parametric family of probability distributions. As far as we know, this is the first time such a method has been applied to real multiple sclerosis data. Our numerical results indicate that the method can achieve AUROC scores exceeding 0.8.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;CLEF</kwd>
        <kwd>multiple sclerosis</kwd>
        <kwd>neurological disease</kwd>
        <kwd>maximum likelihood estimation</kwd>
        <kwd>deep learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>• Task 1: Predicting the risk of disease worsening - predicting the risk of worsening and
ranking subjects based on the risk scores. More specifically, the risk of worsening should</p>
      <p>be a value between 0 and 1 that reflects how early a patient experiences the worsening
event.
• Task 2: Predicting the cumulative probability of worsening - assigning cumulative
probability of worsening at diferent time windows, i.e. between years 0 and 2, 0 and 4, 0 and 6,
0 and 8, 0 and 10.</p>
      <p>In addition, for each task, we consider two diferent subtasks based on two alternative definitions
of worsening. Following clinical standards, worsening is defined on the basis of the Expanded
Disability Status Scale (EDSS):
• Subtask A: the patient crosses the EDSS ≥ 3 threshold at least twice within a one-year
interval.
• Subtask B: the first recorded EDSS value available in clinical records is defined as the
baseline, and worsening occurs according to the following rules:
– if the baseline is EDSS &lt; 1, then worsening occurs when an EDSS increase of 1.5
points is first observed.
– if the baseline is 1 ≤ EDSS &lt; 5.5, then worsening occurs when an EDSS increase
of 1 point is first observed.
– if the baseline is EDSS ≥ 5.5, then worsening occurs when an EDSS increase of 0.5
points is first observed.</p>
      <p>Finally, for each subtask, we are given a separate dataset consisting of general patient
information, as well as a series of observations over 2.5 years.</p>
      <p>The paper is organized as follows: Section 2 introduces related works; Section 3 describes
our approach; Section 4 explains our experimental setup; Section 5 discusses our main findings;
ifnally, Section 6 draws some conclusions and outlooks for future work.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        Applications in various domains involve time-to-event modelling problems in the presence of
censoring. Some examples include healthcare [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ], reliability [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ], finance [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], and other
ifelds.
      </p>
      <p>
        The Cox proportional hazards model [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] is commonly used in survival analysis. However,
it employs a log-linear function to predict the outcome variable from the covariates. In this
sense, it may not be suitable to properly predict a multiple sclerosis patient outcome without
extensive feature engineering.
      </p>
      <p>
        Recently, deep learning has been applied to extend the Cox proportional hazard model. More
specifically, the DeepSurv model [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] employs a multi-layered perceptron to replace the log-linear
function of the Cox proportional hazards model. Similar to the Cox proportional hazard model,
DeepSurv assumes a constant baseline hazard.
      </p>
      <p>
        In addition, DeepHit [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] has proposed discretizing the space of event times, and using a deep
neural network to learn the distribution of survival times. Further, DeepHit does not make any
assumptions about the underlying stochastic process, and has the ability to handle competing
risks.
      </p>
      <p>
        Random survival forests [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] is a non-parametric method that constitutes an extension of
the random forest method approach [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. More specifically, random survival forests learn an
ensemble of trees for the analysis of right-censored survival data.
      </p>
      <p>
        The idea of using a deep learning framework to estimate probability distribution parameters
for maximum likelihood estimation for right-censored data was initially introduced by Nagpal
et al. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] as a part of their deep survival machines (DSM) framework. DSM employs a mixture
of individual parametric survival distributions to fit a set of right-censored survival data. The
work was further extended to recurrent deep survival machines (RDSM) by Nagpal et al. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]
with the introduction of recurrent neural networks in place of the learnt representations. Unlike
DSM and RDSM, we do not use a mixture of parametric distributions but rather focus on fitting
the parameters of a single parametric probability distribution.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology</title>
      <p>In this section we describe the methodology we have adopted.
3.1. Data
We assume that we are given right-censored data which consists of a set of  triples {(x, ,  )}=1.
For each  = 1, . . . , , the real vector x ∈ R denotes the features of the − th entry. Further,
 is either the censoring time, or the time an event occurred. In addition,   is an indicator
variable taking a value of 0 if  is the censoring time, and a value of 1, if  is the time at which
an event took place. It is assumed that for each  = 1, . . . ,  either censoring occurs, or we
observe the event but not both.</p>
      <sec id="sec-3-1">
        <title>3.2. Maximum Likelihood Formulation</title>
        <p>The method of maximum likelihood can be adapted to various applied problems. In this section,
we develop a maximum likelihood estimation approach for intelligent disease progression
prediction. Given a set of right-censored patient data {(x, ,  )}=1, we can assume independence
among patients, and thus define the likelihood function of the observed data as follows:
( ) =
∏︁  (| ) ∏︁ (1 −  (| ))
, =1 , =0
(1)
where
•  (| ) is the probability density function evaluated at time  for distribution parameters
 .
•  (| ) is the cumulative probability density function evaluated at time  for distribution
parameters  .</p>
        <p>
          Please note that we do not assume that the patient data is identically distributed. On the
contrary, we consider diferent distribution parameters   for each patient  = 1, . . . , . Further,
please note that if we knew the distribution functions  and  , and if we could estimate the
distribution parameters   for a previously unseen patient , then we would also be able to
estimate the patient’s probability of worsening over a given time period. In order to achieve
that, we would like to use a parametric family of distributions in the above formulation (1).
Hence, we would need to choose a probability distribution with support over the positive real
line. In this work, we focus on the Weibull distribution [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. It is a continuous probability
distribution that has closed form expressions for both its probability density function  (), and
cumulative density function  ():
connected layers as the function mapping feature inputs x to estimated distribution parameters
  = ( ,  ):
Ψ(A, b, x) :=  (A  (A− 1 (. . . A3 (A2 (A1x + b1) + b2) + b3 . . . ) + b− 1) + b )
where A and b are collection of respectively real matrices A,  = 1, . . . ,  , and real vectors
b,  = 1, . . . ,  , and  is an activation function such as the rectified linear unit. Thus, we
can write the problem of maximizing the likelihood function as follows:
max ∏︁  (| ,  ) ∏︁ (1 −  (| ,  ))
( ,  ) = Ψ(A, b, x),  = 1, . . . , 
  &gt; 0,  = 1, . . . , 
  &gt; 0,  = 1, . . . , 
Ψ(A, b, x) :=  (A (A− 1 (. . . A3 (A2 (A1x + b1) + b2) + b3 . . . ) + b− 1) + b)
In order to improve computational stability and avoid numerical issues, we choose to
instead maximize the log-likelihood function. As the logarithm function is monotonic, we can
(2)
(3)
(4)
where
equivalently write problem (4) as follows:
max ∑︁ log( (| ,  )) + ∑︁ log(1 −  (| ,  ))
Ψ(A, b, x) :=  (A (A− 1 (. . . A3 (A2 (A1x + b1) + b2) + b3 . . . ) + b− 1) + b)
(5)
        </p>
        <p>Please note that even in the case when the activation function  is the identity map, the
proposed formulation is neither convex, nor concave. Thus, in general, we cannot find a global
optimal solution to the proposed model using gradient-type methods. However, sub-optimal
solutions of reasonable quality can be found, as indicated in the next sections.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Experimental Setup</title>
      <p>In this section, we describe our experimental setup.</p>
      <sec id="sec-4-1">
        <title>4.1. Implementation</title>
        <p>
          The training pipeline is implemented in the julia programming language [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]. Further, we use
the Knet library [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] for our neural network implementation. Our code is available on bitbucket.
We run training and testing steps on a Core i7 CPU with 16 GB of RAM. What is more, we
use a validation set to determine the values of our hyper-parameters. Moreover, we use the
Adam optimizer [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] with a learning rate of 0.00001. While the value may seem somewhat
small, we found it to work well in practice. In addition, we do not split the data into batches but
rather use the entire training dataset for each step of Adam. Furthermore, we apply dropout
regularization [
          <xref ref-type="bibr" rid="ref18 ref19">18, 19</xref>
          ] to the input layer with a dropout rate chosen among 0.01 and 0.2. We
choose the number of hidden units in the neural network among 100 and 200. Finally, the
number of training epochs (which also equals the number of training steps of Adam) is chosen
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Datasets</title>
        <p>We are given two diferent datasets, one for subtask A, and one for subtask B. In this way, task
1A and task 2A both use the first dataset, while task 1B and task 2B both use the second dataset.
Each dataset contains 2.5 years of patient visits. In addition, the occurrence of the worsening
event, as well as the time of its occurrence are also given.</p>
        <p>The training and testing data of both datasets (subtask A and subtask B) is partitioned into
static patient data and dynamic patient data. Furthermore, the dynamic patient data includes
information on relapses, EDSS scores, evoked potentials, MRI results and multiple sclerosis
course.</p>
        <p>
          The training dataset for subtask A includes the following: 441 patients, 481 relapses, 2,661
EDSS scores, 1,211 evoked potentials, 960 MRIs, and 310 multiple sclerosis courses. The training
dataset for subtask B includes the following: 511 patients, 553 relapses, 3,069 EDSS scores, 1,522
evoked potentials, 966 MRIs, and 325 multiple sclerosis courses. In addition, the testing dataset
of subtask A includes the following: 111 patients, 95 relapses, 675 EDSS scores, 278 evoked
potentials, 236 MRIs, and 68 multiple sclerosis courses. And the testing dataset of subtask
B includes the following: 129 patients, 125 relapses, 813 EDSS scores, 299 evoked potentials,
266 MRIs, and 75 multiple sclerosis courses. For a detailed description of the datasets and the
evaluation measures, please see the overview papers by the CLEF challenge organizers [
          <xref ref-type="bibr" rid="ref20 ref21">20, 21</xref>
          ].
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Results</title>
      <p>
        The challenge objectives consist of the following:
• Task 1 - predicting the risk of disease worsening
• Task 2 - predicting the cumulative probability of worsening
In order to handle both tasks, we use the available training data to build a model and estimate
a maximum likelihood distribution for each patient given the patient’s covariates (features).
Ideally, for task 1 we would have preferred to use coherent risk measures [22, 23] to estimate
the risk of disease worsening from the patients’ distributions. However, in order to meet the
requirement that risk values are in the range of [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ] we decide to use a cumulative probability
estimate instead of coherent risk measures. The performance of the submitted models is reported
in Figures 1-11. Please note that the name of each model indicates the model’s parameters’
values. For instance, the first model in Figure 1 is named “T1b.0.2.1.0e-5.5000.200”, indicating
that it is a Task 1B model with the following parameters:
• A dropout rate of 0.2 used in the input layer.
• A learning rate of 1.0e-5 used by the Adam optimizer.
• The model is trained for 5000 epochs, i.e. the Adam optimizer performs 5000 steps.
• The number of hidden units is set to 200.
      </p>
      <p>
        We can see that the highest Harrell’s concordance index values fall in the interval [0.6, 0.65].
Ideally, we would like to improve those results in the future. One way we could do that is by
incorporating event ordering into the model training procedure. Another approach we could
try is scaling down classical coherent risk measures to fit into the [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ] interval for the given
patient data. In Figures 2-6 we can see that the highest AUROC exceeds 0.8. In the future, we
can further improve those results by better model optimization. In addition, in Figures 7-11 we
can find the ratio of observed to expected events for all submitted models.
      </p>
      <p>The model with the highest AUROC score T2a.0.01.1.0e-5.10000.100.adj has a couple of
aspects that distinguish it from the rest of the models. First, for each patient dataframe (static
patient data, relapses, EDSS scores, evoked potentials, MRI results and multiple sclerosis course)
it explicitly takes into account the length of the dataframe. And second, it normalizes the
age_at_onset variable using division by fifty. This suggests that current results can be further
improved by additional data pre-processing.</p>
      <p>Finally, in Table 1 we present illustrations of probability density functions for three randomly
chosen patients from the test set for the 2a.0.01.1.0e-5.10000.100.adj model. Please note that the
risk and the probabilities of worsening for each patient depend entirely on the computed values
of the distribution parameters   and  .</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion and Future Work</title>
      <p>The development of predictive models of the disease is a step forward towards better clinical
assessment and an individualized therapeutic approach for multiple sclerosis patients.</p>
      <p>
        In this paper we have developed a maximum likelihood estimation approach for intelligent
disease progression prediction. To the best of our knowledge, this is the first instance of such
a method being applied to real multiple sclerosis data. Our numerical results indicate that
the method can achieve AUROC scores exceeding 0.8. In the future we can explore several
directions of further research. First, we can attempt to incorporate event ordering into the
training procedure in order to improve Harrell’s concordance index scores. Further, we can
attempt to apply (scaled-down) coherent risk measures in order to obtain risk estimates in the
[
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ] interval. Finally, we may also look into improving the quality of the numerical solution
with the use of second-order optimization methods such as K-FAC [24] or L-BFGS [25].
[22] P. Artzner, F. Delbaen, J.-M. Eber, D. Heath, Coherent measures of risk, Mathematical
ifnance 9 (1999) 203–228.
[23] T. Asamov, A. Ruszczyński, Time-consistent approximations of risk-averse multistage
stochastic optimization problems, Mathematical Programming 153 (2015) 459–493.
[24] J. Martens, R. Grosse, Optimizing neural networks with kronecker-factored approximate
curvature, in: International conference on machine learning, PMLR, 2015, pp. 2408–2417.
[25] J. NOCEDAL, J. W. STEPHEN, SPRINGER SERIES IN OPERATONS RESEARCH
NUMERI
      </p>
      <p>CAL OPTIMIZATION., Springer, 2006.</p>
    </sec>
    <sec id="sec-7">
      <title>A. Numerical Results</title>
      <p />
      <p>Probability Density Function

,

|


(

)


,

|


(

)


,

|


(

T2a.0.01.1.0e-5.10000.100.adj model.</p>
      <p>Probability density functions for three randomly selected patients from the testing set for the



T1b.0.2.1.0e-5.5000.200
T1b.0.2.1.0e-5.5000.100
T1b.0.2.1.0e-5.10000.200
T1b.0.2.1.0e-5.10000.100
T1a.0.2.1.0e-5.5000.200
T1a.0.2.1.0e-5.5000.100
T1a.0.2.1.0e-5.10000.200</p>
      <p>T1a.0.2.1.0e-5.10000.100
T1a.0.01.1.0e-5.10000.100.ajd
0.3
0.35
0.4
0.45
0.5
0.55
0.6
0.65
0.7
0.75
0.8
T2b.0.2.1.0e-5.5000.100
T2b.0.2.1.0e-5.10000.200
T2b.0.2.1.0e-5.10000.100
T2a.0.2.1.0e-5.5000.200
T2a.0.2.1.0e-5.5000.100
T2a.0.2.1.0e-5.10000.200</p>
      <p>T2a.0.2.1.0e-5.10000.100
T2a.0.01.1.0e-5.10000.100.adj
0.3
0.4
0.5
0.6
0.7
0.8
0.9
1
T2b.0.2.1.0e-5.5000.100
T2b.0.2.1.0e-5.10000.200
T2b.0.2.1.0e-5.10000.100
T2a.0.2.1.0e-5.5000.200
T2a.0.2.1.0e-5.5000.100
T2a.0.2.1.0e-5.10000.200</p>
      <p>T2a.0.2.1.0e-5.10000.100
T2a.0.01.1.0e-5.10000.100.adj
0.3
0.4
0.5
0.6
0.7
0.8
0.9
1
T2b.0.2.1.0e-5.5000.100
T2b.0.2.1.0e-5.10000.200
T2b.0.2.1.0e-5.10000.100
T2a.0.2.1.0e-5.5000.200
T2a.0.2.1.0e-5.5000.100
T2a.0.2.1.0e-5.10000.200</p>
      <p>T2a.0.2.1.0e-5.10000.100
T2a.0.01.1.0e-5.10000.100.adj
0.3
0.4
0.5
0.6
0.7
0.8
0.9
T2b.0.2.1.0e-5.5000.100
T2b.0.2.1.0e-5.10000.200
T2b.0.2.1.0e-5.10000.100
T2a.0.2.1.0e-5.5000.200
T2a.0.2.1.0e-5.5000.100
T2a.0.2.1.0e-5.10000.200</p>
      <p>T2a.0.2.1.0e-5.10000.100
T2a.0.01.1.0e-5.10000.100.adj
0.3
0.4
0.5
0.6
0.7
0.8
0.9
T2b.0.2.1.0e-5.5000.100
T2b.0.2.1.0e-5.10000.200
T2b.0.2.1.0e-5.10000.100
T2a.0.2.1.0e-5.5000.200
T2a.0.2.1.0e-5.5000.100
T2a.0.2.1.0e-5.10000.200</p>
      <p>T2a.0.2.1.0e-5.10000.100
T2a.0.01.1.0e-5.10000.100.adj
0.2
0.3
0.4
0.5
0.6
0.7
0.8
0.9
T2b.0.2.1.0e-5.5000.100
T2b.0.2.1.0e-5.10000.200
T2b.0.2.1.0e-5.10000.100
T2a.0.2.1.0e-5.5000.200
T2a.0.2.1.0e-5.5000.100
T2a.0.2.1.0e-5.10000.200</p>
      <p>T2a.0.2.1.0e-5.10000.100
T2a.0.01.1.0e-5.10000.100.adj
− 0.2
0
0.2
0.4
0.6
0.8
1
1.2
T2b.0.2.1.0e-5.5000.100
T2b.0.2.1.0e-5.10000.200
T2b.0.2.1.0e-5.10000.100
T2a.0.2.1.0e-5.5000.200
T2a.0.2.1.0e-5.5000.100
T2a.0.2.1.0e-5.10000.200</p>
      <p>T2a.0.2.1.0e-5.10000.100
T2a.0.01.1.0e-5.10000.100.adj
0
0.2
0.4
0.6
0.8
1
1.2
T2b.0.2.1.0e-5.5000.100
T2b.0.2.1.0e-5.10000.200
T2b.0.2.1.0e-5.10000.100
T2a.0.2.1.0e-5.5000.200
T2a.0.2.1.0e-5.5000.100
T2a.0.2.1.0e-5.10000.200</p>
      <p>T2a.0.2.1.0e-5.10000.100
T2a.0.01.1.0e-5.10000.100.adj
0
0.2
0.4
0.6
0.8
1
T2b.0.2.1.0e-5.5000.100
T2b.0.2.1.0e-5.10000.200
T2b.0.2.1.0e-5.10000.100
T2a.0.2.1.0e-5.5000.200
T2a.0.2.1.0e-5.5000.100
T2a.0.2.1.0e-5.10000.200</p>
      <p>T2a.0.2.1.0e-5.10000.100
T2a.0.01.1.0e-5.10000.100.adj
0
0.1
0.2
0.3
0.4
0.5
0.6
0.7
0.8
0.9
1
1.1
0
0.1
0.2
0.3
0.4
0.5
0.6
0.7
0.8
0.9</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <source>[1] Participation guidelines of idpp@ clef</source>
          <year>2023</year>
          ,
          <year>2023</year>
          . URL: https://brainteaser.dei.unipd.it/ challenges/idpp2023/.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D. W.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kwon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Nam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.-H.</given-names>
            <surname>Cha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. J.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <article-title>Deep learning-based survival prediction of oral cancer patients</article-title>
          ,
          <source>Scientific reports 9</source>
          (
          <year>2019</year>
          )
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O. O</given-names>
            <surname>'Donnell</surname>
          </string-name>
          ,
          <string-name>
            <surname>O.</surname>
          </string-name>
          <article-title>O'Donnell, Econometric analysis of health data</article-title>
          ,
          <source>Wiley Online Library</source>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>R. E.</given-names>
            <surname>Barlow</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Proschan</surname>
          </string-name>
          ,
          <article-title>Statistical theory of reliability and life testing: probability models</article-title>
          ,
          <source>Technical Report</source>
          , Florida State Univ Tallahassee,
          <year>1975</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hollander</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. A.</given-names>
            <surname>Peña</surname>
          </string-name>
          ,
          <article-title>Dynamic reliability models with conditional proportional hazards</article-title>
          ,
          <source>Lifetime Data Analysis</source>
          <volume>1</volume>
          (
          <year>1995</year>
          )
          <fpage>377</fpage>
          -
          <lpage>401</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Stepanova</surname>
          </string-name>
          , L. Thomas,
          <article-title>Survival analysis methods for personal loan data</article-title>
          ,
          <source>Operations Research</source>
          <volume>50</volume>
          (
          <year>2002</year>
          )
          <fpage>277</fpage>
          -
          <lpage>289</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D. R.</given-names>
            <surname>Cox</surname>
          </string-name>
          ,
          <article-title>Regression models and life-tables</article-title>
          ,
          <source>Journal of the Royal Statistical Society: Series B (Methodological) 34</source>
          (
          <year>1972</year>
          )
          <fpage>187</fpage>
          -
          <lpage>202</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Katzman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.</given-names>
            <surname>Shaham</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Cloninger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bates</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kluger</surname>
          </string-name>
          ,
          <article-title>Deepsurv: personalized treatment recommender system using a cox proportional hazards deep neural network</article-title>
          ,
          <source>BMC medical research methodology</source>
          <volume>18</volume>
          (
          <year>2018</year>
          )
          <fpage>1</fpage>
          -
          <lpage>12</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>C.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Zame</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yoon</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. Van Der Schaar</surname>
          </string-name>
          ,
          <article-title>Deephit: A deep learning approach to survival analysis with competing risks</article-title>
          ,
          <source>in: Proceedings of the AAAI conference on artificial intelligence</source>
          , volume
          <volume>32</volume>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>H.</given-names>
            <surname>Ishwaran</surname>
          </string-name>
          , U. B.
          <string-name>
            <surname>Kogalur</surname>
            ,
            <given-names>E. H.</given-names>
          </string-name>
          <string-name>
            <surname>Blackstone</surname>
            ,
            <given-names>M. S.</given-names>
          </string-name>
          <string-name>
            <surname>Lauer</surname>
          </string-name>
          , Random survival forests (
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>L.</given-names>
            <surname>Breiman</surname>
          </string-name>
          , Random forests,
          <source>Machine learning 45</source>
          (
          <year>2001</year>
          )
          <fpage>5</fpage>
          -
          <lpage>32</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>C.</given-names>
            <surname>Nagpal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dubrawski</surname>
          </string-name>
          ,
          <article-title>Deep survival machines: Fully parametric survival regression and representation learning for censored data with competing risks</article-title>
          ,
          <source>IEEE Journal of Biomedical and Health Informatics</source>
          <volume>25</volume>
          (
          <year>2021</year>
          )
          <fpage>3163</fpage>
          -
          <lpage>3175</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>C.</given-names>
            <surname>Nagpal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Jeanselme</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dubrawski</surname>
          </string-name>
          ,
          <article-title>Deep parametric time-to-event regression with time-varying covariates, in: Survival Prediction-Algorithms, Challenges and Applications</article-title>
          , PMLR,
          <year>2021</year>
          , pp.
          <fpage>184</fpage>
          -
          <lpage>193</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>A.</given-names>
            <surname>Kızılersü</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kreer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. W.</given-names>
            <surname>Thomas</surname>
          </string-name>
          , The weibull distribution,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bezanson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Karpinski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. B.</given-names>
            <surname>Shah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Edelman</surname>
          </string-name>
          ,
          <article-title>Julia: A fast dynamic language for technical computing</article-title>
          ,
          <source>arXiv preprint arXiv:1209.5145</source>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>D.</given-names>
            <surname>Yuret</surname>
          </string-name>
          ,
          <article-title>Knet: beginning deep learning with 100 lines of julia</article-title>
          ,
          <source>in: Machine Learning Systems Workshop at NIPS</source>
          , volume
          <year>2016</year>
          ,
          <year>2016</year>
          , p.
          <fpage>5</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>D. P.</given-names>
            <surname>Kingma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ba</surname>
          </string-name>
          ,
          <article-title>Adam: A method for stochastic optimization</article-title>
          ,
          <source>arXiv preprint arXiv:1412.6980</source>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>G. E.</given-names>
            <surname>Hinton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Srivastava</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Krizhevsky</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Sutskever</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. R.</given-names>
            <surname>Salakhutdinov</surname>
          </string-name>
          ,
          <article-title>Improving neural networks by preventing co-adaptation of feature detectors</article-title>
          ,
          <source>arXiv preprint arXiv:1207.0580</source>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>P.</given-names>
            <surname>Baldi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. J.</given-names>
            <surname>Sadowski</surname>
          </string-name>
          , Understanding dropout,
          <source>Advances in neural information processing systems</source>
          <volume>26</volume>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>G.</given-names>
            <surname>Faggioli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Guazzo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Marchesin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Menotti</surname>
          </string-name>
          , I. Trescato,
          <string-name>
            <given-names>H.</given-names>
            <surname>Aidos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Bergamaschi</surname>
          </string-name>
          , G. Birolo,
          <string-name>
            <given-names>P.</given-names>
            <surname>Cavalla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Chiò</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dagliati</surname>
          </string-name>
          , M. de Carvalho,
          <string-name>
            <given-names>G. M.</given-names>
            <surname>Di Nunzio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Fariselli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>García Dominguez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gromicho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Longato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. C.</given-names>
            <surname>Madeira</surname>
          </string-name>
          , U. Manera,
          <string-name>
            <given-names>G.</given-names>
            <surname>Silvello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Tavazzi</surname>
          </string-name>
          , E. Tavazzi,
          <string-name>
            <given-names>M.</given-names>
            <surname>Vettoretti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. Di</given-names>
            <surname>Camillo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <source>Intelligent Disease Progression Prediction: Overview of iDPP@CLEF</source>
          <year>2023</year>
          , in: A.
          <string-name>
            <surname>Arampatzis</surname>
            , E. Kanoulas,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Tsikrika</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Vrochidis</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Giachanou</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Aliannejadi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Vlachos</surname>
          </string-name>
          , G. Faggioli, N. Ferro (Eds.),
          <source>Experimental IR Meets Multilinguality, Multimodality, and Interaction. Proceedings of the Fourteenth International Conference of the CLEF Association (CLEF 2023), Lecture Notes in Computer Science (LNCS)</source>
          , Springer, Heidelberg, Germany,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>G.</given-names>
            <surname>Faggioli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Guazzo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Marchesin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Menotti</surname>
          </string-name>
          , I. Trescato,
          <string-name>
            <given-names>H.</given-names>
            <surname>Aidos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Bergamaschi</surname>
          </string-name>
          , G. Birolo,
          <string-name>
            <given-names>P.</given-names>
            <surname>Cavalla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Chiò</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dagliati</surname>
          </string-name>
          , M. de Carvalho,
          <string-name>
            <given-names>G. M.</given-names>
            <surname>Di Nunzio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Fariselli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>García Dominguez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gromicho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Longato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. C.</given-names>
            <surname>Madeira</surname>
          </string-name>
          , U. Manera,
          <string-name>
            <given-names>G.</given-names>
            <surname>Silvello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Tavazzi</surname>
          </string-name>
          , E. Tavazzi,
          <string-name>
            <given-names>M.</given-names>
            <surname>Vettoretti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. Di</given-names>
            <surname>Camillo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <source>Overview of iDPP@CLEF</source>
          <year>2023</year>
          :
          <article-title>The Intelligent Disease Progression Prediction Challenge</article-title>
          , in: M.
          <string-name>
            <surname>Aliannejadi</surname>
            , G. Faggioli,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Ferro</surname>
          </string-name>
          , M. Vlachos (Eds.),
          <source>CLEF 2023 Working Notes, CEUR Workshop Proceedings (CEUR-WS.org)</source>
          ,
          <source>ISSN 1613-0073</source>
          .,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>