<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>HC@AIxIA</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Towards a Transportable Causal Network Model Based on Observational Healthcare Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alice Bernasconi</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alessio Zanga</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Peter J.F. Lucas</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marco Scutari</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fabio Stella</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Data Science and Advanced Analytics, F. Hofmann - La Roche Ltd</institution>
          ,
          <addr-line>Basel</addr-line>
          ,
          <country country="CH">Switzerland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Evaluative Epidemiology Unit, Department of Epidemiology and Data Science, Fondazione IRCCS Istituto Nazionale dei Tumori</institution>
          ,
          <addr-line>Milan</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Models and Algorithms for Data and Text Mining Laboratory (MADLab), University of Milano - Bicocca</institution>
          ,
          <addr-line>Milan</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of Twente</institution>
          ,
          <addr-line>Enschede</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>2</volume>
      <fpage>0000</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>Over the last decades, many prognostic models based on artificial intelligence techniques have been used to provide detailed predictions in healthcare. Unfortunately, the real-world observational data used to train and validate these models are almost always afected by biases that can strongly impact the outcomes validity: two examples are values missing not-at-random and selection bias. Addressing them is a key element in achieving transportability and in studying the causal relationships that are critical in clinical decision making, going beyond simpler statistical approaches based on probabilistic association. In this context, we propose a novel approach that combines selection diagrams, missingness graphs, causal discovery and prior knowledge into a single graphical model to estimate the cardiovascular risk of adolescent and young females who survived breast cancer. We learn this model from data comprising two diferent cohorts of patients. The resulting causal network model is validated by expert clinicians in terms of risk assessment, accuracy and explainability, and provides a prognostic model that outperforms competing machine learning methods.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Causal discovery</kwd>
        <kwd>Causal networks</kwd>
        <kwd>Transportability</kwd>
        <kwd>Missing values</kwd>
        <kwd>Selection bias</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>An important application of artificial intelligence in healthcare is predicting a disease trajectory
conditional on the patient’s history and a projected treatment strategy. In this context, we
aim to develop a model that generalizes well: it must provide accurate predictions not only
on the study cohort (the population from which the data have been collected) but also on the
target cohort, the general population that it is designed for. The model’s validity, accuracy and
usefulness in clinical practice are then of paramount relevance.</p>
      <p>
        In order to establish validity in a rigorous causal framework, we must clearly identify the
scope of the study the data come from, as well as the applicability and limitations of the theory
we are studying. However, many other issues can impact model validity. Firstly, study cohorts
difer from target cohorts in both randomized controlled trials and clinical observational studies
due to their inclusion and exclusion criteria. Secondly, missingness patterns may introduce bias
in the model when the data are missing not at random (MNAR). Thirdly, information can be
fragmented across diferent data sets ( data sparsity) and experts (domain knowledge). Therefore,
it is important to understand the data selection mechanism and the merging process to ensure
the model validity and generalisability [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>Our main contributions in this context are:
• Developing the first causal network model for estimating the risk of cardiovascular
diseases in adolescent and young adults that have been treated and survived breast cancer.
• Developing and adapting a methodological approach to deal with data afected by selection
bias and sufering by MNAR.
• A thoughtful review of the model we developed and of its implications by domain experts
(that is, expert physicians).</p>
      <p>The rest of the paper is organized as follows. We start by introducing our research question
and the topic of the study: cardio-oncology, the subfield of cardiology that aims at significantly
reducing cardiovascular morbidity and mortality and at improving the quality of life in cancer
survivors, here young patients (Section 2). We complete this introduction by providing an
overview of the related work available in the literature (Section 3). We then describe the
data (Section 4.1), the domain knowledge (Section 4.2) and the causal networks methodology
(Section 5) that we use to design and develop our model. Finally, we describe our findings
(Section 6), we contrast them with the available clinical and epidemiological knowledge and we
assess model performance against that of other commonly-used machine learning approaches.
We complete the paper by summarising our conclusions and presenting future work (Section 7).</p>
    </sec>
    <sec id="sec-2">
      <title>2. Young Adult Breast Cancer Survivors</title>
      <p>
        The increasing number and life expectancy of cancer survivors in the last decades has
highlighted the acute and long-term cardio-toxic efects of diferent cancer therapies. The most
common cancer among adolescent and young-adult females (AYAs; 15–39 years at the first cancer
diagnosis) is breast cancer (BC), which is a rare disease with unique genetic and biological
features in this age class. AYA female patients are more likely to be genetically susceptible to
more aggressive form of BC than older females, and their diagnosis is often delayed because of
the lack of BC screening policies for their age group. While lower than in older females with
BC, AYAs BC survival rates are high and continuously increasing, thus making such patients
likely to be long-term survivors [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Moreover, BC is biologically more aggressive in AYAs than
in older women with BC, and requires more aggressive combined neoadjuvant (pre-surgery)
and adjuvant (post-surgery) treatments [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>While cardiovascular diseases (CVDs) are well characterized in some groups of cancer
survivors, only few studies describe CVDs in AYAs because of their rarity in young patients.
Therefore, further evidence is needed to understand and prevent CVDs in AYAs with BC for
helping clinicians to plan personalized and efective follow-up guidelines for these patients.
Thus, with this work we are interested in answering to the question: “To what extent is it possible
to predict and explain individual susceptibility to cardiotoxicity in AYA with BC?”,</p>
    </sec>
    <sec id="sec-3">
      <title>3. Related works</title>
      <p>
        Machine learning (ML) techniques have achieved remarkable results in predicting CVDs in
cancer patients [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Unfortunately, they are still not part of the clinical practice: cardio-oncologists
rely on older, less accurate cardiovascular risk stratification tools such as the Framingham score
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The reasons for their reluctance to use ML methods are as follows:
• Data Availability. The best-performing ML models such as XGBoost and neural networks
rely on biomarkers, laboratory tests, electrocardiograms, echocardiograms,
computerized tomography, and cardiac magnetic resonance imaging data [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. However,
cardiooncologists rarely have immediate and complete access to such data in their everyday
practice.
• Computational Burden and Data Scarcity. ML models are known to be
computationally intensive and to require large amounts of data to achieve optimal performance. This
limits their use in most of healthcare systems, even in rich countries, because healthcare
data sufer from quality issues and are typically scarce. They are also often biased and
afected by missing-not-at-random patterns [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] in ways ML models do not account for.
• High Skills, Knowledge and Expertise. Current state-of-the-art ML methods require
specialized skills, knowledge and expertise to train, validate and deploy especially when
combining diferent types of data, which is almost always the case in healthcare [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
• Lack of Interpretability. The current state-of-the-art ML models are still dificult to
interpret despite recent progress from Explainable AI in addressing this issue [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
• Lack of generalizability. ML models have been applied successfully to childhood cancer
survivors [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], where using a limited number of variables, a relatively simple model, and
an easy-to-understand user-interface ensured that they were incorporated into clinical
practice. However, cancer is a collection of very complex and heterogeneous diseases,
making ML models unlikely to transfer successfully to other cancer survivors cohorts.
• Correlation vs Causation. ML models are typically developed using observational
data, and in general they can achieve excellent predictive accuracy by leveraging the
association between the response and the observed variables. However, clinicians operate
in a causal framework [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] where they need to evaluate what the outcome could be when
prescribing a particular treatment.
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Materials</title>
      <p>
        4.1. Data
This project makes use of data coming from two separate cohorts:
• The population-based cohort (PBC): a retrospective cohort of about 1,500 AYAs BC
patients who completed cancer treatment, that is, that survived at least 1 year after the cancer
diagnosis. In this cohort, BC cases have been identified in population-based cancer
registries and the information from each BC patient has been linked to several administrative
data sets (hospital discharge records, outpatients and drug flows) for CVD follow-ups [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
• The clinical-based cohort (CBC): a retrospective single-institution clinical cohort
consisting of 340 additional BC patients, with additional detailed information on cancer
prognostic factors but with no CVD follow-up information.
      </p>
      <p>
        Around 3% of AYAs BC patients from the PBC had at least one CVD event during the follow-up
(mean follow-up time = 5 years). However, the Framingham score for women with the same
characteristics as those in the PBC ranges between 0 and 9, which translates to a predicted
risk of 0% to 1% of developing a hard coronary heart disease (such as myocardial infraction)
or of dying for a CVD event within 10 years [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Exploring the PBC further, we found that
the most frequent CVDs were due to chemotherapy-induced cardiac damages (arrhythmia and
heart failure, in around 1% of patients) and to ischemic heart diseases that may be related to
hormone therapy (in around 2% of patients). Moreover, 15% of patients developed at least one
major cardiac risk factor (dyslipidemia, diabetes and hypertension) after cancer treatment, thus
resulting in an increase in the corresponding Framingham CVD risk.
      </p>
      <p>The CBC consists of patients treated by a single institution (Fondazione IRCCS Istituto
Nazionale dei Tumori di Milano) and is biased due to the patient selection mechanism in ways
that are apparent from the baseline and treatment variables distribution. Around 8% of AYAs BC
patients from the CBC have at least one major cardiac risk factor before treatment, compared to
only 5% in the PBC. Moreover, there is an higher proportion of patients receiving neoadjuvant
treatments (30% in the CBC vs 23% in the PBC), which are more likely to be prescribed when
dealing with later tumor stages at diagnosis (with lymph nodes involvement, metastasis and high
dimension tumors). These facts reflect a more severe case-mix of both baseline characteristics
of the patients and of tumor aggressiveness in the CBC compared to the PBC, which is expected
since the Institute is a referral expertise cancer center.</p>
      <p>Furthermore, the PBC and the CBC were collected for diferent purposes, resulting in
different missingness mechanisms. In the CBC, the purpose is patient care, so factors that may
influence tumor prognosis (such as tumor grading, staging, etc.) are recorded in greater detail;
whereas cancer registry data, which form the basis of the PBC, are collected for administrative,
epidemiological monitoring and public health purposes. More details about the distribution
of missing values by cohort type is reported in Table S1 in the Supplementary Material. their
percentage of completeness in the cohort of origin.</p>
      <sec id="sec-4-1">
        <title>4.2. Clinical Knowledge on Cardiovascular Diseases</title>
        <p>In this section we elicit the clinical knowledge on cardio-oncology. To make it easier for the
reader to understand how the domain knowledge has been translated from natural language
to the structure of the causal network model, we report in square brackets the label of the
corresponding nodes.</p>
        <p>The description of the selection mechanism behind the selection node [cohort] has already
been described in Section 4.1.</p>
        <p>
          To decide which treatment to give to a young BC patient, clinicians rely on the most
wellknown prognostic factors for 5-year cancer survival [death_in_5y] [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]: age (below or
above 35 years) [age35], tumor grade [grade], tumor histology [histology], ki67+
status [ki67], molecular subtype [receptors], vascular invasion [vascular], lymph nodes
involvement [pN] and tumor dimension [pT]. For example, triple negative BC is a specific
molecular subtype of BC that is more common in AYAs than in other age groups (prevalence
15%–20% [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]) and whose treatment options are limited to chemotherapy and/or radiotherapy
because target therapies are inefective.
        </p>
        <p>
          The risk of CVDs after both neoadjuvant [_neo] and adjuvant [_adju] cancer treatments
is well known. For instance, anthracycline is frequently prescribed as a chemotherapy regimen
[chemo_], either alone or in combinations with other chemotherapical drugs. Its
cardiotoxic efects are well documented and largely attributable to the generation of free radicals:
it ultimately results in left-ventricular dysfunction, arrhythmias [cardiotoxicity] and, in
turn, heart failure in the most severe cases [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] [cvds]. Even though radiotherapy [radio_]
reduces the risk of cancer recurrence and death in BC patients, the heart may be incidentally
exposed to ionizing radiation when the primary cancer is located in the left breast, with in turn
increases the risk of heart diseases [cvds] [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. In addition, target therapy [target_] has
been shown to induce acute cardiac toxicity [cardiotoxicity], especially when administered
as an adjuvant treatment (say, trastuzumab for young BC patients [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]), due to possible changes
in oxidative stress that reflect mainly in left-ventricular-ejection fraction reduction. However, it
is generally believed that the trastuzumab-induced cardiotoxic efects are reversible and that
they do not impact long-term CVD risk [18].
        </p>
        <p>Moreover, 5 years of tamoxifen [hormons_], with or without ovarian suppression/ablation,
is considered the standard hormone therapy in young women with hormone-receptor positive
disease, which represent the vast majority in this age group. However, it has frequent chronic
late efects like type-2 diabetes [t2db], hypertension [hypertension] and dyslipidemia
[dyspidemia] which are all known to be major risk factors for long-term ischemic heart
diseases [ischemic_heart_disease] [19].</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Methods</title>
      <p>Assessing validity is a crucial aspect of causal inference. Specifically, internal validity refers
to generalizing the evidence from a study sample to the underlying study population; external
validity refers to transporting the evidence to a diferent target population. The interest in the
transportability of inference has increased exponentially in the last few years [20, 21] thanks
to methodological breakthroughs in handling distribution shifts and selection bias. In this
work, we leverage both missingness graphs [22, 23, 24] and selection diagrams [25] to learn a
causal Bayesian network by combining observational data together with prior knowledge in
the presence of selection bias and missingness bias.</p>
      <sec id="sec-5-1">
        <title>5.1. Causal Models for Missing Values and Selection Bias</title>
        <p>Model-based approaches rely on the formal specification of a model representing the interactions
we are interested in. In the case of causal inference, causal graphs are the de-facto standard for
encoding causal relationships.</p>
        <p>Definition 1 (Causal Graph). A causal graph  = (V, E) [26] is a directed acyclic graph
where for each directed edge (,  ) ∈ E,  is a direct cause of  and  is a direct efect of .
The vertex set V can be split into two disjoint subsets V = O ∪ U, where O is the set of the fully
observed variables (with no missing values) and U is the set of fully unobserved variables, also
known as latent variables.</p>
        <p>A causal graph allows researches to (i) decide if a consistent estimator exists for a casual
efect and (ii) derive that estimator directly from the graph. Being able to correctly estimate
causal efects is extremely important for policy making. For instance, it allows clinicians to
assess the impact of a given drug on a disease from observational studies when randomized
controlled trials are not available for a particular sub-population [27]. In more general terms,
let  be a causal graph,  a treatment and  an outcome. A consistent estimator for the causal
efect of  on  is given by the do-operator:
 ( =  | ( = )) = ∑︁  ( =  |  = , Z = z) (Z = z)
z
(1)
if there exists a set Z that satisfies the back-door criterion for  [28]. If no consistent estimator
exists, confounding bias makes it impossible to estimate the causal efect correctly. Unbiased
estimators of causal efects are theoretically possible even without a causal graph, but their
assumptions are rarely satisfied in practical applications [ 29]. Therefore, causal graphs remain
the tool of choice to achieve both explainability and consistency.</p>
        <p>The missingness mechanism also plays an important role in causal modeling. Common
pre-processing techniques that deal with missing values such as sample deletion and missing
imputation are often inefective or even detrimental to causal efect estimation. For instance, [ 30,
31] have shown how missing-data handling has a significant impact on the clinical conclusions
that can be drawn from experimental results . Therefore, modeling the reason why a given
value is missing is crucial, especially for MNAR data.</p>
        <p>
          The framework for modeling missingness mechanisms is detailed in Rubin’s foundational
work [
          <xref ref-type="bibr" rid="ref18">32</xref>
          ]. More recently, missingness graphs [
          <xref ref-type="bibr" rid="ref19">33, 22</xref>
          ] have been proposed to reconcile Rubin’s
framework with causal graphs by directly including the missingness indicators in an extended
graph structure.
        </p>
        <p>Definition 2 (Missingness Graph). A missingness graph ℳ = (V, E) [22] is a causal graph
whose vertex set V is partitioned into five disjoint subsets O ∪ U ∪ M ∪ S ∪ R, where: M is the
set of the partially observed variables, that is, the variables with at least one missing value; S is
the set of the proxy variables, that is, the variables that are actually observed; R is the set of the
missingness indicators.</p>
        <p>
          Since missingness graphs are extended causal graphs, d-separation [
          <xref ref-type="bibr" rid="ref20">34</xref>
          ] implies (conditional)
independence. Briefly, a set of variables Z d-separates  from  , denoted by  ⊥⊥  | Z,
if it blocks every path (the combination of all edges and nodes that connect two selected
nodes of interest) between  and  . A path is blocked by Z if and only if it contains: a fork
 ←  →  or a chain  →  →  so that  is in Z, or, a collider  →  ←  so that ,
or any descendant of it, is not in Z. In this framework, there exists a one-to-one correspondence
between the Missing Completely At Random (MCAR), Missing At Random (MAR) and MNAR
patterns and the independence statements implied by the missingness graph:
• MCAR implies O ∪ U ∪ M ⊥⊥ R: missingness is random and independent from the fully
observed variables O and the partially observed variables M;
• MAR implies U ∪ M ⊥⊥ R | O: missingness is random only conditionally on the fully
observed variables O;
• MNAR if neither MCAR nor MAR.
        </p>
        <p>
          This makes it possible to verify whether a consistent estimator for the joint probability  (X)
exists in case of MNAR. When the conditions in [
          <xref ref-type="bibr" rid="ref19">33</xref>
          ] hold,  (X) is recoverable and a consistent
estimator is given by
 (X) = ∏︀
        </p>
        <p>(RX = 0, X)
∈X  ( = 0 | Π  , RΠ = 0)
(2)
where  is the missingness indicator for variable , Π  is the parent set of  and RX is
the union of the  for all the variables in X.</p>
        <p>
          When the data are a collation of multiple data sets, their diferent selection criteria may induce
discrepancies in the distribution of some of the collected variables. In such cases, pooling the
data together without modeling the context from which observations come from could induce
inconsistent estimates [
          <xref ref-type="bibr" rid="ref21">35</xref>
          ]. Selection diagrams are introduced as an extension of causal graphs
for this purpose. We report the definition for completeness.
        </p>
        <p>
          Definition 3 (Selection Diagram). Let Π and Π * be two diferent populations with a common
underlying causal graph . A selection diagram  extends the causal graph  so that:
• E ⊂ E , that is,  is a sub-graph of ;
• ∃S ̸= ∅, S ⊂ V ∧ S ̸⊂ V where ( → ) ∈ E if  ̸= * , that is, the variables S
point to the variables V that difer in their value assignment  in Π and Π * .
The variables S are usually called selection variables and allow us to pool together all the
observations while keeping track of their provenance. In related work on selection bias, these
variables are called context variables and identify the context in which V have been collected.
We refer the reader to [
          <xref ref-type="bibr" rid="ref22">36</xref>
          ] for an extended discussion of the diferences between selection
variables and context variables. In our setting we can use these terms interchangeably. Intuitively,
(3)
selection bias represents the diference between a consistent estimate for Π and a consistent
estimate for Π * : it limits the our ability to transport the inference made on a model for a study
population Π to another target population Π * due to the diferences in their selection criteria.
As for d-separation, when a selection diagram  is modeled, it is possible to identify a consistent
estimator for the given set of selection variables S using the g-transportability [
          <xref ref-type="bibr" rid="ref23">37</xref>
          ] criterion.
        </p>
        <p>
          Once a causal graph  is designed, we need to connect it to the data . To do so, we rely
on Casual networks (CNs) [
          <xref ref-type="bibr" rid="ref24">38</xref>
          ], an extension of the well known Bayesian networks (BNs) [
          <xref ref-type="bibr" rid="ref25">39</xref>
          ],
where the underlying graph is a causal graph.
        </p>
        <p>Definition 4 (Causal Network). Let be  a causal graph and let  (X) be a global probability
distribution with parameters Θ . A causal network  = (, Θ) is a causal model where each
variable of X is a vertex of  and  (X) factorizes into local probability distributions following:
 (X) = ∏︁  ( | Π  )</p>
        <p>∈X
where Π  is the parent set of the variable .</p>
        <p>If we combine a missingness graph and a selection diagram into a single causal graph, we
can recover from confounding bias, inconsistent estimators due to missing values and domain
discrepancies with a single causal network.</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Causal Discovery with Missing Values and Prior Knowledge</title>
        <p>
          Manually constructing the true causal graph is virtually impossible in real-world applications
where there is little to no control over the experimental setting. For instance, there could
be unobserved variables afecting the data generating mechanism or unknown interactions
between the observed variables. Causal discovery [
          <xref ref-type="bibr" rid="ref26 ref27">40, 41</xref>
          ] focuses on recovering the causal
graph from collected data and prior knowledge. The existing literature typically assumes that
there are no missing values or that the missingness mechanism is either MCAR or MAR [
          <xref ref-type="bibr" rid="ref28">42</xref>
          ];
extensions to MNAR have only appeared in recent years [23, 24].
        </p>
        <p>
          In this contribution, we take into account the impact of the missingness mechanism using the
Structural Expectation-Maximization (SEM) algorithm [
          <xref ref-type="bibr" rid="ref29 ref30 ref31">43, 44, 45</xref>
          ] and clinical prior knowledge.
At its core, SEM repeats two steps until convergence is reached:
• Expectation (E): impute the missing values ̂︀ from ̂︀ by estimating Θ ̂︀  using ̂︀.
• Maximization (M): find ̂︀+1 that maximises a given score function for ̂︀.
At the end of the th iteration, SEM produces an estimated causal network ̂︀+1 that will be
the starting point of the ( + 1)th iteration. In the case of MNAR, it is essential to model
the missingness mechanism a priori in order to avoid spurious associations induced by the
missingness pattern. We achieve that by providing an initial 0 that encodes a partially-specified
missingness graph from prior knowledge and by fixing its arcs in place throughout the causal
discovery, efectively assuming that they correctly specify the missingness mechanism. Hence,
the SEM algorithm only extends the initial graph, looking for a super-graph of 0 that better
ifts the collected data .
        </p>
      </sec>
      <sec id="sec-5-3">
        <title>5.3. Causal Model Evaluation</title>
        <p>Literature validation. We validated the edges that are added by the SEM algorithm to the
model using the medical literature to support the results relevance in this domain. For this
purpose, we performed a literature review to identify publications that explain what we observe
in the data.</p>
        <p>Performance metrics. The combined data set comprising the PBC and CBC cohorts was
split in a training set and a test set. As mentioned in Section 4.1, the CBC does not include
information on the outcome variables, so we included it only in the training set. The ratio
between training and test set was 70% (60% PBC + 10% CBC) / 30% (PBC). We considered further
splitting the training set to obtain a validation set, but that would reduce the sets sample size to
the point of making the analysis unfeasible.</p>
        <p>We evaluated the CN we learned from the training set on the test set by estimating the
probability of CVDs and the associated Area Under the receiver operating characteristic Curve
(AUC). In particular, we contrasted the AUC obtained with 0, with the network learned from
SEM (starting from 0), and with other standard ML methods. AUC represents the degree or
measure of separability: it measures how much the model is capable of distinguishing between
classes. The higher the AUC, the better the model is at predicting healthy individuals as healthy
and individuals afected by CVDs as at risk.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Experimental Results</title>
      <p>
        Validation of SEM algorithm results. Figure 1 shows the network learned by SEM: edges
elicited from the domain knowledge are in black, while the new ones, that highlight new medical
insights, are in bold blue. The latter comprise:
• Edges directed from [cohort] to the three adjuvant treatments [radio_adju],
[chemo_adju], [hormons_adju] and to one neoadjuvant treatment [target_neo].
These edges can be explained by the selection mechanism: the treatments nodes act as
proxies of case severity [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] which is more likely to be higher in the CBC than in the PBC.
• Edges directed from the neoadjuvant treatments [chemo_neo] and [target_neo] to
adjuvant treatments [chemo_adju], [radio_adju] and [target_adju]. These edges
can be explained by surgical details not encoded in the [surgery] node, which only
describes whether the surgical approach was radical or conservative. The prioritization of
breast preservation in AYAs results in more conservative surgical approaches and, in turn,
in a higher probability of incomplete surgical margins. As a result, adjuvant treatments
are more frequently used to reduce the risk of both local and distant recurrences [
        <xref ref-type="bibr" rid="ref32">46</xref>
        ].
• Edges directed from target therapy [target_neo] to cardiovascular diseases [cvds].
      </p>
      <p>These edges can be explained by other unstudied efects of oxidative stress and
inflammation that induce CVDs but that are not mediated by cardiotoxicity (for instance,
grade
histology</p>
      <p>cohort
ki67
pt
receptors
vascular</p>
      <p>pn
death_in_5y
radio_neo
chemo_neo
target_neo</p>
      <p>hormons_neo
surgery
radio_adju
chemo_adju
hormons_adju
t2db
target_adju
hypertension
dyslipidemia
cardiotoxicity
ischemic_heart_disease</p>
      <p>
        metabolomic and other unmeasured CV risk factors such as obesity) [
        <xref ref-type="bibr" rid="ref33">47</xref>
        ]. Moreover,
according to the literature, our understanding of the mechanisms of trastuzumab-mediated
cardiotoxicity is still evolving [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. Therefore, they should be further investigated to
better tailor the CVDs follow-up guidelines in young BC survivors.
      </p>
      <p>Model performance. The CN model based on the clinical prior knowledge illustrated in
Section 4.2 showed a strong classification performance, with an AUC of about 83% in the test
set. The new edges introduced by the SEM algorithm improved performance to about 88%.</p>
      <p>It is important to clarify that results obtained from standard ML methods are not directly
comparable to those obtained by the proposed approach, for several reasons. Firstly, standard
ML methods are not designed to accommodate missing values, relying instead on pre-processing
techniques such as sample deletion or imputation. Therefore, only cases coming from the PBC
and with complete information on all variables, less than half of the available cases (≈ 800), can
be used to train and validate ML models. Secondly, these models do not address the selection bias
and lack of external validity. Finally, their results are not causally and clinically interpretable.
Against these limitations, we can see in Table 1 that the proposed CN models outperforms all
standard ML methods in terms of classification performance.</p>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusions and Future work</title>
      <p>The CN shown in Figure 1 is the first causal network model for estimating the CVD risk in
adolescents and young adult females who survived breast cancer. Thanks to the combination of
several causal inference methods, including missingness graphs, selection diagrams and causal
discovery algorithms, we were able to build a model that combines domain expert knowledge
and multiple data sets. The proposed model is able to:
• Deal with the uncertainty and biases that are intrinsic to real-world observational data
and produce valid results, generalizing to other applications or populations;
• Combine data coming from diferent data sources , which is especially relevant when dealing
with small and understudied populations for which data are extremely sparse.
• Incorporate the domain knowledge coming from experts, which has been shown to be
fundamental in maximizing the model’s performance.
• Provide interpretable causal recommendations to identify those patients that are at higher
risk of developing a CVDs and what risk factors are involved, to help physicians in clinical
decision making and patient follow-up management.</p>
      <p>Moreover, the edges added by the causal discovery to the prior-knowledge-based causal
network improved its performance and highlighted new arcs that open medical insights particularly
valuable, especially for such an understudied population.</p>
      <p>According to our experimental results, the proposed CN model outperforms standard
supervised learning methods in terms of classification accuracy: this is unexpected since these
methods are specifically trained to optimise it. The are two possible explanations for this
phenomenon. Firstly, as mentioned in the results section, the data used to train the models are
XGBoost: Extreme Gradient Boosting; AODE: Averaged One-Dependence Estimators.
diferent: standard ML methods can use only cases coming from the PBC and with complete
information on all variables, which means less than half of the cases used to train the CN
models. Secondly, the CN models did not learn the model from scratch thanks to their ability
to incorporate prior domain knowledge; that makes them more eficient at dealing with low
sample sizes and missing-not-at-random issues.</p>
      <p>Despite its relevance and strengths, this work has some limitations. First of all, we focused
on the classification accuracy of a specific variable of interest instead of evaluating the CN as a
whole. This allowed us to compare it with commonly-used ML supervised learning approaches
using standard metrics such as AUC, arguing that the proposed CN is competitive in this
respect. Having shown that, we are now ready to move beyond prediction into more advanced
applications of causal inference which are impossible to carry out with traditional ML models.
Moreover, while the arcs added by the SEM algorithm have been validated by a clinical expert
and a literature review, a further review by a domain expert (that is, a cardio-oncologist) will be
crucial to completely validate the final model.</p>
      <p>In conclusion, we are working on further extending our model to include unmeasured
variables, such as distant metastases; this is necessary to further address the need of a clinically
relevant model, to help cardio-oncologists in their everyday practice.</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgments</title>
      <p>Alice Bernasconi is funded by an AIRC 2020 project, grant number 24864, titled “pRedicting
cardiOvascular diSeAses iN adolescent and young breast caNcer pAtients (ROSANNA)”.</p>
      <p>Alessio Zanga is funded by F. Hofmann-La Roche Ltd.</p>
      <p>This work was partially supported by the MUR under the grant “Dipartimenti di Eccellenza
2023-2027" of the Department of Informatics, Systems and Communication of the University of
Milano-Bicocca, Italy.
[18] J. An, M. S. Sheikh, Toxicology of trastuzumab: An insight into mechanisms
of cardiotoxicity, Current Cancer Drug Targets 19 (2019) 400–407. doi:10.2174/
1568009618666171129222159.
[19] A. Christinat, S. D. Lascio, O. Pagani, Hormonal therapies in young breast cancer patients:
when, what and for how long?, Journal of thoracic disease 5 Suppl 1 (2013) S36–46.
doi:10.3978/j.issn.2072-1439.2013.05.25.
[20] I. Degtiar, S. Rose, A Review of Generalizability and Transportability,
Annual Review of Statistics and Its Application 10 (2023) 501–524. URL: https://
www.annualreviews.org/doi/10.1146/annurev-statistics-042522-103837. doi:10.1146/
annurev-statistics-042522-103837.
[21] K. M. Esterling, D. Brady, E. Schwitzgebel, The Necessity of Construct and External Validity
for Generalized Causal Claims, Technical Report 18, s.l., 2023. URL: http://hdl.handle.net/
10419/268605.
[22] K. Mohan, J. Pearl, Graphical Models for Processing Missing Data, Journal of the American
Statistical Association 116 (2018) 1023–1037. URL: https://www.tandfonline.com/doi/full/
10.1080/01621459.2021.1874961http://arxiv.org/abs/1801.03583. doi:10.1080/01621459.
2021.1874961.
[23] R. Tu, K. Zhang, P. Ackermann, B. C. Bertilson, C. Glymour, H. Kjellström, C. Zhang, Causal</p>
      <p>Discovery in the Presence of Missing Data (2018). URL: http://arxiv.org/abs/1807.04010.
[24] Y. Liu, A. C. Constantinou, Greedy structure learning from data that contain systematic
missing values, Machine Learning 111 (2022) 3867–3896. URL: https://link.springer.com/
article/10.1007/s10994-022-06195-8. doi:10.1007/S10994-022-06195-8/TABLES/9.
[25] S. Lee, J. Correa, E. Bareinboim, General Transportability – Synthesizing Observations
and Experiments from Heterogeneous Domains, Proceedings of the AAAI Conference on
Artificial Intelligence 34 (2020) 10210–10217. URL: https://ojs.aaai.org/index.php/AAAI/
article/view/6582. doi:10.1609/aaai.v34i06.6582.
[26] J. Pearl, Causal diagrams for empirical research, Biometrika 82 (1995) 669–688. doi:10.</p>
      <p>1093/biomet/82.4.669.
[27] C. Fernández-Loría, F. Provost, Causal Decision Making and Causal Efect Estimation Are
Not the Same... and Why It Matters, the inaugural issue of the INFORMS Journal of Data
Science (2021). URL: http://arxiv.org/abs/2104.04103.
[28] J. Pearl, Causality: Models, Reasoning and Inference, 2nd ed., Cambridge University Press,</p>
      <p>USA, 2009.
[29] M. A. Hernán, J. M. Robins, Causal Inference: What If, Boca Raton: Chapman \&amp; Hall/CRC,
2020.
[30] M. R. Stavseth, T. Clausen, J. Røislien, How handling missing data may impact conclusions:
A comparison of six diferent imputation methods for categorical questionnaire data, SAGE
Open Medicine 7 (2019) 205031211882291. URL: https://doi.org/10.1177/2050312118822912.
doi:10.1177/2050312118822912.
[31] A. Zanga, A. Bernasconi, P. J. F. Lucas, H. Pijnenborg, C. Reijnen, M. Scutari, F. Stella,
Causal Discovery with Missing Data in a Multicentric Clinical Study, in: Proceedings of
the 21st International Conference of Artificial Intelligence in Medicine (AIME), volume
13897 LNAI, 2023, pp. 40–44. URL: https://link.springer.com/10.1007/978-3-031-34344-5_5.
doi:10.1007/978-3-031-34344-5{\_}5.</p>
    </sec>
    <sec id="sec-9">
      <title>8. Supplementary Material</title>
      <p>PBC = Population-based cohort; CBC = Clinical-based Cohort; CVDs = Cardiovascular diseases.
CBC</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>L.</given-names>
            <surname>Leung</surname>
          </string-name>
          , Validity, reliability, and
          <article-title>generalizability in qualitative research</article-title>
          ,
          <source>Journal of Family Medicine and Primary Care</source>
          <volume>4</volume>
          (
          <year>2015</year>
          )
          <article-title>324</article-title>
          . doi:
          <volume>10</volume>
          .4103/
          <fpage>2249</fpage>
          -
          <lpage>4863</lpage>
          .
          <fpage>161306</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Trama</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Botta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Foschi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ferrari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Stiller</surname>
          </string-name>
          , E. Desandes,
          <string-name>
            <surname>M. M. Maule</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Merletti</surname>
          </string-name>
          , G. Gatta,
          <article-title>Survival of european adolescents and young adults diagnosed with cancer in 2000-07: population-based data from eurocare-5,</article-title>
          <source>The Lancet Oncology</source>
          <volume>17</volume>
          (
          <year>2016</year>
          )
          <fpage>896</fpage>
          -
          <lpage>906</lpage>
          . doi:
          <volume>10</volume>
          .1016/S1470-2045(
          <volume>16</volume>
          )
          <fpage>00162</fpage>
          -
          <lpage>5</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>R.</given-names>
            <surname>Schafar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bouchardy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. O.</given-names>
            <surname>Chappuis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bodmer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Benhamou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Rapiti</surname>
          </string-name>
          ,
          <article-title>A populationbased cohort of young women diagnosed with breast cancer in geneva, switzerland</article-title>
          ,
          <source>PLOS ONE 14</source>
          (
          <year>2019</year>
          )
          <article-title>e0222136</article-title>
          . doi:
          <volume>10</volume>
          .1371/journal.pone.
          <volume>0222136</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>R.</given-names>
            <surname>Altena</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Hubbert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. A.</given-names>
            <surname>Kiani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wengström</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bergh</surname>
          </string-name>
          , E. Hedayati,
          <article-title>Evidence-based prediction and prevention of cardiovascular morbidity in adults treated for cancer</article-title>
          ,
          <source>CardioOncology</source>
          <volume>7</volume>
          (
          <year>2021</year>
          )
          <article-title>20</article-title>
          . doi:
          <volume>10</volume>
          .1186/s40959-021-00105-y.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>W.</given-names>
            <surname>Law</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Johnson</surname>
          </string-name>
          , M. Rushton,
          <string-name>
            <surname>S. Dent,</surname>
          </string-name>
          <article-title>The framingham risk score underestimates the risk of cardiovascular events in the her2-positive breast cancer population</article-title>
          ,
          <source>Current Oncology</source>
          <volume>24</volume>
          (
          <year>2017</year>
          )
          <fpage>348</fpage>
          -
          <lpage>353</lpage>
          . doi:
          <volume>10</volume>
          .3747/co.24.3684.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>N.</given-names>
            <surname>Madan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lucas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Akhter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Collier</surname>
          </string-name>
          , F. Cheng, A.
          <string-name>
            <surname>Guha</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Sharma</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Hamid</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <string-name>
            <surname>Ndiokho</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Wen</surname>
            ,
            <given-names>N. C.</given-names>
          </string-name>
          <string-name>
            <surname>Garster</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Scherrer-Crosbie</surname>
            ,
            <given-names>S.-A.</given-names>
          </string-name>
          <string-name>
            <surname>Brown</surname>
          </string-name>
          , Artificial intelligence and imaging: Opportunities in cardio-oncology,
          <source>American Heart Journal Plus: Cardiology Research and Practice</source>
          <volume>15</volume>
          (
          <year>2022</year>
          )
          <article-title>100126</article-title>
          . doi:
          <volume>10</volume>
          .1016/j.ahjo.
          <year>2022</year>
          .
          <volume>100126</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ouyang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Baykaner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Jamal</surname>
          </string-name>
          , P. Cheng, J.-W. Rhee,
          <article-title>Artificial intelligence applications in cardio-oncology: Leveraging high dimensional cardiovascular data</article-title>
          ,
          <source>Frontiers in Cardiovascular Medicine</source>
          <volume>9</volume>
          (
          <year>2022</year>
          ). doi:
          <volume>10</volume>
          .3389/fcvm.
          <year>2022</year>
          .
          <volume>941148</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>P.</given-names>
            <surname>Nolan</surname>
          </string-name>
          ,
          <article-title>Artificial intelligence in medicine - is too much transparency a good thing?</article-title>
          ,
          <string-name>
            <surname>Medico-Legal Journal</surname>
          </string-name>
          (
          <year>2023</year>
          )
          <article-title>002581722211412</article-title>
          . doi:
          <volume>10</volume>
          .1177/00258172221141243.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>P.</given-names>
            <surname>Linardatos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Papastefanopoulos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kotsiantis</surname>
          </string-name>
          ,
          <article-title>Explainable ai: A review of machine learning interpretability methods</article-title>
          ,
          <source>Entropy</source>
          <volume>23</volume>
          (
          <year>2020</year>
          )
          <article-title>18</article-title>
          . doi:
          <volume>10</volume>
          .3390/e23010018.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>E. J.</given-names>
            <surname>Chow</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. C.</given-names>
            <surname>Kremer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. E.</given-names>
            <surname>Breslow</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. M. Hudson</surname>
            ,
            <given-names>G. T.</given-names>
          </string-name>
          <string-name>
            <surname>Armstrong</surname>
            ,
            <given-names>W. L.</given-names>
          </string-name>
          <string-name>
            <surname>Border</surname>
            ,
            <given-names>E. A.</given-names>
          </string-name>
          <string-name>
            <surname>Feijen</surname>
            ,
            <given-names>D. M.</given-names>
          </string-name>
          <string-name>
            <surname>Green</surname>
            ,
            <given-names>L. R.</given-names>
          </string-name>
          <string-name>
            <surname>Meacham</surname>
            ,
            <given-names>K. A.</given-names>
          </string-name>
          <string-name>
            <surname>Meeske</surname>
            ,
            <given-names>D. A.</given-names>
          </string-name>
          <string-name>
            <surname>Mulrooney</surname>
            ,
            <given-names>K. K.</given-names>
          </string-name>
          <string-name>
            <surname>Ness</surname>
            ,
            <given-names>K. C.</given-names>
          </string-name>
          <string-name>
            <surname>Oefinger</surname>
            ,
            <given-names>C. A.</given-names>
          </string-name>
          <string-name>
            <surname>Sklar</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Stovall</surname>
            ,
            <given-names>H. J. van der Pal</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>R. E.</given-names>
            <surname>Weathers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. L.</given-names>
            <surname>Robison</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yasui</surname>
          </string-name>
          ,
          <article-title>Individual prediction of heart failure among childhood cancer survivors</article-title>
          ,
          <source>Journal of Clinical Oncology</source>
          <volume>33</volume>
          (
          <year>2015</year>
          )
          <fpage>394</fpage>
          -
          <lpage>402</lpage>
          . doi:
          <volume>10</volume>
          .1200/JCO.
          <year>2014</year>
          .
          <volume>56</volume>
          .1373.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>E.</given-names>
            <surname>Bareinboim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pearl</surname>
          </string-name>
          ,
          <article-title>Causal inference and the data-fusion problem</article-title>
          ,
          <source>Proceedings of the National Academy of Sciences of the United States of America</source>
          <volume>113</volume>
          (
          <year>2016</year>
          )
          <fpage>7345</fpage>
          -
          <lpage>7352</lpage>
          . doi:
          <volume>10</volume>
          .1073/pnas.1510507113.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>A.</given-names>
            <surname>Bernasconi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Barigelletti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Tittarelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Botta</surname>
          </string-name>
          , G. Gatta, G. Tagliabue,
          <string-name>
            <given-names>P.</given-names>
            <surname>Contiero</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Guzzinati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Andreano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Manneschi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Falcini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Castaing</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. A.</given-names>
            <surname>Filiberti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Gasparotti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Cirilli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Mazzucco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Mangone</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Iacovacci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. F.</given-names>
            <surname>Vitale</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Stracci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pifer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Tumino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Carone</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sampietro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Melcarne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Ballotari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Boschetti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pisani</surname>
          </string-name>
          , L.
          <string-name>
            <surname>C. D'Oro</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Cuccaro</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. D'Argenzio</surname>
          </string-name>
          ,
          <string-name>
            <surname>G. D'Orsi</surname>
            ,
            <given-names>A. C.</given-names>
          </string-name>
          <string-name>
            <surname>Fanetti</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Ardizzone</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Candela</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Savoia</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Pascucci</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Castelli</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Storchi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Trama</surname>
          </string-name>
          ,
          <article-title>Adolescent and young adult cancer survivors: Design and characteristics of the first nationwide populationbased cohort in italy</article-title>
          ,
          <source>Journal of Adolescent and Young Adult Oncology</source>
          <volume>9</volume>
          (
          <year>2020</year>
          )
          <fpage>586</fpage>
          -
          <lpage>593</lpage>
          . doi:
          <volume>10</volume>
          .1089/jayao.
          <year>2019</year>
          .
          <volume>0170</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>F.</given-names>
            <surname>Cardoso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kyriakides</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ohno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Penault-Llorca</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Poortmans</surname>
          </string-name>
          , I. Rubio,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zackrisson</surname>
          </string-name>
          , E. Senkus,
          <article-title>Early breast cancer: Esmo clinical practice guidelines for diagnosis, treatment and follow-up</article-title>
          ,
          <source>Annals of Oncology</source>
          <volume>30</volume>
          (
          <year>2019</year>
          )
          <fpage>1194</fpage>
          -
          <lpage>1220</lpage>
          . doi:
          <volume>10</volume>
          .1093/annonc/mdz173.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>H. J.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. A.</given-names>
            <surname>Freedman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. H.</given-names>
            <surname>Partridge</surname>
          </string-name>
          ,
          <article-title>The impact of young age at diagnosis (age &amp;lt;40 years) on prognosis varies by breast cancer subtype: A u.s. seer database analysis</article-title>
          ,
          <source>The Breast</source>
          <volume>61</volume>
          (
          <year>2022</year>
          )
          <fpage>77</fpage>
          -
          <lpage>83</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.breast.
          <year>2021</year>
          .
          <volume>12</volume>
          .006.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M.</given-names>
            <surname>Volkova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Russell</surname>
          </string-name>
          ,
          <article-title>Anthracycline cardiotoxicity: Prevalence, pathogenesis and treatment</article-title>
          ,
          <source>Current Cardiology Reviews</source>
          <volume>7</volume>
          (
          <year>2012</year>
          )
          <fpage>214</fpage>
          -
          <lpage>220</lpage>
          . doi:
          <volume>10</volume>
          .2174/ 157340311799960645.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>C.</given-names>
            <surname>Taylor</surname>
          </string-name>
          , A. Kirby,
          <article-title>Cardiac side-efects from breast cancer radiotherapy</article-title>
          ,
          <source>Clinical Oncology</source>
          <volume>27</volume>
          (
          <year>2015</year>
          )
          <fpage>621</fpage>
          -
          <lpage>629</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.clon.
          <year>2015</year>
          .
          <volume>06</volume>
          .007.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>N.</given-names>
            <surname>Mohan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dokmanovic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. J.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <article-title>Trastuzumab-mediated cardiotoxicity: current understanding, challenges, and frontiers</article-title>
          ,
          <source>Antibody Therapeutics</source>
          <volume>1</volume>
          (
          <year>2018</year>
          )
          <fpage>13</fpage>
          -
          <lpage>17</lpage>
          . doi:
          <volume>10</volume>
          .1093/abt/tby003.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [32]
          <string-name>
            <surname>D. D. Rubin</surname>
          </string-name>
          ,
          <article-title>Inference and missing data</article-title>
          ,
          <source>Biometrika</source>
          <volume>63</volume>
          (
          <year>1976</year>
          )
          <fpage>581</fpage>
          -
          <lpage>592</lpage>
          . URL: https: //academic.oup.com/biomet/article-lookup/doi/10.1093/biomet/63.3.581. doi:
          <volume>10</volume>
          .1093/ biomet/63.3.581.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>K.</given-names>
            <surname>Mohan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pearl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tian</surname>
          </string-name>
          ,
          <article-title>Graphical models for inference with missing data</article-title>
          ,
          <source>in: Advances in Neural Information Processing Systems</source>
          ,
          <year>2013</year>
          . URL: https://proceedings.neurips.cc/ paper/2013/file/0f8033cf9437c213ee13937b1c4c455-Paper.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>D.</given-names>
            <surname>Koller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Friedman</surname>
          </string-name>
          ,
          <article-title>Probabilistic Graphical Models: Principles and Techniques</article-title>
          , The MIT Press,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>P.</given-names>
            <surname>Forré</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Mooij</surname>
          </string-name>
          ,
          <article-title>Causal calculus in the presence of cycles, latent confounders</article-title>
          and
          <source>selection bias, 35th Conference on Uncertainty in Artificial Intelligence</source>
          ,
          <string-name>
            <surname>UAI</surname>
          </string-name>
          <year>2019</year>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [36]
          <string-name>
            <surname>J. M. Mooij</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Magliacane</surname>
          </string-name>
          , T. Claassen,
          <article-title>Joint causal inference from multiple contexts</article-title>
          ,
          <source>Journal of Machine Learning Research</source>
          <volume>21</volume>
          (
          <year>2020</year>
          )
          <fpage>1</fpage>
          -
          <lpage>108</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>E.</given-names>
            <surname>Bareinboim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pearl</surname>
          </string-name>
          , A
          <article-title>General Algorithm for Deciding Transportability of Experimental Results</article-title>
          ,
          <source>Journal of Causal Inference</source>
          <volume>1</volume>
          (
          <year>2013</year>
          )
          <fpage>107</fpage>
          -
          <lpage>134</lpage>
          . URL: https://www.degruyter.com/ document/doi/10.1515/jci-2012-0004/html. doi:
          <volume>10</volume>
          .1515/jci-2012-0004.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>J.</given-names>
            <surname>Pearl</surname>
          </string-name>
          ,
          <article-title>From Bayesian Networks to Causal Networks</article-title>
          ,
          <source>in: Mathematical Models for Handling Partial Knowledge in Artificial Intelligence</source>
          ,
          <string-name>
            <surname>Springer</surname>
            <given-names>US</given-names>
          </string-name>
          , Boston, MA,
          <year>1995</year>
          , pp.
          <fpage>157</fpage>
          -
          <lpage>182</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-1-
          <fpage>4899</fpage>
          -1424-8{\_}
          <fpage>9</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [39]
          <string-name>
            <given-names>J.</given-names>
            <surname>Pearl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Russell</surname>
          </string-name>
          , Bayesian Networks,
          <source>Technical Report</source>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [40]
          <string-name>
            <given-names>A.</given-names>
            <surname>Zanga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Ozkirimli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Stella</surname>
          </string-name>
          ,
          <source>A Survey on Causal Discovery: Theory and Practice</source>
          ,
          <source>International Journal of Approximate Reasoning</source>
          <volume>151</volume>
          (
          <year>2022</year>
          )
          <fpage>101</fpage>
          -
          <lpage>129</lpage>
          . URL: https://linkinghub. elsevier.com/retrieve/pii/S0888613X22001402. doi:
          <volume>10</volume>
          .1016/j.ijar.
          <year>2022</year>
          .
          <volume>09</volume>
          .004.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [41]
          <string-name>
            <given-names>P.</given-names>
            <surname>Spirtes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. N.</given-names>
            <surname>Glymour</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Scheines</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Heckerman</surname>
          </string-name>
          , Causation, prediction, and search, MIT press,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [42]
          <string-name>
            <given-names>M.</given-names>
            <surname>Scutari</surname>
          </string-name>
          ,
          <article-title>Bayesian network models for incomplete and dynamic data</article-title>
          ,
          <source>Statistica Neerlandica</source>
          <volume>74</volume>
          (
          <year>2020</year>
          )
          <fpage>397</fpage>
          -
          <lpage>419</lpage>
          . URL: https://onlinelibrary.wiley.com/doi/10.1111/stan.12197. doi:
          <volume>10</volume>
          .1111/stan.12197.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [43]
          <string-name>
            <given-names>S. L.</given-names>
            <surname>Lauritzen</surname>
          </string-name>
          ,
          <article-title>The EM algorithm for graphical association models with missing data</article-title>
          ,
          <source>Computational Statistics and Data Analysis</source>
          <volume>19</volume>
          (
          <year>1995</year>
          ). doi:
          <volume>10</volume>
          .1016/
          <fpage>0167</fpage>
          -
          <lpage>9473</lpage>
          (
          <issue>93</issue>
          )
          <fpage>E0056</fpage>
          -A.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [44]
          <string-name>
            <given-names>N.</given-names>
            <surname>Friedman</surname>
          </string-name>
          ,
          <article-title>Learning Belief Networks in the Presence of Missing Values and Hidden Variables</article-title>
          ,
          <source>in: Proceedings of the Fourteenth International Conference on Machine Learning</source>
          , ICML '
          <fpage>97</fpage>
          , Morgan Kaufmann Publishers Inc., San Francisco, CA, USA,
          <year>1997</year>
          , pp.
          <fpage>125</fpage>
          -
          <lpage>133</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [45]
          <string-name>
            <given-names>A.</given-names>
            <surname>Zanga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bernasconi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Lucas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Pijnenborg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Reijnen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Scutari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Stella</surname>
          </string-name>
          ,
          <article-title>Risk Assessment of Lymph Node Metastases in Endometrial Cancer Patients: A Causal Approach</article-title>
          , in
          <source>: Proceedings of the 1st Workshop on Artificial Intelligence For Healthcare (HC@AIxIA)</source>
          , volume
          <volume>3307</volume>
          ,
          <year>2022</year>
          . URL: https://www.scopus.com/inward/record.uri?eid=
          <fpage>2</fpage>
          -
          <lpage>s2</lpage>
          .
          <fpage>0</fpage>
          -
          <lpage>85145592167</lpage>
          &amp;partnerID=
          <fpage>40</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [46]
          <string-name>
            <given-names>D. V.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          , S.-W. Kim,
          <string-name>
            <given-names>Y.-T.</given-names>
            <surname>Oh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O. K.</given-names>
            <surname>Noh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Jung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. S.</given-names>
            <surname>Yoon</surname>
          </string-name>
          ,
          <article-title>Local recurrence in young women with breast cancer: Breast conserving therapy vs</article-title>
          .
          <source>mastectomy alone, Cancers</source>
          <volume>13</volume>
          (
          <year>2021</year>
          )
          <article-title>2150</article-title>
          . doi:
          <volume>10</volume>
          .3390/cancers13092150.
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [47]
          <string-name>
            <given-names>G. K.</given-names>
            <surname>Jakubiak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Osadnik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lejawa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kasperczyk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Osadnik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Pawlas</surname>
          </string-name>
          ,
          <article-title>Oxidative stress in association with metabolic health and obesity in young adults</article-title>
          ,
          <source>Oxidative Medicine and Cellular Longevity</source>
          <year>2021</year>
          (
          <year>2021</year>
          )
          <fpage>1</fpage>
          -
          <lpage>19</lpage>
          . doi:
          <volume>10</volume>
          .1155/
          <year>2021</year>
          /9987352.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>