<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>FAIR quantitative imaging in oncology: how Semantic Web and Ontologies will support reproducible science</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>A. Traverso</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Z. Shi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>L. Wee</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>A. Dekker</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Radiation Oncology (MAASTRO), GROW - School for Oncology and Development Biology, Maastricht University Medical Center</institution>
          ,
          <addr-line>Maastricht</addr-line>
          ,
          <country country="NL">the Netherlands</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The automated extraction of quantitative imaging biomarkers from patient's scans, could augment physician decision making in radiation oncology. Unfortunately, lack of reproducibility and robust methodology current limits this promising field to be applied in the clinic. In this paper, we state how the combination of quantitative medical imaging with Semantic Web and Ontologies techniques could speed up the role of quantitative imaging.</p>
      </abstract>
      <kwd-group>
        <kwd>Ontologies</kwd>
        <kwd>Semantic Web</kwd>
        <kwd>Quantitative Imaging</kwd>
        <kwd>Radiation Oncology</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <sec id="sec-1-1">
        <title>A new era of medical imaging: from images to big data</title>
        <p>
          Medical imaging has expanded its fundamental role in radiation oncology since the
advent of the first Computed Tomography (CT) scans in the 70s, followed by PET
(Positron Emission Tomography) and MRI (Magnetic Resonance Imaging).
Radiological examination has moved from purely descriptive to semi-quantitative and
fully automated analysis. In the recent years, the availability of enterprise digital
imaging and the overflowing role of AI (Artificial Intelligence, like Machine Learning)
domain (e.g. machine learning) led to the development of many quantitative imaging
models aimed at assisting and augmenting physician decision-making. The
term “radiomics” was first created in 2012 and it describes the process of advanced
quantitative clinical imaging analysis in medicine. The hypothesis behind radiomics
is that tumor biological properties, often obtained by invasive techniques
such as tissue biopsies, can be measured in a non-invasive fashion via extracting
image-based descriptors (referred as ‘features’) from medical images [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].
        </p>
        <p>
          After 2012, the number of radiomics computational packages has increased [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ].
However, no consensus has been reached: a) on the optimal configuration that should
be used to extract these features for a problem; b) about the robustness of radiomics
features when evaluated in different contexts. Therefore, most of the users
simultaneously extract features using different parameters, leading to an increase
of the number of features. Typical radiomic studies often extract from 500 to 10000
features while starting only from 100 unique features [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] . We are now facing the same
Copyright © 2019 for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0).
“data explosion” defined by Rubin about multi-detector row CT scanners. One
main difference divides the two processes: if the CT data explosion was mainly
driven by an advance in hardware development, producing more images faster than
expected; the new quantitative imaging data explosion is driven by automated imaging
analysis computational pipelines that produce a large amount of processed
data (e.g. radiomic features) from medical images. This data seems mimicking all
the attributes of big data: a) volume: the large amount of data to be processed and
analyzed via machine learning requires now dedicated computational power and
powerful machine learning able to deal with a large hyperspace of parameters; b)
velocity: new data are generated faster as soon as new computational radiomics
software become available, with a larger hyperspace of parameters that can be
tuned for features extraction; c) variety: not only singe features should be stored in
quantitative imaging, but also information about the original source (image, region
of interest, computational details) making the data variety larger; d) veracity: in the
hyperspace determined by features and associated metadata, some information
could be redundant and only meaningful one should be extrapolated [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. For all the
above-mentioned reasons, quantitative imaging strictly connects to the world of big
data. We believe that extending the usage of ontologies and Semantic Web technologies
to quantitative imaging could help solving some of the issues that would be presented
in the next paragraph and further speed up the adoption and acceptance of new image
based quantitative biomarkers in the clinic.
        </p>
      </sec>
      <sec id="sec-1-2">
        <title>1.2 Reproducibility crisis in quantitative imaging</title>
        <p>
          Still a strong unbalance exists between published radiomics-based prediction models
and their real usage as decision support systems in the clinic [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
        </p>
        <p>
          The lack of reproducibility and transparency in radiomics is the major slowdown
of its applicability in the clinic [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. The lack of reproducibility mainly relates to
the fact that most radiomics-based models are built on limited-datasets and often
validated in one single institution, with no guarantee of generalizability power when
applied to multiple centers. This evidence also seems colliding with recommendations
from the TRIPOD (Transparent Reporting of a Multivariable Prediction Model for
Individual Prognosis or Diagnosis), suggesting and encouraging TRIPOD IV-type
models, which are fully validated on completely independent external datasets [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. TRIPOD
IV models are based on the possibility for an external user to fully reproduce and
validate a previously developed model. Unfortunately, this reproducibility crisis reflects
not only on the difficulty for external users to fully reproduce a radiomics experiments
developed in another institution, but also within the same institution.
        </p>
        <p>
          This issue mainly connects to the previously mentioned concept of lack of transparency.
In absence of a standardized and structured way of describing radiomics studies, most
of them only report single feature names or values, with no further details on how the
model was developed, how the features were computed and which where the
computational parameters used (metadata). Even in presence of publications that made available
software and datasets, re-usability and inter-operability remain issues. It is not unlikely
that two software could call a radiomic feature with the same name but meaning a
totally different quantitative descriptor. On the other hand, two features could express the
same quantitative descriptor but show different values when computed with different
software. Without then associated metadata, it is impossible to find the reasons behind
this discrepancy, which probably lie in a different choice of hyperparameters.
It becomes then clear that quantitative imaging is far behind the FAIR principles that
are taking the scene in clinical data science as incentive for reproducible and transparent
science [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. However, the absence of FAIR guiding principles represents a unique
opportunity for the imaging community to propose a new paradigm for a new era
reproducible quantitative imaging. We believe that ontologies and Semantic Web techniques
should guide this effort toward reproducible, transparent quantitative imaging. On the
other side, the imaging community needs to accept the challenge to work closely with
the data science community and re-use as much as possible available tools. A possible
framework and the ongoing actions taken by our group are presented in the following
paragraph.
2.
        </p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Proposed solution</title>
      <p>2.1 Ontologies for quantitative imaging: a dynamic body of knowledge to enhance
consensus</p>
      <p>
        Ontologies represent a formal specification of the terms related to a specific domain
and the relations among them [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. In this specific case, an ontology for quantitative
imaging should mimic the workflow that happens during a radiomic study: from image
pre-processing, region of interest definition, computational settings definition and
finally features extraction, as presented in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Therefore, the ontology not only should
include the main radiomic features and their corresponding units, but also all the
metadata that relate to the above-mentioned workflow. In this view, building this
ontology is a joint exercise between imaging research groups to represent the state of the
art of the knowledge related to the quantitative imaging domain. The ontology acts as
harmonizer and standardizer, eliminating barriers related to different nomenclature or
labels. In fact, each concept in the ontology is universally defined and the whole
community agrees on its meaning. For example, the ontology universally defines the
radiomics features by describing them and associating a unique identifier and their
provenance. In this view, it enhances consensus and creates a shared knowledge domain. It
represents a dynamic body of knowledge that can be expanded with new concepts as
the quantitative imaging field evolved (for example by introducing and defining new
imaging features or computational methods). Our group took the lead in developing an
extensive radiomics ontology (RO), released on the BioPortal
(https://bioportal.bioontology.org/ontologies/RO) as door-opener for FAIR quantitative imaging. Recently, we
published a modular python tool for making radiomics computations FAIR [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Finally,
ontologies express concepts in a machine-readable language and therefore, when data
and metadata are transformed via the ontology, they can be automatically parsed by
machines. This becomes of fundamental utility when comparing results computed from
different software or under different conditions. If each radiomics computational
package is setup to produce ontologies-labelled data and metadata, then automated meta
analyzes can be performed and this will open the path to data-driven standardization
and harmonization. A summary of the concept behind the RO and possible applications
is depicted in Figure 1.
      </p>
      <sec id="sec-2-1">
        <title>Semantic Web: linking quantitative imaging with multiple domains</title>
        <p>Semantic Web has the power to extract knowledge from data labelled via ontologies,
using dedicated SPARQL language.</p>
        <p>
          If radiomics data and metadata are transformed via the Radiomics Ontology and
published on the Semantic Web, then they can be queried using the universal concepts
defined by the ontology, without any prior knowledge on the original labels present in
the original software. Also, the combination of ontologies and Semantic Web
techniques allows parsing and joining data and metadata from multiple sources, such as
different databases. For example, in a typical radiomics-based prediction study it could
be interesting to query a) the value of a certain feature b) computed on an imaging
modality c) referring to a patient with a certain disease; d) finding patients with
similar feature values but different clinical outcomes for comparison. As it is clear from
this example, that type of query requires merging radiomics data (a); DICOM metadata
(b); clinical data (c), and data from other clinics (d). Sooner, additional sources of data
such as for example genomics data or pathology data, for better predictions and for
exploring connections with medical images will be needed. Our group has developed a
portfolio of ontologies for guaranteeing the road to FAIR compliant and transparent
prediction models in radiation oncology: the ROO (Radiation Oncology Ontology) [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ],
the SEDI (Semantic DICOM Ontology) [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] and the presented RO.
        </p>
        <p>
          We successfully showed how this workflow can be used in combination with
Semantic Web for winning barriers related to data sharing and build more accurate models
(distributed learning) [9]. For example, we successfully reproduced a classical
centralized radiomics study [10] in a distributed fashion using the above-mentioned
ontologies combined with Semantic Web [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. By using only SPARQL queries we could
retrieve the model and computational details of the model trained at one local institution
and externally validated on the second one.
        </p>
        <p>We believe the upcoming effort should focus on developing additional ontologies
that could link the quantitative imaging domain with data from multiple sources
presented above.</p>
        <p>Finally, we state that ontologies and Semantic Web are the key for speeding up
reproducible science. Therefore, the quantitative imaging community should work
closely with experts from the semantics, FAIR and data science fields to provide a
sustainable infrastructure for medical imaging and derived big data.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Gillies</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. E.</given-names>
            <surname>Kinahan</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.</given-names>
            <surname>Hricak</surname>
          </string-name>
          , '
          <article-title>Radiomics: Images Are More than Pictures, They Are Data'</article-title>
          ,
          <source>Radiology</source>
          , vol.
          <volume>278</volume>
          , no.
          <issue>2</issue>
          , pp.
          <fpage>563</fpage>
          -
          <lpage>577</lpage>
          , Feb.
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>L. E.</given-names>
            <surname>Court</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Fave</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Mackin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yang</surname>
          </string-name>
          , and L. Zhang, '
          <article-title>Computational resources for radiomics'</article-title>
          ,
          <source>Transl. Cancer Res.</source>
          , vol.
          <volume>5</volume>
          , no.
          <issue>4</issue>
          , pp.
          <fpage>340</fpage>
          -
          <lpage>348</lpage>
          , Aug.
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>V.</given-names>
            <surname>Kumar</surname>
          </string-name>
          et al.,
          <article-title>'Radiomics: the process and the challenges'</article-title>
          ,
          <source>Magnetic Resonance Imaging</source>
          , vol.
          <volume>30</volume>
          , no.
          <issue>9</issue>
          , pp.
          <fpage>1234</fpage>
          -
          <lpage>1248</lpage>
          , Nov.
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>I.</given-names>
            <surname>Buvat</surname>
          </string-name>
          and
          <string-name>
            <given-names>F.</given-names>
            <surname>Orlhac</surname>
          </string-name>
          , '
          <article-title>The Dark Side of Radiomics: On the Paramount Importance of Publishing Negative Results'</article-title>
          ,
          <source>J Nucl Med</source>
          , vol.
          <volume>60</volume>
          , no.
          <issue>11</issue>
          , pp.
          <fpage>1543</fpage>
          -
          <lpage>1544</lpage>
          , Nov.
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Traverso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Wee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dekker</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Gillies</surname>
          </string-name>
          , '
          <article-title>Repeatability and Reproducibility of Radiomic Features: A Systematic Review'</article-title>
          ,
          <source>International Journal of Radiation Oncology*Biology*Physics</source>
          , vol.
          <volume>102</volume>
          , no.
          <issue>4</issue>
          , pp.
          <fpage>1143</fpage>
          -
          <lpage>1158</lpage>
          , Nov.
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>M. D.</surname>
          </string-name>
          Wilkinson et al., '
          <article-title>The FAIR Guiding Principles for scientific data management and stewardship'</article-title>
          ,
          <source>Scientific Data</source>
          , vol.
          <volume>3</volume>
          , p.
          <fpage>160018</fpage>
          ,
          <string-name>
            <surname>Mar</surname>
          </string-name>
          .
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7] 'Ontologies',
          <source>in Ontology Learning and Population from Text</source>
          , Springer US,
          <year>2006</year>
          , pp.
          <fpage>9</fpage>
          -
          <lpage>17</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Shi</surname>
          </string-name>
          et al., '
          <article-title>Distributed radiomics as a signature validation study using the Personal Health Train infrastructure'</article-title>
          ,
          <source>Sci Data</source>
          , vol.
          <volume>6</volume>
          , no.
          <issue>1</issue>
          , p.
          <fpage>218</fpage>
          ,
          <string-name>
            <surname>Dec</surname>
          </string-name>
          .
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>