<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards the Creation of the Cardiovascular Magnetic Resonance Quality Assessment Ontology (CMR-QA)?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ernesto Jime´nez-Ruiz</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Valentina Carapella</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Elena Lukaschuk</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nay Aung</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kenneth Fung</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jose Paiva</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mihir Sanghvi</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stefan Neubauer</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Steffen Petersen</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ian Horrocks</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stefan Piechnik</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Information Systems Group, Department of Computer Science, University of Oxford</institution>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Logic and Intelligent Data Group, Department of Informatics, University of Oslo</institution>
          ,
          <country country="NO">Norway</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Oxford Centre for Clinical Magnetic Resonance Research (OCMR), Radcliffe Department of Medicine, University of Oxford</institution>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>William Harvey Research Institute, NIHR Cardiovascular Biomedical Research Unit at Barts, Queen Mary University of London</institution>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>UK Biobank5 is a large scale population study that started in 2006 aimed at improving the understanding, diagnosis and treatment of a wide range of diseases, such as cancer, stroke or cardiac pathologies [2]. The recruitment of volunteers took place at the national level and reached the considerable number of 500,000 volunteers aged 4069 across UK. All volunteers agreed to go through a series of clinical tests and have their health condition checked on follow-up. Different clinical imaging modalities are included in the protocol applied to the volunteers. In particular, a number of Cardiac Magnetic Resonance Imaging (CMR) sequences are employed to evaluate cardiac function: cine-MRI, tag-MRI, T1-mapping and blood flow imaging [3]. The work presented in this paper is related to a specific pilot study of 5,000 CMR scans limited to cineMRI, the most common CMR modality in clinical practice. The results of this pilot study are about to be released to the public together with a first part of the associated data analysis, Data analysis for the 5,000 cine-MRI was carried out by a team of observers from two centres, OCMR and Barts Hospital, and consisted of two parts: image analysis and assessment of image quality. Image analysis is essentially manual delineation of contours (also known as segmentation) of the four chambers of the heart, which then results in the computation of fundamental parameters of cardiac function. Quality assessment of the Cine-MRI scans was carried out through a combination of free-text comments and numerical quality scores. Figure 1 highlights the two components of the analysis. Quality assessment and general data analysis progress was managed through a shared spreadsheet by the team. For the purposes of our work, we focus only on the quality assessment data and the combination of numerical quality scores and free-text annotation. The quality scores alone (1 = optimal, 2 = suboptimal, 3 = not-analysable or ? This paper represents a short but more technical version of the paper “Towards the Semantic Enrichment of Freetext Annotation of Image Quality Assessment for UK Biobank Cardiac Cine MRI Scans” [1] 5 http://www.ukbiobank.ac.uk/</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Background</title>
      <p>not-reliable) only provide a quick overall classification. The free-text annotation is rich
in information but cannot be processed in an easy and efficient manner as the numerical
scores.</p>
      <p>In this ongoing work we aim at employing tools from the Semantic Web for the
efficient structuring of free-text documentation. The semantics of the free-text
annotations, which describe the quality of the image analysis, will be defined via a structured
vocabulary or ontology, which we are going to call CMR-QA (Cardiovascular
Magnetic Resonance Quality Assessment). The aimed semantic layer (ontology, rules and
data) will provide machine-readable data and will be a powerful tool for (i) fast and
efficient processing of the free-text comments; (ii) automatic image quality assessment
from such comments and generation of quality scores; (iii) evaluation of the quality of
the free-text comments in terms of information completeness, ambiguity and
variability; (iv) training purposes (e.g., showing preferred annotation styles for different types
of images); (v) efficient semantic access (i.e. database querying) to the images by the
UK Biobank target users, such as researchers in the field of automatic segmentation, or
clinical researchers who need a specific subset as a control group in their study.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Definition of the semantic layer</title>
      <p>Free-text annotation are prone to variability. Variability is due to natural human
variability (e.g., used nomenclature, differences in point of view), but it can also reflect different
ranges of professional expertise. For example, different observers can correctly flag as
LA off axis, LA out of plane, wrong LA plane, 2Ch out of plane, or LA foreshortened an
image where the plane chosen to acquire a long axis view was not optimally aligned to
measure left atrial (LA) volumes.</p>
      <p>CMR-QA ontology aims at defining the semantics of the free-text annotations in
order to help identify commonalities and disagreement among different annotations.
An ontology may reveal, as a source of variability, the use of different synonyms, or the
use of narrower, broader or even sibling terms for the same kind of annotations. The
use of an ontology (in combination with rules) may also reveal a source of ambiguity
or incompleteness if some required information is not provided (e.g., the annotation LA
off axis should always refer to the cardiac cycle where it was observed).
2.1</p>
      <sec id="sec-2-1">
        <title>The ontology</title>
        <p>
          The CMR-QA will include general knowledge about the domain,6 for example it will
encode that the concepts CineMRI Scan and T1-Mapping Scan are a kind of MRI Scan.
It will also encode more concrete knowledge about the quality assessment process,
for example, wrong image plane orientation is a kind of technical issue. Similarly,
RA off axis is a specific type of wrong image plane orientation, while mistriggering
can be either a kind of technical issue (as an artefact) or a patient-related issue (as a
pathology). The ontology also encodes relationships between concepts, for example the
property has technical issue relates concepts with types of technical issue. Non-logical
knowledge in the form of lexical information can also be added to the ontology; for
example, the ontology may include that RA out of plane is an alternative label for RA
off axis. Equations (
          <xref ref-type="bibr" rid="ref1">1</xref>
          )-(
          <xref ref-type="bibr" rid="ref6">6</xref>
          ) show the formalization of the above knowledge into ontology
axioms.7
        </p>
        <sec id="sec-2-1-1">
          <title>SubClassOf(CineMRI Scan MRI Scan)</title>
          <p>
            SubClassOf(Wrong Orientation Technical Issue)
SubClassOf(RA o axis Wrong Orientation)
SubClassOf(Misstriggering ObjectUnionOf(Technical Issue Patient Issue)) (
            <xref ref-type="bibr" rid="ref4">4</xref>
            )
ObjectPropertyRange(has technical issue Technical Issue)
AnnotationAssertion(alt label RA o axis RA out of plane)
(
            <xref ref-type="bibr" rid="ref1">1</xref>
            )
(
            <xref ref-type="bibr" rid="ref2">2</xref>
            )
(
            <xref ref-type="bibr" rid="ref3">3</xref>
            )
(
            <xref ref-type="bibr" rid="ref5">5</xref>
            )
(
            <xref ref-type="bibr" rid="ref6">6</xref>
            )
2.2
          </p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>The rules</title>
        <p>
          The ontology is being extended with (manually created) rules to infer implicit
knowledge from the annotations. For example if the free-text annotation includes the comment
RA off-axis then the comment is necessarily referred to the Horizontal long axis (HLA)
view. Analogously, the free-text comment basal slice is missing implies a lack of
coverage associated to the short axis (SA) view. The rules will also be used to infer the
quality assessment scores. For example, Lack of coverage will always lead to a quality
score associated to the right and left ventricle equal or greater than 2 (e.g., suboptimal
or unreliable). In addition, rules may also be used to reveal potential incompleteness or
ambiguity. For example, we can classify the technical issue LA off axis as incomplete if
the cardiac cycle is not indicated. Equations (
          <xref ref-type="bibr" rid="ref7">7</xref>
          )-(9) show the formalization of some of
the aforementioned rules.8
6 We could not find any ontology in BioPortal [4] meeting all our requirements, which evidences
the necessity of a more specific ontology in this particular domain.
7 We use OWL functional-style syntax: https://www.w3.org/TR/owl2-syntax/.
8 We use datalog rules (a subset of Prolog). A rule of the form A(x) B(x) ^ R(x; y) ^ C(y)
means that the combination of the concepts B and C via the relationship R (for the given
individuals ‘x’ and ‘y’) implies that the individual ‘x’ is a member (or a type) of the concept A.
has issue in HLA view (mri; issue)
        </p>
        <p>CineMRI Scan(mri) ^
has technical issue(mri; issue) ^ RA o axis(issue)
has RV quality score(mri; 2)</p>
        <p>CineMRI Scan(mri) ^
has technical issue(mri; issue) ^ Lack coverage(issue)
IncompleteIssueDe nition(issue)</p>
        <p>
          RA o axis(issue) ^
not observed in cardiac cycle(issue; cc)
(
          <xref ref-type="bibr" rid="ref7">7</xref>
          )
(
          <xref ref-type="bibr" rid="ref8">8</xref>
          )
(9)
2.3
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>The data</title>
        <p>The free-text content in the spreadsheet is rich in information but cannot be processed
in an easy and efficient manner. In this work we are developing custom named entity
recognition (NER) techniques to transform the free-text comments into semantically
rich data according to the CMR-QA. For example, the free-text comment “basal slice is
missing. wrong planes ra/la” associated to a CineMRI Scan leads to the following seven
ontology facts:</p>
        <sec id="sec-2-3-1">
          <title>CineMRI Scan(mri)</title>
        </sec>
        <sec id="sec-2-3-2">
          <title>Lack coverage(issue1)</title>
        </sec>
        <sec id="sec-2-3-3">
          <title>RA o axis(issue2)</title>
        </sec>
        <sec id="sec-2-3-4">
          <title>LA o axis(issue3) has technical issue(mri; issue1) has technical issue(mri; issue2) has technical issue(mri; issue3)</title>
          <p>Where mri, issue1, issue2 and issue3 are the identifiers of the ontology data extracted
from the annotations. For example, issue1 is associated to the chunk of text “basal slice
is missing” and represents a concrete individual of an observed Lack of coverage (i.e.,
issue1 is a member of the concept Lack coverage, that is, Lack coverage(issue1)).</p>
          <p>
            These facts together with the ontology and rules introduced in the previous
section will lead via reasoning to new (implicit) facts. For example, using the ontology
axioms in Equations (
            <xref ref-type="bibr" rid="ref1">1</xref>
            )-(
            <xref ref-type="bibr" rid="ref3">3</xref>
            ) we can infer the new facts Technical issue(issue1) and
MRI Scan(mri). The rules in Equations (
            <xref ref-type="bibr" rid="ref7">7</xref>
            )-(9) will also infer new knowledge (e.g.,
has RV quality score(mri; 2) via rule (
            <xref ref-type="bibr" rid="ref8">8</xref>
            ) and has issue in HLA view (mri; issue2)
via rule (
            <xref ref-type="bibr" rid="ref7">7</xref>
            )) or raise potential warnings with regard to ambiguity/incompleteness (e.g.,
IncompleteIssueDe nition(issue2) and IncompleteIssueDe nition(issue3) via
ontology rule (9)).
3
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Discussion</title>
      <p>In this preliminary study we have set the first axioms and rules to define a structured
vocabulary associated to the quality assessment data. Preliminary experiemnts to
automatically extract ontology data from the free-text annotations have also been conducted.</p>
      <p>Related Work. There have also been recent efforts in adding a semantic layer to
describe the information within a biobank. Andrade et al. [5] envisaged the benefits of
using ontologies for querying and searching the information in a biobank and across
biobanks. Muller et al. [6] presents and updated overview of the state of the art and
open challenges for the description and interoperability across biobanks where the use
of Semantic Web technologies will play a key role. Examples of concrete Semantic
Web-based solutions in biobanks can also be found in [7, 8].</p>
      <p>Future Work. As immediate future work, we plan to complete the CMR-QA, define
the necessary ontology rules and finalize the implementation of the techniques to text
mine the comments to extract ontology facts. We also aim at integrating CMR-QA with
other parts of UK Biobank where different ontologies and controlled vocabularies may
be used, and with standards medical vocabularies like SNOMED CT. Furthermore, we
will perform an extensive evaluation to analyse the correctness of our approach. For
example, (automatic) validation will be carried out by comparing the generated automatic
scores using rules with those manually assigned by the observers as part of their quality
assessment.</p>
      <p>Acknowledgements. SEP, SN and SP acknowledge the British Heart Foundation (BHF)
for funding the manual analysis to create a cardiovascular magnetic resonance
imaging reference standard for the UK Biobank imaging resource in 5,000 CMR scans
(PG/14/89/31194, PI Petersen, 6/2015 to 5/2018). SKP, VC and SN were additionally
funded by the National Institute for Health Research (NIHR) Oxford Biomedical
Research Centre based at The Oxford University Hospitals Trust at the University of
Oxford. EJR and IH were funded by the EC FP7 project Optique, and the EPSRC projects
Score!, ED3 and DBOnto. EJR was also funded by the Centre for Scalable Data Access
(SIRIUS) and the RCN project BigMed.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Carapella</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          , et al.:
          <article-title>Towards the Semantic Enrichment of Free-text Annotation of Image Quality Assessment for UK Biobank Cardiac Cine MRI Scans</article-title>
          .
          <source>In: MICCAI Workshop on Large-scale Annotation of Biomedical data and Expert Label Synthesis (LABELS)</source>
          .
          <article-title>(</article-title>
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Petersen</surname>
            ,
            <given-names>S.E.</given-names>
          </string-name>
          , et al.:
          <article-title>Imaging in population science: cardiovascular magnetic resonance in 100,000 participants of UK Biobank - rationale, challenges and approaches</article-title>
          .
          <source>Journal of Cardiovascular Magnetic Resonance</source>
          <volume>15</volume>
          (
          <issue>1</issue>
          ) (
          <year>2013</year>
          )
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Schulz-Menger</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , et al.:
          <article-title>Standardized image interpretation and post processing in cardiovascular magnetic resonance: Society for Cardiovascular Magnetic Resonance (SCMR) board of trustees task force on standardized post processing</article-title>
          .
          <source>Journal of Cardiovascular Magnetic Resonance</source>
          <volume>15</volume>
          (
          <issue>1</issue>
          ) (
          <year>2013</year>
          )
          <fpage>1</fpage>
          -
          <lpage>19</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Fridman</given-names>
            <surname>Noy</surname>
          </string-name>
          ,
          <string-name>
            <surname>N.</surname>
          </string-name>
          , et al.:
          <article-title>BioPortal: ontologies and integrated data resources at the click of a mouse</article-title>
          .
          <source>Nucleic Acids Research</source>
          <volume>37</volume>
          (
          <string-name>
            <surname>Web-Server-Issue)</surname>
          </string-name>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Andrade</surname>
            ,
            <given-names>A.Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kreuzthaler</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hastings</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krestyaninova</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schulz</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Requirements for semantic biobanks</article-title>
          .
          <source>Stud Health Technol Inform</source>
          .
          <volume>180</volume>
          (
          <year>2012</year>
          )
          <fpage>569</fpage>
          -
          <lpage>573</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6. Mu¨ ller, H., et al.:
          <article-title>State-of-the-Art and Future Challenges in the Integration of Biobank Catalogues</article-title>
          . In: Smart Health - Open Problems and Future Challenges. (
          <year>2015</year>
          )
          <fpage>261</fpage>
          -
          <lpage>273</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Pathak</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , et al.:
          <article-title>Applying semantic web technologies for phenome-wide scan using an electronic health record linked Biobank</article-title>
          .
          <source>J. Biomedical Semantics</source>
          <volume>3</volume>
          (
          <year>2012</year>
          )
          <fpage>10</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Brochhausen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , et al.:
          <article-title>Developing a semantically rich ontology for the biobankadministration domain</article-title>
          .
          <source>J. Biomedical Semantics</source>
          <volume>4</volume>
          (
          <year>2013</year>
          )
          <fpage>23</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>