<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Combining Color with Spatial and Temporal Position of the Endoscopic Capsule for Improved Topographic Classification and Segmentation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>M. Coimbra</string-name>
          <email>miguel.coimbra@ieeta.pt</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>J. Kustra</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>P. Campos</string-name>
          <email>pcampos@ieeta.pt</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>J.P. Silva Cunha</string-name>
          <email>jcunha@det.ua.pt</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>IEETA institute and the Fundação para a Ciência e Tecnologia (grant nr. Soares of the gastroenterology department of Santo António General Hospital in Porto</institution>
          ,
          <country country="PT">Portugal (</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>-Capsule endoscopy is a recent technology with a clear need for automatic tools that reduce the long exam annotation times of exams. We have previously developed a topographic segmentation method, which is now improved by using spatial and temporal position information. Two approaches are studied: using this information as a confidence measure for our previous segmentation method, and direct integrating of this data into the image classification process. These allow us not only to automatically know when we have obtained results with error magnitudes close to human errors, but also to reduce these automatic errors to much lower values. All the developed methods have been integrated in the CapView annotation software, currently used for clinical practice in hospitals responsible for over 250 capsule exams per year, and where we estimate that the two hour annotation times are reduced by around 15 minutes.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Index Terms— Endoscopic capsule, image classification,
biomedical engineering, medical imaging</p>
    </sec>
    <sec id="sec-2">
      <title>I. INTRODUCTION</title>
      <p>
        The clinical importance of the endoscopic capsule is now
solidly established in literature: Iddan [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], Oureshi [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], etc.
Due to space limitations, we refer to our previous work [
        <xref ref-type="bibr" rid="ref3 ref4">3,4</xref>
        ],
for more extensive capsule details and clinical importance
information. All this attempts to solve an important limitation
of the endoscopic capsule, excessively long annotation times.
Currently it takes about 2 hours to fully view, annotate an
exam and write its corresponding report. Our clinical studies
show that the task of topographic segmentation is both
difficult (the median error performed by three senior capsule
specialists was about 400 images) and time-consuming
(around 15 minutes can be saved by automation).
      </p>
      <p>The main contribution of this paper is the improvement of
our previous topographic segmentation methods using color
and texture, by incorporating not only temporal but also
spatial position information in the image classification
process.</p>
      <p>The ultimate objective of the presented methods is to
reliably divide the video of the gastrointestinal tract into its 4
constituent parts (entrance, stomach, small intestine, large
intestine) and thus determine its corresponding junctions
(esogastric junction, pylorus, ileo-cecal valve).</p>
      <p>A. Capsule Position and Velocity</p>
      <p>We can theoretically estimate the spatial position of a
capsule via antenna signal triangulation. We have selected 47
capsule exams where a clinical specialist manually annotated
the temporal location of the pylorus (tPYL) and of the ICV (tICV)
in the video using the CapView annotation software. We’ve
then used our automatic topographic segmentation algorithm
to determine these same temporal locations. Using our 2D
position information, we can then obtain the corresponding
spatial locations: xPYL, yPYL, xVIC, yVIC etc. For comparison
purposes, these were normalized. Besides analysing 2D
position information, we have looked at average capsule
displacement velocity (module of the displacement vector
between two points with temporal references t and t+1).</p>
      <p>B. Topographic Segmentation Algorithm</p>
      <p>
        Our previously developed automatic topographic
segmentation method, from now on referred as TSA, is
described in Coimbra [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        C. Spatial Information as a Confidence Measure
Two high-confidence areas were defined, one for the pylorus
and another for the ILC. We have measured the median
segmentation error SEz for all marks (z12 - eso-gastric
junction; z23 – pylorus; z34 – ileo-cecal valve), and for all
exams SE, whose junctions are inside and outside these areas,
D. Integrating Spatial and Temporal Information for
Classification
An alternative way of using this information is to use it
directly for individual image classification. Our previous
method trained 4 SVM classifiers, one for each zone, which
determine the topographic section each image belongs to as
the classifier with the highest positive distance to the SVM
hyperplane (see [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] for details). We can however, use these
distances to build a feature vector for each image, along with
additional information such as spatial and temporal location.
Our new feature vector F is now defined as:
      </p>
      <p>U</p>
      <p>
        F &gt; x, y, Z1 , Z 2 , Z 3 , Z 4 ,V , t@ (1)
where x and y are the normalized spatial location coordinates
(1,2), Z1, Z2 Z3 Z4 are the SVM classifier results [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] (distances
to SVM hyperplanes), V is the spatial velocity, and t the
temporal location in number of frames. The combination of
these different features into a single vector requires that all
coefficients are previously normalized.
      </p>
      <p>A variety of well-known distances was used for classification
(L1 Norm, Euclidean, Mahalanobis). Finally, we have
measured the relevance of each coefficient for the
segmentation process using a step-wise elimination analysis.</p>
    </sec>
    <sec id="sec-3">
      <title>III. RESULTS</title>
      <p>1
0.5
-1
Fig. 1. Spatial distribution of correct (green) and incorrect (blue) estimations.
Points in high confidence areas are highlighted with a black bounding box.
We can observe that most correct detections fall into high-confidence areas
while incorrect ones are more distributed over the whole 2D space.</p>
      <p>Accuracy
58 %
92 %
Mean
Err.
3966
2096</p>
      <p>ICV
and results presented in section 3 have showed that this
information is indeed useful as a confidence measure for
automatic segmentation results.</p>
      <p>An analysis of Table 1 and Figure 1 shows that
highconfidence areas contain almost all correct estimations, and
low-confidence areas mainly contain incorrect estimations.
1
0.5
-0.5
y
Z2
Z3
Z4
V
82.3%
81.8%
81.9%
80.9%
80.8%
83.2%
80.4%
82.3%
83.1%
81.2%
81.0%
82.6%
82.5%
82.4%
79.9%</p>
      <p>Median Segmentation Error
83.5% 82.8% 82.2%
82.1%</p>
      <p>79.2%
83.1%
83.1%
81.7%
82.8%
83.4%
83.5%
83.1%
79.9%
82.8%
82.8%
80.8%
82.8%
79.3%
82.2%
82.8%
79.5%
82.2%
82.2%
79.9%
82.2%
70.3%
82.2%
82.1%
66.4%
82.1%
82.1%
79.2%
82.1%
69.8%
82.1%
82.1%
63.0%
79.2%
79.2%
79.2%
79.2%
68.7%
79.2%
79.2%
54.2%</p>
    </sec>
    <sec id="sec-4">
      <title>IV. DISCUSSION</title>
      <p>Results show that doctors can trust that automatic
segmentation errors in high-confidence areas are as low as
human ones. Including other information has allowed us to
improve segmentation results significantly. Step-wise
elimination analysis has shown us that the most relevant
features for segmentation are capsule temporal position, and
the color recognition of the entrance and the small intestine
topographic sections. It has also shown us that spatial location
is not a relevant factor for individual image classification.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>G.</given-names>
            <surname>Iddan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Meron</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Glukhovsky</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Swain</surname>
          </string-name>
          , “Wireless Capsule Endoscopy”, in Nature,
          <year>2000</year>
          ,
          <volume>405</volume>
          , pp.
          <fpage>417</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2] 2.
          <string-name>
            <given-names>W.A.</given-names>
            <surname>Qureshi</surname>
          </string-name>
          , “
          <article-title>Current and future applications of the capsule camera”</article-title>
          ,
          <source>in Nature</source>
          , vol.
          <volume>3</volume>
          ,
          <issue>2004</issue>
          , pp.
          <fpage>447</fpage>
          -
          <lpage>450</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <fpage>4</fpage>
          .
          <string-name>
            <given-names>M.</given-names>
            <surname>Coimbra</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.P.</given-names>
            <surname>Silva</surname>
          </string-name>
          <string-name>
            <surname>Cunha</surname>
          </string-name>
          , “
          <article-title>MPEG-7 visual descriptors - Contributions for automated feature extraction in capsule endoscopy”</article-title>
          ,
          <source>in IEEE Transaction on Circuits and Systems for Video Processing</source>
          , vol.
          <volume>16</volume>
          /5, 2006, pp.
          <fpage>628</fpage>
          -
          <lpage>637</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <fpage>6</fpage>
          .
          <string-name>
            <given-names>M.</given-names>
            <surname>Coimbra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Campos</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.P.</given-names>
            <surname>Silva</surname>
          </string-name>
          <string-name>
            <surname>Cunha</surname>
          </string-name>
          , “
          <article-title>Topographic Segmentation and Transit Time Estimation for Endoscopic Capsule Exams”</article-title>
          ,
          <source>in Proc. of IEEE ICASSP</source>
          <year>2006</year>
          , Toulouse, France,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>