<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>K. R. Anandan);
bhuvanaj@ssn.edu.in (B. Jayaraman); mirnalineett@ssn.edu.in (M. T. N. T. Thai)
~ https://www.ssn.edu.in/staf-members/dr-j-bhuvana/ (B. Jayaraman);
https://www.ssn.edu.in/staf-members/dr-t-t-mirnalinee/ (M. T. N. T. Thai)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Simple Neural Network based TB Classification</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Anirudh Anand</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Karthik Raja Anandan</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bhuvana Jayaraman</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mirnalinee Thanga Nadar Thanga Thai</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Sri Sivasubramaniya Nadar College of Engineering</institution>
          ,
          <addr-line>Chennai, Tamil Nadu</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>Analysis of images is a vitally important task in medical applications. It helps the prompt detection and categorization of diseases, among others. This paper depicts a intuitive and simple approach to classify the Tuberculosis found in the 3D CT-images of patients' chests as a part of the ImageCLEF2021 challenge. A simple shallow neural network is employed with three layers. The model is trained using augmented images of the dataset. The proposed model is tested for it's accuracy and kappa coeficient to obtain the degree to which the model correctly classifies the chest images.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Computed Tomography</kwd>
        <kwd>Tuberculosis classification</kwd>
        <kwd>Neural Network</kwd>
        <kwd>Tensorflow</kwd>
        <kwd>Image Classification</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        data set. Salient features of Python’s NumPy [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] are implemented to realize the same.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Task and Dataset</title>
      <p>As mentioned above, the broad objective of ImageCLEF2021 is to classify 3D CT-images of
patients’ lungs into one of 5 TB categories, namely: (1) Infiltrative, (2) Focal, (3) Tuberculoma,
(4) Miliary and (5) Fibro-cavernous. A dataset containing chest CT scans of 1338 TB patients is
used. 917 images for the Training (development) data set and 421 for the Test set. Additionally,
metadata is provided for some images.</p>
      <sec id="sec-2-1">
        <title>2.1. Multi-dimensional neuroimaging data</title>
        <p>
          For all patients a single 3D CT image with an image size per slice of 512×512 pixels and number of
slices being around 100 is provided. All the CT images are stored in NIFTI file format with .nii.gz
ifle extension (g-zipped .nii files). This file format stores raw voxel intensities in Hounsfield
units (HU) as well the corresponding image metadata such as image dimensions, voxel size in
physical units, slice thickness, etc. Python’s Nibabel package [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] is used to read the .nii files.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Methodologies</title>
      <sec id="sec-3-1">
        <title>3.1. Data preprocessing</title>
        <p>
          Nibabel library was used to load the zipped Nifti fileformat of CT-scan images and return them
as NumPy arrays. The values of the NumPy array is normalized based on threshold values of
hounsfield units. The given NumPy array contains raw voxel intensities in Hounsfield units
(HU). The Hu for Air is -1000 and Hu for tissues is 500.And those intensities higher than 500
makes up the bones in the image.Here we are taking into those account only those between
-1000 to 500 and used to normalize the voxel(Volume pixel) values of the NumPy array to the
range [0 to 1]. This is scaled down to 128*128*64 image size from 512*512*113. The resulting
scaled down 3D-array is then rotated to randomize the orientation. Since the entire data set can’t
be read into memory in one go,it is read in batches of 20 and is then sent for training. A linear
interpolation [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] operator from SciPy was used to scale down the image sizes.Since the CT-scan
image was already given in higher resolution , it is assumed that the features and edges would
be retained after a simple Linear interpolation.It also results in faster processing.Normalizing
data is said to speed up the learning process and leads to faster convergence.
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Image Augmentation</title>
        <p>To increase the count of training set and to enhance variability into the training set, the obtained
dataset is mapped to functions that rotate the images by degrees of 5 to create augmented data
that is used for training. The training batch size is set as 20 (the maximum possible size without
getting an out of memory error). Validation set contains equal number of all category dataset,
to produce an unbiased accuracy. However, one additional instance each, was added to class 3
and class 4 to tune the model well during the training and hence the validation set size is 27.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Model Architecture</title>
        <p>A simple shallow neural network model has been designed to classify the Tuberculosis found in
the 3D CT-images. The proposed model has three layers in it as shown in Fig. 1 The first layer
accepts the preprocessed images and are flattened before passing them to two fully connected
layers.</p>
        <p>The 128*128*64 3D image is flattened to (128*128*64) single dimensional vector and is passed
through a dense layer of 600 (picked at random after working with lower dimensions) neurons
and the output of this layer is given to the last layer which returns a one hot encoded vector.</p>
        <p>The first fully connected layer uses ’relu’ as activation function, since that activation function
handles the problem of vanishing gradient. The last classification layer has 5 nodes
corresponding to the classes of Tuberculosis and employs sigmoid activation function, that produces
outcome similar to the probabilistic values pertaining to the classes.</p>
        <p>The loss function used is Binary cross entropy which is optimized using the Stochastic
Gradient Descent optimizer. The model performance during training is evaluated using the
accuracy metric.</p>
      </sec>
      <sec id="sec-3-4">
        <title>3.4. Convolutional Neural Network (CNN) Model</title>
        <p>The proposed model for Tuberculosis classification is arrived after exploring another model
using the Convolutional layers called as Convolutional Neural Network (CNN). The CNN model
has been designed with 4 convolutional layers, each followed by max pooling layer to reduce
the spatial dimension of the images. The convolutional layers extract the features from the
input images and are fed to 2 fully connected layers. Batch normalization is done to avoid model
over fitting. The model configuration and parameter details are shown in Fig. 2.</p>
        <p>CNN model used the same pre-processing techniques followed by Neural Network model.
This CNN model did not show any promising results while training when compared to the
Neural Network model explained in section 3.3. The average validation accuracy measured was
only 0.15. We suspect that lack of data can attributed to this poor accuracy. Hence we had to
alter our model to a much simpler neural network that can work with smaller amount of data.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Experiments and Results</title>
      <sec id="sec-4-1">
        <title>4.1. Hardware used</title>
        <p>Google Colab notebook was used to train the model. A general purpose RAM size of 8GB was
alloted with a 2.3GHz Intel Xenon CPU.
4.2. Code</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.3. Result</title>
        <p>
          Implementation is done using Python and the URL to the code is shared below. https://colab.
research.google.com/drive/1wbTPPOn2AF72OMnCkchTcQl2Cd87KeFR?usp=sharing [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ].
The two models are trained for 20 epochs, accuracy metric is used to study the performance of
the model during training. The Metric for both the models are given in Table 1. Complex deep
learning networks namely ResNet, GoogleNet have achieved around 0.4033 in 2017 imageclef.
So keeping them into account, we have tried out to build a simpler Neural Network model for
classifying and have achieved validation accuracy of 0.20. The Neural network model was alone
out of the two experiments, submitted to the ImageCLEFmedical task for evaluation.
        </p>
        <p>Model
Neural Network
CNN</p>
        <p>Training Accuracy
0.647
0.566</p>
        <p>Validation Accuracy
0.20
0.15</p>
        <p>
          The proposed model has obtained a testing accuracy of 0.221 and a kappa value of 0.038
as reported in Table 2. These metrics were used for ranking and placed us in the ninth place
ImageClef 2021 TB classification [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] challenge.
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>The crux of the JBTTM’s submission is based on simple and shallow neural networks (with an
input layer, single hidden layer and an output layer). Other model were such as the 3D CNN
model were also experimented. They weren’t selected due to their low accuracy. The team’s
submission placed it ninth out of a total of eleven participant teams. A rigorous assessment
of the submission showed that the model can be improved by adding more meaningful layers
and/or adding more neurons per layer in such a way that the model doesn’t become intractable.
When compared to previous year’s results, the submission of JBTTM and other teams shows a
steady improvement in the accuracy.</p>
      <p>Figure 2: CNN Architecture Summary</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>N.</given-names>
            <surname>Fogel</surname>
          </string-name>
          ,
          <article-title>Tuberculosis: A disease without boundaries (</article-title>
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <article-title>[2] Stop tb partnership - fact sheet: The missing 3 million (</article-title>
          <year>2019</year>
          ). URL: http://www.stoptb.org/ assets/documents/resources/factsheets/StopTBinfographicMissing3Million.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Assili</surname>
          </string-name>
          ,
          <article-title>A review of tomographic reconstruction techniques for computed tomography https://arxiv</article-title>
          .org/abs/
          <year>1808</year>
          .09172 (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>B.</given-names>
            <surname>Ionescu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Muller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Peteri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. Ben</given-names>
            <surname>Abacha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sarrouti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Demner-Fushman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Hasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kozlovski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Liauchuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Dicente</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Kovalev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Pelka</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. G. S. de Herrera</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Jacutprakart</surname>
            ,
            <given-names>C. M.</given-names>
          </string-name>
          <string-name>
            <surname>Friedrich</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Berari</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Tauteanu</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Fichou</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Brie</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Dogariu</surname>
            ,
            <given-names>L. D.</given-names>
          </string-name>
          <string-name>
            <surname>Ştefan</surname>
            ,
            <given-names>M. G.</given-names>
          </string-name>
          <string-name>
            <surname>Constantin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Chamberlain</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Campello</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>T. A.</given-names>
          </string-name>
          <string-name>
            <surname>Oliver</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Moustahfid</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Popescu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Deshayes-Chossart</surname>
          </string-name>
          ,
          <article-title>Overview of the ImageCLEF 2021: Multimedia retrieval in medical, nature, internet and social media applications, in: Experimental IR Meets Multilinguality</article-title>
          , Multimodality, and
          <string-name>
            <surname>Interaction</surname>
          </string-name>
          ,
          <source>Proceedings of the 12th International Conference of the CLEF Association (CLEF</source>
          <year>2021</year>
          ),
          <source>LNCS Lecture Notes in Computer Science</source>
          , Springer, Bucharest, Romania,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Kozlovski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Liauchuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Dicente Cid</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Kovalev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          , Overview of ImageCLEFtuberculosis 2021 -
          <article-title>CT-based tuberculosis type classification</article-title>
          ,
          <source>in: CLEF2021 Working Notes, CEUR Workshop Proceedings</source>
          , CEUR-WS.org &lt;http://ceur-ws.
          <source>org&gt;</source>
          , Bucharest, Romania,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Santoro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Raposo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. G.</given-names>
            <surname>Barrett</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Malinowski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Pascanu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Battaglia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lillicrap</surname>
          </string-name>
          ,
          <article-title>A simple neural network module for relational reasoning</article-title>
          ,
          <source>arXiv preprint arXiv:1706.01427</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7] What is numpy? https://numpy.org/doc/stable/user/whatisnumpy.html (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <article-title>[8] Nibabel access a cacophony of neuro-imaging file formats https://nipy</article-title>
          .org/nibabel/ (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Amanatiadis</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Andreadis,</surname>
          </string-name>
          <article-title>A survey on evaluation methods for image interpolation</article-title>
          ,
          <source>Measurement Science and Technology</source>
          <volume>20</volume>
          (
          <year>2009</year>
          )
          <fpage>104015</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>H.</given-names>
            <surname>Zunair</surname>
          </string-name>
          ,
          <article-title>3d image classification from ct scans https://keras</article-title>
          .io/examples/vision/3D_ image_classification (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>