<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Detection of Bipolar Disorder Using Machine Learning with MRI</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>R Sujatha</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>K Tejesh</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>H Krithi</string-name>
          <email>krithi.h2017@vitstudent.ac.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>H Rasiga Shri</string-name>
          <email>hrasiga.shri2017@vitstudent.ac.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>International Semantic Intelligence Conference (ISIC)</institution>
          ,
          <addr-line>February</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>School of Information Technology and Engineering, Vellore Institute of Technology</institution>
          ,
          <addr-line>Vellore, Tamilnadu</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Workshop Proce dings</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <fpage>25</fpage>
      <lpage>27</lpage>
      <abstract>
        <p>Bipolar disorder is a mental ailment caused by maximal mood swings with emotional highs and lows. Nowadays, this has become the most common abnormality related to mental health and furthermore it is ignored by people of all age groups. Bipolar disease is generally heritable but not all siblings of the family will be having it though, and will be having same genetics and the factors which can be risky. Here we use random forest algorithm, along with the Mag- netic Resonance Imaging (MRI) information. The utility of these irregularities in recognizing individual bipolar disorder patients from state of mind issue or health controls define patients dependent on their illness. Here we use machine learning algorithms like Random forest algorithm and CNN-mdrp(multimodal disease risk prediction) for the accuracy .We give the risk factor and stage of the healthy patient with the attributes we collected from the MRI. We use a trained dataset and machine learning algorithms mentioned above to get the output. Voxel-Based Morphometry (VBM) will be used to dividing and pre-processing the MRI infor- mation obtained. To see the changes in Gray Matter (GM) and White Matter (WM) of the diferent data groups individually, a simple equation is use and also the Principle Component Analysis will be used and The project gives you the output showing that CNN MDRP with random forest has high accuracy than other algorithms in bipolar disease prediction. Magnetic Resonance Imaging, Random forest algorithm, Gray mat- ter, White matter, Voxel based morphometry, CNN-mdrp system to identify the mental conditions of an individual. provising the algorithms to include more productivity healthy survival. Therefore, even though we can't pre- ac- curacy in fact better than CNN-udrp algorithm.</p>
      </abstract>
      <kwd-group>
        <kwd>(multimodal disease risk prediction)</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <sec id="sec-1-1">
        <title>Mental sickness is one of the most dangerous and lifethreatening ailments. One must always take good care of their mental state of mind. There are various proposed</title>
      </sec>
      <sec id="sec-1-2">
        <title>These systems were developed using some combination</title>
        <p>of machine learning algorithms working with collected
data sets to train and test the model.[2]</p>
      </sec>
      <sec id="sec-1-3">
        <title>However, this existing system has some defects that</title>
        <p>are to be rectified. The collected samples for the data set
is insuficient. The attributes that are used in the data
sets cannot be a fixed one, it changes with person and
the disease they sufer from. [12]</p>
      </sec>
      <sec id="sec-1-4">
        <title>Detection of Bipolar disease at an early stage enables</title>
        <p>patients to have a much higher chances of recovery and
vent it from afecting us, we can certainly detect this at
an early stage so as to provide the appropriate medical
help at the right time to the right people. [5]</p>
      </sec>
      <sec id="sec-1-5">
        <title>The main motive of this paper is to develop a system</title>
        <p>LGOBE</p>
        <p>https://github.com/TEJESH-K (K. Tejesh);
https://github.com/krithi0506 (H. Krithi);
https://github.com/Rasiga-Shri (H. Rasiga Shri)
that detects Bipolar diseases with more accuracy and
less cost using CNN-MDRP algorithm and random forest
classification which is proven to be giving more accurate
results.</p>
      </sec>
      <sec id="sec-1-6">
        <title>There is no accuracy in the present algorithms. Im</title>
        <p>of the framework subsequently improving its working
isn’t finished. There is no appropriate online bipolar
determination framework that are utilized in clinics and
hospitals as of now. [8]</p>
        <p>The proposed random forest Machine Learning
Algorithm for classifying the data not only will predict the
diseases but also its sub diseases, predicting sub diseases
increases the accuracy of the system and the Deep learning
algorithm CNN-mdrp is used for the accuracy prediction
and from the results it shows that CNN-mdrp has high</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Existing System</title>
      <p>The current system for Bipolar disorder disease
prediction system uses random forest algorithm which actually
determines the features present in the data set and makes
tained testing samples are deliberately compared with
the decision attributes at each level to test if a patient
is screened positive or negative. Machine could predict
only the disease but cannot predict sub types and risk
stage of the diseases. It fails to anticipate and predict
all potential states of the people. In the past, System
took care of just organized information. The standing ence Electrode (RE). The detection of Lithium is based on
associations arrange a blend of ML algorithms which are presence of potentiometric Ion-Selective Electrode (ISE)
reasonably good at predicting illnesses or the disease. accompanied by a nanostructured solid contact. [4]
[13] But the limitations with the common framework In this paper, they raised a SVM model using the NumPy
exists. A machine can predict and explain a disease yet library in python. They considered the structural and
can’t speak about the sub kinds of the diseases caused by functional attributes taken from the MRI report of the
the already existing disease. [18] patients. They also used many unique features collected
from the MRI report of the afected individuals. This new
invention of combining the structural and the
function3. Literature Survey als attributes of the brain anatomy improved the overall
eficiency of the system. It also increased the accuracy
This paper solves many problems related to Bipolar dis- of the system.[11]
order detection. Many a times the doctors and patients In this paper, they examined the neurotrophic factors
mistake Major Depressive Disorder as Bipolar Disorder, of the individuals, and analysed if these factors will
inthis paper shows up many solutions for the doctors and lfuence in the timely detection of the Bipolar disorder.
patients to get diagnosed and treated with the disease. They employed the model-based algorithm for the
iden</p>
      <p>In another system, the patients with neuroanatomical tification. As a result of this paper neurotrophic factors
abnormalities were classified using Three- Dimensional were successfully found to assess the disorder.[26]
magnetic resonance imaging (3DMRI). This system also In this paper, they have used a sensitized T-shirt that
focuses on the unafected individuals who are blood- re- records the patient’s Neurotrophic factors. Here, decision
lated to the patients sufering from Bipolar disorder.[6] tree algorithm is employed to classify the patients
suf</p>
      <p>This paper uses machine learning algorithm to screen fering from Bipolar disorder and the healthy ones. This
bipolar disorder by using Mood Disorder Questionnaire T-shirt accurately predicted the Neurotrophic factors of
(MDQ), with decision tree algorithm. the data set is fed the afected individuals and predicted the border value
into the decision tree classifier which determines the range of the factors.[16]
significant features in dataset and make it decision factor In this paper, the bipolar state relapse disorder is
identiat that level of decision tree.[14] ifed among the patients using certain attributes collected</p>
      <p>In this paper, early detection of bipolar disorder is from patient’s smartphones. Nowadays, smartphones
done by using screening question- naire data for around are capable of monitoring the individual’s heart rate,
300 respondents and it served as a knowledge base to be blood pressure rate etc. This disorder was successfully
processed using back propagation algorithm, which was detected.[21]
the drawback in the previous research paper.[9] The other paper, introduced an approach to study the</p>
      <p>In this paper, the patients sufering from BD were ana- mood disorder with respect to the patterns present in
lyzed and the change in their mental states was identified. the emotions of the afected people. They introduced
A dataset was created by monitoring their characteristics LASM - Latent Efective model to locate the connections
and the features was gathered to this disorder. [10] in the emotions of people. They used six videos for the</p>
      <p>In this paper ,1 dimensional time domain data of the detection of the disorder. These videos were the recorded
fNIRS, is acquired while preparing tasks, which is used emotional videos. The study concluded that this method
to train a set of neural networks for the diagnosis of gave more accuracy than the existing one.[15]
common mood disorder, the Bipolar Disorder. With this In this paper, this system detects the mood disorder.
the healthy individuals and the afected ones were classi- The system works with LSTM- based approach for
modifed and deep learning algorithms were employed in the eling the long-range speech of people, it analyses the
detection of the disorder.[7] mood of the person from their speech. The database
con</p>
      <p>In this paper, 65 individuals were isolated and exam- tained of recorded speeches of diferent people mixed
ined. Their activities were rec- orded. Here, Voxel-based with diferent kinds of emotions. The SVM algorithm is
morphometry was used to investigate the brain anatomy also employed for the detection patients sufering mood
from the patient’s MRI. The study included 26 Bipolar disorder.[23]
patients and 38 Healthy individuals. The t-test based sta- This paper identifies the periods of depression of the
tistical method was also employed to group the afected patients sufering from bipolar disorder using the
moveindividuals.[3] ments recorded from mobile location. They have used</p>
      <p>In the following paper, a complete system is been de- the quadratic linear regression model to monitor the
charveloped for the Therapeutic Drug Monitoring (TDM). acteristics of the afected individuals. This also facilitates
Here, the disease is diagnosed from the lithium present the doctors to identify those people who are in the need
in the sweat of people afected by this disorder. This of critical care.[20]
platform incorporates paper fluidics and the stable Refer- In this paper they monitored the patients of BD and
concluded that the report generated from the smartphone
data matched up with the depressive states. This paper
resulted that high methodological rigor along with large
sample of patients afected with</p>
      <p>bipolar disorder having manic symptom addressing
doze factors before being implemented using monitoring
tool.[19]</p>
    </sec>
    <sec id="sec-3">
      <title>4. Methodology</title>
      <p>4.1. DICOM
Expanded as Digital imaging and communications in
medicine, is used as a standard for taking medical data
as input .For example when we try to integrate medical
imaging appliances , we use the DICOM standard. And
here for taking the inputs from the MRI scans, we use
this standard for getting the data and using it for next
step. It is an international standard to send , save and
print the medical imaging data, and The National
Electrical Manufacturers Association(NEMA) holds the all
the copyrights of the above standard After the process of
getting the data from the MRI scanning , we are going
to covert into NIFTI and segment the data obtained into
white matter(WM) and grey matter(GM) which is later be
undergoing normalization process and then the obtained
output will be going through smoothing will generate a
3D Mask finally, which is used in the next process.</p>
      <sec id="sec-3-1">
        <title>4.2. Voxel based Morphometry</title>
        <sec id="sec-3-1-1">
          <title>Voxel based Morphometry technique is used to process</title>
          <p>the identified defects in the brain and prepossess them.
The Voxel-based morphometry computes the changes
that occurs in the structure of the brain that is obtained
from the magnetic resonance imaging (MRI) . This is used
to identify the structural changes that occurs in brain of
the afected individuals. The brain of the afected patients,
consist of greater gray matter volume in the left temporal
lobe and central gray matter structures bilaterally.[1] So,
Voxel Based Morphometry is used to identify structural
changes in voxel-based comparison of multiple brain MRI
images.</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>4.3. Principle Component Analysis</title>
        <p>The Figure 1 represents the principal components of
bipolar that contains diferenti- ated voxels got by using
certain methods are represented as Voxel of Interests. PCA
is used to decrement the proportions of the Voxel that are
obtained from Interests data. Here PCA is a
dimensionality reduction method used in large datasets to reduce
dimen- sions by transforming large set of variables into
smaller set which contains almost every information in
the large sets. We are doing this because smaller datasets
makes the analysis process of data much easier and it
makes it faster for machine learning algo- rithms as there
is no extra variables to process. [22]</p>
      </sec>
      <sec id="sec-3-3">
        <title>4.4. Correlation of Matrix of the Dataset</title>
        <p>The Figure 2 represents the correlation matrix which
gives us the information about the correlation
coeficients between the values taken, above colorful matrix
represents all the values of the considered dataset and
each cell in the image shows the correlation between the
values of the dataset. It is mainly used to congregate the
data which is taken as the input and analyze how it going
to process and work, this matrix is mainly used in the
advanced analysis or as the basic fundamental structure
for the advanced analysis.</p>
      </sec>
      <sec id="sec-3-4">
        <title>4.5. Scatter and Density plot</title>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5. Proposed System</title>
      <p>The above Figure 3 The scatter and density plot
represents above gives the details of the values distributed
in the data set we have taken , and it is very useful in
comparing and plotting the scatter plot and it also shows
us how many dots or the values are being concentrated
on a single d plane or the area considered, by this we
can assess the characteristics of the data and also we can
predict the behavior up to an extent.</p>
      <sec id="sec-4-1">
        <title>The Figure 4 represents the architecture of our system</title>
        <p>and the application is been developed both in front- end
and back-end. SQLite is used for storing all the patient
and doctor information as well as reports. The Random
forest Machine Learning Algorithm not only will
predict the diseases but also its sub diseases. Map Reduce
Algorithm which increases the eficiency of the
operation and also it reduces the retrieval time of the query.
[25] The proposed system is diferent from the
ancestor’s thought of execution It uses random forest machine
learning procedure for calculating diseases and its
respective sub diseases. which in turn increase the eficiency
and performance and query response time is reduced
too. Along with that, it gives separate patterns to each
patient which gives the patient personalized experience.
In addition to that, it provides definite rations for specific
patients to pattern his/her condition. Thus, making our
system broadly open by all at moderate cost. The
prediction accuracy of the CNN-MDRP algorithm reaches 94%
compared to other prediction algorithms.</p>
        <p>All the data obtained about the patients from the
hospital management is stored in this particular module like
the MRI scans and the other types of data, mainly here we
will be using both the structured and unstructured values
of the attributes of brain anatomy, structured data refers
to the data which is obtained from the reports of the
patient and the unstructured data is something which we
get from the patient’s medical history and the informal
conversations with the doctor and etc.</p>
        <p>Training set and Testing data, for producing
sophisticated results we perform training set where initial data
help program how technologies like neural networks are</p>
        <sec id="sec-4-1-1">
          <title>6.1. Creating textual data</title>
          <p>All the data collected from the clinic or the hospital, we
use the word and insert into the first layer to remove the
false wording, the text will be represented as a vector,
we can also call it as the preprocessing. Here, each word
will be outlined as Rd dimensional variable, where dia as
50, thus, a text that incorporates n words can be depicted
as Tx = (tx1, tx2, ···, txn), Tx ∈ Rd × n.</p>
        </sec>
        <sec id="sec-4-1-2">
          <title>6.2. CNN text transforming level</title>
          <p>In each case we assess words. In some words, we incline
towards 2 words from frontal and back of every factor of
words in this issue.</p>
        </sec>
        <sec id="sec-4-1-3">
          <title>6.3. CNN text pool layer</title>
          <p>per- formed. In testing data, the obtained data will be
checking for the execution of test case and verify the
expected output in any of the software applications. DATA
Labelling preparation, in this module we will be
identifying raw data and further adding meaning and informative
to provide the context so that we can prepare the data
that could be enriched further. In this section we try to
append or enrich the obtained data with relevant context
gathered from other additional sources. Further this data
is sent to the random forest classifier for classification
of data ,Will also quantify the values of a out- come by
providing a framework, CNN MDRP, uses both the
structured and unstructured data for the prediction process,
we collect the data form the medical clinics or hospitals
and use CNN-MDRP to process the data and predict the
risk of the disease, whereas CNN-UDRP which is in the
existing system used only the structured data gives less
accuracy than the proposed one.</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>Taking the convolution layer output or the data as the</title>
        <p>grouping or gathering level information, we utilize a
6. CNN MDRP large gathering activity as possible (1-max gathering).
The motivation behind why you pick the maximum or the
Collection CNN-MDRP is performed to foresee who is greatest gathering system is the part of each expression
influenced with the illness in an efective manner by is n’t absolutely equivalent; from this we can decide the
utilizing convolutional neural network that utilizes struc- content or the matter which is useful to the system to
tured and unstructured information from the hospital. To proceed to the next step.
begin with, here we use the potential segment model to
discover lacking data from the clinical records, from the 6.4. Fully connected layer of text CNN
dataset which we consider it as repository. Besides, with
the assistance of factual information, we could manage This particular layer is related to the neural network
the fundamental degenerative illness which is available which is fully connected, we have a process of calculating
in the past. Furthermore, dealing with the structured it in this layer, we have a formula which is Hf3=Wa3
data by talking with clinic or the hospital specialists to Hf2=Bv3,where we say Hf3 is considered as the total
get the reusable information and properties. For unorga- connection level, and we can be considered as the artifact
nized area, properties can be precisely abused through and Be can be considered as the deviations part
the CNN rule. In this way, we infer that the CNN- MDRP
works the best for this specific problem. 7. Result and discussions</p>
        <p>For the assessment in the examination. To start with,
we indicate TP (the number of occurrences accurately Experimental results of this system indicate that among
anticipated as required), FP, TN and FN as evident pos- the existing algorithms CNN-mdrp gives us the best
reitive false positive (the quantity of cases erroneously sults with higher accuracy. In this system we used
CNNanticipated as required), True negative (the quantity of mdrp algorithm and made use of the structured and the
examples accurately anticipated as not needed) and False unstructured data of the patients gathered from the
clininegative (the number of examples mistakenly anticipated cal records of the hospital. No other system worked on
as not needed), separately. At that point, we can get four both structured and unstructured data. Our proposed
alestimations: gorithm gives an accuracy of 94.3%. Thus, the introduced
accuracy, precision, recall and F1-measure as follows: combination of the algorithms gave better results when
1. Accuracy = (TruePos + TrueNeg)/(TruePos + False- compared to the existing systems. So this proposed
sysPos + TrueNeg + FalseNeg) tem reduces the error rate and simultaneously increases
2. Precision = TruePos/(TruePos + FalsePos) the percentage of accuracy. And below we show the
com3. Sensitivity = TruePos/(TruePos + FalseNeg) parison and the explanation of the existing algorithms.
4. Specificity = TrueNeg/(TrueNeg + FalsePos)</p>
        <p>The process could be divided into in to 5 diferent parts:
The existing algorithms which have used previously are
linear regression, SVM, Decision tree and the proposed
one is random forest algorithm. Here what linear model
means is it works completely based on supervised
learning which is a part of machine learning. Prediction of
value is done on independent variables by using
regression models. The obtained accuracy for this algorithm is
63.3</p>
      </sec>
      <sec id="sec-4-3">
        <title>Support vector machine algorithm (SVM )also works</title>
        <p>based on supervised learning for classification of
problems. Here the optimal solution is obtained by
transforming data by using techniques like kernel trick. The
obtained accuracy for this algorithm is 88.9</p>
        <sec id="sec-4-3-1">
          <title>7.3. DECISION TREE</title>
        </sec>
      </sec>
      <sec id="sec-4-4">
        <title>The next algorithm used in the existing system is decision</title>
        <p>tree part of supervised learning used for classification of
problem. Works by following a set of if-else condition
to represent the data and categorize them. The obtained
accuracy for this algorithm is 91.3</p>
        <sec id="sec-4-4-1">
          <title>7.4. RANDOM FOREST</title>
        </sec>
      </sec>
      <sec id="sec-4-5">
        <title>The proposed system uses random forest algorithm since</title>
        <p>it consists many decision trees within them, they use
feature randomness while building each single tree to create
the uncorrelated forest containing trees so that accuracy
produced by this system automatically increases. The
obtained accuracy for this algorithm is 94.3%.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>8. Conclusion and Future work</title>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>