<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>N. Gandhi);
shakti.mishra@sot.pdpu.ac.in (S. Mishra)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Applications of Reinforcement learning for Medical Decision Making</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Neel Gandhi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shakti Mishra</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>School of Technology, Pandit Deendayal Petroleum University Gandhinagar</institution>
          ,
          <addr-line>Gujarat 382007</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0003</lpage>
      <abstract>
        <p>Reinforcement Learning(RL) is used for decision-making by interacting with uncertain/complex environments with the aim of maximizing long-term reward following a certain policy along with evaluative feedback for improvement. RL is advantageous in medical decision making compared to other forms of learning as it focuses on long-term rewards, it is also able to handle long and complex sequential decision-making tasks with sampled, delayed, and exhaustive feedback. It has emerged as a suitable method for developing satisfactory solutions in the healthcare domain. Improvement in the healthcare system can be achieved by integrating traditional health care practices with RL methods by considering health status of a patient. In this paper, we have discussed various applications of RL that would be helpful in providing efective decisions for improving patient health treatment, prognosis, diagnosis, and condition . RL could be efective in the area of healthcare right from medical diagnosis to handling various critical decision-making tasks. The paper provides a broad view of the various applications of RL in the sector of healthcare. The paper illustrates various RL applications that would be efective in improving the existing healthcare sector at same time being eficient in handling complex medical decision-making tasks.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Reinforcement Learning</kwd>
        <kwd>Medical Decision Making</kwd>
        <kwd>Healthcare</kwd>
        <kwd>Medical Diagnosis</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>tially, RL has been used for treatment of
patients in a closed-loop manner having varied
In recent years, Reinforcement Learning has advantages compared to supervised learning.
emerged as one of crucial area in field of arti- Traditionally, supervised learning algorithms
ifcial intelligence impacting the field of health work on labeled data whereas RL has unique
care including diagnosis, prognosis, and other feature of finding the pattern in given
probmedical treatments. Reinforcement learning
methods have been very useful for a long time
lem statement and bound to learn from its
experience. Also, Evolution in RL from the
in sequential decision-making tasks in robotics, past to present has made it capable of
hangaming, and simulation like healthcare that
dling various issues like exploration and
exare able to solve long and complicated decision- ploitation, credit assignment, and at the same
making tasks with the use of policies, aiming
at maximizing reward as their final goal.
Initime maximizing the reward using the
optimal policy for a specific medical decision-making
task. RL has gained popularity among
practitioners dealing with dynamic treatment regimes,
medical diagnosis, and other decision-making
tasks. Reinforcement learning has been
applied for simulations in healthcare domains
like drug dosage, examination time,
assessment of patient’s health status among others.</p>
      <p>Application of RL in medical decision
making has to deal with patient health concern
issues due to the risks involved in the
medical treatment for a particular decision made
by RL method. However, it is imperative to
choose an optimal method for treatment of
specified medical disease or medical
condition.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Applications of Reinforcement Learning in medical decision making</title>
      <p>
        Reinforcement learning in healthcare follows
certain steps like agent (medical
device/computer/equipment/system) that takes
a particular action in medical environment
using defined policy to get a specific reward
and then uses evaluative feedback to improve
its performance[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].RL provides various
methods to solve sequential decision-making
problems with the goal of maximizing reward by
interaction with environment using trial and
error method. Also, exploring and exploiting
the environment for taking decisions by
evaluative feedback from environment and
learning efective strategies during the process.
      </p>
      <p>Reinforcement Learning has emerged as a
prominent solution in decision-making tasks in
medical sector and has application right from
Dynamic Treatment Regimes,medical diagnosis
to various other complicated and cognitive
decision-making tasks.</p>
      <p>
        1. Dynamic Treatment Regimes(DTR)
-DTR designed for sequential
decisionmaking problems by using
reinforcement learning methods for developing
policy with respect to automation for
process of developing treatment regimes
for patients by consideration of long term
health benefits[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
a) Chronic Diseases-Chronic diseases
persist for a long period of time.Hence,
practitioners follow chronic care
model(CCM) sequence of medical
interventions to access patient health
status[3].RL would be helpful to
practitioners for continuous
decisionmaking by helping in treatment
of chronic diseases including
anemia, cancer, diabetes, human
immunodeficiency viruses(HIV) ,
mental illnesses among many other
longlasting diseases.
      </p>
      <p>i. Cancer- Q-learning with
support vector regression and
extremely randomized trees is
used for treatment of cancers[4]
like cell cancer,
chemotherapy efect, and other cancer
conditions.
ii. Diabetes-Proper sequential dosage
of insulin in cases of Type 2
diabetes[5] with specified time
and amount by application of
reinforcement learning
methods for getting long-term health
benefit.
iii. Anemia-Lack of RBC that can
be controlled using by RL method[6]by
applying control input as the
amount of endogenous
erythropoietin and target under
control is hemoglobin level that
also has an impact on iron
storage in the patient’s body with
the state component of hemoglobin
and avoids any damage to
patient body’s by administering
erythropoiesis-stimulating agent.
iv. HIV-HIV/human
immunodeficiency virus[7] are treated
with the combination of
antiHIV drugs that are referred
to as highly active
antiretroviral therapy(HAART) requires
long term treatment using
decisionmaking approach could be
effectively dealt by using
reinforcement learning algorithm
like Batch RL.
v. Mental illness- Mental
illness usually persists for a long
period of time requiring
significant adaptations/changes
in terms of dosage as well as
treatment type involving very
complex decision-making
process. Thus, it can be handled
using our RL approaches to
solve the problems of Depression[8],
Schizophrenia[9] among many
other mental issues.
b) Intensive/Critical Care-RL method
would prove to be helpful in cases
of critical care treatment like
mechanical ventilation[10] as well as
treatment of diseases like sepsis[11]
and other critical care treatment
i. Sepsis - Using model-based
reinforcement learning techniques[11]
with improvised policies has
led to better treatment for the
condition of sepsis in patients.
ii. Anesthesia - Anesthesia is
the process of using specific
drugs to reduce the efect of
sensation in body with the use
of RL-based control methods
[12] like temporal diference
to detect distribution of drug
in patient’s body.
iii. Others Critical Situation</p>
      <p>As RL method are used for
handling decision making
system in uncertain environment
,it would be efective in
dealing with critical situation such
as Mechanical Ventilation[10],
3. Other Medical Decision-making Tasks
for healthcare systems
a) Resource scheduling and task
allocation - Resource allocation
problem in RL are usually
modeled using Markov Decision
Process with reinforcement learning
using appropriate policies to
provide better service to the patient[17].
b) Optimal Process Control -
Healthcare tasks like simulation of
surgical operation, adaptive control
for medical video streaming, and
functional electric simulations
policies control are used with RL
methods like Q-learning, IRL, DRL among</p>
      <p>Heparin Dosing[13] among other
critical care conditions
2. Medical Diagnosis- Medical Diagnosis[14]
is helpful in decision-making using RL
with medical condition data in form of
image and text data.</p>
      <p>a) Computer Vision</p>
      <p>Medical Image- Medical Image
data obtained from various
computer vision techniques are used
for feature extraction, image
segmentation, localization, tracing, and
object detection along with RL algorithm[15].
b) Natural Language Proccesing</p>
      <p>Clinical text data- Clinical text
data has also been used for
treatment of patients using RL method[14]
that are able to diagnose inferences
in RL methods like DQ method.
c) Human-Computer Interface</p>
      <p>Dialogue Systems, Chat-bots,
and Advanced Interfaces-
Multiagent systems were found
efective in monitoring clinical data
using RL method for developing user
interface that is able to adapt
itself for specific user[16].
2006.09.042. [18] J. Shin, T. A. Badgwell, K.-H. Liu, J. H.
[10] N. Prasad, L.-F. Cheng, C. Chivers, Lee, Reinforcement learning–overview
M. Draugelis, B. E. Engelhardt, A rein- of recent progress and implications for
forcement learning approach to wean- process control, Computers &amp; Chemical
ing of mechanical ventilation in in- Engineering 127 (2019) 282–294.
tensive care units, arXiv preprint [19] M. Popova, O. Isayev, A. Tropsha, Deep
arXiv:1704.06300 (2017). reinforcement learning for de novo
[11] A. Raghu, M. Komorowski, S. Singh, drug design, Science advances 4 (2018)
Model-Based Reinforcement Learn- eaap7885.
ing for Sepsis Treatment (2018). [20] J. Mulani, S. Heda, K. Tumdi, J.
PaURL: http://arxiv.org/abs/1811.09602. tel, H. Chhinkaniwala, J. Patel, Deep
arXiv:1811.09602. reinforcement learning based
person[12] B. L. Moore, L. D. Pyeatt, V. Kulkarni, alized health recommendations, in:
P. Panousis, K. Padrez, A. G. Doufas, Deep Learning Techniques for
BiomedReinforcement learning for closed-loop ical and Health Informatics, Springer,
propofol anesthesia: A study in human 2020, pp. 231–255.
volunteers, Journal of Machine
Learning Research 15 (2014) 655–696.
[13] S. Nemati, M. M. Ghassemi, G. D.
Clifford, Optimal medication dosing from
suboptimal clinical examples: A deep
reinforcement learning approach, in:
2016 38th Annual International
Conference of the IEEE Engineering in
Medicine and Biology Society (EMBC),</p>
      <p>IEEE, 2016, pp. 2978–2981.
[14] Y. Ling, S. A. Hasan, V. Datla, A. Qadir,</p>
      <p>K. Lee, J. Liu, O. Farri,
Diagnostic inferencing via improving clinical
concept extraction with deep
reinforcement learning: A preliminary study, in:
Machine Learning for Healthcare
Conference, 2017, pp. 271–285.
[15] F. Sahba, H. R. Tizhoosh, M. M. a.</p>
      <p>Salama, for Medical Image
Segmentation (2006) 1238–1244.
[16] E. M. Shakshuki, M. Reid, T. R. Sheltami,</p>
      <p>An adaptive user interface in
healthcare, Procedia Computer Science 56
(2015) 49–58.
[17] Z. Huang, W. M. van der Aalst, X. Lu,</p>
      <p>H. Duan, Reinforcement learning based
resource allocation in business process
management, Data &amp; Knowledge
Engineering 70 (2011) 127–145.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Deep reinforcement learning: An overview</article-title>
          ,
          <source>arXiv preprint arXiv:1701.07274</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , E. Bareinboim,
          <article-title>Designing optimal dynamic treatment regimes:</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>