<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Challenges for Automation of Public Health Data Analysis Ravi Shankar</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Grenoble</institution>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Advancements in Machine Learning and Data Science are not adequately reflected in how public health data is handled today. There is a visible gap between the advances in computing and medical sciences. In this position paper, we present an example of data science applied to the automation of a repetitive process within a cervical cancer screening program. We discuss the challenges for automating public health data and share our insights to elevate artificial intelligence (AI) in public healthcare.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        More than 80% of the cervical cancer cases
and deaths in a year occur in low medium income
countries (LMICs) where prevention and cervical
screening resources are limited [1][
        <xref ref-type="bibr" rid="ref1">2</xref>
        ]. Recent
research studies have used machine learning
models to support the initial phase of screening for
detection of cancerous lesions using colposcopic
images or cervicography[
        <xref ref-type="bibr" rid="ref2">3</xref>
        ][
        <xref ref-type="bibr" rid="ref3">4</xref>
        ]. These techniques
require tech-savvy healthcare workers who are
very scarce per capita in these countries.
      </p>
      <p>We aim to build a user-friendly automation
that would allow medical experts to diagnose
cancerous tissues of the cervix in a short period of
time while reducing costs and technical
experience required. This idea will work by
combining heath and AI researchers’ expertise
and experiences.</p>
      <p>The main problem we aim to address is
diagnosing biopsied women within a cervical
cancer program. Our motivation is driven by the
importance and time consumption of pathology
process (i.e., pathologists reading histological
slides). In the pathology process, women testing
positive on screening tests are referred to
specialised examination (colposcopy) to collect
biopsy samples from the cervix and then
haematoxylin and eosin (H&amp;E) histological slides
are prepared to be reviewed by pathologists using
process (highlighted in yellow in Figure 1)
excludes these steps for automation (highlighted
in blue in Figure 1).</p>
      <p>While the system works with the current
process, the automation steps are currently done
manually and repeatedly by a group of
pathologists and statisticians. As the ML model
does not exist in the current process, the analysis
reports are produced after 2-3 stages of reviews
involving multiple meetings to concur on the
results. Including our proposed steps for
automation in the current process will lower the
burden of the experts and improve the timeframe
up to 1/20 in comparison to the current process.</p>
      <p>While our proposed project pipeline (Figure 1)
forecasts optimal benefits for cervical cancer
screening, in laying the groundwork, we were
faced with critical challenges encompassing the
realms of – technical, ethical, legal, and (most
importantly) end user facing challenges. In this
Workshop on “Healthy Interfaces (HEALTHI)
2022,” we look forward to discussing our research
on automation of public health data analysis. We
hope to share our current challenges, methods,
and future plans for AI powered healthcare.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Challenges for Automation</title>
    </sec>
    <sec id="sec-3">
      <title>Public Health Data Analysis of</title>
      <p>In this section we generalise the problems we
faced when implementing our project (Figure 1)
to discuss the challenges for automation of public
health data analysis:
1. Trained Data Entry: The first challenge
is the considerable effort needed to change the
conventional data entry practices. To
automate, it is essential to construct database
constraints, design helpful interfaces, and train
non-tech savvy workers to log: complete, error
free, and rightly formatted data.
2. Patient Privacy: Anonymising the data is
important to preserving the privacy of personal
health records of patients who sign up for the
study. If possible, it should be mindfully made
visible at the level of the interface to both the
patients and their clinicians.
3. Data Pre-processing: Data
preprocessing is the cleaning and preparation of
data for the model and analysis tasks. This is a
time-consuming underestimated challenge, if
done improperly, it potentially hinders the
performance and accuracy of the model and
delays the overall study.
4. Handling Large Datasets: As public
health studies are ongoing processes which
include participants on a rolling basis, they can
result in large datasets during overall period of
the study (which might span several years). It
is crucial to prepare for handling the data in
batches for faster training of the model.
5. Cross Validation by Experts:
Validation of results is a necessity with respect
to training ML models. In healthcare-related
data, cross validation by experts is much more
important to prevent fatal diagnosis errors and
to check for any potential biases in the model.
6. Human Control: It is important to have
adequate human control so that the confidence
of the predicted results is higher. Enabling
human control via the automation process’s
interface allows to spot any discrepancies and
malfunctioning.
7. Transparency: The interface should be
made simple and transparent for both
nonmedical and other non-tech savvy stakeholders
involved. The entire automation process
should be comprehensible to all stakeholders
involved for the project to succeed.
8. Legal Efforts and Approval: Last but
the most important challenge is to succeed in
the legal efforts and approvals required for the
automation projects. Developing proof of
concepts with publicly available datasets is
one of the ways to prepare for the challenge of
gaining legal approvals and other grants</p>
    </sec>
    <sec id="sec-4">
      <title>3. Conclusion</title>
      <p>AI powered public healthcare will foster a
health structure in the future where the AI process
drives the speed and accuracy of the diagnosis,
treatment, and recovery. People will get the right
diagnosis at the right time such that their treatment
and recovery chances improve, thus improving
chances of a good life. Furthermore, the cost
efficiency brought by AI techniques will enable
smart healthcare to be adapted to different
healthcare structures in different countries,
specifically in the low-income countries, so that
healthcare becomes accessible and affordable
there. This is a possibility only when AI
researchers combine their expertise and
experiences with health researchers. With this
position paper we aim to contribute by informing
both medical professionals and computer
scientists of the challenges for automation of
public health data analysis.</p>
    </sec>
    <sec id="sec-5">
      <title>4. References</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Bray</surname>
            <given-names>F</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jemal</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grey</surname>
            <given-names>N</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferlay</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Forman</surname>
            <given-names>D.</given-names>
          </string-name>
          <article-title>Global cancer transitions according to the Human Development Index (</article-title>
          <year>2008</year>
          -2030)
          <article-title>: a population-based study</article-title>
          .
          <source>Lancet Oncol</source>
          .
          <year>2012</year>
          ;
          <volume>13</volume>
          (
          <issue>8</issue>
          ):
          <fpage>790</fpage>
          -
          <lpage>801</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Cho</surname>
            ,
            <given-names>BJ.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Choi</surname>
            ,
            <given-names>Y.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
          </string-name>
          , MJ. et al.
          <article-title>Classification of cervical neoplasms on colposcopic photography using deep learning</article-title>
          .
          <source>Sci Rep</source>
          <volume>10</volume>
          ,
          <issue>13652</issue>
          (
          <year>2020</year>
          ). https://doi.org/10.1038/s41598-020-70490-4
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Liming</given-names>
            <surname>Hu</surname>
          </string-name>
          , David Bell,
          <string-name>
            <given-names>Sameer</given-names>
            <surname>Antani</surname>
          </string-name>
          , Zhiyun Xue, Kai Yu, Matthew P Horning, Noni Gachuhi, Benjamin Wilson, Mayoore S Jaiswal,
          <string-name>
            <given-names>Brian</given-names>
            <surname>Befano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L Rodney</given-names>
            <surname>Long</surname>
          </string-name>
          , Rolando Herrero, Mark H Einstein,
          <string-name>
            <surname>Robert D Burk</surname>
            , Maria Demarco, Julia C Gage, Ana Cecilia Rodriguez, Nicolas Wentzensen,
            <given-names>Mark</given-names>
          </string-name>
          <string-name>
            <surname>Schiffman</surname>
          </string-name>
          ,
          <article-title>An Observational Study of Deep Learning and Automated Evaluation of Cervical Images for Cancer Screening</article-title>
          ,
          <source>JNCI: Journal of the National Cancer Institute</source>
          , Volume
          <volume>111</volume>
          ,
          <string-name>
            <surname>Issue</surname>
            <given-names>9</given-names>
          </string-name>
          ,
          <string-name>
            <surname>September</surname>
            <given-names>2019</given-names>
          </string-name>
          , Pages
          <fpage>923</fpage>
          -932, https://doi.org/10.1093/jnci/djy225
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>