<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Development of Genomic Based Diagnostics in Various Application Domains</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>(Extended Abstract)</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>© Zoltan Szallasi</string-name>
          <email>zszallasi@chip.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Bio and Health Informatics, Technical University of Denmark</institution>
          ,
          <addr-line>Kemitorvet 208, 2800 Lyngby, Denmark</addr-line>
          ,
          <institution>Computational Health Informatics Program, Boston Children's Hospital, USA, Harvard Medical School</institution>
          ,
          <addr-line>Boston</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Proceedings of the XIX International Conference “Data Analytics and Management in Data Intensive Domains” (DAMDID/RCDL'2017)</institution>
          ,
          <addr-line>Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <fpage>3</fpage>
      <lpage>4</lpage>
      <abstract>
        <p>We will review the revolution brought about by low cost next generation sequencing in a wide array of diagnostic and industrial applications with a special emphasis on computational requirements and big data challenges.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>Next generation sequencing ((NGS) has
fundamentally changed modern biological research. It is,
in fact, an excellent example of how gradual
improvements on a powerful initial idea, Sanger’s
original dideoxynucleotide sequencing, can lead to such
levels of quantitative increase in data production that
fundamentally changes a given research field.</p>
      <p>Virtually any nucleic acid related research question
can be investigated in a comprehensive, high resolution
fashion free from experimental confounding factors such
as nucleic acid cross hybridization. This has produced a
deluge of data on the scale of hundreds of Terabytes even
for a single research laboratory. This review will survey
both the various application domains of next generation
sequencing and their associated computational and
analytical challenges.</p>
    </sec>
    <sec id="sec-2">
      <title>2 Biochemical considerations of next generation sequencing</title>
      <p>
        It was recognized early on that next generation
sequencing will allow querying both the genome (DNA)
and the transcriptome (RNA) on a wide range of
resolution. The exact sequence of nucleotides (e.g. single
nucleotide polymorphisms, single nucleotide variations)
and the overall architecture of the entire genome can be
determined in a single experiment (one run of whole
genome sequencing) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>A wide array of starting materials can be used for next
generation sequencing. Any form of nucleic acid (DNA
or RNA), from any sources (from inside the cell, from
cell free biological fluids or ancient fragmented DNA)
can be sequenced and quantified. Nucleic acids can be
preselected as in e.g. whole exome sequencing (exon
capture) or by other capture mechanisms such as specific
protein beacons in ChipSeq analysis. The variations are
virtually unlimited and novel approaches are constantly
being added to the toolbox of biological research. This
universality has led to an enormous variety of application
domains.
3 Application domains of next generation
sequencing</p>
      <sec id="sec-2-1">
        <title>3.1 Next generation sequencing in microbiology</title>
        <p>
          Since DNA is universal, next generation sequencing
can be used to detect and investigate any life form from
viruses to humans. This is readily exploited in the various
forms of sequencing based microbial diagnostics and it
has also led to a new research field, metagenomics [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ].
In this, the fact that often a diverse group of microbiota
live together as an “organic whole” has led to the
realization that those species do not need to be isolated
individually before sequencing, but the pooled DNA can
be sequenced together and the sequence tags can be
“sorted out” after the sequencing reaction. This
ingenious idea has led to significant advances in our
understanding of, for example, the microbial community
of our gut flora. This method also allows monitoring
sewage quality and may help to monitor and prevent
disease outbreaks.
3.2 Next generation sequencing in human diseases
        </p>
        <p>Genome wide association studies are a powerful
method to identify germline variants associated with
increased disease risk.</p>
        <p>
          Remarkably, germline DNA of the fetus can be
efficiently detected in maternal blood, leading to the
powerful tool of prenatal testing [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. By this, genomic,
chromosomal aberrations of the fetus can be detected
without virtually any risk to the mother or the fetus.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>3.3 Next generation sequencing in cancer</title>
        <p>
          Cancer is a genetic disease. Accumulating mutations
at various levels of the germline genome lead to
malignant transformation. It is therefore, obvious, that
one of the main targets of next generation sequencing is
cancer diagnostics. Both germline and somatic mutations
are readily identified by NGS. A great number of
oncogenic mutations, many of them targetable by
therapy, have thus been identified [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. NGS data also
allow us to reconstruct the evolutionary history of cancer,
an issue of potentially great significance [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ].
        </p>
        <p>Recently, sequence analysis of liquid biopsies,
essentially cell free DNA obtained from various bodily
fluids, emerged as a minimally invasive tool to obtain
vital information about the presence of cancer in a patient
individual before or during therapy.
4 Computational and analytical challenges
of next generation sequencing data</p>
        <p>
          The uniform nature of the biochemical reactions and
ingenuity of technical development has led to an
unexpected situation. The price drop/throughput increase
of next generation sequencing has significantly outpaced
Moore’s law over the past decade. Therefore, next
generation sequencing has become more of a
computational rather than a biochemical problem. The
speed of data accumulation is so fast that data storage and
data analysis is becoming a more and more challenging
problem in modern sequencing based projects. Whole
genome sequencing on a single cancer sample can easily
take up hundreds of GBs of data storage. Therefore, a
single study, such as the one analyzing the whole genome
of 560 breast cancer cases [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], can easily produce data on
at the level of hundreds of Terabytes. Such amount of
data cannot possibly be downloaded for reanalysis in an
efficient manner, therefore alternative solutions, such as
cloud based computing had to be found.
        </p>
        <p>Management of vast amounts of data is only one,
mainly technical aspect of the challenges at hand.</p>
        <p>While next generation sequencing based genomics
easily qualifies as one of the main areas of big data
science, in many aspects it is also markedly different
from those. While in, e.g., financial data the individual
variables are connected by poorly understood causative
factors and in physics the entire data space is regulated
by well defined, homogenous laws of physics, in
biology, genomics the situation lies somewhere in
between. Variables, such as genes, proteins etc. are
connected by the principles of physical chemistry, but the
actual parameters of those significantly vary across the
various pairs of biological entities. This fact places the
analysis of biological systems in the realm of robust,
complex systems for which the analytical principles are
poorly understood. Therefore, in order to effectively
analyze the massive amounts of genomic information
one needs to “front-load” the computational analysis
with as much biological knowledge as possible.
We will present several strategies along those lines. In
particular, we will discuss how genomics, next
generation sequencing based whole genome analysis
helps us to understand DNA repair pathway aberrations,
and their diagnostic and therapeutic implications in
cancer. We will also discuss how genomics is exploited
to understand the main principles of therapeutic immune
responses against cancer and how genomics, machine
learning and high throughput screening are combined in
an interdisciplinary environment to design effective
vaccines against cancer.
5 The industrial impact of next generation
sequencing</p>
        <p>In order to satisfy the need for NGS based diagnostics
a whole industry has developed during the past decade.
Conferences such as the 2017 Next Generation Dx
Summit, (http://www.nextgenerationdx.com) provide an
excellent overview of the major trends and players.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Mardis</surname>
            <given-names>ER</given-names>
          </string-name>
          (
          <year>2013</year>
          )
          <article-title>Next-generation sequencing platforms</article-title>
          .
          <source>Annu Rev Anal Chem</source>
          (Palo Alto Calif)
          <volume>6</volume>
          :
          <fpage>287</fpage>
          -
          <lpage>303</lpage>
          . doi:
          <volume>10</volume>
          .1146/annurev-anchem-
          <volume>062012</volume>
          - 092628
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Wooley</surname>
            <given-names>JC</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Godzik</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Friedberg</surname>
            <given-names>I</given-names>
          </string-name>
          (
          <year>2010</year>
          )
          <article-title>A primer on metagenomics</article-title>
          .
          <source>PLoS Comput Biol</source>
          <volume>6</volume>
          :
          <fpage>e1000667</fpage>
          . doi:
          <volume>10</volume>
          .1371/journal.pcbi.1000667
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Hui</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bianchi</surname>
            <given-names>DW</given-names>
          </string-name>
          (
          <year>2017</year>
          )
          <article-title>Noninvasive Prenatal DNA Testing: The Vanguard of Genomic Medicine</article-title>
          .
          <source>Annu Rev Med</source>
          <volume>68</volume>
          :
          <fpage>459</fpage>
          -
          <lpage>472</lpage>
          . doi:
          <volume>10</volume>
          .1146/annurev-med-072115-033220
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Lawrence</surname>
            <given-names>MS</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stojanov</surname>
            <given-names>P</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mermel</surname>
            <given-names>CH</given-names>
          </string-name>
          , et al (
          <year>2014</year>
          )
          <article-title>Discovery and saturation analysis of cancer genes across 21 tumour types</article-title>
          .
          <source>Nature</source>
          <volume>505</volume>
          :
          <fpage>495</fpage>
          -
          <lpage>501</lpage>
          . doi:
          <volume>10</volume>
          .1038/nature12912
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Jamal-Hanjani</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wilson</surname>
            <given-names>GA</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McGranahan</surname>
            <given-names>N</given-names>
          </string-name>
          , et al (
          <year>2017</year>
          <article-title>) Tracking the Evolution of Non-SmallCell Lung Cancer</article-title>
          .
          <source>N Engl J Med</source>
          <volume>376</volume>
          :
          <fpage>2109</fpage>
          -
          <lpage>2121</lpage>
          . doi:
          <volume>10</volume>
          .1056/NEJMoa1616288
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Nik-Zainal</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Davies</surname>
            <given-names>H</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Staaf</surname>
            <given-names>J</given-names>
          </string-name>
          , et al (
          <year>2016</year>
          )
          <article-title>Landscape of somatic mutations in 560 breast cancer whole-genome sequences</article-title>
          .
          <source>Nature</source>
          <volume>534</volume>
          :
          <fpage>47</fpage>
          -
          <lpage>54</lpage>
          . doi:
          <volume>10</volume>
          .1038/nature17676
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>