<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Precise Nanopore Signal Modeling Improves Unsupervised Single-Molecule Methylation Detection</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Vladimír Boža</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Eduard Batmendijn</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Peter Perešíni</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Viktória Hodorová</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hana Lichancová</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rastislav Rabatin</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Broňa Brejová</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jozef Nosek</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tomáš Vinař</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Faculty of Mathematics</institution>
          ,
          <addr-line>Physics and Informatics</addr-line>
          ,
          <institution>Comenius University in Bratislava</institution>
          ,
          <addr-line>Mlynská dolina, 842 48 Bratislava</addr-line>
          ,
          <country country="SK">Slovakia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Faculty of Natural Sciences, Comenius University in Bratislava</institution>
          ,
          <addr-line>Ilkovičova 6, 841 05 Bratislava</addr-line>
          ,
          <country country="SK">Slovakia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Base calling in nanopore sequencing is a dificult and computationally intensive problem, typically resulting in high error rates. In many applications of nanopore sequencing, direct analysis of raw signal is a viable alternative. Dynamic time warping (DTW) is an important building block for raw signal analysis. In this paper, we propose several improvements to DTW class of algorithms to better account for specifics of nanopore signal modeling. We have implemented these improvements in a new signal-to-reference alignment tool Nadavca. We demonstrate that Nadavca alignments improve unsupervised methylation detection over Tombo. We also demonstrate that by providing additional information about the discriminative power of positions in the signal, an otherwise unsupervised method can approach the accuracy of supervised models. Availability and implementation: Nadavca is available under MIT license at https://github.com/fmfi-compbio/nadavca. Nanopore sequencing data sets are available from ENA bioproject PRJEB64246. Jaminaea angkorensis reference genome assembly is available from Zenodo https://doi.org/10.5281/zenodo.8145315.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Background</title>
      <sec id="sec-1-1">
        <title>DNA sequencing devices developed by Oxford Nanopore</title>
        <p>Technologies (ONT), including the pocket-sized MinION,
measure changes in electrical current as a DNA strand
passes through a nanopore. The result is a sequence of
signal measurements (also called a squiggle). In this paper,
we provide an improved algorithm to align parts of the
squiggle to a known reference DNA sequence and show
that improved alignment accuracy benefits downstream
analysis in de novo methylation identification.</p>
        <p>
          Typically, the first step in the analysis of nanopore
sequencing data is base calling, which translates a squiggle
into a DNA sequence. The most successful base callers
are based on complex machine learning models, such as
recurrent neural networks [
          <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
          ]. A typical median read
accuracy achieved by a base caller for R9.4.1 nanopores
employed in this study is approx. 97% [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], but when base
called reads are piled up at a deep coverage, the consensus
accuracy can be as high as 99.94% [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. (ONT lists 98.3%
accuracy for R9.4.1 and 99% for newer R10.4.1 nanopores;
note that these numbers are based on accuracy in the
mode of the read accuracy distribution and are generally
higher than median accuracy.)
        </p>
        <p>An alternative class of approaches avoids base
calling and works directly with the raw signal sequence.
The main motivation is to use the additional information
present in the raw signal, which is potentially lost in base
calling. Applications include mainly better modification
calling and methylation detection. Base calling is also
computationally intensive, and thus in some situations,
working with the raw signal can be faster.</p>
        <sec id="sec-1-1-1">
          <title>Basic properties of nanopore signal. The measured</title>
          <p>electrical current depends mostly on the context of 
consecutive nucleotides passing through the nanopore.</p>
          <p>A signal model predicts for each context of size  the
expected signal level. For the case of nanopores in R9.4.1
lfow-cells, which were used in this paper, a typical model
uses the context of size  = 6. The DNA moves through
the pore at the speed of roughly 450 bases per second
and the signal is read at a frequency of 4000 samples per
second. This means that on average we have approx. 9
samples per one DNA context, but the actual number of
samples per context can vary significantly.</p>
          <p>
            Signal-to-reference alignment and dynamic time
warping. One of the basic building blocks of nanopore
signal analysis is the alignment of squiggles to a known
reference sequence. The squiggle can be represented as
a sequence of numbers 1 · · · . For each nucleotide 
of the reference sequence 12 · · · , we need to assign
the corresponding interval  +1 · · ·  of the
squiggle which we call an event. Such squiggle-to-reference
alignment facilitates further analysis, allowing to find
significant diferences between the squiggle and the
reference signal computed based on the signal model applied
to the reference sequence. Identified diferences
typically correspond to sequence variants or modifications
[
            <xref ref-type="bibr" rid="ref5">5</xref>
            ]. Alignments also enable pooling together information
from multiple reads aligned to a single reference location
© 2023 Copyright for this paper by its authors. Use permitted under Creative Commons License and to visualize squiggles in the genomic context [
            <xref ref-type="bibr" rid="ref6">6</xref>
            ].
CPWErooUrckReshdoinpgs ISNc1e6u1r-3w-0s.o7r3g ACttEribUutRion W4.0oInrtekrnsahtioonpal (PCCroBYce4.0e).dings (CEUR-WS.org) The squiggle alignment is typically performed by a
class of algorithms called dynamic time warping (DTW), Supervised and unsupervised modification
detecoriginating in the field of speech processing [
            <xref ref-type="bibr" rid="ref7">7</xref>
            ]. The tion. It is known that modified nucleotides afect the
alignment of a signal level  to a position  in the ref- signal levels when passing through nanopore [
            <xref ref-type="bibr" rid="ref10">10</xref>
            ]. For
erence  is evaluated by a scoring function (, , ), known modification patterns, such as CpG methylation,
which typically takes into account the context around it is possible to prepare artificially methylated sequences
position ; for example, the scoring function (, , ) = and either estimate a signal model for contexts containing
( − expectedsignal(− (− 1)/2 · · · +(− 1)/2))2 uses a methylated nucleotides [
            <xref ref-type="bibr" rid="ref11 ref12 ref14">11, 12, 13</xref>
            ] or to train a machine
signal model based on the context of  nucleotides (as- learning algorithm, such as a neural network, to
recsuming  is odd). Using a scoring function of this type, ognize signals from sequences that contain methylated
we are looking for the alignment minimizing the score nucleotides [
            <xref ref-type="bibr" rid="ref15 ref16 ref17">14, 15, 16</xref>
            ]. We call this approach
supera∑︀dy=n1a∑m︀ic=prog(ram, m, in)g, walhgiocrhitchamn sbiemciolamr ptoutseedquuesnincge lvaibseedlemd ettrhayinlaintigondadteateacrteiorne,qsuinirceedstuobsbtuainldtiasluacmhomuondteslos.f
alignment. In contrast, unsupervised modification detection can be
          </p>
          <p>
            In our experience, the DTW-like algorithms applied in used when labeled training data is not available, or when
the context of MinION squiggle alignment show a variety we are targeting previously uncharacterized DNA
modiof artefacts. The scoring function allows squiggle to be ifcations. One option is to compare signal values in reads
compressed or stretched arbitrarily in time, without any in the studied sample with signals in a control sample
penalty. If several consecutive sequence contexts along devoid of modifications and look for statistically
signifthe reference genome exhibit similar expected signal val- icant diferences [
            <xref ref-type="bibr" rid="ref18 ref6">6, 17</xref>
            ]. However, it is also possible to
ues in the signal model (which is not uncommon), there ifnd putative modifications simply by comparing the
acis a large uncertainty in how the corresponding observed tual signal with the expected signal predicted by a signal
signal should be split between the events. Depending model trained on unmodified sequences (this approach is
on the exact DTW implementation, this may lead to arti- implemented for example in Tombo [
            <xref ref-type="bibr" rid="ref19">18</xref>
            ]). The strength
ifcially short or artificially long events. Small amounts of this method is that it requires only unmodified training
of noise in the signal can easily lead to events split in- data to obtain the signal model and then can be applied
correctly between multiple reference contexts. In fact, directly to biological samples from various organisms
the real squiggles often contain transitional readouts be- without the need for modified training data or matched
tween signal levels corresponding to consecutive signal unmodified dataset for each analysis. The crucial step in
contexts, sometimes even alternating between two con- this approach is accurate alignment of signal to the
refsecutive contexts for some time. If the squiggle does not erence to allow meaningful comparison with the model
correspond exactly to the reference sequence (e.g. in case signal. Using Nadavca signal-to-reference alignments,
of single nucleotide variants or base modifications), DTW we implemented a new tool nanometh for unsupervised
tends to “hide” the diferences between the reference and modification detection. We demonstrate that Nadavca
the squiggle within the neighboring events. This subse- alignments increase the accuracy of unsupervised
modiquently afects the ability to detect such diferences. ifcation detection, and we also evaluate contribution of
          </p>
          <p>
            In this paper, we propose several improvements to the individual features implemented in Nadavca.
DTW framework to overcome these obstacles. Instead
of simple maximization of the scoring function, we pro- 2. Methods
pose to employ alternative methods similar to posterior
decoding in probabilistic models [
            <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
            ], to account for 2.1. Realignment of Signal to the
uncertainty in the event boundaries. We enhance the
signal model underlying the scoring function to take into Reference with Posterior Decoding
account a wider context of the sequence and also specif- The core part of Nadavca aligns a portion of nanopore
ically model transitions between the contexts. We also signal (values 1, · · · , ) to the corresponding part of
employ a two-iteration approach to signal normalization the reference genome (bases 1, · · · , ). The objective
to avoid artefacts resulting from incorrect shifting and is to improve the accuracy of an approximate alignment,
scaling of the squiggle. We have implemented these im- resulting from aligning base-called reads to the reference.
provements in an alignment tool called Nadavca, and we The output alignment assigns an interval of the signal
have evaluated their efect on the accuracy of detection sequence  , · · · ,  (also called event) to each base
of base modifications in nanopore reads. To make the  of the reference. For the basic version of the
algorunning time practical, we use methods similar to banded rithm, we require that the intervals for adjacent bases
and anchored sequence alignment. are adjacent in the signal sequence (i.e. +1 =  + 1).
          </p>
          <p>The standard dynamic time warping (DTW) algorithm
is based on dynamic programming. We define
subproblem , , which represents the score of the best
alignment of the squiggle segment 1 · · ·  to the reference
segment 1 · · ·  , assuming that the signal  is aligned
to the reference base  . , can be computed
gradually using a simple recurrence , = (, , ) +
min(− 1,′ , − 1,− 1), where (, , ) is the score
of aligning signal value  to the reference sequence 
at position . After finishing the computation, the fi- Figure 1: Extended signal model
nal alignment (intervals [, ]) can be easily computed
using standard methods for tracing back the dynamic
programming solution. The alignment can thus be visualized supplied by Oxford Nanopore, which predicts the
exas a path in this dynamic programming matrix. pected signal based on the sequence context of length</p>
          <p>
            Our alignment algorithm is inspired by the posterior 6. Instead, we use a (restricted) 10-mer model which we
decoding approach to sequence alignment [
            <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
            ]. While compute in the following way.
the standard DTW seeks the alignment with the highest First, we align the training dataset (described below) to
overall score, our goal is to consider also a contribution the reference genome using Nadavca with the default
6of sub-optimal alignments. Many of these alignments mer model. Computing the mean signal for each 10-mer
can have scores very close to the optimum, representing would lead to a huge model, requiring a gigantic amount
uncertainty in the true alignment. Posterior decoding of data in order to cover each 10-mer with enough
samalgorithms consider this uncertainty at each position of ples. Therefore, we restrict our 10-mer model by
expressthe alignment. ing it as the sum of three separate sub-models: a base
6
          </p>
          <p>To apply this principle to DTW, we first modify the mer model  for the central part of the context and two
scoring scheme and instead of the distance score (, , ), 4-mer models 1 and 2 that fine-tune the result based
we use probabilistic score (, , ), which gives a prob- on more distant bases. For a context  = 12 . . . 10, the
ability of signal value  being produced from the ref- expected level  () of the signal is computed as follows:
erence  at position . Details of this probabilistic scor-  () = 1(1 . . . 4) +  (3 . . . 8) + 2(7 . . . 10)
ing scheme are described below. Now the score of the (see Figure 1 for illustration). Models 1, 2, and 
alignment becomes the product of probabilities along the are estimated simultaneously for all 4-mers and 6-mers
realignment path ∏︀=1 ∏︀=  (, , ). Higher scores spectively, by fitting the observed signal levels using the
indicate higher alignment quality. least-squares method. The probabilistic score (, , )</p>
          <p>
            The key concept in posterior decoding is the posterior is now computed based on  ∼  ( (− 4 . . . +5),  )
probability , which for each  and  sums the scores of for constant  = 0.3.
all alignments that align signal value  to the reference
base  . All values of , can be computed by Forward- 2.3. Transitional States
Backward algorithm similarly as in the context of pair
hidden Markov models [8, section 4.4]. Finally, we com- The simplest models [
            <xref ref-type="bibr" rid="ref20">19</xref>
            ] of the nanopore signal assume
pute the resulting alignment as a path through the matrix that the signal can be partitioned into events, each event
 maximizing the product of individual posterior scores corresponding to one position in the reference and thus
, (as in [
            <xref ref-type="bibr" rid="ref9">9</xref>
            ]), using a dynamic programming recurrence having a roughly uniform signal level influenced by the
similar to the standard DTW. local sequence context. However, transitions between
          </p>
          <p>
            In the real data applications, we impose additional re- events are not instantaneous, and as a result, the
meastrictions. The minimum length of each event  −  + 1 sured signal may include intermediate values between
is constrained by the parameter min_event_length. The the expected levels typical for two adjacent events or it
very start and the very end of the signal may remain un- may even jump back and forth between these levels (see
aligned. The length of unaligned portions is at most Figure 2).
2bandwidth + 1, where bandwidth is a parameter de- In Nadavca, we handle these phenomena by inserting a
scribed below. Corresponding modifications of the al- special "transitional state” between each pair of adjacent
gorithms are straightforward. reference bases  and +1. The primary score for
aligning signal  to this transitional state is computed from a
2.2. Enhanced Signal Model special distribution as opposed to the -mer distribution
used normally. While this distribution can depend on
Accurate models of signal levels expected in diferent the sequence context, for simplicity we currently use a
sequence contexts are important for precise sequence constant score 0.01. Thus, any signal aligned to these
alignment. Tombo [
            <xref ref-type="bibr" rid="ref19">18</xref>
            ] by default uses the 6-mer model positions is penalized, regardless of the signal level and
the sequence context. This simple heuristic works well
in practice.
          </p>
          <p>The posterior decoding algorithm aligns transitional
states to intervals of signal values, but these intervals are
not constrained to have a minimum required length; they
can even be empty. The intervals aligned to transitional
states are omitted from Nadavca’s output, and thus the
output events corresponding to the reference bases may
have gaps between them.
2.4. Fast Implementation With Banded</p>
          <p>Alignment</p>
          <p>In the next step, we consider all matched bases from
the BWA-MEM alignment. For each base called base 
matched to a reference base , we locate the interval
... in the signal that corresponds to  according to
the Events table provided by Guppy. Interval ... will
be anchored to the reference base  and we call the set
of all anchors the guide alignment.</p>
          <p>We use the guide alignment to determine the portion of
the signal that needs to be aligned to the reference. It will
extend from the start of the interval in the first anchor
to the the end of the interval in tha last anchor, extended
further on both ends by the bandwidth parameter.</p>
          <p>The guide alignment will also define an area of
interest for dynamic programming. Recall that our posterior
alignment uses three passes of dynamic programming
(forward, backward, and posterior), and all three of them
will consider only such alignment paths that are
contained within the area of interest. All values outside of
the area of interest are considered as zeroes. For each
reference base , the area of interest contains the interval
of signal positions between the start of the closest anchor
at or before base  and the end of the closest anchor at
or after base ; this interval is extended by bandwidth
positions on both ends (Figure 4).
2.5. Signal Renormalization</p>
        </sec>
      </sec>
      <sec id="sec-1-2">
        <title>To avoid slow computation of the full dynamic program</title>
        <p>ming matrix, we use the banded alignment heuristic Raw nanopore signal needs to be normalized by applying
around an approximate guide alignment obtained by read-specific shift and scale parameters (denoted  and
aligning the base called sequence to the reference.  respectively). Given a single raw signal level raw and</p>
        <p>
          In particular, we first base call each read, and align the parameters  and  , the normalized signal level is
the base called sequence to the target genomic sequence (raw −  )/ . This normalization can influence
subseusing BWA-MEM [
          <xref ref-type="bibr" rid="ref21">20</xref>
          ]. The BWA-MEM alignment will quent alignment [
          <xref ref-type="bibr" rid="ref22">21</xref>
          ] and other processing. In Nadavca,
determine the window of the target genome to which we initially normalize the signal by using the median
we need to align the signal. This window is padded with signal value in a read as  and the median of |raw −  |
several additional bases from base call to introduce con- in a read as  . This median-based normalization is,
howtext for the signal model; please note that these bases ever, sensitive to various nanopore data artefacts. For
may be diferent from the target due to the adapters or this reason, we improve the normalization by an iterative
barcodes present at the beginning or at the end of the procedure [
          <xref ref-type="bibr" rid="ref19 ref22 ref23">22, 21, 18</xref>
          ].
read (see Figure 3). The resulting sequence will serve as In particular, after the initial normalization, the
rethe reference for the signal alignment. alignment algorithms assigns signal values to particular
reference locations. For each such aligned location, our mM EDTA pH 8.0, resuspended in 100 mM EDTA pH 8.0,
enhanced signal model predicts a particular expected 2 % (v/v) 2-mercaptoethanol and incubated for 15 min at
signal level. In the next iteration, we update normaliza- room temperature. Cells were then pelleted, resuspended
tion parameters  and  to minimize the sum of squared in 0.15 M NaCl, 100 mM EDTA pH 8.0 and disrupted by
errors between the expected and observed signal val- vortexing with glass beads (0.4-0.5 mm, Sigma) three
ues in the read (omitting signal values aligned to tran- times for 30 s, followed by addition of sodium dodecyl
sitional states). Subsequently, we repeat the process by sulfate (SDS) to 0.1 % (w/v). Nucleic acids were obtained
re-computing the alignment and recomputing normaliza- by series of phenol and phenol : chloroform :
isoamylaltion again. We observed, that two iterations are suficient cohol (25 : 24 : 1) extractions, precipitated with an equal
to achieve a good accuracy. volume of 96 % (v/v) ethanol, washed with 70 % (v/v)
ethanol, and air-dried. The precipitate was dissolved in
2.6. Unsupervised Detection of DNA TE bufer (10 mM Tris-HCl, 1 mM EDTA, pH 8.0) and
RNA was removed by RNase A (150 µg/mL digestion) for
Modification Through Anomalies
30 min at 37°C. DNA was extracted by phenol :
chloroDNA modifications typically cause small shifts in sig- form : isoamylalcohol (25 : 24 : 1), precipitated with 0.1
nal values. Given an accurate signal-to-reference align- M NaCl and two volumes of 96 % (v/v) ethanol, washed
ment, this property can be exploited to detect DNA mod- with 70 % (v/v) ethanol, air-dried, dissolved in TE bufer
ifications by comparison of expected signal values pre- and further purified using a Genomic-tip 100/G (Qiagen).
dicted from a signal model to the observed signal values HMW DNA of the yeast Magnusiomyces capitatus NRRL
in individual events produced by the alignment. Here, Y-17686 (CBS 197.35; [
          <xref ref-type="bibr" rid="ref26">25</xref>
          ]) was prepared essentially as
we mostly follow the approach used by Tombo de novo described above, except that prior the phenol extractions,
modification detection module [
          <xref ref-type="bibr" rid="ref19">18</xref>
          ]. However, we sub- yeast cells were resuspended in 1 M Sorbitol, 10 mM
stitute our improved signal-to-reference alignment and EDTA pH 8.0 and converted to spheroplasts using
zyenhanced signal model to demonstrate their usefulness. molyase 20T (0.125 mg/mL; Seikagaku) treatment for
        </p>
        <p>In particular, let [, ] be an event aligned to the refer- 60-90 min at 37°C. Spheroplasts were then lyzed in 0.15
ence context . Let  is the event mean, i.e. the mean of M NaCl, 100 mM EDTA pH 8.0, 0.1 % (w/v) SDS.
signal values  . . . . As an indicator of DNA modifica- Nanopore sequencing was carried out in a MinION
tion, we compute the p-value that value  comes from a Mk1B device using a FLO-MIN106 (R9.4.1, revD) flow
distribution  ( (),  ) defined by the enhanced signal cell (Oxford Nanopore Technologies). The sequencing
model. libraries were constructed essentially as described in</p>
        <p>
          This simple method, however, does not lead immedi- the manufacturer’s instructions. Datasets of control
unately to accurate modification detection. On one hand, modified DNA and DNA methylated in vitro were
prethe shifts in signal values are small and thus individual pared using the PCR barcoding kit (SQK-PBK004,
Oxp-values assigned to events are rarely significant. On the ford Nanopore Technologies). Briefly, HMW DNA was
other hand, a typical modification afects several events sheared in a g-Tube (Covaris) at 2700 ×g or 4300 ×g in
near the modified base. Following the method outlined a MiniSpin Plus centrifuge (Eppendorf). Fragmented
in Tombo, we combine p-values from three adjacent po- DNA (∼ 200 ng) was treated using a NEBNext Ultra II
sitions using the Fisher’s combined probability test [
          <xref ref-type="bibr" rid="ref24">23</xref>
          ] End repair / dA-tailing module (New England Biolabs)
to form a more reliable indicator. and ligated to the Barcode adapters (BCA) from the
se
        </p>
        <p>
          Finally, we output scores for positions where the ref- quencing kit (SQK-PBK004) using a Blunt/TA Ligase
Maserence sequence matches a user-supplied pattern recog- ter Mix (New England Biolabs). DNA samples (∼ 10 ng)
nized by the methyltransferase enzyme potentially active were then amplified using the Rapid Barcode primers
in a given sample. For each position with the pattern, (LWB, SQK-PBK004) and LongAmp Taq DNA polymerase
we consider a window of length 11 centered on the po- (New England Biolabs) and treated by exonuclease I (New
tentially modified base and report the maximum score England Biolabs). DNA aliquots (∼ 1 µg) were modified
within this window. using 160 µM S-adenosyl methionine and
methyltransferases M.EcoRI (40 U), M.BamHI (12 U), or M.HhaI (25
2.7. Nanopore Sequencing U) (New England Biolabs) for 1 hour at 37°C. The
methyltransferase reactions were then stopped at 65°C for 20
High-molecular weight (HMW) genomic DNA was pre- min. The modification by M. TaqI (10 U) was carried out
pared from an overnight culture of Jaminaea angkorensis at 65°C for 1 hour. DNA was purified on AMPure XP
C5b (CBS 10918; [
          <xref ref-type="bibr" rid="ref25">24</xref>
          ]) grown in yeast-peptone-dextrose beads and eluted into 10 mM Tris.Cl pH 8.0, 50 mM NaCl.
(YPD) medium (1 % (w/v) yeast extract, 2 % (w/v) peptone, Methylated and control barcoded DNAs were pooled in
2 % (w/v) glucose) at 28°C with constant aeration. Yeast equal ratio. Rapid adapters (RAP, SQK-PBK004) were
cells were harvested by centrifugation, washed with 20 attached to the ends of pooled DNAs (∼ 400 ng). The
library preparation, flow cell priming, and loading were
completed as described in the protocol (SQK-PBK004,
version PBK_9073_v1_revA_23May2018).
2.8. Training and Testing Data Sets
        </p>
      </sec>
      <sec id="sec-1-3">
        <title>All nanopore reads were base called by Albacore version</title>
        <p>
          2.3.1. Base called reads were aligned to their respective
reference genomes by BWA-MEM [
          <xref ref-type="bibr" rid="ref27">26</xref>
          ]. The reads that did
not align to the reference for at least 80% of their length
were discarded. Table 1 shows basic characteristics of
resulting data sets.
        </p>
        <p>Data set control0 was used for the purpose of training
extended signal model. Mixtures of reads from control1
and met-* data sets were used for the purpose of testing
the application of our methods to DNA methylation
detection. Note that control0 and control1 data sets were
produced from phylogenetically distant species and thus
there is no overlap between training and testing sets.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>3. Results and Discussion</title>
      <p>3.1. Realignment by Nadavca Eliminates</p>
      <p>Apparent Alignment Artefacts</p>
      <sec id="sec-2-1">
        <title>Alignment of the signal to the reference sequence seg</title>
        <p>ments the signal into events, each event corresponding To evaluate the efect of Nadavca alignments on
downto a diferent context read by the pore. If the DNA moved stream analysis, we have implemented a simple tool
through the pore at a constant speed, each event would nanometh for detection of methylation and other DNA
span approximately 9 values. modifications from the nanopore signals. We have used</p>
        <p>
          We have used Tombo [
          <xref ref-type="bibr" rid="ref19">18</xref>
          ], Nanopolish [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], and Na- the following general framework for detection of
modifidavca to realign the signal to the reference sequence, cations (see Methods for details):
using data set control1. Figure 5 shows the comparison
of event length distributions. Only Nanopolish allows 1. Use Nadavca to align signal values of a read to
events of length 0 (skip of the context), Tombo appar- the reference sequence
ently has a minimum event length 3 (even though rarely 2. For each event implied by the alignment, compare
it also outputs events of length 1 and 2), and in Nadavca the mean signal value within the event to the
we require each event to be of length at least 2. With expected signal level from our enhanced signal
Tombo, there is a clear bias against events of length 4, and model, producing a p-value
preference for events of lengths 3 and 6; which appears 3. P-values from several adjacent positions are
comto be an artefact of the alignment process. Interestingly, bined to a single score
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Nanopolish does not report events of length 1 or 2. In case of Nadavca, the only apparent artefact is a higher abundance of events of length 2, which is a consequence of the requirement of minimum length 2 for each event.</title>
        <p>3.2. Unsupervised Single-Molecule
Single-Site DNA Modification
Detection</p>
      </sec>
      <sec id="sec-2-3">
        <title>We compare our results with de novo mode of Tombo.</title>
        <p>Note that unlike most other methods, Nadavca and Tombo
do not require any training, besides training the sequence
context model. To make the results comparable, we use
similar formulas for computing the p-values and scores
as Tombo, however, substituting diferent signal
alignment and diferent signal context model. For Tombo we
have used parameters supplied directly with the tool.</p>
        <p>
          We use a newly produced testing set from the genome
of fungus Jaminaea angkorensis [
          <xref ref-type="bibr" rid="ref25">24</xref>
          ]. Genomic DNA was
amplified by PCR, producing DNA devoid of any native
modifications, and then one part (control) was left
unmodified, and each of four additional parts were modified
by a diferent methyltransferase enzyme. Each enzyme
methylates either adenine or cytosine at a fixed position
within a specific sequence pattern (for example, M.HhaI
methylates the first cytosine in the pattern GCGC). In
each test, we use a mixture of reads from the control
sample and from one of the modified samples. We consider
only positions matching the pattern of the corresponding
enzyme (in the reference genome) and classify them as
positives or negatives based on the sample of origin. Note
that the signal model was trained on a data set from a
diferent organism to avoid overfitting.
        </p>
        <p>We do not attempt to set a threshold for calling a site
methylated, instead we compute AUC score which
evaluates each method over all choices of the threshold. In
the computation of accuracy, each pattern occurrence
within each read is considered separately. Table 2 shows
an overview of results and comparison to Tombo de novo
mode. On all of our methylation data sets, nanometh
significantly outperforms Tombo.
3.4. Contribution of Individual Algorithm
Features
3.3. Comparison to Supervised</p>
        <p>Methylation Detection</p>
      </sec>
      <sec id="sec-2-4">
        <title>Unlike Nadavca/nanometh, most tools for methylation detection are supervised, requiring a separate training set</title>
      </sec>
      <sec id="sec-2-5">
        <title>We were also interested, how individual features of Na</title>
        <p>davca afect the accuracy. Therefore, we started with a
baseline modification detection tool and added individual
features one at a time: posterior alignment,
renormalization, modeling transitional signal states, and using an
extended k-mer model. The results are shown in Table 3.
Replacing Tombo resquiggle with posterior alignment
with renormalization already helps to increase the
ac</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>4. Conclusions</title>
      <sec id="sec-3-1">
        <title>Acknowledgements. The strain of J. angkorensis was</title>
        <p>In this work, we introduced Nadavca, a nanopore sig- kindly provided by Matthias Sipiczki (University of
Denal aligner that incorporates several enhancements to brecen, Hungary). The strain of M. capitatus was kindly
the DTW algorithm. Compared to existing aligners, Na- provided by Cletus P. Kurtzman and James Swezey
(Agridavca’s output exhibits improved accuracy by eliminating cultural Research Service, Peoria, IL, USA)
length distribution artifacts and eliminating the need for
event segmentation as a preliminary step. We demon- References
strated the eficacy of our alignment method by achieving
state-of-the-art results in unsupervised methylation
detection.</p>
        <p>
          Most of the current tools for methylation detection
in nanopore signal are supervised and based on deep
learning (see e.g. Guppy, DeepSignal [
          <xref ref-type="bibr" rid="ref15">14</xref>
          ], DeepMod [
          <xref ref-type="bibr" rid="ref29">28</xref>
          ],
Megalodon, or Remora). A disadvantage of these tools is
that they require a substantial amount of labeled training
data. Such data sets are available for the most common
methylation types (e.g. CpG methylation) for which those
tools were trained, but it is impossible to use these tools
to discover less common or novel DNA modifications,
where an unsupervised approach is needed. While
unsupervised approaches are at a clear disadvantage when
compared to supervised approaches in detection of
common methylation types, by incorporating additional
information such as afected signal positions due to
methylation, our approach yields comparable results to
Nanopolish (non-deeep-learning supervised approach) that has
recently been shown to perform very well among the
supervised methylation detection methods [
          <xref ref-type="bibr" rid="ref28">27</xref>
          ].
        </p>
        <p>In the future, we aim to further refine our alignment
technique by addressing alignment artifacts, particularly
in regions where the expected signals for two events
are similar, resulting in ambiguous alignments. Another
task is to adapt our methods to the new generation R10
nanopores.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>H.</given-names>
            <surname>Teng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. D.</given-names>
            <surname>Cao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. B.</given-names>
            <surname>Hall</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Duarte</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. J.</given-names>
            <surname>Coin</surname>
          </string-name>
          ,
          <article-title>Chiron: translating nanopore raw signal directly into nucleotide sequence using deep learning</article-title>
          ,
          <source>GigaScience</source>
          <volume>7</volume>
          (
          <year>2018</year>
          )
          <article-title>giy037</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>V.</given-names>
            <surname>Boža</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Brejová</surname>
          </string-name>
          , T. Vinař,
          <article-title>DeepNano: deep recurrent neural networks for base calling in MinION nanopore reads</article-title>
          ,
          <source>PloS One</source>
          <volume>12</volume>
          (
          <year>2017</year>
          )
          <article-title>e0178751</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>V.</given-names>
            <surname>Boža</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Perešíni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Brejová</surname>
          </string-name>
          , T. Vinař,
          <article-title>Dynamic pooling improves nanopore base calling accuracy</article-title>
          ,
          <source>IEEE/ACM Transactions on Computational Biology and Bioinformatics</source>
          <volume>19</volume>
          (
          <year>2021</year>
          )
          <fpage>3416</fpage>
          -
          <lpage>3424</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>R. R.</given-names>
            <surname>Wick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. M.</given-names>
            <surname>Judd</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. E.</given-names>
            <surname>Holt</surname>
          </string-name>
          ,
          <article-title>Performance of neural network basecalling tools for oxford nanopore sequencing</article-title>
          ,
          <source>Genome Biology</source>
          <volume>20</volume>
          (
          <year>2019</year>
          )
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>N.</given-names>
            <surname>De</surname>
          </string-name>
          <string-name>
            <surname>Souza</surname>
          </string-name>
          ,
          <article-title>Protein nanopores to detect dna methylation</article-title>
          ,
          <source>Nature Methods</source>
          <volume>11</volume>
          (
          <year>2014</year>
          )
          <fpage>8</fpage>
          -
          <lpage>8</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Stoiber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Quick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Egan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. Eun</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Celniker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. K.</given-names>
            <surname>Neely</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Loman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. A.</given-names>
            <surname>Pennacchio</surname>
          </string-name>
          , J. Brown,
          <article-title>De novo identification of DNA modifications enabled by genome-guided nanopore signal processing</article-title>
          , bioRxiv (
          <year>2017</year>
          ). URL: https://www.biorxiv.org/content/ early/2017/04/10/094672. doi:
          <volume>10</volume>
          .1101/094672.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J. B.</given-names>
            <surname>Kruskal</surname>
          </string-name>
          ,
          <article-title>An overview of sequence comparison: Time warps, string edits, and macromolecules</article-title>
          ,
          <source>SIAM Review 25</source>
          (
          <year>1983</year>
          )
          <fpage>201</fpage>
          -
          <lpage>237</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>R.</given-names>
            <surname>Durbin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. R.</given-names>
            <surname>Eddy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Krogh</surname>
          </string-name>
          , G. Mitchison,
          <article-title>Biological sequence analysis: probabilistic models of proteins and nucleic acids</article-title>
          , Cambridge University Press,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>G.</given-names>
            <surname>Lunter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rocco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Mimouni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Heger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Caldeira</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Hein,</surname>
          </string-name>
          <article-title>Uncertainty in homology inferences: assessing and improving genomic sequence alignment</article-title>
          ,
          <source>Genome research 18</source>
          (
          <year>2008</year>
          )
          <fpage>298</fpage>
          -
          <lpage>309</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>L.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Seki</surname>
          </string-name>
          ,
          <article-title>Recent advances in the detection of base modifications using the Nanopore sequencer</article-title>
          ,
          <source>Journal of Human Genetics</source>
          <volume>65</volume>
          (
          <year>2020</year>
          )
          <fpage>25</fpage>
          -
          <lpage>33</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J. T.</given-names>
            <surname>Simpson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. E.</given-names>
            <surname>Workman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Zuzarte</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>David</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Dursi</surname>
          </string-name>
          , W. Timp,
          <article-title>Detecting DNA cytosine methylation using nanopore sequencing</article-title>
          ,
          <source>Nature Methods</source>
          <volume>14</volume>
          (
          <year>2017</year>
          )
          <fpage>407</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>A. C.</given-names>
            <surname>Rand</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Eizenga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Musselman-Brown</surname>
          </string-name>
          , H. E.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Funding.</surname>
          </string-name>
          <article-title>The research presented in this paper has been supported by funding from the Slovak Research and Development Agency grant APVV-18-0239 (JN), and by</article-title>
          <string-name>
            <surname>Olsen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Akeson</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Paten</surname>
          </string-name>
          ,
          <article-title>Mapping DNA methylation with high-throughput nanopore sequencing</article-title>
          ,
          <source>Nature Methods</source>
          <volume>14</volume>
          (
          <year>2017</year>
          )
          <fpage>411</fpage>
          -
          <lpage>413</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [13]
          <string-name>
            <surname>I. Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Razaghi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Gilpatrick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Molnar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gershman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Sadowski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. J.</given-names>
            <surname>Sedlazeck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. D.</given-names>
            <surname>Hansen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. T.</given-names>
            <surname>Simpson</surname>
          </string-name>
          , W. Timp,
          <article-title>Simultaneous profiling of chromatin accessibility and methylation on human cell lines with nanopore sequencing</article-title>
          ,
          <source>Nature Methods</source>
          <volume>17</volume>
          (
          <year>2020</year>
          )
          <fpage>1191</fpage>
          -
          <lpage>1199</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>P.</given-names>
            <surname>Ni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.-P.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Miao</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.-L. Xiao</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Luo</surname>
            ,
            <given-names>J. Wang,</given-names>
          </string-name>
          <article-title>DeepSignal: detecting DNA methylation state from Nanopore sequencing reads using deep-learning</article-title>
          ,
          <source>Bioinformatics</source>
          <volume>35</volume>
          (
          <year>2019</year>
          )
          <fpage>4586</fpage>
          -
          <lpage>4595</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [15]
          <string-name>
            <surname>A. B. McIntyre</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Alexander</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Grigorev</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Bezdan</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Sichtig</surname>
            ,
            <given-names>C. Y.</given-names>
          </string-name>
          <string-name>
            <surname>Chiu</surname>
            ,
            <given-names>C. E.</given-names>
          </string-name>
          <string-name>
            <surname>Mason</surname>
          </string-name>
          ,
          <article-title>Single-molecule sequencing detection of N 6-methyladenine in microbial reference materials</article-title>
          ,
          <source>Nature Communications</source>
          <volume>10</volume>
          (
          <year>2019</year>
          )
          <fpage>579</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Fang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.-L. Xiao</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Detection of DNA base modifications by deep recurrent neural network on Oxford Nanopore sequencing data</article-title>
          ,
          <source>Nature Communications</source>
          <volume>10</volume>
          (
          <year>2019</year>
          )
          <fpage>2449</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. C.</given-names>
            <surname>Georgieva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Egli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>NanoMod: a computational tool to detect DNA modifications using Nanopore long-read sequencing data</article-title>
          ,
          <source>BMC Genomics 20</source>
          (
          <year>2019</year>
          )
          <fpage>31</fpage>
          -
          <lpage>42</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [18] Oxford Nanopore Technologies, Tombo documentation,
          <year>2018</year>
          . https://nanoporetech.github.io/tombo/.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>N. J.</given-names>
            <surname>Loman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Quick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. T.</given-names>
            <surname>Simpson</surname>
          </string-name>
          ,
          <article-title>A complete bacterial genome assembled de novo using only nanopore sequencing data</article-title>
          ,
          <source>Nature Methods</source>
          <volume>12</volume>
          (
          <year>2015</year>
          )
          <fpage>733</fpage>
          -
          <lpage>735</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>H.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM</article-title>
          ,
          <source>arXiv preprint arXiv:1303.3997</source>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>V.</given-names>
            <surname>Boža</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Brejová</surname>
          </string-name>
          , T. Vinař,
          <article-title>Improving nanopore reads raw signal alignment</article-title>
          ,
          <source>arXiv preprint arXiv:1705.01620</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>M.</given-names>
            <surname>David</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. J.</given-names>
            <surname>Dursi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Yao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. C.</given-names>
            <surname>Boutros</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. T.</given-names>
            <surname>Simpson</surname>
          </string-name>
          ,
          <article-title>Nanocall: an open source basecaller for Oxford Nanopore sequencing data</article-title>
          ,
          <source>Bioinformatics</source>
          <volume>33</volume>
          (
          <year>2016</year>
          )
          <fpage>49</fpage>
          -
          <lpage>55</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [23]
          <string-name>
            <surname>R. A</surname>
          </string-name>
          . Fisher,
          <article-title>Statistical methods for research workers</article-title>
          ,
          <source>Oliver and Boyd (Edinburgh)</source>
          ,
          <year>1925</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>M.</given-names>
            <surname>Sipiczki</surname>
          </string-name>
          , E. Kajdacsi,
          <article-title>Jaminaea angkorensis gen</article-title>
          . nov., sp. nov.,
          <article-title>a novel anamorphic fungus containing an S943 nuclear smallsubunit rRNA group IB intron represents a basal branch of microstromatales</article-title>
          ,
          <source>International Journal of Systematic and Evolutionary Microbiology</source>
          <volume>59</volume>
          (
          <year>2009</year>
          )
          <fpage>914</fpage>
          -
          <lpage>920</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>B.</given-names>
            <surname>Brejová</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Lichancová</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Brázdovič</surname>
          </string-name>
          , E. Hegedűsová,
          <string-name>
            <given-names>M. F.</given-names>
            <surname>Jakubková</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Hodorová</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Džugasová</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Baláž</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zeiselová</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Cillingová</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Neboháčová</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Raclavský</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Ľubomír</given-names>
            <surname>Tomáška</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. F.</given-names>
            <surname>Lang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Vinař</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Nosek</surname>
          </string-name>
          ,
          <article-title>Genome sequence of the opportunistic human pathogen Magnusiomyces capitatus</article-title>
          ,
          <source>Current Genetics</source>
          <volume>65</volume>
          (
          <year>2019</year>
          )
          <fpage>539</fpage>
          -
          <lpage>560</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>H.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM</article-title>
          ,
          <source>arXiv preprint arXiv:1303.3997</source>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>Z. W.</given-names>
            <surname>Yuen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Srivastava</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Daniel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>McNevin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Jack</surname>
          </string-name>
          , E. Eyras,
          <article-title>Systematic benchmarking of tools for CpG methylation detection from nanopore sequencing</article-title>
          ,
          <source>Nature Communications</source>
          <volume>12</volume>
          (
          <year>2021</year>
          )
          <fpage>3438</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Fang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. L.</given-names>
            <surname>Xiao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Detection of DNA base modifications by deep recurrent neural network on Oxford Nanopore sequencing data</article-title>
          ,
          <source>Nature Communications</source>
          <volume>10</volume>
          (
          <year>2019</year>
          )
          <fpage>2449</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>