<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Reading Type Classification based on Generative Models and Bidirectional Long Short-Term Memory</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Seyyed Saleh Mozaffari</string-name>
          <email>mozafari@dfki.uni-kl.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stefan Agne</string-name>
          <email>Stefan.Agne@dfki.de</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Federico Raue</string-name>
          <email>Federico.Raue@dfki.de</email>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Syed Saqib Bukhari</string-name>
          <email>Saqib.Bukhari@dfki.de</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Saeid Dashti Hassanzadeh</string-name>
          <email>saeid.dashti.hassanzadeh@tu-</email>
          <email>saeid.dashti.hassanzadeh@tuclausthal.de</email>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andreas Dengel</string-name>
          <email>dengel@dfki.uni-kl.de</email>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Author Keywords</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Chanijani, TU Kaiserslautern, German Research Center for</institution>
          ,
          <addr-line>Artificial Intelligence</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Eye Tracking; reading type; classification; synthetic data;, generative models; Hierarchical Hidden Markov Models;</institution>
          ,
          <addr-line>Gaussian Mixture Models; LSTM; Recurrent Neural, Networks; reading; skimming; scanning;</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>German Research Center for</institution>
          ,
          <addr-line>Artificial Intelligence</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>German Research Center for</institution>
          ,
          <addr-line>Artificial Intelligence</addr-line>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>TU Clausthal</institution>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>TU Kaiserslautern, German Research Center for</institution>
          ,
          <addr-line>Artificial Intelligence</addr-line>
        </aff>
        <aff id="aff6">
          <label>6</label>
          <institution>TU Kaiserslautern, German Research Center for</institution>
          ,
          <addr-line>Artificial Intelligence</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>Measuring the attention of users is necessary to design smart Human Computer Interaction (HCI) systems. Particularly, in reading, the reading types, so-called reading, skimming, and scanning are signs to express the degree of attentiveness. Eye movements are informative spatiotemporal data to measure quality of reading. Eye tracking technology is the tool to record eye movements. Even though there is increasing usage of eye trackers in research and especially in psycholinguistics, collecting appropriate task-specific eye movements data is expensitive and time consuming. Moreover, machine learning tools like Recurrent Neural Networks need large enough samples to be trained. Hence, designing a generative model in order to have reliable research-oriented synthetic eye movements is desirable. This paper has two main contributions. First, a generative model in order to synthesize reading, skimming, and scanning in reading is developed. Second, in order to evaluate the generative model, a bidirectional Long ShortTerm Memory (BLSTM) is proposed. It was trained with synthetic data and tested with real-world eye movements to classify reading, skimming, and scanning where more than 95% classification accuracy is achieved.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>©2018. Copyright for the individual papers remains with the authors.</p>
      <p>Copying permitted for private and academic purposes.</p>
      <p>
        UISTDA ’18, March 11, 2018, Tokyo, Japan
INTRODUCTION
Reading, is the ability to extract visual information from the
page and comprehend the meaning of underlying text [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
      </p>
      <p>Considering attention, as presented in Figure 1, the reading
types is divided into three categories: reading, skimming, and
scanning. On eye tracking context, the reading is a method of
moving the eyes over the text to comprehend the meaning of
it. The skimming is a rapid eye movement over the document
with the purpose of getting only the main ideas and a general
overview of the document whereas scanning rapidly covers
a lot of contexts in order to locate specific fact or piece of
information.</p>
      <p>
        The fixation progress on words expressed in character
units must be measured in order to detect the reading type,
i.e., deciding whether observed eye movement patterns in
the reading types [
        <xref ref-type="bibr" rid="ref12 ref4">4, 12</xref>
        ]. This approach applies in cases
where the eye tracking accuracy is high enough to provide
word level resolution [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. Example applications include
ScentHighlight [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], which highlights related sentences
during reading; the eyeBook [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], where ambient effects
are to be triggered in proximity of the reading position; or
QuickSkim [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], where non-content words may be faded out in
real time with an increase of skimming speed to make reading
more efficient.
      </p>
      <p>Due to the noisy nature of the eye tracking apparatus
where the point of gaze cannot be determined exactly, it
can be desirable to automatically decide to what extent eye
movements resemble a reading type patterns. In this regard,
a psycholinguist is able to determine what segments of the
scanpath belong to reading, skimming, or scanning, even
though the fixations do not match the underlying text.</p>
      <p>
        Biedert et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] proposed a reading-skimming classifier.
      </p>
      <p>
        Despite the model classifies the reading and the skimming
patterns, it does not cover scanning patterns. It is desirable
to have scanning patterns in order to have better estimation
on the degree of attiontion in reading to designing a proper
Human Document Interaction system. In addition, in the
domain of information retrieval, it has been shown that
acquiring implicit feedback from a reading type detection,
including scanning, can significantly improve search accuracy
through personalization [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        Detecting of the reading types is a sequence classification
problem. The term sequence classification encompasses all
tasks where sequences of data are transcribed with sequences
of discrete labels [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. The discrete labels in the reading
type classification problem are shown in Figure 1. Long
Short-Term Memory (LSTM) [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] is a variation of Recurrent
Neural Networks (RNNs) which is suitable to classify the
sequential data. However, such networks need large enough
samples for training. Unfortunately, the task-specific eye
tracking data size would not big enough to be applied in
the deep networks. Hence, there is a need to synthesize
task-specific eye movement patterns in order to deploy deep
neural networks.
      </p>
      <p>In this paper, we propose a generative model which
synthesizes the reading types patterns. We also designed and
evaluated a BLSTM model. The model trained on both the
original dataset as well as the synthetic dataset.</p>
      <p>The paper is structured as follows. We start with presenting
an experiment conducted to collect real-world eye movements
in reading in order to build a reference for data synthesization
and the reading type classification. Then, a two-layered
Hierarchical Hidden Markov Model (HHMM) for eye
movement data synthesization is proposed. Moreover, we
present a BLSTM-based sequential model to detect and
classify reading, skimming, and scanning. This model built
on features described in section Features and Training. In the
Evaluation and Result section, we evaluate our models and
describe our results. This is followed by our conclusion.</p>
      <p>REAL-WORLD DATA ACQUISITION
The first step of constructing a system which is able to learn
and distinguish eye movement patterns of reading types is
to record the real-world eye movement data during reading
mode. The recorded data must comprise all possible state
categories: reading, skimming, and scanning. To perform the
task, we designed an experiment to record eye movements of
ten participants from the local university. We used this data as
a reference to build up a Hierarchical Hidden Markov Model
(HHMM) for synthetic data generation. Furthermore, this
realworld data partially employed for testing and evaluating the
classifiers.</p>
      <p>
        Apparatus
In this study, we deploy SensoMotoric Instrument iViewX
scientific REDn eye tracker operating at 60Hz. The tracking
error reported by the manufacturer was less than 0.4 degree,
which makes it appropriate for fixation-based eye movement of them were native Engish speaker but they were fluent in
studies. 1 English as the second language. In the first phase of the
experiment, the participants were requested to write a
compreExperimental Setup hensive report on the article they had given to read. Hence,
The experiment was designed in which all the three reading they would read the selected article thoroughly. In the second
types could be obtained. Two articles in plain English were phase, the participants were asked to find the specific
informachosen from Wikipedia. They are about two airplane crashes tion in the second article, e.g., how many crews were in the
took place in Colombia2 and Pakistan3 in 2016. Ten partic- airplane or what was the flight number. Therefore, most of the
ipants from local university participated in our study. None eye movement patterns associated with skimming and
scanning. Consequently, all three reading type patterns recorded
1https://www.smivision.com/eye-tracking/product/redn-scientific- during trials. The trials have been recorded in specialized
eyeeye-tracker/ tracking interface eyeReading [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. The eyeReading facilitates
2https://en.wikipedia.org/wiki/LaMia_Flight_2933 research in reading psychology and provides a framework for
3http://en.wikipedia.org/wiki/Pakistan_International_Airlines_Flight_661 gaze-based Human Document Interaction. Figure 2 shows
top-level architecture of eyeReading.
      </p>
      <p>FEATURES AND ANNOTATION
The recorded raw gaze information must be processed in order
to extract saccadic features associated with reading. The
extracted saccadic features are the length of the saccade, velocity
of the saccadic movement, fixation duration associated with
the saccades, and angularity of the saccade. In this section,
first, we demonstrate the features extraction step and then the
process of making ground-truth will be explained.</p>
      <p>Features
On account of inevitable noise in the eye tracking trials, we
first applied a virtual median filter on the input raw gaze points
E0 = e01; :::; e0n to eliminate any possible outliers.</p>
      <p>ei = (medx(e0i 2; :::; e0i+2); medy(e0i 2; :::; e0i+2))</p>
      <p>
        (1)
In the second step, the fixations were detected using dispersion
method [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. We considered 100ms temporal and 50px spatial
for dispersion parameters. The saccades are considered as two
consecutive fixations with the following features:
1. amplitude(`): the distance between two progressive
      </p>
      <p>fixations in virtual character unit (vc).
2. angularity(q ): the angle of the saccade respect to its
starting point. The a indicates the direction of the saccade
in circle domain: 180 a 179 .
3. velocity(n): the speed of the saccade:
n =</p>
      <p>`
te ts</p>
      <p>(2)
where ts and te are the first timestamps in the start fixation
and the end fixation of a saccade in milliseconds.
4. duration(g): The start fixation duration in each saccade in</p>
      <p>milliseconds respectively.</p>
      <p>Therefore, F(`; q ; n; g) are the features selected for the
saccades in the collected data.</p>
      <p>Data Annotation
After the feature extraction step, we designed a labeling
application to make ground truth from data. The figure 4 shows
the interface of the labeling application. Ground-truthing data
was accomplished in two steps: the saccade labeling and the
sequence labeling.</p>
      <p>
        Saccade Labeling
In the first step of labeling, the expert made a judgment
about the saccades with respect to the provided features
F : `; a; n; q .The saccades grouped into six categories or
labels [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
      <p>MFR(l, θ, ν, γ)
COVFR(l, θ, ν, γ)</p>
      <p>MFR(l, θ, ν, γ)
COVRG(l, θ, ν, γ)</p>
      <p>MFR(l, θ, ν, γ)
COVSW(l, θ, ν, γ)</p>
      <p>State i
LS
SR</p>
      <p>MFR(l, θ, ν, γ)
COVFS(l, θ, ν, γ)</p>
      <p>MFR(l, θ, ν, γ)
COVLS(l, θ, ν, γ)</p>
      <p>MFR(l, θ, ν, γ)</p>
      <p>COVSR(l, θ, ν, γ)</p>
      <p>
        FR: Forward Read (FR) is a progressive saccade associated
with reading. The amplitude is between 7 to 10 characters
[
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
      </p>
      <p>FS: Forward Skim(FS)is also a progressive saccade which
the amplitude is larger than FR but not too large.</p>
      <p>LS: Long Saccades(LS) is those saccades which have the
bigger amplitude than the threshold considered for the
context. The direction of saccade does not apply to LS.</p>
      <p>RG: The regressions(RG) are regressive saccades usually
associated with reading which is the sign of difficulties in
reading. The amplitude is varied and it must target the
passed context.</p>
      <p>SR: Sweeping back to the left on the next line of text is
called the Sweep Returns (SR).</p>
      <p>SW : Unstructured sweeping the text to look up information
are labeled as Sweeps (SW ).</p>
      <p>Sequence Annotation
After all the saccades were labeled in the previous section into
the six categories, in the second step, the sequences made by
these saccades annotated as reading, skimming, or scanning.</p>
      <p>At the end, 396 annotated sequences for reading, 378 for
skimming, and 118 for scanning were acquired.</p>
      <p>SYNTHETIC EYE MOVEMENTS IN READING
It is always desirable to have enough data samples to construct
robust machine learning models. Especially, in deep neural
networks, a very big training set is usually required.
Unfortunately, appropriate data acquisition in eye tracking studies is
Algorithm 1: Algorithm to simulate a HMM states sequence
S given the model l = fP; A; Bg.</p>
      <p>Data: statesground truth = s1; s2; :::; sn and
observationground truth = o1; o2; :::; on where n is the
number of saccades in the ground-truth.</p>
      <p>Result: States sequence Ssequence = ([s1; l1]; [s2; l2]; :::; [sk; lk])
where k is number of sequences, s is the sequence
label si 2 C, and li is the length of si.
1 P = (0:34; 0:33; 0:33);
2 A = fai jji = 1; 2; 3; j = 1; 2; 3g: State transition probability</p>
      <p>3
where ai j = P(st+1 = jjst = i) and å j=1 ai j = 1;
3 B = fbk(ot )jk = 1; 2; 3; t = 1; :::; ng: Observation probability</p>
      <p>where bk(ot ) = P(ot jst = k);
4 Choose an initial state S1 according to the initial state</p>
      <p>distribution p;
5 for time t = f1; :::; ng do
6 Draw ot from the probability distribution Bst;
7 Go to state st + 1 according to the transition probabilities
8</p>
      <p>Ast ;</p>
      <p>Set t = t + 1;
Algorithm 2: Algorithm to generate the emissions sequence
S0 from S.</p>
      <p>Data: S and observationground truth = o1; o2; :::; on where n</p>
      <p>is the number of saccades in the ground-truth data.</p>
      <p>Result: synthetic emissions for all s 2 S
1 P0 = (p10; :::; p60): Initial state probabilities where</p>
      <p>6
pi0 = P(s01 = i) and åi=1 pi0 = 1;
2 A0 = fa0i jji = 1; :::; 6; j = 1; :::; 6g: Emissions transition</p>
      <p>6
probability where a0i j = P(st0+1 = jjst0 = i) and å j=1 a0i j = 1;
3 foreach states reading; skimming; scanning calculate the
covariance matrix COV of the emissions</p>
      <p>FR; FS; RG; LS; SW; SR for the features `; q ; n; g;
4 foreach states reading; skimming; scanning calculate the
mean matrix M of the emissions FR; FS; RG; LS; SW; SR for
the features `; q ; n; g;
5 Choose an initial state S10 according to the initial state</p>
      <p>distribution P0;
6 for time t = f1; :::; ng do
7 Draw ot from the Gaussian distribution with covst and</p>
      <p>meanst ;
8 Go to state st + 1 according to the transition probabilities</p>
      <p>A0st ;
9 Set t = t + 1;
not an easy task. It needs eye trackers which are still
expensive in the market as well as an appropriate experimental setup
designed for a specific goal, i.e., reading type classification.</p>
      <p>
        Hence, the idea is to generate task-specific eye movement data
from the real-world data in which these synthetic data can be
used to construct a better model and even to use in other
applications and research. It motivated us to design a Hierarchical
Hidden Markov Model (HHMM) to generate synthetic eye
movements in reading. Usually, ordinary Markov chains are
often not flexible enough for the analysis of real-world data, as
the state corresponding to a specific event (observation) has to
be known. However, in many problems of interest, this is not
given. Hidden Markov models (HMM) as originally proposed
by Baum et al. (1970) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] can be viewed as an extension of
Markov chains. The only difference compared to common
Markov chains is, that the state sequence corresponding to a
particular observation sequence, i.e., reading types in our case,
is not observable but hidden. In other words, the observation
is a probabilistic function of the state, whereas the underlying
state sequence itself is a hidden stochastic process [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. That
means the underlying state sequence can only be observed
indirectly through another stochastic process that emits an
observable output. Hidden Markov models are extremely popular
when dealing with sequential data, such as speech recognition,
character recognition, gesture recognition as well as
biological sequences. Therefore, the HMM is a right candidate to
handle the eye movement patterns where they are sequential
and by nature stochastic. In order to synthesize our data, the
graphical model should be able to generate both saccadic
sequences and reading state sequences. Therefore, in this paper,
a two-layered Hierarchical HMM is designed. In an HHMM
each state is considered to be a self-contained probabilistic
model [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Briefly, in the first layer as shown in Figure 5 we
modeled the reading, skimming, and scanning as states of the
Markov model and emissions are FR (Forward Reading), FS
(Forward Skimming), SR (Sweep Return), RG (Regression),
SW (Sweeps), LS (Long Saccades). As shown in Figure 6,
each of states in the first level is self-contained mixture
graphical model so-called GMM-HMM (Hidden Markov Model with
Gaussian Mixture Model). This layer responsible to generate
values for the four mentioned saccadic features F : `; a; n; q .
      </p>
      <p>Method
The 2-layered HHMM constructs the probabilistic model that
generates saccades associated with the reading types. It is a
top-down approach to synthesize natural reading types. The
task of the first layer, which is shown in Figure 5, is to generate</p>
    </sec>
    <sec id="sec-2">
      <title>Precision</title>
      <p>0.81
0.89
0.93</p>
    </sec>
    <sec id="sec-3">
      <title>Recall</title>
      <p>0.79
0.90
0.93
the sequence of the states reading, skimming, and scanning.
Algorithm 1 constructs this layer. In order to build the 1st layer
HMM, l (p; A; B), the transition state matrix A and emission
matrix B are built upon the labeled data explained in section
Data Labeling. We considered equal probabilities (33%) for
the states of reading types in p. Then, the states reading,
skimming, and scanning is generated based on multinomial
distribution.</p>
      <p>The algorithm 2 presents the construction of the second layer
of our graphical model so-called GMM-HMM model. Where
the input is the state sequences produced from the first layer.
In contrast to the first layer, the emissions (observations)
associated with each sequence are generated based on Gaussian
distribution. Hence, for each state (reading, skimming, and
scanning), it needs to compute the transition matrix of
observation, the mean of each component (features F : `; q ; n; g) as
well as the covariance matrix of the features. Figure 6 presents
the second layer of the model.</p>
      <p>
        READING TYPE CLASSIFICATION WITH BLSTM
Recurrent neural networks (RNNs) are able to access a wide
range of context and sequences [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. However, standard RNNs
make use of the previous context only whereas bidirectional
RNNs (BRNNs) are able to incorporate emissions on both
sides of every position in the input sequence [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. This is
useful in the problem of reading type detection since it is often
necessary to look at both sides to the right and to the left of
a given sequence in order to identify it. BLSTM is a BRNN
that has hidden layers, which are made up of the so-called
Long Short-Term Memory (LSTM) cells. LSTM is an RNN
architecture specifically designed to bridge long time delays
between relevant input and target events, making it suitable for
problems where long-range emission sequences are required
to disambiguate individual labels [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. In fact, BLSTM
networks suit well for the reading type detection. Figure 7 shows
BLSTM-RNN architecture and
Figure 8 presents the model implemented for our study with
keras [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. It consists of two BLSTMs, one dropout to prevent
overfitting, and two dense networks. Where n is the sequence
length of the input. The loss function is categorical
crossentropy and SoftMax is used as the activation function.
EVALUATION AND RESULTS
In this section, the evaluation of both generative model and the
classifier is presented. Here, two types of the data are available;
the original data recorded with the eye tracker (actual data)
and the synthetic data. The actual data was randomly split into
the train set 60%, validation set 10%, and test set 30%.First
half of the actual data was used in the generative model for
data synthesization. The other half used for testing. In all
cases, the train set first fitted with standard scale function to
scale the mean (m) to 0 and the standard deviation(d ) to 1.
Then the validation and test set transformed respect to fitted
data.
      </p>
      <p>
        The model trained with different sequence length N = 5; 8; 10.
Baseline: Actual data and SVM-RBF classifier
The baseline is to test and evaluate SVM-RBF classifier
proposed by Biedert et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. For the test set, there are only
196 sequences to support the model. 75 for reading, 83 for
skimming, and 38 scanning. By 5 cross-fold validation the
best accuracy acquired was 69% with parameters C = 1000
and gamma = 0:001. With a closer look at the confusion
matrix in figure 9, it is obvious that there is an unacceptable
confusion in the class scanning. This problem is on account
of the less number of supports for the class scanning where
there is just 19% of the class labels. Another reason is about
sequential characteristics of the data. It shows Support Vector
Machines are not the best machine learning model tailored for
such data.
      </p>
      <p>
        Proposed method: Actual data and BLSTM
The model presented in Figure 8 is used to train and test the
original data. The model trained with different sequence length
N = 5; 8; 10. In case, the input sequence has a different length,
the sequence padded or truncated to the fixed length n. Table
1 show the accuracies for the different length. 92.5% accuracy
achieved for sequence length of 10.
roposed method: Synthetic data and BLSTM
Finally, the BLSTM model trained with synthetic data which
was generated with the half of the original dataset. The same
half of the original dataset also was used to validating the
model. The second half of the original dataset used for testing.
Table 2 presents the results for different variation of data size
and sequence length. While the larger sequence windows has
the better results in the model, an instant user interface favors
smaller sequences. The length of a sequence is related to
the number of fixations. If we consider the average fixation
duration in reading 250ms [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], the model must wait for the
input sequence for 2:5s. This result supports the reliability of
the synthetic data. The larger data, the better performance as
expected in deep net frameworks.
      </p>
      <p>CONCLUSION
Recurrent Neural Networks (RNNs) are suitable to model
spatiotemporal data, i.e, eye movements in reading. They usually
need large enough data sample to be trained. On account of
constraints in the experimental setups, accessibility to the
expensive eye tracking apparatus, and finding appropriate and
enough participants, there is usually lack of enough data in
eye tracking research to employ RNNs. In this paper a novel
probabilistic approach for eye movements data synthesization
in reading is proposed. Also a BLSTM model for both the
original recorded data and the synthetic data in order to
classify the reading types: reading, skimming, and scanning is
presented. The RNN-based classifier proposed in this paper
achieved more than 95% accuracy in the reading type
detection which not only outperforms the previous works but also
contains the scanning reading type.</p>
      <p>One important note is on the sterategy on selecting the
sequence length (N) for the model. Even though the longer
sequence length would lead to the higher accuracy, in order to
design instant user interface the shorter length is more
desirable as the model waits for N number of saccades to classify.
Depend on the application, the sequence length could be
selected occasionally.</p>
      <p>The outcome of this research is promising in which
appropriate data synthesization breaks limitations on using RNNs in
eye tracking research in general. It also may offers the
possibilities to provide standard eye movement datasets not only
for the reading type detection but for other research aspects
in reading i.e., in the research about dyslexia. Also, it gives
insight to other eye-tracking research to generate eye
movement transitions in different area of interests which is very
helpful to distinguish experts and novices in several domains
of education. It is also desirable to explore alternative to HMM
for data synthesization. In this regard, using LSTM itself as
a generative model for eye tracking data is in our agenda of
research for the future work.</p>
      <p>Acknowledgment
This work was funded by the Federal Ministry of Education
and Research(BMBF) for the project AICASys.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Leonard</surname>
            <given-names>E Baum</given-names>
          </string-name>
          , Ted Petrie, George Soules, and
          <string-name>
            <given-names>Norman</given-names>
            <surname>Weiss</surname>
          </string-name>
          .
          <year>1970</year>
          .
          <article-title>A maximization technique occurring in the statistical analysis of probabilistic functions of Markov chains</article-title>
          .
          <source>The annals of mathematical statistics 41</source>
          ,
          <issue>1</issue>
          (
          <year>1970</year>
          ),
          <fpage>164</fpage>
          -
          <lpage>171</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Ralf</given-names>
            <surname>Biedert</surname>
          </string-name>
          , Georg Buscher, and
          <string-name>
            <given-names>Andreas</given-names>
            <surname>Dengel</surname>
          </string-name>
          .
          <year>2010a</year>
          .
          <article-title>The eyebook-using eye tracking to enhance the reading experience</article-title>
          .
          <source>Informatik-Spektrum</source>
          <volume>33</volume>
          ,
          <issue>3</issue>
          (
          <year>2010</year>
          ),
          <fpage>272</fpage>
          -
          <lpage>281</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Ralf</given-names>
            <surname>Biedert</surname>
          </string-name>
          , Georg Buscher, Sven Schwarz,
          <source>Jörn Hees, and Andreas Dengel. 2010b. Text 2</source>
          .0.
          <string-name>
            <surname>In</surname>
            <given-names>CHI</given-names>
          </string-name>
          '10
          <source>Extended Abstracts on Human Factors in Computing Systems. ACM</source>
          ,
          <volume>4003</volume>
          -
          <fpage>4008</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Ralf</given-names>
            <surname>Biedert</surname>
          </string-name>
          , Jörn Hees, Andreas Dengel, and
          <string-name>
            <given-names>Georg</given-names>
            <surname>Buscher</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>A robust realtime reading-skimming classifier</article-title>
          .
          <source>In Proceedings of the Symposium on Eye Tracking Research and Applications. ACM</source>
          ,
          <volume>123</volume>
          -
          <fpage>130</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Georg</given-names>
            <surname>Buscher</surname>
          </string-name>
          , Andreas Dengel, Ralf Biedert, and
          <string-name>
            <surname>Ludger</surname>
            <given-names>V</given-names>
          </string-name>
          <string-name>
            <surname>Elst</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Attentive documents: Eye tracking as implicit feedback for information retrieval and beyond</article-title>
          .
          <source>ACM Transactions on Interactive Intelligent Systems (TiiS) 1</source>
          ,
          <issue>2</issue>
          (
          <year>2012</year>
          ),
          <fpage>9</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6. Ed H Chi, Lichan Hong, Michelle Gumbrecht, and Stuart K Card.
          <year>2005</year>
          .
          <article-title>ScentHighlights: highlighting conceptually-related sentences during reading</article-title>
          .
          <source>In Proceedings of the 10th international conference on Intelligent user interfaces. ACM</source>
          ,
          <volume>272</volume>
          -
          <fpage>274</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>François</given-names>
            <surname>Chollet</surname>
          </string-name>
          .
          <year>2015</year>
          . keras. https://github.com/fchollet/keras. (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>Shai</given-names>
            <surname>Fine</surname>
          </string-name>
          , Yoram Singer, and
          <string-name>
            <given-names>Naftali</given-names>
            <surname>Tishby</surname>
          </string-name>
          .
          <year>1998</year>
          .
          <article-title>The hierarchical hidden Markov model: Analysis and applications</article-title>
          .
          <source>Machine learning 32, 1</source>
          (
          <year>1998</year>
          ),
          <fpage>41</fpage>
          -
          <lpage>62</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>Alex</given-names>
            <surname>Graves</surname>
          </string-name>
          and others.
          <source>2012</source>
          .
          <article-title>Supervised sequence labelling with recurrent neural networks</article-title>
          . Vol.
          <volume>385</volume>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>Sepp</given-names>
            <surname>Hochreiter</surname>
          </string-name>
          and
          <string-name>
            <given-names>Jürgen</given-names>
            <surname>Schmidhuber</surname>
          </string-name>
          .
          <year>1997</year>
          .
          <article-title>Long short-term memory</article-title>
          .
          <source>Neural computation 9</source>
          ,
          <issue>8</issue>
          (
          <year>1997</year>
          ),
          <fpage>1735</fpage>
          -
          <lpage>1780</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>Kenneth</given-names>
            <surname>Holmqvist</surname>
          </string-name>
          , Marcus Nyström, Richard Andersson, Richard Dewhurst, Halszka Jarodzka, and Joost Van de Weijer.
          <year>2011</year>
          .
          <article-title>Eye tracking: A comprehensive guide to methods and measures</article-title>
          .
          <source>OUP Oxford.</source>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <given-names>Aulikki</given-names>
            <surname>Hyrskykari</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Eyes in attentive interfaces: Experiences from creating iDict, a gaze-aware reading aid</article-title>
          .
          <source>Tampereen yliopisto.</source>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Seyyed Saleh Mozaffari Chanijani</surname>
          </string-name>
          , Mohammad Al-Naser, Syed Saqib Bukhari, Damian Borth,
          <source>Shanley EM Alleny, and Andreas Denge</source>
          .
          <year>2016a</year>
          .
          <article-title>An eye movement study on scientific papers using wearable eye tracking technology</article-title>
          .
          <source>In Mobile Computing and Ubiquitous Networking (ICMU)</source>
          ,
          <source>2016 Ninth International Conference on. IEEE.</source>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Seyyed Saleh Mozaffari Chanijani</surname>
            , Syed Saqib Bukhari, and
            <given-names>Andreas</given-names>
          </string-name>
          <string-name>
            <surname>Dengel</surname>
          </string-name>
          .
          <year>2016b</year>
          .
          <article-title>eyeReading: Interaction with Text through Eyes</article-title>
          .
          <source>In Proceedings of The Ninth International Conference on Mobile Computing and Ubiquitous Networking</source>
          , Vol.
          <year>2016</year>
          . 1-
          <fpage>2</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Lawrence R Rabiner</surname>
          </string-name>
          .
          <year>1989</year>
          .
          <article-title>A tutorial on hidden Markov models and selected applications in speech recognition</article-title>
          .
          <source>Proc. IEEE 77</source>
          ,
          <issue>2</issue>
          (
          <year>1989</year>
          ),
          <fpage>257</fpage>
          -
          <lpage>286</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Keith</surname>
            <given-names>Rayner</given-names>
          </string-name>
          , Alexander Pollatsek,
          <source>Jane Ashby, and Charles Clifton Jr</source>
          .
          <year>2012</year>
          . Psychology of reading. Psychology Press.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <given-names>Raúl</given-names>
            <surname>Rojas</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Neural networks: a systematic introduction</article-title>
          . Springer Science &amp; Business Media.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <given-names>Mike</given-names>
            <surname>Schuster</surname>
          </string-name>
          and
          <article-title>Kuldip</article-title>
          K Paliwal.
          <year>1997</year>
          .
          <article-title>Bidirectional recurrent neural networks</article-title>
          .
          <source>IEEE Transactions on Signal Processing</source>
          <volume>45</volume>
          ,
          <issue>11</issue>
          (
          <year>1997</year>
          ),
          <fpage>2673</fpage>
          -
          <lpage>2681</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>