<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Visual Analytics Tools for Polycode Stimuli Eye Gaze Tracking in Virtual Reality</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Konstantin Ryabinin</string-name>
          <email>kostya.ryabinin@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Konstantin Belousov</string-name>
          <email>belousovki@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roman Chumakov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Eye Gaze Tracking, Virtual Reality, Visual Analytics, Image Segmentation, Polycode Stimulus</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Perm State University</institution>
          ,
          <addr-line>Bukireva Str., 15, Perm, 614068</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Saint Petersburg State University, University Embankment</institution>
          ,
          <addr-line>7/9, Saint Petersburg, 199034</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper is devoted to extending the previously created unified pipeline for conducting eye-trackingbased experiments in a virtual reality environment. In the previous work, we proposed using SciVi semantic data mining platform, Unreal Engine and HTC Vive Pro Eye head-mounted display to study reading process in the immersive virtual reality. The currently proposed extension enables to handle so-called polycode stimuli: compound visual objects, which consist of individual parts carrying diferent semantics for the viewer. To segment polycode stimuli extracting areas of interest (areas, where the informant's eye gaze is being tracked) we adopt Creative Maps Studio vector graphics editor. To integrate Creative Maps Studio into the existing pipeline we created plugins for SciVi platform to load and handle the segmented stimuli, place them in the virtual reality scenes, collect corresponding eye gaze tracking data and perform visual analysis of the data collected. To analyze the eye gaze tracks, we utilize a circular graph that allows comprehensive visualization of hierarchical areas of interest (mapping them to colorcoded graph nodes grouped into the hierarchy with a help of multilevel circular scale) and corresponding eye movements (mapped to the graph edges). We tested our pipeline on two diferent stimuli: the advertising poster and the painting “The Appearance of Christ Before the People” by A. Ivanov (1857).</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Precise tracking of eye movements is a mature methodology for studying the comprehension
processes in diferent human activities, including reading, viewing of pictures, and even spatial
navigation [
        <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4">1, 2, 3, 4</xref>
        ]. Often, eye tracking is used in the interdisciplinary research involving
specialists in physiology and diferent humanitarian sciences (like linguistics, sociology,
marketing, urbanistics, etc.). The latest advances in computer vision and corresponding hardware
drastically reduced the costs of eye tracking systems allowing them to be used by a wide range
of research groups all over the world [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ].
      </p>
      <p>
        Recently, several vendors introduced head-mounted displays (HMDs) with integrated infrared
eye trackers, enabling to detect eye gaze within virtual reality (VR) scenes [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. This, in turn,
brings eye-tracking-powered research to the new level, allowing to build and study virtual
environments, which are impossible to recreate in the real world [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        The data collected by the eye tracking device are to be analyzed using diferent metrics to
discover the meaningful information about the informants’ perception features. The most
common metrics involve fixations (moments when eyes are stationary), saccades (moments
when eyes rapidly move between viewing positions), and scanpaths (sequences of eye gaze
ifxations and interconnecting saccades in chronological order), which are detected in specific
spatial zones called areas of interest (AOIs) [
        <xref ref-type="bibr" rid="ref10 ref9">9, 10</xref>
        ].
      </p>
      <p>
        To work with the above metrics, visual analytics tools are demanded [
        <xref ref-type="bibr" rid="ref11 ref12">11, 12</xref>
        ]. But while there
are quite a lot of software tools for conducting research with traditional stationary eye trackers,
there is a lack of corresponding tools integrated with VR engines, which allow utilizing the eye
trackers built in the HMDs [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. In our previous work, we proposed a flexible pipeline to track
and analyze the eye gaze in a VR environment, utilizing high-level semantic data mining and
visual analytics platform SciVi [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. The present work is aimed to build new stages upon this
pipeline to allow handling of polycode stimuli in the eye-tracking-based research.
      </p>
      <p>Polycode stimulus is a compound visual object presented to the informant during the
eyetracking-based experiment. This object consists of several subobjects representing diferent
semiotic codes (entities of diferent semantics, for example, elements of diferent nature and/or
meaning). Typical examples of polycode stimuli are modern advertisements, which contain
images and short text blocks. More sophisticated cases are paintings, which typically can be
split up into many areas representing diferent topics, which are combined by the artist into a
solid work, reflecting his/her thoughts, emotions, and world outlook.</p>
      <p>Study of the polycode stimuli comprehension is important both in fundamental and applied
humanitarian sciences. Eye tracking metrics are a powerful way to obtain objective data for
this study, but a problem is in stimuli segmentation to define particular AOIs. In the general
case, such segmentation cannot be performed automatically due to the complexity of stimuli
structures. To solve this problem, user-friendly high-level segmentation tools are needed, which
allow experts to perform the segmentation manually or in an automated yet human-supervised
manner.</p>
      <p>
        In the present work, we propose reusing Creative Maps Studio vector editor [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] to provide
Digital Humanities (DH) research experts with the ergonomic tools for segmenting multimodal
stimuli into hierarchical AOIs, which can then be exported to the SciVi platform and used
in the VR-based eye tracking pipeline. To test the approach proposed, we studied two use
cases: the advertising poster and the painting “The Appearance of Christ Before the People”
by A. Ivanov [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Both use cases include the segmentation phase (performed within Creative
Maps Studio), the VR-based presentation phase (performed with help of Unreal Engine, with
eye tracks being detected and recorded using SciVi), and the visual analytics phase (performed
within SciVi).
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Key Contributions</title>
      <p>
        In this work, we propose an improvement for our previously created visual analytics pipeline
leveraging eye tracking experiments in the VR environment [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. The distinctive new features
proposed are the following:
      </p>
      <sec id="sec-2-1">
        <title>1. The polycode stimuli segmentation tool based on Creative Maps Studio [14]. 2. The corresponding plugins for the SciVi platform, which allow to automatically extract the hierarchical AOIs from the segmented polycode stimuli and visually analyze the corresponding scanpaths during the eye-tracking-based experiments.</title>
      </sec>
      <sec id="sec-2-2">
        <title>The above features allow us to study the perception of polycode stimuli in VR.</title>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Related Work</title>
      <p>
        Eye tracking is a well-documented methodology to study the information perception of humans
since it has been proven that eye movements reflect the comprehension process [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. K. Rayner
presented an elaborate review of the research related to the discovery of information processing
features through eye gaze tracking [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], describing diferent information processing contexts and
providing the most considerable metrics. K. Holmqvist et al. summarized the most popular and
practically significant methods and measures of eye-tracking-based research in a comprehensive
guide [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] that can leverage a quick start for novice scholars. One more concise yet informative
review of eye tracking metrics is provided by Z. Sharafi et al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Later, nearly the same authors
published a longer practical guide on eye tracking research [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. This guide is claimed to be
related to using eye tracking in software engineering, but the mentioned methods and means
can be used far beyond this area. The above guides indicate scanpaths as basic structures to
analyze for revealing the perception features. This is why scanpaths are the focus of our present
work.
      </p>
      <p>
        To handle the eye tracks detected, visual analytics tools are needed. T. Blascheck et al.
collected the state-of-the-art visualization techniques used in eye tracking [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. According to
this review and some other reports of the same author, a handy technique to visualize scanpaths
is a circular heat map [
        <xref ref-type="bibr" rid="ref16 ref17">16, 17</xref>
        ]. It is a graph-based representation, in which AOIs are drawn
as nodes in a circle, and saccades are depicted as arcs between them. Since we already used a
similar representation of linked data in our previous research [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], we decided to adopt this
visualization concept to solve our eye tracking tasks.
      </p>
      <p>
        The recent advances in microelectronics enabled the rapid evolution of mobile eye trackers,
which can now be integrated with VR HMDs. N. Stein et al. carried out some benchmarks among
the most popular VR HMDs with integrated eye trackers and stated that the characteristics of
these devices are fair enough to conduct eye-tracking-based research [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. This corresponds
to the fact that in the last few years, many VR-based research projects emerged involving eye
tracking. For example, D. Sonntag et al. designed an immersive trafic simulation to monitor
the elderly pedestrians’ spatial navigation skills [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. A. Skulmowski et al. used VR to simulate
the critical emotionally charged situation in which the informants encounter a moral choice
and social judgments [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. This kind of experiment would be nearly impossible in the real world,
while it is safe and easy in a VR environment.
      </p>
      <p>
        B. Farnsworth described the most popular software programs used by scholars for eye tracks
analysis [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. However, these programs are mostly designed for stationary eye tracking devices,
and there is a lack of software that would enable seamless integration with VR scenes observed
with eye-tracker-powered HMDs. This is why the above-mentioned research projects involved
ad-hoc software solutions to manage VR scenes and obtain needed eye tracking data instead of
using some unified solutions. To support the spreading of eye-tracking-based experiments in
VR environments, this gap has to be bridged. While there are some attempts to create a unified
software for eye tracking in VR [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], many problems are not yet solved.
      </p>
      <p>In our research, we use so-called polycode stimuli to conduct eye-tracking-based experiments.
A polycode stimulus is a visual object that consists of separable parts having diferent semantics
for the informant. These semantics are denoted as semiotic codes. The codes have their
representation types called the modalities. For example, there are image-based and text-based
codes. In our work we focus on unimodal (single-typed codes) and bimodal (two types of codes)
stimuli.</p>
      <p>
        As an example of unimodal stimulus, we consider a painting “The Appearance of Christ
Before the People” by A. Ivanov [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. One can extract from this painting such codes as diferent
social roles, clothing, and gestures of depicted persons, spatial separation of persons’ groups,
etc. The codes are used by the painting’s author to represent the dynamics of the painting’s
plot, so the viewer (in our case, the informant) has to decode these codes to understand the
story. The ability to decode the semiotic codes is tied to the cultural context. For example, one
can expect completely diferent interpretations of the story told in “The Appearance of Christ
Before the People” of informants familiar and unfamiliar with Christian culture and European
painting (the comparison of the eye gaze tracks of such informants’ sample is beyond of this
work’s scope but appears to be an interesting topic for the future Digital Humanities research).
Polycode nature of paintings is an interesting material to study processes of perception and
symbols’ decoding. Previously, “The Appearance of Christ Before the People” painting was
used to study certain aspects of semiotic codes decoding [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. Currently, VR environment and
eye gaze tracking capabilities can bring this study to a qualitatively new level, covering the
entire comprehension process, from perception up to interpretation.
      </p>
      <p>
        Bimodal stimulus in our research is represented by so-called creolized text: a visual object
containing both picture and text. We took an advertising poster that was previously studied
in [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] to see whether the VR environment introduces new features to the perception process.
      </p>
      <p>
        In our research, we rely on the results obtained in [
        <xref ref-type="bibr" rid="ref22 ref23 ref24">22, 23, 24</xref>
        ], but the distinctive feature of
the present work is a focus on the VR. Although VR provides promising possibilities for Digital
Humanities research, there are only a few works related to the comprehension processes in
the VR environment [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]. The possible reason for this is the absence of high-level software
tools to help the researchers to set up and conduct the related experiments. This is why we aim
to create a high-level user-friendly pipeline integrating SciVi semantic data mining platform,
Creative Maps Studio vector drawing editor, and Unreal Engine graphics rendering tools.
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Background</title>
      <p>
        The reported work is the next step in the development of the unified flexible pipeline for
studying information comprehension processes in a VR environment. Previously, we proposed
an extensible pipeline for composing virtual scenes, collecting the eye gaze tracks of the
informants in these scenes, and analyzing the collected data through the visualization of the
scanpaths [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
      <p>
        The core of the proposed pipeline is the ontology-driven client-server semantic data mining
platform SciVi. The SciVi server is written in Python using Flask, the SciVi thin client is
written in JavaScript using HTML5 and CSS3. This platform provides a high-level graphical
user interface to declare particular data mining pipelines in a form of data flow diagrams
(DFDs), which consists of nodes representing the data handling operators and links representing
data flow. Each operator has its own ontological description stored in the SciVi knowledge
base (using ONT format to represent ontologies and ONTOLIS editor [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] to create/modify
them). This description includes the operator’s typed inputs, outputs, and settings, as well
as a link to the operator’s implementation. Based on this description, SciVi automatically
generates the graphical user interface for tuning and chaining these operators, as well as
internal interoperability mechanisms to execute these operators on the available computing
nodes in the network and interchange data between them. This approach ensures the easy
extensibility and high configurability of the SciVi platform, allowing us to adapt SciVi for solving
diferent data mining tasks in diferent application domains.
      </p>
      <p>
        To adapt SciVi to the eye tracking tasks we created a corresponding set of operators described
in [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. These operators leverage VR scene tuning, eye gaze tracks collecting, and scanpaths
visualization. To render the VR scene, we use Unreal Engine, for which we created a small
WebSocket-based communication module (written in C++) to enable interoperability with SciVi
at runtime. To present the VR scene to the informant and track the informant’s eye gaze we
use the HTC Vive Pro Eye HMD, which is one of the most popular state-of-the-art VR headsets
with an integrated precise infrared eye tracker. To visually analyze the scanpaths, we adopt the
SciVi::CGraph module [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] to draw a circular graph of AOIs and interconnecting saccades.
      </p>
      <p>The focus of our previous work was on the study of the reading process in VR, so we
considered texts as stimuli in the eye gaze tracking experiments. The present work is devoted to
the generalization of the proposed approach enabling to study the comprehension of polycode
stimuli.</p>
      <p>
        For the polycode stimuli segmentation, we use the web-based vector graphic editor Creative
Maps Studio [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. This application was originally developed for creating and editing naive maps.
Its server side is written in Python and client side is written in JavaScript. React, FastAPI and
PostgreSQL are used as the main software Creative Maps Studio is based on. The tools provided
by the Creative Maps Studio are quite versatile and therefore can be used, inter alia, for the
segmentation of images by outlining diferent areas and storing their verbal descriptions.
      </p>
    </sec>
    <sec id="sec-5">
      <title>5. Polycode Stimuli Segmentation</title>
      <p>As mentioned above, automatic segmentation of polycode stimuli by marking individual AOIs
is nearly impossible, so the process has to be supervised by a human expert. In some cases,
semi-automatic methods can be used (for example, to extract areas with text and images).
But, for example, in the case of segmenting the paintings, there are almost no clues to build
a universal segmentation algorithm that could perform automatic AOIs outline. So, either
advanced machine learning algorithms or ergonomic manual segmentation tools are required
to fulfill this task in a general case. In the present work, we start with the manual segmentation
of the polycode stimuli, yet take a very complex stimulus for testing.</p>
      <p>Technically, the polycode stimulus is represented by a raster image. To enable its ergonomic
manual segmentation, we extended the functionality of the Creative Maps Studio web application
by adding the function of setting the raster background to the canvas (the area where the user
draws a vector image). Having a stimulus image as a background, the expert can outline
expected AOIs with the help of vector-based contours (possibly closed sequences of quadratic
Bézier curves and line segments). Contour drawing is a built-in function of Created Maps
Studio. In addition to the graphical representation, the contours can also have verbal metadata:
name, description, and associated emoticons. Verbal metadata helps to identify the AOIs in the
further analysis of scanpaths. Some of the objects in the image can have a complex structure.
An example of such an object is an image of a group of persons, where each person represents
some independent semantics for the viewer but is united with the other persons in a group by
some attribute. In this case, individual outlined regions can be joined into the named groups
representing hierarchical AOIs.</p>
      <p>Fig. 1 shows the segmentation of two images used as stimuli for eye-tracking-based
experiments within this work. Left one represents a unimodal stimulus made of “The Appearance
of Christ Before the People” painting. Right one represents a bimodal stimulus (creolized text)
made of the advertising poster.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Visual Analytics of Scanpaths</title>
      <sec id="sec-6-1">
        <title>6.1. Design of Experiment</title>
        <p>
          Once the segmentation of a stimulus is done, the actual experiment can be conducted to reveal
the specifics of comprehending this stimulus by diferent informants. The general experiment
setup is described in detail in [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ].
        </p>
        <p>Each VR-based experiment with eye gaze tracking involves three diferent software programs
running on diferent computers interconnected via the network. The main program is the SciVi
server that manages the entire experiment process. Being a Python-based software, it can run
on almost any computer. In our case, it runs on the desktop computer in the laboratory. The
second one is the SciVi Web client that provides a control panel for the experiment director. It
can run in any HTML5-capable browser. During the experiment, the director normally uses a
notebook to be mobile in the laboratory. The third software program is a VR application based
on the Unreal Engine. It runs on the desktop computer with NVidia Titan RTX GPU capable of
rendering the VR scene for HTC Vive Pro Eye HMD.</p>
        <p>First, the experiment director explains to the informant the anonymity of data being collected
as well as the safety regulations related to the immersion in VR. If agrees, the informant signs
the informed consent to participate in the experiment (having an option to decline and leave
any time during the experiment).</p>
        <p>The first step of the experiment is the eye tracker calibration leveraged by the standard
procedure provided by the HTC Vive Pro Eye driver. After that, the experiment director sets up
SciVi to show the test pattern to the informant and visually inspect the gaze detection data in
order to check if calibration results are valid and if the informant feels comfortable (the SciVi
setup is described in detail in Section 6.2). The test pattern consists of numbered circles of
diferent sizes and colors the informant is asked to look at in ascending order of numbers.</p>
        <p>Next, an actual stimulus is presented and the tracks of the informant’s eye gaze are collected
and stored (see Section 6.2 6.2). The rendering result of the stimulus image is presented in Fig. 2.</p>
        <p>At the moment, the last step of the experiment is the visualization and analysis of individual
scanpaths (see Section 6.3). But in the future, we plan to further extend the data processing
pipeline by adding the metrics considering eye gaze tracking results of multiple informants to
discover common patterns and search for the comprehension features of diferent informants’
groups.</p>
      </sec>
      <sec id="sec-6-2">
        <title>6.2. Recording of Eye Gaze Tracking Data</title>
        <p>The process of monitoring and recording of eye gaze tracking data for a particular stimulus is
declared in SciVi by the DFD shown in Fig. 3.</p>
        <p>This DFD is composed by the experiment director within the SciVi Web interface provided
by the SciVi thin client. The “Segmented Map” operator allows loading the stimulus image
segmented in the Creative Maps Studio. This operator has two outputs: the raster picture
corresponding to the stimulus and the array of AOIs encoded as named groups of named paths
(arrays of 2D points). The picture (encoded using the Base64 standard) is transmitted to the
“VR Board” operator that is responsible for communication with the VR application. This
communication is based on the WebSocket protocol. The VR application renders the scene and
shows the received image to the informant (see Fig. 2). The gaze of the informant’s eyes is
detected by the built-in eye tracker of the HTC Vive Pro Eye HMD. The detection results are
collected by the SRanipal SDK API that allows to get the scene object the informant is looking
at, as well as 3D coordinates (in the scene space) of the gaze origin and case focal point. If the
object is the plain displaying the stimulus image (see Fig. 2), the gaze focal point is converted
from the scene space to the texture space of this plain in order to get the coordinates related to
the stimulus image.</p>
        <p>The gaze-related data are transmitted back to SciVi via the WebSocket using the following
CSV-based encoding: timestamp, gazeU, gazeV, gazeOriginX, gazeOriginY, gazeOriginZ,
gazeFocusX, gazeFocusY, gazeFocusZ. While gazeU and gazeV are used to perform a hit test
with AOIs, other values are needed to detect saccades and fixations.</p>
        <p>The “Gaze Viewer” is a visualization operator dedicated to monitoring the gaze tracks being
collected. The rendering result of this operator is shown in Fig. 4. The blue circle represents
the gaze focal point. The white contour indicates the AOI the informant is looking at, and the
corresponding AOI name is displayed on the left (the names of groups the AOI belongs to are
prefixed with the slash symbol in hierarchical order).</p>
        <p>The “Gaze Recorder” operator allows recording gaze tracking data to a CSV file, while the
“CSV Table” operator is dedicated to loading these data back to SciVi, and the “Gaze Player”
operator allows playing them back. In Fig. 3 this operator is not connected with the others, but
it can, for example, be linked to the “Gaze Viewer” instead of “VR Board”, so the recorded data
can be viewed again. Using the described operators, the experiment director can easily set up
SciVi for particular eye-tracking-based experiments. It is worth noting that AOIs can be defined
(or redefined) after the data are recorded because the recorded files contain just the raw eye
gaze tracks over a particular stimulus.</p>
      </sec>
      <sec id="sec-6-3">
        <title>6.3. Scanpath Detection and Visualization</title>
        <p>
          Based on the recorded eye gaze tracking data, the scanpaths are calculated and visualized using
our modified version of a radial transition graph [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. The corresponding SciVi DFD is shown
in Fig. 5.
        </p>
        <p>
          The “Eye Movements Detector” operator determines whether the gaze tracking samples
belong to saccades or to fixations. This operator is based on the algorithm proposed by J.
LlanesJurado et al. [
          <xref ref-type="bibr" rid="ref27">27</xref>
          ] (we have ported the reference implementation of this algorithm to JavaScript to
be able to run it on the SciVi client side). The “Scanpath Builder” operator assembles the scanpath
by joining fixations and interchanging saccades, producing the data suitable for displaying with
SciVi::CGraph visualization module represented by the “Circular Graph” operator. The example
of the rendering result is shown in Fig. 6.
        </p>
        <p>
          The AOIs are displayed as graph nodes placed in a circle. The total fixation duration in each
AOI is shown as histogram box behind the name of this AOI, and the number of fixations is
mapped to its color (for this, the array of colors is produced by the “Scanpath Builder” operator
as one of the outputs). The saccades are represented by graph edges rendered as arrows with
the number of saccades mapped to the arrows’ thickness. The AOIs hierarchy is displayed with
a help of the multilevel ring scale drawn around the graph nodes [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ].
        </p>
        <p>
          Compared to our previous work [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ], AOIs hierarchy representation is a new distinctive
feature of scanpaths visual analytics proposed. Together with the filtering capabilities of
SciVi::CGraph adopted to the eye gaze tracking tasks in [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ], this new feature allows eficient
analysis of the scanpaths on the complex polycode stimuli.
        </p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusion</title>
      <p>In the present work, we extended our previously proposed pipeline for visual analytics of
eye gaze tracks collected in a VR environment. The extensions developed allow us to handle
polycode stimuli, enlarging the variety of experiments we can conduct using HTC Vive Pro
Eye HMD, Unreal Engine, Creative Maps Studio vector graphics editor, and SciVi semantic data
mining platform. The tools created were tested on two diferent polycode stimuli: unimodal
one (a painting containing a lot of diferent persons and environmental objects) and bimodal
one (an advertising poster containing picture and text). The preliminary results of eye gaze
tracking analysis were obtained for pipeline testing purposes, while the experiments involving
real informants will be a part of future work. After the tests were conducted, the viability of our
pipeline has been confirmed, and we can conclude that this pipeline possesses great potential
for Digital Humanities research.</p>
      <p>Next, we intend to implement more complex eye gaze tracking metrics as individual SciVi
platform plugins to discover the perception features of informants in a VR environment. Also,
we plan to study possible ways to automate multimodal stimuli segmentation process with a
help of machine learning techniques.</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgments</title>
      <p>This study is supported by the research grant No. ID75288744 from Saint Petersburg State
University.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>K.</given-names>
            <surname>Rayner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. H.</given-names>
            <surname>Chace</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. J.</given-names>
            <surname>Slattery</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ashby</surname>
          </string-name>
          , Eye Movements as Reflections of Comprehension Processes in Reading,
          <source>Scientific Studies of Reading</source>
          <volume>10</volume>
          (
          <year>2006</year>
          )
          <fpage>241</fpage>
          -
          <lpage>255</lpage>
          . doi:
          <volume>10</volume>
          .1207/s1532799xssr1003_
          <fpage>3</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>K.</given-names>
            <surname>Rayner</surname>
          </string-name>
          ,
          <source>Eye Movements in Reading and Information Processing: 20 Years of Research, Psychological Bulletin</source>
          <volume>124</volume>
          (
          <year>1998</year>
          )
          <fpage>372</fpage>
          -
          <lpage>422</lpage>
          . doi:
          <volume>10</volume>
          .1037/
          <fpage>0033</fpage>
          -
          <lpage>2909</lpage>
          .
          <year>124</year>
          .3.372.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D.</given-names>
            <surname>Sonntag</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Orlosky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Weber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Gu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sosnovsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Toyama</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. N.</given-names>
            <surname>Toosi</surname>
          </string-name>
          ,
          <article-title>Cognitive Monitoring via Eye Tracking in Virtual Reality Pedestrian Environments</article-title>
          ,
          <source>in: Proceedings of the 4th International Symposium on Pervasive Displays</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>269</fpage>
          -
          <lpage>270</lpage>
          . doi:
          <volume>10</volume>
          . 1145/2757710.2776816.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>K.</given-names>
            <surname>Holmqvist</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Nyström</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Andersson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Dewhurst</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Jarodzka</surname>
          </string-name>
          , J. van de Weijer, Eye Tracking:
          <string-name>
            <given-names>A Comprehensive</given-names>
            <surname>Guide To Methods</surname>
          </string-name>
          And Measures, OUP Oxford,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Tobii</surname>
          </string-name>
          ,
          <year>2021</year>
          . URL: https://www.tobii.com,
          <source>last accessed 8 Jul</source>
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>B.</given-names>
            <surname>Farnsworth</surname>
          </string-name>
          , 10
          <source>Free Eye Tracking Software Programs [Pros and Cons]</source>
          ,
          <year>2021</year>
          . URL: https://imotions.com/blog/free-eye
          <string-name>
            <surname>-</surname>
          </string-name>
          tracking-software/,
          <source>last accessed 8 Jul</source>
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>N.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. C.</given-names>
            <surname>Niehorster</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Watson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Steinicke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Rifai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Wahl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lappe</surname>
          </string-name>
          ,
          <article-title>A Comparison of Eye Tracking Latencies Among Several Commercial Head-Mounted Displays, i-Perception 12 (</article-title>
          <year>2021</year>
          )
          <fpage>1</fpage>
          -
          <lpage>16</lpage>
          . doi:
          <volume>10</volume>
          .1177/2041669520983338.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Skulmowski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bunge</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kaspar</surname>
          </string-name>
          , G. Pipa,
          <article-title>Forced-Choice Decision-Making in Modified Trolley Dilemma Situations: a Virtual Reality and Eye Tracking Study</article-title>
          ,
          <source>Frontiers in Behavioral Neuroscience</source>
          <volume>8</volume>
          (
          <year>2014</year>
          ). doi:
          <volume>10</volume>
          .3389/fnbeh.
          <year>2014</year>
          .
          <volume>00426</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Poole</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. J.</given-names>
            <surname>Ball</surname>
          </string-name>
          , Eye Tracking in HCI and Usability Research, Encyclopedia of Human Computer Interaction (
          <year>2006</year>
          )
          <fpage>211</fpage>
          -
          <lpage>219</lpage>
          . doi:
          <volume>10</volume>
          .4018/978-1-
          <fpage>59140</fpage>
          -562-7.
          <year>ch034</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Sharafi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Shafer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Sharif</surname>
          </string-name>
          , Y.-G. Guéhéneuc,
          <article-title>Eye-Tracking Metrics in Software Engineering</article-title>
          , in: 2015
          <source>Asia-Pacific Software Engineering Conference (APSEC)</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>96</fpage>
          -
          <lpage>103</lpage>
          . doi:
          <volume>10</volume>
          .1109/APSEC.
          <year>2015</year>
          .
          <volume>53</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Sharafi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Sharif</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.-G.</given-names>
            <surname>Guéhéneuc</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Begel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Bednarik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Crosby</surname>
          </string-name>
          ,
          <string-name>
            <surname>A Practical</surname>
          </string-name>
          <article-title>Guide on Conducting Eye Tracking Studies in Software Engineering</article-title>
          ,
          <source>Empirical Software Engineering</source>
          <volume>25</volume>
          (
          <year>2020</year>
          )
          <fpage>3128</fpage>
          -
          <lpage>3174</lpage>
          . doi:
          <volume>10</volume>
          .1007/s10664-020-09829-4.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>T.</given-names>
            <surname>Blascheck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kurzhals</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Raschke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Burch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Weiskopf</surname>
          </string-name>
          , T. Ertl,
          <article-title>State-of-the-Art of Visualization for Eye Tracking Data</article-title>
          , in: EuroVis - STARs,
          <year>2014</year>
          , pp.
          <fpage>63</fpage>
          -
          <lpage>82</lpage>
          . doi:
          <volume>10</volume>
          .2312/ eurovisstar.20141173.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>K.</given-names>
            <surname>Ryabinin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Belousov</surname>
          </string-name>
          ,
          <article-title>Visual Analytics of Gaze Tracks in Virtual Reality Environment</article-title>
          ,
          <source>Scientific Visualization</source>
          <volume>13</volume>
          (
          <year>2021</year>
          )
          <fpage>50</fpage>
          -
          <lpage>66</lpage>
          . doi:
          <volume>10</volume>
          .26583/sv.13.2.04.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>R.</given-names>
            <surname>Chumakov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ryabinin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Belousov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Duan</surname>
          </string-name>
          , Creative Map Studio:
          <article-title>A Platform for Visual Analytics of Mental Maps</article-title>
          ,
          <source>Scientific Visualization</source>
          <volume>13</volume>
          (
          <year>2021</year>
          )
          <fpage>79</fpage>
          -
          <lpage>93</lpage>
          . doi:
          <volume>10</volume>
          .26583/ sv.13.2.06.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Ivanov</surname>
          </string-name>
          ,
          <source>The Appearance of Christ Before the People</source>
          ,
          <year>1857</year>
          . URL: https: //en.wikipedia.org/wiki/The_Appearance_of_Christ_Before_the_People#/media/File: Александр_Андреевич_Иванов_-_
          <string-name>
            <surname>Явление</surname>
          </string-name>
          _Христа_народу_(Явление_Мессии)_-_
          <string-name>
            <surname>Google</surname>
          </string-name>
          _Art_Project.jpg.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>T.</given-names>
            <surname>Blascheck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Raschke</surname>
          </string-name>
          , T. Ertl, Circular Heat Map Transition Diagram,
          <source>in: Proceedings of the 2013 Conference on Eye Tracking South Africa</source>
          ,
          <year>2013</year>
          , pp.
          <fpage>58</fpage>
          -
          <lpage>61</lpage>
          . doi:
          <volume>10</volume>
          .1145/ 2509315.2509326.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>T.</given-names>
            <surname>Blascheck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Sharif</surname>
          </string-name>
          ,
          <source>Visually Analyzing Eye Movements on Natural Language Texts and Source Code Snippets, in: Proceedings of the 11th ACM Symposium on Eye Tracking Research &amp; Applications</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          . doi:
          <volume>10</volume>
          .1145/3314111.3319917.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>K.</given-names>
            <surname>Ryabinin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Belousov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Chuprina</surname>
          </string-name>
          ,
          <article-title>Novel Circular Graph Capabilities for Comprehensive Visual Analytics of Interconnected Data in Digital Humanities</article-title>
          ,
          <source>Scientific Visualization</source>
          <volume>12</volume>
          (
          <year>2020</year>
          )
          <fpage>56</fpage>
          -
          <lpage>70</lpage>
          . doi:
          <volume>10</volume>
          .26583/sv.12.4.06.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>J. C. E.</given-names>
            <surname>Iacobi</surname>
          </string-name>
          ,
          <article-title>Software for Analyzing User Experiences in Virtual Reality using Eye Tracking</article-title>
          ,
          <year>2018</year>
          . URL: https://kth.diva-portal.org/smash/get/diva2:1231972/FULLTEXT01. pdf,
          <source>last accessed 14 Jul</source>
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>K. I.</given-names>
            <surname>Belousov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. A.</given-names>
            <surname>Podtikhova</surname>
          </string-name>
          ,
          <source>Reflections on Conceptual Composition of Text, Voprosy Kognitivnoy Lingvistiki</source>
          <volume>1</volume>
          (
          <year>2014</year>
          )
          <fpage>62</fpage>
          -
          <lpage>74</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>A.</given-names>
            <surname>Konovalova</surname>
          </string-name>
          , T. Petrova,
          <article-title>Lexical Ambiguity in Slogans: Does It Make a Polycode Text Easier to Recognize?</article-title>
          ,
          <source>in: Proceedings of the 4th International Conference on Neurobiology of Speech and Language</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>54</fpage>
          -
          <lpage>55</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>T.</given-names>
            <surname>Petrova</surname>
          </string-name>
          , E. Riekhakaynen,
          <article-title>Eye-tracking study of processing Russian polycode texts</article-title>
          ,
          <source>in: The 20th European Conference on Eye Movements, ECEM</source>
          <year>2019</year>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>T.</given-names>
            <surname>Petrova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Kalugina</surname>
          </string-name>
          , Eye-Tracking
          <source>Study of Reading Polycode Texts: Evidence From Russian, Perception</source>
          <volume>48</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>T. E.</given-names>
            <surname>Petrova</surname>
          </string-name>
          ,
          <string-name>
            <surname>E. I. Riekhakaynen</surname>
          </string-name>
          ,
          <article-title>Processing of Verbal and Non-verbal Patterns: An Eye-Tracking Study of Russian</article-title>
          , in: Third
          <source>International Congress on Information and Communication Technology. Advances in Intelligent Systems and Computing</source>
          , volume
          <volume>797</volume>
          ,
          <year>2019</year>
          , pp.
          <fpage>269</fpage>
          -
          <lpage>276</lpage>
          . doi:
          <volume>10</volume>
          .1007/
          <fpage>978</fpage>
          -981-13-1165-9_
          <fpage>24</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>J.</given-names>
            <surname>Mirault</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Guerre-Genton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dufau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Grainger</surname>
          </string-name>
          ,
          <article-title>Using Virtual Reality to Study Reading: An Eye-Tracking Investigation of Transposed-Word Efects, Methods in Psychology 3 (</article-title>
          <year>2020</year>
          ). doi:
          <volume>10</volume>
          .1016/j.metip.
          <year>2020</year>
          .
          <volume>100029</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>S.</given-names>
            <surname>Chuprina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Nasraoui</surname>
          </string-name>
          ,
          <article-title>Using Ontology-based Adaptable Scientific Visualization and Cognitive Graphics Tools to Transform Traditional Information Systems into Intelligent Systems, Scientific Visualization 8 (</article-title>
          <year>2016</year>
          )
          <fpage>23</fpage>
          -
          <lpage>44</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>J.</given-names>
            <surname>Llanes-Jurado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Marín-Morales</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Guixeres</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Alcañiz</surname>
          </string-name>
          ,
          <article-title>Development and Calibration of an Eye-Tracking Fixation Identification Algorithm for Immersive Virtual Reality</article-title>
          ,
          <source>Sensors</source>
          <volume>20</volume>
          (
          <year>2020</year>
          ). doi:
          <volume>10</volume>
          .3390/s20174956.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>