<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Breaking the Mona Lisa Efect: Enhancing Eye Contact with Virtual Humans in 2D Display Environments⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sunghun Jung</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Junyeong Kum</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Myungho Lee</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Pusan National University</institution>
          ,
          <country>Republic of Korea</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>With the rise of large language models, virtual humans (VHs) are becoming increasingly prevalent as conversational service agents across various domains. Consequently, kiosks equipped with large 2D displays have become the primary platform for deploying VH systems. However, VHs presented on 2D displays introduce the Mona Lisa efect, which hinders precise eye contact interaction with users. In this paper, we propose a straightforward method to alleviate the ambiguity in VHs' eye gaze direction. To evaluate the efectiveness of our method, we conducted an experiment involving 30 participants. The results revealed a statistically significant improvement in gaze perception accuracy within 2D VH systems when employing our approach. This improvement has the potential to enhance user engagement and the overall quality of human-virtual human interactions.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;virtual human</kwd>
        <kwd>eye contact</kwd>
        <kwd>reduced Mona Lisa efect</kwd>
        <kwd>visual cue</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        while non-verbal communication cues like gestures, eye
contact, and facial expressions are also crucial [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ].
      </p>
      <p>
        Recently, virtual humans (VHs) have been utilized in Among these non-verbal cues, eye contact plays a vital
various fields, including education, counseling, and en- role in expressing interest, concentration, facilitating
diatertainment. This is partially attributed to VH modeling logue in multiparty conversations, and enabling smooth
tools1 that have simplified the creation of VHs resembling turn-taking [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
real people, and the availability of large-scale language However, the representation of 3D VHs on a 2D
dismodels such as GPT-3 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], enabling natural dialogue. Fur- play gives rise to various optical illusions, including the
thermore, advancements in computer vision AI technol- Mona Lisa efect, which poses challenges for establishing
ogy have facilitated the distinction of individuals, user eye contact during communication [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ]. The Mona Lisa
tracking, and analysis of emotions and other essential efect occurs when a user positioned within 5 degrees to
interaction-related information [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. the left or right of the VH on the 2D display perceives the
      </p>
      <p>
        Ongoing studies are utilizing these technologies to in- VH as making eye contact [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Unlike real-life
converteract with VHs [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ], and the use of VH kiosk systems sations occurring in a 3D space, where individuals can
is increasing in various domains, such as banks, govern- accurately perceive each other’s gaze even in confined
ment ofices, and museums. While VH research projects areas, the presence of the Mona Lisa efect can disrupt
typically incorporate specialized devices for the 3D repre- the accurate perception of eye contact between users and
sentation of VHs, those service kiosks often solely employ VHs in 2D kiosk systems, particularly when users are in
a 2D display (see NVIDIA Omniverse Avatar 2), poten- a confined space [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. As a result, rendering VHs in such
tially imposing an issue of delivering non-verbal commu- a 2D kiosk system has the potential to reduce user
ennication cues requiring 3D perception. gagement in conversations and overall satisfaction with
      </p>
      <p>Previous research has demonstrated that people have the VH systems.
a tendency to expect human-like interactions from VHs. In this paper, we propose a technique to mitigate the
In such interactions, verbal expressions hold significance, Mona Lisa efect in 2D display environments without
the need for specialized equipment such as a parallax
APMAR’23: The 15th Asia-Pacific Workshop on Mixed and Augmented barrier or stereoscopic glasses. Our method involves
Reality, Aug. 18-19, 2023, Taipei, Taiwan subtly rotating the virtual camera, which significantly
* Corresponding author. improves the perception of the gaze direction of a VH
($J. Ksuunmg)h; umny@upnughsaon.l.eaec@.krpn(Su..eJudnug()M;j.uLneeeg)old12@pusan.ac.kr displayed on a 2D screen. The rest of the paper discusses
0009-0007-0987-7190 (S. Jung); 0000-0003-3943-3463 (J. Kum); related work, details our proposed method, and presents
0000-0002-9421-8566 (M. Lee) an experiment conducted to validate the efectiveness of
© 2023 Copyright for this paper by its authors. Use permitted under Creative Commons License our approach.
1 hCPWrEooUrctkReshtdoinpgpssIhStpN:/c1e:6u1r3-w/-0s.o7r3g/wACwttErwibUu.trRioenaW4l.0louInrstekironsanhti.oocnpoal m(PCCr/ocBhYce4a.0er).adcintegrs-c(CreEaUtoRr-/WS.org)
2https://www.nvidia.com/en-us/on-demand/session/gtcfall21d31017/</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Works</title>
      <sec id="sec-2-1">
        <title>While we are unaware of previous eforts directly ad</title>
        <p>dressing the mitigation of the Mona Lisa efect, some
relevant areas of work are summarized in this section.</p>
        <p>
          The Mona Lisa efect is a phenomenon in which
individuals perceive that the gaze of a static image, depicting
a person or animal, is following them. This efect is
particularly noticeable within an angle of 5 degrees in both the
left and right directions [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. Moubayed and Beskow [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]
conducted a study to investigate the impact of the Mona
Lisa efect on gaze perception. They compared a face
projected onto a 3D object with a face presented in a 2D
format. Interestingly, while both conditions involved eye
movements that followed the users, the results revealed
a significant decrease in accuracy in the 2D condition.
        </p>
        <p>
          In response to the Mona Lisa efect, researchers have
proposed various methods to enhance eye contact inter- Figure 1: The top view in the Unity 3D virtual environment
action with VHs in 2D display environments. Some of in the 5-degree condition with our proposed method, when
these approaches involve modifying devices and applying VH looking at the left participant
content-based techniques to achieve a more natural eye
contact experience. One such method includes the
addition of an assistive device that rotates the display to align real-world position. However, the position of their faces
with the user’s position, enabling a more natural and im- can vary depending on their sitting posture or height,
mersive eye contact interaction [
          <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
          ]. The use of dual even though we controlled the position of the
particidisplay also has been proposed as a means to mitigate pants’ seats during the experiment. To reduce the
varithe Mona Lisa efect [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. They employed two displays, ability due to these diferences, we limited the camera
one depicting only the eyes and the other featuring the rotation value to a maximum of 1 degree left and right,
VH, to demonstrate that perceiving the character from a respectively. Therefore, the yaw angle of the camera
indiferent angle reduced the Mona Lisa efect. However, creased up to 1 degree, following the participant’s face
these methods have the drawback of requiring special- position. Beyond that, the yaw angle remains at 1 degree
ized equipment or additional displays. Another approach while VH’s gaze continues to follow the participant. It
involves leveraging visual efects within the content to should be noted that the camera position was fixed.
create a sense of natural eye contact. A particular study Figure 1 shows the top view in the Unity 3D virtual
showcased how VHs can establish eye contact with users environment when the VH looks at the left participant.
in passing or standing scenarios by employing head ro- To clearly demonstrate the rotation of the main camera,
tation and changes in pupil position [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. While these we adjusted the camera angle to 5 degrees.
studies focused on the naturalness of eye contact rather
than reducing the Mona Lisa efect, in this paper, we
focused on ways to reduce the Mona Lisa efect using 4. Experiment
visual cues.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Proposed Method</title>
      <p>
        Our method is based on Fish Tank VR, proposed by Ware
et al. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. In Fish Tank VR, the position and rotation
of the viewpoint are continuously updated to match the
user’s head position, so that the stereo image in the 2D
display appears to be 3D. In our proposed method, we
provide a visual cue by simply controlling the rotation
of the main camera of Unity Engine, which displays the
3D virtual environment.
      </p>
      <p>
        We employed YOLO8 [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] with RealSense Depth
Camera D455 to recognize participants’ faces. The estimated
position and depth data were used to calculate the user’s
This experiment examines whether our proposed method
can improve participants’ gaze recognition accuracy in a
narrow range where the Mona Lisa efect can occur. In
other words, when a VH on a 2D display looks at one
of two participants located within a 5-degree, we check
whether the participants can recognize whom a VH is
looking at.
      </p>
      <sec id="sec-3-1">
        <title>4.1. Method</title>
        <p>We used a within-subjects design to examine the
efectiveness of our proposed method. The experiment was
conducted at 5 and 15-degree angles, with all participants
ifrst experiencing the 15-degree condition. In both
condi</p>
      </sec>
      <sec id="sec-3-2">
        <title>4.2. Environment</title>
        <p>(a)
(b)
(c)
(b)
Figure 4: (a) The experimental setup and the schematics of
the experimental space and (b) actual experimental space
tions, the virtual human (VH) appropriately rotated her
spine, head, and eyes to match the position of the target
it was looking at (Figure 2 and Figure 3).</p>
        <p>The 15-degree condition was conducted to verify that
the participants understood the experiment by assessing
the VH’s gaze outside the range where the Mona Lisa
efect occurs. There were no other visual cues except
for the VH’s gesture. Then, the 5-degree condition was
conducted to examine our proposed method in the
narrow range where the Mona Lisa efect could occur. In
both conditions, the VH looked at the left participant,
the middle between both, or the right participant. A set
of 30 gaze judgment tasks consisting of ten of each gaze
direction was used for the 15-degree condition, while in
the 5-degree condition, we used a total of 60 gaze
judgment tasks. In the latter condition, our proposed method
was applied for half the time (see Figure 3(a)), and the
other half was not (Figure 3(b)) The order of VH’s gaze
directions was counterbalanced and randomized.
Additionally, five VHs were used in the experiment, and the
order was also randomized. For the eye gaze of the VH,
we exploited the Wizard of Oz paradigm, where an
experimenter in another space controlled the position of
the target for the VH to look at using a GUI.</p>
        <p>We organized our experimental environment with a
65-inch TV, two tables, and two chairs (Fig 4(b)). The
distance between the 2D display and the center of the two
participants was 2.4 meters. Additionally, the participants
sat at the same distance from side to side based on the VH;
0.64 meters in the 15-degree condition and 0.21 meters
in the 5-degree condition (Fig 4(a)). A desk was placed
in front of the participants so that they could fill out a
questionnaire with their smartphones.</p>
        <p>We used the Unity game engine 3 to render the VH
and an ofice-like environment on the 2D display. For
the VHs, we used rigged 3D human models created by
Character Creator 3 4. To eliminate the gender efect of
the VH, we made all five VHs female. We modified the
appearance of the VHs to have diferent facial features,
hairstyles, and clothes. The VH was placed behind a
table with a monitor and other ofice supplies. In the real
experimental space, we placed an actual desk in front
of the 2D display. We used Final IK 5 for natural gaze
behavior.</p>
        <sec id="sec-3-2-1">
          <title>3https://unity.com</title>
          <p>4https://www.reallusion.com/character-creator/
5https://root-motion.com/
4.3. Participants We found statistically significant results by using the
one-way ANOVA. First, we compared the 15-degree
conDue to the experimental configuration, we recruited two dition and the 5-degree condition (Figure 5(a)). The
reparticipants at a time. In cases where only one participant sults showed that participants were more accurate at
was available, one of the experimenters pretended to guessing where the VH was looking in the 15-degree
be the participant and joined the experiment. The data condition, outside the Mona Lisa efect (  &lt; 0.000).
submitted by the experimenter were excluded from the Additionally, to assess the efectiveness of our
proanalysis. We recruited 30 participants (13 males and 17 posed method, we used a dataset from the 5-degree
confemales) from a local university. The average age of the dition for analysis (Figure 5(b)). In this result, we only
participants was 22.4 (SD=2.10). Before the experiment, utilized the dataset when the VH looked at the left or right
an experimenter tested the participants’ dominant eye participant since there was no camera rotation when the
through a simple test (right eye 18, left eye 12). VH looked at the center. The participants were more
accurate at guessing where the VH was looking in the
4.4. Questionnaire narrow range when our proposed method was applied
The pre-questionnaire included questions about demo- ( = 0.033).
graphics and the dominant eye. During the experiment,
the participants selected whether they thought the VH 6. Discussion&amp;Conclusion
was looking at them, the center, or the other participant.</p>
          <p>We counted the correct responses.</p>
        </sec>
        <sec id="sec-3-2-2">
          <title>In this study, we found that the inclusion of supplemen</title>
          <p>tary visual cues efectively diminished the Mona Lisa
4.5. Procedure efect. Not only did participants recognize eye contact as
a distinguishing feature of the VH, but the subtle camera
After the participants completed the pre-questionnaire, rotation also introduced subtle disparities between the
the experimenter briefly explained the experiment. The desk and chair-like background. Building upon these
participants were then seated in the appropriate seats findings, we anticipate that incorporating multiple visual
based on the conditions, and the experiment began. cues will enhance the likelihood of accurately perceiving</p>
          <p>After a 3-second black screen, the VH was rendered on eye contact.
the display, looking either at the left participant, the cen- For instance, displaying the complete body of the VH
ter, or the right participant, depending on the condition. can provide insight into the character’s body orientation,
After both participants answered where they thought the while incorporating lines or patterns into the background
VH was looking, the experimenter pressed the next but- can aid in establishing eye contact. These additional cues
ton on the GUI to activate the black screen. The direction contribute to a more comprehensive visual context that
the VH looked at was indicated on the GUI in a prede- facilitates the perception of eye contact.
termined order. While the black screen was displayed, However, a notable limitation of this experiment is
the experimenter clicked on the location of the target for its focus on two user cases centered around a specific
the VH to look at in order. The VH and the main camera basis, casting doubt on whether the significant efect
rotation were controlled in Unity automatically. For each observed would be consistent with a larger number of
condition, we repeated this process 30 times. users or under non-central conditions. Consequently,
further investigation is needed to determine whether these
5. Result limitations can be overcome and whether this method
remains efective for more than three users.</p>
          <p>Additionally, our experiments solely considered gaze.</p>
          <p>Further tests are required to ascertain whether these
minute diferences in eye contact can be perceived during
a dialogue with a Virtual Human (VH). To enable these
experiments, we need to equip the VH with the ability to
recognize user speech and maintain eye contact with the
correct user during conversation. Our objective is to test
whether this enhancement enables users to detect these
subtle diferences.</p>
          <p>(a) (b)
Figure 5: Each graph shows the score in (a) 5-degree and
15-degree conditions and in (b) camera rotation on / of in
5-degree condition.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgments</title>
      <sec id="sec-4-1">
        <title>This research is supported Year 2021 Culture Technology</title>
      </sec>
      <sec id="sec-4-2">
        <title>R&amp;D Program by Ministry of Culture, Sports and Tourism and Korea Creative Content Agency(Project Number: R2021040269)</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>T.</given-names>
            <surname>Brown</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Mann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ryder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Subbiah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. D.</given-names>
            <surname>Kaplan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dhariwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Neelakantan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Shyam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sastry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Askell</surname>
          </string-name>
          , et al.,
          <article-title>Language models are few-shot learners</article-title>
          ,
          <source>Advances in neural information processing systems</source>
          <volume>33</volume>
          (
          <year>2020</year>
          )
          <fpage>1877</fpage>
          -
          <lpage>1901</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D.</given-names>
            <surname>Bohus</surname>
          </string-name>
          , E. Horvitz,
          <article-title>Facilitating multiparty dialog with gaze, gesture, and speech</article-title>
          ,
          <source>in: International Conference on Multimodal Interfaces and the Workshop on Machine Learning for Multimodal Interaction</source>
          ,
          <year>2010</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>D. DeVault</surname>
          </string-name>
          , R. Artstein, G. Benn,
          <string-name>
            <given-names>T.</given-names>
            <surname>Dey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Fast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gainer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Georgila</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gratch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hartholt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lhommet</surname>
          </string-name>
          , et al.,
          <article-title>Simsensei kiosk: A virtual human interviewer for healthcare decision support, in: Proceedings of the 2014 international conference on Autonomous agents and multi-agent systems</article-title>
          ,
          <year>2014</year>
          , pp.
          <fpage>1061</fpage>
          -
          <lpage>1068</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Cassell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Bickmore</surname>
          </string-name>
          , L. Campbell,
          <string-name>
            <given-names>H.</given-names>
            <surname>Vilhjalmsson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <article-title>More than just a pretty face: conversational protocols and the afordances of embodiment, Knowledge-based systems 14 (</article-title>
          <year>2001</year>
          )
          <fpage>55</fpage>
          -
          <lpage>64</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Cassell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. R.</given-names>
            <surname>Thorisson</surname>
          </string-name>
          ,
          <article-title>The power of a nod and a glance: Envelope vs. emotional feedback in animated conversational agents</article-title>
          ,
          <source>Applied Artificial Intelligence</source>
          <volume>13</volume>
          (
          <year>1999</year>
          )
          <fpage>519</fpage>
          -
          <lpage>538</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>D.</given-names>
            <surname>Aneja</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Hoegen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>McDuf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Czerwinski</surname>
          </string-name>
          ,
          <article-title>Understanding conversational and expressive style in a multimodal embodied conversational agent</article-title>
          ,
          <source>in: Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>C.</given-names>
            <surname>Oertel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Włodarczak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Edlund</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Wagner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gustafson</surname>
          </string-name>
          ,
          <article-title>Gaze patterns in turn-taking, in: Thirteenth annual conference of the international speech communication association</article-title>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Moubayed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Edlund</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Beskow</surname>
          </string-name>
          ,
          <article-title>Taming mona lisa: communicating gaze faithfully in 2d and 3d facial projections</article-title>
          ,
          <source>ACM Transactions on Interactive Intelligent Systems (TiiS) 1</source>
          (
          <issue>2012</issue>
          )
          <fpage>1</fpage>
          -
          <lpage>25</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>E.</given-names>
            <surname>Boyarskaya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sebastian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Bauermann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Hecht</surname>
          </string-name>
          ,
          <string-name>
            <surname>O. Tüscher,</surname>
          </string-name>
          <article-title>The mona lisa efect: Neural correlates of centered and of-centered gaze</article-title>
          ,
          <source>Human brain mapping 36</source>
          (
          <year>2015</year>
          )
          <fpage>619</fpage>
          -
          <lpage>632</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>K.</given-names>
            <surname>Otsuka</surname>
          </string-name>
          , Mmspace:
          <article-title>Kinetically-augmented telepresence for small group-to-group conversations</article-title>
          .
          <source>in 2016 ieee virtual reality (vr)</source>
          ,
          <source>online)</source>
          ,
          <source>DOI</source>
          <volume>10</volume>
          (
          <year>2016</year>
          )
          <fpage>19</fpage>
          -
          <lpage>28</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>M.</given-names>
            <surname>Vázquez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Milkessa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. M.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Govil</surname>
          </string-name>
          ,
          <article-title>Gaze by semi-virtual robotic heads: Efects of eye and head motion</article-title>
          ,
          <source>in: 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)</source>
          , IEEE,
          <year>2020</year>
          , pp.
          <fpage>11065</fpage>
          -
          <lpage>11071</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>H.</given-names>
            <surname>Mitake</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Ichii</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Tateishi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hasegawa</surname>
          </string-name>
          ,
          <article-title>Wide viewing angle fine planar image display without the mona lisa efect (</article-title>
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>H.-H. Wu</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Mitake</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Hasegawa</surname>
          </string-name>
          ,
          <article-title>Eye-gaze control of virtual agents compensating mona lisa efect</article-title>
          ,
          <source>HAI</source>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>C.</given-names>
            <surname>Ware</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Arthur</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. S.</given-names>
            <surname>Booth</surname>
          </string-name>
          ,
          <article-title>Fish tank virtual reality</article-title>
          ,
          <source>in: Proceedings of the INTERACT'93 and CHI'93 conference on Human factors in computing systems</source>
          ,
          <year>1993</year>
          , pp.
          <fpage>37</fpage>
          -
          <lpage>42</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>G.</given-names>
            <surname>Jocher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Chaurasia</surname>
          </string-name>
          , J. Qiu, YOLO by Ultralytics,
          <year>2023</year>
          . URL: https://github.com/ultralytics/ ultralytics.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>