<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Selfie Guidance System in Good Head Postures</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Naihui Fang</string-name>
          <email>fangnaihui88@hotmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Haoran Xie</string-name>
          <email>xiehr@acm.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Takeo Igarashi</string-name>
          <email>takeo@acm.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dept. of Computer Science, The University of Tokyo</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Taking selfies has become a popular and pervasive activity on smart mobile devices nowadays. However, it is still difficult for the average user to take a good selfie, which is a time-consuming and tedious task on normal mobile devices, especially for those who are not good at selfies. In order to reduce the difficulty of taking good selfies, this work proposes an interactive selfie application developed with multiple user interfaces to improve user satisfaction when taking selfies. Our proposed system helps average users take selfies by providing visual and voice guidance interfaces on the proper head postures to achieve good selfies. Preprocessing through crowdsourcing-based learning is utilized to evaluate the score space of possible head postures from hundreds of virtual selfies. For the interactive application, we adopt a geometric approach to estimate the current head posture of users. Our user studies show that the proposed selfie user interface can help common users taking good selfies and improve user satisfaction.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>INTRODUCTION
The selfie is a type of self-portrait that usually consists of a
photograph taken by mobile camera at an arm’s length. It is
one of the most direct ways to show oneself to others as well as
to record one’s daily life for one’s own benefit. Selfies became
more and more popular with the development of social media.
Now, it plays an important role in our daily life. Along with
its popularity, how to take a good selfie has become an urgent
issue to be solved. People want to look good on social media
but many of them encounter trouble when taking selfies. They
often spend long time and take a bunch of selfies trying to get
one good selfie that makes them look attractive and different
from others. In particular, it is difficult to find a good head
posture for a selfie. In addition, it’s not easy to have the same
head posture for the next one.
©2018. Copyright for the individual papers remains with the authors.
Copying permitted for private and academic purposes.</p>
      <p>SymCollab ’18, March 11, 2018, Tokyo, Japan</p>
      <p>
        With the development of computer-vision technology,
especially face recognition and facial-feature extraction
approaches, various selfie applications have come out to help the
average user take better selfies on mobile platforms. For
example, photo corrections and enhancements from a selfie can
be provided in a post-process way to specific area of the faces
[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Recently researchers proposed an approach to enhance
the selfies by suggesting expression [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] and face geometries
[
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] of the users in real-time. However, a user interface to
guide common users to take good selfies is still absent in all
these previous works.
      </p>
      <p>
        We propose an interactive selfie user interface to give clear
suggestions for users to take selfies based on crowdsourcing
results on mobile platforms. There are many factors in taking
a good selfie, such as the light condition, background, user
expressions, and so on. In contrast to a recent selfie system
considering face position and light condition [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], we chose to
study head posture (orientation), which is the most important
and controllable factor for selfie-taking in this work. Note that
the other factors share commonalities in the design of user
interfaces and can therefore benefit from this work.
In this work, we first conducted crowdsourcing tasks to define
good head postures. In order to avoid factors other than head
posture, we generated many virtual selfies using 3D human
models to receive the selfie score from crowd workers. The
virtual selfie receives higher scores if the users considered it
as more attractive. Based on the crowdsourcing results, we
developed a real-time user interface to be able to extract the
facial features from each frame and estimate the head postures
of the users in the real world. A geometry model approach [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]
was adopted to implement the real-time head-pose estimation
with the facial feature information extracted from each frame.
The system can calculate the closest candidates to good head
postures from crowdsourcing results, so that the system can
instruct a user to achieve good selfie under the guidance of a
reference human head. This work is expected to reduce the
difficulty of taking good selfies, thus improving user satisfaction
while taking selfies.
      </p>
      <p>
        The main contributions of this work include: We proposed
a novel approach to support taking selfies by providing
realtime suggestions on head postures; we collected data of head
postures in selfie-taking from crowdsourcing results obtained
by using hundreds of virtual selfies from 3D models. We
implemented multiple ways of interaction between the user
and the proposed interfaces as illustrated in Figure 1.
RELATED WORK
Selfie Systems. Creamcam [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] and YouCam [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] to the facial
features or add makeup effect onto the human face based on
facial feature extraction algorithms. Smile shutter function of
Sony DSC T300 provides the support on facial expressions.
An interactive selfie application was proposed to suggest face
size, face position and lighting direction [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. In contrast to
these works, our system focuses on this important selfie factor,
head postures, in real time.
      </p>
      <p>
        Photo Editing. An interactive photomontage system was
proposed to create a composite set of images for the same objects,
resulting in a single image for subsequent enhancement [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
A similar approach was proposed to obtain good flash images
using a pair of flash and non-flash images [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        Aesthetics Feedback. Aesthetics feedback was given for
photos taken by mobile users from compositions, color
combinations and the aesthetic rating [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. An aesthetic evaluation
approach utilized a peer-rated online photo-sharing website as
the input datasets based on machine learning [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        Assistive Technology. There are other researches that
considering about the assistive technology to support the interaction
between users and user interface. Voykinska et al. studied
about the way that blind people use to interact with social
network [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Researchers also provided approach to ask
remote paid workers to help the blind people with some camera
devices and the internet [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Similar research asked the crowd
to answer the visual questions to help blind people [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
FRAMEWORK
The framework of this work consists of both offline and online
parts as shown in Figure 2. For the offline part, we explored
good head postures by crowdsourcing results to define good
head posture. We generated hundreds of virtual selfies with
different head postures (two types of virtual selfies used in
this paper are created based on the 3D models from Unity3D).
Among these virtual selfies, we defined “good” head posture
as selfies which got higher than 7. For best selfie position, we
generated virtual selfies with different distances to the camera
to evaluate the attractiveness and chose the distance with the
highest score.
      </p>
      <p>For the online part, we extracted the user’s facial features for
each frame and generated a 3D facial model to calculate the
head posture of users using a geometric approach and in real
time. Our system calculates the closest good head posture to
the user’s current head posture and suggests it as the ideal pose
for the user. To communicate this selfie suggestion, our system
shows a visual guide in the user interface to direct the motion
of the user’s head to achieve the ideal head posture. Voice
instructions are also proposed to help the user to move their
head using the proposed system. After the user achieves the
ideal head posture, the proposed system provides a voice
notification and takes the selfies automatically. This helps the user
to take a good selfie more easily, thus reducing the difficulty
of taking good selfies and enhancing user satisfaction.
Crowdsourcing Tasks for Head Posture
We designed two crowdsourcing tasks to explore good head
postures on CrowdFlower. In the first task, the participants
were asked to score each selfie, and the average score of each
selfie was used to represent its attractiveness. In the first task,
we ask the crwodworkers to score each selfie according to its
attractiveness from 1 to 10 where 10 is stronly attractive while
1 represents not attractive at all. For each selfie, 5 participants
are asked to score it and we adopt the average score of each
selfie to represent its attractiveness. The second task is to ask
users to evaluate the attractiveness of the selfie with different
distances from camera. We adopted the 3D virtual model
instead of real selfies for evaluation due to the accurate control
of head posture that is possible with virtual selfies; meanwhile,
it is difficult to obtain real selfies with exact head postures. The
usage of virtual selfies can reduce other factors influencing
the selfies in order to specifically evaluate the attractiveness of
selfies. To construct a relatively comprehensive selfie database
that includes various head postures, we created a database
using virtual selfies, which was comprised of 486 different
head postures with pitch angle ranging from (-40, 40), roll
angle ranging from (-40, 40) and yaw angle ranging from (-30,
20). We sample virtual selfies in every 10 degrees for each
dimension.</p>
      <p>The results of the crowdsourcing task are shown in the right
of Figure 3. The green points represent the head postures that
were rated higher than 7, the red points represent the head
postures rated higher than 5 but lower than 7 and the grey
points represent the head postures rated lower than 5. In the
first crowdsourcing task, among 486 selfies, the 58 with the
highest score were defined as good head orientations. We
described the head postures that got scores higher than 7 as
“good”, 5-7 as medium and under 5 as bad. The proposed
selfie guidance system aims to take “good” selfies; thus, we
chose the “good” head posture to set the system. We show the
selected samples of the selfies which got good, medium and
bad scores in Figure 3. From these results we found that the
selfies with the front face achieved the highest scores from the
crowdsourcing results. However, the common users would not
be satisfied with single head pose, and more choices of head
poses are preferred.</p>
      <p>In the second crowdsourcing task, we generated 6 virtual
selfies with different distances considering the length of the
arm of human. We utilized the selfie scored as the highest to
define the best distances. We accepted a larger tolerance for
the distances for two reasons: first, we calculated the distance
roughly by calculating the area of the triangle constructed
by the far corners of the eyes and nose (this may be affected
by different head postures, possibly rendering it inaccurate);
from results of this task, we know that the distance doesn’t
affect the selfie as the head posture (selfies in a quite large
range get relatively similar marks).</p>
      <p>
        To obtain continuous assessment for head postures, we
interpolated the crowdsourcing results to compute a score
function for the head postures using trilinear interpolation.
Figure 4 illustrates the results of the interpolation (We
display it using several slices). The multi-colored progress
bar represents the distribution of color from high score to
low score of the head postures; the red region shows the
high-scoring head postures, and the blue region shows the
head postures that get relatively low scores. As a result, both
datasets contain large portions of relatively good selfies as
scored by the score function from the crowdsourcing results.
Since both datasets are comprised of selfies or photographs
that are considered relatively attractive, the experimental
results verified that a good head posture itself has a significant
impact on the aesthetic of a selfie.
Facial Feature Extraction Based on Face Landmark
In this work, facial feature extraction plays an important role
in the estimation of the head posture of the user. Many
researchers have made efforts to solve this problem in order to
attain a high accuracy in the recognition of human faces and
facial features. Also in this work, we extracted eyes, mouth
and nose feature points based on dlib [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] which uses 68 points
to describe human faces. This algorithm is a feature-based
algorithm based on the HOG combined with a linear classifier
and an image pyramid, and sliding window detection scheme.
Head Posture Estimation Based on Geometry Model
Head pose estimation is the key part in this work that can
estimate the head orientation of the user based on the facial
features we extracted. The head-movement suggestions
provided in the user interface are all based on the estimation
results. In this work, we estimate the head posture using a
geometry model [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. To introduce how we estimate the head
posture, we will first introduce how we describe the head
posture. Generally, we define it as the position and the orientation
of the head. However, the position of the head can be
calculated directly from the image; thus, we focus on how to
calculate the orientation from a single image (frame). Head
orientation can be described by roll, yaw and pitch angles. The
three axes are orthogonal, so that we can calculate each angle
respectively. The proposed geometry model assumes that the
head is bilateral symmetry and the two far corner of eyes and
the two mouth corners form a plane and also define a facial
normal to show the direction of the nose. This facial normal
shows the direction of the head. In contrast to other head-pose
estimation algorithms (tracking approach, etc.), the series data
of previous frames are not required by the proposed
geometry approach to obtain the head pose. Another approaches
is direct landmark fitting algorithm which is quite similar to
our approach. However, by establishing a facial model, our
approach enables relatively simple way to calculate each angle
for the head pose which is enough to proof the concept in our
research.
      </p>
      <p>
        Facial Model Conduction
We define L f to be the distance from the far corner of eyes
to the corner of the mouth in the facial model, Ln to be the
distance from the nose tip to the nose base, and Lm to be the
distance from nose base to the mouth. Le denotes the distances
between two far corner of the eyes. Notice that l f , ln, lm, le are
the projected values. We calculate the following three ratios
to define the facial model: Rm, and Rn. Here, Rm is calculated
by Lm=L f and Rn is Ln=L f . Because all these ratios are quite
stable with respect to different faces, constant values can be
adopted to replace the ratios unless a very high accuracy of
the orientation is required. For the typical face, Rm = 0:6,
Rn = 0:4. The Rm and Rn are used to calculate the slant as
shown in the next section, and the more detailed description
can be found in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>Estimating of the User’s Head Posture
We represent a head posture with t, which is the angle between
the symmetry axis and the facial normal (the projected one in
the 2D image) and a slant s , which is the angle between the
normal of the image and the facial normal. The facial normal
can be calculated as Eq. (1).</p>
      <p>nˆ = [sin s cos t; sin s sin t; cos s ]
(1)
Slant s can be calculated by the normal of the image d =
2
[xnormal ; ynormal ; znormal ] as coss =j znormal j (znormal is a
normalized value). znormal can be defined using two values v1 and
v2 which is shown in the Eq. (2) and the projection matrix
(I</p>
      <p>ddT ).
2
znormal =
(2)
(3)
(4)
R2n
ln
l f
2
v1
2v2Rn2 + p(v1
.
Calculating yaw angle and pitch angle involves the normal we
gained before and the x, y, z represents each component of
facial normal vector which is defined in Eq. (1) (x = sin s cos t,
y = sin s sin t, z = cos s , r = px2 + y2 + z2).
These three angles describe the head orientation of the end
user in the real world. Using this head-orientation information
and the crowdsourcing results, the system can provide further
suggestions for head-movement to the user in real time.
USER INTERFACE
We implemented the selfie guidance system on the iOS 9.2
using Swift and Objective-C to instruct the user to achieve the
head posture that was defined to be good by the
crowdsourcing tasks. The system applied the face-landmark algorithm to
process the camera frame in question. We adopted the eyes,
nose and mouth points to estimate the head posture. For the
first frame, the system extracts the facial features to initialize
the face-geometry model. It then calculates the head posture
with the facial features from each frame and the face-geometry
model. The whole system is described in Algorithm 1. The
user interface of our selfie guidance system consists of a
visual user interface and voice instruction. We combined these
two methods for interacting with the user to provide new user
experiences and reduce the difficulty for users in
understanding the instructions when taking selfies. For pilot study, we
implemented two selfie guidance systems: only visual user
interface; both visual and voice user interfaces. From the user
feedback, the participants reported that it is interesting and
convenient to following the voice instructions.</p>
      <p>Visual User Interface
The visual user interface is implemented as shown in Figure 6.
To reduce the operational difficulty of head movement, the user
interface provides suggestions for roll, yaw and pitch angles
separately. Thus, the user can move their head one direction
at a time. With the combination of these head movements,
the user can easily achieve good head posture. The arrows
shown in the top and bottom of the user interface suggest the
pitch angle and ask users to adjust their head by tilting their
head up or down slightly. The arrows in the right and left
Algorithm 1: Real-time processing for taking a selfie
input :Camera frame including a human face.</p>
      <p>output :A good selfie stored in the album.
1 Define good head postures by crowdsourcing approach using
486 virtual selfies with different head postures. This results
in 58 good head posture sets.
2 Extract the eyes, nose and mouth points from the frame based
on the face landmark algorithm which uses HOG feature
combining with a linear classifier to estimate a vector S (S
contains the coordinates information of each facial point).
3 Use the facial features extracted from the frame to establish
geometry facial model F.
4 Based on F and the facial feature information to estimate the
head posture Pc from the current frame which is the current
head posture of the user.
5 Calculate the Euclidean distances between P and each good
head posture Pg (defined by crowdsourcing task) to find a
minimum value which is defined as “the closest good head
posture” Pcg (notice here the system calculates the good
head posture once until the frame count fc to be reset).
6 Compare each angle defined in Pcg and those of Pc to provide
different suggestions
7 Go back to step 1 if the Pc doesn’t equal to the Pcg (with a
tolerance Dd).
8 Terminate the loop when Pc is close enough to the Pcg (with
Dd); take photo and store to the album and reset the frame
count fc.
side of the screen instruct users to turn their head right or left
slightly. The semicircle arrows instruct the user to tilt his or
her head right or left. These arrows will keep showing in the
user interface until the user’s head achieves the right position.</p>
      <p>Voice User Interface
In the voice instruction part, the selfie-supporting system
provides voice instructions to help the user to understand how to
move their head to match the good head posture. There are
four types of voice instructions that indicate suggestions for
proper head orientation and distance, and that check for the
presence of a smile. For the head orientation suggestions, the
system displays “tilt head right slightly!” and “tilt head left
slightly!” to suggest the roll angle; “tilt head up slightly!” and
“tilt head down slightly!” to suggest the pitch angle; and “turn
left slightly!” and “turn right slightly!” for the yaw angle,
and these suggestions correspond to the arrows on the visual
user interface. When the user has the right posture, and is
ready to take selfies, the system will display “Perfect! Let’s
take a photo, 3, 2, 1, cheese!” This notifies the user when to
take the selfie and provides extra time for them to adjust their
expressions.</p>
      <p>USER STUDY
We designed a comparison user study in order to evaluate the
validity of the developed selfie-supporting system. (We also
designed pilot study to compare the system that only have
visual instructions with the system that have visual and voice
instructions). Of the 8 participants, 2 are female and others
are male. The ages of all the participants are between 22 and
28. Some of them spend a great deal of time taking selfies
while others take very few or no selfies. We mainly asked
them to use our developed system and a normal camera to
experience the difference and asked them to fill out a
questionnaire to evaluate the system with regards to several aspects.
There were two tasks, and the participants were asked
to complete the two tasks in a random order. Task 1 asked
the participants to take several selfies using a normal camera
without any suggestions during the process. We asked them to
choose 5 selfies that they were satisfied with. They could take
as many selfies as they wanted until they got 5 satisfactory
selfies. Task 2 asked the participants to use the developed
selfie-supporting system to take 5 or 6 selfies. (The reason
some of the users take more than 5 selfies is that we were
trying to avoiding multiple selfies with the same head posture).
In this task, they followed the suggestions about their head
posture and also heard voice instructions while taking selfies.
In the experiment, 4 of the participants were asked to
complete Task 1 first and then complete the Task 2 while
the other 4 users were asked to complete Task 2 first and
then Task 1. In addition to this comparison user study, we
also asked them to use our selfie-supporting system with the
tripod settings to experience taking selfies with the visual- and
voice-instruction user interface while at a distance further than
human arms can reach. After all the tasks above, they were
also asked to fill a questionnaire to evaluate their satisfaction
with the system, the user interface and also the attractiveness
of the head postures provided by this system.</p>
      <p>RESULTS
We applied the score function according to the crowdsourcing
results to the result of the user study. Figure 7 and Figure 8
illustrate examples in which the user used a normal camera
to take selfies using our proposed system. From the result we
can see that the selfies that taken by our system get higher
scores on average. For example, for the second user, the
average score of selfies taken by a normal camera was 6.94;
meanwhile, the selfies taken by our system were scored 7.7 on
average. Here we divided the user into two groups according
to whether they were good at taking selfies or not. Figure 7
shows the examples of users who were not good at taking
selfies. We can see from this result that the proposed system
helps the users take better selfies overall. In these examples,
the average score of selfies taken by our system increases
around 30% compared to the selfies taken by normal camera.
Figure 8 shows the examples of users who exhibited relatively
better skill in taking selfies. They got higher scores even
using a normal camera. However, our system still helped them
get higher score on average. The proposed system can also
support them by providing more head postures so as to give
them options for different appearances.</p>
      <p>The questions in our questionnaire queried the user regarding
system satisfaction, feedback on head posture provided by
the system, the visual user interface and the voice instruction.
For these statements, users were asked to score the statement
from 1 to 5, where 1 represents “strongly disagree” and the 5
represents “strongly agree.” The evaluation and analysis of the
system have been done according to the result of the user study.
Figure 10 (a) shows the results of the comparison of the normal
camera and our system. As we can see in the Figure 10, 75%
of users think that our system is more convenient and 25% of
users think that the both systems are equivalent. Figure 10
(b) illustrates what users think of the head postures provided
in our system. The results show that 87:5% of users agree
that our selfie supporting system provide more choices for
them when taking selfies (25% of users choose strongly agree;
62:5% of users choose agree); 12:5% of users neither agree
nor disagree.</p>
      <p>Figure 9 (a) shows the user feedback concerning the
developed selfie-support system. We asked users about whether the
application helped them take better selfies and if the function
suggesting head posture was needed, in their opinion. We also
asked about whether they thought it was easier to take good
selfies with the supporting system. The chart shows the
average score of each question and we can see that the users have
quite high satisfaction with the system. Figure 9 (b) shows the
average scores for evaluations about the head posture, which
show that users confirm that the head postures provided from
the application are good. Figure 9 (c) shows the average scores
for evaluations about the visual user interface. In this part we
asked users whether they could understand the meaning of the
arrows and if they could realize which arrows suggested roll,
yaw and pitch angle, respectively. From these results we find
that the users can understand the meaning and how to follow
the suggestions well.</p>
      <p>Figure 11 shows the evaluations of the voice instructions. To
evaluate whether the voice instructions work when the user is
taking selfies, we asked the user about two aspects: whether
they were aware of the voice instruction and whether they
understood the meaning of the voice instructions. We also
asked the users whether the voice instructions were interesting
(Q24), whether they were helpful (Q25) and whether they
were easy to understand (Q26). According to the results, users
were able to notice and understand the voice instruction well.
CONCLUSION
In this work, we proposed a real-time selfie support system
which provides suggestions for the head posture of the user.
This proposed system aimed to help user to take good selfies.
Two crowdsourcing tasks were conducted to explore good
head orientations and the best distance from the camera
(position). In each crowdsourcing task, we generated virtual selfies
to create a database and asked crowdworkers to score each
virtual selfie by its attractiveness. After the crowdsourcing
task, we developed a selfie system which provides real-time
suggestions for head postures. We implemented both the
visual user interface, in which arrows suggest movements of
the head as well as the voice instructions, which also give
guidance to the users on how to move their head and in which
direction. A user study was conducted where participants were
asked to take selfies using two cameras (a normal camera and
the camera with our head-posture support system). According
to the study results, the proposed system can help users to take
good selfies and improve user satisfaction.</p>
      <p>In this work, we explored good head posture using a 3D virtual
model and applied the crowdsourcing results to all the users.
The aesthetic evaluation may vary due to different cultural
backgrounds or knowledge differences, so this could be further
studied by considering more parameters for certain users. The
integration of other selfie factors is not difficult to implement in
our selfie guidance system. In the future, it is worth exploring
the relationship between the substantial amounts of selfie data
and methods for subjective crowdsourcing evaluations using
deep learning to achieve satisfying selfies. An alternative way
to evaluate the quality of selfie can be explored by
examplebased approach to compare the reviews from social network
or ask crowdworkers to score the selfies taken by our system
and other work with functionality. The crowdsourcing-based
evaluation approach can also be applied to other areas such as
evaluating the accuracy of face recognition.</p>
      <p>ACKNOWLEDGMENTS
We thank all participants for joining our case studies, and
Morph 3D for sharing the head model in Figure 3. This work
was supported by JST CREST Grant Number JPMJCR17A1
and JSPS KAKENHI Grant Number JP17H06574, Japan.
Haoran Xie is funded by Epson International Foundation.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Aseem</given-names>
            <surname>Agarwala</surname>
          </string-name>
          , Mira Dontcheva, Maneesh Agrawala, Steven Drucker, Alex Colburn, Brian Curless, David Salesin,
          <string-name>
            <given-names>and Michael</given-names>
            <surname>Cohen</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Interactive Digital Photomontage</article-title>
          .
          <source>ACM Trans. Graph</source>
          .
          <volume>23</volume>
          ,
          <issue>3</issue>
          (Aug.
          <year>2004</year>
          ),
          <fpage>294</fpage>
          -
          <lpage>302</lpage>
          . DOI: http://dx.doi.org/10.1145/1015706.1015718
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Amit</given-names>
            <surname>Agrawal</surname>
          </string-name>
          , Ramesh Raskar,
          <string-name>
            <surname>Shree K. Nayar</surname>
            , and
            <given-names>Yuanzhen</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Removing Photography Artifacts Using Gradient Projection and Flash-exposure Sampling</article-title>
          .
          <source>ACM Trans. Graph</source>
          .
          <volume>24</volume>
          ,
          <issue>3</issue>
          (
          <year>July 2005</year>
          ),
          <fpage>828</fpage>
          -
          <lpage>835</lpage>
          . DOI: http://dx.doi.org/10.1145/1073204.1073269
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Jeffrey</surname>
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Bigham</surname>
            , Chandrika Jayant, Hanjie Ji, Greg Little,
            <given-names>Andrew</given-names>
          </string-name>
          <string-name>
            <surname>Miller</surname>
            ,
            <given-names>Robert C.</given-names>
          </string-name>
          <string-name>
            <surname>Miller</surname>
            , Aubrey Tatarowicz, Brandyn White,
            <given-names>Samuel</given-names>
          </string-name>
          <string-name>
            <surname>White</surname>
          </string-name>
          , and Tom Yeh.
          <year>2010</year>
          .
          <article-title>VizWiz: Nearly Real-time Answers to Visual Questions</article-title>
          .
          <source>In Proceedings of the 2010 International Cross Disciplinary Conference on Web Accessibility (W4A</source>
          )
          <article-title>(W4A '10)</article-title>
          . ACM, New York, NY, USA, Article
          <volume>24</volume>
          , 2 pages. DOI: http://dx.doi.org/10.1145/1805986.1806020
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Perfect</given-names>
            <surname>Corp</surname>
          </string-name>
          .
          <year>2015</year>
          . YouCam Perfect. (
          <year>2015</year>
          ). http://www.perfectcorp.com/stat/product/CyberLink_app/ Perfect_Corp/enu/design/perfectCorp/index.jsp
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Ritendra</given-names>
            <surname>Datta</surname>
          </string-name>
          , Dhiraj Joshi,
          <string-name>
            <given-names>Jia</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and James Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Studying Aesthetics in Photographic Images Using a Computational Approach</article-title>
          .
          <source>In Proceedings of the 9th European Conference on Computer Vision</source>
          - Volume
          <source>Part III (ECCV'06)</source>
          . Springer-Verlag, Berlin, Heidelberg,
          <fpage>288</fpage>
          -
          <lpage>301</lpage>
          . DOI:http://dx.doi.org/10.1007/11744078_
          <fpage>23</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Gee</surname>
          </string-name>
          and
          <string-name>
            <given-names>Roberto</given-names>
            <surname>Cipolla</surname>
          </string-name>
          .
          <year>1994</year>
          .
          <article-title>Determining the gaze of faces in images</article-title>
          .
          <source>Image and Vision Computing</source>
          <volume>12</volume>
          ,
          <issue>10</issue>
          (
          <year>1994</year>
          ),
          <fpage>639</fpage>
          -
          <lpage>647</lpage>
          . DOI: http://dx.doi.org/10.1016/
          <fpage>0262</fpage>
          -
          <lpage>8856</lpage>
          (
          <issue>94</issue>
          )
          <fpage>90039</fpage>
          -
          <lpage>6</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Neel</given-names>
            <surname>Joshi</surname>
          </string-name>
          , Wojciech Matusik,
          <string-name>
            <surname>Edward H. Adelson</surname>
            , and
            <given-names>David J.</given-names>
          </string-name>
          <string-name>
            <surname>Kriegman</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Personal Photo Enhancement Using Example Images</article-title>
          .
          <source>ACM Trans. Graph</source>
          .
          <volume>29</volume>
          ,
          <issue>2</issue>
          ,
          <string-name>
            <surname>Article 12</surname>
          </string-name>
          (
          <year>April 2010</year>
          ),
          <volume>15</volume>
          pages. DOI: http://dx.doi.org/10.1145/1731047.1731050
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>V.</given-names>
            <surname>Kazemi</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Sullivan</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>One millisecond face alignment with an ensemble of regression trees</article-title>
          .
          <source>In 2014 IEEE Conference on Computer Vision and Pattern Recognition</source>
          .
          <year>1867</year>
          -
          <fpage>1874</fpage>
          . DOI: http://dx.doi.org/10.1109/CVPR.
          <year>2014</year>
          .241
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Walter</surname>
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Lasecki</surname>
            , Phyo Thiha, Yu Zhong, Erin Brady, and
            <given-names>Jeffrey P.</given-names>
          </string-name>
          <string-name>
            <surname>Bigham</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Answering Visual Questions with Conversational Crowd Assistants</article-title>
          .
          <source>In Proceedings of the 15th International ACM SIGACCESS Conference on Computers and Accessibility (ASSETS '13)</source>
          . ACM, New York, NY, USA, Article
          <volume>18</volume>
          , 8 pages. DOI:http://dx.doi.org/10.1145/2513383.2517033
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>Qifan</given-names>
            <surname>Li</surname>
          </string-name>
          and
          <string-name>
            <given-names>Daniel</given-names>
            <surname>Vogel</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Guided Selfies Using Models of Portrait Aesthetics</article-title>
          .
          <source>In Proceedings of the 2017 Conference on Designing Interactive Systems (DIS '17)</source>
          . ACM, New York, NY, USA,
          <fpage>179</fpage>
          -
          <lpage>190</lpage>
          . DOI: http://dx.doi.org/10.1145/3064663.3064700
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>LoftLab</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>CreamCam Selfie Filter</article-title>
          . (
          <year>2016</year>
          ). http://creamcamapp.com
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <given-names>Baback</given-names>
            <surname>Moghaddam</surname>
          </string-name>
          and
          <string-name>
            <given-names>Alexander P.</given-names>
            <surname>Pentland</surname>
          </string-name>
          .
          <year>1994</year>
          .
          <article-title>Face recognition using view-based and modular eigenspaces</article-title>
          .
          <source>Proc. SPIE</source>
          <volume>2277</volume>
          (
          <year>1994</year>
          ),
          <fpage>12</fpage>
          -
          <lpage>21</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Violeta</surname>
            <given-names>Voykinska</given-names>
          </string-name>
          , Shiri Azenkot,
          <string-name>
            <surname>Shaomei Wu</surname>
            , and
            <given-names>Gilly</given-names>
          </string-name>
          <string-name>
            <surname>Leshed</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>How Blind People Interact with Visual Content on Social Networking Services</article-title>
          .
          <source>In Proceedings of the 19th ACM Conference on Computer-Supported Cooperative Work &amp; Social Computing (CSCW '16)</source>
          . ACM, New York, NY, USA,
          <fpage>1584</fpage>
          -
          <lpage>1595</lpage>
          . DOI: http://dx.doi.org/10.1145/2818048.2820013
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Lei</surname>
            <given-names>Yao</given-names>
          </string-name>
          , Poonam Suryanarayan, Mu Qiao, James
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Jia</given-names>
            <surname>Li</surname>
          </string-name>
          .
          <year>2012</year>
          . OSCAR:
          <article-title>On-Site Composition and Aesthetics Feedback through Exemplars for Photographers</article-title>
          .
          <source>International Journal of Computer Vision</source>
          <volume>96</volume>
          ,
          <issue>3</issue>
          (
          <year>2012</year>
          ),
          <fpage>353</fpage>
          -
          <lpage>383</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Mei-Chen Yeh</surname>
          </string-name>
          and
          <string-name>
            <surname>Hsiao-Wei Lin</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Virtual Portraitist: Aesthetic Evaluation of Selfies Based on Angle</article-title>
          .
          <source>In Proceedings of the 22Nd ACM International Conference on Multimedia (MM '14)</source>
          . ACM, New York, NY, USA,
          <fpage>221</fpage>
          -
          <lpage>224</lpage>
          . DOI: http://dx.doi.org/10.1145/2647868.2656401
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Jun-Yan</surname>
            <given-names>Zhu</given-names>
          </string-name>
          , Aseem Agarwala, Alexei A Efros,
          <string-name>
            <given-names>Eli</given-names>
            <surname>Shechtman</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Jue</given-names>
            <surname>Wang</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Mirror Mirror: Crowdsourcing Better Portraits</article-title>
          .
          <source>ACM Transactions on Graphics (SIGGRAPH Asia</source>
          <year>2014</year>
          )
          <volume>33</volume>
          ,
          <issue>6</issue>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>