<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Exploring eating behaviours modelling for user
clustering. In Proceedings of Third International Workshop on Health Recom-
mender Systems co-located with Twelfth ACM Conference on Recommender
Systems, Vancouver, BC, Canada, October</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Exploring eating behaviours modelling for user clustering</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sema Akkoyunlu</string-name>
          <email>sema.akkoyunlu@agroparistech.fr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Cristina Manfredotti</string-name>
          <email>cristina.manfredotti@agroparistech</email>
          <email>cristina.manfredotti@agroparistech. fr</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Antoine Cornuéjols</string-name>
          <email>antoine.cornuejols@agroparistech.fr</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nicolas Darcel</string-name>
          <email>nicolas.darcel@agroparistech.fr</email>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fabien Delaere</string-name>
          <email>fabien.delaere@danone.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Danone Nutricia Research</institution>
          ,
          <addr-line>Palaiseau</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>UMR MIA-Paris, AgroParisTech, INRA, Université Paris-Saclay</institution>
          ,
          <addr-line>Paris</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>UMR MIA-Paris, AgroParisTech, INRA, Université Paris-Saclay</institution>
          ,
          <addr-line>Paris</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>UMR MIA-Paris, AgroParisTech, INRA, Université Paris-Saclay</institution>
          ,
          <addr-line>Paris</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>UMR PNCA, AgroParisTech, INRA, Université Paris-Saclay</institution>
          ,
          <addr-line>Paris</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <volume>6</volume>
      <issue>2018</issue>
      <abstract>
        <p>Food based dietary guidelines are not fully adopted by consumers. One of the principal explanations for this failure is that they are too general and do not take into account eating habits. Experts in nutrition believe that providing personalized dietary recommendations via nutrition recommender system can help people improve their eating behaviours. Understanding eating habits is a keystone in order to build a context aware recommender system that delivers personalized dietary recommendations. As a step towards this goal, we propose a method for representing food consumptions based on Doc2Vec for discovering clusters of eating behaviours. We compare our method to the state of the art methods used in the nutrition community.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>CCS CONCEPTS</title>
      <p>• Information systems → Information extraction; •
Humancentered computing → User models;
food recommender systems; user modelling; eating behaviours;
Doc2Vec</p>
    </sec>
    <sec id="sec-2">
      <title>INTRODUCTION</title>
      <p>
        Most chronic diseases such as diabetes, obesity and cardiovascular
diseases are correlated to unhealthy eating habits [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]. In order to
help people to adopt healthier eating habits, public health agencies
have created dietary guidelines targeted to the general population.
These guidelines can be food based, for instance "eat at least 5
fruit or vegetable per day", "limit your consumption of salt" 1 or
∗HealthRecSys’18, October 6, 2018, Vancouver, BC, Canada. ©2018 Copyright for
the individual papers remains with the authors. Copying permitted for private and
academic purposes. This volume is published and copyrighted by its editors."
1http://solidarites-sante.gouv.fr/IMG/pdf/PNNS_2011-2015.pdf
nutrient based "XX gram of iron per day". However, the compliance
to the guidelines are relatively low although the awareness about
food based dietary guidelines is rather good [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Several causes
contribute to this phenomenon: cultural and personal preferences,
dificulty of implementing dietary changes, availability and price of
food items [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. One solution to this problem could be to provide
a food recommender system able to take into account most of
these causes. Early studies showed that web-based personalized
interventions are more efective than standard public health advice
for inducing compliance with healthy eating recommendations
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Moreover changing eating habits is challenging, thus food
based recommendations should better be easy to follow [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. But
for recommendations to be practical, one should first understand
consumers’ eating behaviour.
      </p>
      <p>
        In food related recommender systems, the recommended objects
are recipes [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], food items [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] or menus [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Recipe
recommendation systems take advantage of users’ past recipes ratings to propose
recipes that they might like. Menu based recommendation systems
combine meals that users showed preference for with nutritional
constraints based on the nutritional requirements of users. Food
item based recommendation systems are designed to learn the users’
tastes for food items. Most of them use popular recommendation
algorithms often based on matrix factorizations techniques which
learn an embedding space for representing users and food items
simultaneously. However, this representation does not take into
account that food items are seldom consumed in isolation and that
users’ preferences for food items can change in response to the
other food items consumed (i.e the dietary context) and to the
context of consumption (e.g. eating croissant for breakfast is acceptable,
but it is not for lunch). It seems necessary to take into account these
aspects for increasing the eficacy of food item recommendation
in real-life settings. Context-aware recommender systems seem
therefore to be the appropriate approach. However, modelling the
context is highly dependent on the domain at hand. It is thus
necessary to first model eating behaviours and understand how the
context impacts eating behaviours.
      </p>
      <p>
        Several dietary assessment methods are available: the food
frequency questionnaire (FFQ), 24-hour dietary recall (24HR) and food
diaries. FFQ are easy to implement and cost-efective however, the
questionnaire is tailored by research groups with a specific aim
in mind. Besides, its accuracy is not enough for recommendation
purposes. 24HR method is an interview that requires 30 minutes
rather precise but one day of consumption per user is not suficient
in order to learn preferences. Food diaries are a prospective
openended food consumption assessment method where consumers
write down all the food items and beverages consumed over a
specific time period [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. Quite often, the time periods go from 3 to
7 consecutive days. The main advantages are that no interviewer
is required, the whole process can be automatized adapted for
recommendation purposes and provided several days of consumption,
changes in diet can be captured. Throughout the paper, the toy
food diaries dataset in Table 1 will be used to illustrate the user
modelling methods.
      </p>
      <p>
        Dietary behaviour is modelled using two main types of
methods: theoretical ones and empirical ones [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Theoretical methods
use dietary indexes developed by research groups or agencies in
order to rank the healthiness of eating behaviours. Indexes are
constructed based on the current knowledge in nutrition but can also
include current dietary guidelines and recommendations which
are usually generated from empirical research. However, Newby
et al. [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] stress the fact that there can be conflicts when there is
no scientific consensus about what a healthy behaviour is before
analysis. It results in indexes that measure diferent definitions of
a healthy behaviour. In empirical methods, there is no nutritional
a priori about eating behaviours, i.e there is no definition about
what a healthy behaviour is. Patterns are found with no nutritional
a priori. We only focus on empirical ones as our goal is to learn
dietary behaviours based on consumption data in an unsupervised
way. In the literature, two methods stand out for discovering eating
behaviours: clustering and factor analysis. Cluster analysis aims
at discovering groups of behaviours, while factor analysis seeks
the most relevant factors. Clustering may use factor analysis as a
preprocessing step. Thus, the K-Means algorithm is often applied
to the matrix of consumption of food items directly [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] or after
dimension reduction using e.g. Principal Component Analysis (PCA)
[
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] or Non-Negative Matrix Factorisation [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ].
      </p>
      <p>
        To our knowledge, there is no comprehensive review about
methods used for deriving empirically eating patterns [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Each study
works on its own dataset and, most of the time, only one method
of dimension reduction is applied for deriving eating behaviours.
There is no apparent gold standard method, but the existing
literature seems to favour the use of PCA.
      </p>
      <p>
        These methods are reductionist: they only consider food items
alone. Nutrition experts argue that this reductionist perspective
may not be eficient for recommendation purposes: deeper and more
complex information are needed [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. Opposed to the reductionist
viewpoint, the holistic approach considers the diet as "a dynamic
interaction of the parts of their synthesis" [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Food item interactions
should accordingly be used for modelling eating behaviours.
      </p>
      <p>
        One solution would be to consider dietary data in a meal-based
form. Meal pattern analysis provides more details regarding the way
people compose their meals [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] and could provide more insights for
characterising eating behaviours. This approach takes into account
the complexity of the diet and aims at overcoming the limitations
of the study of foods in isolation [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. A meal based approach for
discovering eating behaviours was introduced by Woolhead et al.[
        <xref ref-type="bibr" rid="ref24">24</xref>
        ].
They used frequent itemsets to generate a generic meal
classification. They derived 63 generic meals across all meal types and
computed mean daily nutrient intakes associated to the generic
meals. For each subject, mean daily intakes of energy percentage
contribution of each generic meal type was computed. Then PCA
was applied to discover eating behaviours. Authors themselves
argue that this methodology induces a subjective classification.
Besides, relying on frequent itemsets to code meals may overlook
infrequent eating patterns at a population level but frequent at an
individual level, discarding these patterns as noise. This shows the
necessity of an adequate representation of meals.
      </p>
      <p>Developing a food recommender system that takes into account
the meals and their context, and not only food items, requires that
two main challenges be met: (1) finding a proper meal description
model in which distances between meals can be computed and
(2) discovering an adequate way of aggregating several meals for
computing distances between users in order to discover clusters of
eating behaviours.</p>
      <p>
        In this paper, our contribution is twofold: we propose a novel
domain of application of word embedding to user profiling and
we compare three approaches to describe eating behaviours. We
propose a new approach to model meal representation by applying
the Doc2Vec algorithm [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] in order to learn a meal embedding
space. This allows, in turn, the use of a cosinus similarity adapted to
matrices to compute similarities between users and infers clusters
of users. Moreover, in food based approaches, we compare the state
in the art methods with Doc2Vec applied on users.
      </p>
      <p>The rest of the paper is organised as follows. Section 2 describes
methods for user modelling. Section 3 reports the results of our
experiments on a real-world dataset. We discuss the results in Section
4 and we finally conclude in Section 5.
2
2.1</p>
    </sec>
    <sec id="sec-3">
      <title>METHODS</title>
    </sec>
    <sec id="sec-4">
      <title>Food-based methods</title>
      <sec id="sec-4-1">
        <title>2.1.1 State of the art methods.</title>
        <p>
          In alimentation behaviour science, researchers work mostly on
food items. They transform food consumption data into matrices
where the columns correspond to the frequency or the quantity
of consumption of food items and the rows to users as shown in
Figure 1. The next step consists in applying Principal Component
Analysis (PCA) or Non-Negative Matrix Factorization (NMF) [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ].
PCA consists in finding a set of linearly independent variables,
called principal components, that capture as much as possible the
variance of the data points. NMF is similar to PCA but imposes
a non-negativity constraint on the parameters of the model. This
is found useful in many domains such as signal processing and
recommender systems, because more amainable to interpretation
by experts [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ].
        </p>
        <p>
          Clusters of eating behaviours are then discovered by applying
K-Means algorithm on the result of PCA or NMF. In order to find
the optimal number of clusters, a popular clustering evaluation
metric is used, the silhouette coeficient [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ].
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>2.1.2 Another food based method: applying Doc2Vec to users.</title>
        <p>
          Word2Vec is a popular model for word embedding. Doc2Vec,
proposed by [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] is an extension of Word2Vec: instead of learning
word embeddings, the model learns distributed representations
of arbitrarily large units of text such as sentences, paragraphs or
documents. It was proposed in two flavours: DBOW (Distributed
Bag Of Words) and DMPV (Distributed Memory version of
Paragraph Vector). DBOW is simpler than DMPV as it does not take
into account the order of the words when learning the embedding
space. It is the version that is suited for our task as the order does
not matter. Besides, empirical evaluations of Doc2Vec showed that
DBOW performs better than DMPV [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ].
        </p>
        <p>The food based approach considers that a user is described by
the frequency of consumption of single food items. Similarly, a user
can be considered as a document where the food items eaten over
a specific amount of time play the role of words.</p>
        <p>Figure 2 is an illustration of what applying Doc2Vec algorithm
on individual eating consumptions means. Individual documents of
consumption are fed in the model. The result is an embedding space
of users based on their eating consumptions which means that each
user is described by a set of coordinates. Users are represented as
vectors in this figure because similarity between users is computed
with cosine similarity, a metric commonly used in document
retrieval. It is basically the angle between two user vectors. At this
stage of the method, we compute the similarity matrix of users. Our
goal is then to cluster users according to their similarity.</p>
        <p>
          Spectral clustering is a method that exploits similarity measures
by considering data points as nodes of a weighted connected graph.
Clusters are found by partitioning this graph based on the
eigenvectors of the Laplacian matrix derived from the similarity matrix.
Choosing the optimal number of clusters is often a problem for
clustering algorithms. There are several heuristics adapted for spectral
clustering. The heuristic advised by [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] is the eigengap heuristic.
The optimal number of clusters k is the number such that the
difference between the eigenvalues λk+1 − λk is large. Justification
for this procedure is provided in [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ].
2.2
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>A novel meal based method using Doc2Vec</title>
      <sec id="sec-5-1">
        <title>2.2.1 Learning an embedding space for meals.</title>
        <p>Meals are defined as combinations of food items simultaneously
consumed by one user at a single moment of consumption on one
survey day. Meals are actually lists of food items. In the meal based
approach, the objective is to be able to compute similarities between
meals in order to compute similarities between users to derive
clusters of users. However, it is not trivial to compute similarity
between two meals, for example between {pasta, beef, fruits} and
{rice, vegetable, fruits}.</p>
        <p>A straightforward idea would be to define first a similarity
between food items and then define a way to summarize those
similarities to compute a similarity between meals. We observe two
problems with this idea. First, there is no domain similarity measure
between food items. One can use classification of food items as a
proxy to a similarity measure. But there are lots of classification
schemes in the literature. Second, this approach is against the
philosophy of the holistic approach as it ignores interactions that may
exist between food items.</p>
        <p>An elegant way of learning such interactions is to learn an
embedding space with Doc2Vec. Indeed, the embedding is learned
in such way that similar meals are closer in the induced space
as showed in Figure 3. Each user is now described by a matrix
where the rows correspond to the meals and the columns to the
coordinates of meals in the Doc2Vec induced space.</p>
      </sec>
      <sec id="sec-5-2">
        <title>2.2.2 Computing distance between users.</title>
        <p>
          Once the meal representation is learned, the challenge becomes
one of computing a similarity between users. In our approach,
this amounts to compute the similarity between two documents
by taking into account the distances between sentences. Indeed,
meals can be considered sentences of users who are documents.
Mathematically speaking, this amounts to compute a similarity
between matrices. Such a similarity was introduced in [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. The
authors of the paper proposed a cosine kernel in order to compute
the similarity between the documents A and B in Equation 1:
⟨A, B⟩
cos(A, B) =
        </p>
        <p>∥A∥F · ∥B ∥F
where ⟨·, ·⟩ is the Frobenius inner product and ∥·∥F the
Frobenius norm. Using the Frobenius inner product enables to compare
the similarity of the sentences to determine the similarity of the
documents. Let us denote sA and sB the number of sentences in
document A and document B respectively. This formula implies that the
cosinus similarity is computed between the first sentences of both
documents then the second ones and so on until the min(sA, sB )-th
sentences. If one document is longer than the other one, the last
sentences of the longer document are not taken into account for
the similarity computation.</p>
        <p>For eating behaviour modelling, this means that two consumers
are similar if they eat similar meals at the same moment of the day
on the same day. This is a rather strong assumption concerning
eating behaviour modelling.
(1)</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>3 EXPERIMENTS</title>
    </sec>
    <sec id="sec-7">
      <title>3.1 INCA2 dataset</title>
      <p>
        The INCA2 dataset 2 consists of individual 7-day food records
collected during 2006-2007 from 2,624 adult French consumers
over several months in order to take into account seasonality. A
close-ended list of 1,342 food items organized in 122 sub-groups
and in 44 groups were used for coding the dietary records. Further
detail about the survey methods can be found in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. We decide to
work on sub-groups because the vocabulary is larger than when
considering groups while having enough repetitions unlike when
considering food items. We do not impose the number of clusters
to be the same for all the methods as we want to see if the number
of clusters that each method discovers is diferent, if the clusters
are overlapping or not.
3.2
      </p>
    </sec>
    <sec id="sec-8">
      <title>PCA and NMF on consumption data</title>
      <p>The state of the art methods require the selection of two parameters:
the number of components C of the reduction of dimensionality
method and the number of clusters K. The number of clusters k is
2https://www.data.gouv.fr/fr/datasets/donnees-de-consommations-et-habitudesalimentaires-de-letude-inca-2-3/
determined by using an internal clustering evaluation score, the
silhouette score. The optimal number of clusters is found when
the silhouette score is maximised. For PCA and NMF, we vary the
number of clusters between 2 and 30 and compute the silhouette
score. The score is maximised for k = 9.</p>
      <p>Loadings of factors of PCA and NMF can be give a hint about
the new representation space of users. Figure 4 shows the loadings
of factors of PCA according to food items. For ease of reading only
food items whose absolute value of contribution to any factor is
superior to 0.005 are displayed. NMF factors are shown in Figure 5.
The food items are displayed if their contribution to any factor is
superior to 0.3.
We constitute the corpus by aggregating the food item
consumptions per user, each user constituting a document. We use the
Gensim implementation of Doc2Vec in order to learn our model. The
corpus contains 2624 documents. After learning the model, we
compute the cosinus similarity of users and perform spectral clustering.
The optimal number of clusters is 5 clusters obtained using the
eigengap heuristic.
3.4</p>
    </sec>
    <sec id="sec-9">
      <title>Doc2Vec on meals</title>
      <p>We gather the corpus of meals by aggregating food items consumed
at the same moment of consumption, at the same day, by the same
user. The corpus is constituted of 37 283 unique meals. A meal
embedding is learned using the Gensim Doc2Vec implementation.
For each user, the vector of each of his meals is computed leading to
user matrices. The similarity matrix between users is obtained by
applying the cosine kernel to user matrices. Spectral clustering is
applied and the number of clusters is determined using the eigengap
heuristic. It yields 3 clusters.</p>
    </sec>
    <sec id="sec-10">
      <title>Comparison of the clustering results</title>
      <p>
        Our goal now is to compare the clustering results and determine
in which cases a food-based approach is adequate and the
contribution of a meal-based approach. In order to compare agreement
between clustering results, we compute the Adjusted Rand Index
(ARI) [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. It is a popular measure which consists in computing the
agreement between two clustering results i.e two partitions. ARI is
recommended for cases where the number of clusters is diferent,
which is our case. ARI takes values in [
        <xref ref-type="bibr" rid="ref1">−1, 1</xref>
        ], 1 meaning that both
clusterings agree, values close 0 mean that clusterings are made at
random.
      </p>
      <p>PCA
NMF
Doc2Vec users
Doc2Vec meals</p>
      <p>PCA
1</p>
      <p>FOOD BASED</p>
      <p>Doc2Vec
NMF users
0,93 0,14
1 0,13
1</p>
      <p>MEAL BASED
Doc2Vec
meals
0,017
0,018
0,013
1</p>
      <p>We also plot in Figure 6 the repartition of users in clusters across
the methods. From one method to another, the number of cluster is
attributed randomly and does not hold meaning.
4
4.1</p>
    </sec>
    <sec id="sec-11">
      <title>DISCUSSION</title>
    </sec>
    <sec id="sec-12">
      <title>Comparison of PCA and NMF for food based user modelling</title>
      <p>No matter the factorization method used before the clustering step,
the clustering results are very similar according to the Adjusted
Rand index. This means that the choice of the factorization method
for clustering users based on their food consumptions is not
primordial. However, as shown in Figures 4 and 5, the eating behaviours
discovered are diferent. The coeficients of PCA can be interpreted
as consumptions when positive and non consumption when
negative. For instance, the eating behaviour 0 consists in drinking tap
water but not spring or mineral water. We can also extract
information such that those who consume cofee do not consume tea
and vice versa. On the opposite, the coecfiients of NMF are strictly
positive hence the interpretation only concerns food consumptions.
For instance, the eating behaviour 0 consists in eating all types of
vegetables. The extracted eating behaviours are diferent according
to the method of reduction of dimensionality. We recommend to
test both methods to compare extracted eating behaviours as the
provided insights of both methods can be interesting.
4.2</p>
    </sec>
    <sec id="sec-13">
      <title>Contribution of Doc2Vec for the food based approach</title>
      <p>We apply Doc2Vec directly to users in order to challenge the state of
the art methods in food based approaches as we want to see how the
NLP method performs on this task. The number of clusters using
the Doc2Vec method on users yields a smaller number of clusters
and clustering results are rather diferent. A major drawback of this
method is that eating behaviours cannot be inspected as easily as in
the state of the art methods. Further analysis is needed in order to
understand why clustering results are so diferent. This method is
adequate if the objective is to extract clusters of consumers, however
in this state, this approach is not really adapted if explanations are
expected about eating behaviours. Being able to identify eating
behaviours is key for recommendation purposes as explanations
may be needed for people to implement the recommendations.
Usually, the performance of a neural language model is computed
on supervised tasks such as document retrieval or analogies. We
are in an unsupervised setting which complicates the assessment
of the performance of the learned meal embedding.</p>
    </sec>
    <sec id="sec-14">
      <title>Comparison state in the art food based approach and meal based approach</title>
      <p>It is in the meal based approach that the number of clusters is the
smallest. This shows that consumers of this dataset with regards
to their way of composing their meals are less diverse as we only
ifnd 3 clusters. This result should be interpreted in the light of the
assumption made about eating behaviours. We consider that two
consumers are similar in the meal based approach if they consume
similar meals on the same moment of the day on the same day,
a strong assumption on 7-day food diary data. This may lead to
more or less low values of similarity overall between users yielding
in lesser clusters. It would be interesting to investigate the
relaxation of this assumption by assuming that users are similar if they
consume similar meals regardless the day of consumption or the
moment of consumption. Again, it is dificult to extract eating
behaviours as the model is not designed for this purpose. Another
language model could be used for modelling food consumption,</p>
      <sec id="sec-14-1">
        <title>Latent Dirichlet Allocation (LDA) model.</title>
        <p>5</p>
      </sec>
    </sec>
    <sec id="sec-15">
      <title>CONCLUSION</title>
      <p>In this paper we explore user modelling in food consumption for
clustering users for recommendation purposes. We compare two
state of the art methods in the nutrition community. Our conclusion
is that both methods yield more or less the same clustering results.
However, the eating behaviours discovered are diferent. Moreover,
we propose a new food-based approach by considering food
consumptions as textual data and learning an embedding model with
Doc2Vec. The application of Doc2Vec to user food consumption
is adequate for user clustering, however it is not adapted for
extracting eating behaviours. We argued the importance of having
a holistic approach toward nutrition in order to make acceptable
recommendations. We propose a new meal based approach which
consists in learning a meal embedding space and then computing
user similarity based on their meals’ similarity. The usage of NLP
for food data analysis is promising. However, if clusters that can
be explained is needed (which is often the case), then it is better to
resort to generative language models such as LDA. Further work
will investigate the use of LDA for modelling eating behaviours.
6</p>
    </sec>
    <sec id="sec-16">
      <title>ACKNOWLEDGEMENT</title>
      <sec id="sec-16-1">
        <title>This study was funded by Danone Nutricia Research.</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Bier</surname>
            ,
            <given-names>D. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Derelian</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>German</surname>
            ,
            <given-names>J. B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Katz</surname>
            ,
            <given-names>D. L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pate</surname>
            ,
            <given-names>R. R.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Thompson</surname>
          </string-name>
          ,
          <string-name>
            <surname>K. M.</surname>
          </string-name>
          <article-title>Improving compliance with dietary recommendations</article-title>
          .
          <source>Nutrition Today</source>
          <volume>43</volume>
          ,
          <issue>5</issue>
          (sep
          <year>2008</year>
          ),
          <fpage>180</fpage>
          -
          <lpage>187</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Dubuisson</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lioret</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Touvier</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dufour</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Calamassi-Tran</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Volatier</surname>
            ,
            <given-names>J.-L.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Lafay</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <article-title>Trends in food and nutritional intakes of french adults from 1999 to 2007: results from the INCA surveys</article-title>
          .
          <source>British Journal of Nutrition</source>
          <volume>103</volume>
          ,
          <issue>07</issue>
          (dec
          <year>2009</year>
          ),
          <fpage>1035</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Elsweiler</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Harvey</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>Towards automatic meal plan recommendations for balanced nutrition</article-title>
          .
          <source>In Proceedings of the 9th ACM Conference on Recommender Systems, RecSys</source>
          <year>2015</year>
          , Vienna, Austria,
          <source>September 16-20</source>
          ,
          <year>2015</year>
          (
          <year>2015</year>
          ), pp.
          <fpage>313</fpage>
          -
          <lpage>316</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Freyne</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Berkovsky</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <article-title>Recommending food: Reasoning on recipes and ingredients</article-title>
          . In User Modeling, Adaptation, and
          <string-name>
            <surname>Personalization</surname>
          </string-name>
          , 18th International Conference, UMAP 2010,
          <string-name>
            <surname>Big</surname>
            <given-names>Island</given-names>
          </string-name>
          ,
          <string-name>
            <surname>HI</surname>
          </string-name>
          , USA, June 20-24,
          <year>2010</year>
          . Proceedings (
          <year>2010</year>
          ), pp.
          <fpage>381</fpage>
          -
          <lpage>386</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Ge</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elahi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fernaández-Tobías</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ricci</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Massimo</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <article-title>Using tags and latent factors in a food recommender system</article-title>
          .
          <source>In Proceedings of the 5th International Conference on Digital Health</source>
          <year>2015</year>
          (New York, NY, USA,
          <year>2015</year>
          ),
          <source>DH '15</source>
          , ACM, pp.
          <fpage>105</fpage>
          -
          <lpage>112</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Hageman</surname>
            ,
            <given-names>P. A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pullen</surname>
            ,
            <given-names>C. H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hertzog</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Boeckner</surname>
            ,
            <given-names>L. S.</given-names>
          </string-name>
          <article-title>Efectiveness of tailored lifestyle interventions, using web-based and print-mail, for reducing blood pressure among rural women with prehypertension: main results of the wellness for women: DASHing towards healthclinical trial</article-title>
          .
          <source>International Journal of Behavioral Nutrition and Physical Activity</source>
          <volume>11</volume>
          ,
          <issue>1</issue>
          (dec
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Hoffmann</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <article-title>Transcending reductionism in nutrition research</article-title>
          .
          <source>The American Journal of Clinical Nutrition</source>
          <volume>78</volume>
          ,
          <issue>3</issue>
          (sep
          <year>2003</year>
          ),
          <fpage>514S</fpage>
          -
          <lpage>516S</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Ivens</surname>
            ,
            <given-names>B. J.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Smith Edge</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>Translating the Dietary Guidelines to Promote Behavior Change: Perspectives from the Food and Nutrition Science Solutions Joint Task Force</article-title>
          .
          <source>J Acad Nutr Diet</source>
          <volume>116</volume>
          ,
          <issue>10</issue>
          (Oct
          <year>2016</year>
          ),
          <fpage>1697</fpage>
          -
          <lpage>1702</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Lau</surname>
            ,
            <given-names>J. H.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Baldwin</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <article-title>An empirical evaluation of doc2vec with practical insights into document embedding generation</article-title>
          .
          <source>In Proceedings of the 1st Workshop on Representation Learning for NLP, Rep4NLP@ACL</source>
          <year>2016</year>
          , Berlin, Germany,
          <year>August 11</year>
          ,
          <year>2016</year>
          (
          <year>2016</year>
          ), pp.
          <fpage>78</fpage>
          -
          <lpage>86</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Le</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <article-title>Distributed representations of sentences and documents</article-title>
          .
          <source>In Proceedings of the 31st International Conference on International Conference on Machine Learning</source>
          - Volume
          <volume>32</volume>
          (
          <year>2014</year>
          ),
          <article-title>ICML'14, JMLR</article-title>
          .org, pp.
          <fpage>II</fpage>
          -1188
          <string-name>
            <surname>-</surname>
          </string-name>
          II-1196.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>D. D.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Seung</surname>
            ,
            <given-names>H. S. Learning</given-names>
          </string-name>
          <article-title>the parts of objects by nonnegative matrix factorization</article-title>
          .
          <source>Nature</source>
          <volume>401</volume>
          (
          <year>1999</year>
          ),
          <fpage>788</fpage>
          -
          <lpage>791</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Luo</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xia</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          <article-title>An eficient non-negative matrixfactorization-based approach to collaborative filtering for recommender systems</article-title>
          .
          <source>IEEE Trans. Industrial Informatics</source>
          <volume>10</volume>
          ,
          <issue>2</issue>
          (
          <year>2014</year>
          ),
          <fpage>1273</fpage>
          -
          <lpage>1284</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Luxburg</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          <article-title>A tutorial on spectral clustering</article-title>
          .
          <source>Statistics and Computing</source>
          <volume>17</volume>
          ,
          <issue>4</issue>
          (Dec.
          <year>2007</year>
          ),
          <fpage>395</fpage>
          -
          <lpage>416</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Mijangos</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sierra</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Montes</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <article-title>Sentence level matrix representation for document spectral clustering</article-title>
          .
          <source>Pattern Recognition Letters</source>
          <volume>85</volume>
          (
          <year>2017</year>
          ),
          <fpage>29</fpage>
          -
          <lpage>34</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Newby</surname>
            ,
            <given-names>P. K.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Tucker</surname>
            ,
            <given-names>K. L.</given-names>
          </string-name>
          <article-title>Empirically derived eating patterns using factor or cluster analysis: a review</article-title>
          .
          <source>Nutr. Rev</source>
          .
          <volume>62</volume>
          ,
          <issue>5</issue>
          (May
          <year>2004</year>
          ),
          <fpage>177</fpage>
          -
          <lpage>203</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Rand</surname>
            ,
            <given-names>W. M.</given-names>
          </string-name>
          <article-title>Objective criteria for the evaluation of clustering methods</article-title>
          .
          <source>Journal of the American Statistical Association</source>
          <volume>66</volume>
          , 336 (dec
          <year>1971</year>
          ),
          <fpage>846</fpage>
          -
          <lpage>850</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Reedy</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wirfalt</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Flood</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mitrou</surname>
            ,
            <given-names>P. N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krebs-Smith</surname>
            ,
            <given-names>S. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kipnis</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Midthune</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leitzmann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hollenbeck</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schatzkin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Subar</surname>
            ,
            <given-names>A. F.</given-names>
          </string-name>
          <article-title>Comparing 3 dietary pattern methods-cluster analysis, factor analysis, and index analysis-with colorectal cancer risk: The NIH-AARP diet and health study</article-title>
          .
          <source>American Journal of Epidemiology 171</source>
          ,
          <issue>4</issue>
          (dec
          <year>2009</year>
          ),
          <fpage>479</fpage>
          -
          <lpage>487</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Rousseeuw</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <article-title>Silhouettes: A graphical aid to the interpretation and validation of cluster analysis</article-title>
          .
          <source>J. Comput. Appl. Math. 20</source>
          ,
          <issue>1</issue>
          (Nov.
          <year>1987</year>
          ),
          <fpage>53</fpage>
          -
          <lpage>65</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Shim</surname>
            ,
            <given-names>J.-S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oh</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Kim</surname>
          </string-name>
          , H. C.
          <article-title>Dietary assessment methods in epidemiologic studies</article-title>
          .
          <source>Epidemiology</source>
          and Health (jul
          <year>2014</year>
          ),
          <year>e2014009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Teng</surname>
          </string-name>
          , C.-Y.,
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>Y.-R.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Adamic</surname>
            ,
            <given-names>L. A.</given-names>
          </string-name>
          <article-title>Recipe recommendation using ingredient networks</article-title>
          .
          <source>In Proceedings of the 4th Annual ACM Web Science Conference</source>
          (New York, NY, USA,
          <year>2012</year>
          ),
          <source>WebSci '12</source>
          , ACM, pp.
          <fpage>298</fpage>
          -
          <lpage>307</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Thorpe</surname>
            ,
            <given-names>M. G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Milte</surname>
            ,
            <given-names>C. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Crawford</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>McNaughton</surname>
            ,
            <given-names>S. A.</given-names>
          </string-name>
          <article-title>A comparison of the dietary patterns derived by principal component analysis and cluster analysis in older australians</article-title>
          .
          <source>International Journal of Behavioral Nutrition and Physical Activity</source>
          <volume>13</volume>
          ,
          <issue>1</issue>
          (feb
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Webb</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Byrd-Bredbenner</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <article-title>Overcoming consumer inertia to dietary guidance</article-title>
          .
          <source>Advances in Nutrition 6</source>
          ,
          <issue>4</issue>
          (jul
          <year>2015</year>
          ),
          <fpage>391</fpage>
          -
          <lpage>396</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Wendel</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dellaert</surname>
            ,
            <given-names>B. G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ronteltap</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , and van Trijp,
          <string-name>
            <surname>H. C.</surname>
          </string-name>
          <article-title>Consumers' intention to use health recommendation systems to receive personalized nutrition advice</article-title>
          .
          <source>BMC Health Services Research</source>
          <volume>13</volume>
          , 1 (apr
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Woolhead</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gibney</surname>
            ,
            <given-names>M. J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Walsh</surname>
            ,
            <given-names>M. C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brennan</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Gibney</surname>
            ,
            <given-names>E. R.</given-names>
          </string-name>
          <article-title>A generic coding approach for the examination of meal patterns</article-title>
          .
          <source>The American Journal of Clinical Nutrition</source>
          <volume>102</volume>
          ,
          <issue>2</issue>
          (jun
          <year>2015</year>
          ),
          <fpage>316</fpage>
          -
          <lpage>323</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25] World Health Organization.
          <article-title>Diet, nutrition and the prevention of chronic diseases: report of a joint who/fao expert consultation</article-title>
          .
          <source>Tech. rep.</source>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Zetlaoui</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Feinberg</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Verger</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Clemençon</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <article-title>Extraction of food consumption systems by nonnegative matrix factorization (NMF) for the assessment of food choices</article-title>
          .
          <source>Biometrics</source>
          <volume>67</volume>
          ,
          <issue>4</issue>
          (
          <year>2011</year>
          ),
          <fpage>1647</fpage>
          -
          <lpage>1658</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>