<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>An Explanatory Matrix Factorization with User Comments Data*</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Donghyun Kim</string-name>
          <email>dhk618@kaist.ac.kr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hayong Shin</string-name>
          <email>hyshin@kaist.ac.kr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Matrix Factorization, Explanatory Analysis, Latent Dirichlet</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Allocation</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Industrial and Systems Engineering, Korea Advanced Institute of Science and Technology</institution>
          ,
          <country country="KR">South Korea</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <abstract>
        <p>Matrix factorization is one of the crucial algorithms of the Recommendation system. It implies that the relationship between user and contents can be explained by hidden latent variables. However, it is not intuitive to understand the meaning of these hidden latent variables. Therefore, this study suggests a way to learn the meaning from supplementary data such as comments and use in matrix factorization. The data used in this study is user comment data from Naver which is the largest web platform and also the largest Webtoons (Web comics) platform in South Korea. We show that the suggest method which uses the supervised latent variable also fits well with users with the distinct tendency compare to conventional matrix factorization.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        Recommendation system (RS) is referred to collecting information
to analyze user’s taste. Numerous methods have been proposed for
the RS, and one of overwhelming method is matrix factorization
(MF) which is predicting a missing value of a score matrix
composed of evaluation for contents given by the user [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. MF
improved the quality of RS significantly, but there are some issues
such as a cold-start problem, insufficient explanatory power, etc.
MF decomposes into low rank matrices with latent features and
make the original score matrix treatable, but it was difficult to
analyze the meaning of each latent features.
      </p>
      <p>
        There are a lot of works that uses user reviews to assist RS. Also,
it is shown that the appropriate latent factor model using topic
selection with LDA is better than the existing model [
        <xref ref-type="bibr" rid="ref2 ref4">2, 4</xref>
        ].
However, previous studies assumes that there exist score matrix
and use reviews to make better while not only this study does not
have score matrix but also this focus on the explanatory power of
MF not the RMSE itself. The methodology itself is not new as part
of research using user reviews to make better RS, but the two
popular methodologies, MF and Latent Dirichlet Allocation (LDA),
have been mixed appropriately and give exploratory power. It also
differ as using user comment data about Webtoons which was not
used previously.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>EXPLANTORY MATRIX FACTORIZATION</title>
      <p>
        In this study, we introduce a method to utilize domain knowledge
by combining LDA and MF. The LDA has explanatory power on
topics, and the MF is explaining the relationship between the user
and the contents with hidden latent variable. However, the meaning
of each latent variable is difficult to grasp. Therefore, we first
derive explanatory power from the supplementary data , =
1 … , such as user comments or the report using LDA. The LDA
assumes that several topics are mixed in each document, and the
analysis results can analyze the themes of documents. There are
many LDA algorithms to infer topic [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], so we do not describe
conventional LDA algorithms in detail. Let , = 1, … , , be a
topic from LDA, then each user can be represented as vector =
, … , , where = . This value indicates how
much a specific user talks about a particular topic, which is an
indirect indicator that shows what the user likes. Therefore, we can
make a user-topic relationship matrix = , , … , for
whole user and use it directly in MF. The core of MF is to divide
an user-contents rating matrix into two low-rank matrices =
which needs to be estimated. However, if can be obtained
sufficiently from supplementary data, MF turned into a simple
matrix inverse problem. Thus, can be obtained simply through
the Moore-Penrose pseudoinverse with = .
3
3.1
      </p>
    </sec>
    <sec id="sec-3">
      <title>EXPERIMENT</title>
    </sec>
    <sec id="sec-4">
      <title>Data Collection and Refinement</title>
      <p>We use user comment data from Naver, which is Korea's largest
web and Webtoon platform. However, since Naver does not
provide any formalized data, the data was collected and refined
through web crawling by Python. The collected data contains 4
features (Title-Episodes-userID-Comment). The raw data has over
100K users, 151 Webtoons, 21927 episodes and over 110 million
comments. Since this raw data needs more than 10TB of capacity,
due to hardware limitations, we limited to small size data. Also,
unlike commonly used reviews, there are a lot of useless data
because comments can be written without any restrictions such as
An Explanatory Matrix Factorization with User Comments Data
multiple comments is allowed in same item. Therefore, in this study,
we chose 3,000 users who kept the grammar as much as possible
and wrote over a reasonable length (more than 70 characters in
average) for certain period consistently (at least 12 weeks).
Compared to all data, 3000 users are quite small numbers, but since
the data used in this paper is very different from the user review
usually used in other papers, it was important to refine useful data
before analysis. This data contains 149 Webtoons, 1.1 million
comments.
3.2</p>
    </sec>
    <sec id="sec-5">
      <title>Experiment Settings and Result</title>
      <p>The topic is modeled through the LDA with selected 1.1 million
comments. In this experiment, the number of topic was set to 10,
20, and 30, and the hidden latent variables of MF were also set to
be the same in each case. Topic selection is very important task, but
it is too vague to use the whole as it is. Therefore, some topics are
collected through each Webtoon, and the some topics are obtained
by whole data. Some of the noticeable topics are listed in Table 1.
Topic 3 is mainly composed of words about stories of comic, and
Topic 7 is made of the drawing style of comic. Even though not all
topics can be identified as the above topics, but there are more
topics that can be interpreted, such as the attitude of the artiest, etc.
=
)</p>
      <sec id="sec-5-1">
        <title>Topic 3</title>
        <p>Sick of</p>
        <p>Crazy
Main character
Story</p>
      </sec>
      <sec id="sec-5-2">
        <title>Topic 7</title>
        <p>Beautiful
Sick of
Drawing
Color</p>
      </sec>
      <sec id="sec-5-3">
        <title>Topic 21</title>
        <p></p>
        <p>Funny
Best comments</p>
        <p>
          Clear
Each user vectors are constructed by the topics we obtained.
We use cosine similarity which is most commonly used. Using this
similarity, we construct a matrix to be used in MF and simply
obtain other matrix . We compare RMSE with original MF. In this
paper we use most basic MF algorithm [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] as conventional MF
algorithm. The score matrix M used in this experiment is composed
of 0 and 1 which indicates whether the user sees a certain comic,
not the score rating. We assume that the user only sees the comics
they commented on. In order to measure the RMSE, about 15% of
each user data was randomly deleted. Therefore, we learned with
85% of the data and observe the difference between the erased 15%
actual data and the predicted data. As can be seen from the results
Figure 1, it cannot be concluded that the overall data performance
is better than conventional MF. However, when compared only for
those with distinct tendencies, whose − &gt;
(in this paper α = 0.4 is used) which means user has at least one
noticeable topic that can be categorized more clearly than other
users, it can be seen that the suggested method using the LDA is
slightly better than the conventional MF method. In other words,
we can see that the unsupervised latent variable which is
conventional MF fits better with users who judge the contents with
a complex view, and the suggested EMF which uses the supervised
latent variable fits well with users with a simple view. This result
cannot be regarded as meaningful for RMSE itself, but it can be
implied that it has a similar RMSE even though it is obtained by
simple matrix inversion using a relatively interpretable latent
variable rather than the existing method. The reason why RMSE is
lower than other studies is because it is not to predict the score, but
to determine whether user sees a specific Webtoon, so we calculate
RMSE with a rounded value which is 0 or 1.
        </p>
        <p>E
S
M</p>
        <p>R</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>CONCLUSIONS</title>
      <p>In this paper, the two popular methodologies, MF and LDA, have
been mixed appropriately and shows some extra synergy. We
conducted experiments with comment data of Webtoons and shows
the suggest method works quite well as much as conventional MF,
and some cases it works better. This study did not fully use the
comment data that is currently available. Webtoon is a content that
is published one episode a week, so we think it will be very
influential in time. Also, this study is domain specific research and
the proposed algorithm is used only in this domain, so the extensive
research with certified data set is needed to generalize the algorithm.</p>
    </sec>
    <sec id="sec-7">
      <title>ACKNOWLEDGMENTS</title>
      <p>This research was supported by Basic Science Research Program
through the National Research Foundation of Korea funded by the
Ministry of Science, ICT &amp; Future Planning (2017R1A2B4006290).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Bobadilla</surname>
          </string-name>
          ,
          <string-name>
            <surname>Jesús</surname>
          </string-name>
          , et al.
          <year>2013</year>
          .
          <article-title>Recommender systems survey</article-title>
          .
          <source>Knowledgebased systems 46</source>
          (pp.
          <fpage>109</fpage>
          -
          <lpage>132</lpage>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Seroussi</surname>
            , Yanir,
            <given-names>Fabian</given-names>
          </string-name>
          <string-name>
            <surname>Bohnert</surname>
            , and
            <given-names>Ingrid</given-names>
          </string-name>
          <string-name>
            <surname>Zukerman</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Personalised rating prediction for new users using latent factor models</article-title>
          .
          <source>Proceedings of the 22nd ACM conference on Hypertext and hypermedia</source>
          (pp.
          <fpage>47</fpage>
          -
          <lpage>56</lpage>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Bobadilla</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ortega</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hernando</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Gutiérrez</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>2013</year>
          .
          <article-title>Recommender systems survey</article-title>
          .
          <source>Knowledge-based systems</source>
          ,
          <volume>46</volume>
          ,
          <fpage>109</fpage>
          -
          <lpage>132</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Chen</surname>
            , Li,
            <given-names>Guanliang</given-names>
          </string-name>
          <string-name>
            <surname>Chen</surname>
            , and
            <given-names>Feng</given-names>
          </string-name>
          <string-name>
            <surname>Wang</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Recommender systems based on user reviews: the state of the art</article-title>
          .
          <source>User Modeling and User-Adapted Interaction</source>
          <volume>25</volume>
          (
          <issue>2</issue>
          ) (pp.
          <fpage>99</fpage>
          -
          <lpage>154</lpage>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Alghamdi</surname>
            , Rubayyi, and
            <given-names>Khalid</given-names>
          </string-name>
          <string-name>
            <surname>Alfalqi</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>A survey of topic modeling in text mining</article-title>
          .
          <source>I. J. ACSA 6</source>
          .
          <issue>1</issue>
          (pp.
          <fpage>147</fpage>
          -
          <lpage>153</lpage>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Koren</surname>
            , Yehuda,
            <given-names>Robert</given-names>
          </string-name>
          <string-name>
            <surname>Bell</surname>
            , and
            <given-names>Chris</given-names>
          </string-name>
          <string-name>
            <surname>Volinsky</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Matrix factorization techniques for recommender systems</article-title>
          .
          <source>Computer</source>
          <volume>42</volume>
          (
          <issue>8</issue>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>