<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Open Source Recommendation Systems for Mobile Application</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Renata Ghisloti De</string-name>
          <email>renata.ghisloti@isep.fr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Raja Chiky</string-name>
          <email>raja.chiky@isep.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zakia Kazi Aoul</string-name>
          <email>zakia.kazi@isep.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>LISITE-ISEP</institution>
          ,
          <addr-line>28 rue Notre Dame Des, Champs, 75006 Paris</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Souza, LISITE-ISEP</institution>
          ,
          <addr-line>28 rue Notre Dame Des, Champs, 75006 Paris</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>The aim of Recommender Systems is to suggest useful items to users. Three major techniques can be highlighted in these systems: Collaborative Filtering, Content-Based Filtering and Hybrid Filtering. The collaborative method proposes recommendations based on what a group of users have enjoyed and it is widely used in Open Source Recommender Systems. The work presented in this paper takes place in the context of SoliMobile Project that aims to design, build and implement a package of innovative services focused on the individual in unstable situation (unemployment, homeless, etc.). In this paper, we present a study of open source recommender systems and their usefulness for SoliMobile. The paper also presents how our recommender system is fed by extracting implicit ratings using the techniques of Web Usage Mining.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Categories and Subject Descriptors</title>
      <p>H.3.4 [Systems and Software]: Performance evaluation
(e ciency and e ectiveness); H.2.8 [Database
applications]: Data mining|Web usage mining
Open source recommender systems, collaborative ltering,
Mahout, Web usage mining</p>
    </sec>
    <sec id="sec-2">
      <title>1. INTRODUCTION</title>
      <p>The amount of information in the web has greatly
increased in the past decade. This phenomenon has
promoted the advance of the recommender systems research
Permission to make digital or hard copies of all or part of this work for
personal or classroom use is granted without fee provided that copies are
not made or distributed for profit or commercial advantage and that copies
bear this notice and the full citation on the first page. To copy otherwise, to
republish, to post on servers or to redistribute to lists, requires prior specific
permission and/or a fee.</p>
      <p>
        Copyright is held by the author/owner(s). Workshop on the Practical Use of
Recommender Systems, Algorithms and Technologies (PRSAT 2010), held
in conjunction with RecSys 2010. September 30, 2010, Barcelona, Spain..
area. These systems intend to help users by providing
useful suggestions to them. They may suggest items in di
erent manners, such as comparing the user taste with other
users tastes or comparing the users preferences with other
items de nitions. These two methods are the so called
collaborative ltering [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and content-based ltering [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. The
collaborative method presents advantages over the
contentbased one. It is more e cient in practice and simpler to
implement. Due to this fact, the majority of open source
projects choose it. Current open source recommender
system projects are usually built on the item-based approach,
a type of collaborative ltering. Their features vary on the
programming language, extent of documentation and
magnitude of the project.
      </p>
      <p>
        We give in this paper an overview on known
recommendation techniques and we analyze open source projects in
this eld of research. Our interest of recommender systems
is justi ed by the fact that we have to choose one of the
studied systems and to integrate it in a complex platform
that includes a Web platform, a personalization system and
a mobile interface. This platform is developed through the
SoliMobile project, funded by ProximaMobile [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>The SoliMobile project in which we are involved, aims at
providing a portal services helping and assisting people who
are in di erent unstable situations. This project provides
end users with information to facilitate the process to access
to charities services from anywhere. The portal has to o er
services adapted to each user pro le, taking into account
their preferences and navigation traces. Our work aims to
provide the user with a recommendation of items (services)
based on the pro le. The recommendation's main function
is to aggregate content from di erent sources and mobile
Web portal and to customize the presentation of services
for each user according to his pro le. It allows classi cation
or restriction of services into a selection that ts the user
pro le.</p>
      <p>The remainder of this paper is organized as follows.
Section 2 describes the context of the work presented in this
paper. In Section 3 , we present the global architecture
of the SoliMobile Project. Section 4 details the analysis of
existing Open Source Recommender Systems. The
recommender system used in the project is explained in section 5.
Section 6 presents the utility of web usage mining in the
recommendation. Finally, Section 7 concludes this paper and
gives an outlook upon our ongoing and future research in
this area.</p>
    </sec>
    <sec id="sec-3">
      <title>WORK CONTEXT</title>
      <p>The work presented in this paper ts in a collaborative
project that aims to design, develop and implement a set of
innovative services focused on persons in situations of
instability or emerging from instability, in order to help them to
nd useful information such as jobs, o ers of housing,
welfare, or medical assistance. The charity association partner
in the project observed that a large majority of people in
unstable situation own a mobile phone that is considered
as a link with family, friends or society. The project aims
to facilitate, for the vulnerable people, processes to access
charities services using their mobile phone from anywhere.
However, di erent services are not suitable for all persons in
a precarious situation. For example, a single mother needs
child services such as pediatrics or nursery while an
unemployed needs services to nd a job or professional training.</p>
      <p>Our role in this project is to enhance and customize
services to users. In this context, we deal with the
implementation of algorithms to adapt to user pro le the platform
services, to lter them and to show only items that may be of
interest. Personnalization according to user pro le is based
both on data available on the platform (eg. databases), the
features and traces of user navigation, and also the social
environment of the user (collaborative ltering approach).</p>
      <p>Typically, adaptation to the user pro le will consolidate
the resources (services) to target only the relevant users.
Conversely, user pro les will also be ordered to form
homogeneous groups in order to assign them to a given resource.
3.</p>
    </sec>
    <sec id="sec-4">
      <title>GLOBAL ARCHITECTURE</title>
      <p>We present in this section the overall architecture of the
application, illustrated in Figure 1, in order to show the role
of the recommendation in the SoliMobile platform. In fact,
end users create an account via the Web platform where
many services are provided. The traces of Web browsing
(also called logs) are collected from servers to feed the
recommender system. These navigation traces will be used to
create the user item ratings matrix. Services play the role
of items. Information regarding the user pro le such as age,
address, occupationo or preferences as well as information
concerning the description of services such as the category
of services (health, employment, child care, etc..) and their
addresses will be provided as input of recommender system.
These inputs will be sent in XML format through Web
services. Once the ratings matrix constructed, the
recommendation is made to categorize and customize the layout of
proposed services on the mobile phone. The
recommendation system will provide as output an XML le that contains
a subset of sorted services to be transmitted to the mobile.
Traces of mobile browsing will also be used as input to the
recommender system to improve results, they can also serve
as a feedback to our system.</p>
      <p>Our goal is structured along the following lines:
Construct a generic model for the user pro le and also
for structural and semantic information of the
application in order to integrate new data when needed;
Select, automatically and dynamically, variables
describing the user, the services and the log navigation
that improve the quality of the recommendation;
Ensure the proper functioning of the recommender
system in case of registration of a new user whose pro le
is poor (or nonexistent) or in case of creating new
services (items) that no one (in our data set) has yet rated
or visited. This problem is well known in the eld of
information ltering and is referred as "Cold Start"
problem. Almost solutions for the cold-start problem
[Lam et al. 2008] are not suitable as they involve users
to rate items.</p>
      <p>
        Develop a generic recommender system, i.e. that adapts
to any application. The challenge is to design a
realtime recommender system that lters resources
dynamically depending on variation in user interests but
also on variation in the environment. The idea is to
associate with each resource a ranking based on the
user pro le and its context. We use for that
incremental learning techniques [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and mining data streams [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]
that requires a limited number of passes on data and
needs to process data on the y. Using these methods
improves computation time and memory space so we
can ensure robustness and scalability of the system;
De ne satisfactory indicators in order to assess the
quality of the recommendation;
Conduct a software platform integrating all the tools
developed during the project.
      </p>
      <p>Given the short duration of the project (18 months),
we decided to study open source recommender
systems. Thus, we present in the following section the
related state of the art.
4.</p>
    </sec>
    <sec id="sec-5">
      <title>OPEN-SOURCE RECOMMENDER SYS</title>
    </sec>
    <sec id="sec-6">
      <title>TEMS</title>
      <p>
        The growth of Web content and the expansion of e-commerce
has deeply increased the interest on recommender systems.
This fact has led to the development of some open source
projects in the area. Among the recommender systems
algorithms available in the Web, we can distinguish the
following: Duine [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], Apache Mahout [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], OpenSlopeOne [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], Co
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], SUGGEST [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] and Vogoo [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. All of these projects
offer collaborative- ltering implementations, in di erent
programming languages.
      </p>
      <p>The Duine Framework supplies also a hybrid
implementation. It is a Java software that presents the content-based
and collaborative ltering in a switching engine: it
dynamically switches between each prediction given the current
state of the data. For example if there aren't many ratings
available, it uses the content-based approach, and switches
to the collaborative when the scenario changes. It also
presents an Explanation API, which can be used to
create user-friendly recommendations and a demo application,
with a Java Client example.</p>
      <p>Mahout constitutes a Java framework in the data mining
area. It has incorporated the Taste recommender system, a
collaborative engine for personalized recommendations.
Vogoo is a PHP framework that implements a collaborative
ltering recommender system. It also presents a Slope-One
code.</p>
      <p>
        A Java version of the Collaborative Filtering method is
implemented in the Co library. It was developed by Daniel
Lemire [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], the creator of the Slope-One algorithms. There is
also a PHP version available in Lemire's webpage.
OpenSlopeOne o ers a Slope One implementation on PHP that cares
about performance.
      </p>
      <p>SUGGEST is a recommendation library made by George
Karkys and distributed in a binary format.</p>
      <p>Analyzing software in the recommendation area is not a
simple task, since it is di cult to de ne measurement
standards. In this work, we propose some criteria of evaluation:
types of recommendation implemented by the project,
programming language, level of documentation and magnitude
of the project.</p>
      <p>The documentation was evaluated based on its volume
and clarity. It is possible to observe that the volume of
documentation presented by Mahout and Duine is remarkably
larger than the other systems. Both o er installation and
utilization guides and come with a demonstration example.
It must be taken into account that OpenSlopeOne and Co
are smaller projects, and thus, their documentation tend to
be smaller. In the Downloads column we have a
representation of the magnitude of the project. It is presented the
number of times the software, in any version, was
downloaded from its source. Although Mahout does not present
its number, its very populated mailing lists shows that it is
a widely used software.</p>
      <p>The two projects that stood out were Apache Mahout and
Duine. We installed and tested them in order to verify which
one was more applicable to our work. Both of them are
based on the Java technology and present a demonstration
example with the Movielens data set. The fact that Mahout
is a greater project and has multiple machine-learning
algorithms made it more interesting to our research. Also, its
module structure encouraged us to choose it.</p>
    </sec>
    <sec id="sec-7">
      <title>APACHE MAHOUT</title>
      <p>The Apache Mahout is a solid project in the Data
Mining area. It is a framework that features various scalable
machine-learning algorithms. It is programmed using the
Java language and runs with Maven project manager. In
April 2008, it has incorporated the Taste Recommender
System, a Java framework for providing personalized
recommendations. Besides Taste, it also o ers clustering
algorithms and a Map Reduce implementation.</p>
      <p>Taste is a very consistent and exible collaborative
ltering engine and supports the user-based, item-based and
Slope-one recommender systems. It can be easily modi ed
due to its well-structured modules abstractions. The
package de nes four interfaces: DataModel, UserSimilarity and
ItemSimilarity, UserNeighborhood and Recommender.</p>
      <p>With these interfaces, it is possible to adapt the
framework to read di erent types of data, personalize the
recommendation or even create new recommendation methods.</p>
      <p>The User Similarity and Item Similarity abstractions are
responsible for calculating the similarity between a pair of
users or items. Their function usually returns a value from
0 to 1 indicating the level of resemblance, being 1 the most
similar possible.</p>
      <p>Trough the DataModel interface is made the access to the
data set. It is possible to retrieve and store the data from
databases or from lesystems (MySQLJDBCDataModel and
FileDataModel respectively). The functions developed in
this interface are used by the Similarity abstraction to help
computing the similarity.</p>
      <p>The main interface in Taste is Recommender. It is
responsible for actually making the recommendations to the user
by comparing items or by determining users with similar
taste (item-based and user-based techniques). The
Recommender access the similarity interface and uses its functions
to compare a pair of users or items. It then collects the
highest similarity values to o er as recommendations.</p>
      <p>
        The UserNeighborhood is an assistant interface that helps
to de ne the neighborhood in the user-based
recommendation technique. It is know that, for greater data sets,
the item-based technique provides better results. For that,
many companies choose to use this approach, such as
Amazon [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. With the Mahout framework, it is not di erent, the
item-based method generally runs faster and provides more
accurate recommendation.
      </p>
      <p>In our project, we choose to adapt the Slope One (a type
of item-based algorithm) approach to our problem. Here
follows a simple Java application example of how to initiate
a recommendation with the Slope One technique:
1. DataModel model =</p>
      <p>new FileDataModel(new File("data.txt"));
2. Recommender recommender =</p>
      <p>new SlopeOneRecommender(model);
3. Recommender cachingRecommender =</p>
      <p>new CachingRecommender(recommender);</p>
      <p>The challenge in adapting this approach to our project
was the fact that our input data le was available in the
XML format, a type not handled by Mahout. It then had
to incorporate another le in the DataModel interface. We
create a program that deals with the XML input les. To
test this new data handler, we used the Movielens data set.
A pack with one million ratings was converted to the XML
type to be used as example. With this data set and the
XML le, the running time of the Slope One algorithm takes
less than one minute.
6.</p>
    </sec>
    <sec id="sec-8">
      <title>WEB USAGE MINING FOR RECOMMEN</title>
    </sec>
    <sec id="sec-9">
      <title>DATION</title>
      <p>
        One objective of the SoliMobile project is to develop a
recommender system that has to be, as much as possible, the
least intrusive. This implies that the system is based only
on information that the user can be free to provide (explicit
data) and must run properly with alternatives such as
implicit data mining. To meet this need, we are studying how
to append Web browsing analysis to the recommender
system as done in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Web browsing analysis becomes almost
necessary for extracting and understanding user behaviors.
      </p>
      <sec id="sec-9-1">
        <title>Language Java Java Java</title>
        <p>PHP
PHP
C</p>
      </sec>
      <sec id="sec-9-2">
        <title>Documentation High High Low</title>
        <p>Low
Medium
Medium</p>
      </sec>
      <sec id="sec-9-3">
        <title>Downloads</title>
        <p>Not available</p>
        <p>1,113
Not available
653
2,128
Not available
In recent years, Web usage mining has become an important
issue in the eld of data mining. The term, Web usage
mining focuses on predicting and learning the users preferences
on the Internet. Generally, the data for Web usage mining
are the user interactions on the web, usually residing on Web
clients, Web servers, and proxy servers. The aim of Web
usage mining is to analyze user behavior through analysis of
its interaction with the Web platform. This analysis is
particularly focused on all the users clicks where visiting the
web application (also known as clickstream analysis). The
interest of Web usage mining in our framework is to enrich
the input of recommender system with user data extracted
from the raw clickstream data, in order to re ne the user
pro les and behavioral patterns. The analysis of Web logs
can also be used as implicit feedback of the user which will
allow to assess the performance of models involved in the
recommender system.</p>
        <p>
          It is obvious that Web logs change over time for several
reasons: an update of the Web application content or
structure, a change in the user preferences, a change in the
execution context, etc. This is why it is important to take into
account the temporal dimension in the analysis of Web usage.
To consider the temporal data in a dynamic way, we plan
to use the techniques of data streams mining. By de nition,
data stream is a real-time, continuous, ordered (implicitly by
arrival time or explicitly by timestamp) sequence of items.
It is impossible to control the order in which items arrive,
nor is it feasible to locally store a stream in its entirety.
Therefore, all the treatment have to be applied in one pass.
Several techniques for mining data streams have emerged as
CluStream for clustering, StreamSamp for sampling, VFDT
for incremental decision trees, etc. The reader may refer to
[
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] for more explanations on these di erent techniques.
        </p>
      </sec>
    </sec>
    <sec id="sec-10">
      <title>CONCLUSIONS</title>
      <p>In this paper, we presented the problem that we deal with
in the SoliMobile project. Then, we presented the global
architecture that is under development in this project. This
architecture includes a recommender system to customize
the services o ered to users based on their pro le and their
browsing history. Given the limited duration of the project,
we opted for an open source recommender system that is
modular in order to easily integrate future developments,
in particular the use of Web usage mining to address the
problem of cold start.</p>
      <p>In this paper, we also discussed several points concerning
the issue of treatment of the temporal dimension in data
analysis. The raised issues demonstrate the need for de
ning new methods or adapting existing methods for
extracting knowledge and monitoring changing and evolutive data.
Although there are many e cient methods for extracting
knowledge, few studies have been devoted to the issue of
temporary evolutive data.
8.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>John</surname>
            <given-names>S. Breese</given-names>
          </string-name>
          , John S. Breese, David Heckerman,
          <article-title>and Carl" Kadie. Empirical analysis of predictive algorithms for collaborative ltering</article-title>
          .
          <source>pages 43{52</source>
          ,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>[2] Co . http://www.nongnu.org/co /.</mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Antoine</given-names>
            <surname>Cornuejols</surname>
          </string-name>
          .
          <article-title>Getting order independence in incremental learning</article-title>
          .
          <source>In ECML '93: Proceedings of the European Conference on Machine Learning</source>
          , pages
          <volume>196</volume>
          {
          <fpage>212</fpage>
          , London, UK,
          <year>1993</year>
          . Springer-Verlag.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Baptiste</given-names>
            <surname>Csernel</surname>
          </string-name>
          , Fabrice Clerot, and
          <string-name>
            <given-names>Georges</given-names>
            <surname>Hebrail</surname>
          </string-name>
          . Streamsamp:
          <article-title>Datastream clustering over tilted windows through sampling</article-title>
          .
          <source>ECML PKDD</source>
          <year>2006</year>
          :
          <article-title>the International Workshop on Knowledge Discovery from Data Streams (IWKDDS-</article-title>
          <year>2006</year>
          ),
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Duine</surname>
          </string-name>
          . http://www.duineframework.org/.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Daniel</given-names>
            <surname>Lemire</surname>
          </string-name>
          and
          <article-title>Anna Maclachlan "</article-title>
          .
          <article-title>Slope one predictors for online rating-based collaborative ltering</article-title>
          .
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Greg</given-names>
            <surname>Linden</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Brent</given-names>
            <surname>Smith</surname>
          </string-name>
          , and Jeremy York. Amazon.
          <article-title>com recommendations: Item-to-item collaborative ltering</article-title>
          .
          <source>IEEE Internet Computing</source>
          ,
          <volume>7</volume>
          (
          <issue>1</issue>
          ):
          <volume>76</volume>
          {
          <fpage>80</fpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Jiahui</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Peter</given-names>
            <surname>Dolan</surname>
          </string-name>
          , and
          <string-name>
            <surname>Elin R nby Pedersen</surname>
          </string-name>
          .
          <article-title>Personalized news recommendation based on click behavior</article-title>
          .
          <source>In IUI '10: Proceeding of the 14th international conference on Intelligent user interfaces</source>
          ,
          <source>pages</source>
          <volume>31</volume>
          {
          <fpage>40</fpage>
          , New York, NY, USA,
          <year>2010</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Mahout</surname>
          </string-name>
          . http://mahout.apache.org/.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Raymond</surname>
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Mooney</surname>
            and
            <given-names>Loriene</given-names>
          </string-name>
          <string-name>
            <surname>Roy</surname>
          </string-name>
          .
          <article-title>Content-based book recommending using learning for text categorization</article-title>
          .
          <source>In DL '00: Proceedings of the fth ACM conference on Digital libraries</source>
          , pages
          <volume>195</volume>
          {
          <fpage>204</fpage>
          , New York, NY, USA,
          <year>2000</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>[11] OpenSlopeOne. http://code.google.com/p/openslopeone/.</mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>[12] ProximaMobile. http://www.proximamobile.fr/.</mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Suggest</surname>
          </string-name>
          . http://glaros.dtc.umn.edu/gkhome/suggest/overview.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Vogoo</surname>
          </string-name>
          . http://www.vogoo-api.com/.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>