<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>2nd Annual International Symposium on Information Management and Big Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Conclusion</string-name>
        </contrib>
      </contrib-group>
      <fpage>212</fpage>
      <lpage>223</lpage>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>PROCEEDINGS
SIMBig 2015 - Organizing Committee
SIMBig 2015 - Program Committee
SIMBig 2015 - List of Contributions
SIMBig 2015 - Keynote Speaker Presentations
SIMBig 2015 - Sessions, Papers
The SIMBig 2015 Organizing Committee con rms that full and concise papers accepted
for this publication:
• Meet the de nition of research in relation to creativity, originality, and increasing
humanity's stock of knowledge;
• Are selected on the basis of a peer review process that is independent, quali ed
expert review;
• Are published and presented at a conference having national and international
signi cance as evidenced by registrations and participation; and
• Are made available widely through the Conference web site.</p>
      <p>Disclaimer: The SIMBig 2015 Organizing Committee accepts no responsibility for omissions and errors.
Organizing Committee of SIMBig 2015
General Organizers
• Juan Antonio</p>
      <p>LIRMM, FRANCE</p>
      <p>LOSSIO-VENTURA, University of Montpellier
• Hugo ALATRISTA-SALAS, Ponti cia Universidad Catolica del Peru</p>
      <p>GRPIAA Labs, PERU</p>
    </sec>
    <sec id="sec-2">
      <title>Local Organizers</title>
      <p>• Cristhian GANVINI VALCARCEL Universidad Andina del Cusco, PERU
• Armando FERMIN PEREZ, Universidad Nacional Mayor de San Marcos,</p>
      <p>PERU</p>
    </sec>
    <sec id="sec-3">
      <title>Track Organizers:</title>
      <p>WTI (Web and Text Intelligence Track)
• Jorge Carlos VALVERDE-REBAZA, University of Sa~o Paulo, Brazil
• Alneu DE ANDRADE LOPES, ICMC { University of S~ao Paulo, Brazil
• Ricardo Bastos CAVALCANTE PRUDENCIO, Federal University of
Pernambuco, Brazil
• Estevam HRUSCHKA, Federal University of Sa~o Carlos, Brazil
• Ricardo CAMPOS, Polytechnic Institute of Tomar - LIAAD{INESC
Technology and Science, Portugal
SIMBig 2015 Program Committee
• Nathalie Abadie, French National Mapping Agency, COGIT, FRANCE
• Elie Abi-Lahoud, University College Cork, Cork, IRELAND
• Salah AitMokhtar, Xerox Research Centre Europa, FRANCE
• Sophia Ananiadou, NaCTeM - University of Manchester, UNITED
KING</p>
      <p>DOM
• Marcelo Arenas, Ponti cia Universidad Catolica de Chile, CHILE
• Jero^me Aze, LIRMM - University of Montpellier, FRANCE
• Pablo Barcelo, Universidad de Chile, CHILE
• Cesar A. Beltran Castan~on, GRPIAA - Ponti cal Catholic University of</p>
      <p>Peru, PERU
• Albert Bifet, Huawei Noah's Ark Research Lab, Hong Kong, CHINA
• Sandra Bringay, LIRMM - Paul Valery University, FRANCE
• Oscar Corcho, Ontology Engineering Group - Polytechnic University of</p>
      <p>Madrid, SPAIN
• Gabriela Csurka, Xerox Research Centre Europa, FRANCE
• Frederic Flouvat, PPME Lab - University of New Caledonia, NEW
CALE</p>
      <p>DONIA
• Andre Freitas, Dept. of Computer Science and Mathematics, University
of Passau, GERMANY
• Adrien Guille, ERIC Lab - University of Lyon 2, FRANCE
• Hakim Hacid, Zayed University, UNITED ARAB EMIRATES
• Sebastien Harispe, LGI2P/EMA Research Centre, Site EERIE, Parc
Scienti que, FRANCE
• Dino Ienco, Irstea, FRANCE
• Diana Inkpen, University of Ottawa, CANADA
• Clement Jonquet, LIRMM - University of Montpellier, FRANCE
• Al pio Jorge, Universidade do Porto, PORTUGAL
• Yannis Korkontzelos, NaCTeM - University of Manchester, UNITED
KINGDOM
• Eric Kergosien, GERiCO Lab - University of Lille 3, FRANCE
• Peter Mika, Yahoo! Research Labs - Barcelone, SPAIN
• Phan Nhat Hai, Oregon State University, UNITED STATES of AMERICA
• Jordi Nin, Barcelona Supercomputing Center (BSC) - BarcelonaTECH,</p>
      <p>SPAIN
• Miguel Nun~ez del Prado Cortez, Intersec Lab - Paris, FRANCE
• Thomas Opitz, Biostatistics and Spatial Processes - INRA, FRANCE
• Yoann Pitarch, IRIT - Toulouse, FRANCE
• Pascal Poncelet, LIRMM - University of Montpellier, FRANCE
• Julien Rabatel, LIRMM, FRANCE
• Jose Luis Redondo Garc a, EURECOM, FRANCE
• Mathieu Roche, Cirad - TETIS - LIRMM, FRANCE
• Nancy Rodriguez, LIRMM - University of Montpellier, FRANCE
• Arnaud Sallaberry, LIRMM - Paul Valery University, FRANCE
• Nazha Selmaoui-Folcher, PPME Labs - University of New Caledonia,</p>
      <p>NEW CALEDONIA
• Maguelonne Teisseire, Irstea - LIRMM, FRANCE
• Paulo Teles, LIAAD-INESC Porto LA, Porto University, PORTUGAL
• Julien Velcin, ERIC Lab - University of Lyon 2, FRANCE
• Maria-Esther Vidal, Universidad Simon Bol var, VENEZUELA
• Boris Villazon-Terrazas, Expert System Iberia - Madrid, SPAIN
• Osmar R. Zaane, Department of Computing Science, University of
Alberta, CANADA
WTI 2015 Program Committee (Web and Text Intelligence Track)
• Adam Jatowt, Kyoto University, Japan,
• Claudia Orellana Rodriguez, Insight Centre for Data Analytics -
University College Dublin, Ireland,
• Ernesto Diaz-Aviles, IBM Research, Ireland,
• Jannik Strotgen, Heidelberg University, Germany,
• Jo~ao Paulo Cordeiro, University of Beira Interior, Portugal,
• Lucas Drumond, University of Hildesheim, Germany,
• Leon Derczynski, University of She eld, UK,
• Miguel Martinez-Alvarez, Signal, UK,
• Brett Drury, USP, Brazil,
• Celso Anto^nio Alves Kaestner, UTFPR, Brazil,
• Edson Matsubara, UFMS, Brazil,
• Flavia Barros, UFPe, Brazil,
• Hercules Antonio do Prado, PUC Bras lia, Brazil,
• Huei Lee, UNIOESTE, Brazil,
• Jo~ao Lu s Rosa, USP, Brazil,
• Jose Lu s Borges, Universidade do Porto, Portugal,
• Jose Paulo Leal, Universidade do Porto, Portugal,
• Maria Cristina Ferreira de Oliveira, USP, Brazil,
• Ronaldo Prati, UFABC, Brazil,
• Thiago A. S. Pardo, USP, Brazil,
• Solange Rezende, USP, Brazil,
• Marcos Aurelio Domingues, USP, Brazil.
List of contributions
1
2</p>
      <p>Overview of SIMBig 2015 10
Overview of SIMBig 2015: 2nd Annual International Symposium
on Information Management and Big Data
spaJuan Antonio Lossio-Ventura and Hugo Alatrista-Salas . . . . . . . 10</p>
      <p>Overview of SIMBig 2015: 2nd Annual International
Symposium on Information Management and Big Data
Juan Antonio Lossio-Ventura Hugo Alatrista-Salas
LIRMM, University of Montpellier Pontificia Universidad Cato´lica del Peru´</p>
      <p>Montpellier, France Lima, Peru
juan.lossio@lirmm.fr halatrista@pucp.pe
Big Data is a popular term used to
describe the exponential growth and
availability of both structured and unstructured
data. The aim of the symposium is to
present the analysis methods for managing
large volumes of data through techniques
of artificial intelligence and data mining.</p>
      <p>Bringing together main national and
international actors in the decision-making field
to state in new technologies dedicated to
handle large amount of information.
1
Big Data is a popular term used to describe the
exponential growth and availability of both
structured and unstructured data. This has taken place
over the last 20 years. For instance, social
networks such as Facebook, Twitter and Linkedin
generate masses of data, which is available to be
accessed by other applications. Several domains,
including biomedicine, life sciences and scientific
research, have been affected by Big Data1. Therefore
there is a need to understand and exploit this data.</p>
      <p>This process can be carried out thanks to “Big
Data Analytics” methodologies, which are based
on Data Mining, Natural Language Processing,
etc. That allows us to gain new insight through
data-driven research (Madden, 2012; Embley and
Liddle, 2013). A major problem hampering Big
Data Analytics development is the need to process
several types of data, such as structured, numeric
and unstructured data (e.g. video, audio, text,
image, etc)2.</p>
      <p>Therefore, the second edition of the Annual
International Symposium on Information
Management and Big Data - SIMBig 20153, aims to present
the analysis methods for managing large volumes
of data through techniques of artificial intelligence
and data mining. Counting with main national and</p>
      <p>1By 2015 the average of data annually generated in
hospitals is 665TB: http://ihealthtran.com/wordpress/2013/
03/infographic-friday-the-body-as-a-source-of-big-data/.</p>
      <p>2Today, 80% of data is unstructured such as images,
video, and notes
3http://simbig.org/SIMBig2015/
international actors in the decision-making field to
state in new technologies dedicated to handle large
amount of information.</p>
      <p>Our first edition, SIMBig 20144 took place in
Cuzco Peru too in September 2015. SIMBig 2014
has been indexed on DBLP5 (Lossio-Ventura and
Alatrista-Salas, 2014) and on CEUR Workshop
Proceedings6.
1.1</p>
      <sec id="sec-3-1">
        <title>Keynote Speakers</title>
        <p>SIMBig 2015 second edition has welcomed five
keynote speakers experts in Big Data, Data
Mining, Natural Language Processing (NLP), and
Social Networks:
• PhD. Pr. Albert Bifet, from HUAWEI</p>
        <p>Noah’s Ark Lab, China;
• PhD. Pr. Diana Inkpen, from University of</p>
        <p>Ottawa, Canada;
• PhD. Pr. Pascal Poncelet, from LIRMM
Laboratory and University of Montpellier,
France;
• PhD. Pr. Mathieu Roche, from Cirad and</p>
        <p>TETIS Laboratory, France;
• PhD. Pr. Osmar R. Za¨ıane, from University
of Alberta, Canada.
1.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Scope and Topics</title>
        <p>To share the new analysis methods for managing
large volumes of data, we encouraged participation
from researchers in all fields related to Big Data,
Data Mining, and Natural Language Processing,
but also Multilingual Text Processing, Biomedical
NLP. Topics of interest of SIMBig 2015 included
but were not limited to:
• Big Data
• Data Mining
• Natural Language Processing
4https://www.lirmm.fr/simbig2014/
5http://dblp2.uni-trier.de/db/conf/simbig/
simbig2014
6http://ceur-ws.org/Vol-1318/index.html
• Bio NLP
• Text Mining
• Information Retrieval
• Machine Learning
• Semantic Web
• Ontologies
• Web Mining
• Knowledge Representation and Linked Open</p>
        <p>Data
• Social Networks, Social Web, and Web Science
• Information visualization
• OLAP, Data Warehousing
• Business Intelligence
• Spatiotemporal Data
• Health Care
• Agent-based Systems
• Reasoning and Logic
• Constraints, Satisfiability, and Search
2</p>
        <sec id="sec-3-2-1">
          <title>Latin American and Peruvian</title>
        </sec>
        <sec id="sec-3-2-2">
          <title>Academic Goals of the</title>
        </sec>
        <sec id="sec-3-2-3">
          <title>Symposium</title>
          <p>The academic goals of the symposium are varied,
among which we can list the following:
• Meet Latin American and foreign researchers,
teachers, and students belonging to several
domains of computer sciences, specially related
to Big Data.
• Promote the production of scientific articles,
which will be evaluated by the international
scientific community, in order to receive a
feedback from experts.
• Foster partnerships between Latin American
universities, local universities and European
universities.
• Promote the creation of alliances between
Peruvian universities, enabling decentralization
of education.
• Motivate students to learn more about
computer sciences research to solve problems
related to the management of information and
Big Data.
• Promote the research in Peruvian universities,
mainly those belonging to the local organizing
committee.
• Create connections, forming networks of
partnerships between companies and universities.
• Promote the local and international tourism,
in order to show to the participants the
architecture, gastronomy and local cultural
heritage.
3</p>
        </sec>
        <sec id="sec-3-2-4">
          <title>Track on Web and Text</title>
        </sec>
        <sec id="sec-3-2-5">
          <title>Intelligence (WTI 2015)</title>
          <p>Web and text intelligence are related areas that
have been used to improve human computer
interaction both in general and in particular to
explore and analyze information that is available on
the Internet. With the advent of social networks
and the emergence of services such as Facebook,
Twitter, and others, research in these areas has
been greatly enhanced. In recent years, shared
knowledge and experiences have established new
and different types of personal and communal
relationships which have been leveraged by social
networks scientists to produce new insights. In
addition there has been a huge increase in community
activities on social networks.</p>
          <p>The Web and Text Intelligence (WTI) track of
SIMBig 2015 have provided a forum that brought
together researchers and practitioners for
exploring technologies, issues, experiences and
applications that help us to understand the Web and to
build automatic tools to better exploit this
complex environment. The WTI track has fostered
collaborations, exchange of ideas and experiences
among people working in a variety of highly
crossdisciplinary research fields such as computer
science, linguistics, statistics, sociology, economics,
and business.</p>
          <p>The WTI track is a follow up of the 4th
International Workshop on Web and Text Intelligence7,
which took place in Curitiba, Brazil, October 2012,
as a workshop of BRACIS 2012; the 3rd
International Workshop on Web and Text Intelligence8,
which took place in S˜ao Bernardo, Brazil, October
2010, as a workshop of SBIA10; the 2nd
International Workshop on Web and Text Intelligence9,
which took place in S˜ao Carlos, Brazil, September
2009, as a workshop of STIL09; the 1st Web and
Network Intelligence10, which took place in Aveiro,
Portugal, October 2009, as a thematic track of
EPIA09; and the 1st International Workshop on
Web and Text Intelligence, which took place in
7http://www.labic.icmc.usp.br/wti2012/
8http://www.labic.icmc.usp.br/wti2010/
9http://www.labic.icmc.usp.br/wti2009/
10http://epia2009.web.ua.pt/wni/
Salvador, Brazil, October 2008, as a workshop of
SBIA08.
We want to thank our wonderful sponsors! We
extend our sincere appreciation to our sponsors,
without whom our symposium would not be
possible. They showed their commitment to making
our research communities more active. We invite
you to support these community-minded
organizations.
4.1</p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>Organizing Institutions</title>
        <p>• Universit´e de Montpellier, France11
• Laboratoire de Informatique, Robotique et</p>
        <p>Micro´electronique de Montpellier, France12
• Universidad Andina del Cusco, Peru´13
11http://www.umontpellier.fr/
12http://www.lirmm.fr/
13http://www.uandina.edu.pe/
4.2</p>
        <p>Collaborating Institutions
• iMedia14
• Bioincuba15
• TechnoPark16
• Grupo de Reconocimiento de Patrones e
In</p>
        <p>teligencia Artificial Aplicada, PUCP, Peru´17
• Universidad Nacional Mayor de San Marcos,</p>
        <p>Peru´18
• Escuela de Post-grado de la Pontificia
Univer</p>
        <p>sidad Cat´olica del Peru´19
14http://www.imedia.pe/
15http://www.bioincuba.com/
16http://technoparkidi.org/
17http://inform.pucp.edu.pe/~grpiaa/
18http://www.unmsm.edu.pe/
19http://posgrado.pucp.edu.pe/la-escuela/
presentacion/
20http://www.icmc.usp.br/Portal/
21http://portal2.ipt.pt/
22http://labic.icmc.usp.br/
23http://www.inesctec.pt/liaad
24http://ppgcc.dc.ufscar.br/pesquisa/
laboratorios-e-grupos-de-pesquisa</p>
        <p>25http://www2.ufscar.br/home/index.php
Real-Time Big Data Stream Analytics
Big Data is a new term used to identify
datasets that we cannot manage with
current methodologies or data mining
software tools due to their large size and
complexity. Big Data mining is the
capability of extracting useful information from
these large datasets or streams of data.</p>
        <p>New mining techniques are necessary due
to the volume, variability, and velocity,
of such data. MOA is a software
framework with classification, regression, and
frequent pattern methods, and the new
APACHE SAMOA is a distributed
streaming software for mining data streams.
1 Introduction
Big Data is a new term used to identify the datasets
that due to their large size, we can not manage
them with the typical data mining software tools.</p>
        <p>Instead of defining “Big Data” as datasets of a
concrete large size, for example in the order of
magnitude of petabytes, the definition is related to the
fact that the dataset is too big to be managed
without using new algorithms or technologies. There
is need for new algorithms, and new tools to deal
with all of this data. Doug Laney (Laney, 2001)
was the first to mention the 3 V’s of Big Data
management:
• Volume: there is more data than ever before,
its size continues increasing, but not the
percent of data that our tools can process
• Variety: there are many different types of
data, as text, sensor data, audio, video, graph,
and more
• Velocity: data is arriving continuously as
streams of data, and we are interested in
obtaining useful information from it in real time</p>
        <p>Nowadays, there are two more V’s:
• Variability: there are changes in the structure
of the data and how users want to interpret
that data
• Value: business value that gives
organizations a competitive advantage, due to the
ability of making decisions based in answering
questions that were previously considered
beyond reach</p>
        <p>For velocity, data stream real time analytics are
needed to manage the data currently generated,
at an ever increasing rate, from such applications
as: sensor networks, measurements in network
monitoring and traffic management, log records
or click-streams in web exploring,
manufacturing processes, call detail records, email, blogging,
twitter posts and others. In fact, all data generated
can be considered as streaming data or as a
snapshot of streaming data, since it is obtained from an
interval of time.</p>
        <p>In the data stream model, data arrive at high
speed, and algorithms that process them must
do so under very strict constraints of space and
time. Consequently, data streams pose several
challenges for data mining algorithm design. First,
algorithms must make use of limited resources
(time and memory). Second, they must deal with
data whose nature or distribution changes over
time.</p>
        <p>We need to deal with resources in an efficient
and low-cost way. In data stream mining, we are
interested in three main dimensions:
• accuracy
• amount of space necessary
• the time required to learn from training
ex</p>
        <p>amples and to predict</p>
        <p>These dimensions are typically interdependent:
adjusting the time and space used by an
algorithm can influence accuracy. By storing more
precomputed information, such as look up tables, an
algorithm can run faster at the expense of space.</p>
        <p>An algorithm can also run faster by processing
less information, either by stopping early or
storing less, thus having less data to process.
2</p>
        <p>MOA
Massive Online Analysis (MOA) (Bifet et al.,
2010) is a software environment for
implementing algorithms and running experiments for
online learning from evolving data streams. MOA
includes a collection of offline and online
methods as well as tools for evaluation. In particular,
it implements boosting, bagging, and Hoeffding
Trees, all with and without Na¨ıve Bayes
classifiers at the leaves. Also it implements regression,
and frequent pattern methods. MOA supports
bidirectional interaction with WEKA, the Waikato
Environment for Knowledge Analysis, and is
released under the GNU GPL license.
3</p>
        <p>APACHE SAMOA
APACHE SAMOA (SCALABLE ADVANCED
MASSIVE ONLINE ANALYSIS) is a platform for
mining big data streams (Morales and Bifet, 2015). As
most of the rest of the big data ecosystem, it is
written in Java.</p>
        <p>APACHE SAMOA is both a framework and a
library. As a framework, it allows the algorithm
developer to abstract from the underlying execution
engine, and therefore reuse their code on
different engines. It features a pluggable architecture
that allows it to run on several distributed stream
processing engines such as Storm, S4, and Samza.</p>
        <p>This capability is achieved by designing a minimal
API that captures the essence of modern DSPEs.</p>
        <p>This API also allows to easily write new bindings
to port APACHE SAMOA to new execution engines.</p>
        <p>APACHE SAMOA takes care of hiding the
differences of the underlying DSPEs in terms of API
and deployment.</p>
        <p>As a library, APACHE SAMOA contains
implementations of state-of-the-art algorithms for
distributed machine learning on streams. For
classification, APACHE SAMOA provides a Vertical
Hoeffding Tree (VHT), a distributed streaming
version of a decision tree. For clustering, it includes
an algorithm based on CluStream. For regression,</p>
        <p>HAMR, a distributed implementation of Adaptive
Model Rules. The library also includes
metaalgorithms such as bagging and boosting.</p>
        <p>The platform is intended to be useful for both
research and real world deployments.
3.1</p>
        <p>High Level Architecture
We identify three types of APACHE SAMOA users:
1. Platform users, who use available ML
algo</p>
        <p>rithms without implementing new ones.
2. ML developers, who develop new ML
algorithms on top of APACHE SAMOA and want
to be isolated from changes in the underlying</p>
        <p>SPEs.
3. Platform developers, who extend APACHE</p>
        <p>SAMOA to integrate more DSPEs into</p>
        <p>APACHE SAMOA.
4</p>
        <p>Conclusions
Big Data Mining is a challenging task, that needs
new tools to perform the most common machine
learning algorithms such as classification,
clustering, and regression.</p>
        <p>APACHE SAMOA is is a platform for
mining big data streams, and it is already
available and can be found online at http://www.
samoa-project.net. The website includes a
wiki, an API reference, and a developer’s manual.</p>
        <p>Several examples of how the software can be used
are also available.</p>
        <p>Acknowledgments
The presented work has been done in collaboration
with Gianmarco De Francisci Morales, Bernhard
Pfahringer, Geoff Holmes, Richard Kirkby, and all
the contributors to MOA and APACHE SAMOA.</p>
        <p>Detecting Locations from Twitter Messages</p>
        <p>Invited Talk</p>
        <p>Diana Inkpen</p>
        <p>University of Ottawa</p>
        <p>School of Electrical Engineering and Computer Science
800 King Edward Avenue, Ottawa, ON, K1N6N5, Canada</p>
        <p>Diana.Inkpen@uottawa.ca
1</p>
        <p>Extended Abstract
There is a large amount of information that can
be extracted automatically from social media
messages. Of particular interest are the topics
discussed by the users, the opinions and
emotions expressed, and the events and the locations
mentioned. This work focuses on machine
learning methods for detecting locations from Twitter
messages, because the extracted locations can be
useful in business, marketing and defence
applications (Farzindar and Inkpen, 2015).</p>
        <p>There are two types of locations that we are
interested in: location entities mentioned in the text
of each message and the physical locations of the
users. For the first type of locations (task 1), we
detected expressions that denote locations and
we classified them into names of cities,
provinces/states, and countries. We approached the task
in a novel way, consisting in two stages. In the
first stage, we trained Conditional Random Field
models with various sets of features. We
collected and annotated our own dataset for training and
testing. In the second stage, we resolved cases
when more than one place with the same name
exists, by applying a set of heuristics (Inkpen et
al., 2015).</p>
        <p>For the second type of locations (task 2), we put
together all the tweets written by a user, in order
to predict his/her physical location. Only a few
users declare their locations in their Twitter
profiles, but this is sufficient to automatically
produce training and test data for our classifiers. We
experimented with two existing datasets
collected from users located in the U.S. We propose a
deep learning architecture for the solving the
task, because deep learning was shown to work
well for other natural language processing tasks,
and because standard classifiers were already
tested for the user location task. We designed a
model that predicts the U.S. region of the user
and his/her U.S. state, and another model that
predicts the longitude and latitude of the user's
location. We found that stacked denoising
autoencoders are well suited for this task, with results
comparable to the state-of-the-art (Liu and
Inkpen, 2015).
2</p>
        <p>Biography
Diana Inkpen is a Professor at the University of
Ottawa, in the School of Electrical Engineering and
Computer Science. Her research is in applications of
Computational Linguistics and Text Mining. She
organized seven international workshops and she was a
program co-chair for the AI 2012 conference. She is
in the program committees of many conferences and
an associate editor of the Computational Intelligence
and the Natural Language Engineering journals. She
published a book on Natural Language Processing for
Social Media (Morgan and Claypool Publishers,
Synthesis Lectures on Human Language Technologies), 8
book chapters, more than 25 journal articles and more
than 90 conference papers. She received many
research grants, including intensive industrial
collaborations.</p>
        <p>Acknowledgments
I want to thank my collaborators on the two tasks
addressed in this work: Ji Rex Liu for his
implementation of the proposed methods for both tasks
and Atefeh Farzindar for her insight on the first
task. I also thank the two annotators of the
corpus for task 1, Farzaneh Kazemi and Ji Rex Liu.</p>
        <p>This work is funded by the Natural Sciences and
Engineering Research Council of Canada.
Atefeh Farzindar and Diana Inkpen. 2015. Natural</p>
        <p>Language Processing for Social Media. Morgan
and Claypool Publishers. Synthesis Lectures on</p>
        <p>Human Language Technologies.</p>
        <p>Ji Liu and Diana Inkpen. 2015. Estimating User
Location in Social Media with Stacked Denoising
Autoencoders. In Proceedings of the 1st Workshop on
Vector Space Modeling for Natural Language
Processing, NAACL 2015, Denver, Colorado, pp.
201210.
Opinion Mining: Taking into account the criteria!</p>
        <p>Pascal Poncelet</p>
        <p>LIRMM - UMR 5506
Campus St Priest - Baˆtiment 5</p>
        <p>860 rue de Saint Priest
34095 Montpellier Cedex 5 - France</p>
        <p>and
UMR TETIS Montpellier - France</p>
        <p>Pascal.Poncelet@lirmm.fr
Today we are more and more
provided with information expressing
opinions about different topics. In the same
way, the number of Web sites giving a
global score, usually by counting the
number of stars for instance, is also growing
extensively and this kind of tools can be
very useful for users interested by having
a general idea. Nevertheless, sometimes
the expressed score (e.g. the number of
stars) does not really reflect what it is
expressed in the text of a review. Actually,
extracting opinions from texts is a
problem that have been extensively addressed
in the last decade and very efficient
approaches are now proposed to extract the
polarity of a text. In this presentation we
focus on a topic related with opinion but
rather than considering the full text we are
interested with the opinions expressed on
specific criteria. First we show how
criteria can be automatically learnt. Second
we illustrate how opinions are extracted.</p>
        <p>By considering criteria we illustrate that it
is possible to propose new recommender
systems but also to evaluate how opinions
expressed on the criteria evolve over time.
1
Extracting opinions that are expressed in a text is
a topic that have been addressed extensively in the
last decade (e.g. (Pang and Lee, 2008)). Usually
proposed approaches mainly focus on the polarity
of a text: this text is positive, negative or even
neutral. Figure 1 shows an example of a review on a
restaurant.</p>
        <p>Actually this review has been scored quite well:
4 stars over 5. Any opinion mining tools will show
that the review is much more negative than
positive. Let us go deeper on this exemple. Even if
We are here on a Saturday night and the food and service was amazing.</p>
        <p>We brought a group back the next day and we were treated so poorly by
a man with dark hair.</p>
        <p>He ignored us when we needed a table for 6 to the
point of us leaving to get takeaway.</p>
        <p>Embarrassing and so disappointing.
the review is negative it clearly illustrates that the
reviewer was mainly disappointed by the service:
he was in the Restaurant and found it amazing. We
could imagine that, at that time, the service was
not so bad. This exemple illustrates the problem
we address in the presentation: we do not focus
on a whole text rather we would like to extract
opinions related to some specific criteria.
Basically, by considering a set of user-specified criteria
we would like to highlight (and obviously extract
opinions) only on the relevant parts of the reviews
focusing on these criteria. The paper is organized
as follows. In Section 2 we give some ideas on
how to automatically learn terms related to a
criterium. We give also some clues for extracting
opinions to the criteria in Section 3. Finally
Section 4 concludes the paper.
2</p>
        <p>Automatic extraction of terms related
to a criterium
First of all we assume that the end user is
interested in a specific domain and some criteria.</p>
        <p>Let us imagine that the domain is movie and the
two criteria are actor and scenario. For each
criterium we only need to have several keywords or
terms of the criterium (seed of terms). For instance
in the movie domain: Actor= {actor, acting,
casting, character, interpretation, role, star} and
Scenario={scenario, adaptation, narrative,
original, scriptwriter, story, synopsis}. Intuitively two
different sets may exist. The first one
corresponding to all the terms that may be used for a
criterium. Such a set is called a class. The second
one corresponds to all the terms which are used in
the domain but which are not in the class. This
set is called anti-class. For instance the term
theater is about movie but is not specific neither to
the class actor nor scenario. Now the problem
is to automatically learn the set of all terms for
a class. Using experts or users to annotate
documents is too expensive and error-prone. By the
way there are many documents available on the
internet having the terms of the criteria that can be
learned. In a practical way by using a research
engine it is easy and possible to get these documents.</p>
        <p>For instance, the following query expressed in
Google: ”+movie +actor -scenario adaptation
narrative original -scriptwriter story -synopsis” will
extract a set of documents of the domain movie
(character +), having actor in the document and
without (character −) scenario, adaptation, etc. In
other terms we are able to automatically extract
movie documents having terms only relative to the
class actor. By performing some text
preprocessing and taking into account a frequency of a term
in a specific window of terms (see (Duthil et al.,
2011) for a full description of the process as well
as the measure that can be used to score the terms)
we can extract quite relevant terms: the higher the
score, the higher the probability of this term
belonging to a class. Nevertheless as the number of
documents to be analyzed is limited, some terms
may not appear in the corpus. Usually these terms
will have more or less the same score both in the
class and the anti-class. They are called
candidates and as we do not know the most
appropriate class, a new query on the Web will extract new
documents. Here again, a new score can be
computed and all the terms with their associated scores
can finally be stored in lexicons. Such lexicon can
then be used to automatically segment a document
for instance.
3</p>
        <p>Extracting opinions
A quite similar process may be adapted for
extracted terms used to express opinions:
adjectives, verbs and even grammatical patterns such as
&lt;adverb + adjective &gt; in order to automatically
learn positive and negative expressions. Then by
using the new opinion lexicon extracted we can
easily detect the polarity of a document. In the
same way by using the segmentation performed in
the previous step it is now possible to focus on
criteria and then extract the opinion for a specific
criterium. Interested reader may refer to (Duthil et
al., 2012).
4
In the presentation we will present more in
detail the main approach. Conducted experiments
that will be presented during the talk will show
that such an approach is very efficient when
considering Precision and Recall measures.
Furthermore some practical aspects will be addressed:
how many documents? how many seed terms?
the quality of the results for different domains?
We will also show that such lexicons could also
be very useful for recommending systems. For
instance we are able to focus on the criteria that
are addressed by newspapers and then recommend
the end user only with a list of newspapers he/she
could be interested in. In the same way, evaluating
how opinions evolve over time on different criteria
is of great interest for many different applications.</p>
        <p>Interested reader may refer to (Duthil, 2012) for
different applications that can be defined.</p>
        <p>Acknowledgments
The presented work has been done mainly during
the Ph.D of Dr. Benjamin Duthil and in
collaboration with Ge´rard Dray, Jacky Montmain, Michel
Plantie´ from the Ecole des Mines d’Ale`s (France)
and Mathieu Roche from the University of
Montpellier (France).</p>
        <p>References
How to deal with heterogeneous data?
The Big Data issue is traditionally characterized
in terms of 3 V, i.e. volume, variety, and velocity.</p>
        <p>This paper focuses on the variety criterion, which
is a challenging issue.</p>
        <p>Data heterogeneity and content
In the context of Web 2.0, text content is
often heterogenous (i.e. lexical heterogeneity).</p>
        <p>For instance, some words may be shortened
or lengthened with the use of specific graphics
(e.g. emoticons) or hashtags. Specific
processing is necessary in this context. For instance,
with an opinion classification task based on the
message SimBig is an aaaaattractive
conference!, the results are generally
improved by removing repeated characters (i.e. a).</p>
        <p>But information on the sentiment intensity
identified by the character elongation is lost with this
normalization. This example highlights the
difficulty of dealing with heterogeneous textual data
content.</p>
        <p>The following sub-section describes the
heterogeneity according the document types (e.g.
images and texts).</p>
        <p>Heterogeneity and document types
Impressive amounts of high spatial resolution
satellite data are currently available. This raises
the issue of fast and effective satellite image
analysis as costly human involvement is still required.</p>
        <p>Meanwhile, large amounts of textual data are
available via the Web and many research
communities are interested in the issue of
knowledge extraction, including spatial information. In
this context, image-text matching improves
information retrieval and image annotation techniques
(Forestier et al., 2012). This provides users with
a more global data context that may be useful
for experts involved in land-use planning
(Alatrista Salas et al., 2014).
2</p>
        <p>Text-mining method for matching
heterogenous data
A generic approach to address the heterogeneity
issue consists of extracting relevant features in
documents. In our work, we focus on 3 types
of features: thematic, spatial, and temporal
features. These are extracted in textual documents
using natural language processing (NLP) techniques
based on linguistic and statistic information
(Manning and Schu¨tze, 1999):
• The extraction of thematic information is
based on the recognition of relevant terms in
texts. For instance, terminology extraction
techniques enable extraction of single-word
terms (e.g. irrigation) or phrases (e.g.
rice crops). The most efficient
state-ofthe-art term recognition systems are based
on both statistical and linguistic information
(Lossio-Ventura et al., 2015).
• Extracting spatial information from
documents is still challenging. In our work, we
use patterns to detect these specific named
entities. Moreover, a hybrid method enables
disambiguation of spatial entities and
organizations. This method combines symbolic
approaches (i.e. patterns) and machine learning
techniques (Tahrat et al., 2013).
• In order to extract temporal expressions in
texts, we use rule-based systems like
HeidelTime (Stro¨tgen and Gertz, 2010). This
multilingual system extracts temporal
expressions from documents and normalizes them.</p>
        <p>HeidelTime applies different normalization
strategies depending on the text types, e.g.</p>
        <p>news, narrative, or scientific documents.</p>
        <p>These different methods are partially used in
the projects summarized in the following section.</p>
        <p>More precisely, Section 3 presents two projects
that investigate heterogeneous data in agricultural
the domain.1
3</p>
        <p>Applications in the agricultural
domain
New and emerging infectious diseases are an
increasing threat to countries. Many of these
diseases are related to globalization, travel and
international trade. Disease outbreaks are
conventionally reported through an organized multilevel
health infrastructure, which can lead to delays
from the time cases are first detected, their
laboratory confirmation and finally public
communication. In collaboration with the CMAEE2 lab, our
project proposes a new method in the epidemic
intelligence domain that is designed to discover
knowledge in heterogenous web documents
dealing with animal disease outbreaks. The proposed
method consists of four stages: data acquisition,
information retrieval (i.e. identification of
relevant documents), information extraction (i.e.
extraction of symptoms, locations, dates, diseases,
affected animals, etc.), and evaluation by different
epidemiology experts (Arsevska et al., 2014).
3.2 Information extraction from</p>
        <p>experimental data
Our joint work with the IATE3 lab and
AgroParisTech4 deals with knowledge engineering issues
regarding the extraction of experimental data from
scientific papers to be subsequently reused in
decision support systems. Experimental data can be
represented by n-ary relations, which link a
studied topic (e.g. food packaging, transformation
process) with its features (e.g. oxygen permeability
in packaging, biomass grinding). This knowledge
is capitalized in an ontological and
terminological resource (OTR). Part of this work consists of
recognizing specialized terms (e.g. units of
measures) that have many lexical variations in
scientific documents in order to enrich an OTR
(Berrahou et al., 2013).</p>
        <p>1http://www.textmining.biz/agroNLP.
html</p>
        <p>2Joint research unit (JRU) regarding the control of exotic
and emerging animal diseases – http://umr-cmaee.
cirad.fr</p>
        <p>3JRU in the area of agro-polymers and emerging
technologies – http://umr-iate.cirad.fr
4http://www.agroparistech.fr
Heterogenous data processing enables us to
address several text-mining issues. Note that we
integrated the knowledge of experts in the core
of research applications summarized in Section 3.</p>
        <p>In future work, we plan to investigate other
techniques dealing with heterogeneous data, such as
visual analytics approaches (Keim et al., 2008).
to integrate declarative query constructs from the
database community into MapReduce to allow
greater data independence.
3</p>
        <p>Relational DBMS
Since it was developed by Edgar Codd in 1970, as
presented in (Shuxin and Indrakshi, 2005), the
relational database (RDBMS) has been the
dominant model for database management. RDBMS is
the basis for SQL, and is a type of database
management system (DBMS) that is based on the
relational model which stores data in the form of
related tables, and manages and queries structured
data. Since the RDBMSs focuse on extending the
database system’s capabilities and its processing
abilities, RDBMSs have become a predominant
powerful choice for the storage of information in
new databases because they are easier to
understand and use. What makes it powerful, is that it is
based on relation between data; because the
possibility of viewing the database in many different
ways since the RDBMS require few assumptions
about how data is related or how it will be
extracted from the database. So, an important feature
of relational systems is that a single database can
be spread across several tables which might be
related by common database table columns.</p>
        <p>RDBMS also provide relational operators to
manipulate the data stored into the database tables.</p>
        <p>However, as discussed in (Hammes et al., 2014),
the lack of the RDBMS model resides in the
complexity and the time spent to design and normalize
an efficient database. This is due to the several
design steps and rules, which must be properly
applied such as Primary Keys, Foreign Keys,
Normal Forms, Data Types, etc. Relational Databases
have about forty years of production experience,
so the main strength to point out is the maturity of
RDBMSs. That ensure that most trails have been
explored and functionality optimized. For the user
side, he must have the competence of a database
designer to effectively normalize and organize the
database, plus a database administrator to
maintain the inevitable technical issues that will arise
after deployment.</p>
        <p>A lot of work has been done to compare the
MapReduce model with parallel relational
databases, such as (Pavlo et al., 2009), where
experiments are conducted to compare Hadoop
MapReduce with two parallel DBMSs in order to
evaluate both parallel DBMS and the MapReduce
model in terms of performance and development
complexity. The study showed that both databases
did not outperformed Hadoop for user-defined
function. Many applications are difficult to
express in SQL, hence the remedy of the
user-defined function. Thus, the efficiency of the
RDBMSs is in regular database tasks, but the
user-defined function presents the main ability
lack of this DBMS type.</p>
        <p>A proof of improvement of the RDBMS model
comes with the introduction of the
Object-Oriented Database Relational Model (ORDBMS). It
aims to utilize the benefits of object oriented
theory in order to satisfy the need for a more
programmatic flexibility. The basic goal presented in
(Sabàu, 2007) for the Object-relational database is
to bridge the gap between relational databases and
the object-oriented modeling techniques used in
programming languages. The most notable
research project in this field is Postgres (Berkeley
University, Californie); Illustra and PostgreSQL
are the two products tracing this research.
4</p>
        <p>MapReduce-RDBMS: integrating
paradigms
In this section, the use of relational DBMS and
MapReduce as complementary paradigms is
considered.</p>
        <p>It is important to pick the right database
technology for the task at hand. Depending on what
problem the organization is trying to solve, it will
determine the technology that should be used.</p>
        <p>Several comparative studies have been co
ducted between MapReduce and parallel DBMS
such as (Mchome, 2011) and (Pavlo et al., 2009).</p>
        <p>MapReduce has been presented as a replacement
for the Parallel DBMS. While each system has its
strengths and its weaknesses. However, an
integration of the two systems is needed, and as
proposed in (Stonebraker et al., 2010), MapReduce
can be seen as a complement to a RDBMS for
analytical applications, because different problems
require complex analysis capabilities provided by
both technologies.</p>
        <p>In this context, we propose a model that
integrates the MapReduce model and a relational
DBMS PostgreSQL presented in (Worsley and
Drake, 2002), in a goal of queries optimization.</p>
        <p>We suggest an OLAP queries process model in a
goal of minimizing Input/Output costs in terms of
the amount of data to manipulate, reading and
writing throughout the execution process.</p>
        <p>The basic idea behind our approach is based on
the cost model to approve execution and
selectivity of solutions based on the estimated cost of
execution. To support the decision making process
for analyzing data and extracting useful
knowledge while minimizing costs, we propose to
compare the estimates of the costs of running a
query on Hadoop MapReduce compared to
PostgreSQL to choose the least costly technology.</p>
        <p>As the detailed analysis of the queries
execution costs showed a gap mattering between both
paradigms, hence the idea of the thorough
analysis of the execution process of each query and the
implied cost. So, to better control the cost
difference between costs of Hadoop MapReduce versus
PostgreSQL on each step of the query's execution
process, we propose to dissect each query for a set
of operations that demonstrates the process of
executing the query and in order to control the
different stages of the execution process of each
query (decomposing the query to estimate the
costs on each system for each individual
operation). In this way, we can check the impact of the
execution of each operation of a query on the
overall cost and we can control the total cost of
the query by controlling the partial cost of each
operation in the information retrieval process. For
this purpose, we suggest to provide a detailed
execution plan for OLAP queries. This execution
plan allows zooming on the sequence of steps of
the process of executing a query. It details the
various operations of the process highlighting the
order of succession and dependence. In addition, it
determines for each operation the amount of data
involved and the dependence implemented in the
succession of phases. These parameters will be
needed to calculate the cost involved in each
operation.</p>
        <p>Having identified all operations performed
during the query execution process the next step is
then to calculate the cost implied in each operation
independently, in both paradigms PostreSQL and
MapReduce with the aim of controlling the
estimated costs difference according to the operations
as well as the total cost of query execution. At this
stage we consider each operation independently to
calculate an estimate of its cost execution on
PostreSQL on one hand then on MapReduce on
the other hand. For this goal, we propose to relay
on a cost model for each system PostgreSQL and
Hadoop MapReduce in order to estimate the I/O
cost of each operation execution on both system
independently.</p>
        <p>We aim to estimate how expensive it is to
preprocess each operation of each query on both
systems. Therefore, controlling the cost implied by
each operation as well as its influence on the total
cost of the query, allows the control of the cost of
each query to support the decision making process
and the selectivity of the proposed solutions based
on the criterion of cost minimization.</p>
        <p>Based on a sample workload of OLAP queries,
and having identifying the cost of each operation
for each query executed on both systems
independently, the results analysis can be useful to
deduct a generalized smart model that integrates the
two paradigms to process the OLAP queries in a
cost minimization way.
5
Given the exploding data problem, the world of
databases has evolved which aimed to escape the
limitations of data processing and analysis. There
has been a significant amount of work during the
last two decades related to the needs of new
supporting technologies for data processing and
knowledge management, challenged by the rise of
data generation and data structure diversity.</p>
        <p>In this paper, we have investigated the
MapReduce model in one hand, then the Relational
DBMS technology in the other hand, in order to
present strengths and weaknesses of each
paradigm.</p>
        <p>Although MapReduce was designed to cope
with large amounts of unstructured data, there will
be advantages in exploiting it in structured data
processing. In this fashion, we have proposed in
this paper a new OLAP queries process model
integrating an RDBMS with the MapReduce
framework in a goal of minimizing Input/Output costs
in terms of the amount of data to manipulate,
reading and writing throughout the execution process.</p>
        <p>Combining MapReduce and RDBMS
technologies has the potential to create very powerful
systems. For this reason, we plan to investigate other
types of integration for different applications.</p>
        <p>Hammes D., Medero, H. and Mitchell H. 2014.
Comparison of NoSQL and SQL Databases in the Cloud.</p>
        <p>Southern Association for Information Systems
(SAIS) Proceedings. Paper 12.</p>
        <p>McClean A., Conceicao RC., and O'Halloran M. 2013.</p>
        <p>A Comparison of MapReduce and Parallel Database
Management Systems. ICONS 2013, The Eighth
International Conference on Systems: 64-68.</p>
        <p>Mchome M.L. 2011. Comparison study between</p>
        <p>MapReduce (MR) and parallel data management
systems (DBMs) in large scale data anlysis. Honors</p>
        <p>Projects Macalester College.</p>
        <p>Nykiel T., Potamias M., Mishra C., Kollios G. and</p>
        <p>Koudas N. 2010. Mrshare: sharing across multiple
queries in mapreduce. Proceedings of the VLDB
Endowment ,volume 3 (1-2). 494-505.</p>
        <p>Ordonez C. 2013. Can we analyze big data inside a</p>
        <p>DBMS?. Proceedings of the sixteenth international
workshop on Data warehousing and OLAP, ACM.</p>
        <p>85-92.</p>
        <p>Olston C., Reed B., Srivastava U., Kumar R. and
Tomkins A. 2008. Pig latin: a not-so-foreign language
for data processing. Proceedings of the 2008 ACM
SIGMOD international conference on Management
of data , ACM. 1099-1110.</p>
        <p>Pavlo A., Rasin A., Madden S., Stonebraker M.,</p>
        <p>DeWitt D., Paulson E., Shrinivas L. and Abadi D.J.
2009. A comparison of approaches to large scale
data analysis. Proceedings of the 2009 ACM
SIGMOD International Conference on Management of
data, ACM. 165-178.</p>
        <p>Shuxin Y. and Indrakshi R. 2005. Relational database
operations modeling with UML. Proceedings of the
19th International Conference on Advanced
Information Networking and Applications. 927-932.</p>
        <p>Sabàu G. 2007. Comparison of RDBMS, OODBMS</p>
        <p>and ORDBMS. Informatica Economic.</p>
        <p>Stonebraker M., Abadi D., DeWitt D.J., Madden S.,</p>
        <p>Paulson E., Pavlo A. and Rasin A. 2010. Mapreduce
and parallel dbmss : friends or foes ?.
Communications of the ACM, volume 53 (1). 64-71.</p>
        <p>Wang G. and Chan CY. 2013. Multi-Query
Optimization in MapReduce Framework. Proceedings of the
VLDB Endowment, 40th International Conference
on Very Large Data Bases, volume 7 (3).</p>
        <p>Worsley J.C. and Drake J.D. 2002. Practical
PostgreSQL. O'Reilly and Associates Inc
Performance of Alternating Least Squares in a distributed approach using</p>
        <p>GraphLab and MapReduce
Elizabeth Veronica Vera Cervantes, Laura Vanessa Cruz Quispe, Jose´ Eduardo Ochoa Luna
National University of San Agustin</p>
        <p>Arequipa, Peru
elizavvc@gmail.com, lcruzq@unsa.edu.pe, eduardo.ol@gmail.com
Automated recommendation systems have
been increasingly adopted by companies
that aim to draw people attention about
products and services on Internet. In
this sense, development of distributed
model abstractions such as MapReduce
and GraphLab has brought new
possibilities for recommendation research tasks
due to allow us to perform Big Data
analysis. Thus, this paper investigates the
suitability of these two approaches for
massive recommendation. In order to do
so, the Alternating Least Squares (ALS),
which is a Collaborative Filtering
algorithm, has been tested using
recommendation benchmark datasets. Results on
RMSE show a preliminary comparative
performance analysis.
1 Introduction
Data on the Internet is increasing, e-commerce
sites, blogs and social networks spread the word
about new products and services everyday. This
social media information overwhelms any user,
who has a given profile and therefore could not
be interested in most of these offers (Koren et al.,
2009).</p>
        <p>In this sense, recommendation systems have
gained momentum, because they “filter”
products and services for users according to
behavior patterns. Traditional approaches for automated
recommendation range from Content-Based,
Collaborative Filtering and Deep Learning systems
(Adomavicius, 2005; Shi et al., 2014). However,
to handle the current amount of available data we
need to resort to frameworks for large-scale data
processing.</p>
        <p>Recently, the machine learning community has
been increasingly interested in the task of
managing Big Data with parallelism (Zhou et al., 2008;
De Pessemier et al., 2011; Xianfeng Yang, 2014).</p>
        <p>However, parallel algorithms are extremely
challenging and traditional approaches, despite of
being powerful like MPI, rely on low levels of
abstraction. On the other hand, distributed models
such as GraphLab(Low et al., 2012) and
MapReduce(Xiao and Xiao, 2014; Dean and Ghemawat,
2008) foster high levels of abstraction and,
therefore, they are more intuitive. The aim of this
paper is to investigate whether these distributed
models are suitable for recommendation tasks. In
order to do so we evaluate the Alternating Least
Squares (ALS) algorithm, a parallel
collaborative filtering approach(Koren et al., 2009;
Schelter et al., 2013), in both GraphLab and
MapReduce frameworks. We evaluate the performance
on the MovieLens and Netflix datasets.
According to preliminary results, GraphLab outperforms
MapReduce in RMSE, when Lambda, iterations
number and latent factor parameters are
considered. Conversely, MapReduce gets a better
execution time than GraphLab using the same
parameters in MoveLens dataset. The paper is organized
as follows. In Section 2, related work is described.</p>
        <p>Background is given in Section 3. Our proposal is
showed in Section 4. Preliminary results are
depicted in section 5. Finally Section 5, concludes
the paper.
2
Several distributed platforms have been used for
studying performance of machine learning
algorithms, for instance, a Matrix Factorization
based on collaborative filtering over MapReduce
model was proposed in (Xianfeng Yang, 2014;
De Pessemier et al., 2011). In Low et al.
(2012), some advantages and disadvantages of
using GraphLab and MapReduce were described.</p>
        <p>For instance, MapReduce fails when there are
computational dependencies on data, but it can be
used to extract features from a massive collection.
In addition, MapReduce is targeted for large data
centers, it is optimized for node-failure and
diskcentric parallelism. Conversely, In GraphLab it is
assumed that processors do not fail, and all data is
stored in shared memory.</p>
        <p>In Low et al. (2012), the Alternating
Least Squares (ALS) algorithm was
implemented over several platforms: GraphLab,
Hadoop/MapReduce and MPI. Comparison results
show that applications created using GraphLab
outperformed equivalent Hadoop/MapReduce
implementations by 20-60 times(Xianfeng Yang,
2014) .</p>
        <p>Our work is most related to Low et al. (2012),
but we focus on the evaluation of different
configurations of ALS algorithm over GraphLab and
MapReduce. Thus, we aim at obtaining optimal
parameters that allow us to improve algorithm
performance. Moreover, comparisons were based on
RMSE and time execution values. The parameters
considered are:
• Lambda, which is the regularization
parame</p>
        <p>ter in ALS
3
3.1
• The number of latent factors
• The number of iterations</p>
        <p>Background
A recommendation system aims at showing items
of interest to a user, considering the context of
where the items are being shown and to whom they
are being shown (Alag, 2008).</p>
        <p>Figure 1, depicts inputs and outputs of a
common recommendation system.
• Content-based Recommendation: Items
similar to the ones he/she has preferred in the
past, are recommended to the user.
• Collaborative Recommendation: Items that
people with similar tastes and preferences
liked in the past, are recommended to the
user.</p>
        <p>– Collaborative Deep Learning: It is a
recent kind of collaborative filtering
using deep learning models Wang et al.</p>
        <p>(2014).
• Hybrid Approach: Recommendations are
made using a combination of Content-based
and Collaborative Recommendation
methods.</p>
        <p>Alternating Least Squares (ALS)
Alternating Least Squares (Low et al., 2012; Zhou
et al., 2008; Koren et al., 2009) is an algorithm
within the collaborative filtering paradigm. Input
of ALS (in Figure 2) is a sparse user by items
matrix R containing the rating of each user. The
algorithm iteratively computes a low-rank matrix
factorization R = U × V where U and V are d
dimensional matrices. The loss function is defined
as the squared error(Zhou et al., 2008), where the
learning objective is to minimize the sum of the
squared errors 1 between values predicted and real
values of rantings.</p>
        <p>(Uˆ , Vˆ ) = argmin X i, j ∈ R(rij − viT uj )2 (1)</p>
        <p>U,V
Complexity and cost depend on the magnitude of
the hidden variables d.</p>
        <p>U=Users</p>
        <p>R
d</p>
        <p>C=Courses
d</p>
        <p>In Adomavicius (2005) three approaches for
building a recommendation system are presented:</p>
        <p>The ALS algorithm is computationally
expensive, every iteration runs on O(d−1[N r + (m +
n)r2] + r3), where m is length of items, and n is
length of users (Schelter et al., 2013; Gemulla et
al., 2011).</p>
        <p>Alternating Least Squares on GraphLab
According to (Low et al., 2012; Gonzalez et al.,
2012) ALS in GraphLab is implemented by
using a bipartite two colorable graph and a
chromatic synchronous engine with an edge
consistency model for serializability.</p>
        <p>Each vertex of the graph has a latent factor
attached, that denotes a user or an item. Thus, they
are linked to a column or a row in the matrix of
ratings R. Each edge of the graph contains
entry data (rating values), and the most recent error
estimated by the algorithm. The goal of ALS
algorithm is to discover values of latent parameters,
such that non-zero entries in R can be predicted by
the dot product of the row and column latent
factor. ALS algorithm for GraphLab is implemented
in the Gather-Apply-Scatter abstraction. ALS
update considers adjacent vertices as X values and
edges as observed y values, and then updates the
current vertex value as a weight w:</p>
        <p>y = X ∗ w + noise
w = inv(X0 ∗ X) ∗ (X0 ∗ y)
In the Gather-Apply-Scatter model, the update is
done as follows:
• Gather: it returns the tuple (X0 ∗ X, X0 ∗ y)
• Apply: it solves inv(X0 ∗ X) ∗ (X0 ∗ y)
• Scatter: it schedules the update of adjacent
vertices if this vertex has changed and the
edge is not well predicted.
3.4</p>
        <p>Alternating Least Squares on</p>
        <p>MapReduce
In Xianfeng Yang (2014; Zhou et al. (2008)
MapReduce implementation is comprised by four
tasks as shown in Figure 3. Each item in dataset
is denoted as a triple (u, j, r). u denotes user, j is
the label of item and r denotes corresponding
rating. In the U-Update step, item matrix V is used
as input and is sent to cluster nodes. Then,
training rating R is used to compute user matrix U ,
including inputs as lambda parameter λ to
regularization, number of latent factors. V-Update does
(2)
(3)
that is accomplished using the following equation:
4
the same as U-Update step, but its input is not an
item matrix. On the contrary, it is a user matrix
computed in U-Update step. Once U and V are
learned, we can compute RMSE values using test
dataset and estimated rating rˆ. So the Parallel ALS
algorithm with Weighted--Regularization is as
follows (Zhou et al., 2008): The objective function in
Algorithm 1 Alternating Least Square(ALS) with
algorithm
1: Initialize V with random values between 0</p>
        <p>and 1
2: Hold V constants, and solve U by minimizing</p>
        <p>the objective function.
3: Hold U constants, and solve M by
minimiz</p>
        <p>ing the objective function.
4: repeat from step 2 and 3 until objective
func</p>
        <p>tion converge.
1 is obtained from equation 4, which is just linear
regression with lambda regularization(λ), to avoid
overfits it penalize large parameters.</p>
        <p>In this paper we evaluate several parameter
configurations (lambda, number of latent factor, number
of iterations) for ALS algorithm over GraphLab
and MapReduce. Our aim is to obtain the best
performance, over clusters of two and four machines,
for the Movielens Dataset, and NetFlix Dataset
(further details will be given in the next section).</p>
        <p>We evaluate performance according to RMSE and
execution time values.</p>
        <p>In order to implement ALS algorithm under the
MapReduce Paradigm, the Mahout 1 API has been
used. ALS algorithm for MapReduce (Zhou et al.,
2008) is shown in 3. User and movie factors have
been computed using equation 4. where nui and
nvj are the numbers of ratings of user i and item
j respectively. When objective function showed in
equation 4 does not change after further iterations,
we attain the final step. Output is the predicted
rating for each user/item pair.</p>
        <p>1http://mahout.apache.org/
f (U, V ) =</p>
        <p>X (rij − uiT vj )2 + λ(X nuik ui k2 +
i,j∈I i
Trainingratingmatrix,
itemmatrix
V-Update</p>
        <p>Trainingratingmatrix,
usermatrix
U-Update</p>
        <p>Testratingmatrix,
itemmatrix,usermatrix
RMSECalculation</p>
        <p>Bestregularizationfactor
ƛ,
itemmatrix,usermatrix
Prediction
estimatedratingsfor
user/itempairs</p>
        <p>In order to evaluate ALS algorithm under
GraphLab, the GraphLab API (Low et al., 2010)
has been used. ALS algorithm for GraphLab (Low
et al., 2012) is shown in Figure 4. User and movie
factors have been computed using equation 5.</p>
        <p>f [i] = argmin
w∈Rd j∈Neighbors(i)</p>
        <p>X</p>
        <p>(rij − wT f [j]) (5)
t)(r
U
s
o
c
a
fr
e
s
U</p>
        <p>C
o
u
r
s
e
f
a
c
t()r
o
s</p>
        <p>C</p>
        <p>Matrix factorization of ALS using
MovieLens is a Web collaborative site that
manages a recommender system for movies. This
recommender system is based on a collaborative
filtering algorithm developed by the GroupLens
research group. The dataset is comprised by 6040
users, 3952 items and 100209 ratings for training.</p>
        <p>The data structure is: user, item, rating.
4.2</p>
        <p>Netflix Dataset
We are using the Small Netflix Dataset. It is also a
data-set for movie recommendation, it has 95526
users, 3561 items and 3298163 ratings. The
structure of the data-set is: user, item, rating.
Setup of the GraphLab cluster is as follows. Two
machines, one working as the master and the other
as the worker node. The master machine
operating system is Ubuntu 14.04, and its processor
is Intel Core i3 CPU M 330@2.13GHzx4. The
worker machine operating system is Ubuntu 13.10
of 64-bit, and its processor is Intel Core i3-2350M
@2.30GHzx4. The cluster was configured using
MPI(Message Passing Interface).
The setup is as follows. Four machines, three
worker nodes and one master. The master machine
operating system is Ubuntu 13.10 of 64-bit, and its
processor is Intel Core i3-2350M @2.30GHzx4.</p>
        <p>Table.1 shows the congfiuration of the worker
machines.</p>
        <p>The cluster was configured using Hadoop, and
Machine Operating System Processor
1 Ubuntu 14.04 Intel Core i3 CPU M 330@2.13GHzx4
2 Ubuntu 13.10 Intel Core i3-2350M @2.30GHzx4
3 Ubuntu 14.04 Intel Core i7-4700MQ @2.40GHzx8
the HDFS(Hadoop Distributed File System). The
ALS(Alternating Least Squares) algorithm
implementation was taken from Mahout Library.
This section shows experimental results
conducted on MovieLens data set aforementioned.</p>
        <p>Experimental setting parameters are described in</p>
        <p>Parameters
Lambda
# Latent factors
# Iterations</p>
        <p>Value
0.01 - 0.09
10-50
2-30
Table 2: Parameters used for ALS algorithms</p>
        <p>In Figure 5a, RMSE values for MapReduce
do not change even if we increase the number
of latent factors thus, RMSE values on
MapReduce are independent on the number of latent
factors. RMSE for MapReduce converges around
0.95. Conversely, RMSE values for GraphLab
decreases while the number of latent factors
increases. When the number of latent factors was
50, RMSE value reaches around 0.25. However,
GraphLab spends more time than MapReduce,
Figure 5b depicts MapReduce times almost as an
horizontal line for MovieLens dataset, the line of
execution time for Netflix dataset is much steeper.</p>
        <p>Between Graphlab and MapReduce lines
representing Movilens dataset execution, Graphlab line
is more pronounced.</p>
        <p>Figure 6a depicts GraphLab and MapReduce
performance according the Lambda parameter.</p>
        <p>While Lambda increases, RMSE decreases
accordingly, i.e., if a greater value of Lambda is used
then algorithm accuracy tends to be better. We also
notice that Graphlab has lower values of RMSE
compared to MapReduce. GraphLab RMSE
values are around 0.5, and MapReduce RMSE values
are around 1. Figure 6b illustrates a better
execution time of MapReduce compare to GraphLab
over Movielens dataset. However, now the
execution time for GraphLab decreases, while the value
of Lambda increases. Figure 6b also shows that It
takes longer to process the data from netflix than
Movielens.</p>
        <p>In Figure 7a we notice that the value of RMSE
is almost invariant to the increase of iterations for
MapReduce execution, given that the number of
iterations are small,nevertheless we notice clearly
that RMSE value for GraphLab decreases as the
number of iterations increases. RMSE value for
Graphlab converges around 0.55. Figure 7b shows
that MapReduce execution time over Movielens
dataset is good, however it increases a lot for
Netflix dataset. Graphlab execution time increases as
the number of iterations grows.
1
0.9
0.8
0.7
ESM 0.6
R
0.5
0.4
0.3</p>
        <p>Netflix(M)
MovieLens(M)</p>
        <p>MovieLens(G)
15
20 25 30 35 40</p>
        <p>Number of Latent Factors
(b) Number of latent factors Vs. Time
Figure 5: Performance of MapReduce and
GraphLab when number of features in ALS
algorithms is increased.</p>
        <p>111...321 MMoovvNiieeeLLteeflnnixss(((MMG)))
1
ES 0.9
RM 0.8
0.7
0.6
0.5
0.40.01 0.02 0.03 0.04 0.05 0.06 0.07 0.08 0.09</p>
        <p>Lambda
(a) Lambda Vs. RMSE
0.004..545 MMoovvNiieeeLLteeflnnixss(((MMG)))
)r 0.35
s 0.3
u
oh 0.25
(
e
iTm 0.2
0.15
0.1
0.05
00.01 0.02 0.03 0.04 0.05 0.06 0.07 0.08 0.09</p>
        <p>Lambda
(b) Lambda Vs. Time</p>
        <p>Netflix(M)
MovieLens(M)</p>
        <p>MovieLens(G)
15</p>
        <p>20
Number of Iterations
25</p>
        <p>30
(a) Number of Iterations Vs. RMSE
15</p>
        <p>20
Number of Iterations
25</p>
        <p>30
(b) Number of Iterations Vs. RMSE
We evaluated the Alternating Least Squares
(ALS) algorithm, a parallel collaborative filtering
in both GraphLab and MapReduce frameworks.</p>
        <p>Experiments were run over the MovieLens and
Netflix datasets. The RMSE between MapReduce
execution in NetFlix dataset and Movielens
dataset in all the experiments was similar, but the
execution time was longer in Netflix dataset.</p>
        <p>Looking at the executions over Moviliens dataset,
we can say, that even though GraphLab only ran in
two machines and MapReduce in 4 machines, the
first one outperformed the second one in RMSE.</p>
        <p>Considering lambda value variation, Figure.6a,the
number of iterations Figure.7a, and the number
of latent factors Figure.5a, GraphLab performed
better (RMSE) than MapReduce. In all previous
three cases MapReduce was faster than GraphLab,
obviously by the difference between the number
of machines in their configuration.</p>
        <p>Thus, when scalability and distribution are
evaluated, MapReduce performs better, because ALS
does not require data dependency for computing.</p>
        <p>Moreover, it took less execution time when more
latent factors were added. In this work we only
used two nodes, however GraphLab demonstrated
best results with few nodes.</p>
        <p>In conclusion, GraphLab performed better
when RMSE was considered but, there are open
issues with shared-memory. GraphLab is also
better for computing recommendations in real time.</p>
        <p>However, for more sophisticated computations
MapReduce performs better so far as to an offline
environment and all data is used.</p>
        <p>References
Tuzhilin A. Adomavicius, G. 2005. Toward the next
generation of recommender systems: a survey of
the state-of-the-art and possible extensions.
Knowledge and Data Engineering, IEEE Transactions on,
17:734 – 749.</p>
        <p>Satnam Alag. 2008. Collective Intelligence in Action.</p>
        <p>Toon De Pessemier, Kris Vanhecke, Simon Dooms,
and Luc Martens. 2011. Content-based
recommendation algorithms on the hadoop mapreduce
framework. In 7th international conference on Web
Information Systems and Technologies, Proceedings,
pages 237–240. Ghent University, Department of
Information technology.</p>
        <p>Jeffrey Dean and Sanjay Ghemawat. 2008.
Mapreduce: Simplified data processing on large clusters.</p>
        <p>Commun. ACM, 51(1):107–113, January.</p>
        <p>Rainer Gemulla, Erik Nijkamp, Peter J Haas, and
Yannis Sismanis. 2011. Large-scale matrix
factorization with distributed stochastic gradient descent.</p>
        <p>In Proceedings of the 17th ACM SIGKDD
international conference on Knowledge discovery and data
mining, pages 69–77. ACM.</p>
        <p>Joseph E. Gonzalez, Yucheng Low, Haijie Gu, Danny</p>
        <p>Bickson, and Carlos Guestrin. 2012. Powergraph:
Distributed graph-parallel computation on natural
graphs. In Proceedings of the 10th USENIX
Conference on Operating Systems Design and
Implementation, OSDI’12, pages 17–30, Berkeley, CA, USA.</p>
        <p>USENIX Association.</p>
        <p>Y Koren, R Bell, and C Volinsky. 2009. Matrix
factorization techniques for recommender systems.
Computer, 42(8):30–37.</p>
        <p>Yucheng Low, Joseph Gonzalez, Aapo Kyrola, Danny</p>
        <p>Bickson, Carlos Guestrin, and Joseph M.
Hellerstein. 2010. Graphlab: A new parallel framework
for machine learning. In Conference on Uncertainty
in Artificial Intelligence (UAI) , July.</p>
        <p>Yucheng Low, Danny Bickson, Joseph Gonzalez,
Carlos Guestrin, Aapo Kyrola, and Joseph M
Hellerstein. 2012. Distributed graphlab: a framework for
machine learning and data mining in the cloud.
Proceedings of the VLDB Endowment, 5(8):716–727.</p>
        <p>Sebastian Schelter, Christoph Boden, Martin Schenck,</p>
        <p>Alexander Alexandrov, and Volker Markl. 2013.</p>
        <p>Distributed matrix factorization with mapreduce
using a series of broadcast-joins. In Proceedings of
the 7th ACM Conference on Recommender Systems,
RecSys ’13, pages 281–284, New York, NY, USA.</p>
        <p>ACM.</p>
        <p>Yue Shi, Martha Larson, and Alan Hanjalic. 2014.</p>
        <p>Collaborative filtering beyond the user-item matrix:
A survey of the state of the art and future challenges.</p>
        <p>ACM Comput. Surv., 47(1):3:1–3:45, May.</p>
        <p>Hao Wang, Naiyan Wang, and Dit-Yan Yeung. 2014.</p>
        <p>Collaborative deep learning for recommender
systems. CoRR, abs/1409.2944.</p>
        <p>Pengfei Liu Xianfeng Yang. 2014. Collaborative
filtering recommendation using matrix factorization: A
mapreduce implementation. International Journal
of Grid and Distributed Computing.</p>
        <p>Zhifeng Xiao and Yang Xiao. 2014. Achieving
accountable mapreduce in cloud computing. Future</p>
        <p>Gener. Comput. Syst., 30:1–13, January.</p>
        <p>Yunhong Zhou, Dennis Wilkinson, Robert Schreiber,
and Rong Pan. 2008. Large-scale parallel
collaborative filtering for the netflix prize. In Proceedings
of the 4th International Conference on Algorithmic
Aspects in Information and Management, AAIM
’08, pages 337–348, Berlin, Heidelberg.
SpringerVerlag.</p>
        <p>Data Modeling for NoSQL Document-Oriented Databases
Harley Vera, Wagner Boaventura, Maristela Holanda, Valeria Guimar a˜es, Fernanda Hondo
Department of Computer Science</p>
        <p>University of Bras´ılia</p>
        <p>Bras´ılia, Brasil.
{harleyve,wagnerbf}@gmail.com, mholanda@cic.unb.br,
{valeriaguimaraes, fernandahondo}@hotmail.com
In database technologies, some of the
new issues increasingly debated are
non-conventional applications, including
NoSQL (Not only SQL) databases, which
were initially created in response to
the needs for better scalability, lower
latency and higher flexibility in an era
of bigdata and cloud computing. These
non-functional aspects are the main reason
for using NoSQL database. However,
currently there are no systematic studies
on data modeling for NoSQL databases,
especially the document-oriented ones.</p>
        <p>Therefore, this article proposes a NoSQL
data modeling standard in the form of
ER diagrams, introducing modeling
techniques that can be used on
documentoriented databases. On the other hand
the purpose of this article is not structure
the data using the model proposed, but
it does helping with the visualization of
data. In addition, to validate the proposed
model, a study case was implemented
using genomic data.
1 Introduction
Huge amounts of data are produced daily. They
are generated by smart phones, social networks,
banks transactions, machines measured by sensors
are part of Internet of Things provide information
that is growing exponentially. The management of
this data is currently performed in most cases by
relational databases that provide centralized
control of data, redundancy control and elimination
of inconsistencies (Elmasri and Navathe, 2010);
but, some of these factors restrict the use of
alternative database models. Consequently, certain
limiting factors have led to alternative models of
databases in these scenarios. Primarily, motivated
by the issue of system scalability, a new
generation of databases, known as NoSQL, is gaining
strength and space in information systems. The
NoSQL databases emerged in the mid-90s, from
a database solution that did not provide an SQL
interface. Later, the term came to represent
solution that promote an alternative to the Relational
Model, becoming an abbreviation for Not Only
SQL.</p>
        <p>The purpose, therefore, of NoSQL solutions is
not to replace the Relational Model as a whole,
but only in cases in which there is a need for
scalability and bigdata. In the recent years, a
variety of NoSQL databases has been developed
mainly by practitioners looking to fit their specific
requirements regarding scalability performance,
maintenance and feature-set. Subsequently, there
have been various approaches to classify NoSQL
databases, each with different categories and
subcategories, such as key-value stores,
columnoriented and graph databases, oriented-document.</p>
        <p>
          MongoDB
          <xref ref-type="bibr" rid="ref3">(MongoDB, 2015)</xref>
          , Neo4j (Partner et
al., 2013), Cassandra (D. Borthakur et al., 2011)
and HBase (F. Chang et al., 2008) are examples
of NoSQL databases. This article only applies to
NoSQL document-oriented databases, because of
the heterogeneous characteristics of each NoSQL
database classification.
        </p>
        <p>Nonetheless, data modeling still has an
important role to play in NoSQL environments. The data
modeling process (Elmasri and Navathe, 2010)
involves the creation of a diagram that represents the
meaning of the data and the relationship between
the data elements. Thus, understanding is a
fundamental aspect of data modeling (R. F. Lans, 2008),
and a pattern for this kind of representation has
few contributions for NoSQL databases.</p>
        <p>Addressing this issue, this article proposes a
standard for NoSQL data modeling. This proposal
uses NoSQL document-oriented databases,
aiming to introduce modeling techniques that can be
used on databases with document features.</p>
        <p>The remainder of the paper is organized as
follows: Section II presents related works. Section
III explores the concepts of modeling for NoSQL
databases based on documents, introducing the
different types of relationships and associations.</p>
        <p>Section IV shows the proposal model in the
context of NoSQL databases based on documents.</p>
        <p>Section V presents the study case to validate the
proposal model. Finally in Section VI, presents
the conclusion of the research and future works.
2
Katsov (H. Scalable, 2015) presents a study of
techniques and patterns for data modeling using
different categories of NoSQL databases.
However, the approach is generic and does not define a
specific modeling engine to each database.</p>
        <p>Arora and Aggarwal (R. Arora and R.
Aggarwal, 2013) propose a data modeling, but restricted
to MongoDB document database, describing a
UML Diagram Class and JSON format to
represent the documents.</p>
        <p>Similarly, Banker (K. Banker, 2011) provides
some ideas of data modeling, but limited to
MongoDB database and always referring to JSON (D.</p>
        <p>Crockford, 2006) format as a modeling solution.</p>
        <p>
          Kaur and Rani
          <xref ref-type="bibr" rid="ref1">(K. Kaur, K.Rani, 2013)</xref>
          present
a work for modeling and querying data in NoSQL
databases, especifically present a case study for
document-oriented and graph based data model.
        </p>
        <p>In the case of document-oriented propose a
data modeling restricted to MongoDB document
database, describing the data model by UML
diagram class to represent documents.
3</p>
        <p>Data Modeling For</p>
        <p>
          Document-Oriented Database
An important step in database implementation is
the data modeling, because it facilitates the
understanding of the project through key features
that can prevent programming and operation
errors. For relational databases, the data
modeling uses the Entity-Relationship Model (Elmasri
and Navathe, 2010). For NoSQL, it depends on
the database category. The focus of this article is
NoSQL document-oriented databases, where the
data format of these documents can be JSON,
BSON, or XML
          <xref ref-type="bibr" rid="ref2">(S. J. Pramod, 2012)</xref>
          .
        </p>
        <p>Basically, the documents are stored in
collections. A parallel is made with relational databases,
the equivalent for a collection is the record
(tuple) and for a document it is the relation (table).</p>
        <p>Documents can store completely different sets of
attributes, and can be mapped directly to a file
format that can be easily manipulated by a
programming language. However, it is difficult to abstract
the modeling of documents for the entity
relationship model (R. F. Lans, 2008).</p>
        <p>Modeling Paradigm for
document-oriented Database
The relational model designed for SQL has some
important features such as integrity, consistency,
type validation, transactional guarantees, schemes
and referential integrity. However, some
applications do not need all of these features. The
elimination of these resources has an important
influence on the performance and scalability of data
storage, bringing new meaning to data modeling.</p>
        <p>
          Document-oriented databases have some
significant improvements, e.g., index management by
the database itself, flexible layouts and advanced
indexed search engines (H. Scalable, 2015). By
associating these improvements (some being
denormalization and aggregation) to the basic
principles of data modeling in NoSQL, it is
possible to identify some generic modeling standards
associated to document-oriented databases.
Analyzing the documentation of the main
documentoriented databases, MongoDB
          <xref ref-type="bibr" rid="ref3">(MongoDB, 2015)</xref>
          and CouchDB
          <xref ref-type="bibr" rid="ref4">(CouchDB, 2015)</xref>
          , similar
representations of data mapping relationships can be
found: References and Embedded Documents,
a structure which allows associating a document
to another, retaining the advantage of specific
performance needs and data recovery standards.
This type of relationship stores the data by
including links or references, from one document to
another. Applications can solve these references to
access the related data in the structure of the
document itself
          <xref ref-type="bibr" rid="ref3">(MongoDB, 2015)</xref>
          . Figure 1 shows
two documents one of them for Fastq files and the
other to Activities.
3.3
This type of relationship stores in a single
document structure, where the embedded documents
are disposed in a field or an array. These
denormalized data models allow data manipulation in
a single database transaction
          <xref ref-type="bibr" rid="ref3">(MongoDB, 2015)</xref>
          .
Figure 2 shows a document of a genome Project
with a Activity embedded document.
4
        </p>
        <p>Proposal For Document-Oriented</p>
        <p>Databases Viewing
Unlike traditional relational databases that have
a simple form in the disposition in rows and
columns, a document-oriented database stores
information in text format, which consists of
collections of records organized in key-value concept, ie,
for each value represented a name (or label) is
assigned, which describes its meaning. This storage
model is known as JSON object, and the objects
are composed of multiple name/value pairs for
arrays, and other objects.</p>
        <p>In this scenario, the number of objects in a
database increases the abstraction complexity of
the logical relationship between the stored
information, especially when objects have references
to other objects. Currently, there is a lack of
solutions to conceptually represent those associated
with a NoSQL document-oriented database. As
described in (R. Arora and R. Aggarwal, 2013),
there is no standard to represent this kind of object
modeling, several diferent manners of modeling
may arise, depending on each data administrator’s
understanding, which makes learning difficult for
those who need to read the database model.</p>
        <p>Therefore, this section proposes a standard for
document-oriented database viewing. Our
proposal has some properties, considering the
conceptual representation modeling type, such as:
• Ensuring a single way of modeling for
the several NoSQL document-oriented
databases.
• Simplifying and facilitating the
understanding of a document-oriented database through
its conceptual model, leveraging the
abstraction and making the correct decisions about
the data storage.
• Providing an accurate, unambiguous and
concise pattern, so that database
administrators have substantial gains in abstraction,
understanding.
• Presenting different types of relationships
between collections defined as References and</p>
        <p>Embedded documents.
• Assisting the recognition and arrangement of
the objects, as well as its features and
relationships with other objects.</p>
        <p>The following subsections present the concepts
and graphing to build a conceptual model for
NoSQL document-oriented databases.
Before starting the discussion about the approach
of each type of the conceptual modeling
representation, it is important to highlight some
basic concepts about objects and relationships in a
document-oriented database:
• A document (or object) describes a set of
attributes that have their properties organized
in a key-value structure.
• Information contained in an document is
described by the identifier (key) and the value
associated with the key.
• Different types of relationships between
documents are defined as References and
Embedded Documents
• Because NoSQL is a non-relational data
database, the concepts of normalization, do
not apply.
• Some concepts of relationships between
objects are similar to ER modeling, such as
cardinality (one-to-one, one-to-many,
many-tomany).
The proposed solution for a conceptual modeling
to the NoSQL document-oriented databases has
two basic concepts: Document and Collections.</p>
        <p>As noted previously, a document is usually
represented by the structure of a JSON object, and as
many fields as needed may be added to the
document. For this solution, a document and a
collection of documents is represented by Figure 3.</p>
        <p>Embedded Documents 1..N
A one-to-many relationship in embedded
documents is represented by the Figure 5. This is the
case when the notation to represent the cardinality
is the same used in UML (F. Booch et al., 2005)
and is placed in the upper right corner of the
embedded documents. According to the cardinality
one-to-many the larger document has embedded
multiple documents within it.</p>
        <p>The following section presents the definitions of
relationship types and degrees for the objects
features.</p>
        <p>Embedded Documents 1..1
This section proposes a model that represents the
one-to-one relationship for documents embedded
in another document. In this case, the proposal
is to use the representation of an individual
Document within another element that represents a
Document. In Figure 4, cardinality is also
suggested to specify the one-to-one relationship type.
A many-to-many relationship in embedded
documents is represented by the Figure 6. According to
the cardinality many-to-many the larger document
has a many to many relationship with the
embedded document. The representation of the
cardinality is the same used in UML (F. Booch et al.,
2005).
A document can reference another, and in this
case, one must use an arrow directed to the
referenced document, as shown in Figure 7. One can
see that the directed arrow makes the left
document references to the right document.
Furthermore, the cardinality of the relationship should be
specified above the arrow. The notation of
cardinality is based on UML (F. Booch et al., 2005).
In NoSQL, a document can reference multiple
documents. To represent this relationship one
should use an arrow directed to the referenced
documents, as shown in Figure 8. The left document
references multiple documents on the right side,
by the directed arrow. Furthermore, the
cardinality of the relationship is represented by the
notation ”1..N” as in UML (F. Booch et al., 2005).
5
In order to evaluate our proposal, part of the
workflow described in (J. C. Marioni et al., 2014) was
used. This workflow aimed at identify and
comparing expression levels of human kidney and liver
RNA samples sequenced by Illumina. The
workflow was designed in three phases (Figure. 10):
• Filtering: all the sequenced transcripts were
filtered, generating new files with good
quality sequences.
• Alignment: transcripts were mapped to the</p>
        <p>human genome used as reference.
• Statistical Analysis: a sort process was first
executed, followed by a statistical
analysis with the mapped transcripts to discover
which genes are mostly expressed both in
kidney and liver samples.
To represent this relationship a bidirectional arrow
is used between reference documents, as shown
in Figure 9. The left document references
multiple documents on the right side and the right
document references multiple documents on the left
side. Furthermore, the cardinality of the
relationship is represented by the notation ”N..N” as in
UML (F. Booch et al., 2005).</p>
        <p>After analyzing the previously mentioned
concepts, we have chosen to create a collection of
documents for each PROV-DM type used to
create a graph node. We also defined a collection
for genomic documents (raw data). The reference
relationship approach was chosen to connect all
PROV-DM components, complementary
information of PROV-DM and genomic documents. Based
on (R. de Paula et al., 2013) we defined the
documents and the attributes. A set of minimum
information related to each one of these entities.</p>
        <p>Figure. 11 shows our document based data
representation, explained as follows:</p>
        <p>• Project: stores different experiments of one
agent. Attributes: Id, name, description,
coordinator, start date, end date and
observation.
• FASTAQ: files used or generated in the
activ</p>
        <p>ity; Attributes: Id, filename and description.
• Activity: represents the execution of a
program; Attributes: Id, name, program, version
program, command line, function, start date,
end date, account ID, used (name FASTQ,
local, size), wasGenerateBy (name FASTQ,
local, size), and wasAssociatedWith (Agent
name).
• Account: represents the performance of an
experiment; Attributes: Id, name,
description, execution place, star date, end date,
observation, version and version date and
project Id.
• Agent: represents the person responsible for
a program or a phase in the workflow.
Attributes: Id, name, login and password.
5.1 Implementation
In this case study, we have considered the
MongoDB NoSQL database to store provenance and
data files. The primary motivation for this choice
was MongoDB’s ability to manipulate large
volumes of data. MongoDB is an open-source
Document-Oriented database designed to store
large amounts of data from multiple servers.</p>
        <p>
          It uses JSON- style documents with dynamic
schemas. The number of fields, content and size
of the document can differ from one document to
another. In practice, however, the documents in
a collection share a similar structure
          <xref ref-type="bibr" rid="ref3">(MongoDB,
2015)</xref>
          and can be mapped directly to a file format
that can be easily manipulated by a programming
language.
        </p>
        <p>
          MongoDB documents have a maximum size of
16MB. This feature is important to ensure that
a single document cannot use excessive amounts
of RAM. In order to store files larger than the
maximum size, MongoDB provides a GridFS API
          <xref ref-type="bibr" rid="ref3">(MongoDB, 2015)</xref>
          . It automatically divides large
data into 256 KB pieces and maintains metadata
for all pieces. GridFS allows for the retrieval of
individual pieces as well as entire documents.
        </p>
        <p>GridFS uses two collections to store the data:
fs.files collections, containing metadata about
files, and fs.chunks collections, which store the
actual 256k data chunks. The collections FS.file
contains the name of the FASTQ file. Thus, it
was possible to implement the relationship
between MongoDB Collection Activity using
Reference Document. In other words, we implemented
the connection between Level 1 and Level 2
through the File Name attribute that was present in
fileprovenance.files and Activity Collection.
Figure. 12 illustrates this particular implementation.
In contrast to relational database management
systems, NoSQL databases are designed to be
schemaless and flexible. Therefore, the challenge
of this work was to introduce a data modeling
standard for NoSQL document-oriented databases, in
contrast to the original idea for NoSQL databases.</p>
        <p>The objective was to build compact, clear and
intuitive diagrams for conceptual data modeling
for NoSQL databases. While the current
studies propose generic techniques and do not define
a specific modeling engine to NoSQL database,
our idea was to present a graphical model for any
NoSQL document-oriented database. Moreover,
while other studies describe techniques based on
UML Diagram Class and JSON format as a
modeling solution, we have a new approach to solve
the conceptual data modeling issue for NoSQL
document-oriented databases.</p>
        <p>Future work includes: verifying our model for
other NoSQL database classifications, such as
key-value and column.
R. Elmasri and S. Navathe. 2010. Fundamentals of</p>
        <p>Database Systems. Pearson Addison Wesley.</p>
        <p>MongoDB. 2015. Document database. [Online]</p>
        <p>Available: http://www.mongodb.org/
[Retrieved:April, 15].</p>
        <p>J. Partner, A. Vukotic, and N. Watt. 2013. Neo4j in</p>
        <p>Action, O’Reilly Media.</p>
        <p>D. Borthakur et al. 2011. Apache hadoop goes
realtime at facebook, in Proceedings of the 2011 ACM
SIGMOD International Conference on Management
of data. ACM, 2011, pp. 1071–1080.</p>
        <p>F. Chang et al. 2008. “Bigtable: A distributed storage
system for structured data, ACM Transactions on
Computer Systems (TOCS), vol. 26, no. 2, 2008, p.</p>
        <p>4.</p>
        <p>R. F. Lans. 2008. Introduction to SQL: mastering the
relational database language, Addison-Wesley
Professional.</p>
        <p>H. Scalable. 2015. Nosql data
modeling techniques. [Online]
Available: http://highlyscalable.
wordpress.com/2012/03/01/
nosql-data-modeling-techniques/
[Retrieved:April, 15].</p>
        <p>R. Arora and R. Aggarwal, 2013. Modeling and
querying data in mongodb, International Journal of
Scientific and Engineering Research (IJSER 2013), vol.</p>
        <p>4, no. 7, Jul. 2013, pp. 141–144.</p>
        <p>K. Banker, 2011. MongoDB in action, Manning
Pub</p>
        <p>lications Co.</p>
        <p>D. Crockford, 2006. RFC 4627 (Informational) The
application json Media Type for JavaScript Object
Notation (JSON), IETF (Internet Engineering Task
Force)
G. Booch, J. Rumbaugh, and I. Jacobson, 2005. The
unified modeling language user guide. , Pearson
Education India.</p>
        <p>J. C. Marioni, C. E. Mason, S. M. Mane, M. Stephens,
and Y. Gilad, 2014. RNA-SEQ: An assessment of
technical reproducibility and comparison with gene
expression arrays, Genome Research, vol. 18, no. 9,
pp. 1509–1517.</p>
        <p>R. de Paula, M. Holanda, L. SA Gomes, S. Lifschitz
and M. E. MT. Walter, 2013. Provenance in
bioinformatics workflows. , BMC Bioinformatics 14
(Suppl 11):S6.</p>
        <p>Hipi, as alternative for satellite images processing
Wilder Nina Choquehuayta</p>
        <p>UNSA / Arequipa
wninac@unsa.edu.pe</p>
        <p>Rene´ Cruz Mu n˜oz</p>
        <p>UNSA / Arequipa
rcruzm@unsa.edu.pe</p>
        <p>Juber Serrano Cervantes</p>
        <p>UNSA / Arequipa
jserranoc@unsa.edu.pe
Alvaro Mamani Aliaga</p>
        <p>UNSA / Arequipa
amamani@unsa.edu.pe</p>
        <p>Pablo Yanyachi</p>
        <p>UNSA / Arequipa
raulpab@unsa.edu.pe</p>
        <p>Yessenia Yari</p>
        <p>UNSA / Arequipa
yyari@unsa.edu.pe
These days, in different fields of both
industry and academia, large amounts of
data is generated. The use of several
frameworks with different techniques is
essential , for processing and extraction of
data. In the remote sensing field, large
volumes of data are generated (satellite
images) over short periods of time.
Information systems for processing these kind
of images were not designed with scalable
features. In this paper, we present an
extension of the HIPI framework (Hadoop
Image Processing Interface) for
processing satellite image formats.</p>
        <p>KeyWords: Big Data, Remote Sensing,
Hadoop, HIPI, MapReduce, Satellite Images
1 Introduction
The field of remote sensing is helpful in different
areas of both industry and academia, because it
uses images of the earth’s surface that are acquired
from different sources like antennas and satellites,
which provide increasingly better image
resolution as technology advances. Nowadays, several
open source frameworks are available to process
big data, such as Hadoop (Shvachko et al., 2010),
H2O (0xdata:H2O, 2015), Spark (Apache:Spark,
2015), etc., which are used for distributed and
parallel processing of large volumes of data. HIPI
(Hadoop Image Processing Interface) is an
image processing library designed to be used with
the Apache Hadoop MapReduce parallel
programming framework. HIPI facilitates efficient and
high-throughput image processing with
MapReduce style parallel programs typically executed on
a cluster. In the present work the HIPI library was
modified, giving additional functionalities to read
and process the GeoTIFF format (format provided
by USGS -United States Geological Survey- for
Landsat satellite images).</p>
        <p>State of the Art
There are different techniques used for image
classification in semantic taxonomy categories such as
vegetation, water, etc. (Codella et al., 2011),
however these methods don’t consider scalability as
part of its solutions, Noel C. F. Codella et. al.. In
Wanfeng Zhang, et. al. (Zhang et al., 2013) An
infrastructure for massive processing of satellite
images in a multi-dataCenter environment,
consisting of a DataCenter, where Access Security,
Information Service and strategy Scheduling for
data management us introduced. It’s important to
consider HIPI (Sweeney et al., 2011) as a state of
the art, extensible library for image processing and
computer vision applications, which helps to avoid
the problem of small files and achieving
improvements in memory and response time.</p>
        <p>Proposal
That is why this paper is a modification of HIPI, to
extends its functionality to work with TIFF images
or GeoTIFF type. To achieve this, we proceeded
as follows:
• It was decided to use the Tiff format from
satellite images obtained from the USGS
since this format unlike others has no
compression or data loss (Adobe, 1992).
• JAI API was chosen to read and write the
chosen format, JAI has more codecs and
features available that can be useful for reading
multiple formats.
• Classes needed to upload, encode and decode</p>
        <p>images of Tiff type were modified.</p>
        <p>Based on the tests, the possibility of using HIPI for
processing multispectral and hyperspectral images
was analyzed. For such images, operations such as
PCA are important to process all spectral bands, it
is proposed to keep them in the same zip file, or
as different images belonging to a tiff format, and
then decode and interpret as a conventional
multiband image.
4</p>
        <p>Experiments
The experiments were performed on a Local
Heterogeneous Cluster depicted in Table 1 where the
characteristics of slaves and master are shown. We
used satellite images from LandSat 7, only
considering the first 4 bands so we compressed a satellite
image in .zip then we used the .zip up to 0.5GB,
1GB, 5GB and 10GB. The algorithm was tested
about the average of channels which is explained
in the official website of HIPI. In each task map,
we iterated over each read of band of satellite
image as FloatImage, added each value of pixel
depending of channel then divided for number of
pixels (width x height) and returned the key of the
satellite image and array of data calculated. In
each reduce task, we only calculated the average
of the average of channels from each satellite
image.</p>
        <p>Node
master
slave 1
slave 2
slave 3
characteristics
Core i7, RAM 8GB, Disk 100GB,
S.O Ubuntu 64 bits
Core i7, RAM 8GB, Disk 100GB,
S.O Ubuntu 64 bits
Core 2 Duo, RAM 4GB, Disk
100GB, S.O Ubuntu 64 bits
Core 2 Duo, RAM 4GB, Disk
100GB, S.O Ubuntu 64 bits</p>
        <p>In Table 2 shows in axis x the amount of data in
GBs and in y axis the execution time. The Hadoop
configuration was 1 replication of data, chunks of
32MB, 64MB and 128MB, 4096MB in memory
for task reduce and map. The java virtual
machine for task reduce and map was configured with
4096MB at the most.
5
Based on the review conducted and theoretical
experimental tests HIPI modified version of the
article concludes as follows: It is possible to perform
various image processing operations, such as
fil4,500
4,000
]S3,500
[
e 3,000
m
ti 2,500
n
tio 2,000
cu 1,500
e
xE1,000
350000
50
1 2 3 4 5 6 7 8 9 10 11 12</p>
        <p>Amount of data [GB]
Table 2: Execution time vs Amount of data spatial
ters, variance, clustering or dimensionality
reduction by using the MapReduce algorithm and also
while the information in compressed format
occupies less space, this does not necessarily mean
faster times when processing, since the matrix
calculations are done on the same decompression.</p>
        <p>References
0xdata:H2O. 2015. H2o@ONLINE.
Website. 0xdata:H2O, In: http://0xdata.com/
product/,accessed(may 2015).</p>
        <p>Apache:Spark. 2015. H2o@ONLINE.
Website. Spark, Lightning-fast cluster
computing , In : https://spark.apache.org,
accessed(accessed(2015-05-20)).</p>
        <p>Noel C.F. Codella, Gang Hua, Apostol Natsev, and</p>
        <p>John R Smith. 2011. Towards large scale land-cover
recognition of satellite images. In Information,
Communications and Signal Processing (ICICS)
2011, 8th International Conference on, pages 1–5.</p>
        <p>IEEE.</p>
        <p>Konstantin Shvachko, Hairong Kuang, Sanjay Radia,
and Robert Chansler. 2010. The hadoop distributed
file system. In Mass Storage Systems and
Technologies (MSST), 2010 IEEE 26th Symposium on, pages
1–10. IEEE.</p>
        <p>Chris Sweeney, Liu Liu, Sean Arietta, and Jason</p>
        <p>Lawrence. 2011. Hipi: a hadoop image
processing interface for image based mapreduce tasks.</p>
        <p>Chris,University of Virginia.</p>
        <p>Wanfeng Zhang, Lizhe Wang, Dingsheng Liu, Weijing</p>
        <p>Song, Yan Ma, Peng Liu, and Dan Chen. 2013.
Towards building a multi-datacenter infrastructure for
massive remote sensing image processing.
Concurrency and Computation: Practice and Experience,
25(12):1798–1812.</p>
        <p>Multi-agent system for usability improvement of a university
administrative system.</p>
        <p>Jorge Leoncio Guerra Guerra</p>
        <p>Universidad Nacional Mayor de San Marcos
Facultad de Ingeniería de Sistemas e Informática
Laboratorio de Robótica e Internet de las Cosas
jguerrag@unmsm.edu.pe</p>
        <p>Félix Armando Fermín Pérez</p>
        <p>Universidad Nacional Mayor de San Marcos
Facultad de Ingeniería de Sistemas e Informática
Laboratorio de Robótica e Internet de las Cosas</p>
        <p>fferminp@unmsm.edu.pe</p>
        <p>Abstract
The implementation of multi-agent systems is one
of the specific ways to integrate heterogeneous
information systems, by creating a software agent
that performs specific tasks within a field of known
action, as in the case of the travelling salesman
problem or an agent searching clinical data of a
patient. In this paper, the development of a mobile
agent is proposed for the implementation of
university administrative services across
heterogeneous servers, interacting in turn with
other intelligent agents in an heterogeneous
communication environment. This will generate a
single-access interface, which will serve to
improve the usability of the system when the client
makes an administrative request regardless of the
server on which the requested procedure is
processed.</p>
        <p>Keywords: Integration, usability, agents,
multiagent systems, heterogeneous services.
1. Introduction
The portal of the University Inca Garcilaso de
la Vega (http://www.uigv.edu.pe) is a web
application that offers different services
available in each of the offices of the
institution, which is accessed through
hyperlinks from Faculties, Research Units,
Distance Education, etc.</p>
        <p>An integrated system by software agents
(Shiao, 2004); in which, with a single
interface, a user can access different services
without having to go through different pages
to finish their request; will make possible to
link these services regardless of the Faculty
where they are and make the process requested
by the user in a transparent manner, so that the
usability of the system is also improved.
2. Proposed implementation
The structure of the proposed multi-agent
system is shown in Figure 1.</p>
        <p>Figure 1 – Proposed multi-agent system
Using the GAIA methodology (Zambonelli et
al, 2003) for modeling a system of agents,
agents that will be built on the prototype are
defined:
a.- Client Agent: captures requirements or
service requests from users; interacts with the
mobile agent to process the request and sends
the response to the client system for viewing
or joining the client process.
b.- Server Agent: interacts with the server of
the corresponding office to transfer the request
from the mobile agent, processes the requested
service, and sends the response to the mobile
agent.
c.- AGMER Agent: mobile agent moving
through JADE middleware to interact with
clients and servers, has the ability to divide
received requests into subtasks performed by
different servers.</p>
        <p>One advantage of using GAIA is that it can be
combined with diagrams from UML as
described in Kang et al (2004), for this reason
UML diagrams and GAIA models
(Wooldridge et al is used., 2000) will be used
for its construction.</p>
        <p>Once developed, the model
implemented as shown in Figure 2.</p>
        <p>will</p>
        <p>be</p>
        <p>In the development process is suggested the
use of Eclipse Luna, JDK 8.0_41 and JADE
4.3.2 for creating agents. In the prototype, it
was initially considered the access to these
basic services for a university student from the
case study:
a. Specialized Library System with Struts 2,
running on Google App Engine cloud, a PAAS
solution.
b. PRODI (English Learning Program)
registration system with Node.JS, OpenShift,
a PAAS solution.</p>
        <p>All modules will communicate via web
services, with interface agents using AXIS 2.
3. Conclusions and future work
The single-interface model for administrative
procedure systems is suitable to improve the
usability of a web portal to reduce the number
of interactions.</p>
        <p>A multi-agent system allows an heterogeneous
implementation of components as well as
asynchronous communication similar to a
request of a university administrative
procedure.</p>
        <p>JADE is a suitable platform for developing
agents due to its ease of use, and adaptability
to development methodologies of multi-agent
systems.</p>
        <p>As future work it is proposed to develop
indicators to measure implemented services,
the frequency of use of the services, and
others, that will be collected and stored in the
cloud for future analysis using Big Data tools.
4. References
Dan Shiao. 2004. Mobile Agent: New Model of
Intelligent Distributed Computing. IBM China,
October, 2004.</p>
        <p>Franco Zambonelli, Nicholas Jennings, Michael
Wooldridge. 2003. Developing Multiagent
Systems: The Gaia Methodology. ACM
Transactions on Software Engineering and
Methodology. 12(3): 317–370.</p>
        <p>Miao Kang, Lan Wang, Kenji Taguchi. 2004.</p>
        <p>Modelling Mobile Agent Applications in UML2.0
Activity Diagrams, Proceedings of the 6th
International Conference on Enterprise
Information Systems. Porto. Portugal. Vol. IV:
519-522.</p>
        <p>Michael Wooldridge, Nicholas Jennings, David
Kinny. 2000. The Gaia Methodology for
AgentOriented Analysis and Design. Journal of
Autonomous Agents and Multi-Agent Systems,
vol. 15. Kluwer Academic Publishers, Boston. The
Netherlands.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>K.</given-names>
            <surname>Kaur</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Rani</surname>
          </string-name>
          ,
          <year>2013</year>
          .
          <article-title>Modeling and querying data in NoSQL databases</article-title>
          , In Big Data, IEEE International Conference on (pp.
          <fpage>1</fpage>
          -
          <lpage>7</lpage>
          ). IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>S. J.</given-names>
            <surname>Pramod</surname>
          </string-name>
          ,
          <year>2012</year>
          .
          <article-title>Nosql distilled: A brief guide to the emerging world of polyglot persistence,</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>MongoDB</surname>
          </string-name>
          ,
          <year>2015</year>
          .
          <article-title>Data modeling introduction</article-title>
          , Online Available: http: //docs.mongodb.org/manual/core/ data-modeling-introduction/ [Retrieved: April,
          <volume>15</volume>
          ].
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>CouchDB</surname>
          </string-name>
          ,
          <year>2015</year>
          .
          <article-title>Modeling entity relationships in couchdb</article-title>
          , [Online]. Available: http://wiki.apache.org/couchdb/ [retrieved: April, 15]
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>