<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>1st Symposium on Information Management and Big Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Cusco - Peru</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Proceeding Editors: J. A. Lossio-Ventura</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>H. Alatrista-Salas</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2008</year>
      </pub-date>
      <fpage>695</fpage>
      <lpage>699</lpage>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>PROCEEDINGS
 
TABLE OF CONTENTS
SIMBig 2014 Conference Organization Committee
SIMBig 2014 Conference Reviewers</p>
    </sec>
    <sec id="sec-2">
      <title>SIMBig 2014 Paper Contents</title>
      <p>SIMBig 2014 Keynote Speaker Presentations</p>
    </sec>
    <sec id="sec-3">
      <title>Full and Short Papers</title>
      <p>The SIMBig 2014 Organization Committee confirms that full and concise papers
accepted for this publication:
• Meet the definition of research in relation to creativity, originality, and increasing
humanity's stock of knowledge;
• Are selected on the basis of a peer review process that is independent, qualified
expert review;
• Are published and presented at a conference having national and international
significance as evidenced by registrations and participation; and
• Are made available widely through the Conference web site.
SIMBig 2014 Organization Committee
 </p>
    </sec>
    <sec id="sec-4">
      <title>GENERAL CHAIR</title>
      <p>• Juan Antonio LOSSIO VENTURA, Montpellier 2 University, LIRMM,</p>
      <p>Montpellier, France
• Hugo ALATRISTA SALAS, UMR TETIS, Irstea, France
LOCAL CHAIR
• Cristhian GANVINI VALCARCEL Andina University of Cusco, Peru
• Armando FERMIN PEREZ, National University of San Marcos, Peru
• Cesar A. BELTRAN CASTAÑON, GRPIAA, Pontifical Catholic University
of Peru
• Marco RIVAS PEÑA, National University of San Marcos, Peru
SIMBig 2014 Conference Reviewers
•
•
•
•
•
•
•
•
•
•
•
•
•
•
•
•
•
•
•
•
•
•
•
•
•
•
•
•
•
•
•</p>
      <p>Salah Ait-Mokhtar, Xerox Research Centre Europa, FRANCE
Jérôme Azé, LIRMM - University of Montpellier 2, FRANCE
Cesar A. Beltrán Castañón, GRPIAA - Pontifical Catholic University of Peru,
PERU
Sandra Bringay, LIRMM - University of Montpellier 3, FRANCE
Oscar Corcho, Ontology Engineering Group - Polytechnic University of Madrid,
SPAIN
Gabriela Csurka, Xerox Research Centre Europa, FRANCE
Mathieu d'Aquin, Knowledge Media institute (KMi) - Open University, UK
Ahmed A. A. Esmin, Federal University of Lavras, BRAZIL
Frédéric Flouvat, PPME Labs - University of New Caledonia, NEW CALEDONIA
Hakim Hacid, Alcatel-Lucent, Bell Labs, FRANCE
Dino Ienco, Irstea, FRANCE
Clement Jonquet, LIRMM - University of Montpellier 2, FRANCE</p>
      <sec id="sec-4-1">
        <title>Eric Kergosien, Irstea, FRANCE</title>
        <p>Pierre-Nicolas Mougel, Hubert Curien Labs - University of Saint-Etienne,
FRANCE
Phan Nhat Hai, Oregon State University, USA
Thomas Opitz, University of Montpellier 2 - LIRMM, FRANCE
Jordi Nin, Barcelona Supercomputing Center (BSC) - Technical University of
Catalonia (BarcelonaTECH), SPAIN
Gabriela Pasi, Information Retrieval Lab - University of Milan Bicocca, ITALY
Miguel Nuñez del Prado Cortez, Intersec Labs - Paris, FRANCE
José Manuel Perea Ortega, European Commission - Joint Research Centre
(JRC)- Ispra, ITALY
Yoann Pitarch, IRIT - Toulouse, FRANCE
Pascal Poncelet, LIRMM - University of Montpellier 2, FRANCE</p>
      </sec>
      <sec id="sec-4-2">
        <title>Julien Rabatel, LIRMM, FRANCE</title>
        <p>Mathieu Roche, Cirad - TETIS and LIRMM, FRANCE
Nancy Rodriguez, LIRMM - University of Montpellier 2, FRANCE
Hassan Saneifar, Raja University, IRAN
Nazha Selmaoui-Folcher, PPME Labs - University of New Caledonia, NEW
CALEDONIA</p>
      </sec>
      <sec id="sec-4-3">
        <title>Maguelonne Teisseire, Irstea - LIRMM, FRANCE</title>
        <p>Boris Villazon-Terrazas, iLAB Research Center - iSOCO, SPAIN
Pattaraporn Warintarawej, Prince of Songkla University, THAILAND
Osmar Zaïane, Department of Computing Science, University of Alberta,
CANADA
SIMBig 2014 Paper Contents</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Title and Authors</title>
      <sec id="sec-5-1">
        <title>Keynote Speaker Presentations</title>
        <p>Pascal Poncelet
Mathieu Roche
Hakim Hacid
Mathieu d’Aquin</p>
      </sec>
      <sec id="sec-5-2">
        <title>Long Papers</title>
        <p>Identification of Opinion Leaders Using Text Mining Technique in
Virtual Community, Chihli Hung and Pei-Wen Yeh
Quality Metrics for Optimizing Parameters Tuning in Clustering
Algorithms for Extraction of Points of Interest in Human Mobility,
Miguel Nunez Del Prado Cortez and Hugo Alatrista-Salas
A Cloud-based Exploration of Open Data: Promoting Transparency
and Accountability of the Federal Government of Australia, Richard
Sinnott and Edwin Salvador
Discovery Of Sequential Patterns With Quantity Factors, Karim
Guevara Puente de La Vega and César Beltrán Castañón
Online Courses Recommendation based on LDA, Rel Guzman
Apaza, Elizabeth Vera Cervantes, Laura Cruz Quispe and Jose
Ochoa Luna</p>
      </sec>
      <sec id="sec-5-3">
        <title>Short Papers</title>
        <p>ANIMITEX project: Image Analysis based on Textual Information,
Hugo Alatrista-Salas, Eric Kergosien, Mathieu Roche and
Maguelonne Teisseire
A case study on Morphological Data from Eimeria of Domestic
Fowl using a Multiobjective Genetic Algorithm and R&amp;P for
Learning and Tuning Fuzzy Rules for Classification, Edward
Hinojosa Cárdenas and César Beltrán Castañón
SIFR Project: The Semantic Indexing of French Biomedical Data
Resources, Juan Antonio Lossio-Ventura, Clement Jonquet,
Mathieu Roche and Maguelonne Teisseire</p>
      </sec>
      <sec id="sec-5-4">
        <title>Spanish Track</title>
        <p>Mathematical modelling of the performance of a computer system,
Felix Armando Fermin Perez
Clasificadores supervisados para el análisis predictivo de muerte y
sobrevida materna, Pilar Hidalgo
Page
7
8
14
22
33
42
49
53
58
62
67
SIMBig 2014 Keynote Speakers
Mining Social Networks: Challenges and Directions
Pascal Poncelet, Lirmm, Montpellier, France
Professor Pascal Poncelet talk was focused on analysis of data associate to
social network. In his presentation, Professor Poncelet summarized several
techniques through knowledge discovery on databases process.
NLP approaches: How to Identify Relevant Information in</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Social Networks?</title>
      <p>Mathieu Roche, UMR TETIS, Irstea, France
In this talk, doctor Roche outlined last tendencies in text mining task. Several
techniques (sentiment analysis, opinion mining, entity recognition, ...) on
different corpus (tweets, blogs, sms,...) were detailed in this presentation.
Social Network Analysis: Overview and Applications
Hakim Hacid, Zayed University, Dubai, UAE
Doctor Hacid presented a complete overview concerning techniques
associated to explote social web data. Techniques as social network analysis,
social information retrieval and mashups full completion were presented.
Putting intelligence in Web Data With Examples Education
Mathieu d’Aquin, Knowledge Media Institute, The Open University, UK
Doctor d'Aquin presented several web mining approaches addressed to
education issues. In this talk doctor d'Aquin detailed themes like intelligent
web information and knowledge processing, the semantic web, among others.
Identification of Opinion Leaders Using Text Mining Technique in</p>
      <p>Virtual Community</p>
      <p>Chihli Hung
Department of Information Management</p>
      <p>Chung Yuan Christian University</p>
      <p>Taiwan 32023, R.O.C.</p>
      <p>Pei-Wen Yeh
Department of Information Management</p>
      <p>Chung Yuan Christian University</p>
      <p>Taiwan 32023, R.O.C.
Word of mouth (WOM) affects the buying
behavior of information receivers stronger than
advertisements. Opinion leaders further affect
others in a specific domain through their new
information, ideas and opinions. Identification
of opinion leaders has become one of the most
important tasks in the field of WOM mining.</p>
      <p>Existing work to find opinion leaders is based
mainly on quantitative approaches, such as
social network analysis and involvement.</p>
      <p>Opinion leaders often post knowledgeable and
useful documents. Thus, the contents of WOM
are useful to mine opinion leaders as well. This
research proposes a text mining-based approach
to evaluate features of expertise, novelty and
richness of information from contents of posts
for identification of opinion leaders. According
to experiments in a real-world bulletin board
data set, this proposed approach demonstrates
high potential in identifying opinion leaders.
1</p>
      <p>Introduction
This research identifies opinion leaders using the
technique of text mining, since the opinion leaders
affect other members via word of mouth (WOM)
on social networks. WOM defined by Arndt (1967)
is an oral person-to-person communication means
between an information receiver and a sender, who
exchange the experiences of a brand, a product or a
service based on a non-commercial purpose.</p>
      <p>Internet provides human beings with a new way of
communication. Thus, WOM influences the
consumers more quickly, broadly, widely,
significantly and consumers are further influenced
by other consumers without any geographic
limitation (Flynn et al., 1996).</p>
      <p>Nowadays, making buying decisions based on
WOM becomes one of collective decision-making
strategies. It is nature that all kinds of human
groups have opinion leaders, explicitly or
implicitly (Zhou et al., 2009). Opinion leaders
usually have a stronger influence on other
members through their new information, ideas and
representative opinions (Song et al., 2007). Thus,
how to identify opinion leaders has increasingly
attracted the attention of both practitioners and
researchers.</p>
      <p>
        As opinion leadership is relationships between
members in a society, many existing opinion leader
identification tasks define opinion leaders by
analyzing the entire opinion network in a specific
domain, based on the technique of social network
analysis (SNA)
        <xref ref-type="bibr" rid="ref11 ref23">(Kim, 2007; Kim and Han, 2009)</xref>
        .
      </p>
      <p>This technique depends on relationship between
initial publishers and followers. A member with
the greatest value of network centrality is
considered as an opinion leader in this network
(Kim, 2007).</p>
      <p>However, a junk post does not present useful
information. A WOM with new ideas is more
interesting. A spam link usually wastes readers'
time. A long post is generally more useful than a
short one (Agarwal et al., 2008). A focused
document is more significant than a vague one.</p>
      <p>That is, different documents may contain different
influences on readers due to their quality of WOM.</p>
      <p>WOM documents per se can also be a major
indicator for recognizing opinion leaders. However,
such quantitative approaches, i.e. number-based or
SNA-based methods, ignore quality of WOM and
only include quantitative contributions of WOM.</p>
      <p>
        Expertise, novelty, and richness of information
are three important features of opinion leaders,
which are obtained from WOM documents
        <xref ref-type="bibr" rid="ref11 ref23">(Kim
and Han, 2009)</xref>
        . Thus, this research proposes a text
mining-based approach in order to identify opinion
leaders in a real-world bulletin board system.
      </p>
      <p>Besides this section, this paper is organized as
follows. Section 2 gives an overview of features of
opinion leaders. Section 3 describes the proposed
text mining approach to identify opinion leaders.</p>
      <p>Section 4 describes the data set, experiment design
and results. Finally, a conclusion and further
research work are given in Section 5.
2</p>
      <p>Features of Opinion Leaders
The term “opinion leader”, proposed by Katz and
Lazarsfeld (1957), comes from the concept of
communication. Based on their research, the
influence of an advertising campaign for political
election is lesser than that of opinion leaders. This
is similar to findings in product and service
markets. Although advertising may increase
recognition of products or services, word of mouth
disseminated via personal relations in social
networks has a greater influence on consumer
decisions (Arndt, 1967; Khammash and Griffiths,
2011). Thus, it is important to identify the
characteristics of opinion leaders.</p>
      <p>According to the work of Myers and Robertson
(1972), opinion leaders may have the following
seven characteristics. Firstly, opinion leadership in
a specific topic is positively related to the quantity
of output of the leader who talks, knows and is
interested in the same topic. Secondly, people who
influence others are themselves influenced by
others in the same topic. Thirdly, opinion leaders
usually have more innovative ideas in the topic.</p>
      <p>Fourthly and fifthly, opinion leadership is
positively related to overall leadership and an
individual’s social leadership. Sixthly, opinion
leaders usually know more about demographic
variables in the topic. Finally, opinion leaders are
domain dependent. Thus, an opinion leader
influences others in a specific topic in a social
network. He or she knows more about this topic
and publishes more new information.</p>
      <p>Opinion leaders usually play a central role in a
social network. The characteristics of typical
network hubs usually contain six aspects, which
are ahead in adoption, connected, travelers,
information-hungry, vocal, and exposed to media
more than others (Rosen, 2002). Ahead in adoption
means that network hubs may not be the first to
adopt new products but they are usually ahead of
the rest in the network. Connected means that
network hubs play an influential role in a network,
such as an information broker among various
different groups. Traveler means that network hubs
usually love to travel in order to obtain new ideas
from other groups. Information-hungry means that
network hubs are expected to provide answers to
others in their group, so they pursue lots of facts.</p>
      <p>Vocal means that network hubs love to share their
opinions with others and get responses from their
audience. Exposed to media means that network
hubs open themselves to more communication
from mass media, and especially to print media.</p>
      <p>Thus, a network hub or an opinion leader is not
only an influential node but also a novelty early
adopter, generator or spreader. An opinion leader
has rich expertise in a specific topic and loves to be
involved in group activities.</p>
      <p>
        As members in a social network influence each
other, degree centrality of members and
involvement in activities are useful to identify
opinion leaders
        <xref ref-type="bibr" rid="ref11 ref23">(Kim and Han, 2009)</xref>
        . Inspired by
the PageRank technique, which is based on the link
structure (Page et al., 1998), OpinionRank is
proposed by Zhou et al. (2009) to rank members in
a network. Jiang et al. (2013) proposed an
extended version of PageRank based on the
sentiment analysis and MapReduce. Agarwal et al.
(2008) identified influential bloggers through four
aspects, which are recognition, activity generation,
novelty and eloquence. An influential blog is
recognized by others when this blog has a lot of
inlinks. The feature of activity generation is
measured by how many comments a post receives
and the number of posts it initiates. Novelty means
novel ideas, which may attract many in-links from
the blogs of others. Finally, the feature of
eloquence is evaluated by the length of post. A
lengthy post is treated as an influential post.
      </p>
      <p>Li and Du (2011) determined the expertise of
authors and readers according to the similarity
between their posts and the pre-built term ontology.</p>
      <p>However both features of information novelty and
influential position are dependent on linkage
relationships between blogs. We propose a novel
text mining-based approach and compare it with
several quantitative approaches.
3</p>
      <p>Quality Approach-Text Mining
Contents of word of mouth contain lots of useful
information, which has high relationships with
important features of opinion leaders. Opinion
leaders usually provide knowledgeable and novel
information in their posts (Rosen, 2002; Song et al.,
2007). An influential post is often eloquent (Keller
and Berry, 2003). Thus, expertise, novelty, and
richness of information are important
characteristics of opinion leaders.
3.1</p>
      <p>Preprocessing
This research uses a traditional Chinese text
mining process, including Chinese word
segmenting, part-of-speech filtering and removal
of stop words for the data set of documents. As a
single Chinese character is very ambiguous,
segmenting Chinese documents into proper
Chinese words is necessary (He and Chen, 2008).</p>
      <p>This research uses the CKIP service
(http://ckipsvr.iis.sinica.edu.tw/) to segment
Chinese documents into proper Chinese words and
their suitable part-of-speech tags. Based on these
processes, 85 words are organized into controlled
vocabularies as this approach is efficient to capture
the main concepts of document (Gray et al., 2009).
3.2</p>
      <p>Expertise
This can be evaluated by comparing their posts
with the controlled vocabulary base (Li and Du,
2011). For member i, words are collected from his
or her posted documents and member vector i is
represented as fi=(w1, w2, …wj, …, wN), where wj
denotes the frequency of word j used in the posted
documents of user i. N denotes the number of
words in the controlled vocabulary. We then
normalize the member vector by his or her
maximum frequency of any significant word. The
degree of expertise can be calculated by the
Euclidean norm as show in (1).
expi 
fi ,
mi
where  is Euclidean norm.
3.3</p>
      <p>Novelty
We utilize Google trends service
(http://www.google.com/trends) to obtain the
firstsearch time tag for significant words in documents.</p>
      <p>Thus, each significant word has its specific time
tag taken from the Google search repository. For
example, the first-search time tag for the search
term, Nokia N81, is 2007 and for Nokia Windows
Phone 8 is 2011. We define three degrees of
novelty evaluated by the interval between the
firstsearch year of significant words and the collected
year of our targeted document set, i.e. 2010. This
significant word belongs to normal novelty if the
interval is equal to two years. A significant word
with an interval of less than two years belongs to
high novelty and one with an interval greater than
two years belongs to low novelty. We then
summarize all novelty values based on significant
words used by a member in a social network. The
equation of novelty for a member is shown in (2).</p>
      <p>e  0.66  em  0.33 el ,
novi  h</p>
      <p>eh  em  el
where eh , em and el is the number of words that
belong to the groups of high, normal and low
novelty, respectively.</p>
      <p>(2)
In general, a long document suggests some useful
information to the users (Agarwal et al., 2008).</p>
      <p>Thus, richness of information of posts can be used
for the identification of opinion leaders. We use
both textual information and multimedia
information to represent the richness of
information as (3).
ric=d + g,</p>
      <p>(3)
where d is the total number of significant words
that the user uses in his or her posts and g is the
total number of multimedia objects that the user
posts.
(1)
distribution and range, we normalize each feature
to a value between 0 and 1. Thus, the weights of
opinion leaders based on the quality of posts
become the average of these three features as (4).
Due to lack of available benchmark data set, we
crawl WOM documents from the Mobile01
bulletin board system (http://www.mobile01.com/),
which is one of the most popular online discussion
forums in Taiwan. This bulletin board system
allows its members to contribute their opinions
free of charge and its contents are available to the
public. A bulletin board system generally has an
organized structure of topics. This organized
structure provides people who are interested in the
same or similar topics with an online discussion
forum that forms a social network. Finding opinion
leaders on bulletin boards is important since they
contain a lot of availably focused WOM. In our
initial experiments, we collected 1537 documents,
which were initiated by 1064 members and
attracted 9192 followers, who posted 19611
opinions on those initial posts. In this data set, the
total number of participants is 9460.
As we use real-world data, which has no ground
truth about opinion leaders, a user centered
evaluation approach should be used to compare the
difference between models (Kritikopoulos et al.,
2006). In our research, there are 9460 members in
this virtual community. We suppose that ten of
them have a high possibility of being opinion
leaders.</p>
      <p>As identification of opinion leaders is treated to
be one of important tasks of social network
analysis (SNA), we compare the proposed model
(i.e. ITM) with three famous SNA approaches,
which are degree centrality (DEG), closeness
centrality (CLO), betweenness centrality (BET).</p>
      <p>
        Involvement (INV) is an important characteristic
of opinion leaders
        <xref ref-type="bibr" rid="ref11 ref23">(Kim and Han, 2009)</xref>
        . The
number of documents that a member initiates plus
the number of derivative documents by other
members is treated as involvement.
      </p>
      <p>Thus, we have one qualitative model, i.e. ITM,
and four quantitative models, i.e. DEG, CLO, BET
and INV. We put top ten rankings from each model
in a pool of potential opinion leaders. Duplicate
members are removed and 25 members are left.</p>
      <p>We request 20 human testers, which have used and
are familiar with Mobile01.</p>
      <p>In our questionnaire, quantitative information is
provided such as the number of documents that the
potential opinion leaders initiate and the number of
derivative documents that are posted by other
members. For the qualitative information, a
maximum of three documents from each member
are provided randomly to the testers. The top 10
rankings are also considered as opinion leaders
based on human judgment.
4.3</p>
      <p>Results
We suppose that ten of 9460 members are
considered as opinion leaders. We collect top 10
ranking members from each models and remove
duplicates. We request 20 human testers to identify
10 opinion leaders from 25 potential opinion
leaders obtained from five models. According to
experiment results in Table 1, the proposed model
outperforms others. This presents the significance
of documents per se. Even INV is a very simple
approach but it performs much better than social
network analysis models, i.e. DEG, CLO and BET.</p>
      <p>One possible reason is the sparse network structure.</p>
      <p>Many sub topics are in the bulletin board system so
these topics form several isolated sub networks.</p>
      <p>Recall</p>
      <p>Precision
DEG
CLO
BET
INV
ITM</p>
      <p>Conclusions and Further Work</p>
      <p>Word of mouth (WOM) has a powerful effect
on consumer behavior. Opinion leaders have
stronger influence on other members in an opinion
society. How to find opinion leaders has been of
interest to both practitioners and researchers.</p>
      <p>Existing models mainly focus on quantitative
features of opinion leaders, such as the number of
posts and the central position in the social network.</p>
      <p>This research considers this issue from the
viewpoints of text mining. We propose an
integrated text mining model by extracting three
important features of opinion leaders regarding
novelty, expertise and richness of information,
from documents. Finally, we compare this
proposed text mining model with four quantitative
approaches, i.e., involvement, degree centrality,
closeness centrality and betweenness centrality,
evaluated by human judgment. In our experiments,
we found that the involvement approach is the best
one among the quantitative approaches. The text
mining approach outperforms its quantitative
counterparts as the richness of document
information provides a similar function to the
qualitative features of opinion leaders. The
proposed text mining approach further measures
opinion leaders based on features of novelty and
expertise.</p>
      <p>In terms of possible future work, some
integrated strategies of both qualitative and
quantitative approaches should take advantages of
both approaches. For example, the 2-step
integrated strategy, which uses the text
miningbased approach in the first step, and uses the
quantitative approach based on involvement in the
second step, may achieve the better performance.</p>
      <p>Larger scale experiments including topics, the
number of documents and testing, should be done
further in order to produce more general results.
Quality Metrics for Optimizing Parameters Tuning in Clustering
Algorithms for Extraction of Points of Interest in Human Mobility
Miguel Nue˜z del Prado Cortez</p>
      <p>Peru I+D+I</p>
      <p>Technopark IDI
miguel.nunez@peruidi.com
Hugo Alatrista-Salas
GRPIAA Labs., PUCP</p>
      <p>Peru I+D+I
halatrista@pucp.pe
Clustering is an unsupervised learning
technique used to group a set of elements into
nonoverlapping clusters based on some predefined
dissimilarity function. In our context, we rely
on clustering algorithms to extract points of
interest in human mobility as an inference
attack for quantifying the impact of the privacy
breach. Thus, we focus on the input
parameters selection for the clustering algorithm,
which is not a trivial task due to the direct
impact of these parameters in the result of the
attack. Namely, if we use too relax
parameters we will have too many point of interest
but if we use a too restrictive set of
parameters, we will find too few groups. Accordingly,
to solve this problem, we propose a method to
select the best parameters to extract the
optimal number of POIs based on quality metrics.
1
The first step in inference attacks over mobility
traces is the extraction of the point of interest (POI)
from a trail of mobility traces. Indeed, this phase
impacts directly the global accuracy of an inference
attack that relies on POI extraction. For instance,
if an adversary wants to discover Alice’s home and
place of work the result of the extraction must be as
accurate as possible, otherwise they can confuse or
just not find important places. In addition, for a more
sophisticated attack such as next place prediction, a
mistake when extracting POIs can decrease
significantly the global precision of the inference. Most
of the extraction techniques use heuristics and
clustering algorithms to extract POIs from location data.</p>
      <p>On one hand, heuristics rely on the dwell time, which
is the lost of signal when user gets into a building.</p>
      <p>Another used heuristic is the residence time, which
represents the time that a user spends at a
particular place. On the other hand, clustering algorithms
group nearby mobility traces into clusters.</p>
      <p>In particular, in the context of POI extraction, it is
important to find a suitable set of parameters, for a
specific cluster algorithm, in order to obtain a good
accuracy as result of the clustering. The main
contribution of this paper is a methodology to find a
“optimal” configuration of input parameters for a
clustering algorithm based on quality indices. This
optimal set of parameters allows us to have the
appropriate number of POIs in order to perform another
inference attack. This paper is organized as follows.</p>
      <p>First, we present some related works on parameters
estimation techniques in Section 2. Afterwards, we
describe the clustering algorithms used to perform
the extraction of points of interests (POIs) as well as
the metrics to measure the quality of formed clusters
in sections 3 and 4, respectively. Then, we introduce
the method to optimize the choice of the parameters
in Section 5. Finally, Section 6 summarizes the
results and presents the future directions of this paper.
2</p>
      <p>Related works
Most of the previous works estimate the parameters
of the clustering algorithms for the point of interest
extraction by using empirical approaches or highly
computationally expensive methods. For instance,
we use for illustration purpose two classical
clustering approaches, K-means (MacQueen et al., 1967)
and DBSCAN (Ester et al., 1996). In the former
clustering algorithm, the main issue is how to
determine k, the number of clusters. Therefore, several
approaches have been proposed to address this issue
(Hamerly and Elkan, 2003; Pham et al., 2005). The
latter algorithm relies on OPTICS (Ankerst et al.,
1999) algorithm, which searches the space of
parameters of DBSCAN in order to find the optimal
number of clusters. The more parameters the clustering
algorithm has, the bigger the combinatorial space of
parameters is. Nevertheless, the methods to calibrate
cluster algorithm inputs do not guarantee a good
accuracy for extracting meaningful POIs. In the next
section, we described the cluster algorithms used in
our study.
3</p>
      <p>Clustering algorithms for extraction of
points of interest
To perform the POI extraction, we rely on the
following clustering algorithms:
3.1</p>
      <p>Density Joinable Cluster (DJ-Cluster)
DJ-Cluster (Zhou et al., 2004) is a clustering
algorithm taking as input a minimal number of points
minpts, a radius r and a trail of mobility traces
M . This algorithm works in two phases. First, the
pre-processing phase discards all the moving points
(i.e. whose speed is above , for a small value)
and then, squashes series of repeated static points
into a single occurrence for each series. Next, the
second phase clusters the remaining points based
on neighborhood density. More precisely, the
number of points in the neighborhood must be equal or
greater than minpts and these points must be within
radius r from the medoid of a set of points. Where
medioid is the real point m that minimizes the sum of
distances from the point m to the other points in the
cluster. Then, the algorithm merges the new cluster
with the clusters already computed, which share at
least one common point. Finally, during the
merging, the algorithm erases old computed clusters and
only keeps the new cluster, which contains all the
other merged clusters.
3.2</p>
      <p>Density Time Cluster (DT-Cluster)
DT-Cluster (Hariharan and Toyama, 2004) is an
iterative clustering algorithm taking as input a distance
threshold d, a time threshold t and a trail of
mobility traces M . First, the algorithm starts by building
a cluster C composed of all the consecutive points
within distance d from each other. Afterwards, the
algorithm checks if the accumulated time of
mobility traces between the youngest and the oldest ones
is greater than the threshold t. If it is the case, the
cluster is created and added to the list of POIs.
Finally as a post-processing step, DT-Cluster merges
the clusters whose medioids are less than d/3 far
from each other.
Introduced in (Gambs et al., 2011), TD-Cluster is
a clustering algorithm inspired from DT Cluster,
which takes as input parameters a radius r, a time
window t, a tolerance rate τ , a distance threshold d
and a trail of mobility traces M . The algorithm starts
by building iteratively clusters from a trail M of
mobility traces that are located within the time window
t. Afterwards, for each cluster, if a fraction of the
points (above the tolerance rate τ ) are within radius
r from the medoid, the cluster is integrated to the list
of clusters outputted, whereas otherwise it is simply
discarded. Finally, as for DT Cluster, the algorithm
merges the clusters whose medoids are less than d
far from each other.
The objective of the begin and end location finder
inference attack (Gambs et al., 2010) is to take as
meaningful points the first and last of a journey.</p>
      <p>More precisely, this heuristic considers that the
beginning and ending locations of a user, for each
working day, might convey some meaningful
information.</p>
      <p>Since we have introduced the different clustering
algorithms to extract points of interest, we present in
the next section the indices to measure the quality of
the clusters.
4</p>
      <p>Cluster quality indices
One aspect of the extraction of POIs inference
attacks is the quality of the obtained clusters, which
impacts on the precision and recall of the attack.</p>
      <p>In the following subsection we describe some
metrics to quantify how accurate or “how good“ is the
outcome of the clustering task. Intuitively, a good
clustering is one that identifies a group of clusters
that are well separated one from each other, compact
and representative. Table 1 summarizes the notation
used in this section.</p>
      <p>Symbol</p>
      <p>C
ci
nc
mi
d(x, y)
|ci|
m0
m00
|C|</p>
      <p>Definition
An ensemble of clusters.</p>
      <p>The ith cluster of C.</p>
      <p>The number of clusters in C.</p>
      <p>The medoid point of the ith cluster.</p>
      <p>The Euclidean distance between x and y.</p>
      <p>The number of points in a cluster ci.</p>
      <p>The closest point to the medoid mi.</p>
      <p>The second closest point to the medoid mi.</p>
      <p>The total number of points in a set of C.
The intra-inter cluster ratio (Hillenmeyer, 2012)
measures the relation between compact (Equation 1)
and well separated groups (Equation 3). More
precisely, we first take the inter-cluster distance, which
is the average distance from each point in a cluster
ci to its medoid mi.</p>
      <p>DIC(ci) =
d(xj , mi)</p>
      <p>(1)
1
|ci|</p>
      <p>X
|ci| − 1 xj∈ci,xj6=mi
Then, the average intra-cluster distance (DIC) is
computed using Equation 2.</p>
      <p>AV G DIC(C) =
Afterwards, the mean distance among all medoids
(DOC) in the cluster C is computed, using Equation
3.</p>
      <p>|C|
X
|C|</p>
      <p>X
1
DOC(C) =
|nC | × (|nC | − 1) ci∈C cj∈C,i6=j
(3)
Finally, the ratio intra-inter cluster rii is given by the
Equation 4 as the relationship between the average
intra cluster distance divided by the inter-cluster
distance.</p>
      <p>AV G DIC(C)
rii(C) = (4)</p>
      <p>DOC(C)
The intra-inter ratio has an approximate linear
complexity in the number of points to be computed and
gives low values to well separated and compact
cluster.
Inspired by the Ben-David and Ackerman
(BenDavid and Ackerman, 2008) k-additive Point
Margin (K-AM) metric , which evaluates how well
centered clusters are. We measure the difference
between the medoid mi and its two closest points m0
and m00 of a given group ci belonging to a cluster C
(Equation 5).</p>
      <p>K − AM (ci) = d(mi, m0i0) − d(mi, m0i)
(5)
Since the average of the k-additive point margins for
all groups ci in a cluster C is computed, we take the
ratio between the average k-additive Point Margin
and the minimal inter-cluster distance (Equation 1)
as shown in Equation 6.</p>
      <p>1 Pnc
AM (C) = minci∈C nc ci∈C K − AM (ci)</p>
      <p>DIC(ci)
(6)
The additive margin method has a linear complexity
in the number of clusters. This metric gives a high
value for a well centered clusters.
The information loss ratio is a metric inspired by the
work of Sole and coauthors (Sole´ et al., 2012). The
basic idea is to measure the percent of information
that is lost while representing original data only by
a certain number of groups (e.g., when we represent
the POIs by the cluster medoids instead of the whole
set of points). To evaluate the percent of information
loss, we compute the sum of distance of each point
represented by xi to its medoid mi for all clusters
ci ∈ C as we shown in Equation 7.</p>
      <p>nc |c|
SSE(C) = X X d(xj , mi)</p>
      <p>ci∈C xj∈ci
Then, we estimate the accumulated distance of all
points of a trail of mobility traces in the cluster C to
a global centroid (GC) using the following equation
Equation 8.</p>
      <p>|C|
SST (C) = X d(xi, GC)</p>
      <p>xi∈C
Finally, the ratio between aforementioned distances
is computed using Equation 9, which results in the
(7)
(8)
information loss ratio.</p>
      <p>IL(C) =</p>
      <p>SSE(C)
SST (C)
(9)</p>
      <p>Afterwards, a similarity measure between two
clusters ci and cj , called R-similarity, is estimated, based
on Equation 14.
diam(ci) = maxx,y∈ci,x6=yd(x, y)
(10)</p>
      <p>DBI (C) =
The computation of this ratio has a linear
complexity. The lowest is the value of this ratio, the more
representative the clusters are.
This quality index (Dunn, 1973; Halkidi et al., 2001)
attempts to recognize compact and well-separated
clusters. The computation of this index relies on
a dissimilarity function (e.g. Euclidean distance d)
between medoids and the diameter of a cluster (c.f,
Equation 10) as a measure of dispersion.</p>
      <p>Then, if the clustering C is compact (i.e, the
diameters tend to be small) and well separated (distance
between cluster medoids are large), the result of the
index, given by the Equation 11, is expected to be
large.</p>
      <p>DIL(C) = minci∈C [mincj∈C,j=i+1
[
maxck∈C diam(ck)
]]</p>
      <p>(11)
The greater is this index, the better the performance
of the clustering algorithm is assumed to be. The
main drawbacks of this index is the computational
complexity and the sensitivity to noise in data.
4.5</p>
      <p>Davis-Bouldin index
The objective of the Davis-Bouldin index (DBI)
(Davies and Bouldin, 1979; Halkidi et al., 2001) is to
evaluate how well the clustering was performed by
using properties inherent to the dataset considered.</p>
      <p>First, we use a scatter function within the cluster ci
of the clustering C (Equation 12).</p>
      <p>v</p>
      <p>n</p>
      <p>S(ci) = utu n1c xXj∈ci d(xj , mi)2
Then, we compute the distance between two
different clusters ci and cj , given by Equation 13.</p>
      <p>q
M (ci, cj ) =
(14)
(16)
R(ci, cj ) =</p>
      <p>S(ci) + S(cj )</p>
      <p>M (ci, cj )
After that, the most similar cluster cj to ci is the
one maximizing the result of the function Rall(ci),
which is given by Equation 15 for i 6= j.</p>
      <p>Rall(ci) = maxcj∈C,i6=j R(ci, cj )</p>
      <p>(15)
Finally, the DBI is equal to the average of the
similarity between clusters in the clustering set C
(Equation 16).
Ideally, the clusters ci ∈ C should have the
minimum possible similarity to each other. Accordingly,
the lower is the DB index, the better is the
clustering formed. These indices would be used to
maximize the number of significant places a cluster
algorithm could find. More precisely, in the next section
we evaluate the cluster algorithm aforementioned as
well as the method to extract the meaningful places
using the quality indices.
5</p>
      <p>Selecting the optimal parameters for
clustering
In order to establish how to select the best set
of parameters for a given clustering algorithm, we
have computed the precision, recall and F-measure
of all users of LifeMap dataset (Chon and Cha,
2011). One of the unique characteristic of this
dataset is that the POIs have been annotated by
the users. Consequently, given a set of clusters
ci ∈ C such that C = {c1, c2, c3, . . . , cn} and a
set of points of interest (POIs) defined by the users
Ppoi = {ppoi 1, ppoi 2, ppoi 3, . . . , ppoi n} we were
able to compute the precision, recall and f-measure
as we detail in the next subsection.</p>
      <p>Precision, recall and F-measure
To compute the recall (c.f. Equation 17), we take as
input a clustering set C, the ground truth represented
by the vector Ppoi (which was defined manually by
each user) as well as a radius to count all the
clusters c ∈ C that are within the radius of ppoi ∈ Ppoi,
which represents the ”good clusters”. Then, the ratio
of the number of good clusters compared to the total
number of found clusters is computed. This measure
illustrates the ratio of extracted cluster that are POIs
divided by the total number of extracted clusters.</p>
      <p>Characteristics</p>
      <p>Total nb of users
Collection period (nb of days)</p>
      <p>Average nb of traces/user</p>
      <p>Total nb of traces
Min #Traces for a user
Max #Traces for a user</p>
      <p>To compute the recall (c.f. Equation 18), we take
as input a clustering set C, a vector of POIs Ppoi as
well as a radius to count the discovered POIs ppoi ∈
Ppoi within a radius of the clusters c ∈ C, which
represents the ”good POIs”. Then, the ratio between
the number of good POIs and the total number of
POIs is evaluated. This metric represents the percent
of the extracted unique POIs.</p>
      <p>(18)
(19)
Recall =</p>
      <p>good P OIs
total number of P OIs</p>
      <p>Finally, the F-measure is defined as the weighted
average of the precision and recall as we can see in
Equation 19.</p>
      <p>precision × recall</p>
      <p>F − measure = 2 × precision + recall
We present the dataset used for our experiments in
the next subsection.
5.2 Dataset description
In order to evaluate our approach, we use the
LifeMap dataset (Chon et al., 2012), which is
composed of mobility traces of 12 user collected for a
year in Seoul, Korea. This dataset comprises
location (latitude and longitude) collected with a
frequency between 2 to 5 minutes with the user defined
point of interest as true if the mobility trace is
considered as important or meaningfull for each user.</p>
      <p>Table 2 summarizes the main characteristics of this
dataset, such as the collect period, the average
number of traces per user, the total number of mobility
traces in the dataset, the minimal and maximal
number of users’ mobility traces.</p>
      <p>Since we have described our dataset, we present
the results of our experiments in the next subsection.
5.3 Experimental results
This section is composed of two parts, in the first
part we compare the performance of the previously
described clustering algorithms, with two
baseline clustering algorithms namely k-means and
DBSCAN. In the second part, a method to select the
most suitable parameters for a clustering algorithm
is presented.</p>
      <p>Input parameters
Tolerance rate (%)
Tolerance rate (%)
Minpts (points)
Eps (Km.)
Merge distance (Km.)
Time shift (hour)
K (num. clusters)</p>
      <p>Possible values
{0.75, 0.8, 0.85, 0.9}
{0.75, 0.8, 0.85, 0.9}
{3, 4, 5, 6, 7, 8, 9, 10, 20, 50}
{0.01, 0.02, 0.05, 0.1, 0.2}
{0.02, 0.04, 0.1, 0.2, 0.4}
{1, 2, 3, 4, 5, 6}
{5, 6, 7,8, 9}</p>
      <p>In order to compare the aforementioned clustering
algorithms, we have take into account the precision,
recall, F-measure obtained, average execution time,
number of input parameters and time complexity.</p>
      <p>To evaluate these algorithms, we used the LifeMap
dataset with POIs annotation and a set of different
parameters configurations for each algorithm, which
are summarized in Table 3. After running these
con0.5
0.0
-0.5
e-1.0
u
l
a
V
0.5
0.0
-0.5
-1.0
figurations, we obtained the results shown in Table
4 for the different input values.</p>
      <p>It is possible to observe that the precision of
DJCluster out performs better than the other
clustering algorithms. In terms of recall DBSCAN and
TD-Cluster perform the best but DJ-Cluster is just
behind them. Moreover, DJ-Cluster has the best
F-measure. Regarding the execution time,
DTClustering the fastest one while DJ-Cluster is the
slowest algorithm due to the preprocessing phase.</p>
      <p>Despite the high computational time of DJ-Cluster,
this algorithm performs well in terms of F-measure.</p>
      <p>In the following, we describe our method to
choose “optimal” parameters for obtaining a good
F-measure. We have used the aforementioned
algorithms with a different set of input parameters
configurations for users with POIs annotations in the
LifeMap dataset (Chon and Cha, 2011). Once
clusters are built, we evaluate the clusters issued from
different configurations of distinct algorithms
using the previously described quality indices.
Afterwards, we were able to estimate the precision, recall
and F-measure using the manual annotation of POIs
by the users in the LifeMap dataset.</p>
      <p>Regarding the relation between the quality indices
and the F-measure, we studied the relationship
between these factors, in order to identify the indices
that are highly correlated with the F-measure, as
can be observed in Figure 1. We observe that the
two best performing indices, except for k-means,
are IL and DBI. The former shows a negative
correlation with respect to the F-measure. While the
latter, has a positive dependency to the F-measure.</p>
      <p>Our main objective is to be able to identify the
relationship between quality and F-measure among the
previous evaluated clustering algorithms.
Accordingly, we discard the inter-intra cluster ratio (RII)
and the adaptive margin (AM), which only perform
well when using k-means and the DJ clustering
algorithms. Finally, we observe that the Dunn index
has a poor performance. Based on these
observations, we were able to propose an algorithm to
automatically choose the best configuration of input
parameters.
5.4</p>
      <p>Parameter selection method
Let us define a vector of parameterspi ∈ P and P a
set of vectors, such that P = {p1, p2, p3, . . . , pn},
a trail of mobility traces M of a user. From
previous sections we have the clustering function
C(pi) and the quality metrics Information Loss
IL(C) and Davis-Bouldin index DBI (C). Thus,
for each vector of parameters we have a tuple
composed of the trail of mobility traces, the result
of the clustering algorithm and the quality metrics
(pi, M, Cpi , ILCpi , DBICpi ). When we compute
the clustering algorithm and the quality metrics for
each vector of parameter for a given user u. We
define also aχ0u matrix, which the matrix χu sorted by
IL ascending. Finally, the result matrix χu is of the
form:
χu =
p1
p2
p3
. . .
pn
DBICpn
k means</p>
      <p>Resume</p>
      <p>TD cluster</p>
      <p>Heuristic</p>
      <p>DBI
IL
IL_DBI
MAX
u1 u2 u3 u4 u5 u6 u7 u1 u2 u3 u4 u5 u6 u7 u1 u2 u3 u4 u5 u6 u7</p>
      <p>User</p>
      <p>Therefore, the parameter selection function
S(χu) could be defined as:</p>
      <p>(pi, if maxpi (DBI ) &amp; minpi (IL)
S(χu) =
p0i, if maxp0i (DBI ) in 1st quartile</p>
      <p>(20)</p>
      <p>In detail, the function S takes as input a χ matrix
containing the parameters vector pi, a trail of
mobility traces M , the computed clustering C(pi, M ) as
well as the quality metrics, such as Information loss
(IL(C)) and the Davis-Bouldin index (DBI (C)).</p>
      <p>Once all these values have been computed for each
evaluated set of parameters, two cases are possible.</p>
      <p>In the first case, both IL and DBI agree on the same
set of input parameters. In the second situation, both
IL and DBI refer each one to a different set of
parameters. In this case, the algorithm sorts the values
by IL in the ascending order (i.e., from the
smallest to the largest information loss value). Then, it
chooses the set of parameters with the greatest DBI
in the first quartile.</p>
      <p>For the sake of evaluation, our methodology was
tested using the LifeMap dataset to check if the
chosen parameters are optimal. We have tested the
method with the seven users of LifeMap that have
annotated manually their POIs. Consequently, for
every set of settings of each clustering algorithm, we
have computed the F-measure because we have the
more efficient, elitist and doesnt need to define
additional parameters.</p>
      <p>In the NSGA-II, the population Qt (size N ) is
generated using the parent population Pt (size N ).</p>
      <p>
        After this, the two populations are combined for
generating the population Rt (size 2N ). The
population Rt is sorted according the dominance of the
solutions in different Pareto fronts
        <xref ref-type="bibr" rid="ref21">(Pareto, 1896)</xref>
        and the crowding distance. A new population Pt+1
(size N ) is generated with the bests Pareto fronts F1,
F2, F3 and so forth, until the Pt+1 size equals to the
value of N . The solutions in the Pareto fronts under
this limit are removed. After Pt+1 is a new Pt and
the process is repeated until a conditions is satisfied.
      </p>
      <p>
        The figure 1 shows the process of evolutions of the
solutions in the NSGA-II. More details on NSGA-II
can be found at
        <xref ref-type="bibr" rid="ref12">(Deb et al., 2002)</xref>
        .
      </p>
      <p>Proposed Method
This section presents the proposed methods for
learning fuzzy rules using the IRL approach and a
MOGA, and tuning the weights of the fuzzy rules
using a R&amp;P method. In the next subsections each
method is detailed.
4.1</p>
      <p>
        Learning Fuzzy Rules
The proposed method for learning fuzzy rules is
based in the iterative multiobjective genetic method
described in
        <xref ref-type="bibr" rid="ref22">(Hinojosa and Camargo, 2012)</xref>
        and uses
a MOGA for learning a single fuzzy rule in each
iteration of the MOGA. The main difference with the
proposed method is the module for defining the
order of the class for learning. This method proposed
is illustrated in the Figure 2.
      </p>
      <p>A set of examples is used as the set of training.</p>
      <p>
        The proposed IRL method starts when is defined the
order of the class for learning. After that, a class is
selected and the module for generate the best rule
that used a MOGA is executed. The MOGA
considers two objectives for minimization: accuracy and
interpretability. The accuracy is determined by the
integrity and consistency of each rule
        <xref ref-type="bibr" rid="ref9">(Gonzalez and
Perez, 1999)</xref>
        and the interpretability is defined by the
quantity of conditions of each rule. When the best
rule in the Pareto front improves the rate of
classification of the RB, this rule is inserted into the RB,
some examples are marked and the process of
learning a fuzzy rule starts again. When the best rule in
the Pareto front doesnt improve the rate of
classification of the RB, the process verifies that all the
sequence of class was learned, if the sequence is not
learned a new class is selected and the process of
learning a fuzzy rule starts again, else the process
finishes and the set of the best rules is the RB.
      </p>
      <p>In the process detailed above all rules has a weight
equals to one. These weights can be tuning for
improve the rate classification. This tuning is detailed
in the next subsection.</p>
      <p>
        Tunning Weights of the Fuzzy Rules
We use the method proposed in
        <xref ref-type="bibr" rid="ref13">(Nozaki et al., 1996)</xref>
        for this task. This method rewards or increases the
weight the a fuzzy rule Ri when a example eq is
correctly classified by this rule according to the next
equation:
      </p>
      <p>CFinew = CFiold + n1 1 − CFiold</p>
      <p>(1)</p>
      <p>And this method punishes or decreases the weight
of the fuzzy rule Ri when a example eq is
misclassifed by this rule according to the next equation:</p>
      <p>CFinew = CFiold − n2CFiold</p>
      <p>(2)</p>
      <p>In the experimental study detaild in Section 5 we
used the values n1=0.001 and n2=0.1 and the tuning
procedure for 500 iterations.
5</p>
      <p>Experimental Study
The experimental study is aimed to show the
application of the proposed method and the comparation
with the classification with non-parametric method
for classifying cells of the Eimeria of Domestic Fowl
based on Morphological Data. The Emeira genus
comprises a group of protozoan parasites that infect
a wide range of hosts, seven different Emeira species
infect the domestic fowl, causing enteritis with
several economic losses. This protozoan morphology
was represented by 13 features: mean of curvature,
standard deviation of curvature, entropy of
curvature, major axis (lenght), minor axis (width),
symmetry through major axis, symmetry through minor
axis, area, entropy of internal structure, second
angular moment, contrast, inverse difference moment,
entropy of co-occurrence matrix; these features are
used as the input pattern for the classification
process.</p>
      <p>
        The Table 1 shows the class and the number of
examples or instances of each class. More detail
how the features were extracted or about the Eimeria
genus can be found at
        <xref ref-type="bibr" rid="ref14">(Beltran, 2007)</xref>
        .
      </p>
      <p>The Table 2 shows the parameters for learning
fuzzy rules and the NSGA-II (the MOGA used in
this paper).</p>
      <p>After the process of learning fuzzy rules, the
process of tuning the weights starts. The Table 3 shows
the result of the classification or dispersion matriz
after the process tuning the weights.</p>
      <p>After the proposed method, the result is a set of
rules similar of the set shows in the Figure 3. This
Table 1: Distribution of Classes
Parameter Value
Size the population 50.0
Crossover rate 1.0
Mutation rate 0.2
Number of generations 500.0</p>
      <p>Mark value 0.3
rules has a high level of interpretability for the expert
users.</p>
      <p>
        We compared the proposed method with the
method non-parametric classifier proposed in
        <xref ref-type="bibr" rid="ref14">(Beltran, 2007)</xref>
        with the same set of examples. The
Table 4 shows the classification rate by each class of
both classifiers. These results shows that the
proposed method (PM) has a similar rate
classification (overall 77.17) that the non-parametric method
(NPM) (overall 80.24), but with a high degree of
interpretability. The non-parametric method does not
consider the interpretability.
6
      </p>
      <p>Conclusions
In this article, we proposed a iterative multiobjective
genetic method to learn fuzzy classification rules.</p>
      <p>The fuzzy rules are learned in each iteration depend
of the sequencia of class. After that, the weights of
the each fuzzy rules are tuned using a R&amp;P method.</p>
      <p>The results obtained have indicated that FRBCSs
ACE
MAX
BRU
MIT
PRA
TEN
NEC
have better interpretability and similar accuracy than
a non-parametric method for classify the Eimeria of
domestic fowl.</p>
      <p>SIFR Project: The Semantic Indexing of French Biomedical Data</p>
      <p>Resources
Juan Antonio Lossio-Ventura,</p>
      <p>Clement Jonquet
LIRMM, CNRS, Univ. Montpellier 2</p>
      <p>Montpellier, France
fName.lName@lirmm.fr</p>
      <p>Mathieu Roche,</p>
      <p>Maguelonne Teisseire
TETIS, Cirad, Irstea, AgroParisTech</p>
      <p>
        Montpellier, France
fName.lName@teledection.fr
The Semantic Indexing of French
Biomedical Data Resources project proposes to
investigate the scientific and technical
challenges in building ontology-based services
to leverage biomedical ontologies and
terminologies in indexing, mining and
retrieval of French biomedical data.
1 Introducci o´n
Hoy en d´ıa la gran cantidad de datos disponibles
en l´ınea suele componerse de texto no
estructurado, por ejemplo reportes cl´ınicos,
informes de reportes adversos, historiales cl´ınicos
electro´nicos
        <xref ref-type="bibr" rid="ref30">(Lossio-Ventura et al., 2013)</xref>
        .
Regularmente estos textos son escritos usando un
lenguaje espec´ıfico (expresiones y t e´rminos)
usados por una comunidad. Es por eso existe la
necesidad de formalizar e indexar te´rminos o
conceptos te´cnicos. Lo cual implica un gran consumo
de tiempo.
      </p>
      <p>Los te´rminos relevantes son u´tiles para obtener
una mayor comprensio´n de la estructura
conceptual de un dominio. Estos pueden ser:
(i) te´rminos de una sola palabra (sencillo a
extraer), o (ii) te´rminos de varias palabras (dif´ıcil).</p>
      <p>
        En el a´mbito biome´dico, hay una gran
diferencia entre los recursos existentes (ontolog´ıas) en
ingle´s y france´s. En Ingle´s hay cerca de 7 000 000
de te´rminos asociados a 6 000 000 de conceptos,
tales como los de UMLS1 o BioPortal
        <xref ref-type="bibr" rid="ref33">(Noy et al.,
2009)</xref>
        . Mientras que, en france´s so´lo hay
alrededor de 330 000 te´rminos asociados a 160 000
conceptos
        <xref ref-type="bibr" rid="ref31">(Neveol et al., 2014)</xref>
        . Por lo tanto, hay
una necesidad de enriquecer terminolog´ıas u
ontolog´ıas en franc e´s. Por lo tanto, nuestro
trabajo se compone de dos pasos principales: (i) la
extraccio´n de te´rminos biome´dicos, y (ii) el
en1http://www.nlm.nih.gov/research/umls
riquecimiento de ontolog´ıas, con el fin de poblar
ontolog´ıas con los t e´rminos extra´ıdos.
      </p>
      <p>El art´ıculo es organizado como sigue. Primero
discutimos sobre la sobre la metodolog´ıa puesta
en marcha para este proyecto en la Seccio´n 2.</p>
      <p>La evaluacio´n de la precisio´n es presentada en
la Seccio´n 3 seguida de las conclusiones en la
Seccio´n 4.
2</p>
      <p>Metodolog´ıa
Nuestro trabajo se divide en dos procesos
principales: (i) la extraccio´n de te´rminos biome´dicos, y
(ii) el enriquecimiento de ontolog´ıas, explicados a
continuacio´n.</p>
      <p>Extraccio´n Automa´tica de Te´rminos</p>
      <p>Biome´dicos
La extraccio´n de te´rminos es una tarea esencial
en la adquisicio´n de conocimiento de un dominio.</p>
      <p>
        En este trabajo presentamos las medidas creadas
para este objetivo. Medidas que se basan en
varios criterios como lingu¨´ıstico, estad ´ıstico, grafos
y web para mejorar el resultado de extraccio´n
de te´rminos biome´dicos. Las medidas
presentadas a continuacio´n son puestas a disposicio´n de
la comunidad, bajo la aplicacio´n llamada
BIOTEX
        <xref ref-type="bibr" rid="ref26 ref27 ref28 ref29">(Lossio-Ventura et al., 2014)</xref>
        .
2.1.1 Ling u¨´ıstica
Estas te´cnicas intentan recuperar te´rminos gracias
a la formacio´n de patrones. La idea principal es
la construccio´n de reglas para describir las
estructuras de los te´rminos de un dominio
mediante el uso de caracter´ısticas ortogr a´ficas, le´xicas
o morfo-sinta´cticas. La idea principal es la
construccio´n de reglas, normalmente de forma
manual, que describen las estructuras comunes de
te´rminos para ciertos campos. En muchos
casos tambie´n, diccionarios conteniendo te´rminos
te´cnicos (e.g., prefijos, sufijos y acor´nimos
espec´ıficos) son usados para ayudar a extraer
te´rminos
        <xref ref-type="bibr" rid="ref25">(Krauthammer et al., 2004)</xref>
        .
2.1.2 Estad´ıstica
Las te´cnicas estad´ısticas se basan en la
evidencia presentada en el corpus a trave´s de la
informacio´n contextual. Tales enfoques abordan
principalmente el reconocimiento de te´rminos
generales
        <xref ref-type="bibr" rid="ref34">(Van Eck et al., 2010)</xref>
        . La mayor´ıa de
medidas se basan en la frecuencia. La mayor parte
de trabajos combinan la informacio´n lingu¨´ıstica
y estad´ıstica, tal es el caso de C-value (Frantzi
et al., 2000) combina la informacio´n estad´ıstica y
lingu¨´ıstica tanto para la extracci o´n de te´rminos de
varias palabras como de te´rminos largos y
anidados. Es la medida ma´s conocida en la
literatura. En el trabajo de (Zhang et al., 2008),
demostraron que C-value obtiene los mejores
resultados comparado a otras medidas. Adema´s del
ingle´s, C-value tambie´n ha sido aplicado a otros
idiomas tales como japone´s, serbio, esloveno,
polaco, chino
        <xref ref-type="bibr" rid="ref24">(Ji et al., 2007)</xref>
        , espan˜ol
(Barro´nCedeno et al., 2009), a´rabe. Es por eso, en
nuestro primer trabajo
        <xref ref-type="bibr" rid="ref30">(Lossio-Ventura et al., 2013)</xref>
        ,
la modificamos y adaptamos para el france´s.
      </p>
      <p>
        A partir de C-value, hemos creados otras
medidas, como F-TFIDF-C, F-OCapi, C-OKapi,
C-TFIDF
        <xref ref-type="bibr" rid="ref26 ref27 ref28 ref29">(Lossio-Ventura et al., 2014)</xref>
        , estas
medidas obtienen mejores resultados que
Cvalue. Finalmente una nueva medida basada
en la informacio´n lingu¨´ıstica y estad ´ıstica es
LIDF-value
        <xref ref-type="bibr" rid="ref26 ref27 ref28 ref29">(Lossio-Ventura et al., 2014)</xref>
        (patrones Lingu¨´ısticos, IDF, and C-value
information), que mejora con gran diferencia los
resultados obtenidos por las medidas antes citadas.
2.1.3 Grafos
El modelo de grafos es una alternativa al modelo
de informacio´n, muestra claramente las relaciones
entre los nodos gracias a las aristas. Gracias a los
algoritmos de centralidad se puede aprovechar los
grupos de informacio´n en grafos. Existen
aplicaciones de grafos para la Recuperacio´n de
Informacio´n (RI) en el contexto de las redes sociales, de
colaboracio´n y sistemas de recomendacio´n
        <xref ref-type="bibr" rid="ref32">(Noh
et al., 2009)</xref>
        .
      </p>
      <p>
        Una medida basada en grafos creada para este
proceso es TeRGraph
        <xref ref-type="bibr" rid="ref26 ref27 ref28 ref29">(Lossio-Ventura et al., 2014)</xref>
        (Terminology Ranking based on Graph
information). Esta medida tiene como objetivo mejorar
la precisio´n de los primeros k te´rminos extra´ıdos
despue´s de haber aplicado LIDF-value. El grafo
es construido con la lista de te´rminos obtenidos
con LIDF-value, donde los nodos representan los
te´rminos relacionados con otros te´rminos gracias
a la co-ocurrencia en el corpus.
2.1.4 Web
Diferentes estudios de Web Mining se enfocan en
la similitud sema´ntica, relacio´n sema´ntica. Esto
significa para cuantificar el grado en el que
algunas palabras esta´n relacionadas, teniendo en
cuenta no so´lo similitud sino tambie´n cualquier
posible relacio´n sema´ntica entre ellos. La
primera medida web creada fue WebR
(LossioVentura et al., 2014), finalmente la mejora
llamada WAHI
        <xref ref-type="bibr" rid="ref26 ref27 ref28 ref29">(Lossio-Ventura et al., 2014)</xref>
        (Web
Association based on Hits Information). Nuestra
medida basada en la Web tiene por objetivo volver
a clasificar la lista obtenida previamente
conTeRGraph. Demostramos con esta medida que la
precisio´n de los k primeros te´rminos extra´ıdos
superan los resultados de las medidas arriba
mencionadas (ver Seccio´n 3).
      </p>
      <p>Enriquecimiento de Ontolog´ıas
El objetivo de este proceso es enriquecer las
terminolog´ıas u ontolog ´ıas con los t e´rminos nuevos
extra´ıdos en el proceso anterior. Los tres grandes
pasos a seguir en este proceso son:
(1) Determinar si un te´rmino es polise´mico:
con la ayuda del Meta-Learning, hemos
podido predecir con una confianza de 97% si un
te´rmino es polise´mico. Esta contribucio´n sera´
valorizada en la conferencia ECIR 2015.
(2) Identificar los posibles significados si el
te´rmino es polise´mico: es nuestro trabajo
actual, con la ayuda de clustering,
clustering sobre los grafos tratamos de resolver este
problema.
(3) Posicionar el te´rmino en una ontolog´ıa.
3</p>
      <p>Experimentaciones</p>
      <p>Datos, protocolo y validacio´n
En nuestros experimentos, hemos usado el corpus
esta´ndar GENIA2, el cual es compuesto de 2 000
t´ıtulos y res u´menes de art´ıculos de revistas que han
sido tomadas de la base de datos Medline,
contiene ma´s de 400 000 palabras. GENIA corpus
contiene expresiones lingu¨´ısticas que se refieren a
entidades con intere´s en biolog´ıa molecular tales
como prote´ınas, genes y c e´lulas.</p>
      <p>2http://www.nactem.ac.uk/genia/
genia-corpus/term-corpus
Los resultados son evaluados en te´rminos de
precisio´n obtenidos sobre los primeros k te´rminos
extra´ıdos ( P @k) para las medidas propuestas y
las medidas base (referencia) para la extraccio´n
de te´rminos compuestos de varias palabras. En
las subsecciones siguientes, limitamos los
resultados para la medida basada en grafos con so´lo los
primeros 8 000 te´rminos extra´ıdos y los
resultados para la medida basada en la web con so´lo los
primeros 1 000 te´rminos.</p>
      <p>Resultados ling u¨´ısticos y estad ´ısticos</p>
      <p>Resultados basados en grafos
Este art´ıculo presenta la metodolog ´ıa propuesta
para el proyecto SIFR. Este proyecto consta de dos</p>
      <p>Figure 3: Comparacio´n de la precisio´n de WAHI y
TeRGraph
grandes procesos.</p>
      <p>El primer proceso Extraccio´n Automa´tica de
Te´rminos Biome´dicos, terminado y siendo
valorizado en varias publicaciones citadas
anteriormente. En este proceso demostramos que las
medidas propuestas mejoran la precisio´n de la
extraccio´n automa´tica de te´rminos en comparacio´n
a las medidas ma´s populares de extraccio´n de
te´rminos.</p>
      <p>El segundo proceso Enriquecimiento de
Ontolog´ıas , a la vez dividio en 3 etapas, es
nuestra tarea actual, solo la primera etapa ha sido
finalizada. En este proceso buscamos encontrar la
mejor posicio´n de un te´rmino en una ontolog´ıa.</p>
      <p>Como trabajo futuro, pensamos acabar el
segundo proceso. Adema´s, planeamos probar
estos enfoques generales sobre otros dominios,
tales como ecolog´ıa y agronom ´ıa. Finalmente,
planeamos aplicar estos enfoques con corpus en
espan˜ol.</p>
      <p>Agradecimientos
Este proyecto es apoyado en parte por la Agencia
Nacional de Investigacio´n de Francia bajo el
programa JCJC, ANR-12-JS02-01001, as´ı como por
la Universidad de Montpellier 2, el CNRS y el
programa de becas FINCyT, Peru´.</p>
      <p>References
Barro´n-Cedeno, A., Sierra, G., Drouin, P., Ananiadou,</p>
      <p>S. 2009. An improved automatic term recognition
method for Spanish. Computational Linguistics,
Intelligent Text Processing, pp. 125-136. Springer.</p>
      <p>Frantzi K., Ananiadou S., Mima, H. 2000. Automatic
recognition of multiword terms: the
C-value/NCvalue Method. International Journal on Digital
Libraries, (3):115-130.</p>
      <p>Mathematical modeling of the performance of a computer system.</p>
      <p>Félix Armando Fermín Pérez
Universidad Nacional Mayor de San Marcos
Facultad de Ingeniería de Sistemas e Informática</p>
      <p>Lima, Perú
fferminp@unmsm.edu.pe
Generally the management of computer systems is
manual; autonomic computing by self-management
tries to minimize human intervention using
autonomic controllers; for this, first the
mathematical modeling of the system under study
is performed, then is designed a controller that
governs the behavior of the system. In this case,
the determination of the mathematical model of a
web server is based on a black box model by
system identification, using data collected and
stored in its own log during operation of the
computer system under study.</p>
      <p>Keywords: autonomic computing,
management, system identification.</p>
      <p>self1
La tecnología influye en cada aspecto de la vida
cotidiana; las aplicaciones informáticas, por
ejemplo, son cada vez más complejas,
heterogéneas y dinámicas, pero también la
infraestructura de información, como la Internet
que incorpora grandes cantidades de recursos
informáticos y de comunicación, almacenamiento
de datos y redes de sensores, con el riesgo de que
se tornen frágiles, inmanejables e inseguras. Así, se
hace necesario contar con administradores de
servicios y de servidores informáticos,
experimentados y dedicados, además de
herramientas software de monitoreo y supervisión,
para asegurar los niveles de calidad de servicio
pactados (Fermín, 2012).</p>
      <p>Diao et al. (2005) promueven la utilización de la
computación autonómica mediante sistemas de
control automático en lazo cerrado, reduciendo la
intervención humana, que Fox y Patterson (2003)
también han identificado como parte principal del
problema, debido al uso de procedimientos ad hoc
donde la gestión de recursos aún depende
fuertemente del control y administración manual.</p>
      <p>Sobre la computación autonómica, Kephart y
Chess (2003) mencionan que la idea es que un
sistema informático funcione igual que el sistema
nervioso autonómico humano cuando regula
nuestra temperatura, respiración, ritmo cardíaco y
otros sin que uno se halle siempre consciente de
ello, esto es, promueve la menor intervención
humana en la administración de la performance de
los sistemas informáticos tendiendo hacia la
autoadministración de los mismos. Tal
autoadministración se caracteriza por las propiedades
de auto-configuración (self-configuration),
autocuración (self-healing), auto-optimización
(selfoptimization) y auto-protección (self-protection).</p>
      <p>En la visión de la computación autonómica los
administradores humanos simplemente especifican
los objetivos de alto nivel del negocio, los que
sirven como guía para los procesos autonómicos
subyacentes. Así los administradores humanos se
concentran más fácilmente en definir las políticas
del negocio, a alto nivel, y se liberan de tratar
permanentemente con los detalles técnicos de bajo
nivel, necesarios para alcanzar los objetivos, ya
que estas tareas son ahora realizadas por el sistema
autonómico mediante un controlador autonómico
en lazo cerrado, que monitorea el sistema
permanentemente, utiliza los datos recolectados del
propio sistema en funcionamiento, compara estas
métricas con las propuestas por el administrador
humano, y controla la performance del sistema, por
ejemplo, manteniendo el tiempo de respuesta del
sistema dentro de niveles prefijados.</p>
      <p>De acuerdo con la teoría de control, el diseño de un
controlador depende de un buen modelo
matemático. En el caso de un sistema informático,
primero debe tenerse un modelo matemático para
luego diseñar un controlador en lazo cerrado o
realimentado, pero sucede que los sistemas
informáticos son bastante complicados de modelar
ya que su comportamiento es altamente
estocástico. Según Hellerstein et al (2004) se ha
utilizado la teoría de colas para modelarlos,
tratándolos como redes de colas y de servidores,
bastante bien pero principalmente en el modelado
del comportamiento estacionario y no cuando se
trata de modelar el comportamiento muchas veces
altamente dinámico de la respuesta temporal de un
sistema informático en la zona transitoria, donde la
tarea se complica.</p>
      <p>De manera que en el presente artículo se trata el
modelado matemático de un sistema informático
mediante la identificación de sistemas, enfoque
empírico donde según Lung (1987) deben
identificarse los parámetros de entrada y salida del
sistema en estudio, basándose en los datos
recolectados del mismo sistema en
funcionamiento, para luego construir un modelo
paramétrico, como el ARX por ejemplo, con las
técnicas estadísticas de autoregresión. La sección 2
describe alguna teoría básica sobre las métricas
para el monitoreo de la performance de un sistema
informático. En la sección 3 se trata el modelado
matemático, y en la sección 4 se describe el
experimento realizado; finalmente en la sección 5
se describen las conclusiones y trabajos futuros.
2</p>
      <p>Monitoreo de la performance.</p>
      <p>En la computación autonómica los datos obtenidos
del monitoreo de la performance del sistema en
estudio contribuye fundamentalmente en la
representación del estado o del comportamiento del
sistema, esto es, en el modelo matemático del
mismo. Según Lalanda et al. (2013), conocer el
estado del sistema desde las perspectivas
funcionales y no funcionales es vital para llevar a
cabo las operaciones necesarias que permitan
lograr los objetivos en el nivel deseado y el
monitoreo de la performance permite saber cuan
bien lo está logrando. Generalmente los datos de la
performance se consiguen vía el log del sistema en
estudio, con herramientas de análisis utilizando
técnicas de análisis estadísticas, principalmente.</p>
      <p>Entre las métricas de performance inicialmente se
encontraba la velocidad de procesamiento, pero al
agregarse más componentes a la infraestructura
informática, surgieron nuevas métricas, siendo las
principales las que proporcionan una idea del
rendimiento o trabajo realizado en un periodo de
tiempo, la utilización de un componente, o el
tiempo en realizar una tarea en particular como por
ejemplo el tiempo de respuesta. Lalanda et al.
(2013) mencionan que las métricas de performance
más populares son las siguientes:
- Número de operaciones en punto flotante por
segundo (FLOPS), representa una idea del
rendimiento del procesamiento, realiza
comparaciones entre máquinas que procesan
complejos algoritmos matemáticos con punto
flotante en aplicaciones científicas.
- Tiempo de respuesta, representa la duración en
tiempo que le toma a un sistema llevar a cabo
una unidad de procesamiento funcional. Se le
considera como una medición de la duración
en tiempo de la reacción a una entrada
determinada y es utilizada principalmente en
sistemas interactivos. La sensibilidad es
también una métrica utilizada especialmente en
la medición de sistemas en tiempo real,
consiste en el tiempo transcurrido entre el
inicio y fin de la ejecución de una tarea o hilo.
- Latencia, medida del retardo experimentado en
un sistema, generalmente se le utiliza en la
descripción de los elementos de comunicación
de datos, para tener una idea de la performance
de la red. Toma en cuenta no solo el tiempo de
procesamiento de la CPU sino también los
retardos de las colas durante el transporte de
un paquete de datos, por ejemplo.
- Utilización y carga, métricas entrelazadas y
utilizadas para comprender la función de
administración de recursos y proporcionan una
medida de cuan bien se están utilizando los
componentes de un sistema y se describe como
un porcentaje de utilidad. La carga mide el
trabajo realizado por el sistema y usualmente
es representado como una carga promedio en
un periodo de tiempo.</p>
      <p>Existen muchas otras métricas de performance:
- número de transacciones por unidad de costo.
- función de confiabilidad, tiempo en el que un
sistema ha estado funcionando sin fallar.
- función de disponibilidad, indica que el sistema
está listo para ser utilizado cuando sea necesario.
- tamaño o peso del sistema, indica la portabilidad.
- performance por vatio, representa la tasa de
cómputo por vatio consumido.
- calor generado por los componentes ya que en
sistemas grandes es costoso un sistema de
refrigeración.</p>
      <p>Todas ellas entre otras más permiten conocer
mejor el estado no funcional de un sistema o
proporcionar un medio para detectar un evento que
ha ocurrido y ha modificado el comportamiento no
funcional. Así, en el presente caso se ha elegido
inicialmente como métrica al tiempo de respuesta,
ya que el sistema informático en estudio es un
servidor web, por esencia de comportamiento
interactivo, de manera que lo que se mide es la
duración de la reacción a una entrada determinada.
3</p>
      <p>Modelado matemático.</p>
      <p>En la computación autonómica, basado en la teoría
de control realimentado, para diseñar un
controlador autonómico primero debe hallarse el
modelo matemático del sistema en estudio, tal
como se muestra en la Figura N° 1. El modelo
matemático de un servidor informático se puede
hallar mediante dos enfoques: uno basado en las
leyes básicas, y otro en un enfoque empírico
denominado identificación de sistemas. Parekh et
al. (2002) indican que en trabajos previos se ha
tratado de utilizar principios, axiomas, postulados,
leyes o teorías básicas para determinar el modelo
matemático de sistemas informáticos, pero sin
éxito ya que es difícil construir un modelo debido a
su naturaleza compleja y estocástica; además es
necesario tener un conocimiento detallado del
sistema en estudio, más aún, cuando cada cierto
tiempo se va actualizando las versiones del
software, y finalmente que en este enfoque no se
considera la validación del modelo.</p>
      <p>Figura N° 1. Modelado matemático basado en la
teoría de control. (Adaptado de Hellerstein, 2004).</p>
      <p>En contraste, según Ljung (1987) la identificación
de sistemas es un enfoque empírico donde debe
identificarse los parámetros de entrada y salida del
sistema en estudio, para luego construir un modelo
paramétrico, como el ARX por ejemplo, mediante
técnicas estadísticas de autoregresión. Este modelo
o ecuación paramétrica relaciona los parámetros de
entrada y de salida del sistema de acuerdo a la
siguiente ecuación:
y(k+1) = Ay(k) + Bu(k)</p>
      <p>(1)
donde y(k)
u(k)
A, B
k
: variable de salida
: variable de entrada
: parámetros de autoregresión
: muestra k-ésima.</p>
      <p>Este enfoque empírico trata al sistema en estudio
como una caja negra, de manera que no afecta la
complejidad del sistema o la falta de conocimiento
experto, incluso cuando se actualicen las versiones
del software bastaría con estimar nuevamente los
parámetros del modelo. Así, para un servidor web
Apache, la ecuación paramétrica relaciona el
parámetro entrada, Max Clients (MC), un
parámetro de configuración del servidor web
Apache que determina el número máximo de
conexiones simultáneas de clientes que pueden ser
servidos; y el parámetro de salida Tiempo de
respuesta (TR), que indica lo rápido que se
responde a las solicitudes de servicio de los
clientes del servidor, ver Figura N° 2.</p>
      <p>Figura N° 2. Entrada y salida del sistema a
modelar. (Elaboración propia).</p>
      <p>En (Hellerstein et al, 2004) se propone realizar la
identificación de sistemas informáticos, como los
servidores web, de la siguiente manera:
1.- Especificar el alcance de lo que se va a modelar
en base a las entradas y salidas consideradas.
2.- Diseñar experimentos y recopilar datos que
sean suficientes para estimar los parámetros de la
ecuación diferencia lineal del orden deseado.
3.- Estimar los parámetros del modelo utilizando
las técnicas de mínimos cuadrados.
4.- Evaluar la calidad de ajuste del modelo y si la
calidad del modelo debe mejorarse, entonces debe
revisarse uno o más de los pasos anteriores.
4</p>
      <p>Experimento.</p>
      <p>La arquitectura del experimento implementado, se
muestra en la Figura Nº 3.</p>
      <p>Figura N° 3. Arquitectura para la identificación del
modelo de un servidor web (Elaboración propia).</p>
      <p>La computadora personal PC1 es el servidor web
Apache, asimismo contiene el sensor software que
recoge los tiempos de inicio y fin de cada http que
ingresa al servidor web, datos que se almacenan en
el log del mismo servidor, y luego se realiza el
cálculo del tiempo de respuesta de cada http
completado en una unidad de tiempo y el tiempo
de respuesta promedio del servidor. El servidor
web por sí mismo no hace nada, por ello, la
computadora personal PC2 contiene un generador
de carga de trabajo, que simula la actividad de los
usuarios que desean acceder al servidor web; en
este caso se utilizó el JMeter una aplicación
generadora de carga de trabajo y que forma parte
del proyecto Apache.</p>
      <p>La operación del servidor web, no solo depende de
la actividad de los usuarios, sino también de la
señal de entrada MaxClients (MC), que toma
forma de una sinusoide discreta variable que excita
al servidor web junto con las solicitudes http del
generador de carga de trabajo (PC2). De manera
que con la actividad del servidor web almacenada
en su log, un sensor software calcula los valores de
la señal de salida Tiempo de Respuesta (TR),
mostrados en la Figura N° 4.</p>
      <p>Figura N° 4. Señal de entrada MaxClients MC y
señal de salida Tiempo de respuesta TR.</p>
      <p>Con los datos de MC y TR obtenidos se estiman
los parámetros de regresión A y B de la ecuación
paramétrica ARX, haciendo uso de las técnicas de
mínimos cuadrados, implementadas en el ToolBox
Identificación de Sistemas del Matlab, logrando la
siguiente ecuación paramétrica de primer orden:</p>
      <p>TR(k+1) = 0.06545TR(k) + 4.984MC(k+1)
Se puede observar que el Tiempo de respuesta
actual depende del Tiempo de respuesta anterior y
del parámetro de entrada MaxClients. El modelo
hallado es evaluado utilizando la métrica r2,
calidad de ajuste, del ToolBox utilizado, que indica
el porcentaje de variación respecto a la señal
original. En el caso de estudio, el modelo hallado
tiene una calidad de ajuste del 87%, lo que se
puede considerar como un modelo aceptable. En la
Figura N° 5 puede observarse la gran similitud
entre la señal de salida medida y la señal de salida
estimada.</p>
      <p>Figura N° 5. TR medida y TR estimada.
5</p>
      <p>Conclusiones.</p>
      <p>La identificación del comportamiento de un
sistema informático tratado como una caja negra es
posible, para ello debe simularse el funcionamiento
del mismo con hardware y herramientas software
como el Jmeter y Matlab, por ejemplo.</p>
      <p>El modelo matemático determinado en base a los
datos recopilados del log del sistema en estudio se
aproxima bastante al modelo real, en este caso se
obtuvo un 87% de calidad de ajuste.</p>
      <p>El sensor software implementado ha permitido
calcular el tiempo de respuesta en base a los datos
almacenados en el log del mismo servidor.</p>
      <p>De los datos tomados se observa que los sistemas
informáticos poseen un comportamiento cercano al
lineal solo en tramos por lo que se sugiere
experimentar con modelos matemáticos no lineales
para comparar la calidad de ajuste de ambos.</p>
      <p>Como trabajos futuros se plantea diseñar un
controlador autonómico basado en el modelo lineal
ARX, aunque más adelante se planteará el diseño
de controladores autonómicos no lineales para
sistemas informáticos en general, ya que el
comportamiento temporal no lineal, hace adecuada
la utilización de técnicas de inteligencia artificial
como la lógica difusa y las redes neuronales.
6</p>
      <p>Referencias bibliográficas.</p>
      <p>Diao, Y.; Hellerstein, J.; Parekh, S.; Griffith, R.;
Kaiser, G.; Phung, D. (2005) A Control Theory
Foundation for Self-Managing Computing
Systems, IEEE Journal on Selected Areas in
Communications, Vol. 23, Issue 12, pp. 2213 –
2222. IEEE. DOI: 10.1109/JSAC.2005.857206.</p>
      <p>Fermín F. (2012). La Teoría de Control y la
Gestión Autónoma de Servidores Web. Memorias
del IV Congreso Internacional de Computación y
Telecomunicaciones. ISBN 978-612-4050-57-2.</p>
      <p>Fox, A.; Patterson, D. (2003). Self-repairing
computers, Scientific American, Jun2003, Vol. 288
Issue 6, pp. 54-61.</p>
      <p>Hellerstein, J.; Diao, Y.; Parekh, S.; Tilbury, D.
(2004). Feedback Control of Computing Systems.</p>
      <p>Hoboken, New Jersey: John Wiley &amp; Sons, Inc.</p>
      <p>ISBN 0-471-26637-X
Kephart, J.; Chess, D. (2003). The vision of
autonomic computing. Computer, Jan2003, Vol.
36, Issue 1, pp. 41-50. IEEE. DOI:
10.1109/MC.2003.1160055
Lalanda, P.; McCann, J.; Diaconescu, A. (2013).</p>
      <p>Autonomic Computing. Principles, Design and
Implementation. London: Springer-Verlag. ISBN
978-1-4471-5006-0
Ljung L. (1987). System Identification: Theory for the
User. Englewood Cliffs, New Jersey: Prentice Hall.</p>
      <p>Parekh, S.; Gandhi, N.; Hellerstein, J.; Tilbury, D.;
Jayram, T.; Bigus, J. (2002). Using control theory
to achieve service level objectives in performance
management. Real-Time Systems. Jul2002, Vol.
23, Issue 1/2, pp. 127-141. Norwell,
Massachusetts: Kluwer Academic Publishers. DOI:
10.1023/A:1015350520175.
Clasificadores supervisados para el análisis predictivo de muerte y
sobrevida materna</p>
      <p>Pilar Vanessa Hidalgo León
Universidad Andina del Cusco/San Jerónimo, Cusco-Perú</p>
      <p>Resumen
El presente trabajo se basa en el análisis de los
clasificadores supervisados que puedan generar
resultados aceptables para la predicción de la muerte
y sobrevida materna, según características de
pacientes complicadas durante su gestación
determinada por los expertos salubristas. Se describe
la metodología del desarrollo, las particularidades de
la muestra, además los instrumentos utilizados para el
procesamiento de los datos. Los resultados de la
investigación luego de la evaluación de cada
clasificador y entre ellos el que mejores resultados
arroja. Los histogramas acerca de cada atributo de las
pacientes, y la inclusión en la muestra. Los
parámetros determinantes para su correcta
clasificación. La comparación de cada resultado entre
cada tipo de clasificador dentro de la familia a la que
pertenece para después de identificado, implementar,
el algoritmo de Naive-Bayes con estimador de
Núcleo activado (KERNEL=TRUE), en un software
que contribuya a la toma de decisiones certera y
respaldada para los profesionales de la salud. En
conclusión se encontró un clasificador supervisado
que responde positivamente a dar cambio y mejora de
la problemática que abarca a la sobrevida materna a
pesar de sus complicaciones.</p>
      <p>Palabras Clave: Mortalidad materna,
clasificadores supervisados, Redes bayesianas,
Aprendizaje supervisado.</p>
      <p>1
El objetivo en el uso de clasificadores
supervisados (Araujo, 2006), es construir
modelos que optimicen un criterio de
rendimiento, utilizando datos o experiencia
previa. En ausencia de la experiencia humana,
para resolver una disyuntiva que requiere
explicación precisa, los sistemas implementados
por modelos clasificadores han sido parte
importante en la toma de decisiones. A parte,
cuando este problema requiere prontitud por su
naturaleza, los clasificadores transforman los
datos en conocimiento y aportan aplicaciones
exitosas.</p>
      <p>En el caso de factores de riesgo para la salud
materna, existen estudios estadísticos y
aplicaciones salubristas para determinarla, mas
no integrados simultáneamente como parte de
una probabilidad clasificatoria como modelo.</p>
      <p>Por ello, este estudio determinará, el clasificador
supervisado más eficiente en tiempo y resultado
que establezca la diferencia de clases entre
pacientes gestantes complicadas durante su
embarazo que pueden llegar a presentar síntomas
fatales y las que no, así apoyar al personal de
salud a tomar la decisión más optima y a
prevenir futuras alzas en el índice de mortalidad
de su comunidad.</p>
      <p>Los problemas que generan el alza de este
indicador ya son conocidos y puestos en valor en
esta investigación:
“Hoy en día existe suficiente evidencia que
demuestra que las principales causas de la
muerte materna son la hemorragia posparto, la
preclampsia o la sepsis y los problemas
relacionados con la presentación del feto.</p>
      <p>Asimismo, sabemos cuáles son las medidas más
eficaces y seguras para tratar estas emergencias
obstétricas. Para poder aplicarlas, es necesario
que la gestante acceda a un establecimiento de
salud con capacidad resolutiva, pero
lamentablemente muchas mujeres indígenas no
acuden a este servicio por diversas razones, tanto
relacionadas con las características geográficas,
económicas, sociales y culturales de sus grupos
poblacionales, como por las deficiencias del
propio sistema de salud.</p>
      <p>En los últimos años se han hecho muchos
esfuerzos para revertir esta situación, tanto
mediante proyectos promovidos por el estado
como ejecutados por organismos no
gubernamentales de desarrollo. Estos esfuerzos
han tenido, sin embargo, resultados desiguales
debido principalmente a la poca adecuación de
los proyectos al contexto geográfico y de
infraestructura en el que vive gran parte de la
población indígena, a sus dificultades
económicas para acceder al servicio, su cultura,
sus propios conceptos de salud y enfermedad, y
su sistema de salud.”(Cordero, 2010)
2 Contenido
Problema
¿Cuáles son los clasificadores supervisados que
predicen la muerte o la sobrevida materna con
mayor efectividad?
Entonces para determinar adecuadamente la
efectividad de cada algoritmos nos
cuestionamos:
¿Cuál es la especificidad, la
clasificación correcta y el error absoluto
y sensibilidad del clasificador
supervisado en relación a los datos de
mortalidad materna?</p>
      <p>Redes neuronales
Regresión Logística
Arboles de Decisión</p>
      <p>Algoritmos Basados en Distancias
Mediante la herramienta Weka, se determinó la
sensibilidad, la certeza más cercana de cada uno
de estos algoritmos, y cuya conclusión sugerirá
el más eficiente.</p>
      <p>Actualmente la problemática en mortalidad
materna es un indicador determinante de
desarrollo en los países Latinoamericanos.</p>
      <p>Siendo no solo un indicador de pobreza y
desigualdad sino de vulnerabilidad de los
derechos de la mujer. (OMS, 2008)
Limitaciones de la Investigación
Los datos recolectados para este estudio con
respecto a pacientes fallecidas que tuvieron
complicaciones durante el embarazo fueron 48,
ya que el estado del registro de las historias
clínicas correspondientes a los casos de
fallecimiento no son legibles, ni están
conservadas en las mejores condiciones en los
archivos de la Dirección Regional de Salud
Cusco, esto hace que la muestra no pueda ser
nutrida con mayor diversidad de datos.
(DIRESA, 2007)
Existe poca investigación acerca del tema
relacionado con el uso de clasificadores
supervisados y otras correspondientes a
diagnósticos médicos que tienen similitud con la
muestra pero ninguna que se relacione
directamente.</p>
      <p>La relevancia de los datos se limito a los
antecedentes sobre estudios en mortalidad
materna (edad, estado civil, analfabeta,
ocupación, procedencia, anticoncepción, entorno
(estrato social), controles pre-natales, ubicación
domiciliaria, tiempo de demora en atención,
atención profesional, antecedentes familiares,
espacio intergenésico (en años), paridad (número
de hijos), complicaciones no tratadas,
fallecimiento), conservando el anonimato de
cada paciente.</p>
      <p>Objetivos</p>
      <p>Determinar el clasificador supervisado
que brinde mejores resultados para el
análisis predictivo de muerte y
sobrevida materna.</p>
      <p>Luego para lograr este objetivo se debe:
Determinar la especificidad, la
clasificación correcta, el error absoluto,
la sensibilidad del clasificador
supervisado en relación a los datos de
muerte y sobrevida materna.</p>
      <p>Redes neuronales
Regresión Logística
Arboles de Decisión
Algoritmos Basados en</p>
      <p>Distancias
Hipótesis General.
H. nula: No existen clasificadores supervisados
que predicen la muerte o la sobrevida materna
con efectividad
Definición de Variables:
Variable principal:</p>
      <p>Clasificadores supervisados
Variables Implicadas:</p>
      <p>Variable
Dimensión
Indicador/
Criterios de
Medición</p>
      <p>Clase
Instrumento</p>
      <p>Clasificadores supervisados
Las Redes Neuronales, Algoritmos</p>
      <p>supervisados, Las Redes
Bayesianas, Arboles de decisión,
Regresión Logística, Algoritmos</p>
      <p>basados en instancias
Clasificación correcta,
Clasificación incorrecta,
Sensibilidad,Especificidad, Tiempo
de ejecución
Mean absolute error
Kappa statistic
Root mean squared erro, Relative
absolute error
Root relative squared error</p>
      <p>Numérica discreta</p>
      <p>Weka 3.5.7, Explored</p>
      <p>Tabla I: Variables implicadas
Metodología de investigación
Cuasi Experimental, Aplicada, Inductiva
Descripción
recolección
de la
muestra y
método</p>
      <p>de
Los datos recolectados en todas 100 historias
clínicas (HC) de casos de sobrevida y de casos
de muerte materna.</p>
      <p>Las características de las gestantes son muy
similares entre si y corresponden a la población
de la ciudad del Cusco, del archivo en la Red
Sur de la Dirección Regional de Salud sobre el
control sanitario de mortalidad materna de la
Región Cusco.</p>
      <p>El número es limitado, pues las historias desde
1992 al 2011, no han sido redactadas ni
conservadas en el mejor estado haciendo difícil
la tarea de interpretar los datos suficientes para
ser analizados.</p>
      <p>Estas pacientes no fueron necesariamente
atendidas desde el inicio en estos
establecimientos, sino que debido a sus
complicaciones durante el parto y e embarazo
fueron derivadas a las capitales y luego a los
establecimientos de mayor capacidad
resolutiva para su atención.</p>
      <p>Los datos de las historias clínicas que
incluyeran:
•
•
•
•
•
•
•
•
•
•
•
•
•
•
•
•</p>
      <p>Edad
Estado civil
Analfabeta
Ocupación
Procedencia
Anticoncepción
Entorno (estrato social)
Controles pre-natales
Ubicación domiciliaria
Tiempo de demora en atención
Atención profesional
Antecedentes familiares
Espacio intergenésico (en años)
Paridad (#de hijos)
Complicaciones no tratadas</p>
      <p>Fallecimiento
Entre 52 sobrevivientes y 48 fallecidas, ambos
grupos con similares características, siendo
factores determinantes: (Ramírez, 2009)</p>
      <p>Ubicación domiciliaria /tiempo demora en</p>
      <p>atención</p>
      <p>Controles &gt; 2: n (49.0/2.0)</p>
      <p>PRODECENCIA = rural: s (6.0/1.0)
Técnica e instrumentos de investigación
Se utilizó la herramienta Weka Explorer para
la interpretación de los datos.</p>
      <p>Las opciones de clasificación supervisada y los
algoritmos que propone esta herramienta.
(Corso, 2009)
Se evaluaron los siguientes clasificadores
supervisados:
• Las Redes Neuronales
- MultilayerPerceptron
- RBFNetwork
• Las Redes Bayesianas
- BayesNet
- Bayes simple estimator
- BMA bayes
- Naive-Bayes
- BayesNet Kernel
- Naive-Bayes
- Discretizacion Supervisada
• Arboles de decisión
- J48
- DecisionTable
• Regresión Logística
- MultiClassClassifier
- Logistic
• Algoritmos basados en Distancias
- IBK
- LWL
- KStar
Estos resultados fueron comparados con las
reglas de clasificación que nos proporcionará
los algoritmos de predicción inmediata:</p>
      <p>OneR</p>
      <p>ZeroR
•
•
•
•
•
•
•
Procedimiento de recolección de datos
Las Historias Clínicas se insertaron en un
fichero CSV (delimitado por comas) cuya
cabecera contiene las etiquetas de cada
atributo, y la última columna se refiere a la
clase a la pertenecen.</p>
      <p>Con respecto a los indicadores particulares en
cada uno de los atributos de los sujetos de la
clase tenemos los siguientes valores: Tabla II</p>
      <p>La calidad de la estructuración
Comparar todas mediciones en cada
clasificador por algoritmo.</p>
      <p>Evaluar los resultados.</p>
      <p>Interpretar los resultados.
2.9. Plan de análisis de la información</p>
      <p>Determinación de objetivos
Preparación de datos
Selección de datos:
Edad
Estado civil
Analfabeta
Ocupación
Procedencia
Anticoncepc
ión
Entorno
(estrato
social)
Controles
Ubicación
domiciliaria/
tiempo
demora en
atención
Personal de
atención
profesional
Antecedente
s familiares
Espacio
intergenésic
o (en años)
Paridad (#de
hijos)
Complicacio
nes no
tratadas
Fallecida
o Identificación de las fuentes de
información externas e internas y
selección del subconjunto de
datos necesario.</p>
      <p>Pre procesamiento: estudio de la calidad
de los datos y determinación de las
operaciones de minería que se pueden
realizar.</p>
      <p>Transformación de datos: conversión de
datos en un modelo analítico.</p>
      <p>Análisis de datos interpretación de los
resultados obtenidos en la etapa anterior,
generalmente con la ayuda de una
técnica de visualización.</p>
      <p>Asimilación de conocimiento
descubierto(Calderón, 2006)
Minería de datos: tratamiento
automatizado de los datos seleccionados
con una combinación apropiada de
algoritmos. (Ramirez, 2011).</p>
      <p>NEGATIVO
Menor a 19 y
mayor 35
Soltera
Analfabeta
No
remunerada
Rural
No
Bajo</p>
      <p>POSITIVO
Entre 19-35
Pareja
Primaria
Remunerada
Urbana
Si
Medio-alto</p>
      <p>VALORES
14-48
Soltera-pareja
Analfabetaprimariasecundariasuperior
Remunerada-no
remunerada
Rural-urbana
Si-no</p>
      <p>Baja-alta
De 1 a 5
Más de 2
horas
6 a mas</p>
      <p>0-12
Menos 3 del
ee.ss</p>
      <p>Menos de 1 hora,
1-2,3-5,6 a mas
No
No
Menos de 2
mayor a 4
Primípara o
mas de 4
Complicacion
es antes y
durante
Si</p>
      <p>Si
Si
Entre 2-4
Entre 2-4
Sin
complicacion
es
No</p>
      <p>No-si
No-si
0-10
No-si
No-si
Primera
gesta,13,4-6, menos 1
Tabla II: Rango de valores de los atributos en las
pacientes de la muestra
Cada Atributo es evaluado visualmente por los
histogramas que arroja Weka (en total 15), por
ejemplo con respecto a la EDAD de las pacientes
de la muestra.</p>
      <p>Gráfico 1: Histograma de la Edad de las
pacientes de la muestra
Este histograma nos muestra el intervalo de edad
de las pacientes de la muestra, la diferencia de
colores determinan la clase a la que pertenece
cada intervalo. Podemos observar lo siguiente:</p>
      <p>Las pacientes del intervalo 14-20 años
pertenecen en su mayoría a la clase
“sobreviviente”
Las pacientes del intervalo 21-31 años
tienen un mayor porcentaje de sobrevida,
coincide con el promedio de edad
adecuado y de la muestra.</p>
      <p>El intervalo de paciente entre 32-40 años
tiene mayor porcentaje de muerte.</p>
      <p>Las pacientes mayores a 40 años
pertenecen a la clase “fallecida” en gran
porcentaje.</p>
      <p>Resultados por algoritmo testeado:
Los algoritmos usados para evaluar la base de
datos en mortalidad materna dieron como
resultado cifras continuas indicando los
siguientes sucesos: Tabla 2.</p>
      <p>Especificidad: es la probabilidad de que
pacientes complicadas y de riesgo
pertenezcan a la clase Sobreviviente. Es
decir los verdaderos negativos.</p>
      <p>Fracción de verdaderos negativos (FVN).</p>
      <p>Demuestra la cantidad de pacientes que
realmente pertenecen a la clase
Sobreviviente. Quiere decir que si el
algoritmo estudiado tiene alto porcentaje
de especificidad determina con gran éxito
la probabilidad de sobrevida en pacientes
complicadas durante su embarazo según
los datos proporcionados en la ficha de
antecedentes.
• Clasificación correcta: de la totalidad de
datos, entre los que 52 que pertenecen a la
clase Sobreviviente, y los 48 que pertenece
a la clase Fallecida, determina dentro de
cada clase cuantas instancias luego de la
construcción del clasificador cuantas si
pertenecen a la clase determinada.</p>
      <p>En el caso de pertenecer a la clase
sobreviviente o a la fallecida de las 100
instancias cuantas fueron clasificadas
correctamente.</p>
      <p>Clasificación incorrecta: del mismo
modo la cantidad de instancias que no
fueron clasificadas de manera correcta,
son las que de manera supervisada se sabe
que pertenecen a una u otra clase y fueron
incluidas dentro de la cual no eran. Si el
indicador emite un número mayor al 50%
de la cantidad total de instancias, no se
debe considerar como eficiente.</p>
      <p>Sensibilidad: es la capacidad del
algoritmo de clasificar a las pacientes
complicadas dentro de la clase Fallecidas.</p>
      <p>Es decir que si el clasificador tiene un alto
porcentaje tiene mejor curva de corte y
discernimiento entre los sujetos que
pertenecen o no a la clase fallecida, es así
que si la cifra de sensibilidad es del 90%,
existe entonces esa probabilidad de que la
paciente fallezca.</p>
      <p>Mean absolute error: Se define error
absoluto de una medida la diferencia entre
el valor medio obtenido y el hallado en esa
medida todo en valor absoluto.</p>
      <p>Entonces el promedio de error absoluto, es
la suma de los errores absolutos de
clasificación en cada uno de los sujetos
llevados a promedio. El clasificador que
arroje mayor cifra (mayor a 0.1) define un
error de clasificación alto, por lo cual no
se debe considerar por sobre los que
arrojen una cifra menor.</p>
      <p>Tiempo de ejecución: medido en
segundos es la cantidad de tiempo que
demora en construir la arquitectura del
clasificador y en arrojar resultados.</p>
      <p>Puede que un clasificador se defina como
eficiente si el tiempo que emplea en emitir
resultados es menor a 5 segundos, aun así
depende de los demás indicadores para
valerse de esta característica.</p>
      <p>Kappa statistic: el Kappa statistic es la
concordancia de comparación que tienen
los observadores de clasificación. Quiere
decir en una matriz de clasificación, el
índice esperado entre el diagonal principal
esperada (Xii elemento clasificado en la
misma clase por ambos observadores) y el
índice real luego de la clasificación
efectuada por la arquitectura seleccionada
(sea regresión lineal, backpropagation,
Naive-bayes, etc.), es la diferencia en
porcentaje de su lejanía a este valor.</p>
      <p>Si por ejemplo, la matriz esperada clasifica
el valor en 25.00 y el resultado de la
arquitectura es 26.7, la diferencia seria, 1.7
equivale al 90.32%. Entonces cuanto mas
grande sea el porcentaje, estará más cerca
de ser considerado eficiente.</p>
      <p>Root mean squared error: error
cuadrático medio, es una medida de uso
frecuente de las diferencias entre los
valores pronosticados por un modelo o un
estimador y los valores realmente
observados. RMSD es una buena medida
de precisión, pero sólo para comparar
diferentes errores de predicción dentro de
un conjunto de datos y no entre los
diferentes, ya que es dependiente de la una
escala muestra. Estas diferencias
individuales también se
denominan residuos, y la RMSD sirve para
agregarlos en una sola medida de la
capacidad de predicción.</p>
      <p>Relative absolute error: es el error
relativo a cada característica de la clase,
por ejemplo el error relativo de tener de
Espacio Intergenésico 0.5 años y
pertenecer o no a la clase fallecida, la
clasificación indicaría que si pertenece,
por ser el valor indicado para aquellas
pacientes que están en peligro.</p>
      <p>En este caso el valor positivo para
pertenecer al clase sobreviviente es de
entre 2-4 años o primera gesta: 0.5 años
incluido en la clase fallecido si el error
entre los valores determinados por la clase
y el valor ingresado es menor a 1.</p>
      <p>Root relative squared error: La raíz
relativa E de error al cuadrado i de un
programa individual i es evaluado por la
ecuación:</p>
      <p>donde P (ij) es el valor predicho por el
programa para el individuo i j muestra de
casos (de los casos de la muestra n), T j es el
valor objetivo para la muestra j caso,
y</p>
      <p>está dada por la fórmula:
Para un ajuste perfecto, el numerador es
igual a 0 y E i = 0. Así, el E i índice varía de
0 a infinito, con el ideal que corresponde a
0.</p>
      <p>INDICADORES OneR
Especificidad
Clasificación
correcta
Clasificación
incorrecta
Sensibilidad
Mean
error</p>
      <p>absolute
0.836
0.91</p>
      <p>RBFNetwork
0.9
0.868</p>
      <p>0.914
Kappa statistic
Root mean
squared error
Relative absolute
error
Root relative
squared error
Tabla IV: Algoritmos de clasificación
supervisada con mejores resultados por familia
Descripción de la metodología propuesta
Se propone entonces que evaluando cada uno
de los indicadores más importantes en la
construcción del clasificador, se descarten
aquellos que no cumplen las siguientes
características:
• Especificidad &gt; 90%
• Clasificación correcta &gt; 90 instancias
• Clasificación incorrecta &lt; 10 instancias
• Sensibilidad &gt;90%
• Mean absolute error &lt; 0.1 ideal
• Kappa statistic &gt;0.79, &gt;0.9 ideal
• Root mean squared error &lt; 0.3, &lt;1 ideal
• Relative absolute error &lt;25%, &lt;1 ideal
• Root relative squared &lt;50%, 0 ideal
El clasificador que cumpla con estas
especificaciones, se considera como optimo
para la integración en un sistema que evaluará
la base de datos que contengan los datos de las
pacientes a comparar con un registro nuevo de
entrada.</p>
      <p>Etapa de Evaluación:
Cada uno de los clasificadores estudiados, que
analizaron la muestra de pacientes, arrojaron los
indicadores que muestra la tabla, en nivel de
rendimiento y buenos resultados, aceptables para
el estudio, la arquitectura que lleva los
porcentajes mas óptimos dentro de los
parámetros, es el algoritmo de Naive bayes con
estimador de núcleo.</p>
      <p>Descartando así el resto de arquitecturas como
apropiadas para este tipo de muestras. Además
de visualizar de manera más clara y correcta los
resultados comunes entre cada una de las
arquitecturas dejando atrás aquel paradigma que
incluye a la regresión logística como la más
adecuada a la hora de realizar diagnósticos
preventivos en salud.</p>
      <p>Kernel estimator, al ser un método de
construcción no paramétrico, es mas flexible que
los clasificadores que incluyen parámetros.</p>
      <p>Dividiendo la muestra en 40 grupos en la
variable edad, paridad 15, controles 17, para la
exploración descubre nuevas alternativas de
clasificación en los atributos de la clase,
mostrándolos como relevantes.</p>
      <p>El problema más resaltante es que requiere
mayor tamaño de muestra para mantener estos
resultados, ya que utilizado el agrupamiento para
la clasificación (clustering) que podría ser una
desventaja si de clasificadores supervisados
hablamos. Esto se denomina the curse of
dimensionality, pues dependen de la elección de
un parámetro suavizado, en este caso los valores
la edad, número de controles, y número de hijos,
transformándola en no objetiva.</p>
      <p>Cuando Naive-Bayes actúa con la estimación no
paramétrica, las estructuras que se construyen a
partir de esta arquitectura se obtienen a partir de
un árbol formado con variables potencialmente
predictoras multidimensionales. (Serrano, 2010)
Con respecto a la respuesta de las hipótesis
podemos afirmar lo siguiente:</p>
      <p>Encontramos que las arquitecturas
propuestas por esta memoria, totas
conservan una efectividad bastante
aceptable e cada uno de sus indicadores
que en conjunto hacen un 80% de
eficacia. Siendo el más optimo el de
Naive-Bayer Kernel.</p>
      <p>Descarta estas
trabajo realizado.</p>
      <p>hipótesis luego del
Las Redes Bayesianas brindan al estudio
de los datos en mortalidad y sobrevida
materna una especificidad 91%,
clasificación correcta 91%., error
absoluto 0.1142, y sensibilidad 0.914
recomendada.</p>
      <p>Las Redes Neuronales brindan al estudio
de los datos en mortalidad y sobrevida
materna una especificidad 90%,
clasificación correcta 90%, error
absoluto 0.1468 y sensibilidad 90%
recomendada.</p>
      <p>Los Arboles de Decisión brindan al
estudio de los datos en mortalidad y
sobrevida materna una especificidad
90%, clasificación correcta 90%, error
absoluto 0.1174 y sensibilidad 90%
recomendada.</p>
      <p>Los Algoritmos basados en Distancias
brindan al estudio de los datos en
mortalidad y sobrevida materna una
especificidad 88.9%, clasificación
correcta 89%, error absoluto 0.168 y
sensibilidad 90.2% recomendada.</p>
      <p>La Regresión Logística brinda al estudio
de los datos en mortalidad y sobrevida
materna una especificidad 82%,
clasificación correcta 82%, error
absoluto 0.1753 y sensibilidad 82%
recomendada.</p>
      <p>Gráfico 2: Indicadores estadísticos de
todos los algoritmos de clasificación
supervisada
4 Recomendaciones</p>
      <p>Se recomienda usar datos discretizados,
si se desea usar Naive-Bayes para su
evaluación.</p>
      <p>Agrupar los datos de la manera
propuesta en la Tabla 1, así podremos
respetar los parámetros acerca de
mortalidad materna que establecen los
expertos salubristas.</p>
      <p>Es importante comparar los resultados
con la regla de decisión OneR para
próximos experimentos, pues nos da la
mejor noción de veracidad y efectividad
a la hora de analizar la información.</p>
      <p>Se pretende implementar un sistema en
R Project, para ingresar nuevos registros
y que lleve en la memoria la base de
datos recolectada a través de este trabajo.</p>
      <p>Para ayudar a la muestra numeraria que
requiere Naive- Bayes Kernel, es
necesario ingresar los registros de las
muertes maternas y complicaciones
diarias de las gestantes a nivel Nacional
en la base de datos.
Agradezco al Dr. Basilio Araujo y al Dr.</p>
      <p>Yosu Yuramendi, por su apoyo y
perseverancia para la culminación de este
proyecto de fin de master.</p>
      <p>Al personal de salud que proporciono la
información clínica que se usó en la
investigación.
5 Síntesis curricular de la autora:
6 Referencias bibliográficas</p>
      <p>Clasificadores supervisados: el objetivo es
obtener un modelo clasificatorio valido
para permitir tratar casos futuros.( 2006)
Araujo, B. S. Aprendizaje Automatico:
conceptos básicos y avanzados. Madrid,</p>
      <p>España: Pearson Prentice Hall,
Cordero Muñoz, L., Luna Flórez, A., &amp;</p>
      <p>Vattuone Ramírez, M. (2010) Salud de la
mujer indígena : intervenciones para reducir
la muerte materna. © Banco Interamericano
de Desarrollo.</p>
      <p>Antonio Serrano, E. S., (2010) Redes</p>
      <p>Neuronales Artificiales. Universidad de</p>
      <p>Valencia.</p>
      <p>Calderón Saldaña, J., &amp; Alzamora de los</p>
      <p>Godos Urcia, L. (2009) Regresión Logística
Aplicada A La Epidemiología. Revista</p>
      <p>Salud, Sexualidad y Sociedad, 1[4].</p>
      <p>Calderón, S. G, (2006) Una Metodología</p>
      <p>Unificada para la Evaluación de Algoritmos
de Clasificación tanto Supervisados como
No-Supervisados. México D. F.: resumen de
tesis doctoral.</p>
      <p>Corso, C. L. (2009) Aplicación de algoritmos
de clasificación supervisada usando Weka.</p>
      <p>Córdoba: Universidad Tecnológica
Nacional, Facultad Regional Córdoba.</p>
      <p>María García Jiménez, &amp; Aránzazu Álvarez</p>
      <p>Sierra. (2010) Análisis de Datos en WEKA
– Pruebas de Selecitividad. Articulo.</p>
      <p>María N. Moreno García* L. A. (2005)</p>
      <p>Obtención Y Validación De Modelos De
Estimación De software Mediante Técnicas
De Minería De Datos. Revista colombiana
de computación, 3[1], 53-71.</p>
      <p>Msp Mynor Gudiel M., 1. E. (2001-2002)</p>
      <p>Modelo Predictor De Mortalidad Materna.</p>
      <p>Mortalidad Materna, 22-29.</p>
      <p>Organización Mundial De La Salud.</p>
      <p>Mortalidad Materna En 2005, (2008) :
Estimaciones Elaboradas Por La Oms, El
Unicef, El Unfpa Y El Banco. Ginebra 27:</p>
      <p>Ediciones de la OMS.</p>
      <p>Porcada, V. R. (2003) Clasificación</p>
      <p>Supervisada Basada En Redes Bayesianas.</p>
      <p>Aplicación En Biología Computacional.</p>
      <p>Madrid: Tesis Doctoral.</p>
      <p>Ramírez Ramírez, R., &amp; Reyes Moyano, J. ,
(2011). Indicadores De Resultado
Identificados En Los Programas
Estratégicos. Lima: Encuesta Demográfica</p>
      <p>Y De Salud Familiar – Endes
Ramirez, C. J. (2009)Caracterización De</p>
      <p>Algunas Técnicas Algoritmicas De La
Inteligencia Artificial Para El
Descubrimiento De Asociaciones Entre
Variables Y Su Aplicación En Un Caso De
Investigación Específico. Medellin: Tesis</p>
      <p>Magistral.</p>
      <p>Reproductive Health Matters, (2009)</p>
      <p>Mortalidad Y Morbilidad
Materna:Gestación Más Segura Para Las
Mujeres. Lima: © Reproductive Health</p>
      <p>Matters.</p>
      <p>Vilca, C. P. (2009) Clasificación De Tumores</p>
      <p>De Mama Usando Métodos. Lima: Tesis.</p>
      <p>Dirección Regional De Salud Cusco, (2007)</p>
      <p>Análisis De Situación De La Mortalidad</p>
      <p>Materna Y Perinatal, Región Cusco.</p>
      <p>Ministerio De Salud, Dirección General De</p>
      <p>Epidemiología Situación De Muerte
Materna, (2010-2011), Perú.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Zadeh L. A.</surname>
          </string-name>
          <year>1965</year>
          .
          <string-name>
            <given-names>Fuzzy</given-names>
            <surname>Sets</surname>
          </string-name>
          .
          <source>Information and Control</source>
          ,
          <volume>8</volume>
          (
          <issue>3</issue>
          ):
          <fpage>338</fpage>
          -
          <lpage>353</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Herrera F.</surname>
          </string-name>
          <year>2005</year>
          .
          <article-title>Genetic fuzzy systems: Status, critical considerations and future directions</article-title>
          .
          <source>International Journal of Computational Intelligence Research</source>
          ,
          <volume>1</volume>
          (
          <issue>1</issue>
          ):
          <fpage>59</fpage>
          -
          <lpage>67</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Herrera F.</surname>
          </string-name>
          <year>2008</year>
          .
          <article-title>Genetic fuzzy systems: taxonomy, current research trends and prospects</article-title>
          .
          <source>Evolutionary Intelligence</source>
          ,
          <volume>1</volume>
          (
          <issue>1</issue>
          ):
          <fpage>27</fpage>
          -
          <lpage>46</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Nojima</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Ishibuchi</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <year>2013</year>
          .
          <article-title>Multiobjective genetic fuzzy rule selection with fuzzy relational rules</article-title>
          .
          <source>IEEE International Workshop on Genetic and Evolutionary Fuzzy Systems (GEFS)</source>
          ,
          <volume>1</volume>
          :
          <fpage>60</fpage>
          -
          <lpage>67</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>S.-M</given-names>
          </string-name>
          , Chang,
          <string-name>
            <given-names>Y.-C.</given-names>
            ,
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.-S.</surname>
          </string-name>
          <year>2013</year>
          .
          <article-title>Fuzzy Rules Interpolation for Sparse Fuzzy Rule-Based Systems Based on Interval Type-2 Gaussian Fuzzy Sets and Genetic Algorithms</article-title>
          .
          <source>IEEE Transactions on Fuzzy Systems</source>
          ,
          <volume>21</volume>
          (
          <issue>3</issue>
          ):
          <fpage>412</fpage>
          -
          <lpage>425</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Jalesiyan</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yaghubi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Akbarzadeh</surname>
            ,
            <given-names>T.M.R.</given-names>
          </string-name>
          <year>2014</year>
          .
          <article-title>Rule selection by Guided Elitism genetic algorithm in Fuzzy Min-Max classifier</article-title>
          .
          <source>Conference on Intelligent Systems</source>
          ,
          <volume>1</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>De Jong</surname>
            <given-names>KA</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spears</surname>
            <given-names>WM</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gordon DF</surname>
          </string-name>
          .
          <year>1993</year>
          .
          <article-title>Using genetic algorithms for concept learning</article-title>
          .
          <source>Mach Learn</source>
          ,
          <volume>13</volume>
          :
          <fpage>161</fpage>
          -
          <lpage>188</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Holland</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Reitman</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <year>1978</year>
          .
          <source>Cognitive Systems Based on Adaptive Algorithms</source>
          , ACM SIGART Bulletin
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Gonzalez</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Perez</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <year>1999</year>
          .
          <article-title>SLAVE: A genetic learning system based on an iterative approach</article-title>
          .
          <source>IEEE Transactions on Fuzzy Systems</source>
          ,
          <volume>7</volume>
          :
          <fpage>176</fpage>
          -
          <lpage>191</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Giordana</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Neri</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <year>1995</year>
          .
          <article-title>Search-intensive concept induction</article-title>
          .
          <source>Evol Comput</source>
          ,
          <volume>3</volume>
          :
          <fpage>375</fpage>
          -
          <lpage>416</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Casillas</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Carse</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <year>2009</year>
          .
          <article-title>Special issue on Genetic Fuzzy Systems: Recent Developments and Future Directions</article-title>
          .
          <source>Soft Comput.</source>
          ,
          <volume>13</volume>
          (
          <issue>5</issue>
          ):
          <fpage>417</fpage>
          -
          <lpage>418</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Deb</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Pratap</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Agarwal</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Meyarivan</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <year>2002</year>
          .
          <article-title>A Fast and Elitist Multiobjective Genetic Algorithm: NSGA-II</article-title>
          .
          <source>Trans. Evol</source>
          . Comp.,
          <volume>6</volume>
          (
          <issue>2</issue>
          ):
          <fpage>182</fpage>
          -
          <lpage>197</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Nozaki</surname>
            ,
            <given-names>k.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ishibuchi</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tanaka</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <year>1996</year>
          .
          <article-title>Adaptive fuzzy rule-based classification systems</article-title>
          .
          <source>IEEE Trans. Fuzzy Systems</source>
          ,
          <volume>4</volume>
          (
          <issue>3</issue>
          ):
          <fpage>238</fpage>
          -
          <lpage>250</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Beltran</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <year>2007</year>
          .
          <article-title>Anlise e reconhecimento digital de formas biolgicas para o diagnstico automtico de parasitas do gnero Eimeria</article-title>
          .
          <source>PhD.</source>
          Teses - USP - Brazil.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Kumar</surname>
            ,
            <given-names>S.U.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Inbarani</surname>
            ,
            <given-names>H.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kumar</surname>
            ,
            <given-names>S.S.</given-names>
          </string-name>
          <year>2013</year>
          .
          <article-title>Bijective soft set based classification of medical data</article-title>
          .
          <source>International Conference on Pattern Recognition, Informatics and Mobile Engineering</source>
          ,
          <volume>1</volume>
          :
          <fpage>517</fpage>
          -
          <lpage>521</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>Yongzhi</surname>
          </string-name>
          , Ma.,
          <string-name>
            <surname>Hong</surname>
            <given-names>Gao</given-names>
          </string-name>
          , Yi Ding, Wei Liu.
          <year>2013</year>
          .
          <article-title>Logistics market segmentation based on extension classification</article-title>
          .
          <source>International Conference on Information Management, Innovation Management and Industrial Engineering</source>
          ,
          <volume>2</volume>
          :
          <fpage>216</fpage>
          -
          <lpage>219</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>Silla</surname>
            ,
            <given-names>C.N.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Kaestner</surname>
            ,
            <given-names>C.A.A.</given-names>
          </string-name>
          <year>2013</year>
          .
          <article-title>Hierarchical Classification of Bird Species Using Their Audio Recorded Songs</article-title>
          .
          <source>IEEE International Conference on Systems, Man, and Cybernetics</source>
          , 1:
          <fpage>1895</fpage>
          -
          <lpage>1900</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Lakhmi C.</given-names>
            and
            <surname>Martin</surname>
          </string-name>
          <string-name>
            <surname>N.</surname>
          </string-name>
          <year>1998</year>
          .
          <article-title>Fusion of Neural Networks</article-title>
          ,
          <source>Fuzzy Systems and Genetic Algorithms: Industrial Applications (International Series on Computational Intelligence)</source>
          . CRC Press
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <surname>Cordon</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Herrera</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hoffmann</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Magdalena</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <year>2001</year>
          .
          <article-title>Genetic Fuzzy Systems: Evolutionary Tuning and Learning of Fuzzy Knowledge Bases World Scientific</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <surname>Srinivas</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Deb</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <year>1994</year>
          .
          <article-title>Multiobjective Optimization Using Nondominated Sorting in Genetic Algorithms</article-title>
          .
          <source>Evolutionary Computation</source>
          ,
          <volume>2</volume>
          :
          <fpage>221</fpage>
          -
          <lpage>248</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <surname>Pareto</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <year>1896</year>
          .
          <article-title>Cours d'Economie Politique</article-title>
          . Droz, Genve.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <surname>Hinojosa</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Camargo</surname>
            ,
            <given-names>H.A.</given-names>
          </string-name>
          <year>2012</year>
          .
          <article-title>Multiobjective genetic generation of fuzzy classifiers using the iterative rule learning</article-title>
          .
          <source>International Conference on Fuzzy Systems</source>
          ,
          <volume>1</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <surname>Gonzalez</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Perez</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <year>2009</year>
          .
          <article-title>Improving the genetic algorithm of SLAVE</article-title>
          .
          <source>Mathware and Soft Computing</source>
          ,
          <volume>16</volume>
          :
          <fpage>59</fpage>
          -
          <lpage>70</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <surname>Ji</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sum</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <year>2007</year>
          .
          <article-title>Chinese Terminology Extraction Using Window-Based Contextual Information</article-title>
          .
          <source>Proceedings of the 8th International Conference on Computational Linguistics, Intelligent Text Processing (CICLing07)</source>
          , pp.
          <fpage>62</fpage>
          -
          <lpage>74</lpage>
          . Springer-Verlag, Mexico City, Mexico.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <surname>Krauthammer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nenadic</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <year>2004</year>
          .
          <article-title>Term Identification in the Biomedical Literature</article-title>
          .
          <source>Journal of Biomedical Informatics</source>
          , vol.
          <volume>37</volume>
          , pp.
          <fpage>512</fpage>
          -
          <lpage>526</lpage>
          . Elsevier Science, San Diego, USA.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <surname>Lossio-Ventura</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jonquet</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roche</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Teisseire</surname>
            <given-names>M.</given-names>
          </string-name>
          <year>2014</year>
          .
          <article-title>BIOTEX: A system for Biomedical Terminology Extraction, Ranking, and Validation</article-title>
          .
          <source>Proceedings of the 13th International Semantic Web Conference (ISWC'14)</source>
          . Trento, Italy.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <surname>Lossio-Ventura</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jonquet</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roche</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Teisseire</surname>
            <given-names>M.</given-names>
          </string-name>
          <year>2014</year>
          .
          <article-title>Integration of linguistic and Web information to improve biomedical terminology ranking</article-title>
          .
          <source>Proceedings of the 18th International Database Engineering and Applications Symposium (IDEAS'14)</source>
          , ACM. Porto, Portugal.
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <string-name>
            <surname>Lossio-Ventura</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jonquet</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roche</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Teisseire</surname>
            <given-names>M.</given-names>
          </string-name>
          <year>2014</year>
          .
          <article-title>Yet another ranking function to automatic multi-word term extraction</article-title>
          .
          <source>Proceedings of the 9th International Conference on Natural Language Processing (PolTAL'14)</source>
          , Springer LNAI. Warsaw, Poland.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <surname>Lossio-Ventura</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jonquet</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roche</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Teisseire</surname>
            <given-names>M.</given-names>
          </string-name>
          <year>2014</year>
          .
          <article-title>Biomedical Terminology Extraction: A new combination of Statistical, Web Mining Approaches</article-title>
          . Proceedings of Journe´
          <article-title>es internationales d'Analyse statistique des Donne´es Textuelles (JADT2014)</article-title>
          . Paris, France.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <surname>Lossio-Ventura</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jonquet</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roche</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Teisseire</surname>
            <given-names>M.</given-names>
          </string-name>
          <year>2013</year>
          .
          <article-title>Combining C-value, Keyword Extraction Methods for Biomedical Terms Extraction</article-title>
          .
          <source>Proceedings of the Fifth International Symposium on Languages in Biology, Medicine (LBM13)</source>
          , pp.
          <fpage>45</fpage>
          -
          <lpage>49</lpage>
          , Tokyo, Japan.
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          <string-name>
            <surname>Neveol</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grosjean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Darmoni</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zweigenbaum</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <year>2014</year>
          .
          <article-title>Language Resources for French in the Biomedical Domain</article-title>
          .
          <source>Proceedings of the 9th International Conference on Language Resources and Evaluation (LREC'14)</source>
          . Reykjavik, Iceland
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          <string-name>
            <surname>Noh</surname>
          </string-name>
          , TG., Park, SB.,
          <string-name>
            <surname>Yoon</surname>
          </string-name>
          , HG.,
          <string-name>
            <surname>Lee</surname>
          </string-name>
          , SJ., Park, SY.
          <year>2009</year>
          .
          <article-title>An Automatic Translation of Tags for Multimedia Contents Using Folksonomy Networks</article-title>
          .
          <source>Proceedings of the 32nd International ACM SIGIR Conference on Research, Development in Information Retrieval SIGIR '09</source>
          , pp.
          <fpage>492</fpage>
          -
          <lpage>499</lpage>
          . Boston, MA, USA, ACM.
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          <string-name>
            <surname>Noy</surname>
            ,
            <given-names>N. F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shah</surname>
            ,
            <given-names>N. H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Whetzel</surname>
            ,
            <given-names>P. L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dai</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dorf</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Griffith</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jonquet</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rubin</surname>
            ,
            <given-names>D. L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Storey</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chute</surname>
            ,
            <given-names>C.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Musen</surname>
            ,
            <given-names>M. A.</given-names>
          </string-name>
          <year>2009</year>
          .
          <article-title>BioPortal: ontologies and integrated data resources at the click of a mouse</article-title>
          .
          <source>Nucleic acids research</source>
          , vol.
          <volume>37</volume>
          (
          <issue>suppl 2</issue>
          ), pp
          <fpage>170</fpage>
          -
          <lpage>173</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          <string-name>
            <surname>Van Eck</surname>
            ,
            <given-names>N.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Waltman</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Noyons</surname>
            ,
            <given-names>E.CM.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buter</surname>
            ,
            <given-names>R.K.</given-names>
          </string-name>
          <year>2010</year>
          .
          <article-title>Automatic term identification for bibliometric mapping</article-title>
          .
          <source>Scientometrics</source>
          , vol.
          <volume>82</volume>
          , pp.
          <fpage>581</fpage>
          -
          <lpage>596</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>