<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Derivation of Complex Linguistic Summaries from Databases for a More Human Consistent Descriptive Data Analysis</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Gerda Bortsova</string-name>
          <email>gerdabortsova@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pakizar Shamoi</string-name>
          <email>pakita.shamoi@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, Kazakh-British Technical University</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Words provide a natural way of perceiving and manipulating information by humans. Based on that premise, a concept of linguistic description of phenomena arose, and emerged into a large and growing field. In this paper, we present the approach for derivation of meaningful linguistic summaries from databases. We make an emphasis on description of relationships between attributes of a dataset, which can find applications in a number of domains. To demonstrate that, we took a dataset obtained during a research of cervical osteochondrosis among miners and developed a prototype system, which can produce important and interesting conclusions from data, such as “Muscle strength of miners does not change significantly with an increase of work experience from about 10 years to around 15. However, when experience exceeds 20 years, it decreases dramatically.”</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>An extensiveness of the use of information technologies in
almost all areas of human activity has grown drastically in
the past few decades, enabling a collection of huge amounts
of various data, which have to undergo some kind of
analysis for the benefit of individuals, organizations, or
nations. There are many types of data analyses, such as
descriptive, exploratory, inferential, predictive, causal,
mechanistic, with all of them having different goals and
approaches. In this paper we will concentrate our attention
on the most frequently used one, that is, descriptive data
analysis.</p>
      <p>Descriptive statistics and visualization are commonly
used methods in the descriptive analysis. However, they
constitute only a first step in the process of analysis: the
remaining step is interpreting the statistical figures and
graphs obtained during the first part. Usually, the
interpretation is given in the form of a summary in natural
language, such as “men are exposed to alcohol dependence
much more frequently, than women”. Human-generated
summaries often contain subjective and fuzzy terms, such as
“much more”, “significantly”, “large”, “young”, etc., which
eases perception of this information. For example, a
sentence “last week, 15,000 units of the product were sold”
contains less information for an analyst and sounds less
natural compared to “last week, the sales were quite weak”.</p>
      <p>
        Based on the fact that natural language is easier to
perceive for humans than numbers, a concept of linguistic
description, or summarization, of data arose. One of the
good examples of work in this area is [
        <xref ref-type="bibr" rid="ref6 ref7">6-7</xref>
        ], which presents
a method that employs fuzzy logic to discover relevant,
nontrivial dependencies in multivariate datasets in the form
of short sentences that signify some patterns in the data. We
believe that this approach, as well as those we mention
further in the paper, can be enhanced so as to rich a higher
degree of efficacy and widen the area of application of them.
Particularly, we argue that such characteristics of a system
for linguistic summarization as ability to identify how a
variable of interest behaves in relation to the other, or to
compare groups of objects in the dataset with respect to a
chosen parameter would be very helpful for many
applications. Specifically, we believe they are crucial when
applied to a field of medical research.
      </p>
      <p>In the next section of this paper we outline the reasons for
choosing this research topic. In the part called “Overview of
Existing Approaches” we familiarize the reader with
modern developments in the field of linguistic summaries
derivation and explain novelty of our work. Then we
describe our methodology in detail, after which we present
results of its implementation in “Application” section and
make concluding remarks.</p>
    </sec>
    <sec id="sec-2">
      <title>Motivation</title>
      <p>
        In this paper we demonstrate a method to derive important
and interesting conclusions from data, which were obtained
during a research of cervical osteochondrosis among miners.
We examined goals of the original study [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], and found that
a system for analysis of data collected by the author should
have the following capacity to be truly helpful:
1. Give an easily perceivable and accurate description of a
desired piece of data, i.e. of a collection of records in the
database, selected by applying certain constraints.
2. Be able to identify differences between distinct subsets of
a whole set of records, which may or may not exhibit
dissimilar qualities, and provide this comparative analysis in
a comprehensive and convenient way.
3. Depict a relation between attributes of a dataset in
sufficiently laconic, but condensed form.
      </p>
      <p>We believe that linguistic summarization methods can
offer a lot to this field of application and can be used
effectively to complement conservative data analytics tools,
such as statistics and visualization. However, a
methodology we have developed is quite universal and can
be applied in any situation, where foregoing functionality is
needed.</p>
    </sec>
    <sec id="sec-3">
      <title>Overview of Existing Approaches</title>
      <p>
        Emergence of Zadeh's fuzzy sets and logic theory made
research towards linguistic description of phenomena
possible, with many researchers turning their attention to
this field [
        <xref ref-type="bibr" rid="ref13 ref4 ref9">4, 9, 13</xref>
        ].
      </p>
      <p>
        One of the commonly used methods of linguistic
summarization is Yager's approach [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Suppose we have:
•  is a quality (attribute) of interest, e.g., age
•  = { 1, … ,   } is a set of objects (records) that manifest
quality  , e.g., a set of employees; hence  (  ) are values
of quality  for object  
• D = { V(y1), . . . , V (y ) } is a set of data (“database”)
      </p>
      <sec id="sec-3-1">
        <title>A summary of a data set consists of:</title>
        <p>• A summarizer  (e.g., young)
• A quantity in agreement  (e.g., most)
• Truth (validity)  – e.g., 0.7
as, e.g.,  (most of employees are young) = 0.7.</p>
        <p>In addition, a set of  's to be described may be restricted
by a set of constraints, called filter. Taking into an account
all of the above, typical summary looks like:</p>
        <p>'s are  ,
e.g., “most ( ) of single ( ) employees are young ( )”.</p>
        <p>
          Truth value of a summary may be determined in several
ways; the basic way is Zadeh's calculus of linguistically
quantified propositions [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]:
where µ ( ),   ( ) and   ( ) are membership functions of
the quantifier, filter and summarizer respectively.
        </p>
        <p>
          This idea was implemented by Kacprzyk and Zadrozny as
an add-on to Microsoft Access, called FQUERY [
          <xref ref-type="bibr" rid="ref5 ref8">5, 8</xref>
          ]. As
fully automatic derivation of linguistic summaries from a
sufficiently large database would be very time consuming,
they proposed an interactive approach to summarization, in
which a user should guess some of the summary's
parameters.
        </p>
        <p>
          In the works mentioned above, as well as in other research
works concerning the application of linguistic summaries to
data mining, e.g., [
          <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
          ], fuzzy summaries are employed to
elicit hidden dependencies in data (fuzzy rules), i.e., in a
fashion of exploratory data analysis.
        </p>
        <p>In our research work, we focus on using fuzzy summaries
for purposes of describing pieces of data, which are
particularly interesting to a user. Firstly, this includes
finding a best fitting linguistic label (which may be a
compound of labels) for a subset of records in a dataset, e.g.,
“most of young employees have middle or high income”.</p>
        <p>In addition, we provide an advanced fuzzy comparison for
identifying distinction between subsets of data. This is
necessary because, while many approaches to linguistic
summarization suggest a good way to assign linguistic
labels to groups of records in the database (defined by
different filters), it appears quite frequently that we have to
directly compare two or more collections of objects'
parameters and identify how significant is their difference.
For instance, the proposed fuzzy summarization mechanism
is able to produce summaries like “Muscle strength
asymmetry coefficient of the main group is dramatically
greater than that of the control group.”</p>
        <p>As a third novelty, we provide a way to depict relations
between variables in a dataset. By a relationship we mean
how a variable of interest behaves in relation to the other,
for example, “Muscle strength of the miners drops a little
from roughly between 725 and 775 to nearly between 675
and 750 with an increase of age from around 40 or less to
roughly between 45 and 50. Then it falls drastically till
approximately between 600 and 625 with an increase of age
till about 55 or greater.”</p>
        <p>Also, it is of interest to note that linguistic summaries are
most frequently used in the area of business (sales database
is a popular example). In this paper we will try to examine
what benefits this methodology can offer when applied to
the field of medical research.</p>
        <p>E-book sales drop sharply from the beginning of the week,
when they are particularly strong, till midweek, when they
become rather weak. However, there is a substantial increase
towards the end of the week, when sales are quite strong.</p>
        <sec id="sec-3-1-1">
          <title>Linguistic</title>
        </sec>
        <sec id="sec-3-1-2">
          <title>Summarization Engine</title>
          <p>- constructs sentence(s) in
natural language using a set
of templates
How e-book sales change throughout
the week?</p>
        </sec>
        <sec id="sec-3-1-3">
          <title>Fuzzy querying interface</title>
          <p>a set of constraints
appropriate linguistic labels for
attributes’ values</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Methodology</title>
      <p>Over the course of our research work, we have developed a
system for derivation of linguistic summaries from
databases with an emphasis on description of relationships
between variables (fuzzy or crisp in nature), or, technically
speaking, values of dataset attributes. In particular, the
system is able to:
• give an accurate linguistic descriptor for an attribute of
subset of data, selected by applying some constraints (fuzzy
or exact), for instance, “strong”, “roughly between 600 and
650”;
• compare groups of objects in a dataset by a chosen
parameter, e.g., “this week’s sales have significantly
exceeded that of the previous week”;
• discover trends in behavior of a variable in relation to the
other one (see example on the Fig. 1).</p>
      <p>Fig. 1 reflects main components of the system and their
interaction. Data model, which is an abstraction of the
database, is responsible for a direct interaction with it.
Additionally, it contains a set of linguistic labels (fuzzy
partition), associated with each attribute, and their definition
(corresponding fuzzy sets). For instance, linguistic variable
sales has the values: weak, average, strong. Data model is
accessed by summarization engine through fuzzy querying
interface, which allows obtaining sufficiently good
linguistic descriptors (labels) for attributes, manifested by
subsets of interest, and difference between these subsets.
Summarization engine then uses these pieces of knowledge
to build a summary in natural language using built-in
sentence structures.</p>
      <p>Now let us describe the most important pieces of
functionality of our system.</p>
    </sec>
    <sec id="sec-5">
      <title>Linguistic Description of a Subset of Data</title>
      <p>One of the common tasks in data analysis is to characterize
values of an attribute of interest of its subgroups. For
instance, a researcher may be interested in values of nerve
conduction velocity among stope miners with extensive
work experience. Although there are standard procedures
for numerical description of data, such as finding mean and
standard deviation, we advocate using linguistic labels and
approximate numbers and intervals, technically defined by
fuzzy sets, as they provide a more human consistent way of
information representation.</p>
      <p>One of the difficulties we faced in the process of
implementation of this feature is concerned with how to
define fuzzy partitions for fuzzy variables so as to be able to
describe and compare different subsets of the data by a
certain parameter. Generally, various groups of objects
(selected using different filters) might have highly diverse
values of an attribute of interest, or, in contrast, highly
narrow variance. On the Fig. 2 you may see a frequency
diagram of muscle strength coefficient of asymmetry of
miners of a main group (miners of specializations with
harmful working conditions, which involve factors that
influence development of cervical osteochondrosis) and a
control group (miners of specializations with less harmful
working conditions). As can be noticed, the main group is
“wider” and one would describe it as “approximately in the
interval between 1.3 and 1.7”, and “roughly between 1.1 and
1.3” for the control group, which is clearly a shorter interval.
It is obvious that it is impossible to find an ideal width of a
fuzzy interval and a number of fuzzy sets in partition
(granularity). Therefore, we either have to define a large
number of fuzzy sets which signify approximate intervals of
a different length, or find a way to build fuzzy sets on
demand, depending on the data itself, but without loosing
their descriptive power, which means they have to be
standard in a certain way.</p>
      <p>
        We decided to follow the second path, and designed an
algorithm for finding a fuzzy set, combined from predefined
fuzzy sets in a partition using disjunction, which fits a subset
of data best, with a degree of truth being higher than some
threshold value (we took 0.7). In this algorithm we used
Zadeh's formula [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] for finding validity of the summary.
Conjunction and disjunction (“and” and “or”) operators in
filter and summarizer parts of the expressions are simple
minimum and maximum operators introduced in [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>Let us demonstrate the results of running this algorithm
on some examples beneath. See Fig. 2 and Fig. 3 for
comparison (fuzzy intervals identified for the example
groups approximately correspond to intervals of high
density on the frequency diagrams).</p>
      <p>Muscle strength asymmetry coefficient of the main group
is approximately between 1.4 and 1.6.</p>
      <p>Muscle strength asymmetry coefficient of the control
group is about 1.3 or less.</p>
      <p>Muscle strength of stope miners is approximately between
675 and 775.</p>
      <p>Muscle strength of timbermen is roughly between 650 and
750.</p>
      <p>Muscle strength of loader operators is nearly between
625 and 750.</p>
      <p>This is important to note that the method described above
can be equally efficiently used for crisp and fuzzy subsets of
data (selected using fuzzy filter, e.g., “middle-aged” for
age).</p>
    </sec>
    <sec id="sec-6">
      <title>Fuzzy Comparison of Subsets</title>
      <p>Another main component of the system is a mechanism of
fuzzy comparison, which is used to identify a degree to
which two groups of objects are different. Comparison of
groups is very useful for many applications. Specifically, in
our application it is necessary for such tasks as: determining
distinction between the main and control group of workers,
as well as between different professions of miners (also with
varying labor conditions) to prove negative impact of certain
workplace factors on the development of cervical
osteochondrosis and overall health, as well as find core
influencing factors; compare medical test results before and
after treatment, and results of several different therapies.</p>
      <p>A standard way to assess difference between two samples
is t-test. However, we did not use it for a number of reasons.
Firstly, numbers are perceived and interpreted by human
worse, than words, and, secondly, there is a problem
designing weighted t-test (for cases when subsets are fuzzy).</p>
      <p>Instead, we invented our own approach for finding a
coefficient of distinction based on linguistic summaries. In
brief, we take the best fit fuzzy sets (let’s say,  1 and  2)
corresponding to two subsets of interest (lets call them  and
 ), and find an average of validity of two summaries: “ is
NOT  2 ” and “ is NOT  1” (“not” is in traditional Zadeh’s
sense). This method also identifies a direction of this
difference (less or greater).</p>
      <p>A linguistic label for a coefficient of distinction is found
as a fuzzy set in the partition (defined subjectively) having
the largest degree of membership of this coefficient.</p>
      <p>Some of the summaries, generated by the program, that
implements fuzzy comparison (refer Fig. 2 and 3 for visual
comparison) can be seen beneath:</p>
    </sec>
    <sec id="sec-7">
      <title>Structure of a Linguistic Summary</title>
      <p>The simple kinds of summaries given in the previous
subsections do not worth much attention per se. However, if
combined into more complex structures, they can become a
really powerful tool, that enables to depict not only single
subgroup or relation between two subgroups' parameters,
but a relationship between variables, i.e. attributes of a
dataset, which can be fuzzy or crisp in nature. For instance,
we are interested in a relationship between age of the
workers and muscle strength, which can be represented by a
couple of sentences in a natural language (see Fig. 4. for
comparison; red dots are means of muscle strength of
workers with age intervals loosely corresponding to fuzzy
sets in the partition):</p>
      <p>Muscle strength of miners drops a little from roughly
between 725 and 775 to nearly between 675 and 750 with
an increase of age from around 40 or less to roughly
between 45 and 50. Then it falls drastically from
approximately between 675 and 750 to approximately
between 600 and 625 with an increase of age from roughly
between 45 and 50 to about 55 or greater.</p>
      <p>of miners   from  1 and  2 with an increase of 
from  1 to  2,
where:
•  and  are an independent and dependent variable
respectively, age and muscle strength according to the
example;</p>
      <p>•  1 and  2 are two linguistic labels associated with 
domain, that are neighboring if to put all labels in ascending
order (approximately 40 or less and between 45 and 50,
between 45 and 50 and 55 or greater in the sentence);
•  1 and  2 are their respective linguistic descriptors, that
were calculated by a method explained earlier in the text
(e.g., roughly between 725 and 775 and nearly between 675
and 750);</p>
      <p>•  is a relation between subsets  1 and  2 (increase,
decrease or no change), exemplified by drop and fall
(synonyms) in the above summary;</p>
      <p>• D is a linguistic label for a coefficient of distinction
(slightly, dramatically, etc.), found using a method
described in the previous subsection.</p>
      <sec id="sec-7-1">
        <title>Adding a Linguistic Flexibility</title>
        <p>In order to make our summaries more realistic and
humangenerated like, we introduced synonyms to the functional
parts of a summary and the linguistic labels (refer to Tab. 1).
So, for instance, to specify that  is significantly greater
than  , system may also take any one of its synonyms, such
as substantially, noticeably.</p>
        <p>Actually, these sentences were generated by a computer
program, that compared subsets of data corresponding to all
linguistic labels associated with an independent variable
(age in the example) by the value of the second parameter
(muscle strength), after which it putted the results of each
comparison into a simple template:</p>
      </sec>
      <sec id="sec-7-2">
        <title>Ingredients</title>
        <sec id="sec-7-2-1">
          <title>Approximate number</title>
        </sec>
      </sec>
      <sec id="sec-7-3">
        <title>Direction of difference</title>
        <sec id="sec-7-3-1">
          <title>a) a downward trend</title>
        </sec>
        <sec id="sec-7-3-2">
          <title>b) an upward trend</title>
        </sec>
      </sec>
      <sec id="sec-7-4">
        <title>Coefficient of distinction</title>
        <sec id="sec-7-4-1">
          <title>b) minor difference</title>
        </sec>
        <sec id="sec-7-4-2">
          <title>c) significant difference</title>
        </sec>
        <sec id="sec-7-4-3">
          <title>d) major difference</title>
          <p>Synonyms
about X, around X, roughly
greater than X,
approximately less than X,
nearly between A and B
decrease, fall, drop, decline
increase, rise, climb, grow
quite significantly, slightly,
a little
significantly, substantially,
noticeably
very significantly,
dramatically, drastically
a) not significant difference almost the same, identical</p>
        </sec>
      </sec>
      <sec id="sec-7-5">
        <title>Enriching Summary Structure</title>
        <p>While the kind of summaries demonstrated above is quite
human consistent, it is possible to introduce a number of
refinements in the future:</p>
        <p>a) Add linking words. Words like but, furthermore,
nevertheless, likewise, again would allow to increase quality
of summaries by highlighting a contrast in trends or their
repetition. Examples:</p>
        <p>Electromyography test results for the groups of workers
with less than 10 years of experience and with 10 -15 years
of experience do not differ significantly. However, that for
the group of workers with more than 20 years of experience
fall sharply from 700-750 to 600-625.</p>
        <p>Nerve conduction velocity falls gradually from the age of
35-40 till the age of 40-45. Then it decreases sharply till the
age of 45-50 and after that it starts to decrease gradually
again.</p>
        <p>b) Creating a structure and organization for sentences,
summarizing distinctions between groups subsetted using a
parameter which is crisp in nature and cannot undergo
qualitative comparison, for instance, profession or a method
of treatment a worker undergone:</p>
        <p>Nerve conduction velocity (NCV) of workers, who have
undergone basic therapy treatment, has significantly
increased due to the therapy. NCV of workers, who have
undergone DENS treatment, has also improved, but to the
lesser degree.</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>Application</title>
      <p>In order to enliven the developed methodology and identify
its advantages and disadvantages, we implemented some of
the aforementioned functionality of the proposed system in
Ruby programming language, together with Ruby on Rails
web framework, that we utilized for easier database
manipulation and user interface design. Let us now
demonstrate our results.</p>
      <p>Example 1. Age and muscle strength coefficient of
asymmetry (refer to Fig. 5 for comparison).</p>
      <p>Muscle strength asymmetry of miners rose quite
significantly from approximately 1.5 or less to roughly
between 1.4 and 1.6 with an increase of age from
approximately 40 or less to approximately between 45 and
50. Then it rose drastically from approximately between 1.4
and 1.6 to roughly 1.6 or greater with an increase of age
from roughly between 45 and 50 to roughly 55 or greater.</p>
      <p>Example 2. Age and muscle strength coefficient of
asymmetry with a finer scale for age (see Fig. 6 for
comparison).</p>
      <p>Muscle strength asymmetry of miners climbed quite
significantly from roughly less than 1.5 to approximately
between 1.4 and 1.6 with an increase of age from
approximately less than 40 to roughly 45. Then it climbed a
little from approximately between 1.4 and 1.6 to roughly
between 1.5 and 1.6 with an increase of age from about 45
to roughly 50. Then it climbed very significantly from
roughly between 1.5 and 1.6 to approximately greater than
1.6 with an increase of age from around 50 to around 55.
Then it was almost the same with an increase of age from
roughly 55 to approximately greater than 55.</p>
      <p>
        Earlier in this paper, we mentioned several use cases of
our system. Let us provide a more comprehensive list of its
potential applications with accordance to the goals of the
research [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Namely, the system can be used to:
1. Compare health parameters of miners of different
specializations with varying labor conditions to determine
the most and the least harmful occupations; as well as to
compare main and control group (workers, which labor
conditions are determined as very heavy and those with not
heavy working conditions).
      </p>
      <p>2. Find relationships between various parameters of
patients, such as age and work experience, and results of
medical tests, so as to prove negative impact of miner's
workplace factors on the development of cervical
osteochondrosis.</p>
      <p>3. Analyze a relationship of potentially unfavorable
workplace factors that influence development of cervical
osteochondrosis, like heavy physical labor, vibration
(coming from drilling machine, loader, etc.), microclimate,
noise, and degradation of functional parameters of the body.</p>
    </sec>
    <sec id="sec-9">
      <title>Concluding remarks</title>
      <p>In this paper we presented a system for derivation of
linguistic summaries from databases, which possesses
functionality of linguistically describing relationships
between attributes of a dataset and distinction in properties
of subsets of data. Main challenges we have faced were
construction of a good descriptor for a subset, deciding on
the metric for significance of difference between subsets,
and designing the summary structure. In our future works
we plan to enrich summary structure with new templates and
components like linking words. Also, we plan to examine a
potential of our methodology for application in areas other
than medical research, as we believe that it is universal and
can be applied in any context where linguistic summaries
can aid in decision support or other purposes, involving data
analysis.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Bortsova</surname>
            ,
            <given-names>S.R.</given-names>
          </string-name>
          <year>2010</year>
          .
          <article-title>Клинико-функциональная оценка компрессионно-корешковых нарушений шейного остеохондроза у горнорабочих и немедикаментозная коррекция</article-title>
          .
          <source>Candidate of Sciences dissertation, Karagandy</source>
          , Kazakhstan.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Fiot</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Laurent</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Teisseire</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Laurent</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <year>2006</year>
          .
          <article-title>Why Fuzzy Sequential Patterns can Help Data Summarization: An Application to the INPI Trade Mark Database</article-title>
          .
          <source>In Proceedings of the IEEE International Conference on Fuzzy Systems</source>
          ,
          <volume>3596</volume>
          -
          <fpage>3603</fpage>
          . Vancouver, Canada.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Laurent</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bouchon-Meunier</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Doucet</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>2001</year>
          .
          <article-title>Towards Fuzzy-OLAP Mining</article-title>
          ,
          <source>In Proceedings of Workshop PKDD 'Database Support for KDD'</source>
          ,
          <volume>51</volume>
          -
          <fpage>62</fpage>
          . Freiburg, Germany.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Liétard</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <year>2012</year>
          .
          <article-title>A functional interpretation of linguistic summaries of data</article-title>
          .
          <source>Information Sciences</source>
          <volume>188</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>16</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Kacprzyk</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yager</surname>
            ,
            <given-names>R.R.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Zadrożny</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <year>2001</year>
          .
          <article-title>Fuzzy Linguistic Summaries of Databases for an Efficient Business Data Analysis and Decision Support. In Knowledge discovery for business information systems</article-title>
          ,
          <volume>129</volume>
          -
          <fpage>152</fpage>
          . Boston, Mass.: Kluwer Academic Publishers.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Kacprzyk</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yager</surname>
            ,
            <given-names>R.R.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Zadrożny</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <year>2000</year>
          .
          <article-title>A Fuzzy Logic Based Approach to Linguistic Summaries of Databases</article-title>
          .
          <source>International Journal of Applied Mathematics and Computer Science</source>
          <volume>10</volume>
          :
          <fpage>813</fpage>
          -
          <lpage>834</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Kacprzyk</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Yager</surname>
            ,
            <given-names>R.R.</given-names>
          </string-name>
          <year>2001</year>
          .
          <article-title>Linguistic Summaries of Data Using Fuzzy Logic</article-title>
          .
          <source>International Journal of General Systems</source>
          <volume>30</volume>
          :
          <fpage>133</fpage>
          -
          <lpage>154</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Kacprzyk</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Zadrożny</surname>
            <given-names>S.</given-names>
          </string-name>
          <year>2001</year>
          .
          <article-title>SQL and FQUERY for Access</article-title>
          .
          <source>In Proceedings of IFSA/NAFIPS</source>
          ,
          <fpage>2464</fpage>
          -
          <lpage>2469</lpage>
          . Vancouver, Canada: IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Pei</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ruan</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Qin</surname>
          </string-name>
          , K..
          <year>2009</year>
          .
          <article-title>Extracting complex linguistic data summaries from personnel database via simple linguistic aggregations</article-title>
          .
          <source>Information Sciences</source>
          <volume>179</volume>
          (
          <issue>14</issue>
          ):
          <fpage>2325</fpage>
          -
          <lpage>2332</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Yager</surname>
            ,
            <given-names>R.R.</given-names>
          </string-name>
          <year>1982</year>
          .
          <article-title>A new approach to the summarization of data</article-title>
          .
          <source>Information Sciences</source>
          <volume>28</volume>
          :
          <fpage>69</fpage>
          -
          <lpage>86</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Zadeh</surname>
            ,
            <given-names>L.A.</given-names>
          </string-name>
          <year>1983</year>
          .
          <article-title>A computational approach to fuzzy quantifiers in natural languages</article-title>
          .
          <source>Computer Mathematics with Applications</source>
          <volume>9</volume>
          :
          <fpage>149</fpage>
          -
          <lpage>183</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Zadeh</surname>
            ,
            <given-names>L.A..</given-names>
          </string-name>
          <year>1965</year>
          .
          <article-title>Fuzzy sets</article-title>
          .
          <source>Information and Control</source>
          <volume>8</volume>
          :
          <fpage>338</fpage>
          -
          <lpage>353</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Zadrożny</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Kacprzyk</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <year>2011</year>
          .
          <article-title>From a static to dynamic analysis of weblogs via linguistic summaries</article-title>
          .
          <source>In Proceedings of 2011 IFSA World Congress AFSS International Conference</source>
          ,
          <volume>110</volume>
          -
          <fpage>119</fpage>
          . Surabaya and Bali, Indonesia.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>