<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>On Concept Lattices as Information Channels</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Francisco J. Valverde-Albacete</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Carmen Pelaez-Moreno</string-name>
          <email>carmen@tsc.uc3m.es</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anselmo Pen~as</string-name>
          <email>anselmog@lsi.uned.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Departamento de Lenguajes y Sistemas Informaticos Universidad Nacional de Educacion a Distancia, c/ Juan del Rosal</institution>
          ,
          <addr-line>16. 28040 Madrid</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Departamento de Teor a de la Sen~al y de las Comunicaciones Universidad Carlos III de Madrid</institution>
          ,
          <addr-line>28911 Leganes</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper explores the idea that a concept lattice is an information channel between objects and attributes. For this purpose we study the behaviour of incidences in L-formal contexts where L is the range of an information-theoretic entropy function. Examples of such data abound in machine learning and data mining, e.g. confusion matrices of multi-class classi ers or document-term matrices. We use a wellmotivated information-theoretic heuristic, the maximization of mutual information, that in our conclusions provides a avour of feature selection providing and information-theory explanation of an established practice in Data Mining, Natural Language Processing and Information Retrieval applications, viz. stop-wording and frequency thresholding. We also introduce a post-clustering class identi cation in the presence of confusions and a avour of term selection for a multi-label document classi cation task.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Information Theory (IT) was born as a theory to improve the e ciency of
(manmade) communication channels [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ], but it soon found wider application [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
This paper is about using the model of a communication channel in IT to explore
the formal contexts and concept lattices of Formal Concept Analysis as
realisations of information channels between objects and attributes. Given the highly
unspeci ed nature of both the latter abstractions such a model will bring new
insights into a number of problems, but we are speci cally aiming at machine
learning and data mining applications [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ].
      </p>
      <p>
        The metaphor of a concept lattice as a communication channel between
objects and attributes is already implicit in [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ]. In there, adjoint sublattices were
already considered as subchannels in charge of transmitting individual
acoustical features, and some e orts were done to model such features explicitly [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ],
? FJVA and AP are supported by EU FP7 project LiMoSINe (contract 288024) for
this work. CPM has been supported by the Spanish Government-Comision
Interministerial de Ciencia y Tecnolog a project TEC2011-26807.
but no conclusive results were achieved. The di culty rose from a thresholding
parameter ' that controls the lattice-inducing technique and was originally xed
by interactive exploration, a procedure hard to relate to the optimization of a
utility or cost function, as required in modern machine learning.
      </p>
      <p>
        In this paper we set this problem against the backdrop of direct mutual
information maximization|using techniques and insights developed since [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ]|for
matrices whose entries are frequency counts. These counts appear frequently
in statistics, data mining and machine learning, for instance, in the form of
document-term matrices in Information Retrieval [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], confusion matrices for
classi ers in perceptual studies, data mining and machine learning [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], or simply
two-mode contingency tables with count entries. Such matrices are called
aggregable in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], in the sense that any group of rows or columns can be aggregated
together to form another matrix whose frequencies are obtained from the data
of the elements in the groups. We will use this feature to easily build count and
probability distributions whose mutual information can be maximized, following
the heuristic motivated above, to improve classi cation tasks. Note that
maximizing mutual information (over all possible joint distributions) is intimately
related to the concept of channel capacity as de ned by Shannon [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>For this purpose, in Sec. 2 we cast the problem of analysing the transfer of
information through the two modes of contingency tables as that of analysing a
particular type of formal context. First we present in Sec. 2.1 the model of the
task to be solved, then we present aggregable data, as usually found in machine
learning applications in Sec. 2.2, and then introduce the entropic encoding to
make it amenable to FCA. As an application, in Sec. 3.1 we explore the particular
problem of supervised clustering as that of transferring the labels from a set of
input patterns to the labels of the output classes. Speci cally we address the
problem of assigning labels to mixed clusters given the distribution of the input
labels in them. We end with a discussion and a summary of contributions and
conclusions.
2
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>Theory</title>
      <sec id="sec-2-1">
        <title>Classi cation optimization by mutual information maximization</title>
        <p>Consider the following, standard supervised classi cation setting: we have two
domains X and Y , m instances of i.i.d. samples S = f(xi; yi)gim=1 X Y , and
we want to learn a function h : X ! Y , the hypothesis, with certain \good"
qualities, to estimate the class Y from X , the measurements of Y , or features.</p>
        <p>A very productive model to solve this problem is to consider two probability
spaces Y = hY; PY i and X = hX; PX i with Y PY and X PX , and suppose
that there exists the product space hX Y; PXY i wherefrom the i.i.d. samples
of S have been obtained. So our problem is solved by estimating the random
variable Y^ = h(X), and a \good" estimation is that which obtains a low error
probability on every possible pair P (Y^ 6= Y ) ! 0 .</p>
        <p>Since working with probabilities might be di cult, we might prefer to use
a (surrogate) loss function that quanti es the cost of this di erence L(y^ =
h(x); y) and try to minimize the expectation of this loss, called the risk R(h) =
E[L(h(x); y)] over a class of functions h 2 H, h = minh2H R(h) . Consequently,
this process is called empirical risk minimization.</p>
        <p>
          An alternate criterion is to maximize the mutual information between Y and
Y^ [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. This is clearly seen from Fano's inequality [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], serving as a lower bound,
and the Hellman-Raviv upper bound [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ],
        </p>
        <p>HPY^</p>
        <p>IPY Y^
HU ^</p>
        <p>Y
1</p>
        <p>P (Y^ 6= Y )
1
2 HPY^ jY
where UY^ is the uniform distribution on the support of Y^ , HPXX denotes the
di erent entropies involved and IPY Y^ is the mutual information of the joint
probability distribution.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Processing aggregable data</title>
        <p>If the original rows and columns of contingency tables represent atomic events,
their groupings represent complex events and this structure is compatible with
the underlying sigma algebras that would transform the matrix into a joint
distribution of probabilities, hence these data can be also interpreted as joint
probabilities, when row- and column-normalized.</p>
        <p>When insu cient data is available for counting, the estimation of empirical
probabilities from this kind of data is problematic, and complex probability
estimation schemes have to be used. Even if data galore were available, we still
have to deal with the problem of rarely seen events and their di cult probability
estimation. However, probabilities are, perhaps, the best data that we can plug
onto data mining or machine learning techniques, be they for supervised or
unsupervised tasks.</p>
      </sec>
      <sec id="sec-2-3">
        <title>The weighted Pointwise Mutual Information. Recall the formula for</title>
        <p>the mutual information between two random variables IPXY = EPXY [IXY (x; y)]</p>
        <p>PXY (x;y)
where IXY (x; y) = log PX (x) PY (y) is the pointwise mutual information, (PMI).</p>
        <p>Remember that 1 IXY (x; y) &lt; 1 with IXY (x; y) = 0 being the case
where X and Y are independent. The negative values are caused by phenomena
less represented in the joint data than in independent pairs as captured by the
marginals. The extreme value IXY (x; y) = 1 is generated when the joint
probability is negative even if the marginals are not. These are instances that
capture \negative" association whence to maximize the expectation we might
consider disposing of them.</p>
        <p>On the other hand, on count data the PMI has an unexpected and unwanted
e ect: it is very high for hapax legomena phenomena that are encountered only
once in a tallying, and in general it has a high value for phenomena with low
counts of whose statistical behaviour we are less certain.</p>
        <p>However, we know that</p>
        <p>IPXY =</p>
        <p>X PXY (x; y) IXY (x; y) =
x;y</p>
        <p>X PXY (x; y) log
x;y</p>
        <p>PXY (x; y)
PX (x) PY (y)
and this is always a positive quantity, regardless of the individual values of
IXY (x; y). This suggests calling weighted pointwise mutual information, (wPMI)
the quantity
wPMI(x; y) = PXY (x; y) log
(1)</p>
        <p>PXY (x; y)
PX (x) PY (y)
and using it as the subject of optimization or exploration to do so. Note that
pairs of phenomena whose joint probability are close to independent, as judged
by the pointwise information, will be given a very low value in the wP M I , and
that the deleterious character of hapaxes on IPXY is lessened by the in uence of
the joint probability.
2.3</p>
      </sec>
      <sec id="sec-2-4">
        <title>Visualizing mutual information maximization</title>
        <p>
          For a joint distribution PY Y^ (y; y^), [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] introduced a balance equation binding the
mutual information between two variables IPY Y^ , the sum of their conditional
entropies V IPY Y^ = HPY jY^ + HPY^ jY and the sum of their entropic distance between
their distributions and uniformity
        </p>
        <p>HPY Y^ = (HUY</p>
        <p>HPY ) + (HUY^</p>
        <p>HPY^ ),
log(HUY ) + log(HUY^ ) =</p>
        <p>HPY Y^ + 2 IPY Y^ + V IPY Y^ :
By normalizing in the total entropy log(HUY ) + log(HUY^ ) we may obtain the
equation of the 2-simplex that can be represented as a De Finetti diagram like
that of Fig. 2.(a), as the point in the 2-simplex corresponding to coordinates</p>
        <p>F (PY Y^ ) = [ HP0Y Y^ ; 2 IP0Y Y^ ; V IP0Y Y^ ]
where the primes represent the normalization described above.</p>
        <p>
          The axis of this representation were chosen so that the height of the
2simples|an equilateral triangle|is proportional to the mutual information
between the variables so a maximization process is extremely easy to represent (as
in Fig. 2): given a parameter ' whereby to maximize IPY Y^ (as a variable), draw
the trace of the evaluation of the coordinates in the ET of the distributions that
it generates, and choose the ' that produces the highest point in the triangle.
This technique is used in Sec. 3.1, but other intuitions can be gained from this
representation as described in [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ].
2.4
        </p>
      </sec>
      <sec id="sec-2-5">
        <title>Exploring the space of joint distributions</title>
        <p>
          Since the space of count distributions is so vast, we need a technique to explore
it in a principled way. For that purpose we use K-Formal Concept Analysis
(KFCA). This is a technique to explore L-valued contexts where L is a complete
idempotent semi eld using a free parameter called the threshold of existence [
          <xref ref-type="bibr" rid="ref13 ref15">15,
13</xref>
          ].
        </p>
        <p>We proceed in a similar manner to Fuzzy FCA: For L-context hY; Y^; Ri,
con^
sider two spaces LY and LY , representing, respectively, L-valued sets of objects
and attributes. Pairs of such sets of objects and attributes that ful l certain
polars equation have been proven to de ne dually-ordered lattices of closed L-sets
in the manner of FCA 3.</p>
        <p>Since the actual lattices of object sets and attributes are so vast, KFCA
uses a simpli ed representation for them: for the singleton sets in each of the
spaces y, for y 2 Y and y^, for y^ 2 'Y^ , we use the L-polars to generate their
object- Y' (y) and attribute-concept Y^ (y^), respectively, and obtain a structural
'-context K' = hY; Y^; R'i, where yR'y^ () Y' (y) Y'^ (y^) 4.</p>
        <p>In this particular case we consider the min-plus idempotent semi eld and
the L-context hY; Y^; wP M Ii where wPMI is the weighted Pointwise Mutual
Information relation between atomic events in the sigma lattices of Y and Y^ of
Sec. 2.2, whence the degree or threshold of existence is a certain amount of
entropy required for concepts to surpass for them to be considered.</p>
        <p>The following step amounts to an entropy conformation of the joint
distribution, that is, a redistribution of the probability masses in the joint distribution
to obtain certain entropic properties. Speci cally, we use the (binary) '-formal
context to lter out certain counts in the contingency table to obtain a
conformal contingency table NY'Y^ (y; y^) = NY Y^ (y; y^) K', where represents here the
Hadamard (pointwise) product. For each conformal NY'Y^ (y; y^) we will obtain a
certain point F (') in the ET to be represented as described in Sec. 2.3.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Application</title>
      <p>We next present two envisaged applications of the technique of MI Maximization.
3.1</p>
      <sec id="sec-3-1">
        <title>Cluster identi cation</title>
        <p>
          Confusion matrices are special contingency tables whose two modes refer to the
same underlying set of labels[
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. We now put forward a procedure to maximize
the information transmitted from a set of \ground truth" patterns acting as
objects with respect to \perceived patterns" which act as attributes. As noted
in the introduction, this is just one of the possible points of view about this
problem.
        </p>
        <p>
          Consider the following scenario, there is a clustering task for which extrinsic
evaluation is possible, that is, there is a gold standard partitioning of the
input data. One way to evaluate the clustering solution is to obtain a confusion
3 Refer to [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] for an in-depth discussion of the mathematics of idempotent semi elds
and the di erent kinds of Galois connections that they generate.
4 And a structural '-lattice B'(K') as its concept lattice, but this is not important
in the present application
matrix out of this gold standard, in the following way: If the number of classes
is known|a realistic assumption in the presence of a gold standard|then the
MI optimization procedure can be used to obtain the assignments between the
classes in the gold standard and the clusters of the procedure, resulting in cluster
identi cation.
        </p>
        <p>
          For the purpose of testing the procedure, we used the segmented numeral
data from [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. This is a task of human visual confusions between numbers as
displayed by seven-segment LED displays, as shown in Fig. 1.(a). The entry in
the count matrix NCK (c; k) = nck counts the number times that an instance of
class c was confused with class k . Figure 1.(b) shows a heatmap presentation
of the original confusion matrix and column-reshu ed variants. Note that the
confusion matrix is diagonally-dominant, that is nii &gt; Pj;j6=i nij and likewise
for column i .
        </p>
        <p>(a)
(b)</p>
        <p>To test the MI optimization procedure, we randomly permuted the
confusion matrix columns: the objective was to recover the inverse of this random
permutation from the MI optimization process so that the original order could
be restored. This amounts to an assignment between classes and induced
clusters, and we claim that it can be done by means of the mutual information
maximization procedure sketched above.</p>
        <p>For that purpose, we estimated PCK (c; k) using the empirical estimate
P^CK (c; k)</p>
        <p>NCK (c; k)
n
where n is the number of instances to be clustered n = Pck NCK (c; k) , and
then we obtained its empirical PMI</p>
        <p>I^CK (c; k) = log P^CK (c; k)
and its weighted PMI</p>
        <p>wP M ICK (c; k) = P^CK (c; k) I^CK (c; k) :</p>
        <p>Next, we used the procedure of Sec. 2.4 to explore the empirical wPMI and
select the threshold value which maximizes the MI. Figure 2.(a) shows the
trajectory of the di erent conformed confusion matrices as ' ranges in [0; 1) on
the ET: we clearly see how for this balanced task dataset the exploration results
in a monotonous increase in MI in the thresholding range until a value that
produces the maximum MI, at wP M I = 0:1366 . The discrete set of points stems
from the limited range of counts in the data.</p>
        <p>We chose this value as threshold and obtained the binary matrix which is
the assignment from classes to clusters and vice-versa shown in Fig. 2.(b). Note
that in this particular instance, the ' -concept lattice is just a diamond
lattice re ecting the perfect identi cation of classes and clusters. In general, with
contingency tables where modes have di erent cardinalities, this will not be the
case.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Entropy conformation for count matrices</title>
        <p>The case where the contingency matrix is squared and diagonally dominant, as
in the previous example, is too speci c: we need to show that for a generic,
rectangular count contingency matrix, entropy maximization is feasible and
meaningful.</p>
        <p>
          The rst investigation should be on how to carry the maximization
process. For that purpose, we use a modi ed version of the Reuters-21578 5 that
has already been stop-listed and stemmed. This is a multi-label classi cation
dataset [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] describing each document as a bag-of-terms and some
categorizations labels, the latter unused in our present discussion.
        </p>
        <p>We considered the document-term matrix for training, a count distribution
with D = 7 770 documents and T = 5 180, terms. Its non-conformed entropy
coordinates are F (NDT ) = [0:1070; 0:3584; 0:5346] as shown in the deep blue circle
to the left of Fig. 3. We carried out a joint-mutual information maximization
process by exploring at the same time a max-plus threshold|the count has to
be bigger thant the threshold to be considered|and a min-plus threshold|the
count has to be less than the threshold. The rationale for this is a well-tested
hypothesis in the bag-of-term model: very common terms (high frequency) do not
select well for documents, while very scarce terms (low frequency) are too
speci c and biased to denote the general \aboutness" of a document. Both should
be ltered out of the document-term matrix.
5 http://www.daviddlewis.com/resources/testcollections/reuters21578/
readme.txt. Visited 24/06/2014.
Fig. 2: Trajectory of the evolution of MI transmission for the segmented numeral
data as the exploration threshold is raised in the wPMI matrix (a), and maximal
MI cluster assignment matrix at wPMI = 1:366 bits (b) for column-shu ed
Segmented Numerals. The resulting concept lattice is just a diamond lattice
identifying classes and clusters and not shown.</p>
        <p>Instead of count-based individual term ltering we carry a joint term-document
pair selection process: for a document-matrix, we calculate its overall weighted
PMI matrix, and only those pairs (d; t) whose wPMI lies in between a lower and
an upper ' thresholds are considered important for later processing. For each
such pairs, we created an indicator matrix I(d; t) that is 1 i wM I(d; t) ',
and we used the Kronecker multiplication to lter out non-conforming pairs from
the nal entropy calculation,</p>
        <p>M^ IP0DT =</p>
        <p>X wP M IDT (d; t) I(d; t)
d;t</p>
        <p>We can see how M^ IP0DT reaches a maximum over two values and then
decreases again, going even below the original mutual information value. We read
two di erent facts in this illustration: that the grid used is e ective in obtaining
approximations to and ' for MI maximization, and that not every possible
pair of values is a good solution for the process.</p>
        <p>All in all, this procedure shows that MI maximization is feasible by tracking
its in the ET. We do not present any results in this paper as to the e ectiveness
of the process for further processing tasks, which should be evaluated on the
extrinsic measures on the Reuters multi-labelling task.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Discussion</title>
      <p>
        We now discuss the applications selected in a wider context. Although less
pervasive than its unsupervised version, the basic task of supervised clustering has
application, for instance, in tree-induction for supervised classi cation [
        <xref ref-type="bibr" rid="ref18 ref5">5, 18</xref>
        ] or
unsupervised clustering evaluation using a gold-set [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. Cluster identi cation
in Sec. 3.1 is a sometimes-fussy sub-procedure in clustering which our proposal
solves elegantly.
      </p>
      <p>The feasibility study on mutual information conformation of Sec. 3.2 is a
necessary step for further processing|binary or multi-labelling classi cation|
but as of this paper unevaluated. Further work should concentrate on leveraging
the boost in mutual information to lower the classi cation error, as suggested in
the theoretical sections.</p>
      <p>Besides, the use of two simultaneous, thresholds on di erent algebras makes
it di cult to justify the procedure on FCA terms: this does not conform to the
de nition of any lattice-inducing polars that we know of, so this feature should
be looked into critically. Despite this fact, the procedure of conformation \makes
sense", at least for this textual classi cation task.</p>
      <p>Note that the concept of \information channel" that we have developed in
this paper is not what Communication Theory usually considers. In there, \input
symbols" enter the channel and come out as \output symbols", hence input has
a sort of ontological primacy over output symbols in that the former cause the
latter. If there is anything particular about FCA as an epistemological theory is
that it does not prejudge the ontological primacy of objects over attributes or vice
versa. Perhaps the better notion is that a formal concept is an information
cochannel between objects and attributes, in the sense that the information \ ows"
both from objects to attributes and vice versa, as per the true symmetric nature
of mutual information: receiving information about one of the modes decreases
the uncertainty of the other.</p>
      <p>
        The previous paragraph notwithstanding, we will often nd ourselves in
application scenarios in which one of the modes will be primary with respect to
the other, in which case the analogies with communication models will be more
evident. This is one of the cases that we explore in this paper, and that rst
pointed at in [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ].
      </p>
      <p>Contingency tables are an instance of aggregable data tables [4, x0.3.4]. It
seems clear that not just counts, but any non-negative entry aggregable table can
be treated with the tools here presented, e.g. concentrations of solutes. In that
case, the neat interpretation related to MI maximization will not be available,
but analogue ones can be found.</p>
      <p>
        A tangential approach to the de nition of entropies in (non-Boolean) lattices
has been taken by [
        <xref ref-type="bibr" rid="ref20 ref21 ref22 ref23 ref24">20, 21, 22, 23, 24</xref>
        ]. These works approach the de nition of
measures, and in particular entropy measures, in general lattices instead of nite
sigma algebras (that is, Boolean lattices). [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] and [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] speci cally address the
issue of de ning them in concept lattices, but the rest provide other heuristic
foundations for the de nition of such measures which surely must do without
some of the more familiar properties of the Shannon (probability-based) entropy.
5
      </p>
    </sec>
    <sec id="sec-5">
      <title>Conclusions and further work</title>
      <p>We have presented an incipient model of L-formal contexts of aggregable data
and their related concept lattices as information channels. Using KFCA as the
exploration technique and the Entropy Triangle as the representation and
visualization technique we can follow the maximization procedure on confusion
matrices in general, and in confusion matrices for cluster identi cation in
particular.</p>
      <p>We present both the basic theory and two proof-of-concept applications in
this respect: a rst one cluster identi cation, fully interpretable in the framework
of concept lattices, and another, entropy conformation for rectangular matrices
more di cultly embeddable in this framework.</p>
      <p>Future applications will extend the analysis of count contingency tables, like
document-term matrices, where our entropy-conformation can be likened to
feature selection techniques.</p>
      <p>Bibliography</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Shannon</surname>
            ,
            <given-names>C.E.</given-names>
          </string-name>
          :
          <source>A mathematical theory of Communication</source>
          .
          <source>The Bell System Technical Journal XXVII</source>
          (
          <year>1948</year>
          )
          <volume>379</volume>
          {
          <fpage>423</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Shannon</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weaver</surname>
            ,
            <given-names>W.:</given-names>
          </string-name>
          <article-title>A mathematical model of communication</article-title>
          . The University of Illinois Press (
          <year>1949</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Brillouin</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          : Science and
          <string-name>
            <given-names>Information</given-names>
            <surname>Theory. Second Edition. Courier Dover Publications</surname>
          </string-name>
          (
          <year>1962</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Mirkin</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Mathematical Classi cation and Clustering</article-title>
          . Volume
          <volume>11</volume>
          of Nonconvex Optimization and
          <string-name>
            <given-names>Its</given-names>
            <surname>Applications</surname>
          </string-name>
          . Kluwer Academic Publishers (
          <year>1996</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Mirkin</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Core Concepts in Data Analysis: Summarization, Correlation and Visualization</article-title>
          . Summarization, Correlation and Visualization. Springer, London (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Pelaez-Moreno</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garc</surname>
            a-Moral,
            <given-names>A.I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Valverde-Albacete</surname>
            ,
            <given-names>F.J.:</given-names>
          </string-name>
          <article-title>Analyzing phonetic confusions using Formal Concept Analysis</article-title>
          .
          <source>Journal of the Acoustical Society of America</source>
          <volume>128</volume>
          (
          <year>2010</year>
          )
          <volume>1377</volume>
          {
          <fpage>1390</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Pelaez-Moreno</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Valverde-Albacete</surname>
            ,
            <given-names>F.J.:</given-names>
          </string-name>
          <article-title>Detecting features from confusion matrices using generalized formal concept analysis</article-title>
          . In Corchado, E.,
          <string-name>
            <surname>Grana-Romay</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Savio</surname>
          </string-name>
          , A.M., eds.:
          <source>Hybrid Arti cial Intelligence Systems. 5th International Conference, HAIS</source>
          <year>2010</year>
          , San Sebastian, Spain, June 23-25,
          <year>2010</year>
          . Proceedings,
          <string-name>
            <surname>Part</surname>
            <given-names>II</given-names>
          </string-name>
          . Volume
          <volume>6077</volume>
          of LNAI., Springer (
          <year>2010</year>
          )
          <volume>375</volume>
          {
          <fpage>382</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Manning</surname>
            ,
            <given-names>C.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Raghavan</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , Schutze, H.: Introduction to Information Retrieval. Cambridge University Press (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Japkowicz</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shah</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Evaluating Learning Algorithms: A Classi cation Perspective</article-title>
          . Cambridge University Press (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Frenay</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Doquire</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Verleysen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Theoretical and empirical study on the potential inadequacy of mutual information for feature selection in classi cation</article-title>
          .
          <source>NEUROCOMPUTING</source>
          <volume>112</volume>
          (
          <year>2013</year>
          )
          <volume>64</volume>
          {
          <fpage>78</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>M</given-names>
            <surname>Fano</surname>
          </string-name>
          ,
          <string-name>
            <surname>R.</surname>
          </string-name>
          :
          <source>Transmission of Information: A Statistical Theory of Communication</source>
          . The MIT Press (
          <year>1961</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Feder</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Merhav</surname>
          </string-name>
          , N.:
          <article-title>Relations between entropy and error probability</article-title>
          .
          <source>IEEE Transactions on Information Theory</source>
          <volume>40</volume>
          (
          <year>1994</year>
          )
          <volume>259</volume>
          {
          <fpage>266</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Valverde-Albacete</surname>
            ,
            <given-names>F.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pelaez-Moreno</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Two information-theoretic tools to assess the performance of multi-class classi ers</article-title>
          .
          <source>Pattern Recognition Letters</source>
          <volume>31</volume>
          (
          <year>2010</year>
          )
          <volume>1665</volume>
          {
          <fpage>1671</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Valverde-Albacete</surname>
            ,
            <given-names>F.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pelaez-Moreno</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <volume>100</volume>
          %
          <article-title>classi cation accuracy considered harmful: the normalized information transfer factor explains the accuracy paradox</article-title>
          .
          <source>PLOS ONE</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Valverde-Albacete</surname>
            ,
            <given-names>F.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pelaez-Moreno</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Galois connections between semimodules and applications in data mining</article-title>
          . In Kusnetzov, S.,
          <string-name>
            <surname>Schmidt</surname>
          </string-name>
          , S., eds.:
          <article-title>Formal Concept Analysis</article-title>
          .
          <source>Proceedings of the 5th International Conference on Formal Concept Analysis, ICFCA</source>
          <year>2007</year>
          ,
          <article-title>Clermont-Ferrand, France</article-title>
          . Volume
          <volume>4390</volume>
          of LNAI., Springer (
          <year>2007</year>
          )
          <volume>181</volume>
          {
          <fpage>196</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Keren</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baggen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Recognition models of alphanumeric characters</article-title>
          .
          <source>PERCEPT PSYCHOPHYS 29</source>
          (
          <year>1981</year>
          )
          <volume>234</volume>
          {
          <fpage>246</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Tsoumakas</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <article-title>Katakis, I.: Multi-label classi cation: An overview</article-title>
          .
          <source>International Journal of Data Warehousing and . .</source>
          . (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Hall</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Frank</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Holmes</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pfahringer</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reutemann</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Witten</surname>
            ,
            <given-names>I.H.</given-names>
          </string-name>
          :
          <article-title>The WEKA data mining software: An update</article-title>
          .
          <source>SIGKDD Explorations</source>
          <volume>11</volume>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Meila</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Comparing clusterings|an information based distance</article-title>
          .
          <source>Journal of Multivariate Analysis</source>
          <volume>28</volume>
          (
          <year>2007</year>
          )
          <volume>875</volume>
          {
          <fpage>893</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Knuth</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Valuations on Lattices and their Application to Information Theory</article-title>
          .
          <source>Fuzzy Systems</source>
          , IEEE International Conference on (
          <year>2006</year>
          )
          <volume>217</volume>
          {
          <fpage>224</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Grabisch</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Belief functions on lattices</article-title>
          .
          <source>International Journal Of Intelligent Systems</source>
          <volume>24</volume>
          (
          <year>2009</year>
          )
          <volume>76</volume>
          {
          <fpage>95</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Kwuida</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schmidt</surname>
            ,
            <given-names>S.E.</given-names>
          </string-name>
          :
          <article-title>Valuations and closure operators on nite lattices</article-title>
          .
          <source>Discrete Applied Mathematics</source>
          <volume>159</volume>
          (
          <year>2011</year>
          )
          <volume>990</volume>
          {
          <fpage>1001</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Simovici</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          : Entropies on Bounded Lattices.
          <string-name>
            <surname>Multiple-Valued Logic</surname>
          </string-name>
          (ISMVL),
          <year>2011</year>
          41st IEEE International Symposium on (
          <year>2011</year>
          )
          <volume>307</volume>
          {
          <fpage>312</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Simovici</surname>
            ,
            <given-names>D.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fomenky</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kunz</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          :
          <article-title>Polarities, axiallities and marketability of items</article-title>
          .
          <source>In: Proceedings of Data Warehousing and Knowledge Discovery - DaWaK</source>
          . Volume
          <volume>7448</volume>
          <source>of LNCS</source>
          . Springer (
          <year>2012</year>
          )
          <volume>243</volume>
          {
          <fpage>252</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>