<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Representing Informational Harmoniums using Semifield-Valued Formal Concept Analysis</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Francisco J. Valverde-Albacete</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Carmen Pel´aez-Moreno</string-name>
          <email>carmen@tsc.uc3m.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Depto. Teor ́ıa de Sen ̃al y Comunicaciones, Univ. Carlos III de Madrid</institution>
          ,
          <addr-line>Madrid</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <fpage>159</fpage>
      <lpage>170</lpage>
      <abstract>
        <p>In this paper we apply semifield-based Formal Concept Analysis (FCA) to the analysis of Information Harmoniums. These are unsupervised graphical models that, similarly to energy-based models, are written in terms of information functions, but whose expression is linear in an information semifield. In particular we concentrate in the case of the limit information semifields that are the completed max-plus and min-plus semifields for which strong representation theorems in relation to Galois connections and semifield-valued FCA have been proven. We center our contribution in the analysis of the representation spaces for visible and hidden nodes of Information Harmoniums and the lattices related to them.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Restricted Boltzmann Machines (RBM) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] are one of the basic techniques that
revolutionized Artificial Neural Networks some 10 years ago transforming them
into Deep Neural Networks [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        RMBs are, in fact, a type of harmonium [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] at the intersection of Boltzmann
Machines [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]—a kind of Energy-Based Model [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]—and a Product-of-experts
(PoE) [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. PoE are readily trained by means of Contrastive Divergence, a better
approach for Maximum Likelihood Estimation (MLE) than Gibbs sampling [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ],
which explains their efficiency.
      </p>
      <p>In this paper we put in evidence some discrepancies in the definition of
harmoniums (Section 2.1) that suggest that a more natural point of view is to
consider their energy functions as based in an information semifield (Section
2.2). If this is the case, a particular instance of the information semifields are
the Rmax,+ and Rmin,+ semifields over which a multi-valued generalization of
FCA can be defined, Rmax,+-FCA (Section 2.3). In Section 3 we present our
results and contend that FCA in general, and Rmax,+-FCA in particular, provides
a framework for the visualization and understanding of information harmoniums
that could also provide clues for other types of harmoniums. Finally we provide
some conclusions.
? Corresponding author.</p>
    </sec>
    <sec id="sec-2">
      <title>Theory and Methods</title>
      <sec id="sec-2-1">
        <title>RBM and Harmoniums</title>
        <p>The basic model. Technically speaking a harmonium is an undirected
graphical model for the generation of a joint probability distribution. Their graphical
model can be seen in Fig. 1.</p>
        <p>state
vectors
~
h
~v</p>
        <sec id="sec-2-1-1">
          <title>K hidden variables</title>
        </sec>
        <sec id="sec-2-1-2">
          <title>L observed variables</title>
          <p>~c
~
b</p>
        </sec>
        <sec id="sec-2-1-3">
          <title>W parameters</title>
          <p>One type of nodes tagged with vl labels are the set of L input or observed
nodes. The other type is the set of K output or hidden nodes, tagged with hk.
Conventionally we will use indices l ∈ [1 . . . L] and k ∈ [1 . . . K] to go over them.</p>
          <p>Harmoniums are generative models: consider random vectors V = {Vi | i ∈ I}
and H = {Hj | j ∈ J } with component random variables Vi and Hj associated
to the visible and hidden nodes, respectively. We posit that for a particular pair
of random vector (values) (~v, ~h) ∈ V × H the joint probability distribution pV H
of Figure 1 takes the form of a Boltzmann distribution,
pV H (~v, ~h) =</p>
          <p>e−βE(~v,~h)
P
~v∈V</p>
          <p>P~h∈H e−βE(~v,~h)
where β ∈ (0, ∞] is a formal parameter, the coldness, and E(~v, ~h) is a provided
energy function, given in terms of bias vectors ~b ∈ V , ~c ∈ H and the weight
matrix W :</p>
          <p>E(~v, ~h) = ~vt~c + ~bt~h + ~vtW~h</p>
          <p>
            Several points are worth making here:
– The coldness parameter β is in the context of Physics often written as its
inverse, the temperature T = 1/β respecting the original form of the
Boltzmann’s function. It was not originally introduced in the Harmonium [
            <xref ref-type="bibr" rid="ref3">3</xref>
            ],
but was always considered as part of the training procedure of Boltzmann
machines as a relaxation parameter [
            <xref ref-type="bibr" rid="ref4">4</xref>
            ].
– It is easy to see, e.g. [
            <xref ref-type="bibr" rid="ref8">8</xref>
            ], that the energy function can be interpreted as a
scalar product in the standard field of reals,
1
E(~v, ~h) = h~v0|W 0|~h0i =4 1 ~vt · W 0 · ~h
t
a ~b
~c W
where ~v0 = 1 ~vt t and ~h0 = h1 ~htit and a = 0.
– By the properties of Boltzmann distributions, the constant in the upper
left hand corner of the matrix, say a, is arbitrary, since it would appear as
a multiplying factor e−βa in both numerator and denominator of (1). It
amounts to a minimum level attainable by the energy function,
          </p>
          <p>E0(~v, ~h | W 0) = a + E(~v, ~h | W |a=0)
and it may be used, for instance, to ensure that the joint distribution is
positive, that is, non-null for any ~v ∈ V, ~h ∈ H. We will not distinguish
between these two forms and consider the W = W 0|a=0, including ~b and ~c,
but not a, to be the set of parameters of the model pV H (~v, ~h | W 0|a=0).
– Z(W ) = P~v∈V P~h∈H e−E(~v,~h|W ) is the well-known partition function that
ensures the normalization of pV H|W . As argued, we will use sometimes
Z(W 0) = e−βa · Z(W ) to make it explicit that we include all possible
parameters, including the minimal offset. In fact, the partition function can be
included in the energy function by defining a = Fβ (W |k=0) = β1 loge Z(W |k=0)
in which case</p>
          <p>
            Z(W 0) = e−βFβ(W ) · Z(W |k=0) = Z(W |k=0)−1 · Z(W |k=0) = 1
so that pV H (~v, ~h) = e−βE(~v,~h|W 0) with an explicit normalization.
– With this encoding, the number of free parameters of this model is (L +
1)(K + 1) − 1 [
            <xref ref-type="bibr" rid="ref9">9</xref>
            ].
          </p>
          <p>
            Inference in harmoniums. As mentioned, RBM are harmoniums that can
also be seen as PoE [
            <xref ref-type="bibr" rid="ref1 ref10">1, 10</xref>
            ] where certain joint distributions are product of
individual components. In the case of the harmonium as a PoE, the components
are the conditional distribution of each of the hidden nodes given the input, or
the conditional distribution of each of the input nodes given the output:
pH|V (~h | ~v0, W ) =
          </p>
          <p>Y pHj|V (hj | ~v0, W ) pV |H (~v | ~h0, W ) =
j
i
Y pVi|H (vi | ~h0, W )
(3)</p>
          <p>
            In fact, RBMs are harmoniums with Bernouilli-distributed binary visible and
hidden variables (see below) [
            <xref ref-type="bibr" rid="ref1">1</xref>
            ]. But note that harmoniums have been generalized
to visible and hidden variables with distributions in the exponential family, which
include the most used distributions in data models [
            <xref ref-type="bibr" rid="ref11">11</xref>
            ].
2.2
          </p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>Information Semifields</title>
        <p>
          Positive semifields. A technical requisite for the energy function is that it
is always positive [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. This suggests investigating energy functions with positive
semifields [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ].
Example 1 (Multiplicative-product real semifields [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]). Consider a free
parameter r ∈ [−∞, 0) S(0, ∞] in the following operations
u ⊕r v =
        </p>
        <p>1
ur + vr r
u ⊗r v =</p>
        <p>
          1
ur × vr r
1
ur
1
r
u∗ =
= u−1
(4)
where the basic operations are to be interpreted in R≥0 and the dotted notation
is adopted from that used by Moreau for convex analysis [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ], where 0 × ∞ = 0
but 0 × ∞ = ∞. Notice the following properties:
– if r ∈ (0, ∞] then u ⊗r v = u ⊗r v = u × v, ⊥r = 0, er = 1, and &gt;r = ∞,
and the complete positive semifield generated, order-aligned with R≥0, is:
(R≥0)r = h[0, ∞], ⊕r, ⊗r, ·∗, ⊥r = 0, e, &gt;r = ∞i
and (R≥0)r−1 =
        </p>
        <p>R−≥01 r . In particular:
– if r ∈ [−∞, 0) then u⊗r v = u ⊗r v = u × v, ⊥r = ∞, er = 1, and &gt;r = 0, and
the complete positive semifield generated, dually order-aligned with R≥0, is:
(R≥0)−r = h[0, ∞], ⊕r, ⊗r, ·∗, ⊥r∗ = ∞, e, &gt;r∗ = 0i
Therefore, (R≥0)r and (R≥0)−r are inverse, completed positive semifields,
(R≥0)1 = R≥0
lim (R≥0)r = Rmax,×
r→∞
(R≥0)−1 = R−≥01
lim (R≥0)r−1 = Rmin,×
r→−∞</p>
        <p>Note that:
– these semifields are able to capture the usual concept of a “positive quantity”,
like a mass, length, etc.
– All these semifields have the same product, and the same “extreme” points,
{0, 1, ∞} . Their only difference lies in the addition. Sometimes, when only
the product is important in an application, the addition remains in the
background and we are not really sure in which algebra we are working on.
– Instead of using the abstract notation for the inversion ·∗, since R≥0 is the
paragon originating all other behaviours, we have decided to use the original
notation for the inversion in the (incomplete) semifield.</p>
        <p>The need for information functions. If we were to use the semifields (R≥0)r
to build the energy function of the harmonium model (1), despite the fact that
these semifields are positive, and they generate models that are products of
positive terms, from a physical point of view there is a mismatch between them
(5)
(6)
(7)
(8)
tu
and the numbers required in the model: these semifields describe natural
measurements of mass, length, etc., but model (1) demands that we use logarithmic
magnitudes.</p>
        <p>
          A first step is to use an information function to transform normalized masses,
that is, something akin to probabilities, into (quantity of ) information. Hartley’s
information function h(p) = − log2 p, p ∈ [
          <xref ref-type="bibr" rid="ref1">0, 1</xref>
          ] was extended by R´enyi to include
a free parameter α ∈ R±∞, the R´enyi order [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ], and this has later been shifted
into r = α − 1 [12, § 3.1]—where r ∈ [−∞, ∞] is the (shifted) R´enyi order —as:
ϕ0(h) = b−rh
ϕ0−1(p) = −1 logb p
r
Informational or Entropy semifields. The shifted R´enyi order allows us to
obtain new semifields from the R±∞ using R´enyi’s information function [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]:
Example 2 (Entropy semifields [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]). Let r ∈ [−∞, ∞]\{0} and b ∈ (1, ∞). Then
the algebra h[−∞, ∞], ⊕r, ⊗r, ·−1, ⊥ = ∞, e = 0i obtained from the semifield of
positive reals by R´enyi’s information function and whose basic operations are:
u ⊕r v = − 1r logb b−ru + b−rv
u ⊗r v = u + v
u∗ = −u
can be completed to two dually-ordered positive semifields
        </p>
        <p>Hr = h[−∞, ∞], ⊕r, ⊗r, −·, ⊥ = −∞, e = 0, &gt; = ∞i
−Hr = h[−∞, ∞], ⊕r, ⊗r, −·, −⊥ = ∞, e = 0, −&gt; = −∞i
whose elements can be considered as generalized information values and operated
accordingly. We will typically consider b = e1 so that loge(m) = log(m) and the
informations and entropies are measured in nats.</p>
        <p>It is easy to see that addition is a very complicated operation in information
semifields in general: For r ∈ R/{0} we expand the notation for addition as:
Xr</p>
        <p>i
X
i</p>
        <p>hi , hi ⊕r h2 ⊕r . . . = − 1r log(X e−rhi )
r hi , hi ⊕r h2 ⊕r . . . = − 1r log(X e−rhi )
i
i
r ∈ (0, ∞]
r ∈ [−∞, 0)</p>
        <p>Xr
It is easy to see that
Xri hi smoothly approximates the maximum, since:</p>
        <p>i hi is smooth approximation to the minimum while
i
X• hi , lim
r→∞</p>
        <p>Xr
i
hi = min hi
i</p>
        <p>i •
X hi , Xr hi = max hi</p>
        <p>i
i
In this case, the dually ordered complete positive semifields have an idempotent
addition, and are normally called the (completed) max-plus Rmax,+ and min-plus
Rmin,+ semifields, or tropical semirings.
(9)
(10)
(11)
(12)
(13)
(14)
h~x|W |~yir =4 ~xt ⊗r W ⊗r ~y =
xi ⊕r wij ⊕r yj = −1 log 
r</p>
        <p>X e−r(xi + wij + yj)
i,j
h~x|W |~yir =4 ~xt ⊗r W ⊗r ~y = X
1
r xi ⊕r wij ⊕r yj = r log </p>
        <p>X er(xi + wij + yj)</p>
        <p>Note that when ~h is a vector of generalized informations, the additions in (13)
are clearly entropies—more precisely, cross entropies—hence the name assigned
to this type of semifield. tu</p>
        <p>
          It is difficult to ascertain in what algebra the computations inside the
logsum-exp functions in (13) are carried out. The following result was proven in [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]:
Proposition 1. The generalized R´enyi information function is an isomorphism
of positive semifields between R≥0 and Hr. In particular, when r = 1 (the case
of Hartley’s function), it is an isomorphism between R≥0 and H.
Therefore, all computations can actually be carried out in R≥0. In expressions
we maintain, however, the more abstract notation, so that we write, for scalar
products over an informational semifield, in general
In a complete positive semifield K, an element ϕ is invertible if and only if it is
not extremal, ϕ ∈ K \ {⊥, &gt;}. We have then the four possible types of Galois
connections due to scalar products in the semifield K [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ].
        </p>
        <p>Theorem 1 (The four-fold connection). Let G and M be sets of formal
objects and formal attributes, respectively, with |G| = g and |M | = m . Let
(cGom,Mple,tRe )iKdebmepaotefonrtmsaelmciofineltdexKt.wChoosnesiidnecridtehnecveecRto∈r sKpagc×ems Xtak=esKvgalauneds Yin =a
Km, an invertible element ϕ = γ ⊗ μ ∈ K and the scaled spaces Xeγ = γ−1 ⊗ X
and Yeμ = μ−1 ⊗ Y . Then
1. The bracket hx | R | yioi = x∗ ⊗ R ⊗ y−1 induces a Galois connection
(·↑R, ·↓R) : Xeγ (* Yeμ between the scaled spaces through the polars
x↑R = Rt ⊗ x−1
yR↓ = R ⊗ y−1
which define two bijective sets, the system of extents BγG and the system of
μ
intents BM</p>
        <p>BγG = (Ye μ)↓</p>
        <p>R</p>
        <p>BM = (Xe γ )↑
μ</p>
        <p>R
(15)
and whose composition generate closure operators:
πR(x) = (x↑R)↓R = R ⊗(R∗ ⊗ x)
πRt (y) = (yR↓)↑R = Rt ⊗(R−1 ⊗ y)
which are the identities on BγG and BμM , respectively.
2. The bracket hx | R | yiio = xt ⊗ R ⊗ y induces a co-Galois connection
(·↑R−1 , ·↓R−1 ) : Xeγ +) Yeμ between the scaled spaces through the maps:
x↑R−1 = R∗ ⊗ x−1
yR↓−1 = R−1 ⊗ y−1
γ
which define two bijective sets the systems of neighbourhoods of objects NG
and neighbourhoods of attributes NμM
(17)
(18)
(19)
(20)
(21)
(22)
(23)
and whose composition generate interior operators:</p>
        <p>NγG = (Ye μ)↓R−1</p>
        <p>NμM = (Xe γ )↑R−1
κR−1 (x) = (x↑R−1 )↓R−1 = R−1 ⊗(Rt ⊗ x)</p>
        <p>↑
κR∗ (y) = (yR↓−1 )R−1 = R∗ ⊗(R ⊗ y)
which are the identities on NγG and NμM , respectively.
3. The bracket hx | R | yioo = x∗ ⊗ R ⊗ y induces a left adjunction (·∃R, ·∀R) :
Xeγ Yeμ between the scaled spaces through the left adjunct pair of maps:
x∃R = R∗ ⊗ x
yR∀ = R ⊗ y
which define another bijection between the systems of extents BγG and
neighμ
bourhoods of attributes NM</p>
        <p>BγG = (Ye μ)∀</p>
        <p>R
πR(x) = (x∃R)∀R</p>
        <p>NμM = (Xe γ )∃</p>
        <p>R
κR∗ (y) = (yR∀)∃R
and whose compositions are the closure of extents and interior of attributes:
4. The bracket hx | R | yiii = xt ⊗ R ⊗ y−1 induces an adjunction on the right
(·∀Rt , ·∃Rt ) : Xeγ Yeμ between the scaled spaces through the pair of adjunct
maps:
x∀Rt = Rt ⊗ x
yR∃t = R−1 ⊗ y
which define anothers bijection between the systems of neighbourhood of
objects NγG and intents BM</p>
        <p>μ
NγG = (Ye μ)∃Rt</p>
        <p>BM = (Xe γ )∀Rt
μ
and whose compositions are the interior of objects and the closure of
attributes:
κR−1 (x) = (x∀Rt )∃Rt
πRt (y) = (yR∃t )∀Rt
(26)
Therefore, it makes sense to define the following (meta) concept:
Definition 1. For a formal coγntext (G, M, R)K, thγe 4-formal conμcept (a, b, c, d)
is a 4-tuple such that a ∈ BG, b ∈ BμM , c ∈ NG, and d ∈ NM and all the
following relations hold:
a = (b)↓R = (d)∀R
d = (c)↑R−1 = (a)∃R
b = (a)↑R = (c)∀Rt
c = (d)↓R−1 = (b)∃Rt
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <p>
        (25)
(27)
(28)
(29)
Given the distinction between probability-based and information-based
semifields in Section 2.2, and considering that many sensorial magnitudes are
“logarithmically perceived”—despite the criticism to the Weber-Fenchner law [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]—
we would like to “fix” the magnitudes used in harmoniums while maintaining
their main design considerations.
3.1
      </p>
      <p>Information Measures from Mass Measures
We may further generalize R´enyi’s information function ϕ0−1(·) and its inverse
ϕ0(·) to vectors, so if we consider a mass measure m ∈ [0, ∞]I , we have its related
information measure:
hr(·) : [0, ∞]I → [−∞, ∞]I
−1 logb mi}i∈I
m = {mi}i∈I 7→ hr(m) = {ϕ0−1(mi)}i∈I = { r
whereas if we consider an information measure h ∈ [−∞, ∞]n, then we have its
related mass measure:
mr(·) : [−∞, ∞]I → [0, ∞]I
h = {hi}i∈I 7→ mr(h) = {e−rhi }i∈I
Note that the mr(·) and hr(·) thus defined are mutually inverse bijections of
tuple spaces between the semifields. In particular we state without proving:
Proposition 2. Let r ∈ [−∞, ∞] \ {0}. Then hr(·) : (R≥0)n → (Hr)n is a dual
isomorphism of semivector spaces over their corresponding semifields, with
inverse mr(·) : (Hr)n → (R≥0)n.</p>
      <p>The duality mentioned in Proposition 2 concerns their natural orders. The
Hartley information function, e.g. R´enyi’s function with r = 1 is a very special case:
Corollary 1. The Hartley information function is a dual isomorphism of
semivectors spaces h1(·) : (R≥0)n → (H)n .</p>
      <p>The following is much more interesting and admits the previous one as a
particular case:
Proposition 3. The Hartley information function is a dual isomorphism of
semivectors spaces hr(·) : ((R≥0)r)n → (Hr)n .</p>
      <p>rj
Proof. Let {αj ∈ R≥0 | j ∈ J } and {~vj ∈ (R≥0)rn
− log αj and ~xj = − log ~vj . Then, a linear combination o|f jve∈ctoJrs} inwi(tRh≥a0j)rn=,
X
αj ⊗r ~vj is transformed into a linear combination of vectors in (Hr)n as

− log  X
j


1/r

r αj ⊗r ~vj  = − log X αjr × ~vir</p>
      <p>j

= −1 log 
r</p>
      <p>j
 
X e−r(aj + ~xj) = −1 log 
r</p>
      <p>X e−raj × e−r~xj  =</p>
      <p>
X e−r(aj ⊗ ~xj) =</p>
      <p>j

= −1 log </p>
      <p>r
=</p>
      <p>Xr
j</p>
      <p>j
aj ⊗r ~xj
which is a linear combination of vectors with modified scalars. Since the
transformations are equalities and biunivocal it describes an isomorphism which is
inverted by taking the negative exponential function on each of the terms. tu</p>
      <p>In fact, we have also the following corollary:
Corollary 2. Let {β~, ~v} ⊂ ((R≥0)r)n and {~b, ~x} ⊂ (Hr)n where β~ = exp(−~b)
and ~v = exp(−~x). Then:
~b∗ ⊗ ~x = − log(β~∗ ⊗ ~v)
β~∗ ⊗ ~v = exp(−~b∗ ⊗ ~x)
(30)
Proof. Just collect n of the scalars αj on a single vector β~ = {αj−1}in=1 and
multiply by ~v as β~∗ ⊗ ~v. Then by the previous procedure we get ~b∗ ⊗ ~x = − log(β~∗ ⊗ ~v).
The second equality comes from the isomorphism. tu
3.2</p>
      <sec id="sec-3-1">
        <title>Informational Harmoniums</title>
        <p>The relation between mass measures and information measures suggests we
define harmoniums using information:
Definition 2. Let X ≡ Hrn and Y ≡ Hrm be semivector spaces over the
informational semifield Hr. Let ~x ∈ X and ~y ∈ Y be generalized information
functions. We call an informational harmonium one whose joint information
function is a scalar product between the spaces, with W ∈ (Hr)n×m,
r
hXY (·, · | W ) : X × Y → Hr</p>
        <p>r
(~x, ~y) 7→ hXY (~x, ~y | W ) = h~x|W |~yir</p>
        <p>If the information functions are related to the mass distributions mX = m(~x)
and mY = m(~y) by Hartley’s function, then the information harmonium is the
associated mass function over the whole product space (~x, ~y) ∈ X × Y that is
mXY (~x, ~y) = m(hXY (~x, ~y | W )) = {e−hrXY (~x,y~|W ) | ~x ∈ X , ~y ∈ Y}
r
with associated partition function and distribution
kmXY k1 =</p>
        <p>X</p>
        <p>X e−hrXY (~x,y~|W )
~x∈X y~∈Y
q1(mXY ) =
e−hrXY (~x,y~|W )
kmXY k1
(31)
(32)
tu
Proposition 4. In the conditions of Definition 2, let ~v = exp(−~x), ~h = exp(−~y)
and U = exp(W ) whereby we mean the entry-wise exponentiation. Then:
r r
mV H (~v, ~h | U ) = exp(−hXY (~x, ~y | W )) hXY (~x, ~y | W ) = − log(mV H (~v, ~h | U ))
(33)
Proof. This is just a repeating of the proof for Proposition 3.</p>
        <p>When r → ∞ we have the following corollary:
Corollary 3. In the conditions of Definition 2, let ~v = exp(−~x), ~h = exp(−~y)
and U = exp(W ) whereby we mean the entry-wise exponentiation. Then:
(~x)∗ ⊗ W ⊗ ~y = − log((~v)∗ × U × ~h)
(~v)∗ × U × ~h = exp(−(~x)∗ ⊗ W ⊗ ~y)
where the dotted notation refers to the Rmin,+ semifield.
3.3</p>
        <p>The FCA in Informational Harmoniums
From Theorem 1 in Section 2.3 we have:
Theorem 2. Informational harmoniums over the Rmax,+ semifields are
cryptomorphic to Rmax,+-formal contexts.
Proof. It is clear that harmoniums are in general bipartite graphs, as shown in
Fig. 1. This is one of the isomorphisms of standard formal contexts. With the
provisions we made after (1), we can gather the bias information into one extra
visible and one extra hidden node vo and ho extending the sets of nodes to L0
and K0. Therefore we extend W as suggested by including the biases into the
incidence W 0 so that (L0, K0, W 0) is a weighted bipartite graph, that is, a formal
context with entries in the carrier set of Rmax,+ or Rmin,+, which is the
lookedfor cryptomorphism.
tu
If the informational harmonium uses the particular form in (31),
h~x | W | ~yioo = (~x)∗ ⊗ W ⊗ ~y)
this allows us to borrow the results from part 3 of Theorem 1 and we know that
there is a left adjunction (·∃W , ·∀W ) : Xe Ye between the spaces through the left
adjunct pair of maps:
~x∃W = W ∗ ⊗ ~x
~y∀W = W ⊗ ~y
which define a bijection between the systems of visible nodes BL0 (L0, K0, W 0)
and neighbourhoods of hidden nodes NK0 (L0, K0, W 0)</p>
        <p>BL0 (L0, K0, W 0) = (Ye )∀W</p>
        <p>NK0 (L0, K0, W 0) = (Xe )∃W</p>
        <p>Regarding learning the harmonium, this type seems to resemble a
heteroassociative memory, for if (a, d) is a neighbourhood concept of the pair of lattices
then by the definition of the polars we have W ≥ ~a ⊗ d~∗, and in general the
harmonium can be built (not efficiently) as the join:</p>
        <p>W ≥</p>
        <p>_
(~a,d~)∈N(L0,K0,W 0)
~a ⊗ d~∗
(34)
4</p>
        <p>Conclusions and Further Research
We have introduced the information harmoniums as a way to “patch” the
magnitude problems of harmoniums: the energy function is not written in an algebra of
logarithmic quantities, whence the “dimensions” in natural units (probabilities)
of the harmonium generative model are incorrect.</p>
        <p>By means of defining generalized information and mass functions we have
been able to related the informational harmoniums to non-normalized
harmoniums defined in positive semifields obtained from the basic R≥0 semifield.</p>
        <p>If we further concentrate on the informational semifields of order r = ±∞
we recover the Rmax,×, Rmin,×, Rmax,+ and Rmin,+. When the harmoniums are
described with these semirings they relate to one of the four types of Galois
connection definable over the matrix of weights (W or U ): they are the scalar
products that defined the Galois connections, so that the posterior “mass
measures” and “information measures” are simply the polars.</p>
        <p>This opens up a number of avenues of research into the representation of the
spaces associated to r = ±∞-harmoniums, and suggest that K-FCA has points
to make relating to inference and learning of these, including the consideration
of the other three types of connections between spaces. Also, more efficient ways
to build harmoniums, as well as approximate building in the absence of the
supervision provided by the intents, d~ will be explored in future work.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Freund</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haussler</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Unsupervised learning of distributions on binary vectors using two layer networks</article-title>
          .
          <source>In: Advances in Neural Information Processing Systems</source>
          <volume>4</volume>
          . (
          <year>1992</year>
          )
          <fpage>912</fpage>
          -
          <lpage>919</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>LeCun</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bengio</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hinton</surname>
          </string-name>
          , G.:
          <article-title>Deep learning</article-title>
          .
          <source>Nature</source>
          <volume>521</volume>
          (
          <year>2015</year>
          )
          <fpage>436</fpage>
          -
          <lpage>444</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Smolensky</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <source>Information Processing in Dynamical Systems: Foundations of Harmony Theory. In: Parallel Distributed Processing</source>
          . MIT Press (
          <year>1986</year>
          )
          <fpage>194</fpage>
          -
          <lpage>281</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Ackley</surname>
            ,
            <given-names>D.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hinton</surname>
            ,
            <given-names>G.E.</given-names>
          </string-name>
          , Sejnowski, T.J.:
          <article-title>A learning algorithm for Boltzmann machines</article-title>
          .
          <source>Cognitive Science</source>
          (
          <year>1985</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>LeCun</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chopra</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hadsell</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ranzato</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>F.J.:</given-names>
          </string-name>
          <article-title>A tutorial on energy-based learning</article-title>
          . In Bakir, G.,
          <string-name>
            <surname>Hofman</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          , Scho¨lkopf,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Smola</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Taskar</surname>
          </string-name>
          , B., eds.: Predicting Structured Output. MIT Press (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Hinton</surname>
            ,
            <given-names>G.E.</given-names>
          </string-name>
          :
          <article-title>Training Products of Experts by Minimizing Contrastive Divergence</article-title>
          .
          <source>Neural Computation</source>
          <volume>14</volume>
          (
          <year>2002</year>
          )
          <fpage>1771</fpage>
          -
          <lpage>1800</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Hinton</surname>
            ,
            <given-names>G.E.</given-names>
          </string-name>
          :
          <article-title>A Practical Guide to Training Restricted Boltzmann Machines</article-title>
          . In Montavon, G.,
          <string-name>
            <surname>Orr</surname>
            ,
            <given-names>G.B.</given-names>
          </string-name>
          , Mu¨ller, K.R.,
          <source>eds.: Neural Networks: Tricks of the Trade. Volume 7700 of LNCS. 2nd edn</source>
          . Springer, Berlin, Heidelberg (
          <year>2012</year>
          )
          <fpage>599</fpage>
          -
          <lpage>619</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Murphy</surname>
            ,
            <given-names>K.P.:</given-names>
          </string-name>
          <article-title>Machine Learning. A Probabilistic Perspective</article-title>
          . MIT Press (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Cueto</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Morton</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sturmfels</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Geometry of the restricted Boltzmann machine</article-title>
          .
          <source>Contemporary Mathematics</source>
          <volume>516</volume>
          (
          <year>2010</year>
          )
          <fpage>135</fpage>
          -
          <lpage>153</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Freund</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haussler</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Unsupervised Learning of Distributions of Binary Vectors Using Two Layer Networks</article-title>
          .
          <source>Technical Report 94 25</source>
          ,
          <string-name>
            <surname>UCSC CRL</surname>
          </string-name>
          (
          <year>June 1994</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Welling</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosen-Zvi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hinton</surname>
          </string-name>
          , G.E.:
          <article-title>Exponential Family Harmoniums with an Application to Information Retrieval</article-title>
          .
          <source>In: Advances in Neural Information Processing Systems</source>
          . (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Valverde-Albacete</surname>
            ,
            <given-names>J.F.</given-names>
          </string-name>
          ,
          <article-title>Pela´ez-</article-title>
          <string-name>
            <surname>Moreno</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>The R´enyi Entropies Operate in Positive Semifields</article-title>
          .
          <source>Entropy</source>
          <volume>21</volume>
          (
          <issue>8</issue>
          ) (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Mesiar</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pap</surname>
          </string-name>
          , E.:
          <article-title>Idempotent integral as limit of g-integrals</article-title>
          .
          <source>Fuzzy Sets And Systems</source>
          <volume>102</volume>
          (
          <issue>3</issue>
          ) (
          <year>1999</year>
          )
          <fpage>385</fpage>
          -
          <lpage>392</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Moreau</surname>
            ,
            <given-names>J.J.</given-names>
          </string-name>
          :
          <article-title>Inf-convolution, sous-additivit´e, convexit´e des fonctions num´eriques</article-title>
          .
          <source>J. Math. Pures et Appl</source>
          .
          <volume>49</volume>
          (
          <year>1970</year>
          )
          <fpage>109</fpage>
          -
          <lpage>154</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Renyi</surname>
            ,
            <given-names>A.: Probability</given-names>
          </string-name>
          <string-name>
            <surname>Theory. Courier Dover Publications</surname>
          </string-name>
          (
          <year>1970</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Valverde-Albacete</surname>
            ,
            <given-names>F.J.</given-names>
          </string-name>
          ,
          <article-title>Pela´ez-</article-title>
          <string-name>
            <surname>Moreno</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>The Linear Algebra in Extended Formal Concept Analysis Over Idempotent Semifields</article-title>
          . In Bertet,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Borchmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Cellier</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.</surname>
          </string-name>
          , Ferr´e, S., eds.:
          <source>Formal Concept Analysis</source>
          . Springer Berlin Heidelberg, Rennes (
          <year>June 2017</year>
          )
          <fpage>211</fpage>
          -
          <lpage>227</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Mackay</surname>
            ,
            <given-names>D.M.:</given-names>
          </string-name>
          <article-title>Psychophysics of perceived intensity: A theoretical basis for Fechner's and Stevens' laws</article-title>
          .
          <source>Science</source>
          <volume>139</volume>
          (
          <issue>3560</issue>
          ) (
          <year>1963</year>
          )
          <fpage>1213</fpage>
          -
          <lpage>1216</lpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>