<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>GO Membrane Part</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Formal Concept Analysis Applied to Transcriptomic Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mehwish Alam</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Adrien Coulet</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Amedeo Napoli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Malika Sma¨ıl-Tabbone</string-name>
          <email>malika.smail@inria.fr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CNRS, LORIA, UMR 7503, Vandoeuvre-le`s-Nancy</institution>
          ,
          <addr-line>F-54506</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Inria, Villers-le`s-Nancy</institution>
          ,
          <addr-line>F-54600</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Universit ́e de Lorraine, LORIA, UMR 7503, Vandoeuvre-le`s-Nancy</institution>
          ,
          <addr-line>F-54506</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <volume>27</volume>
      <issue>0</issue>
      <fpage>339</fpage>
      <lpage>344</lpage>
      <abstract>
        <p>Identifying functions shared by genes responsible for cancer is a challenging task. This paper describes the preparation work for applying Formal Concept Analysis (FCA) to complex biological data. We present here a preliminary experiment using these data on a core context with the addition of domain knowledge. The resulting concept lattices are explored and some interesting concepts are discussed. Our study shows how FCA can help the domain experts in the exploration of complex data.</p>
      </abstract>
      <kwd-group>
        <kwd>Formal Concept Analysis</kwd>
        <kwd>Knowledge Discovery</kwd>
        <kwd>Transcriptomic Data</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Over past few years, large volumes of transcriptomic data were produced but
their analysis remains a challenging task because of the complexity of the
biological background. Some earlier studies aimed at retrieving sets of genes sharing
the same transcriptional behavior with the help of Formal Concept Analysis [
        <xref ref-type="bibr" rid="ref1 ref2">1,
2</xref>
        ]. Further studies analyze gene expression data by using gene annotations to
determine whether a set of differentially expressed genes is enriched with
biological attributes [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ]. Several efforts have been made for integrating heterogeneous
data [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. For example, at the Broad Institute, biological data were recently
gathered from multiple resources to get thousands of predefined genesets stored in
the Molecular Signature DataBase (MSigDB) [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. A predefined geneset is a set
of genes known to have a specific property such as their position on the genome,
their involvement in a molecular pathway etc.
      </p>
      <p>This paper focuses on the preparation of biological data to data mining
guided by domain knowledge. The objective is to apply knowledge discovery
techniques for analyzing a list of differentially expressed genes and
identifying functions or pathways shared by these genes assumed to be responsible for
cancer. Section 2 explains the proposed approach for FCA-based analysis of
biological data. Section 3 focuses on the conducted experiment. Section 4 discusses
the results. Section 5 concludes the paper.
c 2012 by the paper authors. CLA 2012, pp. 339–344. Copying permitted only for
private and academic purposes. Volume published and copyrighted by its editors.
Local Proceedings in ISBN 978–84–695–5252–0,
Universidad de M´alaga (Dept. Matem´atica Aplicada), Spain.</p>
    </sec>
    <sec id="sec-2">
      <title>The Proposed Framework</title>
      <p>
        We rely on the standard definition of FCA fully described in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] and adapt it
according to the current problem. Let G be the set of genes {g1, g2, g3, ..., gn}, and
M be a set of attributes of MSigDB for describing genes. M will be considered
as a partition of three points of view, M = M1 ∪ M2 ∪ M3, with Mi ∩ Mj = ∅
whenever i 6= j.
      </p>
      <p>The first set of attributes M1 refers to four types of attributes, “Location”,
“Pathway”, “Transcription Factors” and “GO Terms” (see Table 1). For our
convenience we have named MSigDB categories as types of attributes and used
only C1, C2, C3 and C5. The category C4 was not used as it keeps information
on sets of genes related to a certain kind of cancer, which is not useful for the
current problem. Thus we have a first context K1 = (G, M1, I1) where I1 denotes
the relation stating that gene gi has an attribute mj in M1.</p>
      <p>Types of Attributes
C1: Positional Gene Sets</p>
      <sec id="sec-2-1">
        <title>C2: Curated Gene Sets</title>
        <p>Description Data Provenance
Location of the gene on the chro- Broad Institute
mosome.</p>
        <p>Pathway
C3: Motif Gene Sets Transcription Factors
C4: Computational Gene Cancer Modules
Sets
C5: Gene Ontology (GO) Biological Process, Cellular Com- AmiGO
Gene Sets ponents, Molecular Functions</p>
      </sec>
      <sec id="sec-2-2">
        <title>KEGG, REACTOME, BIOCARTA Broad Institute Broad Institute</title>
        <p>The second set of attributes M2 is related to the so-called “categories” where
a category makes reference to a set of attributes with the “Pathway” type. For
example, “Cell Growth and Death” is an example of category (see Figure 1).
The categories in M2 determine a second context, K2 = (G, M2, I2) where M2 is
the set of categories and I2 denotes the relation between a gene and a category.
It can be noticed that the categories are only related to the “Pathway” type and
that they can be considered as domain knowledge.</p>
        <p>
          Moreover, the third set of attributes, namely M3, refers to the so-called
“upper categories”, which are defined as groupings of categories. Actually, we
have for the type “Pathway” a hierarchy of categories with two levels, categories
and upper categories (see Figure 1). The upper categories in M3 define a third
context K3 = (G, M3, I3) where M3 is the set of upper categories and I3 denotes
the relation between gene gi and an upper level category mj . Upper categories
are also related to the “Pathway” type and as categories, they can be considered
as domain knowledge too.
The framework described above was applied on three published sets of genes
corresponding to Cancer Modules defined in [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. Our test data are composed of
three lists of genes corresponding to the so-called “Cancer Module 1” (Ovary
Genes), “Cancer Module 2” (Dorsal Root Ganglia Genes), and “Cancer Module
5” (Lung Genes). For example, “PSPHL” is one gene with “Pathway” attribute
as “PPAR Signaling” which belongs to category “kc:Endocrine System” and
upper category “kuc:Organismal System”. Considering the three lists of genes
given by “Cancer Module 1”, “Cancer Module 2” and “Cancer Module 5”, we
built three different contexts having the same form as the context in Table 2).
Then we obtained three associated concept lattices with the help of the Coron
Plate-form (http://coron.loria.fr). The concept lattice for Table 2 is given
in Figure 2. The global characteristics of the three concept lattices are given in
Table 3.
        </p>
        <p>The exploration of a given concept lattice is carried out following the “Iceberg
metaphor”, i.e., the lattice is explored level by level according to the support of</p>
        <p>g
din )
n m
i r</p>
        <p>B e
P T
T O
A (G
s
r
o
t
p
e
c
e</p>
        <p>R
in y)
Sero(tPonathwa</p>
        <p>g
alin
n
ig )</p>
        <p>S y
R wa</p>
        <p>A th
PP Pa
(
)
r
o
t
c
2 Fa
0 n</p>
        <p>F2 tio
U3 crip
O s</p>
        <p>P an
V$ Tr
(</p>
        <p>ponenTterm)</p>
        <p>Com(GO
lar bly
llu em
e s
C s</p>
        <p>A
2 on)
1 i
chr5(qLocat
PBGSTenPBeH0s3L × × × × ×× × ×
CQMCNYTGC6PAT1 ×× ×× × × ×
Table 2. A toy example of formal context including domain knowledge.
e
rin
c
o
d
n m
E te
c: s
k y</p>
        <p>S
l
a
ism
n
a
rg s
uc:Ostem
k y</p>
        <p>
          S
each concept, where the the support of a concept is the cardinality of the extent.
In addition, we also used stability for extracting interesting frequent and stable
concepts [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ].
Data Sets No. of Genes No. of Attributes No. of Concepts Levels
Module 1 361 3496 9,588 12
Module 2 378 3496 6,508 11
        </p>
        <p>Module 5 419 3496 5,004 12
In this study, biologists are interested in links between the input genes in terms
of pathways in which they participate, relationships between genes and their
positions etc. We obtained concepts with shared transcription factors, pathways,
locations of genes and GO terms. After the selection of concepts with a high
support (≥ 10), we observed that there were some concepts with pathways either
related to cell proliferation or apoptosis (expert interpretation). The addition
of domain knowledge gives an opportunity to obtain the pathway categories
shared by larger sets of genes (as categories and upper categories are there for
maximizing the grouping of objects, see below).</p>
        <p>Table 4 shows the top-ranked concepts found in each module. For
example, in Table 4, we have the concept C4938:(KEGG Cytokine Cytokine Receptor
Interaction, kc:Signaling Molecules and Interaction, kuc:Environmental
Information Processing) and the concept C4995:(kc:Signaling Molecules and
Interaction, kuc:Environmental Information Processing). These two concepts are such
as C4938 ≤ C4995, meaning that C4995 has greater support than C4938.
Moreover, we observed that the introduction of categories and upper categories in
the global context allows us to consider concepts that otherwise would not be
frequent. Actually, the role of categories and upper level categories is to facilitate
the observation of sets of related genes.</p>
        <p>This is a general way of obtaining larger sets of objects to interpret. When
available, one can introduce a hierarchy of attributes –this is domain knowledge–
and then insert the levels of each attribute in this hierarchy as a new attribute
in the context. As a result, some classes of objects, that could not emerge before,
will appear based on these hierarchical indications. Given the test data sets, the
preliminary results obtained here constitute an interesting and positive control,
and confirm that FCA-based analysis offers an efficient and practical procedure
to explore complex and large sets of genes.
5</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Conclusion</title>
      <p>The preliminary study presented here shows how FCA can be applied to complex
biological data and can give flexibility in using various types of attributes for
analyzing a list of genes. In addition, domain knowledge can be introduced and
guide the analysis.
Dataset Concept Intents</p>
      <p>ID
Module 1 9585
9571
9566
9402
9078</p>
      <p>As for future work, we plan to take into account relationships between genes
and between terms (Gene Ontology relationships) and use the framework of
relational concept analysis.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Kaytoue-Uberall</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Duplessis</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuznetsov</surname>
            ,
            <given-names>S.O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Napoli</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Two FCA-Based Methods for Mining Gene Expression Data</article-title>
          . In Ferr´e,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Rudolph</surname>
          </string-name>
          , S., eds.
          <source>: ICFCA</source>
          . Volume
          <volume>5548</volume>
          of Lecture Notes in Computer Science., Springer (
          <year>2009</year>
          )
          <fpage>251</fpage>
          -
          <lpage>266</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Rioult</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boulicaut</surname>
            ,
            <given-names>J.F.</given-names>
          </string-name>
          , Cr´emilleux,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Besson</surname>
          </string-name>
          , J.:
          <article-title>Using Transposition for Pattern Discovery from Microarray Data</article-title>
          . In: DMKD. (
          <year>2003</year>
          )
          <fpage>73</fpage>
          -
          <lpage>79</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Berriz</surname>
            ,
            <given-names>G.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>King</surname>
            ,
            <given-names>O.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bryant</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sander</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roth</surname>
            ,
            <given-names>F.P.</given-names>
          </string-name>
          :
          <article-title>Characterizing gene sets with FuncAssociate</article-title>
          .
          <source>Bioinfo</source>
          .
          <volume>19</volume>
          (
          <issue>18</issue>
          ) (
          <year>2003</year>
          )
          <fpage>2502</fpage>
          -
          <lpage>2504</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Doniger</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salomonis</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dahlquist</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vranizan</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lawlor</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Conklin</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>MAPPFinder: using Gene Ontology and GenMAPP to Create a Global Geneexpression Profile from Microarray Data</article-title>
          .
          <source>Genome Biology</source>
          <volume>4</volume>
          (
          <issue>1</issue>
          ) (
          <year>2003</year>
          ) R7
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Galperin</surname>
            ,
            <given-names>M.Y.</given-names>
          </string-name>
          ,
          <article-title>Ferna´ndez-</article-title>
          <string-name>
            <surname>Suarez</surname>
            ,
            <given-names>X.M.:</given-names>
          </string-name>
          <article-title>The 2012 Nucleic Acids Research Database Issue and the online Molecular Biology Database Collection</article-title>
          .
          <source>Nucleic Acids Research</source>
          <volume>40</volume>
          (
          <year>2012</year>
          )
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Liberzon</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Molecular Signatures Database (MSigDB) 3.0</article-title>
          . Bioinfo.
          <volume>27</volume>
          (
          <issue>12</issue>
          ) (
          <year>2011</year>
          )
          <fpage>1739</fpage>
          -
          <lpage>1740</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Ganter</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wille</surname>
          </string-name>
          , R.:
          <source>Formal Concept Analysis: Mathematical Foundations</source>
          . Springer, Berlin/Heidelberg (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Segal</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Friedman</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koller</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Regev</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>A Module Map Showing Conditional Activity of Expression Modules in Cancer</article-title>
          .
          <source>Nat.Genet</source>
          .
          <volume>36</volume>
          (
          <year>2004</year>
          )
          <fpage>1090</fpage>
          -
          <lpage>8</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Kuznetsov</surname>
            ,
            <given-names>S.O.</given-names>
          </string-name>
          :
          <article-title>On stability of a Formal Concept</article-title>
          . Ann. Math. Artif. Intell.
          <volume>49</volume>
          (
          <issue>1-4</issue>
          ) (
          <year>2007</year>
          )
          <fpage>101</fpage>
          -
          <lpage>115</lpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>