<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Textual Information Extraction in Document Images Guided by a Concept Lattice</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Cynthia Pitou</string-name>
          <email>cpitou@gaa.fr</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jean Diatta</string-name>
          <email>jean.diatta@univ-reunion.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>EA2525-LIM, Saint-Denis de La Reunion</institution>
          ,
          <addr-line>F-97490</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Groupe Austral Assistance</institution>
          ,
          <addr-line>16 rue Albert Lougnon, 97490 Sainte Clotilde</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Text Information Extraction in images is concerned with extracting the relevant text data from a collection of document images. It consists in localizing (determining the location) and recognizing (transforming into plain text) text contained in document images. In this work we present a textual information extraction model consisting in a set of prototype regions along with pathways for browsing through these prototype regions. The proposed model is constructed in four steps: (1) produce synthetic invoice data containing the textual information of interest, along with their spatial positions; (2) partition the produced data; (3) derive the prototype regions from the obtained partition clusters; (4) build the concept lattice of a formal context derived from the prototype regions. Experimental results, on a corpus of 1000 real-world scanned invoices show that the proposed model improves signi cantly the extraction rate of an Optical Character Recognition (OCR) engine.</p>
      </abstract>
      <kwd-group>
        <kwd>textual information extraction</kwd>
        <kwd>concept lattice</kwd>
        <kwd>clustering</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Document processing is the transformation of a human understandable data in a
computer system understandable format. Document analysis and understanding
are the two phases of document processing. Considering a document containing
lines, words and graphical objects such as logos, the analysis of such a
document consists in extracting and isolating the words, lines and objects and then
grouping them into blocks. The subsystem of document understanding builds
relationships (to the right, left, above, below) between the blocks. A document
processing system must be able to: locate textual information, identify if that
information is relevant comparatively to other information contained in the
document, extract that information in a computer system understandable format.
For the realization of such a system, major di culties arise from the variability
of the documents characteristics, such as: the type (invoice, form, quotation,
report, etc.), the layout (font, style, disposition), the language, the typography and
the quality of scanning. In the literature, works in pattern recognition [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] and
character recognition [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ] provide solutions for textual information extraction in
a computer system understandable format. Works in automatic natural language
processing [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] contribute to solving the problem about the detection of relevant
information. This paper is concerned with scanned documents, also known as
document images. We are particularly interested in locating textual information
in invoice images. Invoices are largely used and well regulated documents, but
not uni ed. They contain mandatory information (invoice number, unique
identi er of the issuing company, VAT amount, net amount, etc.) which, depending
on the issuer, can take various locations in the document. For instance, it seems
di cult to identify a trend as to the position of the date and invoice number.
However, similarities may occur locally for one or many information. To take
an example, the amount is usually positioned at bottom-right in the French and
English systems. Recent approaches such as those presented in [
        <xref ref-type="bibr" rid="ref3 ref4 ref9">3, 4, 9</xref>
        ] are
specifically concerned with the extraction of information in administrative documents
such as invoices. These works have in common the search, within a base, for a
document similar to an input document. Each document of this base is assigned
a template that lists some attributes (position, type, keywords) to be used in
order to locate information contained in similar input documents. Bartoli et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
propose a system of selecting, for an input document, the nearest wrapper based
on a distance measure. A wrapper is an object containing information about
geometric properties and textual content of elements to extract. Belaid et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]
propose a case-based-reasoning approach for invoice processing. Cesarini et al.
[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] propose a system to process documents that can be grouped into classes. The
system comprises three phases: (1) document analysis, (2) document classi
cation, (3) document understanding.
      </p>
      <p>
        The present paper is in the framework of region-based textual information
localization and extraction [
        <xref ref-type="bibr" rid="ref29 ref30">29, 30</xref>
        ]. We present a textual information extraction
model consisting in a set of prototype regions along with pathways for
browsing through these prototype regions. The proposed model is constructed in four
steps:
1. produce synthetic invoice data from real-world invoice images containing the
textual information of interest, along with their spatial positions;
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. partition the produced data;</title>
    </sec>
    <sec id="sec-3">
      <title>3. derive the prototype regions from the obtained partition clusters;</title>
      <p>4. derive pathways for browsing through the prototype regions, from the
concept lattice of a suitably de ned formal context;
The paper is organized as follows. Section 2 is devoted to the construction of
prototype regions. The formal context de ned using the obtained prototype
regions, and the determination of paths from the concept lattice of that formal
context are described in Section 3. Section 4 presents our approach for textual
information extraction, using the de ned paths. Finally, some experimental
results are presented in Section 5 and the paper is closed with a conclusion and
perspectives.</p>
      <sec id="sec-3-1">
        <title>Construction of prototype regions</title>
        <sec id="sec-3-1-1">
          <title>Construction of a synthetic data set</title>
          <p>The present work is motivated by the request of a company interested in
developing its own system enabling to automatically extract some textual information
from scanned invoices. The company has provided us with a corpus of 1000
real-world scanned invoices, emitted by 18 service providers whose business is
around car towing and auto repair. All the images are one page A4 documents.
The whole set of information the company is interested in, comprises: invoice
number, invoice date, net amount, VAT amount, customer reference, the type of
service provided, the issuer identity. In our study, we consider only the following
ve information:
{ I1: the key word of the service provided: towing or auto repair,
{ I2: customer reference: a string of 9-13 characters,
{ I3: the plate number of the assisted vehicle,
{ I4: the invoice issuer unique identi er: a string of 14 digits,
{ I5: the net amount of money requested for the provided service.
Each of the textual information is located in a region delimited by a rectangle
de ned by the coordinates (x; y) (in pixels) of its top left corner and the
coordinates (z; q) of its bottom right corner. In the sequel, by the term region will
be meant a rectangular area in an invoice image. Hence, a region may be
represented by the four coordinates (x; y; z; q) of its top left and bottom right corners.
As the information to be extracted are located in (rectangular) regions we adopt
a region-based extraction approach. The regions which the proposed approach
is based on are prototypes obtained from the more speci c regions containing,
each, a single information. Now, the coordinates of the regions containing the
needed information are not available for the real-world scanned invoices at hand.
To cope with this, we develop a JAVA program, with a graphical interface,
enabling to create synthetic invoice data simulating the real-world scanned invoices
along with the approximate coordinates of the speci c rectangles containing the
needed information. For instance, from a real-world scanned invoice an initial
synthetic invoice is manually created. This synthetic invoice is a single black
and white A4 page. This page will contain a string corresponding to a plate
number approximately at the same location as the plate number information
appears in the real-world invoice. Additionally, the string will be inserted
approximately with the same size and the same font as in the real-world invoice
in order to look like it. The string is inserted manually in the initial synthetic
invoice as one can do with a text editor. However, many strings contained in
the real-world invoice are not reproduced in the synthetic invoice. For instance,
the information about the emitter and the receiver (address, phone number,
...) are not reproduced because they are not relevant for the study. Thus, the
set of textual information I1 to I5 is placed manually on the synthetic invoice.
Then, from the obtained initial synthetic invoice a xed number of synthetic
invoice images may be created automatically. In such synthetic invoice images,
the information locations are maintained identically to the initial synthetic
invoice but the contained string may vary. Finally, for each distinct emitter of the
real-world invoice images corpus, one initial synthetic invoice image is created
manually and a xed number of synthetic invoice images is created
automatically from the initial synthetic invoice. A corpus of 1000 synthetic invoice images
is thus produced and, for each synthetic invoice, both the textual information
and the coordinates of the respective rectangles containing them, are stored in
a database. The original distribution per emitter of the real-world invoices is
preserved in the synthetic corpus of images. Synthetic data sets can then be
generated from this database for closer insight. An example of such data sets
is a set of (synthetic invoice) records described by 20 variables representing 5
blocks of 4 coordinates (x; y; z; q), each block being associated with one of the 5
considered information I1 to I5. It should be noted that this possibility to
produce a synthetic representation of a real-world scanned invoice is an important
step for updating the proposed model, namely when one has to extract
information from previously unseen scanned invoice. This point will be discussed later
in Section 6.
2.2</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>Clustering of the synthetic data</title>
          <p>As we mentioned in Section 2.1, the regions which our proposed approach is
based on are the so-called prototype regions, obtained from the speci c regions
that contain, each, a single information. More precisely, a prototype region
associated with a given information should be a region containing a homogeneous
set of speci c regions related to various positions of this information in di erent
invoice images. This makes cluster analysis methods, (such as the partitioning
ones) good candidates for capturing such homogeneous sets of speci c regions.
Then, in the next section, the construction of prototype regions from such
homogeneous sets of speci c regions is explained.</p>
          <p>
            The synthetic data obtained from the previous phase can be partitioned either:
(a) in an overall view taking into account all of the 20 variables, or (b) in 5
independent views, each corresponding to one of the 5 information I1 to I5 and
taking into account, for each view, the 4 associated variables. The approach in
ve independent views consists in creating ve data sets: D1, D2, D3, D4 and
D5. A record in Di is described by the four coordinates of regions containing
information Ii. For both approaches we adopted the K-means [
            <xref ref-type="bibr" rid="ref23">23</xref>
            ] clustering
with Euclidean distance. K-means is a popular, simple and e cient algorithm
for cluster analysis. To determine the number of clusters, we conducted, on the
one hand, an (agglomerative) ascending hierarchical clustering with Ward
criterion (clustering method based on a classical sum-of-squares criterion, producing
groups that minimize within-group dispersion) [
            <xref ref-type="bibr" rid="ref24">24</xref>
            ] and, on the other hand,
executions of K-means for values of k between 2 and 18. Several validity criteria,
such as within cluster sum of squares, silhouette and Calinkski-Harabasz [
            <xref ref-type="bibr" rid="ref33">33</xref>
            ],
provided by the package clusterCrit of R software, were used to determine the
optimal number of clusters for each data set. It turns out that considering ve
independent views leads to better clusters w.r.t. each of the considered validity
criteria. Therefore, we adopt the option consisting in partitioning each of the 5
views. The best values of k obtained for the respective ve views are shown in
Table 1.
As we mentioned in the previous section, a prototype region associated with
an information should contain a homogeneous set of speci c regions related to
various positions of this information in di erent invoice images.
          </p>
          <p>Recall that for each information Ii, the associated data set Di is partitioned into
some number of clusters (see Table 1). Then, we associate to each of these
clusters, say C, a prototype region de ned as the smallest rectangle R containing each
of the speci c rectangles in C. Thus, we obtain 43 prototype regions R1,...,R43,
with the rst 10 related to information I1, the next 10 to I2, the next 10 to I3,
the next 3 to I4 and the last 10 to I5. Figure 1 shows the prototype regions
related to information I4. The next step of the construction of our proposed
model is to set up pathways for e ciently browsing through the set of de ned
prototype regions. Such pathways will be obtained from the concept lattice of a
suitably de ned formal context.
3</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>Determination of pathways for browsing through the prototype regions</title>
        <p>So far, we indicated how we determine prototype regions containing the textual
information to be extracted. So we come to the fourth step in our approach,
namely, de ne pathways for e ciently browsing through the set of these
prototype regions. For this, Formal Concept Analysis (FCA) appears very appropriate.
Indeed, the pathways we seek to determine may be obtained from the concept
lattice of a suitably de ned formal context.
3.1</p>
        <sec id="sec-3-2-1">
          <title>Construction of the concept lattice</title>
          <p>
            Recall that a formal context is a triple K = (O; A; R), where O is a set of objects,
A a set of attributes and R O A a binary relation from O to A. A formal
concept of K is a pair (X; Y ) such that Y = X0 = fa 2 A : xRa for all x 2 Xg
and X = Y 0 = fx 2 O : xRa for all a 2 Y g. Note that the double application of
the derivation operator (:)0 is a closure operator, i.e. (:)00 is extensive, idempotent
and monotone. Sets X O; Y A, such that X = X00 and Y = Y 00 are said to
be closed. The subset X O is called the extent of the concept (X; Y ) and Y its
intent. The concept lattice of the formal context K [
            <xref ref-type="bibr" rid="ref35">35</xref>
            ], also known as the Galois
lattice of the binary relation R [
            <xref ref-type="bibr" rid="ref2">2</xref>
            ], is the (complete) lattice (L(K); ), where
L(K) is the set of formal concepts of K and the subconcept/superconcept
partial order. Thus, a concept lattice contains a minimum (resp. a maximum)
element according to the relation , called the bottom (resp. the top). In this
work, we consider the formal context where the objects are the invoice images
and the attributes the predicates Ii = j, where Ii, i = 1; : : : ; 5 denotes the ve
textual information mentioned in Section 2.1, and j = 1; : : : ; 43 denotes the
ID of the 43 prototype regions R1,...,R43. An invoice on is in relation with a
predicate Ii = j if the textual information Ii is located at prototype region Rj
in the invoice on. A summary of this formal context is shown in Table 2.
          </p>
        </sec>
        <sec id="sec-3-2-2">
          <title>Determination of paths from the concept lattice</title>
          <p>
            Recall that, given a formal context K = (O; A; R), an association rule is a pair
(X; Y ), denoted as X ! Y , where X and Y are disjoint subsets of A [
            <xref ref-type="bibr" rid="ref1">1</xref>
            ]. The set
X is called the antecedent of the rule X ! Y and Y its consequent. The support
of an association rule X ! Y is the proportion of objects that contain all the
attributes in X [ Y , i.e. j(X[Y )0j . The con dence of X ! Y is the proportion of
jOj
objects that contain Y , among those containing X. A (support,con dence)-valid
o1
...
o895
          </p>
          <p>...
o1000</p>
          <p>
            X
association rule is an association rule whose support and con dence are at least
equal to a xed minimum support threshold and a xed minimum con dence
threshold, respectively. An approximate association rule is an association rule
whose con dence is less than 1. When the minimum support threshold is set to
0, the Luxenburger basis of approximate association rules is the set of rules of
the form X ! Y n X where X = X00, Y = Y 00, X Y and there is no Z such
that Z00 = Z and X Z Y [
            <xref ref-type="bibr" rid="ref21">21</xref>
            ].
          </p>
          <p>
            The Luxenburger basis can be visualized directly in the line diagram of a concept
lattice. Each approximate rule in the Luxenburger basis corresponds exactly to
one edge in the line diagram. The line diagram of a lattice contains paths by
which one can move from the top concept to the bottom one. The pathways
we adopt for browsing through the set of prototype regions are exactly those
corresponding to sequences of association rules of the Luxenburger basis, i.e.
top-down consecutive edges in the concept lattice. In other words, a pathway
is a sequence Y0 ! Y1 ! ::: ! Yn, where Y0 is the intent of the top formal
concept and for all 0 i &lt; n, Yi ! Yi+1 is an association association rule
of the Luxenburger basis. Given a node of the concept lattice, there are as
many approximate association rules of the Luxenburger basis whose antecedent
is the intent of this node, as are the children nodes of this node in the concept
lattice. Between two approximate association rules having the same antecedent,
the one with highest support is considered rst. For instance, let a pathway
p1: I5=42 ! fI1=9, I3=25g holds with a support of 4% and a pathway p2:
I5=42 ! fI2=19, I3=28g holds with a support of 6%. In the aim to extract
information I1 to I5 from a candidate invoice image, and supposing that I5=42
is the lattice top node's direct child node which holds the highest support value,
prototype region R42 is visited rst in order to nd information I5. Then, using
pathway p2, prototype regions R19 and R28 are visited for nding information
I2 and I3 respectively. When, an information Ii is not found in a prototype
region given by pathway p2, so p1 may be used to nd it. Thus, all approximate
association rules given by the Luxenburger basis are used for information I1 to
I5 localization and extraction. In systems, such as, CREDO [
            <xref ref-type="bibr" rid="ref8">8</xref>
            ] and SearchSleuth
[
            <xref ref-type="bibr" rid="ref12">12</xref>
            ] the browsing strategy consists in focusing on a concept and its neighbors.
The e ectiveness and performance for using this type of strategy in Web search
have been demonstrated in [
            <xref ref-type="bibr" rid="ref12 ref8">8, 12</xref>
            ].
4
          </p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>Textual information extraction</title>
        <p>To extract the textual information of interest, we perform an optical character
recognition (OCR) engine on prototype regions, using the pathways determined
in the previous step. Recall that a pathway is a sequence Y0 ! Y1 ! ::: ! Yn,
where Y0 is the intent of the top formal concept and for all 0 i &lt; n, Yi ! Yi+1
is an approximate association rule of the Luxenburger basis. It should be noted
that each node Y in such a sequence represents a set of predicates \Ii=j"
indicating that information Ii belongs to prototype region Rj . First, the set of
pathways is ordered by descending support value of the intents. Then, in each
node given by a pathway, an OCR engine is performed on each prototype region
in order to extract the corresponding information Ii. In formal language theory,
a regular expression is a sequence of characters that de nes a search pattern,
mainly for use in pattern matching with strings. For each sought information Ii,
a regular expression is built and then used to check whether the extracted string
(by the OCR engine) matches with the given information.</p>
        <p>
          In the literature, approaches such as in [
          <xref ref-type="bibr" rid="ref18 ref19 ref34">34, 18, 19</xref>
          ] are based on concept lattice
classi er and use a concept lattice for a classi cation task. Such approaches aim
to improve the task of character or symbol recognition in images. In [
          <xref ref-type="bibr" rid="ref34">34</xref>
          ], the
authors developed a recognition system named Navigala and tted to recognize
noisy graphical objects and especially symbols images in technical documents
such as architectural plans or electrical diagrams. The authors noted that
Navigala is somewhat generic and can be successfully applied to other types of data.
In [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ], the authors proposed modi cations of some classi ers (naive Bayes,
nearest neighbor and random forest classi ers) in order to use the modi ed classi ers
as a part of the ABBYY OCR Technologies recognition schema for its
performance improving. The authors note that their approach based on random forest
can be applied to combine results of concept lattice classi ers.
        </p>
        <p>In this paper, the task of text extraction is done with a free OCR engine named
Tesseract OCR (https://github.com/tesseract-ocr). Tesseract OCR was chosen
because it is a free software providing a JAVA API. In this paper, we focus
on the textual information localization task in administrative document images.
Indeed, OCR engine such as ABBYY OCR has a better recognition rate than
free OCR such as Tesseract OCR, but both are not able to localize or pick out a
given information such as the net amount in invoice images. Their task is just to
transform, as e ciently as possible, the text contained in images into plain text.
In this work, we propose to combine the proposed localization approach (based
on clustering analysis and navigation in a concept lattice) with any OCR engine
in order to extract a given information in document images without browsing
and recognizing the entire images.</p>
      </sec>
      <sec id="sec-3-4">
        <title>Experimental results</title>
        <p>We achieved an experiment in order to test the proposed model for textual
information extraction in real-world invoice images. The experiment consists
in extracting information I1 to I5, in the set of 1000 real-word invoice images
(Section2.1). Despite the fact that the approach was trained and tested with
good results on the synthetic data, in this section we present test results of the
approach on real-world invoice images. Indeed, the corpus of real-world invoice
images contains some noise which is not present in the synthetic invoice images.
The real-world invoice images may contain colored images such as a logo, shadow
areas and handwritten text. Additionally, they may be scanned with poor
quality and may present distortion. Thus, the corpus of real-world invoice images
seems to us to be quite interesting for testing the proposed textual information
localization and extraction model. We performed two types of extraction:
1. from full images: OCR is performed on the entire page images regardless to
speci c regions;
2. from prototype regions, using the pathways presented in Section 3: OCR is
performed only on image sub-regions, using the pathways.</p>
        <p>To perform OCR on images, the JAVA library of the free OCR engine named
Tesseract in its 3.02 version is used. We considered two measures:
1. the rate of correct information among the total number of sought information
(recall),
2. the rate of correct information among the total number of detected
information (precision).</p>
        <p>
          On the one hand, a sought information is considered detected, if a string which
matches the corresponding regular expression is found. On the other hand, a
sought information is considered correctly extracted, if the extracted textual
information corresponds exactly to the visual information that should be read
in the image. For instance, let 'net amount' be a sought information and assume
that the net amount is 107e in the invoice image. During the process, if the
retrieved information is \101e", the net amount will not be considered correctly
extracted because the real net amount mentioned in the original invoice image is
\107e". The results are presented in Table 3. On the one hand, despite the fact
that according to [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ], the Tesseract OCR engine has accuracy of 70% for text
extraction in gray scale number plate images, we observe that the accuracy of the
OCR engine is weak for text extraction in real-world invoice images. On the other
hand, these results show that our proposed model improves signi cantly the
performance of the OCR engine. Note that a p-value of 8.799e-05 was obtained
for this experiment, which means that the results are signi cant.
6
        </p>
      </sec>
      <sec id="sec-3-5">
        <title>Conclusion and perspectives</title>
        <p>
          We presented a (prototype) region-based model for localizing and extracting
textual information in document images. Experimental results show that the
proposed model improves signi cantly the correctness of a textual information
extraction process based on an OCR engine. This model is constructed in four
steps:
1. produce synthetic invoice data from real-world invoice images containing the
textual information of interest, along with their spatial positions;
2. partition the produced data;
3. derive the prototype regions from the obtained partition clusters;
4. derive pathways for browsing through the prototype regions, from the
concept lattice of a suitably de ned formal context;
The Step 1 is important when one has to extract information from a previously
unseen invoice image. Indeed, if some information of such an invoice image are
not retrieved, then a synthetic representation of the considered invoice can be
produced, and this triggers incremental updates of the synthetic data sets, the
prototype regions, and the concept lattice. Our future work will focus on these
update processes, and compare them with those proposed in [
          <xref ref-type="bibr" rid="ref3 ref4 ref9">3, 4, 9</xref>
          ]. We also
plan to develop a classi cation model, as in [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], that will enable to predict the
invoice emitter based on the ve textual information I1 to I5 considered in the
present paper. This will allow to easier retrieve the other textual information
the company is interested in: invoice number, invoice date, tax rate, tax due.
        </p>
      </sec>
      <sec id="sec-3-6">
        <title>Acknowledgment References</title>
        <p>This work was partially carried out within the project ClustOverlap supported
by Reunion Island Region - grant DIRED 20140704. The authors are also very
grateful to the 3 anonymous reviewers for their valuable comments.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Agrawal</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Imielinski</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Swami</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Mining association rules between sets of items in large databases</article-title>
          .
          <source>ACM SIGMOD Record</source>
          ,
          <volume>22</volume>
          .2, pp.
          <volume>207</volume>
          {
          <issue>216</issue>
          (
          <year>1993</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Barbut</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Monjardet</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Ordre et classi cation: algebre et combinatoire</article-title>
          .
          <source>Hachette</source>
          (
          <year>1970</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bartoli</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Davanzo</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Medvet</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sorio</surname>
          </string-name>
          , E.:
          <article-title>Semisupervised wrapper choice and generation for print-oriented documents</article-title>
          .
          <source>IEEE Transactions on Knowledge and Data Engineering 26.1</source>
          , pp.
          <volume>208</volume>
          {
          <issue>220</issue>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4. Belad, A.,
          <string-name>
            <surname>D'Andecy</surname>
            ,
            <given-names>V. P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hamza</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          , Belad, Y.:
          <article-title>Administrative document analysis and structure</article-title>
          .
          <source>Learning Structure and Schemas from Documents</source>
          , pp.
          <volume>51</volume>
          {
          <issue>71</issue>
          (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Birkho</surname>
          </string-name>
          , G.:
          <source>Lattice Theory. American Mathematical Society</source>
          <volume>25</volume>
          .3 (
          <year>1967</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Cambria</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>White</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Jumping NLP Curves: A Review of Natural Language Processing Research</article-title>
          .
          <source>IEEE Computational Intelligence Magazine 9.2</source>
          , pp.
          <volume>48</volume>
          {
          <issue>57</issue>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Carpineto</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Michini</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nicolussi</surname>
            ,
            <given-names>R.:</given-names>
          </string-name>
          <article-title>A concept lattice-based kernel for SVM text classi cation</article-title>
          .
          <source>Formal Concept Analysis</source>
          , pp.
          <volume>237</volume>
          {
          <issue>250</issue>
          (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Carpineto</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Romano</surname>
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>Expoiting the Potential of Concept Lattices for Information Retrieval with CREDO</article-title>
          .
          <source>Journal of Universal Computer Science</source>
          ,
          <volume>10</volume>
          .8, pp.
          <volume>985</volume>
          {
          <issue>1013</issue>
          (
          <year>2004</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Cesarini</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Francesconi</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gori</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soda</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>Analysis and understanding of multi-class invoices</article-title>
          .
          <source>International Journal on Document Analysis and Recognition 6.2</source>
          , pp.
          <volume>102</volume>
          {
          <issue>114</issue>
          (
          <year>2003</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Charikar</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chekuri</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Feder</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <article-title>Rajeev Motwani: Incremental clustering and dynamic information retrieval</article-title>
          .
          <source>In: Proceedings of the twenty-ninth annual ACM symposium on Theory of computing. ACM</source>
          , pp.
          <volume>626</volume>
          {
          <issue>635</issue>
          (
          <year>1997</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Cho</surname>
            ,
            <given-names>W.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Richards</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Improvement of Precision and Recall for Information Re- trieval in a Narrow Domain: Reuse of Concepts by Formal Concept Analysis</article-title>
          .
          <source>In Proceedings of the 2004 IEEE/WIC/ACM International Conference on Web Intelligence (WI '04)</source>
          , pp.
          <volume>370</volume>
          {
          <issue>376</issue>
          (
          <year>2004</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Ducrou</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eklund</surname>
            ,
            <given-names>P. W.:</given-names>
          </string-name>
          <article-title>SearchSleuth: The conceptual neighbourhood of an web query</article-title>
          . In J. Diatta,
          <string-name>
            <given-names>P.</given-names>
            <surname>Eklund</surname>
          </string-name>
          , &amp; M. Liquire (Eds.),
          <source>Proc. CLA</source>
          <year>2007</year>
          , LIRMM &amp; University of Montpellier II (
          <year>2007</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Eisenbarth</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koschke</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Simon</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Locating Features in Source Code</article-title>
          .
          <source>IEEE Transactions on software engineering 29.3</source>
          , pp.
          <volume>210</volume>
          {
          <issue>224</issue>
          (
          <year>2003</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Ganter</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wille</surname>
          </string-name>
          , R.:
          <source>Contextual attribute logic</source>
          . Springer Berlin Heidelberg (
          <year>1999</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Ganter</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wille</surname>
          </string-name>
          , R.:
          <source>Formal Concept Analysis: Mathematical Foundations</source>
          . Springer Science &amp; Business
          <string-name>
            <surname>Media</surname>
          </string-name>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Guyon</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cawley</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dror</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <article-title>Sa ari, A.: Hands-on Pattern Recognition. Challenges in Machine Learning 1</article-title>
          .
          <string-name>
            <given-names>Isabelle</given-names>
            <surname>Guyon</surname>
          </string-name>
          , Gavin Cawley, Gideon Dror, and Amir Sa ari editors (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Hyontai</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>An E ective Sampling Method for Decision Trees Considering Comprehensibility and Accuracy</article-title>
          .
          <source>WSEAS Transactions on Computers 8.4</source>
          , pp.
          <volume>631</volume>
          {
          <issue>640</issue>
          (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Itskovich</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuznetsov</surname>
            ,
            <given-names>S.O.</given-names>
          </string-name>
          <article-title>: Machine Learning Methods in Character Recognition</article-title>
          .
          <source>Proc. 13th International Conference on Rough Sets, Fuzzy Sets, Data Mining, and Granular Computing (RSFDGrC</source>
          <year>2011</year>
          ), Lecture Notes in Computer Science 6743, pp.
          <volume>322</volume>
          {
          <issue>329</issue>
          (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Kaytoue</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuznetsov</surname>
            ,
            <given-names>S.O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Napoli</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Duplessis</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Mining gene expression data with pattern structures in Formal Concept Analysis</article-title>
          .
          <source>Inf. Sci</source>
          .
          <volume>181</volume>
          (
          <issue>10</issue>
          ):
          <fpage>1989</fpage>
          -
          <lpage>2001</lpage>
          (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Langley</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          : Practical statistics simply explained.
          <source>Courier Corporation</source>
          (
          <year>1971</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Kuznetsov</surname>
            ,
            <given-names>S.O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Makhalova</surname>
            ,
            <given-names>T.P.</given-names>
          </string-name>
          :
          <article-title>Concept interestingness measures: a comparative study</article-title>
          .
          <source>CLA</source>
          <year>2015</year>
          , pp.
          <volume>59</volume>
          {
          <issue>72</issue>
          (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Luxenburger</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Implications partielles dans un contexte</article-title>
          .
          <source>Mathematiques et Sciences Humaines 113</source>
          , pp.
          <volume>35</volume>
          {
          <issue>55</issue>
          (
          <year>1991</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>MacQueen</surname>
          </string-name>
          , J. et al.:
          <article-title>Some methods for classi cation and analysis of multivariate observations</article-title>
          .
          <source>Proceedings of the fth Berkeley symposium on mathematical statistics and probability 1</source>
          .14, pp.
          <volume>281</volume>
          {
          <issue>297</issue>
          (
          <year>1967</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Murtagh</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Legendre</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Wards hierarchical agglomerative clustering method: Which algorithms implement wards criterion?</article-title>
          .
          <source>Journal of Classi cation 31.3</source>
          , pp.
          <volume>274</volume>
          {
          <issue>295</issue>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Patel</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patel</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patel</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Optical character recognition by open source OCR tool tesseract: A case study</article-title>
          .
          <source>International Journal of Computer Applications</source>
          ,
          <volume>55</volume>
          .10, pp.
          <volume>50</volume>
          {
          <issue>56</issue>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Poelmans</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ignatov</surname>
            ,
            <given-names>D.I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuznetsov</surname>
            ,
            <given-names>S.O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dedene</surname>
          </string-name>
          , G.:
          <article-title>Formal Concept Analysis in knowledge processing: A survey on applications</article-title>
          .
          <source>Expert systems with applications 40</source>
          .16, pp.
          <volume>6538</volume>
          {
          <issue>6560</issue>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <given-names>R</given-names>
            <surname>Core Team</surname>
          </string-name>
          :
          <article-title>A language and environment for statistical computing</article-title>
          .
          <source>R Foundation for Statistical Computing</source>
          , Vienna, Austria.
          <source>ISBN 3-900051-07-0</source>
          , (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Singh</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Optical Character Recognition Techniques: A survey</article-title>
          .
          <source>International Journal of Advanced Research in Computer Engineering &amp; Technology (IJARCET) 4.6</source>
          , pp.
          <volume>545</volume>
          {
          <issue>550</issue>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Sumathi</surname>
            ,
            <given-names>C.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Santhanam</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Devi</surname>
            ,
            <given-names>G.G.</given-names>
          </string-name>
          :
          <article-title>A survey on various approaches of text extraction in images 3.4</article-title>
          , pp.
          <volume>27</volume>
          {
          <issue>42</issue>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Sumathi</surname>
            ,
            <given-names>C.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Santhanam</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Priya</surname>
          </string-name>
          , N.:
          <article-title>Techniques and challenges of automatic text extraction in complex images: A survey</article-title>
          .
          <source>Journal of Theoretical and Applied Information Technology, 35.2</source>
          , pp.
          <fpage>225</fpage>
          -
          <lpage>235</lpage>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31.
          <string-name>
            <surname>Taouil</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pasquier</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bastide</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lakhal</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Mining bases for association rules using closed sets</article-title>
          .
          <source>ICDE2000 International Conference</source>
          , pp.
          <volume>307</volume>
          (
          <year>2000</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          32. van De Vel,
          <string-name>
            <surname>M.L.J</surname>
          </string-name>
          :
          <article-title>Theory of convex structures</article-title>
          .
          <source>Elsevier</source>
          (
          <year>1993</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          33.
          <string-name>
            <surname>Vendramin</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Campello</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hruschka</surname>
          </string-name>
          , E.:
          <article-title>Relative Clustering Validity Criteria: A Comparative Overview</article-title>
          .
          <source>Statistical Analysis and Data Mining 3</source>
          , pp.
          <volume>209</volume>
          {
          <issue>235</issue>
          (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          34.
          <string-name>
            <surname>Visani</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bertet</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ogier</surname>
            ,
            <given-names>J.-M.:</given-names>
          </string-name>
          <article-title>NAVIGALA: An original symbol classi er based on navigation through a galois lattice</article-title>
          .
          <source>International Journal of Pattern Recognition and Arti cial Intelligence</source>
          <volume>25</volume>
          .4, pp.
          <volume>449</volume>
          {
          <issue>473</issue>
          (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          35.
          <string-name>
            <surname>Wille</surname>
          </string-name>
          , R.:
          <source>Restructuring lattice theory: an approach based on hierarchies of concepts</source>
          .
          <source>Springer Netherlands</source>
          (
          <year>1982</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>