<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>ConIText: An Improved Approach for Contextual Indexation of Text Applied to Classification of Large Unstructured Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mohamed Salim El Bazzi</string-name>
          <email>elbazzi.mohamedsalim@edu.uiz.ac.ma</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          ,
          <addr-line>Abdelatif Ennaji</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>IRF-SIC Laboratory, Ibn Zohr University</institution>
          ,
          <addr-line>Agadir</addr-line>
          ,
          <country country="MA">Morocco</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>LITIS Laboratory, University of Rouen</institution>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <fpage>80</fpage>
      <lpage>93</lpage>
      <abstract>
        <p>Managing text documents of a large size can be particularly challenging. For this reason, the indexation process is a crucial and decisive step for retrieving relevant textual features. Therefore, it is essential to design methods that effectively take into account complex data. In this work, we define a new contextual automatic indexation approach. Thus, we present ConIText System, a context-based approach for document indexation and its application. Also, we study its tolerance to large amounts of documents. In other words, we propose a new large corpus of texts and assess the performance gradually from 1.000 to 20.000 documents. It is to observe the behavior of the indexation system while data is getting big. To compare classification results, we used KNN and SVM classifiers. ConIText system outperforms the conventional statistical indexation, based on TFIDF method. Nevertheless, our contextualization system is generic, is not based on external resources. Although we have tested ConIText System on an Arabic dataset, it is not limited to one and unique language.</p>
      </abstract>
      <kwd-group>
        <kwd>ConIText</kwd>
        <kwd>Text Mining</kwd>
        <kwd>Indexation</kwd>
        <kwd>Context</kwd>
        <kwd>Data Analysis</kwd>
        <kwd>Classification</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The indexation of texts is a crucial step in text processing. It allows to represent
documents by their most relevant features. Several approaches are used for this
purpose. However, extracting knowledge from textual data is an important issue,
especially for large amounts of data.</p>
      <p>Consequently, we proposed a contextual approach for the automatic
indexation of texts. Indeed, in order to explore big data and to disclose hidden semantic
information in unstructured documents, such as texts, an efficient indexation
system is required. Consequently, we propose a new approach for text indexation
based on semantic proximity and taking into account the contexts contained in
each document.</p>
      <p>
        This lead to our second proposition of a new approach for document
modeling. To test the performance of our approach, a large corpus was needed.
Therefore, we built our dataset from Arabic online Encyclopedias. It contains 20,000
documents labeled and categorized into 7 classes. The tests will be done gradually
in 1000, 5000, 10000 and 20000 documents to study the robustness of ConIText
system. Once the input is integrated into the system, a certain level of syntax
processing is required. After the preprocessing step, ConIText System identifies
common sentences of each document, and classify them according to their
semantic proximity. Then, the system identify contexts and model the document,
for classification aim [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>We propose, in this work, a new approach for context discovery, based on
sentence clustering techniques. Moreover, we introduce an efficient document
modeling method. This model illustrates the most dominant context in a given
document. Nevertheless, to assess the robustness of our proposed system, we have
conducted experimental tests to compare its results to conventional statistical
methods, as TF IDF.</p>
      <p>The organization of this paper is as follows. In part 2, we introduce related
works. Part 3 details our proposed ConIText System. In part 4, we highlight the
experiments and results. Finally, we conclude by synthesizing the contributions
of this work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related works</title>
      <p>Most of the researches in the field of unsupervised information extraction focuses
on keyword extraction Few of them offer methods to extract the contextual
relations present in a document.</p>
      <p>The contextual approaches aim, on the one hand, to remove the ambiguity
of the meaning of texts. On the other hand, they highlight the semantic
relations between these words. Semantic relationships can also be calculated using
methods that evaluate the quantity of information shared between n-to-n words.</p>
      <p>
        A study of Named Entity Recognition (NER) is presented in [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], for
identifying different classes on NER in social media. They use words similarity besides
several text mining technics for named entity class discovery.
      </p>
      <p>
        A survey of documents clustering using semantic approaches is introduced
in [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. A comparison between LSI, Graph, Ontology, and Lexical Chain is
presented.
      </p>
      <p>
        The authors of [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] propose a text classifier called Supervises Meaning
Classification. They introduced Helwheltz principle to measure meaning. It is about
noticing unexpected events in a particular context. They compare the results to
SVM classifier that has been outperformed.
      </p>
      <p>
        Mohamed and al. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] used LSA method, to evaluate each term in a document,
and then applied an Evidential Reasoning method. It is to attribute the new
document to a category based on the corpus. Experiments showed that ER-LSA
is more efficient than ER-TFIDF.
      </p>
      <p>
        The authors of [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] presents graph medialization for text. The nodes of a
graph correspond to the terms of a document. The relationship between two
nodes represents the semantic relationship between two words. The proposed
approach outperforms the traditional Bag-of-words (BOW) approach.
      </p>
      <p>
        Same for Herskovic [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], they propose MedRank algorithm to reorder the
ranks of the concepts extracted from a medical base. First, the MetaMap
program extracts these concepts. Then, new scores are assigned to the concepts
using the TextRank algorithm. The best results are obtained using the MedRank
approach.
      </p>
      <p>
        In [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], the authors try to classify a large amount of texts. Each text is
modeled by a vector that contains a big number of words. The authors highlight
the importance of selecting the most relevant feature for a classification aim.
This was the goal of their feature extraction algorithm. Also, three different
feature selection methods are used: Information Gain, Correlation and
k-BestDiscriminative-Terms (k-BDT).
      </p>
      <p>
        Multivariate Relative Discriminative Criterion (MRDC) is proposed in [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ],
to perform text classification. First, stopwords removal, stemming and term
weighting are applied before the classification step. Second, a multivariate
features ranking criterion to evaluate features is proposed for text classification.
Then, a subset of features is evaluated using a supervised learning algorithm.
      </p>
      <p>
        The objective of the authors in [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] is dimensionality reduction, without
compromising the performance of a classifier. After forming the document-term
matrix, they apply data mining techniques to solve this problem. Their research
introduces a method for document classification by performing dimensionality
reduction with PCA.
      </p>
      <p>
        The authors of [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] propose fuzzy logic based on a multi-document
summarization system to extract relevant sentences to generate a non-redundant
summary. This approach is based on a generic summarization system.
      </p>
      <p>
        Authors in [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] treat the problem of classification of Arabic text. They use
SVM, NB and MLP-NN algorithms and apply tests on in-house made dataset.
This study aims to apply those algorithms on an Arabic dataset and proceed for
a comparative study. The average measures show that SVM algorithm
outperformed NB and MLP-NN.
      </p>
      <p>
        In [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], authors discuss a proposition to perform text classification using a
space-independent text classification algorithm. This method depends on Markov
chain theory. Each document is represented using a sequence of characters
cooccurrences in the document. Each category of the corpus is used to create a
single probability transition matrix that will be used in the classification process.
      </p>
      <p>
        TF-IDF with dimensionality reduction can improve the precision in the
process of lexical matching, for identification of domain categories referencing to
a document, as proposed in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Higher level of accuracy is possible to perform
based on the reduction approach that can be adopted for documents
classification.
      </p>
      <p>
        The study in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] reports the results of an improved feature selection
algorithm combined with decision three and SVM on text classification. This study
compares the impact of this approach with results of text classification using
Chisquare, Mutual Information, and Gini Index. The results show that ImpCHI and
SVM in text classification outperform the use of Chi-square, MI and GI.
      </p>
      <p>
        The authors of [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] show that the application of TF-IDF-ICF (Term
FrequencyInverse Document Frequency – Inverse Class Frequency) method with
dimensionality reduction technique can be more powerful in precision for classification of
documents.
      </p>
      <p>
        In [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], the authors introduce a method based on continuous distributed
representation of words. The proposed Arabic taxonomy, which is independent of
the model used to classify Arabic questions, provides promising results in Arabic
question classification.
      </p>
      <p>
        Authors in [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] propose a research that raises the effectiveness of unsupervised
learning, semi-supervised learning, and semi-supervised learning with
dimensionality reduction algorithms using k-means, incremental k-means, Threshold
- kmeans and k-means with dimensionality reduction, to calculate the accuracy
of SVM.
      </p>
      <p>Each of the presented methods highlight certain criteria. The approach we
propose takes advantage of the existing advances and introduces a new concept
of contextualization for a more refined indexation process.
3</p>
    </sec>
    <sec id="sec-3">
      <title>ConIText : Contextual Indexation of Text</title>
      <p>In this part, we introduce the architecture of the automatic system for
contextual indexation. It is a set of complex text mining methods, which forms
an autonomous process of extracting contexts, then document features, for a
context-based indexation (Fig. 1).</p>
      <p>In fact, a dataset may contain different categories of documents that are
either homogeneous or heterogeneous. Thus, the categories of politics and
economics can be homogeneous. In contrast, the categories of new technologies and
literature can be heterogeneous. Moreover, a document can express several
contexts. For example, a document that describes a political decision and its impact
on the economy will be difficult to be classed within the appropriate category.
This is the kind of ambiguity that the ConIText System overcomes, by selecting
the most appropriate context for each document.</p>
      <p>We define context as the linguistic environment of a textual element (word,
sequence of words, etc.) within the utterance in which it appears. That is to
say the series of text units that precede and follow it. The term context refers
to all the circumstances in which an act of enunciation takes place, as cultural
and psychological situations, experiences and knowledge of the world, trade and
promotion in economics, etc.</p>
      <p>Furthermore, we define a sentence as the minimal element, which can express
a context. A sentence is a set of words giving a complete meaning. Therefore,
sentences is the first unit detected by ConIText System. Then, sentences close
semantically are gathered to form a context. From each context of a document,
we extract relevant features to form contexts vectors. Finally, the vector
corresponding to the most dominant context models the whole document. To perform
ConIText System, three steps are essential. The first step is the segmentation of
texts. The second step is building context from which the features will be
extracted. Finally, we model documents with our proposed principle of dominance.
3.1</p>
      <sec id="sec-3-1">
        <title>Segmentation</title>
        <p>The works that study the semantic grouping of sentences to extract relevant
information inspire this proposition. As matter of fact, a sentence tends to express
an idea, a context in our case, in an affine way. To go further in our data
processing, the phase of sentences splitting is essential. In fact, words are organized
into sequences, sentences or paragraphs, to define the meaning of the document.
Therefore, the exploration of the relationship between the different components
of the document is important to understand the document in depth.</p>
        <p>Hence, this step consists on defining units of a text that will form the
contexts. Segmentation of the document is the process of dividing the textual
documents into meaningful sentences. Humans naturally understand the sentence
when reading the text. Intuitively, we instill this power of understanding to our
algorithm.</p>
        <p>Texts have markers of explicit sentence boundaries. We use punctuation
marks to delineate a sentence. In this work, we test ConIText system on an
Arabic corpus. Since the Arabic language does not have a capital letters, our
sentence segmentation is based on points, exclamation points and question marks
(".", "!", "?").
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Building contexts</title>
        <p>This phase consists on grouping the sentences obtained in clusters. Each cluster
represents a context. Obviously, each document will have at least one context.
To perform sentence clustering, we used Iterative K-means with a mean square
error metric. This method makes possible to find, for each document, the optimal
k number of the clustering method. Thus, each document is subdivided into k
clusters. Therefore, we group together all the sentences to obtain k contexts in
each document in the corpus (Fig. 2).</p>
        <p>To conceive an efficient clustering process, the weights of words must be
standardized based on their apparition in the document and their distribution
in the entire corpus. In general, a common representation used for text processing
is the TF-IDF representation.</p>
        <p>The TFIDF weighting method is widely used by researchers. It is a frequency
associated to the Vector Space Model (VSM), which involves associating a weight
vector to each document. TF represents the number of occurrences of a word
in the document and IDF is the absolute inverse frequency of the word in the
corpus.</p>
        <p>This method reduces the importance of common terms in the collection while
ensuring that the matching of documents is more influenced by most
discriminating words, which have a relatively high frequency in the document and low
frequencies in the corpus.</p>
        <p>In this work, since the form of the documents has changed, considering the
generated contexts, we have introduced a slight modification for the TFIDF
formula. Named TF-ICF (Term Frequency – Inverse Contexts Frequency), it is
expressed as follows:</p>
        <p>T F</p>
        <p>ICF context(i) (t) = T F context(i)(t)x log(
tf (t)context (i)</p>
        <p>ICF (t)
)</p>
        <p>Where t is a term of the context i, TF context(i) (t) is the frequency of t
within the context(i) and ICF(t) is the occurrence of t in all contexts of the
corpus.</p>
        <p>The obvious advantage of using this method is to calculate the relevance of
a term according to all contexts of the corpus. This induces to express the value
of the terms judged irrelevant in the conventional TFIDF system, whereas they
have a powerful discrimination role in the document.
3.3</p>
      </sec>
      <sec id="sec-3-3">
        <title>Document Modeling with Dominance Principle</title>
        <p>The dilemma in text mining is to select the appropriate representation of the
textual information that will be able to represent the semantic content of the
text. To model the document, we use the VSM representation. After building
contexts, we calculate the score of each word in the context using the TF-ICF
method, in order to model each context by a corresponding vector of words
weights. Thus, each document is represented by a set of vectors (Fig. 3).</p>
        <p>A constraint occurs, it is to represent each document by a single vector of
weight. To perform this step, we define the principle of dominate context. After
the contextualization, each document is divided into one or more contexts. Each
context is modeled by one vector of weights. The Dominant Context is the vector
of the strongest weights. Formally, n vectors model contexts of a given document
D, as:</p>
        <p>V (D) = fM axV i (D); i2[1; k]g</p>
        <p>Where V(D) is the unique vector that models the document D, Vi is the
vector modelling the context i, and k the number of contexts discovered in D.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Data</title>
      <p>During our research, we often face the problem of a lack of significant corpus.
To approve the efficiency of an indexation system, it is essential to test it on a
large amount of data. The best structured corpus are often not open access (see
Table 2).</p>
      <p>In lots of works on text mining, the authors build their own dataset. They
choose the number of categories and themes to use. For each category, the
documents are collected manually and those belonging to several categories are
eliminated. Nonetheless, the size of datasets is relatively small to assess a system
power, and the areas covered are geared towards specific issues.</p>
      <p>This problem led us to create a new labeled corpus, in Arabic language, of
20,000 texts, with 27,605,263 words after document pretreatment (stemming and
stopwords removing), and 7 classes labeled as presented in Table 3:</p>
      <p>This corpus is collected from Arabic encyclopedias, and have the particularity
of containing homogeneous themes and other heterogeneous to better assess the
precision of systems. We make this data freely available to researchers.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Experiments</title>
      <p>In this work, we proposed ConIText System, a context-based system for
automatic text indexation. To test our approach, we opted for a large dataset
to evaluate its robustness and reliability. We have tested the proposed system
gradually on a large amount of data (Table 3).</p>
      <p>For comparative reasons, we conducted the tests using two classifiers, KNN
and SVM. The different models of these classifiers will show the tolerance of our
indexation approach to classification systems.</p>
      <p>The experimental evaluation of the classifier is the final step in the
classification process. It usually tries to evaluate the effectiveness of a classifier, namely
its ability to make categorization decisions.</p>
      <p>Fig. 4 and Fig. 5 show the results of KNN and SVM classification of
documents using TFIDF and ConIText systems. These results are expressed by the
f measure. The performance of our system is obvious. This is due to the
complexity of the techniques used for the context-based indexation. However, the
classification parameters are stationary for both classifiers.</p>
      <p>The strong point that can be drawn from these experiments is the behavior of
ConIText on a wide range of documents including more than 10000 documents.
This advantage is more visible in the following figures.</p>
    </sec>
    <sec id="sec-6">
      <title>Discussion</title>
      <p>The first test was performed on 1000 documents, which is a proportion widely
used in the literature. Then, we have passed the test on 5000 documents. This
number represents the maximum of the documents used in similar works (Table
2). We have pushed the tests on 10,000 documents to see how the system will
react with such a mass of documents. Finally, we have performed a test on 20000
documents to study the stability of the system.</p>
      <p>Table 4 and 5 presents the results of the comparison between TFIDF
system and ConIText system, expressed by recall, precision and F-measure. Table 4
presents the result using KNN classifier and Table 5, SVM Classifier. In
particular, those results show the relevance of using contextual indexing that effectively
improves classification performance.</p>
      <p>The results are illustrated in Fig. 6 and Fig. 7 in terms of precision and recall
of classifiers KNN and SVM. The curves indicate that the performance of the
TFIDF method drops dramatically as soon as the database takes more and more
documents. However, ConITexte’s results are not only more satisfying but also
tolerable to large datasets. We can see clearly that the performances are almost
constant between 10,000 and 20,000 documents. This enhances the effectiveness
of the context-based indexation system and confirms our theory.</p>
      <p>We can deduce many conclusions from our experimental results. First, the
contextual model showed its performance to be the appropriate representation
for large datasets. Indeed, the context has several advantages over which it is
possible to act to refine the extraction of the keywords. Second, the relations
between words are expressed by maintaining the shared information of the context.
This will certainly lead to better results.
7</p>
    </sec>
    <sec id="sec-7">
      <title>Conclusion</title>
      <p>In this paper, we have introduced ConIText, a contextual indexation System
for texts. The integration of a semantic measure between sentences in this
approach is necessary. For this reason, we have introduced our sentence grouping
contribution to formalize the adaptation of the model to the semantic
proximity. The advantage of this approach is that it does not need any preliminary
specific knowledge to identify terms in order to assign them weights since the
identification of terms is done from an automatic document processing.</p>
      <p>We also proposed contextual modeling for document to increase the accuracy
of indexation. In fact, the semantic proximity between words must be
emphasized when we are dealing with complex and unstructured documents such as
texts. For this reason, it is essential to broaden our thinking to models of
representation adapted to the nature of our resources. To this end, we have introduced
a contextual modeling for documents based on the principle of dominance. The
advantage of this model is that it reduces the space representation of features
and reduce the whole modeling of a document to its most significant context.
8</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgements</title>
      <p>This work was funded by LITIS laboratory, and the University of Rouen
Normandy, France.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Al-Anzi</surname>
            ,
            <given-names>F.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>AbuZeina</surname>
          </string-name>
          , D.:
          <article-title>Beyond vector space model for hierarchical arabic text classification: A markov chain approach</article-title>
          .
          <source>Information Processing &amp; Management</source>
          <volume>54</volume>
          (
          <issue>1</issue>
          ),
          <fpage>105</fpage>
          -
          <lpage>115</lpage>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Alami</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Meknassi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ouatik</surname>
            ,
            <given-names>S.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ennahnahi</surname>
          </string-name>
          , N.:
          <article-title>Impact of stemming on arabic text summarization</article-title>
          .
          <source>In: 2016 4th IEEE International Colloquium on Information Science and Technology (CiSt)</source>
          . pp.
          <fpage>338</fpage>
          -
          <lpage>343</lpage>
          . IEEE (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bahassine</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Madani</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Al-Sarem</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kissi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Feature selection using an improved chi-square for arabic text classification</article-title>
          .
          <source>Journal of King Saud UniversityComputer and Information Sciences</source>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Dhar</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dash</surname>
            ,
            <given-names>N.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roy</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Application of tf-idf feature for categorizing documents of online bangla web text corpus</article-title>
          .
          <source>In: Intelligent Engineering Informatics</source>
          , pp.
          <fpage>51</fpage>
          -
          <lpage>59</lpage>
          . Springer (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Dhar</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dash</surname>
            ,
            <given-names>N.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roy</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Categorization of bangla web text documents based on tf-idf-icf text analysis scheme</article-title>
          .
          <source>In: Annual Convention of the Computer Society of India</source>
          . pp.
          <fpage>477</fpage>
          -
          <lpage>484</lpage>
          . Springer (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>El</given-names>
            <surname>Bazzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.S.</given-names>
            ,
            <surname>Mammass</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Ennaji</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Zaki</surname>
          </string-name>
          ,
          <string-name>
            <surname>T.</surname>
          </string-name>
          :
          <article-title>Toward a complex system for context discovery to index arabic documents</article-title>
          .
          <source>JCP</source>
          <volume>13</volume>
          (
          <issue>8</issue>
          ),
          <fpage>955</fpage>
          -
          <lpage>962</lpage>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Ganiz</surname>
            ,
            <given-names>M.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tutkan</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Akyokuş</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>A novel classifier based on meaning for text classification</article-title>
          .
          <source>In: 2015 International Symposium on Innovations in Intelligent SysTems and Applications (INISTA)</source>
          . pp.
          <fpage>1</fpage>
          -
          <lpage>5</lpage>
          . IEEE (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Gonçalves</surname>
            ,
            <given-names>C.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Iglesias</surname>
            ,
            <given-names>E.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Borrajo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Camacho</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vieira</surname>
            ,
            <given-names>A.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gonçalves</surname>
          </string-name>
          , C.T.:
          <article-title>Comparative study of feature selection methods for medical full text classification</article-title>
          .
          <source>In: International Work-Conference on Bioinformatics and Biomedical Engineering</source>
          . pp.
          <fpage>550</fpage>
          -
          <lpage>560</lpage>
          . Springer (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Hamza</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>En-Nahnahi</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zidani</surname>
            ,
            <given-names>K.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ouatik</surname>
            ,
            <given-names>S.E.A.</given-names>
          </string-name>
          :
          <article-title>An arabic question classification method based on new taxonomy and continuous distributed representation of words</article-title>
          .
          <source>Journal of King</source>
          Saud University-Computer and Information Sciences (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Herskovic</surname>
            ,
            <given-names>J.R.</given-names>
          </string-name>
          , Cohen,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Subramanian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Iyengar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.S.</given-names>
            ,
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.W.</given-names>
            ,
            <surname>Bernstam</surname>
          </string-name>
          ,
          <string-name>
            <surname>E.V.</surname>
          </string-name>
          :
          <article-title>Medrank: Using graph-based concept ranking to index biomedical texts</article-title>
          .
          <source>International journal of medical informatics 80(6)</source>
          ,
          <fpage>431</fpage>
          -
          <lpage>441</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Kumar</surname>
            ,
            <given-names>B.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ravi</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Text document classification with pca and one-class svm</article-title>
          .
          <source>In: Proceedings of the 5th International Conference on Frontiers in Intelligent Computing: Theory and Applications</source>
          . pp.
          <fpage>107</fpage>
          -
          <lpage>115</lpage>
          . Springer (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Labani</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moradi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ahmadizar</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jalili</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>A novel multivariate filter method for feature selection in text classification problems</article-title>
          .
          <source>Engineering Applications of Artificial Intelligence</source>
          <volume>70</volume>
          ,
          <fpage>25</fpage>
          -
          <lpage>37</lpage>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Mohamed</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Watada</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>An evidential reasoning based lsa approach to document classification for knowledge acquisition</article-title>
          .
          <source>In: 2010 IEEE International Conference on Industrial Engineering and Engineering Management</source>
          . pp.
          <fpage>1092</fpage>
          -
          <lpage>1096</lpage>
          . IEEE (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Mohammad</surname>
            ,
            <given-names>A.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alwada'n</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Al-Momani</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Arabic text categorization using support vector machine, naïve bayes and neural network</article-title>
          .
          <source>GSTF Journal on Computing (JoC) 5</source>
          (
          <issue>1</issue>
          ),
          <volume>108</volume>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Patel</surname>
            ,
            <given-names>D.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shah</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chhinkaniwala</surname>
            ,
            <given-names>H.R.:</given-names>
          </string-name>
          <article-title>Fuzzy logic based multi document summarization with improved sentence scoring and redundancy removal technique</article-title>
          .
          <source>Expert Systems with Applications</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Saiyad</surname>
            ,
            <given-names>N.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prajapati</surname>
            ,
            <given-names>H.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dabhi</surname>
            ,
            <given-names>V.K.</given-names>
          </string-name>
          :
          <article-title>A survey of document clustering using semantic approach</article-title>
          . In: 2016 International Conference on Electrical,
          <source>Electronics, and Optimization Techniques (ICEEOT)</source>
          . pp.
          <fpage>2555</fpage>
          -
          <lpage>2562</lpage>
          . IEEE (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Sangaiah</surname>
            ,
            <given-names>A.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fakhry</surname>
            ,
            <given-names>A.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Abdel-Basset</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>El-henawy</surname>
          </string-name>
          , I.:
          <article-title>Arabic text clustering using improved clustering algorithms with dimensionality reduction</article-title>
          .
          <source>Cluster</source>
          Computing pp.
          <fpage>1</fpage>
          -
          <lpage>15</lpage>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Taşpınar</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ganiz</surname>
            ,
            <given-names>M.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Acarman</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>A feature based simple machine learning approach with word embeddings to named entity recognition on tweets</article-title>
          .
          <source>In: International Conference on Applications of Natural Language to Information Systems</source>
          . pp.
          <fpage>254</fpage>
          -
          <lpage>259</lpage>
          . Springer (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>