<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Feature selection and graph representation for an analysis of science fields evolution: an application to the digital library ISTEX</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jean-Charles LAMIREL</string-name>
          <email>jean-charles.lamirel@loria.fr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pascal CUXAC</string-name>
          <email>pascal.cuxac@inist.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Inist-CNRS</institution>
          ,
          <addr-line>2,allée du parc de Brabois, 54519 Vandoeuvre-lès-Nancy</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>SYNALP Team-LORIA, INRIA Nancy-Grand Est</institution>
          ,
          <addr-line>Vandoeuvre-lès-Nancy</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <fpage>88</fpage>
      <lpage>99</lpage>
      <abstract>
        <p>This paper presents an original approach based on a recent metric called feature maximization for developing accurate diachronic analysis tools. In such process, querying of bibliographic databases is firstly exploited to provide a thematic corpus of scientific publications covering a large time period. In a second step, two strategies based on contrast graphs generated by the use of feature maximization metric are proposed. The first one is based on the direct use contrast graphs who relates time periods and publication contents. The second strategy combines a preliminary step of clustering with the use of contrast graph generated by feature maximization applied on cluster contents to highlight the relation between topics represented in clusters as well as to embed them in a temporal path. Both techniques are parameter-free and knowledge agnostic. We illustrate the efficiency and the complementarity of the proposed technique by experimenting then on a dataset related to gerontology research extracted from the data collected by the ISTEX project, a project whose aims is to construct a general purpose and open access database of scientific documents.</p>
      </abstract>
      <kwd-group>
        <kwd>Feature selection</kwd>
        <kwd>graph-based approach</kwd>
        <kwd>diachronic analysis</kwd>
        <kwd>visualization</kwd>
        <kwd>big data management</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The ISTEX1 project’s main objective is to provide the whole French higher education
and research community with on-line access to retrospective collections of scientific
literature in all disciplines.</p>
      <p>On the basis of the initial platform services, we are currently working towards
proposing new added-value services. One of our central concern is then to develop tools
for highlighting the dynamics of the collection. Hence, the development of dynamic
information analysis methods, like incremental clustering and novelty detection
techniques, is becoming a central concern in a bunch of applications whose main goal is to
deal with large volume of textual information whose content is varying over time, such
as ISTEX. The purpose of the analysis and diachronic mapping is to track, for a given
domain, changes in contexts (sub-themes) and the evolution of vocabularies and actors
http://www.istex.fr/
that materialize these changes in terms of appearances, disappearances, divergence or
convergence. The applications relate to very various and highly strategic domains,
including web mining, technological and scientific survey.</p>
      <p>In order to identify and analyze the emergence, or to detect changes in the data, we
have previously proposed two different and complementary approaches:
1. Performing static classifications at different periods of time and analyzing changes
between these periods (time-step approach or diachronic analysis);
2. Developing methods of classification that can directly track the changes:
incremental clustering methods (incremental clustering) and novelty detection methods
(incremental supervised classification).</p>
      <p>
        The development of direct methods being still an ongoing research, we present
hereafter two original word-based methods relying on the first approach and using a metric
called feature maximization we have recently developed
        <xref ref-type="bibr" rid="ref1">(Lamirel and al. 2013)</xref>
        . The
goal of the two methods that are based on contrast graphs derived from this metric is to
tackle with document belonging to the same scientific field in order to detect significant
topic differences between documents related to different time periods:
1. Unlike common approaches based on graph analysis
        <xref ref-type="bibr" rid="ref11">(Porter and Rafols 2009)</xref>
        <xref ref-type="bibr" rid="ref12">(Sayama and Akaishi 2012)</xref>
        , our first approach is a supervised approach that
establishes a bipartite contrast graph between documents time stamps and documents
salient keywords, those latter being extracted through a feature selection process
based on feature maximization.
2. Our second approach is an unsupervised approach based on clustering. Thanks to
this approach optimal number of clusters (i.e. topics) is extracted from the whole
document dataset and relations between extracted topics and selected salient
keywords are used to form the bipartite contrast graph. Documents timestamps are
exploited in a second step to highlight diachronic changes and diachronic path
between topics.
      </p>
      <p>We first present our feature maximization metrics and related contrast graph
exploited throughout our approach. In a next step, we describe our experimental data and
associated preprocessing. Lastly, we highlight our results with the two proposed
approaches and our conclusion.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Feature maximization, feature selection and contrast graph</title>
      <sec id="sec-2-1">
        <title>2.1 Feature maximization</title>
        <p>
          Feature maximization (F-max) is an unbiased cluster quality metrics that exploits the
properties of the data associated to each cluster without prior consideration of clusters
profiles. This metrics has been initially proposed in
          <xref ref-type="bibr" rid="ref5">(Lamirel and al 2004)</xref>
          . Its main
advantage is to be independent altogether of the clustering methods and of their
operating mode.
        </p>
        <p>Consider a partition  which results from a clustering method applied to a dataset 
represented by a group of features  . The feature F-measure   ( ) of a feature 
associated with a cluster  is defined as the harmonic mean of the feature recall 
 ( )
and of the feature predominance   ( ), which are themselves defined as follows:
with
  ( ) =
  ( ) =</p>
        <p>∈   
 ∈  ∈   
 ∈</p>
        <p>′∈</p>
        <p>,∈   ′
  ( ) = 2 (  ( )×  ( ))
  ( )+  ( )
(1)
(2)
(3)
where    represents the weight of the feature  for the data  and   represents all the
features present in the dataset associated with the cluster  . Feature Predominance
measures the ability of  to describe cluster  . In a complementary way, Feature Recall
allows to characterize  according to its ability to discriminate  from other clusters.</p>
        <p>
          Feature recall is a scale independent measure but feature predominance is not. We
have however shown experimentally in
          <xref ref-type="bibr" rid="ref8">(Lamirel, Cuxac, et al. 2015)</xref>
          that the
F-measure which is a combination of these two measures is only weakly influenced by feature
scaling. Nevertheless, to guaranty full scale independent behavior for this measure, data
must be standardized. Furthermore, the choice of the weighting scheme for data is not
really constrained by the approach, but it is necessary to deal with positive values. Such
scheme is supposed to figure out the significance (i.e. semantic and importance) of the
feature for the data2.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2 Feature selection</title>
        <p>
          In supervised context, feature maximization measure can be exploited to generate a
powerful feature selection process
          <xref ref-type="bibr" rid="ref8">(Lamirel, Cuxac, et al. 2015)</xref>
          . In our unsupervised
(clustering) context, the selection process can be used to describe or label clusters
according to the most typical and representative features. This process is a
non-parametrized process that uses both the capacity of F-measure to discriminate between
clusters (  ( ) index) and its ability to faithfully represent the cluster data (  ( ) index).
The set   of features that are characteristic of a given cluster  belonging to a partition
 is translated by:

 = { ∈   ∨   ( ) &gt; 
( ) ∧   ( ) &gt;   }
(5)
2
part.
        </p>
        <p>A feature having some negative values can be separated in 2 different positive sub-features,
the first one representing the positive part of original feature and the second one, its negative
(5)
(6)
(7)


 =∪∈</p>
        <p>() =  ′∈
 ′()
| / |

  =  ∈
 ()
||
where   represents the subset of  in which the feature  occurs.</p>
        <p>Finally, the set of all selected features   is the subset of  defined by:
In other words, the features judged relevant for a given cluster are those whose
representations are better than average in this cluster, and better than the average
representation of all the features in the partition, in terms of Feature F-measure. Features which
never respect the second condition in any cluster are discarded.
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Contrast</title>
        <p>A specific concept of contrast   ( ) can be defined to calculate the performance of a
retained feature  for a given cluster  . It is an indicator value which is proportional to
the ratio between the F-measure   ( ) of a feature in the cluster  and the average
Fmeasure</p>
        <p>of this feature for the whole partition.Contrast of a feature  for a cluster
 is expressed as:
  ( ) =   ( )⁄</p>
        <p>( )
The active features of a cluster are those for which the contrast is greater than 1.
Moreover, the higher the contrast of a feature for one cluster, the better its performance in
describing the cluster content.
2.4</p>
      </sec>
      <sec id="sec-2-4">
        <title>Contrast graphs</title>
        <p>
          In the mathematical field of graph theory, a bipartite graph (or bigraph) is a graph whose
vertices can be divided into two disjoint and independent sets U and V such that every
edge connects a vertex in U to one in V. Contrast graphs are bipartite graphs based on
the relations between a set of features S and a set of labels L
          <xref ref-type="bibr" rid="ref1">(Cuxac and Lamirel 2013)</xref>
          .
Theoretically, the set of labels L could represent any kind of information to which
features can be related with and the set of features S is a subset of a global feature set F
(i.e. he original feature space on which rely the data of a dataset) that has been obtained
through a feature selection process, like feature maximization presented above. In the
case of the use of feature maximization, the weight  (, ) of an edge (, 
),  ∈ ,  ∈
 represents the contrast of feature u for a label v as, it is defined by equation 7.
        </p>
        <p>Such kind of graphs have many interesting properties. First, they reduce the
cognitive overload produced with classical graphs representation because of the associated
feature selection process that reduces the number of potential connections. Second, they
can be used to indirectly highlight relationships between labels, whenever features have
contrasted interaction with several labels. Third, the combination of this approach with
weighted force-directed model (Kobourov 2012) for graph representation permits
altogether to highlight central or most influent labels of the L set and to easily identify the
labels that are the most densely connected through associated features, these latter
appearing in close neighborhood position in the graph.</p>
        <p>
          We have proposed a first original use of contrast graph in the case of the analysis of
the transdisciplinarity between different research domains and time periods in
          <xref ref-type="bibr" rid="ref1">(Cuxac
and Lamirel 2013)</xref>
          .
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Data</title>
      <p>Our experimental data is a collection of 9801 scientific papers in English language
related to gerontology domain published between 1995 and 2010, extracted from ISTEX
database by INIST documentary engineers specialized in the medical domain. After a
tokenization step, the keywords are extracted from the abstracts by a part-of-speech
method developed in Python. However, the NLP treatments are minimalist (just
morpho-lexical and syntactic) and we thus don’t use any vocabulary resource except a stop
word dictionary.</p>
      <p>We present on the following experimental section the two different approaches we
have applied on the extracted metadata of our dataset, that are, the GRAFSEL approach
which is a supervised approach based on a direct exploitation of the relations between
document content extracted form keywords and document publication year to build a
contrast graph and the CLUSTSEL approach that exploit a clustering process on the
extracted document content and build up a contrast graph highlighting relation between
cluster content with a further use of document publication years to highlight diachronic
changes.
4
4.1</p>
    </sec>
    <sec id="sec-4">
      <title>Experimental results</title>
      <sec id="sec-4-1">
        <title>The GRAFSEL approach</title>
        <p>To clarify the principle of our supervised approach, we named GRAFSEL, we follow
three steps that are schematically presented in figure 1:
1. The papers of the experimental dataset are assigned to a class that represents their
publication year;
2. The papers being represented by their extracted keywords, we select keywords
related to each year and compute the strength of the relations (i.e. the contrast)
between selected keywords and years exploiting the feature maximization metric
shortly described in section 3;
3. The last step is to build the graph highlighting the relationships between the years
and the selected keywords by weighting the links of the graph with the formerly
obtained contrast values.</p>
        <p>Figure 2a shows the global result obtained by taking into account altogether the whole
corpus and the fifteen considered years. All of the following graphs are obtained with
a force-directed algorithm (Spring algorithm) taking into account as weight of a
yearword edge the contrast of the word for the considered year.</p>
        <p>We perform hereafter an attempt of interpretation of our results, mostly to illustrate
the potential of the method. Such attempt is not substituting an in-deep expert validation
that is planned in a near future.</p>
        <p>As show in the figure 2, which is a zoom on the 2000s, we can observe in the years
2003-2006 the emergence of terms like "nurse”, “nursing”, “home care”, “medicare”,
“family caregivers”, “home”, ”satisfaction”, ... that denote the development of home
help services to maintain autonomy of elderly people.</p>
        <p>In a complementary way, we can also exploit radar chart representation for each year
in order to facilitate the interpretation of the results. In such representation each
extracted word is a radius of a circle whose length depends on its contrast value whilst
the description space is the same for all considered years allowing in such a way to
detect changes.</p>
        <p>Considering all the above-mentioned years, it is possible to detect invariant
directions during each related period (figure 3). Furthermore, sudden changes of said
directions suggest new scientific domains but also changes in professionals’ practices.
As we have formerly observed in figures 2, the direction corresponding to “nursing”,
“home”, “care” appears in the year 2002-2003, and terms as “risk”, “cancer”,
”mortality” and thereafter “exposure”, “stress” emerge in the years 2006. Additionally,
if the first years were solely marked by the term “women”, the use of “people”,
“personality” in years 2003-2006 indicates that a humanization of care might appear.
On its own side, the term “mice” is often used until 2001 and disappear after: it might
figure out an indicator of changes in experimentation protocols.</p>
        <p>This short discussion shows that the use complementary modes of representation
obviously enables a quick and simple view of the evolution of a thematic corpus
through time.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>The CLUSTSEL approach</title>
        <p>The overall principle of our unsupervised approach, we named CLUSTSEL is presented
in figure 4:
1. The papers being represented their extracted keywords, we cluster their descriptions
using a clustering algorithm. Several experiments are achieved by varying the
number of expected clusters and consequently obtaining clustering models of various
sizes.
2. A clustering quality measurement based on feature maximization is exploited to
find out the optimal model among all the ones that have been generated.
3. The clusters of the optimal model being represented by the keywords extracted from
their associated papers, we select keywords related to each cluster and compute the
strength of the relations (i.e. the contrast) between selected keywords and clusters
exploiting the feature maximization metric shortly described in section 3;
4. The graph highlighting the relationships between the clusters (i.e. the topics) and
the selected keywords is built by weighting its links with the formerly obtained
contrast values;
5. Papers publication years are used to find out dominant period of the topics as well
as to build up a diachronic chart figuring out the comparative influence of topics
during each year.</p>
        <p>
          For clustering, we exploit 2 different usual clustering methods, namely k-means
          <xref ref-type="bibr" rid="ref10">(MacQueen 1967)</xref>
          , a winner-take-all method, and GNG (Fritske 1995), a
winner-takemost method with Hebbian learning. The GNG method proved to be superior to
kmeans method because of (altogether) Hebbian,incremental and winner take-most
learning process providing better independence to initial conditions and avoiding
producing degenerated clustering results. Similar results have been already reported in
          <xref ref-type="bibr" rid="ref6">(Lamirel, Mall, et al. 2011)</xref>
          . The selection of optimal model relies on feature
maximization metrics presented in the former section. Our former experiments on reference
datasets show that most of the usual quality estimators do not produce satisfactory
results in a realistic data context, are sensitive to noise and perform poorly with high
dimensional data
          <xref ref-type="bibr" rid="ref4">(Kassab and Lamirel 2008)</xref>
          . A more accurate method is thus to exploit
feature maximization and more especially information related to the activity and
passivity of selected features in clusters to define clustering quality indexes identifying an
optimal partition. This kind of partition is expected to maximize the contrast described
by eq. 7. The method is more precisely detailed in
          <xref ref-type="bibr" rid="ref9">(Lamirel, Dugué, et al. 2016)</xref>
          .
        </p>
        <p>In the specific case of your experiment we propose to build up a contrast graph
between a set of clusters representing the main research topics of the domain that have
been extracted by the clustering process and the most contrasted features issued from
the clusters’ descriptions. This approach that combines clustering and contrast graph in
an original way highlighting the most connected topics.</p>
        <p>In the case of our experiment we focus on one type of external labels that are papers’
publication years. Papers’ publication years are exploited to perform a diachronic
analysis of the topics’ activity, highlighting the importance of each topic in each time
period, either this activity is considered individually (see figure 7) or relatively to the
other topics (see figure 6). As it is shown in the next part related to the analysis of the
results, this approach helps to precisely understand the chronology of the research
activity of a global research domain, like in our specific case.</p>
        <p>In the context of our dataset we obtained an optimal model comprising 12 clusters
(i.e. topics). The spatial distribution of 12 topics presented on the graph of figure 6
highlights clearly interpretable structure of the domain. Such graph provides generic
although detailed representation of the domain-related research topics whilst
highlighting the main relationships between the said topics whenever those topics appears as
close neighbors on the graph. As an example, the ‘Homecare” topic is directly related
to logically connected topics like “Physical performance” and “Health condition”.
Similarly, “Menopause related problems”, an early topic, appears to be accurately related
to “Cancer studies” and “Gene senescence” that figure out more recent and more
general research topics.</p>
        <p>On its own side, diachronic representations that are presented in figure 6-7 can then
be used to get a better understanding of the gerontology development from early
research (“Menopause related problems”, Hearing loss, “Age change” general studies) to
more up to date research (“Home care”, “Risks factors”, “Physical performance”,
“Health condition”, “Sociology of health”) that fits well with the global changes
regarding health politics. In that context research on “Physical performance” becomes the
most prominent in the recent years and seems thus clearly represent a central focus
because of its obvious influence on the other recent research areas.</p>
        <p>Last but not least, research on “Neurodegenerative diseases” (Alzheimer,
Parkinson, …) seems to have split into two parts by generating a new specialized area related
to “Memory performance”.</p>
        <p>In such a way, results provide by CLUSTSEL approach appears clearly
complementary to the ones obtained by the GRAFSEL approach. Hence, the two methods
provide similar results although they highlight those ones with different levels of
generality.</p>
        <p>We have presented an original overall methodology for the diachronic analysis of
large and heterogeneous text collections based on feature maximization and associated
contrast graphs. The originality of that approach comes from the fact that the nodes of
the obtained graphs result from the combination of a feature selection processes and a
classification or a clustering process, depending on the chosen option. Thus, one main
advantage of the approach is to avoid cognitive overload in the current case of
management of high dimensional data. Another of its main advantage is to be altogether
parameter-free and knowledge/language-agnostic. Our first experimental results
obtained from the analysis of a realistic dataset extracted from the ISTEX bibliographic
database are promising. Hence, they prove to be easily interpretable by an expert of the
analyzed domain. Moreover, the supervised and unsupervised options of our approach
provide similar results that can be considered of different levels of generality.</p>
        <p>One further and encouraging domain of investigation would concern to check the
scalability of our approach to the context of massive data analysis.
6</p>
        <p>Acknowledgments</p>
        <p>ISTEX receives assistance from the French state managed by the National Research
Agency under the program "Future Investments" bearing the reference
ANR-10-IDEX0004-12.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Cuxac</surname>
            <given-names>P.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Lamirel</surname>
            <given-names>J.C.</given-names>
          </string-name>
          (
          <year>2013</year>
          ).
          <article-title>Analysis of evolutions and interactions between science fields: the cooperation between feature selection and graph representation</article-title>
          .
          <source>14th COLLNET Meeting, August 15-17</source>
          ,
          <year>2013</year>
          Tartu, Estonia
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Dubey</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ho</surname>
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Williamson S</surname>
          </string-name>
          . and
          <string-name>
            <surname>Xing</surname>
            <given-names>E. P.</given-names>
          </string-name>
          (
          <year>2014</year>
          ),
          <article-title>Dependent nonparametric trees for dynamic hierarchical clustering</article-title>
          ,
          <source>NIPS</source>
          <year>2014</year>
          :
          <fpage>1152</fpage>
          -
          <lpage>1160</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Fritzke</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          (
          <year>1995</year>
          ).
          <article-title>A growing neural gas network learns topologies</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          (pp.
          <fpage>625</fpage>
          -
          <lpage>632</lpage>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Kassab</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Lamirel</surname>
            ,
            <given-names>J.-C.</given-names>
          </string-name>
          (
          <year>2008</year>
          ).
          <article-title>Feature-based cluster validation for high-dimensional data</article-title>
          .
          <source>In Proceedings of the 26th IASTED International Conference on Artificial Intelligence and Applications</source>
          (pp.
          <fpage>232</fpage>
          -
          <lpage>239</lpage>
          ). ACTA Press.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Lamirel</surname>
            ,
            <given-names>J.-C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Al Shehabi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Francois</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Hoffmann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2004</year>
          ).
          <article-title>New classification quality estimators for analysis of documentary information: application to patent analysis and web mapping</article-title>
          ,
          <source>Scientometrics</source>
          , vol.
          <volume>60</volume>
          , n°
          <issue>3</issue>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Lamirel</surname>
          </string-name>
          , J.-C.,
          <string-name>
            <surname>Mall</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cuxac</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Safi</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          (
          <year>2011</year>
          ).
          <article-title>Variations to incremental growing neural gas algorithm based on label maximization</article-title>
          .
          <source>In Neural Networks (IJCNN)</source>
          , The 2011 International Joint Conference on (pp.
          <fpage>956</fpage>
          -
          <lpage>965</lpage>
          ). IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Lamirel</surname>
            ,
            <given-names>J.-C.</given-names>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>A new approach for automatizing the analysis of research topics dynamics: application to optoelectronics research Scientometrics (</article-title>
          <year>2012</year>
          )
          <volume>93</volume>
          :
          <fpage>151</fpage>
          -
          <lpage>166</lpage>
          ,
          <year>October 01</year>
          ,
          <year>2012</year>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Lamirel</surname>
          </string-name>
          , J.-C.,
          <string-name>
            <surname>Cuxac</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chivukula</surname>
            ,
            <given-names>A. S.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Hajlaoui</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>Optimizing text classification through efficient feature selection based on quality metric</article-title>
          .
          <source>Journal of Intelligent Information Systems</source>
          ,
          <volume>45</volume>
          (
          <issue>3</issue>
          ),
          <fpage>379</fpage>
          -
          <lpage>396</lpage>
          . doi:
          <volume>10</volume>
          .1007/s10844-014-0317-4
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Lamirel</surname>
          </string-name>
          , J.-C.,
          <string-name>
            <surname>Dugué</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Cuxac</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          (
          <year>2016</year>
          ).
          <article-title>New efficient clustering quality indexes</article-title>
          .
          <source>In Neural Networks (IJCNN)</source>
          , 2016 International Joint Conference on (pp.
          <fpage>3649</fpage>
          -
          <lpage>3657</lpage>
          ). IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>MacQueen</surname>
          </string-name>
          , J. (
          <year>1967</year>
          ).
          <article-title>Some methods for classification and analysis of multivariate observations</article-title>
          .
          <source>In Proceedings of the fifth Berkeley symposium on mathematical statistics and probability (Vol. 1</source>
          , pp.
          <fpage>281</fpage>
          -
          <lpage>297</lpage>
          ). Oakland, CA, USA.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Porter</surname>
            ,
            <given-names>A. L.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Rafols</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          (
          <year>2009</year>
          )
          <article-title>Is science becoming more interdisciplinary? Measuring and mapping six research fields over time</article-title>
          ,
          <source>Scientometrics</source>
          , vol.
          <volume>81</volume>
          , no 3, p.
          <fpage>719</fpage>
          -
          <lpage>745</lpage>
          ,
          <year>2009</year>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Sayama</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <article-title>and</article-title>
          <string-name>
            <surname>Akaishi</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2012</year>
          )
          <article-title>Characterizing Interdisciplinarity of Researchers and Research Topics Using Web Search Engines, Plos One</article-title>
          , vol.
          <volume>7</volume>
          , no 6, p.
          <fpage>e38747</fpage>
          ,
          <year>2012</year>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>