<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Factorial Correspondence Analysis Applied to Citation Contexts</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Marc Bertin</string-name>
          <email>bertin.marc@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Iana Atanassova</string-name>
          <email>iana.atanassova@univ-fcomte.fr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Centre Interuniversitaire de Rercherche sur la Science et la Technologie (CIRST), Universite du Quebec a Montreal (UQAM)</institution>
          ,
          <country country="CA">Canada</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Centre Tesniere, University of Franche-Comte</institution>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper, we analyze citation contexts and characterize the di erent sections of scienti c articles in terms of the verbs that appear in citation contexts. We have performed Factorial Correspondence Analysis (CA) using the four sections of the IMRaD (Introduction, Methods, Results and Discussion) structure as categories. Our dataset contains about 80,000 research articles published in the six PLOS journals. The results of this approach show that the sections in the rhetorical structure of research articles have very di erent characteristics when we take into consideration the occurrences of verbs, and more generally, their lexical content. Our results demonstrate a strong relation between verbs used around citations and the positions in the rhetorical structure.</p>
      </abstract>
      <kwd-group>
        <kwd>Factorial Correspondence Analysis</kwd>
        <kwd>Bibliometrics</kwd>
        <kwd>Citation Analysis</kwd>
        <kwd>Information Retrieval</kwd>
        <kwd>Citation Context Analysis</kwd>
        <kwd>IMRaD</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The study of in-text citations is a very old subject, but it still remains an
active area of research. This problem interests many researchers and covers several
di erent disciplines and research areas: Bibliometrics, Information science,
Sociology of science, Computer science. Despite numerous studies, there do not yet
exist automated approaches for the analysis of in-text citations, mainly because
of the di culties in constructing a complete model of the function of citations.
As Cronin [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] points out, the need to have a citation theory and citation context
analysis has already been investigated (e.g. [
        <xref ref-type="bibr" rid="ref11 ref14 ref15">11, 14, 15</xref>
        ]). But to this day, we still
do not have a model explaining the actions of citations. This task is extremely
complex because a lot of factors need to be taken into consideration.
      </p>
      <p>
        In this paper, we address this problem from the point of view of textual
statistics. Recent works have shown that the distribution of in-text citations in
scienti c articles are strongly correlated to their IMRaD structure (see [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]). In
addition, other studies investigate this issue using a lexically-based approach.
The study of verbs found in citation contexts is an important step towards a
better de nition of the meaning of citations acts (see [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]). One of the results
presented in this work is that the rhetorical sections in scienti c papers do not
have the same status according the verb frequency distributions. This means
that the position of a text segment in the rhetorical structure governs to a large
extent the lexical items (verbs) that are used by the author.
      </p>
      <p>In this paper, we report on a set of experiments to classify sentences
containing in-text citations, according to their position in the rhetorical structure. We
believe that the study of citation contexts implies observing the use of citations
in a certain amount of textual data. In this exploratory study, we will focus on
this perspective that seems relevant for the understanding of citation acts.</p>
      <p>
        For this purpose, our study uses a multivariate statistical method, namely
Correspondence Factorial Analysis (CFA) (see [
        <xref ref-type="bibr" rid="ref2 ref3 ref9">9, 2, 3</xref>
        ]) to propose an analysis
of a dataset of about 8,000 textual contexts of bibliographical references
(intext citations). Correspondence analysis is a technical description of contingency
tables and is mainly used in the eld of text mining (e.g. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]).
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Method</title>
      <p>In this study, we investigate the relationship that exists between the rhetorical
structure of papers and the text structure and more speci cally the lexicon. By
analysing occurrence frequencies of di erent lexical items, the main objective of
this method is to achieve an optimal projection of the multidimensional system
on a factorial plot. The hypothesis that we want to verify can be formulated
as follows: the lexical forms present in the contexts of in-text citations are not
randomly distributed. They are strongly dependent on the particular positions
of the rhetorical structure.
2.1</p>
      <sec id="sec-2-1">
        <title>Dataset</title>
        <p>Our dataset consists of six peer-reviewed academic journals published in Open
Access by the Public Library of Science (PLOS): six domain-speci c journals
(PLOS Biology, PLOS Computational Biology, PLOS Genetics, PLOS Medicine,
PLOS Neglected Tropical Diseases ) and PLOS ONE, a general journal that
covers all elds of science and social sciences. We have processed the entire dataset
of about 80,000 research articles published up to September 2013.</p>
        <p>We have identi ed the section structure in each article by analyzing the
section titles. All six journals use similar publication models, where authors are
explicitly encouraged to follow the IMRaD (Introduction, Methods, Results, and
Discussion) structure. As a result, more than 97% of all research articles contain
these four section types, although not always in the same order.</p>
        <p>We have identi ed and extracted all textual segments that contain in-text
citations. To do this, the text was segmented into sentences and for each of the
four sections we have considered the set of sentences containing in-text citations.
As a result, we have obtained a total of 3,314,884 sentences, 31.52% out of
which belong to Introduction sections, 19.50% to Methods, 14.66% to Results
and 34.33% to Discussion sections.</p>
        <p>Next, we propose to use these sets to examine the characteristics of
citations in the di erent sections and the ways citations are used according to their
position in the rhetorical structure of articles.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Protocol</title>
        <p>
          We study the sets of sentences from sections that are identi ed according to the
IMRaD structure and the presence of in-text citations. For this purpose, we use
two text analysis tools. The rst one is an R Commander plugin (see temis [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ])
which provides integrated tools of text mining tasks. Corpora can be imported in
raw text. The second one, is a python application called IRaMuTeQ [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] which
uses the R libraries. These tools were used in order to produce the outputs for
the correspondence analysis and tables.
        </p>
        <p>The set of sentences have been split into words and lemmatized: all di erent
forms of a lexical item are identi ed and associated with the same lexical item.
IRaMuTeQ performs stemming from dictionaries, without disambiguation, also
called endogenous lemmatization. After lemmatization, we have ltered all verb
forms and ranked them by occurrence frequency for each section. This allowed us
to produce a map displaying proximity among variables (Lexical vs Rhetorical
Structure).</p>
        <p>To perform the analysis, we have created a subset of sentences that consists
of about 2,000 randomly extracted sentences for each section of the rhetorical
structure and for each journal. This amounts to a total of 48,000 sentences for
this analysis. As shown in Table 1, this corpus contains 47,714 unique terms,
that have 1,569,201 occurrences.</p>
        <sec id="sec-2-2-1">
          <title>Number of terms</title>
        </sec>
        <sec id="sec-2-2-2">
          <title>Number of unique terms</title>
        </sec>
        <sec id="sec-2-2-3">
          <title>Percent of unique terms</title>
        </sec>
        <sec id="sec-2-2-4">
          <title>Number of hapax legomena</title>
        </sec>
        <sec id="sec-2-2-5">
          <title>Percent of hapax legomena</title>
        </sec>
        <sec id="sec-2-2-6">
          <title>Number of words</title>
        </sec>
        <sec id="sec-2-2-7">
          <title>Number of long words</title>
        </sec>
        <sec id="sec-2-2-8">
          <title>Percent of long words</title>
        </sec>
        <sec id="sec-2-2-9">
          <title>Number of very long words</title>
        </sec>
        <sec id="sec-2-2-10">
          <title>Percent of very long words</title>
        </sec>
        <sec id="sec-2-2-11">
          <title>Average word length</title>
        </sec>
        <sec id="sec-2-2-12">
          <title>Introduction</title>
        </sec>
        <sec id="sec-2-2-13">
          <title>Methods</title>
        </sec>
        <sec id="sec-2-2-14">
          <title>Results Discussion</title>
          <p>Total
327,506
21,389</p>
          <p>6.5
8,809</p>
          <p>2.7
327,506
124,618</p>
          <p>38.1
43,499
13.3
5.7
295,798 337,379
22,848 22,316</p>
          <p>7.7 6.6
10,539 9,326</p>
          <p>3.6 2.8
295,798 337,379
102,422 114,858</p>
          <p>34.6 34.0
34,032 38,668
11.5 11.5
5.5 5.4
608,518 1,569,201
28,694 47,714</p>
          <p>4.7 3.0
11,689 18,082</p>
          <p>1.9 1.2
608,518 1,569,201
222,373 564,271</p>
          <p>36.5 36.0
76,932 193,131
12.6 12.3
5.6 5.5
We have performed Factorial Correspondence Analysis (CA) using the four
sections of the IMRaD structure as categories. The interpretation of this analysis
is a graph representation of associations between rows and columns. Columns
express the sections while the rows correspond to all forms of occurrences. As
word meanings strongly depend on their contexts, we consider only sentences
containing in-text citations, thus limiting the possible ambiguities.</p>
          <p>The interpretation of an axis in CA in a linguistic context is de ned by
the opposition between the extreme points. Figure 1 presents the projection of
the four sections. This gure shows that, for example, the Methods section is
in opposition to all other sections. Similarly, the Results and the Introduction
sections are opposed to each other on the vertical axis.</p>
          <p>Figure 2 presents the projection of a sample of the most frequent verbs. For
the same verbs, table 2 presents the values for the relative frequencies in the
di erent sections: Introduction - Methods - Results - Discussion. If certain verbs
are more or less homogeneous among the sections, some of the verbs, as observed,
are only predominantly found in the Discussion and Results sections.</p>
          <p>For example, we can see on gure 2 that the verbs performs and calculate
(see i) are mainly present in the Methods section. For the same verbs, table 2
indicates the values [perform - 52.16] and [calculate - 22.47]. The position of the
verbs cause and review (see ii) on gure 2 show that they are characteristic
of the Introduction section. This is con rmed by the values given in table 2,
respectively [case - 16.79] and [review - 14.06]. Another example are the verbs
analyse
approach
assume
calculate (i)
cause (ii)
carry
characterize
compare
con rm (iii)
consider
contrast
de ne
demonstrate
detect
determine
develop
encode
establish
examine
expect (iii)
identify
include
indicate
involve
mean
measure
note
observe
obtain
perform (i)
predict
present
propose
provide
regulate
remain
represent
reveal
review (ii)
study
suggest
6.38
9.18
3.7
1.32
9.35
2.98
4.38
11.22
5.74
5.87
9.18
4.93
19.3
10.2
6.04
9.18
7.91
5.61
4.04
5.57
18.02
25.67
8.33
14.71
4.72
7.86
7.18
27.03
4.76
3.87
8.8
15.6
12.45
11.01
12.37
5.65
6.8
6.89
7.95
91.72
40.16
6.33
11.47
3.09
1.18
16.79
4.87
7.96
8.19
2.96
5.6
6.19
7.19
13.52
7.69
7.46
15.61
12.29
6.51
3.6
2.96
23.16
32.86
5.55
17.93
3.23
7.24</p>
          <p>2
11.29
4.32
3.82
7.46
11.83
10.06
11.92
12.42
9.01
5.64
7.15
14.06
70.54
25.67
con rm and expect (see iii) that mainly belong to the Results section. Their
values for this section in table 2 are [expect - 7.96] and [con rm - 9.01].</p>
          <p>The results of this approach show that the sections in the rhetorical
structure of research articles have very di erent characteristics when we take into
consideration the occurrences of verbs, and more generally, their lexical content.
Our results demonstrate a strong relation between verbs used around citations
and the positions in the rhetorical structure. In addition, gure 1 shows some
proximity between the Results and the Discussion sections, as well as between
the Discussion and the Introduction sections.
4</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Conclusion</title>
      <p>
        This study con rms the results of previous work (see [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]) around the lexical
analysis of citation contexts. It also shows that citation contexts are strongly
dependent on the rhetorical structure and this is an important factor for the
analysis of citation contexts.
      </p>
      <p>
        The results can also be considered in the perspective of other studies on the
distribution of references [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], according to which the distribution of in-text
citations is strongly related to the rhetorical structure of articles. They show, for
example, the great speci city of the Methods section, because it has a relatively
low frequency of in-text citations. In addition, by considering the most frequent
verbs in the di erent sections, our results imply functionality contexts which
are speci c to the rhetorical structure. Indeed, this study shows that citation
contexts in the di erent sections can be characterized in terms of their
lexical content, and more speci cally the verbs that appear near in-text citations.
Inversely, the function of citations is strongly related to the rhetorical
structure and the position of the citation in the article. Taking into consideration
the rhetorical structure is therefore necessary for the analysis of citation acts.
By studying the di erent verbs that are present in citation contexts and their
relation to the rhetorical structure, we will be able to determine the semantic
relations that authors use when they cite other work.
      </p>
      <p>
        The results of our study have numerous applications, especially in
Information Retrieval ([
        <xref ref-type="bibr" rid="ref10 ref8">8, 10</xref>
        ]) and Bibliometrics (e.g. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]). The next step is to improve
Information Extraction and the analysis of citation networks. Taking into
account these results and analyzing more closely the characteristics of citation
contexts is an essential step in the understanding the functions of citations and
citation acts.
5
      </p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgments</title>
      <p>We thank Benoit Macaluso of the Observatoire des Sciences et des Technologies
(OST), Montreal, Canada, for harvesting and providing the PLOS data set.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bastin</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bouchet-Valat</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>RcmdrPlugin. temis, a Graphical Integrated Text Mining Solution in R</article-title>
          . The
          <source>R Journal</source>
          <volume>5</volume>
          (
          <issue>1</issue>
          ),
          <volume>188</volume>
          {
          <fpage>196</fpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Benzecri</surname>
            ,
            <given-names>J.P.:</given-names>
          </string-name>
          <article-title>L'analyse des donnees: L'analyse des correspondances</article-title>
          .
          <source>Dunod</source>
          (
          <year>1973</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Benzecri</surname>
            ,
            <given-names>J.P.</given-names>
          </string-name>
          :
          <article-title>Correspondence Analysis Handbook. (translated from: Pratique de l'analyse des donnees, 1</article-title>
          . Expose elementaire. Dunod, Paris). (
          <year>1992</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Bertin</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Atanassova</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>A study of lexical distribution in citation contexts through the IMRaD standard</article-title>
          .
          <source>In: Proceedings of the First Workshop on Bibliometric-enhanced Information Retrieval co-located with 36th European Conference on Information Retrieval (ECIR</source>
          <year>2014</year>
          ). pp.
          <volume>5</volume>
          {
          <fpage>12</fpage>
          . Amsterdam,
          <source>The Netherlands (April 13</source>
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Bertin</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Atanassova</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lariviere</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gingras</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>The distribution of references in scienti c papers: an analysis of the imrad structure</article-title>
          .
          <source>In: 14th International Society of Scientometrics and Informatics Conference. International Society for Scientometrics and Infometrics</source>
          , Vienna, Austria (
          <volume>15</volume>
          -
          <issue>19th</issue>
          <year>July 2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Cronin</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>The need for a theory of citing</article-title>
          .
          <source>Journal of Documentation</source>
          <volume>37</volume>
          (
          <issue>1</issue>
          ),
          <volume>16</volume>
          {
          <fpage>24</fpage>
          (
          <year>1981</year>
          ), http://dx.doi.org/10.1108/eb026703
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Cronin</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Bibliometrics and beyond: some thoughts on web-based citation analysis</article-title>
          .
          <source>Journal of Information science 27(1)</source>
          , 1{
          <issue>7</issue>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8. Glanzel, W.:
          <article-title>Bibliometrics-aided retrieval: where information retrieval meets scientometrics</article-title>
          .
          <source>Scientometrics</source>
          <volume>102</volume>
          ,
          <issue>2215</issue>
          {
          <fpage>2222</fpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Hirschfeld</surname>
            ,
            <given-names>H.O.:</given-names>
          </string-name>
          <article-title>A connection between correlation and contingency</article-title>
          .
          <source>In: Mathematical Proceedings of the Cambridge Philosophical Society</source>
          . vol.
          <volume>31</volume>
          , pp.
          <volume>520</volume>
          {
          <fpage>524</fpage>
          . Cambridge Univ Press (
          <year>1935</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Mayr</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Scharnhorst</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Scientometrics and information retrieval: weak-links revitalized</article-title>
          .
          <source>Scientometrics</source>
          <volume>102</volume>
          (
          <issue>3</issue>
          ),
          <volume>2193</volume>
          {
          <fpage>2199</fpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Moravcsik</surname>
            ,
            <given-names>M.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Murugesan</surname>
            ,
            <given-names>P.:</given-names>
          </string-name>
          <article-title>Some results on the function and quality of citations</article-title>
          .
          <source>Social studies of science 5(1)</source>
          ,
          <volume>86</volume>
          {
          <fpage>92</fpage>
          (
          <year>1975</year>
          ), http://sss.sagepub.com/content/5/1/86.full.pdf
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Morin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Intensive use of factorial correspondence analysis for text mining: application with statistical education publications</article-title>
          .
          <source>Statistics Educational Research Journal ( SERJ</source>
          ) pp.
          <volume>1</volume>
          {
          <issue>6</issue>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Ratinaud</surname>
          </string-name>
          , P.: IRaMuTeQ:Interface de R pour les Analyses Multidimensionnelles de Textes et de Questionnaires (
          <year>2009</year>
          ), http://www.iramuteq.org
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Small</surname>
          </string-name>
          , H.:
          <article-title>Citation Context Analysis</article-title>
          .
          <source>Progress in Communication Sciences</source>
          <volume>3</volume>
          ,
          <issue>287</issue>
          {
          <fpage>310</fpage>
          (
          <year>1982</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>White</surname>
          </string-name>
          , H.D.:
          <article-title>Citation analysis and discourse analysis revisited</article-title>
          .
          <source>Applied linguistics 25(1)</source>
          ,
          <volume>89</volume>
          {
          <fpage>116</fpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>