<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>InTeReC: In-text Reference Corpus for Applying Natural Language Processing to Bibliometrics</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Marc Bertin</string-name>
          <email>marc.bertin@univ-lyon1.fr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Iana Atanassova</string-name>
          <email>iana.atanassova@univ-fcomte.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CRIT-Centre Tesnière, Université de Bourgogne Franche-Comté</institution>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>ELICO Laboratory</institution>
          ,
          <addr-line>Université Claude Bernard Lyon 1</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <fpage>54</fpage>
      <lpage>62</lpage>
      <abstract>
        <p>Bibliometrics is more and more interested in the full text processing and the study of the structure of scientific papers. The contexts of in-text references present in articles are particularly relevant for such studies. This work describes the construction of the InTeReC dataset, which is an in-text reference corpus that aims to promote experimental reproducibility and to provide a standard dataset for further research. The InTeReC dataset is a set of sentences containing in-text references together with all the data necessary for their recontextualization in papers using standard CSV format. This should encourage the implementation of natural language processing tools for Bibliometric studies and related research in information retrieval and visualization.</p>
      </abstract>
      <kwd-group>
        <kwd>Bibliometrics</kwd>
        <kwd>Citation Analysis</kwd>
        <kwd>Citation Context Analysis</kwd>
        <kwd>Information Analysis</kwd>
        <kwd>Natural Language Processing</kwd>
        <kwd>IMRaD</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>The assumption that the contexts of the bibliographic references present in a
scientific article play an important role in characterizing the relationship between
citing works and cited works have been accepted for many decades.
Publications are connected to each other by citations and citations contexts categorize
the semantic relations that exist between them. Whether the study of citation
contexts relies mainly on linguistic clues or machine learning techniques,
citation contexts for each research experiment need to be extracted from scientific
corpora. Also, the extraction of citation contexts is a preliminary step to any
statistical, distributional, syntactic or semantic analysis.</p>
      <p>
        Sentences containing in-text references may contain relevant information
about the cited research and cited author’s research areas. Recently, He and
Chen [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] provides an approach to understanding citation contexts which
characterizes the complex roles of a publication. If we are interested in the intellectual
structure of a discipline, the analysis of co-citations has been widely studied.
However, recent work highlights the interest of taking into account the full text
and more precisely the paragraphs [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] of papers. Other approaches are
interested in extracting information from publications and adding semantic attributes
to in-text references that can be defined as traditional. For example, Parinov [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]
focuses on papers’ references, in-text references and citation contexts with the
purpose to visualize citations relationships, their semantic attributes and related
statistics as annotations.
      </p>
      <p>
        In this paper, we propose a large scale dataset of citation contexts and
explain the methods that were used for its construction. Other similar initiatives
exist with various objectives. For example, the ESWC-14 Challenge: Semantic
Publishing3 – Assessing the Quality of Scientific Output (see [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]) focus on the
extraction and assessment of workshop proceedings.The recent activity of
research based on full text and the analysis of in-text references has lead to a race
in the size of datasets. If we look at the size of the corpora used by the different
actors of our community, we can see that the values are very heterogeneous. The
corpus for the CL-SciSumm task4 deal with automatic paper summarization in
Computational Linguistics and is extracted from the ACL Anthology corpus and
its citing papers [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. In 2017, Hu et al. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] worked with 350 articles from Journal
of Informetrics. In 2013 Ding et al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] analyse 866 articles from JASIST. The
largest study of in-text reference distributions in 2016 was proposed by Bertin
et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] who analysed 45,000 papers published in the PLOS journals. Recently,
Boyack et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] focused on the PubMed Central Open Access Subset and
Elsevier journals with five million full text records for in-text reference analysis.
It is clear that we observe an increase in the size of the textual data but also
a methodological evolution in the processing capabilities, advocating for larger
datasets and the use of statistical tools. In general, the construction of a corpus
is a heavy task and requires skills and means that can be important.
      </p>
      <p>
        In this paper we describe the creation of the InTeReC dataset [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], which is
a corpus of sentences containing in-text references extraction from papers
published by PLOS. This dataset is available at https://zenodo.org/record/1203737.
Here we will not detail or define what a corpus is. For this we can refer for
example to Lüdeling and Kytö (see [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]).
      </p>
      <p>The aim of this paper is to present the method of the construction of the
InTeReC corpus, which is a standard in-text reference corpus taking into account
the different elements relevant to the implementation of experimental protocols.
The overall objective is to facilitate citation context analyses and various
distributional analyses by providing a large dataset to the community. The InTeReC
dataset also serves the purposes of reproducibility, interoperability and
cumulative research.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Method</title>
      <p>The construction of the InTeReC dataset is based on several analyzes and
experiments carried out in the recent years. In this section we summarize the methods
that were used in order to propose a dataset that is reusable by the community
for studying citation contexts.</p>
      <sec id="sec-2-1">
        <title>3 http://challenges.2014.eswc-conferences.org/index.php/SemPub</title>
      </sec>
      <sec id="sec-2-2">
        <title>4 http://wing.comp.nus.edu.sg/ cl-scisumm2018/</title>
        <p>
          Working with the full text of papers, we first classify the section titles in
order to identify the four major section types in the IMRaD sequence
(Introduction, Methods, Results and Discussion). This categorization aims to verify
the coherence of the corpus with the IMRaD structure. In many articles, the
four section types exist but not in the same order. For the InTeReC dataset, we
focused only on paper that follows the IMRaD structure, i.e. papers that contain
the four section types in the correct order. We then process the text content of
all paragraphs and segment them into sentences. In our approach, sentences are
considered as the basic textual units and are used to express the positions of
references in the article and in the section. This approach allows for example
to assign relative positions of all references and to obtain the distribution of
references along the text [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. Finally, we count the number of references in each
sentence. The InTeReC dataset contains only sentences are to have one single
in-text reference.
        </p>
        <p>The links between the in-text references and the cited papers or bibliography
items are preserved throughout the processing.
2.1</p>
        <sec id="sec-2-2-1">
          <title>Data: source and structure</title>
          <p>For this corpus, we have used the entire set of research articles published by
PLOS5 up to September 2013. This initial corpus contains 90,071 articles.</p>
          <p>As these 7 journals follow the same publication model but are in different
scientific fields, our aim is to observe the different uses of bibliographic references
in these fields and their relation to the structure of the articles.</p>
          <p>PLOS provides access to the articles in the XML format. The set of XML
elements and attributes that are used for the representation of journal articles are
known as Journal Article Tag Suite (JATS), which is an application of
Z39.962012. Technology evolves quickly and we have to take into consideration that
JATS is a continuation of the NLM Archiving and Interchange DTD works by
NCBI6. As this format is also used by PubMed, this work can easily be extended
to processing the PubMed Open Access Subset which is a larger dataset. The
JATS structure of an article consists of three main elements: front – body – back,
and the textual content of the article is in the body element. It is further
divided into sections and paragraphs. The front element contains some traditional
metadata fields (title, authors, etc.) as well as the article type.</p>
          <p>Different types of articles are present in the corpus, such as "Research
article", "Synopsis", "Primer", "Essay", and the typology is given in the article’s
metadata. We have focused on the "Research article" type, obtaining a total of
85,660 articles out of the initial 90,071 articles in the corpus.</p>
        </sec>
      </sec>
      <sec id="sec-2-3">
        <title>5 Founded in 2001, the Public Library of Science (PLOS) is an Open Access pub</title>
        <p>lisher of seven peer-reviewed academic journals, mostly in the fields of Biology and
Medicine. PLOS ONE, the publishers’ general journal covers, however, all fields of
science and social sciences.</p>
      </sec>
      <sec id="sec-2-4">
        <title>6 http://dtd.nlm.nih.gov</title>
        <p>2.2</p>
        <sec id="sec-2-4-1">
          <title>Segmentation and section title processing</title>
          <p>One of our objectives is to identify the rhetorical structure of the articles. The
use of the IMRaD sequence (Introduction, Methods, Results and Discussion) is
part of the editorial requirements of the PLOS journals and the large majority
of articles include these four sections. In some articles however, the sections are
not always in the same order.</p>
          <p>
            Sections are represented as separate elements in the original XML files. The
research articles in the corpus contain a total of 404 311 sections. We categorized
them automatically by analyzing the section titles in order to match the existing
sections with one of the section types in the IMRaD structure [
            <xref ref-type="bibr" rid="ref1 ref2">1,2</xref>
            ]. In fact,
variations can exist in the ways authors choose to title the sections, e.g. the
Methods section can have titles such as "Materials and Methods", "Method and
Model", etc. We have constructed a set of regular expressions in order to classify
the sections automatically. Table 1 presents some basic statistics of the result of
this classification. The last two classes, (MR) and (RD), appear in some articles
where one section merges two of the main section types and thus the article
contains only three main sections.
          </p>
          <p>Class Section type
I Introduction
M Methods
R Results
D Discussion
(MR) Methods and Results
(RD) Results and Discussion
Total</p>
          <p>We further restricted the set of sentences to be included in the InTeReC
dataset by selecting only sentences that have a single citation and that contain
at least one occurrence of the most frequent verbs that have been attested in
citation contexts. These steps are explained in the following subsections.</p>
          <p>Each paragraph was segmented into sentences by analyzing the punctuation
of the text following a set of typographic rules. All the occurrences of symbols
denoting sentence boundaries (point, exclamation mark, etc.) were examined and
disambiguated. In fact, the occurrence of a point in a text does not necessarily
mean a sentence end, because in many cases it can be part of an abbreviation,
references, genus species, numeric values, etc. We used a set of finite-state
automata in order to determine the contexts in which the points signal sentence
ends.
2.3</p>
        </sec>
        <sec id="sec-2-4-2">
          <title>Article structures</title>
          <p>Once we classified the sections, we examined the sequence of sections present
in each article. To produce the InTeReC dataset, we focused only on articles
where the order of the four sections is: "I,M,R,D". Considering merged sections,
there are three possible article structures, that are listed in table 2. The last two
columns of this table give the total number of sentences in the articles and the
number of sentences that contain at least one in-text reference.</p>
          <p>Article structure Articles Sentences Sentences with references
I,M,R,D 44,370 7,656,518 1,704,326
I,M,(RD) 2,971 504,246 113,237
I,(MR),D 28 5,300 937
Total 47,369 8,166,064 1,818,500</p>
          <p>The following processing was done on these 47,369 articles from which was
selected the InTeReC dataset.
2.4</p>
        </sec>
        <sec id="sec-2-4-3">
          <title>Reference processing</title>
          <p>Our algorithm examines each sentence and counts the number of references
present in the text. In fact, the input data is in the XML format where the
references are represented in &lt;xref&gt; tags. Our algorithm covers all possible
typographic variations for reference ranges and infers the missing data from the
input XML. As a result we obtain the list of sentences in the text, where to
each sentence we have associated a reference count as well as a list of reference
identifiers corresponding to the bibliography entries.</p>
          <p>
            We note that counting the &lt;xref&gt; tags are not a reliable method to obtain
the reference counts, especially if one is interested in multiple in-text references
(MIR) [
            <xref ref-type="bibr" rid="ref5">5</xref>
            ]. When in-text references are in a numeric form, reference ranges are
often present in sentences containing MIR. For example, in-text references such
as ”[
            <xref ref-type="bibr" rid="ref16">16</xref>
            ]–[24]” are represented by two xref elements that point to the
corresponding bibliography items, while in fact there are 9 different citations, 7 of which
are not present in the XML markup. In order to identify correctly MIR and their
number in sentences it is important to detect in-text reference ranges.
          </p>
          <p>For this first version of the InTeReC dataset we have chosen to include only
sentences that contain one single reference. These citation contexts establish
links between only two works, the cited work and the citing article, and thus
we can consider them as the simplest cases to study in terms of citation context
analysis.
2.5</p>
        </sec>
        <sec id="sec-2-4-4">
          <title>Part-Of-Speech tagging and verb phrase extraction</title>
          <p>
            A series of experiments have been published around verbs occurring in citation
contexts and their distributions [
            <xref ref-type="bibr" rid="ref3 ref4">3,4</xref>
            ]. In general, verbs give important
information about the nature of the relation between the article and the cited work.
Polysemy is one possible problem when dealing with verbs, but in our case this
phenomenon is reduced as we work specifically on citation contexts. Our
hypothesis is that the semantic meaning of the relation that exists between the cited
work and the citing article is often expressed, to some extent, by the verb phrase
in the sentence containing the in-text reference. For this reason, we examined
verb phrases that appear in the sentences that contain citations and included
them in the dataset.
          </p>
          <p>
            Bertin et al. [
            <xref ref-type="bibr" rid="ref3 ref4">3,4</xref>
            ]. published lists of most frequent verbs that appear in
citation contexts with respect to section types in the IMRaD structure. We
considered the most frequent verbs in each section that are given on table 3.
show use include
suggest identify find
require associate involve
lead perform follow
obtain generate base
determine contain calculate
carry report observe
express see
          </p>
          <p>
            All sentences were processed using the Part-Of-Speech tagger of python
NLTK7 [
            <xref ref-type="bibr" rid="ref8">8</xref>
            ]. In the output verb forms are tagged by labels such as VB, VBD,
VBG, VBN, VBP, VBZ that stands for base form, past tense, present participle,
etc. We then identified verb phrases by producing parse trees using a simple
grammar.
          </p>
          <p>The InTeReC dataset contains sentences that contain occurrences of the verbs
in table 3 and their verb phrases have been identified. By keeping only sentences
that contain these verbs, we eliminate many sentences that contain perfunctory
citations because they only mention the cited work without explicitely
identifying the its relation with the article. Thus we obtained the final set of 314,023
sentences for the dataset.
3</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>InTeReC dataset structure</title>
      <p>The InTeReC dataset contains a list of sentences in full text. Information is
given on the position of the sentences relative to the article and the section</p>
      <sec id="sec-3-1">
        <title>7 https://www.nltk.org</title>
        <p>in which they appear, the section type with respect to the four main types of
the IMRaD structure, as well as verb phrases that occur in the sentence. Each
sentence contains one single in-text reference.</p>
        <p>
          The dataset is published in the CSV format [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], with the following column
list:
journal: journal title
doi: DOI of the article from which the sentence was extracted
article-length: size of the article, as number of sentences
article-pos: position of the sentence in the article, as number of sentences from
the beginning of the article
section-length: size of the section, as number of sentences
section-pos: position of the sentence in the section, as number of sentences
from the beginning of the section
section-type: section type (one of: I, M, R, D, MR, RD)
sentence-text: full text of the sentence
verb-phrases: a list of verb phrases that occur in the sentence, comma
separated
        </p>
        <p>The format of the dataset has been chosen to facilitate the exploitation of the
data and make it compatible and easily reusable for most types of processing.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Discussion and Conclusion</title>
      <p>Although this corpus takes many aspects into account, it is not exhaustive and
has characteristics that underline the inherent limitations of this approach. For
example, we have limited this work to level of the sentence, and thus do not take
into account relevant citation contexts that span across sentence boundaries
through the use of anaphora.</p>
      <p>Many improvements are possible and they are currently receiving our
attention in the evolution of our corpus. The first concerns the quantitative aspect
with the extension of the corpus to new sources. For this, we can take into
account for example PubMed8, arXiv9 and the CEUR Workshop Proceedings10.</p>
      <p>
        The second evolution concerns a more qualitative aspect of the dataset and
adding semantic annotations. The main idea is to implement semantic annotation
in order to provide an training dataset for supervised learning tools. For this
purpose, we investigate tools to annotate segments with precision values close to
those of human annotators. Indeed, one challenge would be proposing semantic
annotations for ontologies such as CiTO (see [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]).
      </p>
      <p>Finally, the InTeReC dataset can be accessed and visualized using R-shiny11
interface in order to provide users a way to interact with the data and observe
the distributional phenomena.</p>
      <sec id="sec-4-1">
        <title>8 www.nlm.nih.gov/databases/download/pubmed_medline.html</title>
      </sec>
      <sec id="sec-4-2">
        <title>9 https://arxiv.org 10 http://ceur-ws.org 11 https://shiny.rstudio.com</title>
        <p>This dataset aims to facilitate the reproducibility of future research on in-text
citation analysis and thus provide a common foundation for the development of
a unified model of citation context analysis.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <p>We thank Benoit Macaluso of the Observatoire des Sciences et des Technologies
(OST), Montreal, Canada, for harvesting and providing the PLOS dataset.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Atanassova</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bertin</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Semantic Web Evaluation Challenge: SemWebEval 2014 at ESWC 2014, chap</article-title>
          .
          <source>Semantic Facets for Scientific Information Retrieval</source>
          , pp.
          <fpage>108</fpage>
          -
          <lpage>113</lpage>
          . Communications in
          <source>Computer and Information Science (Book 475)</source>
          , Springer, Anissaras, Crete,
          <source>Greece (May</source>
          <volume>25</volume>
          -29
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bertin</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Atanassova</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Semantic Web Evaluation Challenge: SemWebEval 2014 at ESWC 2014, chap</article-title>
          .
          <source>Extraction and Characterization of Citations in Scientific Papers</source>
          , pp.
          <fpage>120</fpage>
          -
          <lpage>128</lpage>
          . Communications in
          <source>Computer and Information Science (Book 475)</source>
          , Springer, Anissaras, Crete,
          <source>Greece (May</source>
          <volume>25</volume>
          -29
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bertin</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Atanassova</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>A study of lexical distribution in citation contexts through the imrad standard</article-title>
          .
          <source>In: Proceedings of the First Workshop on Bibliometric-enhanced Information Retrieval co-located with 36th European Conference on Information Retrieval (ECIR</source>
          <year>2014</year>
          ). vol.
          <volume>1143</volume>
          , pp.
          <fpage>5</fpage>
          -
          <lpage>12</lpage>
          . CEUR Workshop Proceedings, Amsterdam,
          <source>The Netherlands (April 13</source>
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Bertin</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Atanassova</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Factorial correspondence analysis applied to citation contexts</article-title>
          .
          <source>In: Proceedings of the First Workshop on Bibliometric-enhanced Information Retrieval co-located with 37th European Conference on Information Retrieval (ECIR</source>
          <year>2015</year>
          ). Vienna,
          <source>Austria (March 29</source>
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Bertin</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Atanassova</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Multiple in-text reference aggregation phenomenon</article-title>
          .
          <source>In: Proceedings of the 3rd Workshop on Bibliometric-enhanced Information Retrieval co-located with 38th European Conference on Information Retrieval (ECIR</source>
          <year>2016</year>
          ). pp.
          <fpage>14</fpage>
          -
          <lpage>22</lpage>
          . Padua,
          <string-name>
            <surname>Italy</surname>
          </string-name>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Bertin</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Atanassova</surname>
          </string-name>
          , I.:
          <article-title>InTeReC: In-text Reference Corpus - Single References Dataset (Mar</article-title>
          <year>2018</year>
          ), https://doi.org/10.5281/zenodo.1203737
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Bertin</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Atanassova</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gingras</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Larivière</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>The invariant distribution of references in scientific articles</article-title>
          .
          <source>Journal of the Association for Information Science and Technology</source>
          <volume>67</volume>
          (
          <issue>1</issue>
          ),
          <fpage>164</fpage>
          -
          <lpage>177</lpage>
          (
          <year>2016</year>
          ), http://dx.doi.org/10.1002/asi.23367
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Bird</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klein</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Loper</surname>
            ,
            <given-names>E.: Natural</given-names>
          </string-name>
          <string-name>
            <surname>Language Processing with Python. O'Reilly Media</surname>
          </string-name>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Boyack</surname>
            , K.W., van Eck,
            <given-names>N.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Colavizza</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Waltman</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Characterizing intext citations in scientific articles: A large-scale analysis</article-title>
          .
          <source>Journal of Informetrics</source>
          <volume>12</volume>
          (
          <issue>1</issue>
          ),
          <fpage>59</fpage>
          -
          <lpage>73</lpage>
          (
          <year>2018</year>
          ), http://www.sciencedirect.com/science/article/pii/ S1751157717303516
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Ding</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guo</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cronin</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>The distribution of references across texts: Some implications for citation analysis</article-title>
          .
          <source>Journal of Informetrics</source>
          <volume>7</volume>
          (
          <issue>3</issue>
          ),
          <fpage>583</fpage>
          -
          <lpage>592</lpage>
          (
          <year>2013</year>
          ), http://www.sciencedirect.com/science/article/pii/ S1751157713000230
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Dragoni</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Solanki</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blomqvist</surname>
          </string-name>
          , E.:
          <source>Semantic Web Challenges: 4th SemWebEval Challenge at ESWC</source>
          <year>2017</year>
          , Portoroz, Slovenia, May 28-June 1,
          <year>2017</year>
          , Revised Selected Papers, vol.
          <volume>769</volume>
          . Springer (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>He</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Understanding the changing roles of scientific publications via citation embeddings</article-title>
          .
          <source>arXiv preprint arXiv:1711.05822</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Hsiao</surname>
            ,
            <given-names>T.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <year>h</year>
          .:
          <article-title>Yet another method for author co-citation analysis: A new approach based on paragraph similarity</article-title>
          .
          <source>Proceedings of the Association for Information Science and Technology</source>
          <volume>54</volume>
          (
          <issue>1</issue>
          ),
          <fpage>170</fpage>
          -
          <lpage>178</lpage>
          (
          <year>2017</year>
          ), http://dx.doi.org/ 10.1002/pra2.
          <year>2017</year>
          .14505401019
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Hu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hou</surname>
          </string-name>
          , H.:
          <article-title>Understanding multiply mentioned references</article-title>
          .
          <source>Journal of Informetrics</source>
          <volume>11</volume>
          (
          <issue>4</issue>
          ),
          <fpage>948</fpage>
          -
          <lpage>958</lpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Jaidka</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chandrasekaran</surname>
            ,
            <given-names>M.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rustagi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kan</surname>
          </string-name>
          , M.Y.:
          <article-title>Overview of the clscisumm 2016 shared task</article-title>
          .
          <source>In: In Proceedings of Joint Workshop on Bibliometricenhanced Information Retrieval and NLP for Digital Libraries (BIRNDL</source>
          <year>2016</year>
          )
          <article-title>(</article-title>
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Lüdeling</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kytö</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Corpus linguistics: An international handbook</article-title>
          .
          <source>Citeseer</source>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Parinov</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Semantic attributes for citation relationships: Creation and visualization</article-title>
          . In: Garoufallou,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Virkus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Siatri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Koutsomiha</surname>
          </string-name>
          ,
          <string-name>
            <surname>D</surname>
          </string-name>
          . (eds.)
          <source>Metadata and Semantic Research</source>
          . pp.
          <fpage>286</fpage>
          -
          <lpage>299</lpage>
          . Springer International Publishing,
          <string-name>
            <surname>Cham</surname>
          </string-name>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Shotton</surname>
          </string-name>
          , D.:
          <article-title>CiTO, the citation typing ontology</article-title>
          .
          <source>Journal of biomedical semantics 1(Suppl 1)</source>
          ,
          <source>S6</source>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>