<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Optimal Citation Context Window Sizes for Biomedical Retrieval</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Boris Lykke Nielsen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stefan Lavlund Skau</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Florian Meier</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Birger Larsen</string-name>
          <email>birgerg@hum.aau.dk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Science, Policy and Information Studies Department of Communication and Psychology Aalborg University</institution>
          ,
          <addr-line>Copenhagen</addr-line>
          ,
          <country country="DK">Denmark</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <fpage>51</fpage>
      <lpage>63</lpage>
      <abstract>
        <p>We investigate the TREC-CDS 2016 test collections as a new resource for citation context and citation-based IR experiments. The collection contains more than 1.25 million biomedical full-text articles in XML. We nd that a citation index can easily be extracted, and citation contexts easily be identi ed. We conduct initial experiments to determine the optimal citation context window size in this domain and collection. Surprisingly We nd that quite long citation contexts of more than 250 word yield the best performance when combined linearly with the fulltext and with moderate weight on the citation contexts.</p>
      </abstract>
      <kwd-group>
        <kwd>Citation contexts for IR</kwd>
        <kwd>TREC-CDS 2016</kwd>
        <kwd>Citation Con- text Windows</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>In this work, we investigate the feasibility of extracting citation contexts from
citing articles and using them in the retrieval of scienti c documents. Adding
citation contexts of a document to its full-text may help the retrieval process
by providing additional relevant keywords for indexing. Keywords from citation
contexts may be valuable as they provide a di erent perspective on the cited
text | that of the authors citing and using the cited document.</p>
      <p>
        Prior research on this idea indicates that citation contexts can indeed
improve retrieval performance, e.g. [
        <xref ref-type="bibr" rid="ref1 ref14 ref16 ref3">14,1,16,3</xref>
        ]. Most previous work, however, has
been carried out on small collections of documents of no more than a few
thousand documents or in specialized domains. In the present work, we take
advantage of the increased availability of Open Access publications in full-text to
extract and study the usefulness of citation contexts for scienti c retrieval in
the hitherto largest publicly available test collection that supports this type of
retrieval. Speci cally, we work with 1.25 million documents from the biomedical
domain taken from the Open Access Subset of the PubMed Central collection.
These documents were included in the 2016 Text REtrieval Conference Clinical
Decision Support track (TREC-CDS) which produced a test collection that in
addition to the documents also includes information needs and associated
relevance assessments by medical professionals [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. This allows us to carry out
experiments where we explore the feasibility of extracting citation contexts from
such a vast collection and their potential for improving retrieval performance
in this domain. In particular as an initial e ort we investigate the following
research question: What is the optimal citation context window size with respect
to improving retrieval performance?
      </p>
      <p>The paper is structured as follows: Section 2 discusses related work. Section
3 presents the methods we used including analysis of the TREC-CDS 2016 test
collection, and details on the extraction of citation contexts. Section 4 describes
our experimental setup and our ndings. Section 5 presents discussion and
conclusions.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related work</title>
      <p>In establishing the Science Citation Index in 1964 Eugene Gar eld created a
retrieval system solidly based on the idea that citations form explicit links
between papers that have particular points in common and can be used to search
the scienti c literature [7, p.1]. A researcher would rely on the author's
judgement to include references to other publications with shared subjects or topics
and thus, through the network established from bibliographic references, identify
other citing or cited papers with similar or new topics or subjects in common [7,
p.2]. Continued as Clarivate's Web of Science, Gar eld's citation indexes is still
based on the links between papers, but ignores the nature and meaning of these
links. With the increasing availability of scienti c literature in full-text and as
open source, we can now begin to investigate the nature of the links by studying
the text where a given paper is mentioned. Existing research has demonstrated
that the text surrounding citations often contains descriptions of the cited
paper, the reason or function of the citation, or the disposition towards the cited
paper [10, p.201]. As such, it may be possible to glimpse what the cited paper
is about, or how the author has used the paper, by examining the surrounding
associated text of a speci c citation. In other words, by examining the textual
content of a citation, it is potentially possible to identify how and why a citation
is made.</p>
      <p>
        Analysis of Citation Contexts The purpose of citation context analysis, rst
proposed by Small[
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], is in general to examine the contextual relationship
between the citing and cited papers. White [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] reviews work in the area and notes
three lines of research: (1) Classifying citations, which "are attempts to
understand what people are doing when they cite" and involve citation classi cation
schemes of which he identi es and compares more than 20, (2) Content analysis
of citation contexts, in which words occurring in citation contexts are used to
describe the cited paper as basis for analysis, and (3) studies of Citer
motivations, which examines the deeper reason for "why authors make references". In
the present paper we do content analysis of citation contexts on a large scale.
      </p>
      <p>
        Citation context analysis has been studied from a number of perspectives,
including citation summarisation [
        <xref ref-type="bibr" rid="ref13 ref15 ref2 ref5">13,5,15,2</xref>
        ], creation of personalised citation
recommendations [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], discovery of new knowledge [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] or to manually or
automatically classify the motivations and functions that lies behind citations and
measure their impact [
        <xref ref-type="bibr" rid="ref22 ref9">9,22</xref>
        ].
      </p>
      <p>
        Analysis of Citation Contexts and IR Performance Of particular interest
to the present work is attempts to enhance and improve retrieval of scienti c
documents. An early example is O'Connor who proposed to extract additional
indexing terms from citation contexts that cite a given paper [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Bradshaw
follows a similar approach, but argues that terms from citation contexts can
su ce as document representation [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Both authors demonstrate that terms
from citation contexts can indeed improve retrieval performance | in particular
precision-based measures. More recent works that have used citation contexts
to improve retrieval of literature and have further motivated our project, are
the works by Ritchie [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] and Dabrowska and Larsen [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Ritchie conducted
retrieval experiments on a small test collection of 9800 full-text documents within
the scienti c area of computational linguistics [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. Their collection contained
approximately 20.000 citations pointing to about 3200 documents within the
collection. In her work, Ritchie de ned citation contexts, similar to Bradshaw,
within xed windows, but pursued more variations of windows using both
sentences and words (50, 75, 100 on each side of the citation), to be able to compare
the e ectiveness of the window sizes relative to one another.
      </p>
      <p>
        Following a similar approach to Ritchie but at a larger scale, Dabrowska
and Larsen performed retrieval experiments on a test collection of over 430.000
full-text papers from a subset of the iSearch (Integrated Search) document
collection [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. This collection contains approximately 3.7 million citations, with about
260.000 unique documents being cited by other documents within the collection
[
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. They found that retrieval, with their larger collection of physics papers, was
improved with the addition of citation contexts as index terms to the full-text
documents [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. More speci cally, they found improvements with xed windows
of words (25, 50, 75 &amp; 100 on each side) with the best results having a moderate
weight (of around 25%) on the citation contexts relative to the full-text.
3
3.1
      </p>
    </sec>
    <sec id="sec-3">
      <title>Method</title>
      <sec id="sec-3-1">
        <title>Data-set</title>
        <p>
          For our experiments we use the TREC 2016 Clinical Decision Support Track 1
data-set [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. Like most data-sets released by TREC, it consists of (i) a collection
of documents (ii) topics or user search tasks (iii) relevance assessments on which
documents of the collection are relevant for the topics. The document collection
is made up of 1.25 million full-text biomedical articles representing a snapshot
of the Open Access Subset of PubMed Central (PMC)2. PMC launched in 2000
        </p>
        <sec id="sec-3-1-1">
          <title>1 http://www.trec-cds.org/2016.html 2 https://www.ncbi.nlm.nih.gov/pmc/</title>
          <p>and is a free archive for full-text biomedical and life sciences journal articles. It
contains at present more than 5 million full-text articles, most of which are also
included in as bibliographical references in the 25+ million records of PubMed.
Each article in the collection is represented as an NXML le and identi ed by
an o cial PMCID.</p>
          <p>
            The topics of the track simulate actual information needs of physicians and
are divided in three di erent types representing the most common generic clinical
questions [
            <xref ref-type="bibr" rid="ref6">6</xref>
            ]. These types are: (i) Diagnosis, (ii) Treatment and (iii) Test. A
search task consists of (i) An admission note, (ii) A case description based on the
note, (iii) A summary of the case. These summaries were often written as
shortened versions of the case description. The search queries used in the experiments
of this study were based on the summary sections (iii), and the summaries were
not edited in any way when before being inserted into Indri as search queries.
The relevance assessments were conducted after the submission of retrieval runs
by participants of the track [
            <xref ref-type="bibr" rid="ref17">17</xref>
            ]. The relevance assessment was done by pooling
results from the submitted retrieval runs. The pooled documents were then
assessed as De nitely relevant (2), Possibly relevant (1), and Not relevant (0) [
            <xref ref-type="bibr" rid="ref17">17</xref>
            ].
A rating of Possibly relevant was given to documents that were not themselves
relevant to the topic but could prove relevant in the context of a broader
literature review [
            <xref ref-type="bibr" rid="ref17">17</xref>
            ]. The released relevance assessments (so-called QRELs) contain
28,349 unique documents, 5,461 of which are De nitely or Possibly relevant.
3.2
          </p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>Extracting Citation Contexts</title>
        <p>
          Citation contexts are de ned as the textual passages or sentences surrounding
or containing the citations [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]. In this work we utilize the cross-reference tags,
to determine the position and target reference for each in-text reference and
extract citation contexts. Although the position of the citation in the text is
known, identifying the start and end, i.e. the optimal context length is a di
cult and complex problem [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. Several researchers have used windows of xed
sizes starting with O'Connor and Bradshaw. However, by using xed windows
to de ne the contexts, it is possible that the context does not adequately
characterise the relationship to the referred citation. The context might exclude words
or sentences that implicitly refers to the citation, or include words or sentences
that do the exact opposite. In other words, the de ned citation context should
only contain the text that describes the cited paper. This has been attempted by
considering the linguistic features of the text to de ne contexts within windows
that contains the full scope of descriptive text to identify the optimal context
window size around the citation [
          <xref ref-type="bibr" rid="ref16 ref8">16,8</xref>
          ].
        </p>
        <p>We downloaded the TREC-CDS 2016 collection and rst investigated the raw
XML les to determine if (i) internal citations between the documents in the
data-set can be readily identi ed, and to answer the question: (ii) how feasible
is it to identify and extract citation contexts?</p>
        <p>The documents in the TREC-CDS 2016 test collection are encoded in the
Journal Archiving and Interchange Tag Set (JATS)3. All full-text documents in
the test collection use the XML le type and use the tags de ned in the JATS
DTD. The JATS include special tags for bibliographic references (&lt;xref&gt;) which
are wrapped around cross-references in the full-text document. The conducted
experiments utilised the cross-reference tags to determine the position and
target reference for each in-text reference. Additionally, the JATS include special
tags for the reference list to ease parsing of the list. This list was used to
determine the target document for each in-text reference. The documents in the test
collection have two unique and di erent identi ers that we could use to map
documents to each other: (i) PubMed Central Identi er (PMCID) This
identi er is assigned to all documents included in the PMC. As the test
collection is a snapshot of a subset of the PMC, all documents will have this identi er.
(ii) PubMed Identi er (PMID) This identi er is assigned to any record also
included in PubMed. This identi er (and not the PMCID) has also been added
by PMC sta in the JATS reference lists when pointing to documents included
in PubMed.</p>
        <p>As each document has a PMCID as well as a PMID (if it is in PubMed)
in its header, we were able to match these two IDs as a basis for extracting a
citation network and citation contexts. It is important to note that because not
all documents in the PMC are also present in PubMed, not all documents in
the test collection have a PMID. 87.622 (7%) of the documents did not have a
PMID.</p>
        <p>Table 1 gives a summary of the statistics of the extracted data. Of the 1.25
million documents just over a million have references (87.3%). 58,845 of the
documents without references are abstract-only documents without full-text, and
the remaining may be publication types that do not contain references (e.g.
editorials, news items, etc.) The 1+ million documents with references contain
more than 43 million references (40 references on average per article). Of the
1,255,260 documents in the collection, 567,650 documents (45.2%) received at
least one citation from within the collection. 370,426 of these were cited between
2 and 100 times | see Table 1 for other ranges. More than 60 million citation
contexts were identi ed (the same document can be mentioned more than once
in the full-text in one document). 46 million of these (76.6%) have a target PMID
pointing to a PubMed document, and 4.8 million (8%) could be matched to a
PMCID in the TREC-CDS 2016 collection. These 4,833,813 citation contexts
point to the 567,650 cited documents and form the core data-set used in our
experiments. This means that each cited document has 8.5 linked citation contexts
on average. A total of 37,707 documents were assessed for relevance in relation to
the 30 topics. Some were retrieved and assessed for multiple topics. The number
of unique PMCIDs in the QRELs is 28,349. Of these, 10,132 (36.1%) were cited
and had at least one associated citation context. However, only 2,019 of these
were assessed as De nitely relevant or Possibly relevant. This means that only
37% of the relevant documents also had citation contexts added.</p>
        <sec id="sec-3-2-1">
          <title>3 https://jats.nlm.nih.gov/index.html</title>
          <p>Contexts total count
Contexts with target PMID
Usable contexts (i.e. with a PMCID)
Documents with appended contexts
Length of QREL
Unique PMCIDs in QREL
Unique and relevant documents in QREL
Documents with appended contexts in QREL
Relevant documents with contexts in QREL
# of docs/refs/contexts</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experiments</title>
      <sec id="sec-4-1">
        <title>Experimental Setup</title>
        <p>
          Citation Context lengths We limit ourselves to the simplistic, however,
widely used approach of using xed window sizes surrounding the citation. This
is done by considering the citation as a starting point and then extend the
context to include text around the citation with some determined xed or variable
distance on each side of the citation. Citation context text can be variable and
range from just a few characters, over phrase and clauses to sentences and
entire paragraphs. The task of identifying the optimal context length of citations
is di cult and complex, as the citing behaviour, characteristics and processes
of citations is di erent within papers and sections and across authors, scienti c
discourses, elds and domains [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. However, one can argue that the best context
length is the one that yields the best results for the given purposes. In the present
work we investigate two simple approaches: (i) extracting a number of sentences
before and/or after the sentence in which the in-text reference occurs, and (ii)
extracting a number of words before and/or after the in-text reference. We did
experiments with 1-5 sentences before the in-text reference, and/or 1 sentence
after the reference and with 50-300 words on the left and/or right of the in-text
reference (See columns one and two of Table 2 for an overview. 300 25 means
that we used 300 words, with 25% = 75 words on the right). We include more
text before the in-text reference as these often occur at the end of a sentence,
and as we expect that most of the text commenting on that reference occurs
before it. Ritchie [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] and Dabrowska and Larsen [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] used up to 100 words. The
maximum of 300 words correspond to almost a page of full-text in the format
of the present paper, and we deemed it to be well beyond the upper-bound for
what could be bene cial for retrieval.
        </p>
        <p>
          Retrieval experiments As we are working with a test collection, our
experiments are solidly within the Cran eld tradition. We use the Indri IR system
to conduct the retrieval experiments, with Language Modeling and Dirichlet
smoothing [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ]. Initial experimentation in relation to constructing a baseline
investigated 24 values of the tuning parameter of the Dirichlet smoothing
between 1 and 45,000. Values in the range 20,000 - 35,000 provided the better
results with this collection so we choose to tune in this range: 20000, 22000,
24000, 27000, 30000, 33000, 36000, 39000. The baseline was tuned separately for
P@10, MAP and NDCG - the resulting baseline values can be seen in Table 2 4.
        </p>
        <p>To integrate citation contexts seamlessly into retrieval we took advantage of
two features in Indri. First, we placed the full-text of the original article in a
separate eld and added the di erent versions of extracted citation contexts in
elds of their own. Second, we used the Indri query language to create a
linear combination of the two. We used the #weight operator to assign weights to
the full-text and citation context respectively in each query, with the weights
adding up to 1. The following is an example of the query formatting (topic 1
from TREC-CDS 2016):</p>
        <p>#weight (0.6 #combine [fulltext] (A 78 year old male presents with
frequent stools and melena) 0.4 #combine [250 25] (A 78 year old male
presents with frequent stools and melena))</p>
        <p>where the same query string is matched against the full-text eld with a
weight of 0.6, and against one of the citation context elds with a weight of 0.4
(the #combine operator is standard operator for combining beliefs in Indri ).
In this way we can add the citation contexts as additional index terms to the
cited document, and control their in uence relative to the full-text. We tested
weights of 20, 40, 60, 80 and 100% in the main experiment. Retrieval results
were evaluated using trec eval. For P@10 and MAP both De nitely relevant and
Possibly relevant were counted as relevant | for NDCG De nitely relevant had
a gain value of 2 and Possibly relevant a gain value of 1. P@10 represents a
purely precision-oriented perspective on results whereas MAP and NDCG gives
a perspective that balances precision and recall. It should be noted that
retrieved documents that were not assessed were counted as non-relevant in all
the measures.
4 The smoothing parameter for the baseline run was 24,000 for MAP and 30,000 for
MAP and NDCG.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Findings</title>
        <p>The aim of the experiments was to learn how much context to include to reach
optimal retrieval performance, and to determine how much weight to put on
them relative to the full-text. Table 2 shows the overall results for P@10, MAP
and NDCG. The table shows the best result for each run with the optimal
smoothing parameter and the weight between full-text and citation context that
performs the best. Figures 1 and 2 below demonstrate the e ect of weighting
and smoothing.</p>
        <p>
          From Table 2 we see that the citation context runs all outperform the baseline
to some degree. The runs with more context added in general perform better.
The best performing run for P@10 is the one with 250 words added (19% over
the baseline), and for MAP and NDCG the run with 300 words added (4.6%
and 2.8% over the baseline respectively). The best performing word-based runs
outperform the sentence-based runs. An explanation may be that the
sentencebased runs in general would be shorter in term of the number of words - with
the longest of 5 sentences corresponding to 150 words or less. 5
5 We did not examine sentence length of the 4.8 million citation contexts, but other
studies of academic biomedical text nd around 25 words per sentence on average
[
          <xref ref-type="bibr" rid="ref24">24</xref>
          ].
        </p>
        <p>With regards to whether adding text to the right of the in-text reference
is bene cial, a more complex picture emerges: for early precision (P@10) runs
without right-hand text performs better in almost all cases | especially for
the better performing runs with more words. For MAP and NDCG the actual
di erences are very small | for the runs with more context in the 100-250 word
range performance is marginally better with right-hand text, but as more context
gets added, right-hand context doesn't seem to have much in uence, as can be
seen for the MAP performance (4.6% over baseline) in the 250-300 word range.</p>
        <p>Overall, all results point to the fact that more context added leads to better
performance. Given this nding even larger context windows should be
experimented with to ensure that an upper-bound has indeed been reached at 300
words, which was the maximum context length studied in the present
experiment.</p>
        <p>The results in Table 2 represent the best smoothing and weight
combinations. With respect to the linear combination the top 8 runs in relation to P@10
all have 40% weight on citation contexts - for MAP and NDCG the top runs
all have a 20% weight on the citation contexts (not shown in Table 2). Figure
1 demonstrates the e ect of the linear combination of full-text and the added
citation contexts for MAP for weights of 10, 15, 20, 25 and 40%. The shading of
each line indicates the variation across the parameter tuning range, and the
dashed line is the baseline. It can be seen that performance is very stable, and
increases steadily as more weight is placed on the citation contexts from 10% up
to 25%. This also holds across the di erent context length combinations with the
0.2633</p>
        <p>MAP
top performance being reached at 300 words as discussed above, with scores that
are consistently over the baseline except for the shortest citation context
windows. At 40% however, performance overall drops below the baseline and shows
great diversity across contexts lengths, demonstrating that too much weight on
the citation contexts can hurt performance and lead to erratic behaviour. It can
also be noted that variation across the smoothing parameter range (shading) is
not prohibitively large with little overlap between each type of weighting.</p>
        <p>Figure 2 further illustrates the e ect of smoothing for contexts of 250 words
(the best performing context length for P@10). Results are plotted for weights
of 15, 20, 25 and 40% and for P@10, MAP and NDCG. A low or moderate
level of variation across the tuning range is desirable as such an approach is less
dependent on setting the tuning parameter correctly and can thus be considered
more stable. It can be observed that performance is quite stable across the chosen
smoothing range, with a bit more variation for P@10, which can be expected as it
is a less stable measure. It can be clearly seen that a weight of 40% outperforms
other weights for P@10 across the smoothing range. At a weight of 20% MAP
and NDCG is very stable across the smoothing range. Higher performance can
be achieved at 25% for MAP and NDCG but only at the lower smoothing values.</p>
        <p>Finally, it should be noted that as expected, the citation distribution is
skewed (1) - with 63% receiving no citations, the remaining 37% receiving more
than 1 - and 18 documents receiving more than 1000 citations. With the
experimental setup used this means that some documents have no additional
representation, and that a sizable proportion have a great deals of text added to their
representation - in some cases tens of thousands of words. The e ect of this on
retrieval is at present unknown.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Discussion and Conclusion</title>
      <p>Generating a citation index and extraction of citation contexts It was
unproblematic to identify internal citations documents in the TREC-CDS 2016
collection due to explicit tagging in the XML and the added PMIDs in the
reference lists. Not all documents have references, but more than 1 million do |
providing a rich testbed for citation-based IR. Further, 45% of the documents in
the collection received at least one internal citation. The resulting citation index
from RQ1 and the XML and JATS formats made it possible to identify citing
documents and to identify the location of in-text references. 4,8 million citation
contexts could thus be extracted and linked to the 567,650 cited documents. It is
worth noting that only 37% of the relevant documents had one or more citation
contexts added - this reduces the impact that citation contexts can have on
IR performance in the TREC-CDS collection. This underlines that fact that
retrieval based on citation contexts is inherently dependent on documents being
cited, and that a test collection which has been created with citation and citation
context based runs in the pooling process is really needed to fully investigate
the true potential of these approaches.</p>
      <p>We limited ourselves to a simple de nition of citation contexts extracting a
number of words or sentences before and/or after each in-text reference. Tests
of more advanced linguistically-based de nitions would be interesting, but can
be a challenge given the large number of contexts. Somewhat to our surprise
the best performance was found among the longest citation windows of 250-300
words - both for precision and recall-oriented measures. As argued this a large
amount of text (a full page) that almost certainly goes beyond where a given
cited document is discussed. This may indicate that identifying the exact extent
of the actual citation context may not be of prime importance - and on the
other hand leads to the question of why so much text from citing documents is
bene cial for retrieval, and if even larger windows will be bene cial?</p>
      <p>Compared to previous research this is much longer citation context windows
than previously tested. With regards to how much weight to put on the contexts
our ndings are in line with previous work of e.g. Ritchie (2009) and Dabrowska
(2014). Best performance is achieved with moderate weight on the citation
contexts of around 20% relative to the full-text - much more leads to decreased
performance and erratic behaviour. As regards stability, results are quite
stable across the smoothing range, indicating that this approach does not depend
critically on getting the smoothing parameter right.</p>
      <p>
        The present study mainly serves to introduce the TREC-CDS 2016 test
collection as an attractive ressource for the BIR community and those interested
in citation context analysis - and to conduct initial tests of the feasibility of IR
experiments using such features on this collection. Much more interesting work,
where it is tested if the context of citations are useful for semantically
categorizing a relationship and perhaps even an intention or an opinion between two
publications, can be built on top of this, e.g. along the lines of Ritchie [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
5.1
      </p>
      <sec id="sec-5-1">
        <title>Acknowledgments</title>
        <p>We wish to thank the organisers of TREC-CDS for creating a great resource,
and the four anonymous reviewers for insightful comments.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bradshaw</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Reference directed indexing: Redeeming relevance for subject search in citation indexes</article-title>
          . In: Koch, T., S lvberg, I.T. (eds.) Research and
          <article-title>Advanced Technology for Digital Libraries</article-title>
          . pp.
          <volume>499</volume>
          {
          <fpage>510</fpage>
          . Springer, Berlin, Heidelberg (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Cohan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goharian</surname>
          </string-name>
          , N.:
          <article-title>Scienti c document summarization via citation contextualization and scienti c discourse</article-title>
          .
          <source>Int. J. Digit. Libr</source>
          .
          <volume>19</volume>
          (
          <issue>2-3</issue>
          ),
          <volume>287</volume>
          {
          <fpage>303</fpage>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Dabrowska</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Larsen</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Exploiting citation contexts for physics retrieval</article-title>
          .
          <source>Proc. of BIR Workshop @</source>
          ECIR pp.
          <volume>14</volume>
          {
          <issue>21</issue>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Ding</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
          </string-name>
          , G.,
          <string-name>
            <surname>Chambers</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Song</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhai</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Content-based citation analysis: The next generation of citation analysis</article-title>
          .
          <source>JASIST</source>
          <volume>65</volume>
          (
          <issue>9</issue>
          ),
          <year>1820</year>
          {
          <year>1833</year>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Elkiss</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fader</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Erkan</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>States</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Radev</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Blind men and elephants: What do citation summaries tell us about a research article? JASIST 59(1</article-title>
          ),
          <volume>51</volume>
          {
          <fpage>62</fpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Ely</surname>
            ,
            <given-names>J.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oshero</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gorman</surname>
            ,
            <given-names>P.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ebell</surname>
            ,
            <given-names>M.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chambliss</surname>
            ,
            <given-names>M.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pifer</surname>
            ,
            <given-names>E.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stavri</surname>
            ,
            <given-names>P.Z.:</given-names>
          </string-name>
          <article-title>A taxonomy of generic clinical questions: classi cation study</article-title>
          .
          <source>BMJ</source>
          <volume>321</volume>
          (
          <issue>7258</issue>
          ),
          <volume>429</volume>
          {
          <fpage>432</fpage>
          (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7. Gar eld, E.:
          <string-name>
            <surname>Citation</surname>
          </string-name>
          indexing
          <article-title>- Its Theory and</article-title>
          Application in Science, Technology, and Humanities. Wiley, New York (
          <year>1979</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Hernandez-Alvarez</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gomez</surname>
            ,
            <given-names>J.M.:</given-names>
          </string-name>
          <article-title>Survey about citation context analysis: Tasks, techniques, and resources</article-title>
          .
          <source>Natural Language Engineering</source>
          <volume>22</volume>
          (
          <issue>3</issue>
          ),
          <volume>327</volume>
          {
          <fpage>349</fpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Hernndez-Alvarez</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Gomez</given-names>
            <surname>Soriano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.M.</given-names>
            ,
            <surname>Martinez-Barco</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.</surname>
          </string-name>
          :
          <article-title>Citation function, polarity and in uence classi cation</article-title>
          .
          <source>Natural Language Engineering</source>
          <volume>23</volume>
          (
          <issue>4</issue>
          ),
          <volume>561</volume>
          {
          <fpage>588</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thoma</surname>
            ,
            <given-names>G.R.:</given-names>
          </string-name>
          <article-title>Machine learning with selective word statistics for automated classi cation of citation subjectivity in online biomedical articles</article-title>
          .
          <source>In: Proc. of ICAI'17</source>
          . pp.
          <volume>201</volume>
          {
          <issue>207</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yan</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yan</surname>
          </string-name>
          , H.:
          <article-title>Guess what you will cite: Personalized citation recommendation based on users' preference</article-title>
          . In: Banchs, R.E.e.a. (ed.) IR Technology,
          <string-name>
            <surname>LNCS</surname>
          </string-name>
          , Volume
          <volume>8281</volume>
          . pp.
          <volume>428</volume>
          {
          <fpage>439</fpage>
          . Springer, Berlin, Heidelberg (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Lykke</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Larsen</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lund</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ingwersen</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Developing a test collection for the evaluation of integrated search</article-title>
          . In: Gurrin, C.e.a. (ed.)
          <source>Advances in Information Retrieval</source>
          . pp.
          <volume>627</volume>
          {
          <fpage>630</fpage>
          . Springer, Berlin, Heidelberg (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Mei</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhai</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Generating impact-based summaries for scienti c literature</article-title>
          .
          <source>In: In Proc. of ACL'08 HLT</source>
          . pp.
          <volume>816</volume>
          {
          <fpage>824</fpage>
          .
          <string-name>
            <surname>ACL</surname>
          </string-name>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <given-names>O</given-names>
            <surname>'Connor</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          :
          <article-title>Citing statements: Computer recognition and use to improve retrieval</article-title>
          .
          <source>Information Processing &amp; Management</source>
          <volume>18</volume>
          (
          <issue>3</issue>
          ),
          <volume>125</volume>
          {
          <fpage>131</fpage>
          (
          <year>1982</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Qazvinian</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Radev</surname>
            ,
            <given-names>D.R.</given-names>
          </string-name>
          :
          <article-title>Scienti c paper summarization using citation summary networks</article-title>
          .
          <source>In: Proc. of the 22Nd ICCL - Volume 1</source>
          . pp.
          <volume>689</volume>
          {
          <fpage>696</fpage>
          . ACL, Stroudsburg, PA, USA (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Ritchie</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Citation context analysis for information retrieval</article-title>
          . University of Cambridge Computer Laboratory
          <source>Technical Report 744</source>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Roberts</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Demner-Fushman</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Voorhees</surname>
            ,
            <given-names>E.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hersh</surname>
            ,
            <given-names>W.R.</given-names>
          </string-name>
          :
          <article-title>Overview of the trec 2016 clinical decision support track</article-title>
          .
          <source>In: Proc. of TREC 25</source>
          . pp.
          <volume>1</volume>
          {
          <issue>14</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Small</surname>
          </string-name>
          , H.:
          <article-title>Citation context analysis</article-title>
          . In: Dervin,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Voigt</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.J</surname>
          </string-name>
          . (eds.) Progress in Communication Sciences, pp.
          <volume>287</volume>
          {
          <fpage>310</fpage>
          .
          <string-name>
            <surname>Ablex</surname>
            , Norwood,
            <given-names>N.J.</given-names>
          </string-name>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Small</surname>
          </string-name>
          , H.:
          <article-title>Interpreting maps of science using citation context sentiments: a preliminary investigation</article-title>
          .
          <source>Scientometrics</source>
          <volume>87</volume>
          (
          <issue>2</issue>
          ),
          <volume>373</volume>
          {
          <fpage>388</fpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Small</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tseng</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patek</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Discovering discoveries: Identifying biomedical discoveries using citation contexts</article-title>
          .
          <source>Journal of Informetrics</source>
          <volume>11</volume>
          (
          <issue>1</issue>
          ),
          <volume>46</volume>
          {
          <fpage>62</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Strohman</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Metzler</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Turtle</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Croft</surname>
          </string-name>
          , W.B.:
          <article-title>Indri: a language-model based search engine for complex queries</article-title>
          .
          <source>Tech. rep.</source>
          ,
          <source>In Proc. of the International Conference on Intelligent Analysis</source>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Teufel</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Siddharthan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tidhar</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Automatic classi cation of citation function</article-title>
          .
          <source>In: Proc. of EMNLP'06</source>
          . pp.
          <volume>103</volume>
          {
          <fpage>110</fpage>
          . ACL, Stroudsburg, PA, USA (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>White</surname>
          </string-name>
          , H.D.:
          <article-title>Citation analysis and discourse analysis revisited</article-title>
          .
          <source>Applied Linguistics</source>
          <volume>25</volume>
          (
          <issue>1</issue>
          ),
          <volume>89</volume>
          {
          <fpage>116</fpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Zorita</surname>
            ,
            <given-names>C.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moreno-Sandoval</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Sentence length and np complexity of general and medical written academic and media texts. an analysis using a trained syntactic parser</article-title>
          . In: Ortiz,
          <string-name>
            <given-names>A.M.</given-names>
            ,
            <surname>Perez-Hernandez</surname>
          </string-name>
          ,
          <string-name>
            <surname>C</surname>
          </string-name>
          . (eds.)
          <source>In Proc. of CILC2016. 8th International Conference on Corpus Linguistics. EPiC Series in Language and Linguistics</source>
          , vol.
          <volume>1</volume>
          , pp.
          <volume>181</volume>
          {
          <issue>190</issue>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>