<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Are Cited References Meaningful? Measuring Semantic Relatedness in Citation Analysis</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Hassan Alam</string-name>
          <email>hassana1@bcltechnologieshttp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aman Kumar</string-name>
          <email>amank2@bcltechnologieshttp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tina Werner</string-name>
          <email>twerner3@bcltechnologieshttp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Manan Vyas</string-name>
          <email>mvyas4@bcltechnologieshttp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>BCL Technologies (hassana</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this proof-of-concept study we use standard cosine similarity measure to calculate the semantic similarity between two pieces of text - the citing document and the cited text. Three subject matter experts then evaluate the citing and the cited text based on the cosine score to give their judgement on the semantic similarity between the two pieces of text.</p>
      </abstract>
      <kwd-group>
        <kwd>Bibliometrics</kwd>
        <kwd>citation analysis and network analysis for IR</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Researchers and scientists in both academia and industry present and publish their
research work in a variety of places and platforms. Because of career pressure and other
factors, they are encouraged to publish and present more and more. The large and
rapidly increasing amount of scientific literature online and otherwise (book, journals) has
triggered intensified-research into understanding the effectiveness and quality of this
research work.</p>
      <p>When reading a research paper we often glance at the bibliography or the list of
references for additional information. An author cites the references when they look up for
information while preparing for their research paper and they want to acknowledge all
the sources they have used in the process of writing that paper. Ideally, authors are
expected to report the sources even though they do not quote directly from that source.
Readers can use the referenced list to check for the accuracy of the published material
and that establishes credibility for the author. But as a reader, we may not have time to
do further consultation because of the sheer size of the cited material.
In this study we are building a proof-of-concept system that looks only at the relevant
parts of the cited material that is appropriate for evaluating the claims made by the
original author in a given paragraph. This system analyzes the text around the cited
sentences or text in the original article and tries to find the cited material from the
referenced articles to check if the cited text is semantically related to the citing text in the
original document.
Compared to humans, this tool cuts down the time considerably in reading and
analyzing the cited material.</p>
      <p>Our goal in this study is do a proof-of-concept study to evaluate the relationship
between citing and cited documents, by examining measures of cosine similarity between
the citing sentences and the text of the cited scientific articles. Since both the citing and
the cited documents discuss the same topics, we assume that the concepts that are
relevant to one another will be more similar than those that are not. If effective, this will
allow identification of the material in the cited article that is relevant to the citing text.
Once we establish that this similarity metrics for this specific task gives satisfactory
results, we will implement other semantic similarity measures such as Latent Semantic
Indexing and evaluate the results.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        Author in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] explores reasons why citing and cited works may be related. The analysis
indicates that factors such as sources of the cited document, citing work, frequency of
a work cited, and type of citing articles predict closer relatedness between citing and
cited works. Authors in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and references there-in discuss several measures of
similarity and relatedness, such as the Pearson correlation and conclude that the cosine
index performs the best.
      </p>
      <p>In this preliminary work, first, we use standard similarity index – cosine similarity score
to establish similarity between two pieces of text, and then we use manual judgment to
understand what type of citing and cited texts closely match semantically with each
other.</p>
      <p>This paper is organized as follows. In section 2 we describe the methodology we adopt
to do the empirical analysis to establish semantic relatedness. In section 3 we describe
the experiment set up for this task, followed by a discussion of the evaluations, and
conclusions.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Methodology</title>
      <p>In this study we want to investigate the degree to which automated methods can reflect,
match, or even predict human judgments and to understand the semantic relationship
between the original article and the cited text.</p>
      <p>The automated system calculates the cosine similarity between all sentence pairs, which
is then compared with the Subject Matter Expert’s (SME) relevancy judgment. The idea
is that can we correlate the semantic similarity of two sentences and ascertain the
relationship of relevance between the citing and the cited text.</p>
      <p>
        For Data Preprocessing, we used the stop word list [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] to get rid of the stop words for
further processing.
For stemming, we used Porter stemming algorithm [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] which is a Java implementation.
The motivation for stemming is that if we do not do stemming, the tfidf counts will
yield false results. tf as such is not sufficient for our goals of predicting an article’s
relevancy or establishing similarity between two pieces of texts. Using the inverse
document frequency lowers the weight of common terms. A weight is created by the tfidf
for each term. This establishes a balance between how often a term appears in an
individual document and with how many documents use the term. In this model, a common,
more frequent term is weighted lightly and an unfamiliar or rare term is weighted more
heavily. This results in identifying discriminative terms. Mathematically, tfidf weight
is calculated using the standard formula:
,
∗ log
      </p>
      <p>Where, is the term, and j is the document.</p>
      <sec id="sec-3-1">
        <title>Normalization</title>
        <p>The term frequencies can be influenced by difference in the length of the article. A
more frequent term in a long article will skew the results. Also, it's likely that in short
article a term gets repeated a number of times which may lead to misleading results. In
order to mitigate the effect due to article length and term frequencies we need to
normalize the term weight for each article. The normalization of term weight is expressed
mathematically as:</p>
        <p>norm(D) = √(∑w(j)2)</p>
        <p>Here j is the document</p>
      </sec>
      <sec id="sec-3-2">
        <title>Cosine Similarity Score</title>
        <p>In order to compute the similarity of each pair of the compared items, the cosine
similarity gives a numerical value that describes by how much the two compared items are
close to each other. A group of cosine similarity score creates a natural ordering of
comparisons in which the highest values are the most similar and the lowest values are
the least.
This cosine similarity score computes a value, adjusted for article length, to depict the
similarity for each sentence pair, based on the values of shared terms. Mathematically,
it is represented as follows.</p>
        <p>Cosine (D1, D2) = ∑(wD1(j)*wD2(j)/norm(D1)*norm(D2))
To summarize, the algorithm we implemented for this proof-of-concept system is as
follows.</p>
        <p>Step 1: Term frequencies and inverse document frequencies is calculated for
each individual stemmed term.</p>
        <p>Step 2: The term frequencies are combined to create a TF*IDF score.
Step 3: The TF*IDF score is then normalized to account for varying lengths
between sentences.</p>
        <p>Step 4: This normalization is then be used to calculate cosine similarity
between each citing sentence and every sentence in the cited article.</p>
        <p>Step 5: The similarity score is compared with manual assessments of
whether the paired sentences from the citing and citing articles cite or
support one another.
3.1</p>
      </sec>
      <sec id="sec-3-3">
        <title>Data</title>
        <p>
          We wrote a tool to extract data from NLM/NCBI [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. The NLM index includes the full
title for each journal, as well as each journal’s accepted abbreviations, making it
possible to disambiguate and group the varied forms of each journal title under the same
identifier. Each article is indexed, and has a unique identifying number, the Pubmed
ID, or PMID. The NLM offers a Batch Citation Matcher at
www.ncbi.nlm.nih.gov/entrez/getids.cgi. This citation matcher provides the PMID for each known citation.
Here’s a snapshot of the NCBI Batch Citation Matcher.
Using this interface at NCBI we can submit extracted citations in batches that could
range from fifty thousand to one hundred thousand at a time, and load the responses
from the NLM back into our database. This allows us to link the articles to their PMID
using the title, date, journal, etc. from each citation.
and the corresponding number in the reference section contained the full details of the
citation.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experiment</title>
      <p>We extracted 50 journal articles from PubMed. For each citation in an article the tool
extracted the corresponding paper. The tool then extracted two sentences before and
two sentences after the citation in the original document and tried to match the words
in those sentences with the target document using the cosine similarity metric. This
process generated a cosine similarity index for each citation in the original document.
Once we have the cosine similarity measurements, we picked up the pairs (citation in
the original sentence and relevant parts in the cited document) that scored higher than
0.90. Three subject matter experts (SME - clinical experts in this case) then manually
evaluated the citations sentences and the cited documents, and decided which of the
correlated documents matched the most. The human experts based their judgement
mainly on semantic matching of the sentences in the two documents and not just on
matched strings.
4.1</p>
      <sec id="sec-4-1">
        <title>Results</title>
        <p>For manual evaluations the three SMEs looked at 100 matched set that scored higher
than 0.90 cosine similarity score. SMEs rated their assessment on a scale of 1-100, 100
being the best match. For example, SME-1 found that out of the 100 paired texts, 62
talked about the same concept. In this preliminary study we did not analyze the
disagreements between the SMEs. Table 1 gives a summary of this evaluation process.
SME
SME-1
SME-2
SME-3
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>In this proof-of-concept study we analyzed the textual similarity between citation text
in an original research paper from PubMed and the corresponding text in the cited
document. We tried to understand how close the author was in citing the cited paper. We
first used cosine similarity measure to come up with a paired list of citation text and
cited text. We then looked at 100 such pairs with a cosine similarity score of over 0.90.
The system recorded an average accuracy of 64.33% based on the evaluations of the
three SMEs. For future work, we plan to extend the similarity metrics using the
WordNet synset hierarchy and distributional similarity and Latent Semantic Analysis (LSA)
index.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bonzi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Characteristics of a literature as predictors of relatedness between cited and citing works</article-title>
          .
          <source>Journal of the American Society for Information Science</source>
          ,
          <volume>33</volume>
          (
          <issue>4</issue>
          ):
          <fpage>208</fpage>
          -
          <lpage>216</lpage>
          (
          <year>1982</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Boyack</surname>
            ,
            <given-names>K. W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Small</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Klavans</surname>
          </string-name>
          , R.:
          <article-title>Improving the Accuracy of Co-citation Clustering Using Full Text</article-title>
          .
          <source>In Proceedings of 17th International Conference on Science and Technology Indicators</source>
          . (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Corley</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Mihalcea</surname>
          </string-name>
          , R.:
          <article-title>Measuring the Semantic Similarity of Text</article-title>
          ,
          <source>in Proceedings of the ACL workshop on Empirical Modeling of Semantic Equivalence and Entailment</source>
          , pp.
          <fpage>13</fpage>
          −
          <lpage>18</lpage>
          . (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Klavans</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boyack</surname>
            ,
            <given-names>K.W.</given-names>
          </string-name>
          :
          <article-title>Identifying a Better Measure of Relatedness for Mapping Science</article-title>
          .
          <source>Journal of the American Society for Information Science and Technology</source>
          .
          <volume>57</volume>
          (
          <issue>2</issue>
          ) pp.
          <fpage>251</fpage>
          -
          <lpage>263</lpage>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Madylova</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Oguducu</surname>
            ,
            <given-names>S.G.</given-names>
          </string-name>
          :
          <article-title>A taxonomy based semantic similarity of documents using the cosine measure</article-title>
          ,
          <source>in Proceeding of International Symposium on Computer and Information Sciences</source>
          ,
          <year>2009</year>
          , pp.
          <fpage>129</fpage>
          −
          <lpage>134</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6. van Eck,
          <string-name>
            <surname>N. J.</surname>
          </string-name>
          , Waltman.:
          <article-title>Appropriate Similarity Measures for Author Co-citation Analysis</article-title>
          .
          <source>Journal for the American Society for Information Science and Technology</source>
          .
          <volume>59</volume>
          (
          <issue>10</issue>
          ) pp.
          <fpage>1653</fpage>
          -
          <lpage>1661</lpage>
          . (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>7. https://www.ncbi.nlm.nih.gov/books/NBK3827/table/pubmedhelp.T.stopwords/</mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>8. http://www.nltk.org/howto/stem.html</mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>9. https://www.ncbi.nlm.nih.gov/pubmed/batchcitmatch)</mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>