<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>CE-KG: Citation-enhanced Knowledge Graph via Citation Sentiment Fusion and Evidence Tracing</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yalan Huang</string-name>
          <email>huang.yalan@imicams.ac.cn</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Xuemei Yang</string-name>
          <email>yang.xuemei@imicams.ac.cn</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bin Zhang</string-name>
          <email>zhang.bin@imicams.ac.cn</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Xiaoli Tang</string-name>
          <email>tang.xiaoli@imicams.ac.cn</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute of Medical Information, Chinese Academy of Medical Sciences</institution>
          ,
          <addr-line>Beijing</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This study proposes a knowledge graph (KG) construction method integrating citation information to address credibility tracing of knowledge claims. By fine-tuning large language models (LLMs), SPO triples are precisely extracted from the "Conclusion" sections of literature abstracts. A dual-level fusion mechanism is developed based on sentiment analysis (Positive/Neutral/Negative) of citation contexts. This approach injects academic evaluation attributes at the paper level and links them to SPO triples. An innovative Neo4j-MySQL heterogeneous storage architecture is designed, enabling fine-grained evidence tracing from KG relationships to citation contexts through uniform identifiers. The constructed KG simultaneously provides knowledge claims and supports credibility tracing from the academic community perspective.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Knowledge Graph</kwd>
        <kwd>Citation Sentiment Analysis</kwd>
        <kwd>Evidence Tracing</kwd>
        <kwd>Credibility Verification</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Biomedical literature serves as a critical carrier of scientific research achievements, documenting major
breakthroughs, knowledge discoveries, and direct empirical evidence from clinical studies. Current
studies typically construct domain knowledge graphs by extracting entity relationships from titles and
abstracts[
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1, 2, 3</xref>
        ]. However, this approach exhibits two key limitations: First, while abstracts summarize
academic knowledge production processes, the Conclusion section contains investigated knowledge
claims, arguments, and assertions[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], where other sections may interfere with core claims. Second,
identical entity-relationship pairs from diferent sources have unequal credibility. Notably, citations
are vital pathways for disseminating knowledge claims and establishing credibility[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The sentiment
polarity (e.g., Positive/Neutral/Negative) in citation contexts reflects citing authors’ academic attitudes
toward cited claims[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], serving as a key credibility source.
      </p>
      <p>To overcome these limitations and leverage citation sentiment, this study proposes a KG framework
integrating evidence from citation information (including sentiment and context). Core innovations
include:</p>
      <p>1. Structured Evidence Screening: SPO triples are strictly extracted from "Conclusion" sections of
scientific abstracts, ensuring precise reflection of research claims while minimizing information noise.</p>
      <p>2. Citation Information Fusion: Deep learning models parse sentiment polarity in original citation
contexts to quantify academic evaluations between papers. Citation sentiments, contexts, and SPO
triples are fused to enable credibility tracing. These mechanisms aim to transform static KGs into tools
supporting knowledge claim credibility tracing.</p>
      <p>These mechanisms transform static KGs into tools supporting knowledge claim credibility tracing.
The code and data used in our Knowledge Graph all available at https://github.com/wang-lch/CE-KG.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Methodology</title>
      <sec id="sec-2-1">
        <title>2.1. Data Collection and Processing</title>
        <p>
          This study establishes a multimodal biomedical knowledge base integrating three heterogeneous data
sources: Metadata of 9,201 clinical research papers on breast cancer drug therapy (including trials
and guidelines) published from 2000-2024 were retrieved from PubMed; 83,510 medical SPO triples
were extracted from the SemMedDB[
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] using PMIDs; 127,927 valid citation contexts were collected
via the Colil[
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] database API. To enhance extraction reliability, the fine-tuned LLM Qwen3-0.6B[
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]
was adopted for abstract structuring. This model automatically classifies content into Background,
Objective, Methods, Results, and Conclusion sections. Training utilized 100,000 biomedical abstracts
with structured labels, with optimization parameters set to learning rate 1e-5, batch size 4, and random
seed 42. After three iterations, the model achieved a 0.94 F1-score for Conclusion sections on the test
set. SPO triples were strictly extracted from "Conclusion" sections to focus on core research claims and
minimize noise. This process yielded 14,223 deduplicated triples.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Citation-enhanced Knowledge Graph Construction</title>
        <p>
          An evidence-oriented filtering strategy was adopted, using only Conclusion-derived SPO triples to
ensure claims reflect empirical findings. For confidence quantification, the Dict-Senti-BERT[
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] model
identified sentiment polarity (Positive/Neutral/Negative) in citation contexts. Sentiments were injected
through dual-level fusion: Paper-level calculations determined Positive/Negative sentiment proportions;
SPO Triples-level integration embedded source literature and sentiment distributions as provenance
attributes. This enables knowledge claim tracing and evidence queries. Model training employed
optimization parameters of learning rate 5e-6, batch size 32, and weight decay 0.01 over 50 epochs. After
training, the model achieved an overall F1-score of 0.93 and an accuracy of 0.94 on the held-out test set,
demonstrating stable performance in conclusion-oriented triple extraction and sentiment-enhanced
integration
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. KG Storage and Update Mechanism</title>
        <p>A Neo4j-MySQL heterogeneous storage architecture enables dynamic knowledge network storage and
ifne-grained evidence tracing. Neo4j stores normalized entities (e.g., drugs, diseases) as nodes and
SPO predicates (e.g., TREATS, CAUSES) as directed relationships with embedded evidence support
(PMID lists and sentiment distributions). MySQL maintains structured evidence: The Papers table
records metadata (titles, structured abstracts), while the Citations table stores citation types and context
snippets. A unified identifier system (Triple_ID/PMID/Citation_ID) enables cross-database mapping
from KG relationships to specific citation contexts, forming an integrated solution supporting knowledge
evolution and claim verification.</p>
      </sec>
      <sec id="sec-2-4">
        <title>2.4. Knowledge Graph Visualization</title>
        <p>This study designs a hierarchical knowledge graph (KG) visualization scheme enabling traceable analysis
from knowledge claims to microscopic evidence through three progressive interfaces. The workflow
begins at the semantic triple layer where clicking a predicate edge of any SPO triple triggers a detailed
attribute panel. This panel displays the PMID list of associated literature. The system then dynamically
renders a literature overview layer (central expansion panel). Each paper presents structured data
including title, PMID, publication year, and sentiment polarity distribution (e.g., positive/negative
citation counts). Finally, the citation layer comprehensively displays all citation contexts of individual
papers: Total citation statistics, sentiment classification data, and sentiment-labeled citation contexts
with associated metadata (source PMID, title, year) are systematically presented. This forms a complete
evidence tracing chain, facilitating credibility verification of knowledge claims.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Conclusion and Future Work</title>
      <p>This study proposes a knowledge graph (KG) construction method integrating citation sentiment
and contextual evidence. The evidence-oriented mechanism—strictly extracting SPO triples only
from abstract "Conclusion" sections—significantly improves knowledge claim accuracy. Innovatively
incorporating citation information enables credibility tracing for academic knowledge discovery. The
constructed KG links knowledge claims to source literature’s citation data, allowing examination of
academic community evaluations and facilitating access to high-credibility claims.</p>
      <p>Future work will employ deep learning tools to extract multi-source citation data, complementing
the Colil database to enrich academic evaluations. An alignment mechanism between citation contexts
and SPO triples will be explored for precise evidence tracing. Additionally, applicability assessment in
other domains will be conducted.</p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgments</title>
      <p>This research was funded by the Innovation Fund for Medical Sciences of the Chinese Academy of
Medical Sciences (Grant number: 2021-I2M-1-033)</p>
    </sec>
    <sec id="sec-5">
      <title>Declaration on Generative AI</title>
      <p>During the preparation of this manuscript, the authors employed ChatGPT and Grammarly to support
grammar correction, spelling refinement, and overall language polishing. These tools were used to
enhance clarity and professionalism in writing. The authors subsequently reviewed and revised all
content, and took full responsibility for the accuracy and integrity of the final publication.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Nian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Feng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Tao</surname>
          </string-name>
          ,
          <article-title>Mining on alzheimer's diseases related knowledge graph to identify potential AD-related semantic triples for drug repurposing</article-title>
          ,
          <source>BMC Bioinformatics 23</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Feng</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y. Jing,</surname>
          </string-name>
          <article-title>PlagueKD: a knowledge graph-based plague knowledge database</article-title>
          ,
          <year>Database 2022</year>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Jin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Knowledge graph for breast cancer prevention and treatment: Literature-based data analysis study</article-title>
          ,
          <source>JMIR Medical Informatics</source>
          <volume>12</volume>
          (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>X.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Du</surname>
          </string-name>
          , E. Dong,
          <article-title>Extracting and measuring uncertain biomedical knowledge from scientific statements</article-title>
          ,
          <source>Journal of Data and Information Science</source>
          <volume>7</volume>
          (
          <year>2022</year>
          )
          <fpage>6</fpage>
          -
          <lpage>30</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Sarol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ming</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Radhakrishna</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Schneider</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Kilicoglu</surname>
          </string-name>
          ,
          <article-title>Assessing citation integrity in biomedical publications: corpus annotation and NLP models</article-title>
          ,
          <source>Bioinformatics</source>
          <volume>40</volume>
          (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <article-title>Review on progress of citation sentiment identification</article-title>
          ,
          <source>Information Studies: Theory &amp; Application</source>
          <volume>47</volume>
          (
          <year>2023</year>
          )
          <fpage>173</fpage>
          -
          <lpage>181</lpage>
          +
          <fpage>189</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <article-title>[7] SemMedDB database details</article-title>
          , https://ii.nlm.nih.gov/SemRep_SemMedDB_SKR/dbinfo_FML.shtml,
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Colil:</surname>
          </string-name>
          <article-title>Comments on literature in literature</article-title>
          , https://colil.dbcls.jp/browse/papers/,
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Qwen</surname>
          </string-name>
          , https://github.com/QwenLM,
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A.</given-names>
            <surname>Daowd</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Barrett</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Abidi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. S. R.</given-names>
            <surname>Abidi</surname>
          </string-name>
          ,
          <article-title>Building a knowledge graph representing causal associations between risk factors and incidence of breast cancer</article-title>
          ,
          <source>in: Public Health and Informatics</source>
          , IOS Press,
          <year>2021</year>
          , pp.
          <fpage>724</fpage>
          -
          <lpage>728</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>