<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Work done prior to joining Amazon
$ fabio.giachelle@unipd.it (F. Giachelle)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>CoreKB: A Web-based Platform for Searching Reliable Facts over a Medical Knowledge Base⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Fabio Giachelle</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stefano Marchesin</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gianmaria Silvello</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Omar Alonso</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Amazon</institution>
          ,
          <addr-line>Santa Clara, California</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Information Engineering, University of Padua</institution>
          ,
          <addr-line>Padua</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>CoreKB is a web-based platform enabling users to search for reliable scientific facts concerning gene expression-cancer associations over a medical Knowledge Base (KB). CoreKB provides a streamlined interface for searching either using natural language queries or by exploiting structured facets providing autocomplete facilities. It is designed to simplify information access and search of scientific facts targeting healthcare stakeholders (i.e., clinicians, physicians, and researchers). CoreKB aims at presenting the user a comprehensive overview of the scientific evidence supporting a medical fact, fully connected with ontology-based entities and well-defined literature resources. In addition, CoreKB provides the user a quantitative comparison of the possible gene-cancer associations related to a specific fact, thus enabling users to assess the degree of agreement among the evidence support.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Knowledge Discovery</kwd>
        <kwd>Fact Search</kwd>
        <kwd>Gene Cancer Associations</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The World Health Organization has identified cancer prevention as a critical public health
challenge of the 21st century, with an estimated 32% increase in cancer cases by 20401. To
address this challenge, cancer research has increasingly relied on microarray and next-generation
sequencing technologies, generating vast amounts of experimental data on gene
expressioncancer interactions [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ], which is crucial for cancer diagnostics, prognosis, and therapies [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ].
However, the huge amounts of data and their heterogeneous nature poses hindrances to efective
analyses and secondary data re-use. In this regard, scientific literature is a critical source to
complement and validate this data. Nevertheless, extracting, combining, and interpreting
multiple scientific claims to produce an explanatory overview concerning a fact is a challenging
and time-consuming task [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Yet, it is crucial to shed a light on the subject and provide useful
insights – thus simplifying its comprehension, providing deeper understanding, and making it
more accessible.
      </p>
      <p>
        In this context, eficient and reliable methods to store and organize knowledge from various
sources have become increasingly important, and KBs have emerged as essential resources for
cancer researchers [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ]. KBs are designed to arrange structured information that can be used
to support data-driven research. To this aim, a graph-based representation is usually adopted;
it consists of nodes representing concepts/entities whereas the edges represent relationships
among them. Specifically, the Resource Description Framework ( RDF) format can be used to
provide an accessible, interoperable, and machine-readable representation of the underlying
data that can be queried with SPARQL [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. However, searching within KBs can be a challenging
task for multiple reasons. First, KBs may lack a fixed schema, making it dificult to comprehend
the organization of data and formulate SPARQL queries to extract the required information.
Furthermore, the absence of a fixed schema may result in uncertainty and confusion while
interpreting search results. Secondly, even for expert users, looking for specific information in
a KB via SPARQL queries can be cumbersome and time-consuming [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <p>
        Moreover, the massive amount of information contained in a KB can pose hindrances in
locating and extracting specific data of interest. This is compounded by the technical nature
of the queries needed to retrieve the desired information, resulting in an additional layer of
complexity, especially for non-technologically savvy users [
        <xref ref-type="bibr" rid="ref10 ref11 ref12">10, 11, 12</xref>
        ].
      </p>
      <p>Cancer research also relies on specialized KBs adopting domain-specific jargon, ontologies,
and terminologies resources with a large presence of complex terms, acronyms, and
morphosyntactic variants, thus making query formulation and efective search even more challenging [ 13].</p>
      <p>
        As a consequence, researchers may encounter dificulties when attempting to retrieve
information of interest from KBs, thus limiting efective data discovery, re-use, and new insights.
Therefore, there is an urgent need for intuitive and user-friendly search applications [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] that
can facilitate and speed up the retrieval of relevant data over KBs.
      </p>
      <p>To address these issues, we propose CoreKB, an easy-to-use web-based search platform
that builds upon one of the largest KBs containing fine-grained facts about gene
expressioncancer associations [14]. The KB comprises over 230,000 reliable and unreliable facts extracted
semi-automatically from the scientific literature by mining PubMed.</p>
      <p>CoreKB is publicly available at http://gda.dei.unipd.it along with a demonstration video
(http://gda.dei.unipd.it/static/videos/demo.mp4) presenting its features. The platform provides
a streamlined interface for searching and exploring knowledge about gene expression-cancer
associations. CoreKB allows users to query the system using either natural language queries
or facets, enabling search for both entities and facts easily. In this regard, CoreKB provides
several features (e.g., infometrics and entity cards) to support specialized users, such as medical
researchers and clinicians, who are interested in searching for knowledge about gene
expressioncancer associations quickly and without requiring the exact terminology or entity identifiers.</p>
      <p>As a distinctive trait, CoreKB embraces a fact-oriented approach, which aims at providing
users with a comprehensive overview of the key information concerning each fact. The overview
provided consists of aggregated data including (i) the gene class distribution among the sentences
extracted from the scientific literature; (ii) the number of supporting and conflicting sentences
(iii) and the number of supporting publications per year, enabling users to assess the consensus
supporting a fact. In summary, CoreKB ofers a powerful platform for comprehensive knowledge
discovery in precision medicine.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>Despite the importance of providing a fact-level overview for a scientific claim – i.e., showing
aggregated information describing a scientific fact overall, by combining data coming from
several scientific publications and resources – most of the state-of-the-art methods and tools
only focuses on providing point-wise sentence-level information – thus not presenting the
big-picture understanding. CoreKB aims to fill this gap by providing a comprehensive,
literaturesupported overview for each scientific fact. In this way, CoreKB difers from the approaches
described below.</p>
      <p>DEXTER [15], a text mining method for extracting associations from scientific literature,
ofers a basic search endpoint that requires users to search for gene-cancer pairs and does not
support natural language queries. OncoMX [16] provides a unified and user-friendly interface for
various datasets, but its literature mining capabilities are limited. OncoSearch [17] enables users
to search for sentences that mention gene expression changes in cancer and ofers advanced
faceted search features. However, it lacks natural language search and does not provide entity
cards or infometrics to support evidence. DisGeNET [18] has a well-designed faceted navigation
system, but it only returns gene-disease associations at the sentence level, making it dificult
to establish reliable associations for gene-disease pairs at the fact level. Finally, BioKB [19]
provides access to the semantic content of biomedical literature but does not support natural
language search and focuses on exploring and visualizing the structure of the underlying KB.</p>
    </sec>
    <sec id="sec-3">
      <title>3. COREKB Search Platform</title>
      <sec id="sec-3-1">
        <title>3.1. Knowledge Base Construction</title>
        <p>To construct the KB that serves as the foundation for CoreKB, we employed a system that
collects text from scientific literature and processes it to obtain sentences. These sentences then
undergo Named Entity Recognition and Disambiguation (NERD) annotation, which identifies
gene-cancer pairs. The NERD annotations are then subjected to bootstrapping and deployment
processes. During bootstrapping, fine-grained relationships between entities are manually
annotated and used to (i) train Relation Extraction (RE) methods and (ii) populate the KB. In
the deployment process, the trained RE methods are used to automatically extract facts from
sentences and populate the KB.</p>
        <p>Each fact in the KB is associated with a probability distribution that reflects the likelihood
of a specific gene-cancer association. These probabilities are used to perform reliability tests,
which determine whether facts are reliable or unreliable based on the quantity and level of
consensus of collected evidence. By adopting a probabilistic approach, the system can capture
the inherent uncertainty in scientific discourse and help users understand the strength of the
evidence supporting a particular gene-cancer association. The KB contains 23,879 genes and</p>
        <p>Researchers &amp; Clinicians</p>
        <p>CORE System
NERD + RE</p>
        <p>Reliability Testing + KB
population</p>
        <p>Index</p>
        <p>Request
11,530 cancer diseases, encompassing a total of 230,000 fine-grained facts. These facts are
supported by 1,037,845 sentences taken from 251,038 research articles [14].</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Architecture</title>
        <p>CoreKB consists of several components including (i) a React-based web interface for the
frontend; (ii) a Django-based business logic for the back-end providing support for REST APIs and
services; (iii) a PostgreSQL database aligned with an instance of the Virtuoso RDF triple store for
storing the KB data; (iv) a Redis server broker for eficient in-memory data store and access and
(v) a Python-based search component that performs NERD on entity mentions in user queries
and then performs a structured search in the database to retrieve the matching facts and the
related information. To this end, the search component relies on a Redis in-memory dictionary
of entities returning for each entity name and synonyms the corresponding unique identifier.</p>
        <p>When a user-provided query is entered in the system, the search component assigns a score to
each entity that is maximum in case of an exact match or proportional to the number of matching
terms in case of a partial match. To avoid favoring long-named entities at the disadvantage
of those with short names, the score is discounted proportionally to the entity name’s length.
If a single entity is recognized, all the related facts are ordered by scientific evidence support.
Similarly, if multiple entities are identified, the facts regarding the most matching gene-cancer
pairs are promoted in the ranking of results. It’s worth noting that CoreKB is not limited to
human gene-cancer pairs, instead, it contains facts regarding any living organism. Nevertheless,
human-specific genes are ranked on top.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. User Interaction and Interface features</title>
        <p>The CoreKB platform enables users to search for factual information related to gene-cancer
associations supported by scientific literature, as well as entities such as genes and cancer
diseases. Users can search using free-text or structured search facilities with autocomplete
features. The interface allows users to switch between the two search modes via the settings
menu and provides clickable sample queries to demonstrate the system’s capabilities. The search
results are presented in cards that display gene class, cancer label, statistics on supporting and
conflicting sentences, key examples of supporting and conflicting sentences, and bibliometrics
on the number of supporting publications per year. The reliability of the fact is indicated by
a circle that is colored diferently according to the informativeness and reliability of the fact
i.e., green for reliable facts, gray for facts missing proper support, and red for unreliable facts.
On the right side of each fact card, specific information about the gene and cancer involved is
presented, including links to the corresponding entries in the National Center for Biotechnology
Information (NCBI)2 and Linked Life Data3 platforms. Users can access additional quantitative
and aggregated information on dedicated landing pages, including related facts and supporting
and conflicting sentences. The landing page for a particular gene displays all the available
information about the gene, including its symbol, full name, type, synonyms, designations,
summary, and gene class distribution for diferent cancer diseases. The sentences involving the
gene are presented in a table that users can filter and sort arbitrarily. Action buttons enable
users to expand or collapse each card and choose the columns to be shown in the sentence table.</p>
        <p>Figure 3 shows the landing page for a given entity, that is, in this case the human gene ERBB2.
The interface comprises two major cards (A, B). The first card (A) displays comprehensive
information about the gene, including its symbol (1), full name (3), type (4), synonyms (5),
designations (6), last modified date (7), summary (8), and the gene class distribution – i.e., the
role of the gene in a specific gene-cancer association which can assume one of alternatives
2https://www.ncbi.nlm.nih.gov/gene/
3http://linkedlifedata.com/</p>
        <p>B
oncogene, biomarker, tumor suppressor gene – for several cancer diseases presented both as a
textual statement (9) and depicted also in a horizontal bar chart (10). It is worth noting that all
the numbers reported in the textual statement are clickable and allows the users to visualize
on a new tab the specific gene-cancer associations of interest (e.g., for which cancer diseases
the ERBB2 gene is indicated to be an oncogene). To save space on the interface, long textual
information such as the gene summary is truncated by default, but it can be expanded or
collapsed by clicking on the dedicated button after the ellipsis.</p>
        <p>The second card (B) provides a list of sentences related to the gene. The sentences are
presented in a dynamic table that allows users to filter and sort them according to their preferences.
For instance, in Figure 3, only the sentences where ERBB2 is reported as a biomarker are
displayed because of the user-provided value biomarker for the Gene class column (13). In case of
long sentences, users can either (i) resize columns; (ii) going over with the mouse on a sentence
to view the full sentence text in a small tool-tip or (iii) click on the sentence to visualize it in a
separate pop-up. Moreover, there are two buttons reported on the top-right corner of each card
so that users can expand/collapse each card alternatively to fit the full-screen size (2, 12) and
choose the columns to be shown in the sentence table (11).</p>
        <p>Most of the information reported in the sentence table of Figure 3.B are clickable, including
(i) the sentence; (ii) the scientific fact; (iii) the gene; (iv) the disease; (v) the PubMed link to the
publication from which the sentence has been extracted and (vi) the publication journal. When
the user clicks on a sentence, a pop-up appears to show all the relevant information concerning
the sentence such as the sentence identifier, its textual content, the gene, the cancer disease,
and their relationship. Figure 4 shows the informative pop-up for the sentence’s fact “ERBB2 is
a biomarker for Stomach Carcinoma”, which is reported on top of the window card. Then the
pop-up reports other information concerning the sentence, the gene, the cancer disease, and
their association in separated collapsible menus. Each sentence and fact in CoreKB is univocally
identified with a permanent URL enabling point-wise access to the information specific for
an individual resource. Indeed, when the user clicks on the sentence identifier in Figure 4.A
the corresponding landing page for the sentence is shown, as depicted in Figure 4.B. Note that
in the landing page, users can access provenance information about the sentence, that is, the
PubMed URL to the original publication and a link to the corresponding journal entry in the
SCImago Portal4. Similarly, when users click on a fact the corresponding landing page for the
scientific claim is shown providing a holistic perspective enabling to assess overall the literature
consensus among the supporting/conflicting evidences.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusions</title>
      <p>CoreKB is a web-based search platform designed to provide users with reliable scientific facts on
gene expression-cancer associations. The platform ofers natural language query support as well
as structured facet search features integrated with autocomplete facilities. As a distinguishing
feature, CoreKB focuses on presenting fact-oriented information rather than adopting a
sentencelevel approach, where only single, independent results are presented. To this end, CoreKB
combines multiple aggregated information from several literature resources either supporting
or conflicting with the scientific claim of interest. This approach ofers a comprehensive
overview of the scientific evidence associated with a given fact. Moreover, CoreKB provides
users with a quantitative comparison of the possible gene-cancer associations, making it easier
to determine if there is a consensus on a specific gene’s role. Although CoreKB focus is on gene
expression-cancer associations, its model can be adapted and reused to accommodate other
types of relationships with the appropriate modifications.</p>
      <p>The platform aims to support clinicians and researchers by providing fast search, access,
and consultation of reliable scientific findings along with validating literature evidence. Hence,
simplifying knowledge discovery and promoting a serendipity-oriented perspective. As future
work, we plan to improve CoreKB according to the requisites and feedback of clinicians, which
we plan to collect by conducting a user study.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <p>The work was supported by the ExaMode project as part of the EU H2020 program under Grant
Agreement no. 825292.
[13] M. Agosti, S. Marchesin, G. Silvello, Learning Unsupervised Knowledge-Enhanced
Representations to Reduce the Semantic Gap in Information Retrieval, ACM Trans. Inf. Syst. 38
(2020) 38:1–38:48. URL: https://doi.org/10.1145/3417996. doi:10.1145/3417996.
[14] S. Marchesin, L. Menotti, G. Silvello, O. Alonso, CORE: Gene Expression-Cancer
Knowledge Base, 2023. URL: https://doi.org/10.5281/zenodo.7577127. doi:10.5281/zenodo.
7577127.
[15] S. Gupta, H. Dingerdissen, K. E. Ross, Y. Hu, C. H. Wu, R. Mazumder, K. Vijay-Shanker,
DEXTER: disease-expression relation extraction from text, Database J. Biol. Databases
Curation 2018 (2018) bay045.
[16] H. M. Dingerdissen, F. Bastian, K. Vijay-Shanker, M. Robinson-Rechavi, A. Bell, N. Gogate,
S. Gupta, E. Holmes, R. Kahsay, J. Keeney, H. Kincaid, C. H. King, D. Liu, D. J. Crichton,
R. Mazumder, OncoMX: A Knowledgebase for Exploring Cancer Biomarkers in the Context
of Related Cancer and Healthy Data, JCO Clin. Cancer Inform. (2020) 210–220.
[17] H. J. Lee, T. C. Dang, H. Lee, J. C. Park, OncoSearch: cancer gene search engine with
literature evidence, Nucleic Acids Res. 42 (2014) 416–421.
[18] J. P. González, J. M. Ramírez-Anguita, J. Saüch-Pitarch, F. Ronzano, E. Centeno, F. Sanz, L. I.</p>
      <p>Furlong, The DisGeNET knowledge platform for disease genomics: 2019 update, Nucleic
Acids Res. 48 (2020) D845–D855.
[19] M. Biryukov, V. Grouès, V. P. Satagopam, BioKB - Text mining and semantic technologies
for the biomedical content discovery, in: Proc. of the 10th International Conference
on Semantic Web Applications and Tools for Health Care and Life Sciences (SWAT4LS
2017), Rome, Italy, December 4-7, 2017, volume 2042 of CEUR Workshop Proceedings,
CEUR-WS.org, 2017. URL: http://ceur-ws.org/Vol-2042/paper5.pdf.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>F.</given-names>
            <surname>Giachelle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Marchesin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Silvello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Alonso</surname>
          </string-name>
          ,
          <article-title>Searching for reliable facts over a medical knowledge base</article-title>
          ,
          <source>in: Proc. of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          ,
          <string-name>
            <surname>SIGIR</surname>
          </string-name>
          <year>2023</year>
          , Taipei, Taiwan,
          <source>July 23-27</source>
          ,
          <year>2023</year>
          , ACM,
          <year>2023</year>
          . (in print).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>C.</given-names>
            <surname>Manzoni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. A.</given-names>
            <surname>Kia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Vandrovcova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hardy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. W.</given-names>
            <surname>Wood</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. A.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ferrari</surname>
          </string-name>
          , Genome,
          <source>Transcriptome and Proteome: the Rise of Omics Data and Their Integration in Biomedical Sciences, Briefings in Bioinformatics</source>
          <volume>19</volume>
          (
          <year>2016</year>
          )
          <fpage>286</fpage>
          -
          <lpage>302</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>P.</given-names>
            <surname>Borry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. B.</given-names>
            <surname>Bentzen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Budin-Ljøsne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. C.</given-names>
            <surname>Cornel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. C.</given-names>
            <surname>Howard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Feeney</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L</given-names>
            .
            <surname>Jackson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Mascalzoni</surname>
          </string-name>
          , Á. Mendes,
          <string-name>
            <given-names>B.</given-names>
            <surname>Peterlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Riso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Shabani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Skirton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sterckx</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Vears</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wjst</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Felzmann</surname>
          </string-name>
          ,
          <article-title>The Challenges of the Expanded Availability of Genomic Information: an Agenda-Setting Paper</article-title>
          ,
          <source>J. Community Genet</source>
          .
          <volume>9</volume>
          (
          <year>2018</year>
          )
          <fpage>103</fpage>
          -
          <lpage>116</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>B.</given-names>
            <surname>Neary</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Qiu</surname>
          </string-name>
          ,
          <article-title>Identifying Gene Expression Patterns Associated with DrugSpecific Survival in Cancer Patients</article-title>
          ,
          <source>Scientific Reports</source>
          <volume>11</volume>
          (
          <year>2021</year>
          )
          <fpage>1</fpage>
          -
          <lpage>12</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Dugger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Platt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Goldstein</surname>
          </string-name>
          ,
          <article-title>Drug development in the era of precision medicine</article-title>
          ,
          <source>Nat. Rev. Drug. Discov</source>
          .
          <volume>17</volume>
          (
          <year>2018</year>
          )
          <fpage>183</fpage>
          -
          <lpage>196</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S.</given-names>
            <surname>Marchesin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Giachelle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Marini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Atzori</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Boytcheva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Buttafuoco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ciompi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. M.</given-names>
            <surname>Di Nunzio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Fraggetta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Irrera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Primov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Vatrano</surname>
          </string-name>
          , G. Silvello,
          <article-title>Empowering Digital Pathology Applications through Explainable Knowledge Extraction Tools</article-title>
          ,
          <source>Journal of Pathology Informatics</source>
          <volume>13</volume>
          (
          <year>2022</year>
          )
          <article-title>100139</article-title>
          . doi:https://doi.org/10. 1016/j.jpi.
          <year>2022</year>
          .
          <volume>100139</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Warner</surname>
          </string-name>
          ,
          <article-title>A Review of Precision Oncology Knowledgebases for Determining the Clinical Actionability of Genetic Variants, Front</article-title>
          .
          <source>Cell Dev. Biol</source>
          .
          <volume>8</volume>
          (
          <year>2020</year>
          ). URL: https://www. frontiersin.org/articles/10.3389/fcell.
          <year>2020</year>
          .
          <volume>00048</volume>
          . doi:
          <volume>10</volume>
          .3389/fcell.
          <year>2020</year>
          .
          <volume>00048</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S.</given-names>
            <surname>Marchesin</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>Silvello, TBGA: a large-scale gene-disease association dataset for biomedical relation extraction</article-title>
          ,
          <source>BMC Bioinform</source>
          .
          <volume>23</volume>
          (
          <year>2022</year>
          )
          <fpage>111</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>G.</given-names>
            <surname>Weikum</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X. L.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Razniewski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. M.</given-names>
            <surname>Suchanek</surname>
          </string-name>
          ,
          <article-title>Machine Knowledge: Creation and Curation of Comprehensive Knowledge Bases, Found</article-title>
          .
          <source>Trends Databases</source>
          <volume>10</volume>
          (
          <year>2021</year>
          )
          <fpage>108</fpage>
          -
          <lpage>490</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>W.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <article-title>Proactive natural language search engine: tapping into structured data on the web</article-title>
          ,
          <source>in: Proc. of the Joint</source>
          <year>2013</year>
          EDBT/ICDT Conferences, EDBT '13,
          <string-name>
            <surname>Genoa</surname>
          </string-name>
          , Italy, March
          <volume>18</volume>
          -22,
          <year>2013</year>
          , ACM,
          <year>2013</year>
          , pp.
          <fpage>143</fpage>
          -
          <lpage>148</lpage>
          . URL: https://doi.org/10.1145/2452376.2452394. doi:
          <volume>10</volume>
          . 1145/2452376.2452394.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>H.</given-names>
            <surname>Bast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Buchhold</surname>
          </string-name>
          , E. Haussmann,
          <article-title>Semantic Search on Text and Knowledge Bases, Found</article-title>
          .
          <source>Trends Inf. Retr</source>
          .
          <volume>10</volume>
          (
          <year>2016</year>
          )
          <fpage>119</fpage>
          -
          <lpage>271</lpage>
          . URL: https://doi.org/10.1561/1500000032. doi:
          <volume>10</volume>
          . 1561/1500000032.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>A.</given-names>
            <surname>Badan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Benvegnù</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Biasetton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Bonato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Brighente</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Cenzato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Ceron</surname>
          </string-name>
          , G. Cogato,
          <string-name>
            <given-names>S.</given-names>
            <surname>Marchesin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Minetto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Pellegrina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Purpura</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Simionato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Soleti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Tessarotto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Tonon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Vendramin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <article-title>Towards Open-Source Shared Implementations of Keyword-Based Access Systems to Relational Data, in: Proc. of the Workshops of the EDBT/ICDT 2017 Joint Conference</article-title>
          (EDBT/ICDT 2017), Venice, Italy, March
          <volume>21</volume>
          -24,
          <year>2017</year>
          , volume
          <volume>1810</volume>
          <source>of CEUR Workshop Proceedings, CEUR-WS.org</source>
          ,
          <year>2017</year>
          . URL: https://ceur-ws.
          <source>org/</source>
          Vol-1810/KARS_paper_01.pdf.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>