<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Entity Co-occurrences Derived from Biomedical Literature in the PubChemRDF</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Bolton</string-name>
          <email>bolton@ncbi.nlm.nih.gov</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Qingliang Li</string-name>
          <email>qingliang.li@nih.gov</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sunghwan Kim</string-name>
          <email>kimsungh@ncbi.nlm.nih.gov</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Leonid Zaslavsky</string-name>
          <email>zaslavsk@ncbi.nlm.nih.gov</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tiejun Cheng</string-name>
          <email>chengt2@ncbi.nlm.nih.gov</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bo Yu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Evan E.</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Bethesda</institution>
          ,
          <addr-line>MD 20894</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Chemical</institution>
          ,
          <addr-line>Disease, Gene, Protein, PubChemRDF, Co-Occurrence</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health</institution>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>PubChem</institution>
          ,
          <addr-line>RDF, Semantic Web, Linked Data, Triplestore, SPARQL, Modeling, Drug</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>Named entities, such as chemicals/drugs, diseases, and genes/proteins, and their associations are not only important components of biomedical literature, but also the foundation of creating biomedical knowledge bases/graphs. This work details the expression of co-occurrence associations among chemicals, genes, proteins, and diseases in the Resource Description Framework (RDF) format within the PubChemRDF resource, which is freely accessible and publicly available. The co-occurrence model is populated into a triplestore with named entities and their associations that are derived from text mining of about 35 million biomedical references in PubMed. Use cases are provided to demonstrate the utility of the model. Together with meta-data modeling of the references including the information about the author, journal, grant, and funding agency, this data model can address pertinent biomedical questions through SPARQL queries and help exploit bio-medical knowledge in various user perspectives and use cases.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        PubChem (https://pubchem.ncbi.nlm.nih.gov) [
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1–3</xref>
        ] is an open chemical database at the National
Center for Biotechnology Information (NCBI), the National Library of Medicine (NLM), the U.S.
National Institutes of Health (NIH). Among many tools and services provided by PubChem are the
literature knowledge panels [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], which assist users in quickly finding important relationships between
chemicals, genes, proteins, and diseases. The literature knowledge panels for a given entity (e.g., a
chemical, gene, or protein) display a selected set of “co-occurrence neighbors”, which are defined as
any chemicals, genes, proteins, and diseases mentioned together with that entity in the biomedical
literature. In addition, the literature knowledge panel provides a sample of PubMed records that
comention the entity and those selected co-occurrence neighbors. The list of the co-occurrence neighbors
and relevant PubMed records can be downloaded for further analysis through the download button in
the panel. Note that the use of the term “co-occurrence neighbors” avoids confusion with the existing
2-dimensional (2-D) and 3-dimensional (3-D) chemical structure-based neighbor relationships [
        <xref ref-type="bibr" rid="ref5 ref6 ref7">5–7</xref>
        ].
The 14th International Semantic Web Applications and Tools for Health Care and Life Sciences (SWAT4HCLS) Conference, February 13–
EMAIL:
(S.
      </p>
      <p>Kim);</p>
      <p>2023 Copyright for this paper by its authors.</p>
      <p>
        The underlying data for the knowledge panel is derived from text mining the 35 million biomedical
references in PubMed, using the LeadMine software [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Chemicals, genes, proteins, and diseases are
extracted from the titles and abstracts of the PubMed records and the most relevant co-occurrence
neighbors are identified using statistical analysis and relevance-based sampling. A detailed explanation
of the method used to develop the knowledge panel is given in our previous paper [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        The present paper describes a data model that expresses the named entities and their co-occurrence
associations in the Resource Description Framework (RDF) format [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], along with the meta-data
modeling of the reference information (including the author, journal, grant, and funding agency). The
data model augments the existing RDF-formatted data (also known as PubChemRDF) [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] and helps
find answers to biomedical questions (as demonstrated in a recent study [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]) through SPARQL
Protocol and RDF Query Language queries [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Model design</title>
      <p>It is important to note that the co-occurrence neighboring relationships between entities (e.g.,
compound-disease) are not symmetrical due to the asymmetry of the co-occurrence scores and the
truncation of the neighbor lists. For example, when a disease is one of the top 1,000 co-occurrence
neighbors of a chemical, there is no guarantee that the chemical is among the top 1,000 co-occurrence
neighbors of the disease. This asymmetry is reflected in the directed graph in the RDF by means of
designating the subject and the object.</p>
      <p>
        Literature data is modeled in reference nodes, which are linked to the nodes representing their
metadata (e.g., journal, author, publication date, grant number, funding agency, and MeSH terms [22]),
using the relevant terms from FRAPO [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], DCMI [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], and PRISM [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] as predicates.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Applications</title>
      <p>In this section, we present three use cases of the co-occurrence RDF resource. Here, these use cases
assume that the co-occurrence RDF resource and PubChemRDF data are loaded into the triplestore of
Virtuoso [23]. It is possible to load these data into other triplestores or RDF-aware graph databases,
such as Apache Jena.
3.1.</p>
    </sec>
    <sec id="sec-4">
      <title>Use Case 1: Diseases Co-occurring with a Chemical</title>
      <p>
        The simplest use case of the co-occurrence RDF is to retrieve named entities commonly mentioned
with a query entity in PubMed articles (e.g., diseases or genes/proteins co-occurring with a chemical).
As an example, Figure 2 shows the SPARQL query that retrieves the top 25 diseases mentioned together
with indomethacin (CID 3715) in terms of their co-occurrence scores [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The diseases returned from
the query are listed in Table 1, along with their co-occurrence scores and preferred disease names. The
preferred disease name was retrieved from the PubChemRDF disease subdomain (with the term
prefLabel in the Simple Knowledge Organization System (SKOS) [24] as a predicate).
a Derived from the values computed from Formula (3) in Reference [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] by
multiplying by 100 and rounding to the nearest integer.
      </p>
      <p>The most commonly occurring disease with indomethacin in literature is “inflammation”, followed
by “ulcer”. It reflects the fact that indomethacin is a non-steroidal anti-inflammatory drug (NSAID)
and that NSAIDs’ common side effects include stomach ulcer. In essence, this query performs the same
data retrieval task used to create the chemical-disease co-occurrence knowledge panel, available on the
Summary page of indomethacin (https://pubchem.ncbi.nlm.nih.gov/compound/3715
#section=Chemical-Disease-Co-Occurrences-in-Literature). It is noteworthy that the co-occurrence
RDF allows the user to get an arbitrary number of co-occurrence neighbors (up to 1000, as explained
in the Model Design section), while the knowledge panel in the Compound Summary page shows
provides a maximum of 25 co-occurrence neighbors with the current settings.</p>
    </sec>
    <sec id="sec-5">
      <title>3.2. Use Case 2: References that Co-mentions a Particular Chemical and</title>
    </sec>
    <sec id="sec-6">
      <title>Disease Pair</title>
      <p>It is noteworthy that the diseases in Table 1 are related to the input chemical (indomethacin) in
different contexts. For example, indomethacin is used to treat inflammatory diseases like arthritis, while
it is also known to cause stomach ulcer (as a side effect). To understand the context of the relationship
between two entities, it is often necessary to get relevant articles that mention them together. The
SPARQL query for this task is shown in Figure 3, with the indomethacin–inflammation pair as an
example.</p>
      <p>
        In the query for Use Case 2, while the named entities are specified using the PubChem-specific
vocabulary (pcvocab:discussesAsDerivedByTextMining), the external vocabularies from
DMCI [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] and PRISM [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] are used to get the metadata for the reference (i.e., date, title, and journal).
The result of the query is shown in Table 2.
3.3.
      </p>
    </sec>
    <sec id="sec-7">
      <title>Use Case 3: Diseases Implicitly Related to a Chemical via Genes</title>
      <p>Use Case 3 intends to identify diseases related to a chemical via genes, by first identifying genes
commonly mentioned with the query chemical and then retrieving diseases co-occurring with those
genes. In this Use Case, while some of the resulting diseases may already be mentioned together with
the query chemical in scientific articles, others may not. This implicit relationship can serve as a good
starting point to formulate a new hypothesis to test in future studies. Figure 4 shows the SPARQL
query with maribavir (CID 471161) as an example.</p>
      <p>Maribavir is an antiviral drug approved in 2021 by the U.S. Food and Drug Administration (FDA)
for the treatment of posttransplant cytomegalovirus (CMV) infection. Because of its short history, this
drug has not been mentioned in many papers, compared to old drugs introduced in the market decades
ago (e.g., indomethacin). Table 3 shows the diseases retrieved from the query. The gene most
mentioned together with maribavir is the protein kinase, X-linked (PRKX) gene. While some of the
diseases co-occurring with PRKX are directly co-mentioned with maribavir in literature, other diseases,
including “genetic translocation”, “depressive disorder”, “ischemia”, and “neurodegenerative
diseases”, have not appeared with maribavir in PubMed records, implying implicit associations between
Maribavir and these diseases via the PRKX gene.</p>
    </sec>
    <sec id="sec-8">
      <title>4. Conclusions</title>
      <p>In this paper, we described the RDF data model that expresses the co-occurrence associations
between chemicals, genes, and diseases derived from biomedical literature. This data model allows
users to quickly identify chemicals, genes/proteins, and diseases mentioned together with a given named
entity (Use Case 1). In addition, the model can be used to get references that mention two entities
together, helping one to understand the context of the co-occurrence association between the entities
(Use Case 2). It can also be used to find an implicit link between entities that are not mentioned
together, through the common entities associated with them (Use Case 3).</p>
      <p>
        The underlying data used in the co-occurrence RDF was derived from text mining of 35 million
references available in PubMed as shown in our previous study [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. While this data is also used to
generate the PubChem literature knowledge panel, the co-occurrence RDF enables additional tasks. For
example, with the co-occurrence RDF, the user can work with a large set of relevant co-occurrence
neighbors (up to 1,000) for a given entity and automate this data retrieval task using a computer program
or script. The co-occurrence RDF model is an enhancement to the PubChemRDF ecosystem that can
facilitate exploring biomedical knowledge and seeking new discoveries in a semantic way. More
importantly, it is naturally connecting to other linked data resources in various scientific communities
to greatly enhance the usability and accessibility of biomedical data. The co-occurrence RDF data
generated in this study is freely available at Zenodo [25].
      </p>
    </sec>
    <sec id="sec-9">
      <title>5. Acknowledgements</title>
      <p>This work was supported by the National Center for Biotechnology Information of the National
Library of Medicine (NLM), National Institutes of Health.</p>
    </sec>
    <sec id="sec-10">
      <title>6. References</title>
      <p>[19] Akiko Aizawa, An information-theoretic perspective of tf–idf measures, Information Processing
&amp; Management 39 (2003) 45–65. doi:10.1016/S0306-4573(02)00021-3.
[20] Stephen Robertson, Understanding inverse document frequency: on theoretical arguments for</p>
      <p>IDF, Journal of Documentation 60 (2004) 503–520. doi:10.1108/00220410410560582.
[21] Anand Rajaraman and Jeffrey David Ullman, Mining of Massive Datasets. Cambridge University</p>
      <p>Press, Cambridge (2011). doi:10.1017/CBO9781139058452.
[22] Medical Subject Headings, https://www.nlm.nih.gov/mesh/, last accessed 2022/11/12.
[23] OpenLink Software: Virtuoso Homepage, https://virtuoso.openlinksw.com/, last accessed
2022/11/12.
[24] Thomas Baker, Sean Bechhofer, Antoine Isaac, Alistair Miles, Guus Schreiber, and Ed Summers,
Key choices in the design of Simple Knowledge Organization System (SKOS), Journal of Web
Semantics 20 (2013) 35–49. doi:10.1016/j.websem.2013.05.001.
[25] Qingliang Li, Sunghwan Kim, Leonid Zaslavsky, Tiejun Cheng, Bo Yu, and Evan Bolton,
Resource Description Framework (RDF) Modeling of Named Entity Co-occurrences Derived
from Biomedical Literature in the PubChemRDF [Data set], Zenodo (2023).
doi:10.5281/zenodo.7521846.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Sunghwan</given-names>
            <surname>Kim</surname>
          </string-name>
          , Exploring Chemical Information in PubChem,
          <source>Current Protocols</source>
          <volume>1</volume>
          (
          <year>2021</year>
          )
          <article-title>e217</article-title>
          . doi:
          <volume>10</volume>
          .1002/cpz1.
          <fpage>217</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Sunghwan</given-names>
            <surname>Kim</surname>
          </string-name>
          , Jie Chen, Tiejun Cheng, Asta Gindulyte, Jia He, Siqian He,
          <string-name>
            <given-names>Qingliang</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Benjamin A</article-title>
          .
          <string-name>
            <surname>Shoemaker</surname>
            ,
            <given-names>Paul A.</given-names>
          </string-name>
          <string-name>
            <surname>Thiessen</surname>
          </string-name>
          , Bo Yu, Leonid Zaslavsky, Jian Zhang, and Evan E. Bolton, PubChem in 2021:
          <article-title>new data content and improved web interfaces</article-title>
          ,
          <source>Nucleic Acids Research</source>
          <volume>49</volume>
          (
          <year>2021</year>
          )
          <fpage>D1388</fpage>
          -
          <lpage>D1395</lpage>
          . doi:
          <volume>10</volume>
          .1093/nar/gkaa971.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Sunghwan</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <article-title>Getting the most out of PubChem for virtual screening</article-title>
          ,
          <source>Expert Opinion on Drug Discovery</source>
          <volume>11</volume>
          (
          <year>2016</year>
          )
          <fpage>843</fpage>
          -
          <lpage>855</lpage>
          . doi:
          <volume>10</volume>
          .1080/17460441.
          <year>2016</year>
          .
          <volume>1216967</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Leonid</given-names>
            <surname>Zaslavsky</surname>
          </string-name>
          , Tiejun Cheng, Asta Gindulyte, Siqian He, Sunghwan Kim,
          <string-name>
            <given-names>Qingliang</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <surname>Paul Thiessen</surname>
          </string-name>
          , Bo Yu, and Evan E. Bolton, Discovering and Summarizing Relationships Between Chemicals, Genes, Proteins, and Diseases in PubChem,
          <source>Frontiers in Research Metrics and Analytics</source>
          <volume>6</volume>
          (
          <year>2021</year>
          )
          <article-title>689059</article-title>
          . doi:
          <volume>10</volume>
          .3389/frma.
          <year>2021</year>
          .
          <volume>689059</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Evan</surname>
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Bolton</surname>
          </string-name>
          , Sunghwan Kim, and
          <string-name>
            <surname>Stephen H. Bryant</surname>
          </string-name>
          ,
          <article-title>PubChem3D: Similar conformers</article-title>
          ,
          <source>Journal of Cheminformatics</source>
          <volume>3</volume>
          (
          <year>2011</year>
          )
          <article-title>13</article-title>
          . doi:
          <volume>10</volume>
          .1186/
          <fpage>1758</fpage>
          -2946-3-13.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Sunghwan</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Evan E.</given-names>
            <surname>Bolton</surname>
          </string-name>
          , and
          <string-name>
            <surname>Stephen H. Bryant</surname>
          </string-name>
          ,
          <article-title>Similar compounds versus similar conformers: complementarity between PubChem 2-D and 3-D neighboring sets</article-title>
          ,
          <source>Journal of Cheminformatics</source>
          <volume>8</volume>
          (
          <year>2016</year>
          )
          <article-title>62</article-title>
          . doi:
          <volume>10</volume>
          .1186/s13321-016-0163-1.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <article-title>[7] PubChemRDF neighbor subdomain</article-title>
          , https://pubchem.ncbi.nlm.nih.gov/docs/rdf-neighbor,
          <source>last accessed</source>
          <year>2023</year>
          /01/11.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Daniel</surname>
            <given-names>M.</given-names>
          </string-name>
          <article-title>Lowe and Roger A. Sayle, LeadMine: a grammar and dictionary driven approach to entity recognition</article-title>
          ,
          <source>J Cheminform</source>
          <volume>7</volume>
          (
          <year>2015</year>
          )
          <article-title>S5</article-title>
          . doi:
          <volume>10</volume>
          .1186/
          <fpage>1758</fpage>
          -2946-7-S1-S5.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <source>[9] RDF 1.1 Concepts</source>
          and Abstract Syntax, https://www.w3.org/TR/rdf11-concepts/,
          <source>last accessed</source>
          <year>2022</year>
          /11/12.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Gang</surname>
            <given-names>Fu</given-names>
          </string-name>
          , Colin Batchelor, Michel Dumontier, Janna Hastings, Egon Willighagen, and Evan Bolton,
          <article-title>PubChemRDF: towards the semantic annotation of PubChem compound and substance databases</article-title>
          ,
          <source>Journal of Cheminformatics</source>
          <volume>7</volume>
          (
          <year>2015</year>
          )
          <article-title>34</article-title>
          . doi:
          <volume>10</volume>
          .1186/s13321-015-0084-4.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Begoña</given-names>
            <surname>Talavera</surname>
          </string-name>
          <string-name>
            <given-names>Andújar</given-names>
            , Dagny Aurich, Velma T. E. Aho,
            <surname>Randolph R. Singh</surname>
          </string-name>
          , Tiejun Cheng, Leonid Zaslavsky, Evan E. Bolton, Brit Mollenhauer, Paul Wilmes, and Emma L.
          <article-title>Schymanski, Studying the Parkinson's disease metabolome and exposome in biological samples through different analytical and cheminformatics approaches: a pilot study</article-title>
          ,
          <source>Anal Bioanal Chem</source>
          <volume>414</volume>
          (
          <year>2022</year>
          )
          <fpage>7399</fpage>
          -
          <lpage>7419</lpage>
          . doi:
          <volume>10</volume>
          .1007/s00216-022-04207-z.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <source>[12] SPARQL 1</source>
          .
          <article-title>1 Query Language</article-title>
          , https://www.w3.org/TR/sparql11-query/,
          <source>last accessed</source>
          <year>2022</year>
          /11/12.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Michel</surname>
            <given-names>Dumontier</given-names>
          </string-name>
          ,
          <source>Christopher JO Baker</source>
          , Joachim Baran, Alison Callahan, Leonid Chepelev,
          <source>José Cruz-Toledo, Nicholas R. Del Rio</source>
          ,
          <string-name>
            <given-names>Geraint</given-names>
            <surname>Duck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Laura I. Furlong</given-names>
            , Nichealla Keath, Dana Klassen,
            <surname>Jamie P. McCusker</surname>
          </string-name>
          ,
          <string-name>
            <surname>Núria</surname>
          </string-name>
          Queralt-Rosinach, Matthias Samwald,
          <string-name>
            <surname>Natalia</surname>
            <given-names>VillanuevaRosales</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Mark D.</given-names>
            <surname>Wilkinson</surname>
          </string-name>
          , and
          <article-title>Robert Hoehndorf, The Semanticscience Integrated Ontology (SIO) for biomedical research and knowledge discovery</article-title>
          ,
          <source>J Biomed Semant</source>
          <volume>5</volume>
          (
          <year>2014</year>
          )
          <article-title>14</article-title>
          . doi:
          <volume>10</volume>
          .1186/2041-1480-5-14.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14] FRAPO, the Funding, Research Administration and Projects Ontology, https://sparontologies.github.io/frapo/current/frapo.html,
          <source>last accessed</source>
          <year>2022</year>
          /11/12.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>Dublin</given-names>
            <surname>Core Metadata Initiative (DCMI) Metadata</surname>
          </string-name>
          <string-name>
            <surname>Terms</surname>
          </string-name>
          , https://www.dublincore.org/specifications/dublin-core/dcmi-terms/,
          <source>last accessed</source>
          <year>2022</year>
          /11/12.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>Publishing</given-names>
            <surname>Requirements For Industry Standard Metadata (PRISM) Specification</surname>
          </string-name>
          <string-name>
            <surname>Package</surname>
          </string-name>
          , https://www.w3.org/Submission/prism/,
          <source>last accessed</source>
          <year>2022</year>
          /11/12.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Jon</surname>
            <given-names>Ison</given-names>
          </string-name>
          , Matúš Kalaš, Inge Jonassen, Dan Bolser, Mahmut Uludag,
          <string-name>
            <surname>Hamish</surname>
            <given-names>McWilliam</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>James</given-names>
            <surname>Malone</surname>
          </string-name>
          , Rodrigo Lopez, Steve Pettifer, and
          <string-name>
            <given-names>Peter</given-names>
            <surname>Rice</surname>
          </string-name>
          ,
          <article-title>EDAM: an ontology of bioinformatics operations, types of data and identifiers, topics and formats</article-title>
          ,
          <source>Bioinformatics</source>
          <volume>29</volume>
          (
          <year>2013</year>
          )
          <fpage>1325</fpage>
          -
          <lpage>1332</lpage>
          . doi:
          <volume>10</volume>
          .1093/bioinformatics/btt113.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Christopher</surname>
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Manning</surname>
          </string-name>
          , Prabhakar Raghavan, and Hinrich Schütze, Introduction to Information Retrieval. Cambridge University Press, Cambridge (
          <year>2008</year>
          ). doi:
          <volume>10</volume>
          .1017/CBO9780511809071.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>