<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Open Science Portal Based on Knowledge Graph</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>sily Bun</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>STFC UKRI</institution>
          ,
          <addr-line>Harwell Campus, Didcot OX11 0QX</addr-line>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
      </contrib-group>
      <fpage>134</fpage>
      <lpage>141</lpage>
      <abstract>
        <p>The work outlines an ongoing effort of building the Open Science Portal for STFC UKRI based on the knowledge graph assembled from various records of science. The graph is a result of interplay across the records that the organization maintains and the external quality records in reference databases supported elsewhere. The twofold role of persistent identifiers is illustrated as, on one hand, facilitators of building the knowledge graph and, on the other hand, as a specific means of the information enrichment within the graph once it is built. The business case and the implementation detail of the Portal is reported, as well as the projections for its further development and possible applications. The work is one of the outcomes of FREYA which is a Horizon 2020 project developing the persistent identifiers infrastructure and recommendations for Open Science.</p>
      </abstract>
      <kwd-group>
        <kwd>Open Science</kwd>
        <kwd>knowledge graph</kwd>
        <kwd>persistent identifiers</kwd>
        <kwd>EU project</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        The FREYA project [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] is developing the infrastructure for persistent identifiers (PIDs)
and recommendations for their effective use as part of the European and global
environment for Open Science. PIDs landscape is rich nowadays and includes identifiers
for digital objects such as research publications or datasets, as well as for real-world
entities such as people or organizations. The grand vision of FREYA is the "PID Graph"
that creates relationships across objects and entities with PIDs and provides a basis for
new services within research disciplines and across them.
      </p>
      <p>The FREYA partners develop pilot applications that exploit parts of the PID Graph
relevant to their respective research disciplines, also these applications in turn provide
contributions to the PID Graph by opening up the information assets within the
organizations using a common set of recommendations and where reasonable common
technological solutions, too.</p>
      <p>
        This work outlines a particular pilot application by STFC UKRI [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] in support of the
Open Science agenda. First, the business case for the Open Science Portal is described,
then the detail is given about data sources, technology used and examples of the Portal
content, then plans for the further development of the Portal are discussed.
      </p>
    </sec>
    <sec id="sec-2">
      <title>Business Case for STFC Open Science Portal</title>
      <p>
        STFC (Science and Technology Facilities Council) is a part of UKRI (UK Research
and Innovation) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] which is the national funding agency investing in science and
research in the UK. STFC is a funder of science and of the post-graduate education in the
UK, and a funder for the research conducted by the UK scientists on large-scale
scientific instruments abroad. STFC is also a research organization that operates or co-owns
large-scale scientific instruments (facilities) and high-performance computing across a
few UK locations, granting access to them to the UK and overseas visitor scientists who
conduct their own experiments.
      </p>
      <p>All the mentioned streams of STFC funding and funding-in-kind (facility time and
computation time awarded) result in research artefacts such as journal papers and
preprints, PhD theses, data and software. Tracking down these artefacts back to the
instruments, organizations and people involved is important for the evaluation of STFC role
in a number of research fields such as biomedicine, chemistry, materials science,
engineering, particle physics, astronomy, also of its role in higher education.</p>
      <p>Research artefacts that have been produced with the support of STFC funding or
funding-in-kind are reflected in records of science, e.g. bibliographic information for
journal papers, or records of data deposited in certain reference databases. These
records of science and the artefacts behind them are handled by a variety of information
systems within STFC, also some well-curated STFC-related records of science are
managed by external providers as parts of their larger collections (examples being
crystallography or biomedical databases or national services for PhD theses).</p>
      <p>To fully account for STFC funding and funding-in-kind, there is a need to
systematically collect and manage these records of science, including the discovery of
connections across them. Apart from this objective of having a better accountability for public
spending on science, there is an important aspect of knowledge preservation and
knowledge discovery in the spirit of Open Science that encourages and supports reuse
of research outcomes beyond the point of their origin. Open Science contributes to new
research by other research organizations and individuals, raises the public awareness of
science and its applications, also the organization itself can benefit from the more
explicit and context-rich representation of its records of science with a single point of
entry for their discovery.</p>
      <p>Accountability and Open Science aspects are interrelated and can be supported by
the same research information infrastructure that STFC are now building, with the first
prototype of such infrastructure receiving support of FREYA project and having the
focus on aspects that are specifically relevant to FREYA scope, i.e. persistent identifiers
as a means of knowledge discovery and integration.</p>
      <p>This new research information infrastructure is provisionally coined with the name
of STFC Open Science Portal. This is going to be a publically available resource for
the discovery of records of science that have been produced with the support of STFC
funding or other flavours of sponsorship such as facility time awarded to visitor
scientists. The records of science include information about research outcomes in any form
(publications, data, etc.), records of STFC funding or other flavours of sponsorship,
organizational context of STFC-supported research, as well as research attribution to
particular large-scale instruments (facilities and their beamlines).</p>
      <p>The Portal is going to be a multi-purpose research information infrastructure that can
be used by STFC and by external stakeholders. The Portal can contribute to knowledge
preservation and knowledge discovery, raise visibility of STFC-sponsored research,
demonstrate STFC adherence to the principles of Open Science, and support practical
applications of these principles to research impact studies, professional engagement
with other research organizations and funders, and to public engagement.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Data Sources and Technology for STFC Open Science Portal</title>
    </sec>
    <sec id="sec-4">
      <title>Implementation</title>
      <p>The Portal ingests metadata from a few sources, leaving data, full-text publications and
other research artefacts in their current respective locations, and represents the
integrated metadata as the knowledge graph. The metadata is subject to a moderate level of
harmonization when integrated, yet there is no intention to support higher levels of the
metadata harmonization or unification. Records of science in the external quality
sources are not ingested in the Portal but linked from it through the use of persistent
identifiers or in some cases by other record matching techniques.</p>
      <p>
        The sources where the Portal ingests metadata from:
 STFC publications repository [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]
 STFC data repository [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]
 Diamond Light Source bibliographic database [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]
 DataCite (records there that have been produced by STFC) [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]
 Unpaywall [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] to discover Open Access versions of publications
 (under consideration) Gateway to Research [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]
 (under consideration) Crossref COCI and other sources of citations [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]
The sources that the Portal links to:
 The Cambridge Structural Database [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]
 The British Library EThOS service [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]
 Europe PMC [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]
 Protein Data Bank in Europe [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]
 (under consideration) Zenodo [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]
For managing the integrated metadata, a community edition of the neo4j graph database
[
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] is used that is currently hosted on the developers’ computers and in the STFC
Intranet. The current size of the database is within tens of thousands of nodes and tens
of thousands of relationships; this has a potential to grow to hundreds of thousands or
a few millions of nodes, and hundreds of thousands or a few millions of relationships.
The inflation of the database beyond hundreds of thousands or low millions of nodes
and relationships is not expected at the moment.
      </p>
      <p>For the records ingestion, OAI-PMH endpoints, bespoke APIs or bulk export
features of the aforementioned sources have been used, which resulted in tabular or XML
files. These have been further processed using XSLT transformations and Unix shell
scripts, then uploaded in the graph database using standard neo4j tools.</p>
      <p>
        Relationships across records are produced using elements of them associated with
persistent identifiers where possible, otherwise fuzzy matching techniques are used,
such as measuring distance between corresponding metadata elements, as was in the
case of doctoral theses records [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. The enrichment of scholar communication with
persistent identifiers is a gradual process, and the broader PIDs proliferation should
allow more efficient records matching in the future, saving the effort of building
research knowledge graphs. This is a good example when best practices (of persistent
identifiers minting and use) can augment and in certain cases replace technology (of
fuzzy records matching).
      </p>
      <p>
        Exploration and visualizations of the resulted graph and its parts (subgraphs) are
made using queries in Cypher language [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] and standard Web components of neo4j.
Further experiments with visualization are going to involve Apache ECharts [20] and
JavaScript Cytoscape [21].
      </p>
      <p>An example of a subgraph extracted from the integrated graph that is going to
support the Open Science Portal is illustrated by Fig. 1. The metadata sources where these
records are harvested from do not necessarily know about each other, but their
integration in STFC Open Science Portal provides connections and allows cross-walks
between any of the records of science involved.</p>
      <p>
        The vision of the Portal is for it to become a single entry point for searching across
all records of science that could be publications, dataset descriptions, grants etc. The
full-text search across all these records of science are supported by cross-records
indexes created in the graph database and powered by Apache Lucene [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. At the
moment, an out-of-the box indexing algorithms are used that are implemented by default
in the database engine and that do not use the terminological structure of any research
domain. It would be interesting to explore the efficiency of discipline-specific indexing,
yet this requires a dedicated study which can be costly, too, especially if it involves
human evaluators. Also in the case of the STFC knowledge graph, there would be
natural obstacles to the application of discipline-specific text indices, as STFC research is
multi-discipline by its nature and thus the knowledge graph may incorporate records of
science from many branches of science, as different as biology or physics or
occasionally even archeology or history of art when samples are supplied by these disciplines
for investigation on synchrotron or neutron sources.
      </p>
      <p>
        The focus of the Open Science Portal for the foreseeable future is going to remain
the metadata, yet data visualizations are going to be incorporated opportunistically
where effort to produce them is moderate. The Fig. 2 illustrates a rich context around a
record in the INSDB (Inelastic Neutron Scattering Database) [22] which is published
as a part of the earlier mentioned STFC data repository [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The orange circle on the
top right with the “RecordedAs” relation to it represents a record in the STFC ePubs
repository that is currently unaware of the most of the information context in this graph
(apart from the corresponding journal paper). Some of the entities in this graph, e.g.
research instrument or organizations involved, do not currently bear persistent
identifiers but have a good potential for acquiring them through the further refinement of the
graph. The Fig. 3 contains visualization of the data file associated with the same INSDB
record. Visualization (which is interactive when accessed in a Web browser) is made
using Plotly JavaScript library [23]. URLs on the top refer to the data record and to the
paper that reports on the data, so that this visualization gives a good idea of both what
the data is and where further detail of it can be found through the resolution of persistent
identifiers that the paper and the data URLs are based on.
      </p>
    </sec>
    <sec id="sec-5">
      <title>Further Development of the Open Science Portal</title>
      <p>The development of the Open Science Portal has been focused so far on the data sources
integration in the back-end graph database, with illustrative visualizations to see what
usage scenarios are principally possible. The ultimate goal is to make the Portal a fully
functional Web application with features of a full-text search across a variety of entities
(publications, dataset and software descriptions, grants information) and with
contextual visualization of metadata (subgraphs). The target audience for the graphical user
interface are going to be STFC staff, visitor scientists, other funders and policy makers.</p>
      <p>Apart from the graphical user interface, an API based on the GraphQL technology
[24] is going to be implemented. Implementing a simple GraphQL endpoint to a graph
database is straightforward and can be realized with standard plugins. A more
sophisticated approach may be required though owing to the diverse nature of the metadata
sources integrated; this might be based on measuring the popularity of certain attributes
of the graph database nodes and relations [25]. The open API should allow third party
developers to build their own applications around the STFC records of science.</p>
      <p>The existing graph will benefit from disambiguation of certain entities in it, and
assigning them with persistent identifiers that are not immediately available as they were
not present in the source records that the graph is composed of. Organization names are
the prime candidate for such disambiguation; GRID.AC [26] or ROR [27] services can
be sources of persistent identifiers for organizations, also Crossref Funder Registry [28]
for organizations that are funders of science.</p>
      <p>It will be most productive to think of the Open Science Portal as a new piece of the
research information infrastructure that the organization itself and other parties can use
for multiple purposes. Some of the possible applications of the Portal are mentioned in
the “Business case” section, yet it is in the nature of every infrastructure to find its uses
that are beyond the initial thinking of the stakeholders needs. So when developing the
Open Science Portal, a good attention should be given to non-functional requirements
such as maintainability and extendibility of the underpinning knowledge graph.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements</title>
      <p>The work is supported by FREYA project funded by the European Commission under
the Horizon 2020 programme (Grant Agreement number 777523). The views expressed
are those of the author and not necessarily of the project or the funder.</p>
      <p>The author thanks Natalie Johnson (Cambridge Crystallography Data Centre) for
matching the Cambridge Structural Database records with those in STFC repositories.</p>
      <p>The author thanks Johanna McEntyre and Christine Ferguson (European
Bioinformatics Institute) for their advice on using Europe PMC and Protein Data Bank APIs.</p>
      <p>The author thanks Stewart Parker (STFC ISIS Neutron and Muon Source) for the
example of the Inelastic Neutron Scattering database record connected to the record of
the original experiment on ISIS facility.</p>
      <p>The author thanks his colleagues in FREYA project for their feedback on the
prototype demonstrations.
20. Apache ECharts data visualization framework, https://echarts.apache.org/, last accessed
2020/09/08.
21. Cytoscape JavaScript library, https://js.cytoscape.org/, last accessed 2020/04/29.
22. Inelastic Neutron Scattering Database,
https://www.isis.stfc.ac.uk/Pages/INSdatabase.aspx, last accessed 2020/04/29.
23. Plotly JavaScript Open Source Graphing Library, https://plotly.com/javascript/, last
accessed 2020/04/29.
24. GraphQL query language, https://graphql.org/, last accessed 2020/04/29.
25. Bunakov, V. Metadata Integration with Labeled-Property Graphs. In: Garoufallou, E.,
Fallucchi, F., William De Luca, E. (eds.) Metadata and Semantic Research. Communications
in Computer and Information Science, vol. 1057, pp. 441-448. Springer International
Publishing, Cham (2019).
26. GRID: Global Research Identifier Database, https://grid.ac/, last accessed 2020/04/29.
27. ROR: Research Organization Registry, https://ror.org/, last accessed 2020/04/29.
28. Crossref Funder Registry, https://www.crossref.org/services/funder-registry/, last accessed
2020/04/29.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1. FREYA project, https://www.project-freya.eu/,
          <source>last accessed</source>
          <year>2020</year>
          /09/08.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2. Science and Technology Facilities Council, https://stfc.ukri.org/,
          <source>last accessed</source>
          <year>2020</year>
          /05/31.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3. UK Research and Innovation, https://www.ukri.org/,
          <source>last accessed</source>
          <year>2020</year>
          /05/31.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4. ePubs:
          <article-title>STFC publications repository</article-title>
          , https://epubs.stfc.ac.uk/,
          <source>last accessed</source>
          <year>2020</year>
          /05/31.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5. eData:
          <article-title>STFC “Long Tail” data repository</article-title>
          , https://edata.stfc.ac.uk/,
          <source>last accessed</source>
          <year>2020</year>
          /05/31.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Diamond</given-names>
            <surname>Light</surname>
          </string-name>
          <article-title>Source bibliographic database</article-title>
          , https://publications.diamond.ac.uk/pubman/searchpublicationsquick, last accessed
          <year>2020</year>
          /05/31.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7. DataCite search, https://search.datacite.org/,
          <source>last accessed</source>
          <year>2020</year>
          /04/29.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Unpaywall</surname>
          </string-name>
          :
          <article-title>An open library of scholarly articles</article-title>
          , https://unpaywall.org/,
          <source>last accessed</source>
          <year>2020</year>
          /04/29.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9. Gateway to Research:
          <article-title>UKRI gateway to publicly funded research</article-title>
          and innovation https://gtr.ukri.org/,
          <source>last accessed</source>
          <year>2020</year>
          /04/29.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <article-title>COCI, the OpenCitations Index of Crossref open DOI-to-DOI citations</article-title>
          , http://opencitations.net/index/coci, last accessed
          <year>2020</year>
          /04/29.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11. The Cambridge Structural Database, https://www.ccdc.cam.ac.uk/solutions/csd-system/components/csd/,
          <source>last accessed</source>
          <year>2020</year>
          /04/29.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <article-title>The British Library EThOS service</article-title>
          , https://ethos.bl.uk/,
          <source>last accessed</source>
          <year>2020</year>
          /04/29.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Europe</surname>
            <given-names>PMC</given-names>
          </string-name>
          portal, https://europepmc.org/,
          <source>last accessed</source>
          <year>2020</year>
          /04/29.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Protein</surname>
          </string-name>
          Data Bank in Europe, https://www.ebi.ac.uk/pdbe/,
          <source>last accessed</source>
          <year>2020</year>
          /04/29.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15. Zenodo repository, https://zenodo.org/,
          <source>last accessed</source>
          <year>2020</year>
          /04/29.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <article-title>neo4j graph database</article-title>
          , https://neo4j.com/,
          <source>last accessed</source>
          <year>2020</year>
          /04/29.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Bunakov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Madden</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Integration of a National E-Theses Online Service with Institutional Repositories</article-title>
          .
          <source>Publications</source>
          <volume>8</volume>
          (
          <issue>2</issue>
          ),
          <volume>20</volume>
          (
          <year>2020</year>
          ). doi:
          <volume>10</volume>
          .3390/publications8020020
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18. Apache Lucene, https://lucene.apache.org/,
          <source>last accessed</source>
          <year>2020</year>
          /04/29.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Cypher</surname>
          </string-name>
          Query Language, https://neo4j.com/developer/cypher-query-language/, last accessed
          <year>2020</year>
          /04/29.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>