<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Semantically Mapping Science (SMS) Platform</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ali Khalili</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Peter van den Besselaar</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Al Koudous Idrissou</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Klaas Andries de Graaf</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Frank van Harmelen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, Vrije Universiteit Amsterdam</institution>
          ,
          <addr-line>NL</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Organization Sciences, Vrije Universiteit Amsterdm</institution>
          ,
          <addr-line>NL</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>Up to now, STI (Science, Technology, Innovation) studies are either rich but small scale (qualitative case studies) or large scale and under-complex { because they generally use only a single dataset like Patstat, Scopus, WoS (Web of Science), OECD STI indicators, etc., and therefore deploying only a few variables { determined by the data available. However, progress in the STI research eld (and the social sciences in general) depends in our view on the ability to do large-scale studies with often many variables speci ed by relevant theories. There is a need for studies which are at the same time big and rich. The aim of the Semantically Mapping Science (SMS) platform is to enable enriching and integration of heterogeneous data, ranging from tabular statistical data to unstructured data found on the Web, in order to exploit the huge amount of data that are `out there' in an innovative and meaningful way.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Social phenomena generally are complex, and understanding those phenomena
requires integrating and analyzing data from multiple sources. Up to now, STI
(Science, Technology, Innovation) studies are either rich but small scale
(qualitative case studies) or large scale and under-complex { because they generally
use only a single dataset like Patstat, Scopus, WoS (Web of Science), OECD
STI indicators, etc., and therefore deploying only a few variables { determined
by the data available. However, progress in the STI research eld (and the social
sciences in general) depends in our view on the ability to do large-scale
studies with often many variables speci ed by relevant theories. There is a need for
studies which are at the same time big and rich.</p>
      <p>In this paper, we present the Semantically Mapping Science (SMS) platform
as a means to enable enriching and integrating heterogeneous public and private
data, ranging from tabular statistical data to unstructured data found on the
Web, in an innovative and meaningful way. SMS is built as an open source
platform3 and is available online at http://sms.risis.eu.
As shown in Figure 1, the SMS platform consists of three main layers: data
layer, services layer and application layer. Data layer deals with data conversion,
storage and access plans. Service layer provides a set of Web services on top of the
created Linked Data to allow developing innovative applications. The application
layer is the terminal for end-users who interact with the SMS platform. In this
system paper, we brie y describe the main services and applications provided
by the SMS platform:
SMS platform at its conceptual model employs an entity-centric approach to
interlink heterogeneous datasets in the STI domain. As shown in Figure 2, the
following entity types are extracted after analysis of existing RISIS datasets and
their related open datasets: Funding Programs, Projects, Publications, Patents,
Persons, Organizations, Organization Rankings, Geo locations, Geo boundaries
and Geo statistical data. It is also possible to add new entity types based on
the research questions which need to be answered by the SMS infrastructure.
The main idea is creating a data network by linking and enriching the data, a
network which the social science user can access through the faceted browser.
By selecting the required entities and properties from the data network, the user
gets an overview of the data he/she is interested in. The platform produces in
the background the required SPARQL queries to retrieve the selected data from
multiple datasets in a required format for further analysis.
3 https://github.com/risis-eu/sms-platform</p>
      <p>Fig. 2: The Main Entity Types Involved in the SMS System.</p>
    </sec>
    <sec id="sec-2">
      <title>2.2 Data Curation</title>
      <p>
        Metadata helps potential users of a dataset to decide whether the dataset is
appropriate for their purposes or not. SMS platform has a collection of various
heterogeneous datasets that are not always publicly accessible due to privacy
issues, and often require a researcher to be physically at the dataset location. To
access these datasets, one needs to be granted an access request. This
administrative detour that a researcher has to endure prior to detecting which dataset
to use for a particular research question can reduce the number of SMS datasets
visitors. It has been shown that research publications that provide access to their
base data yield consistently higher citation rates than those that do not.
Therefore, to attract more users, to visit and cite RISIS datasets, SMS provides a
dataset metadata service and application - modeled using the Resource
Description Framework (RDF) - that allows researchers to search for data, and have
an in-depth understanding of the data without the need to directly access it [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
Metadata service powered by an intuitive UI allows dataset holders to describe
their datasets in a detailed, consistent and uniform way, store the description
and if needed modify the stored metadata. 4 The curated metadata are then
re ected on RISIS dataset's portal available at http://datasets.risis.eu.
      </p>
    </sec>
    <sec id="sec-3">
      <title>2.3 Browsing and Querying Datasets</title>
      <p>One of the objectives in developing the SMS platform was to enable non-Linked
Data experts to query and browse RDF datasets without having the knowledge
of SPARQL query language. There are currently two main approaches to make
information retrieval from SPARQL endpoints more usable: user interaction and
natural language (NL). In the category of user interaction-based query
generation, faceted browsing user interfaces are well-known techniques which provide</p>
      <sec id="sec-3-1">
        <title>4 see an screencast of the SMS metadata editor at https://youtu.be/p_2D3ydcx1U</title>
        <p>
          a convenient and user-friendly way to navigate through a wide range of data
collections [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. Faceted browsing UIs allow users to nd information without
apriori knowledge of its schema [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. A faceted interface has several advantages
over keyword search or NL queries: it allows exploration of an unknown dataset
since the system suggests restriction values at each step; it is a visual interface,
removing the need to write explicit queries; and it prevents dead-end queries, by
only o ering restriction values that do not lead to empty results [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
        </p>
        <p>
          SMS provides an adaptive component-based faceted browser environment5 on
top of the LD-R framework [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] to allow end-users explore STI related datasets
in an integrated way and to incorporate additional features for serendipitous
knowledge discovery (see Figure 3 for a screenshot).
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>2.4 Semantic Enrichment of Data</title>
      <p>SMS provides a set of services and applications that allow users to enrich their
data by adding complementary data to their current data. There are three
categories of data-enrichment services provided:
Named Entity Recognition. Named-entity recognition (NER) (also known as
entity identi cation and entity extraction) is a subtask of information extraction
that seeks to locate and classify named entities in text into prede ned categories
such as the names of persons, organizations, locations, expressions of times,
quantities, monetary values, percentages, etc. Given a dataset which has one
or more attributes with textual values, SMS NER service can extract named
entities from the text and more importantly connect the extracted entities to a
knowledge graph or taxonomy (which can then provide more data about those
entities). By default, SMS employs DBpedia Spotlight service for NER. However,</p>
      <sec id="sec-4-1">
        <title>5 see an screencast of the SMS faceted browser at https://youtu.be/9TMLKdGZExY</title>
        <p>any arbitrary NER service can be plugged into SMS NER service as long as the
output of service is reconciled to SMS named entities annotation model. SMS
provides an interactive UI to annotate a dataset using the NER service6.
Geo-enrichment. Geo-enrichment is an instrument to enrich data by linking
through geo-location. Many (open) datasets provide variables that are measured
at some level of geographical aggregation: e.g., environmental data, educational
data, or socio-economic data. In order to exploit these linking and enriching
possibilities, the SMS platform provides a variety of geo-services. The geo-services
are based on a series of open geo-resources, such as GADM, OpenStreetMap and
Flickr geotagged data. By integrating these geo-resources, the service can give
for an entity's address the geo-location up to 11 di erent levels. One practical
application we built for batch processing of addresses is a Google spreadsheet
addon7 which chains Google Geocoding API with our geo-boundary services. Given
addresses in a spreadsheet are enriched with di erent levels of administrative
boundaries and FUAs. The users are then able to export the extracted
boundaries and process them in geodata analysis tools such as CartoDB.8 We have
also developed a user interface for automatic geo-enrichment of linked datasets
in the SMS platform. The interface allows users to select an existing dataset and
geocode the whole dataset by selecting the right attributes in the dataset9.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>2.5 Data Linking</title>
      <p>Linking between entities in di erent datasets is a crucial element of the SMS
platform. Whether or not two entities should be considered equal depends not
only on their intrinsic properties, but also on the purpose or task for which the
entities are used. As an example, to study the success of scienti c
organizations, STI researchers need to align research organizations across datasets such
as GRID10 and OrgRef11 that describe organisations across various countries
including public and private research organisations. The 3M corporation, a large
multinational organisation with a substantial patent portfolio, occurs in both
datasets. GRID distinguishes between national 3M branches across six countries
3M (Canada), 3M (France), 3M (Germany), 3M (Israel), 3M(United Kingdom)
and 3M(United States), while OrgRef only refers to a single 3M entity. Should
these entities be designated as \the same" across these datasets? It depends.
For a study that aims to compare organizations at a global level, all branches of
`3M' should be considered the same. Whereas, for a study that compares
organizations for a comparison across countries, the Canadian and U.S. branches of
`3M' should be considered separately.</p>
      <p>SMS provides a novel approach called \Lenticular Lens" for building
contextspeci c links between entities of interest. These links are decorated with rich
metadata describing how, why, when and by whom they were generated. As
6 see an screencast of the NER UI at https://youtu.be/OcYNpVRP9_Q
7 https://docs.google.com/document/d/1JoJM7VF_ZaaAPbSjtgpydzRDYLvr-tROzhITGj0cH3w
8 see an screencast of the SMS Google spreadsheet add-on at https://youtu.be/qZGDD5RN7pI
9 see an screencast of the geo-enrciher UI at https://youtu.be/PFalWjluMR8
10 See https://grid.ac/
11 See http://www.orgref.org/web/download.htm
shown in Figure 4, SMS exposes an intuitive UI12 to allow end-users create their
own lenticular lenses available at http://lenticular-lens.risis.eu.
3 Use Cases
In order to demonstrate how the SMS platform can be used for research, we
describe several use cases at http://sms.risis.eu/usecases . The use cases
demonstrate di erent features of the platform in connection to addressing certain
challenges covering topics such as investigating network structure of research
organisations, browsing research data for temporal evolution of higher education,
analyzing the geography of innovation and the structure of research portfolio and
predicting Leiden Ranking from University environment factors.
12 see an screencast of the</p>
      <p>PLo4YbUaRFSnwJ9XJvp6rlIMsaw_rfKT9C
linking</p>
      <p>UI</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>M.</given-names>
            <surname>Hildebrand</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. van Ossenbruggen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Hardman</surname>
          </string-name>
          .
          <article-title>/facet: A browser for heterogeneous semantic web repositories</article-title>
          .
          <source>ISWC</source>
          , pages
          <volume>272</volume>
          {
          <fpage>285</fpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>A. K. Idrissou</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Khalili</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Hoekstra</surname>
          </string-name>
          , and P. V. den Besselaar.
          <article-title>Managing metadata for science, technology and innovation studies: The RISIS case</article-title>
          . In A.
          <string-name>
            <surname>Adamou</surname>
          </string-name>
          , E. Daga, and L. Isaksen, editors,
          <source>WHiSe</source>
          , volume
          <volume>1608</volume>
          <source>of CEUR Workshop Proceedings</source>
          , pages
          <volume>15</volume>
          {
          <fpage>20</fpage>
          . CEUR-WS.org,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>A.</given-names>
            <surname>Khalili</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Loizou</surname>
          </string-name>
          , and
          <string-name>
            <surname>F. van Harmelen</surname>
          </string-name>
          .
          <article-title>Adaptive linked data-driven web components: Building exible and reusable semantic web interfaces</article-title>
          .
          <source>In ESWC</source>
          , volume
          <volume>9678</volume>
          of Lecture Notes in Computer Science, pages
          <volume>677</volume>
          {
          <fpage>692</fpage>
          . Springer,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>E.</given-names>
            <surname>Oren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Delbru</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Decker</surname>
          </string-name>
          .
          <article-title>Extending faceted navigation for rdf data</article-title>
          .
          <source>In International semantic web conference</source>
          , volume
          <volume>4273</volume>
          , pages
          <fpage>559</fpage>
          {
          <fpage>572</fpage>
          . Springer,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>