<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>cLODg - Conference Linked Open Data Generator</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Anna Lisa Gentile</string-name>
          <email>a.gentile@sheffield.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrea Giovanni Nuzzolese</string-name>
          <email>andrea.nuzzolese@istc.cnr.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, University of She eld</institution>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Semantic Technology Lab, ISTC-CNR.</institution>
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper we describe cLODg (conference Linked Open Data generator), a set of tools to collect, re ne and produce Linked Data to describe a scienti c conference and its publications, participants and events. cLODg is an open source project, which has the aim to encourage conference metadata publication and foster collaborative e orts in this direction between researchers and publishers.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        A good practise in the semantic Web community is to encourage the publication
\eating our own dog food" [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. The main example is the Semantic Web Dog Food 3
(SWDF), a corpus that collects linked data about papers, people, organizations,
and events related to academic conferences. Currently, all main semantic Web
conferences and related events publish their data as linked data on SWDF, but
for many other conferences, events and publication venues information is still
not available in a structured and linked form.
      </p>
      <p>There are two main challenges to pursue this task: (i) the knowledge of
available vocabularies to represent the data and (ii) the availability of tools
to ease the task of data acquisition, conversion, augmentation, veri cation and
nally publication.</p>
      <p>In this work we present cLODg (conference Linked Open Data generator),
a tool that provides a formalized process for the conference metadata
publication work ow. cLODg is an Open Source solution currently under
development4, which has been used to gather and publish metadata for ESWC20145
and ESWC20156.</p>
      <p>The main contributions of cLODg is a formalized, open source work ow which
provides: (i) Facilities to gather/convert conference data from di erent source
formats. (ii) The possibility to represent such data using di erent vocabularies.
? Part of this research has been sponsored by the EPSRC funded project LODIE:</p>
      <p>Linked Open Data for Information Extraction, EP/J019488/1
3 SWDF: http://data.semanticweb.org
4 https://github.com/AnLiGentile/cLODg
5 http://2014.eswc-conferences.org/
6 http://2015.eswc-conferences.org/
(iii) Facilities to involve conference participants in the loop and allow the
collection of additional information and the veri cation of automatically generated
data. (iv) Facilities to represent additional and non-conventional information,
not currently captured by existing vocabularies, using Ontology Patterns7. (v)
The serialization of data in di erent output formats, including e cient
representations for mobile app consumption.</p>
      <p>The main advantage is the open source and modular nature of the work,
with the primary goal to encourage the collaboration between researchers and
publishers towards increasing the availability of structured scholarly data.
2</p>
    </sec>
    <sec id="sec-2">
      <title>State of the art</title>
      <p>
        The rst considerable e ort to o er comprehensive semantic descriptions of
conference events is represented by the metadata projects at ESWC 2006 and ISWC
2006 conferences [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], with the Semantic Web Conference Ontology [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] being the
vocabulary of choice to represent such data.
      </p>
      <p>
        Increasing number of initiatives are pursuing the publication about
conferences data as Linked Data, mainly promoted by publishers such as Springer8
or Elsevier9 amongst many others. For example, the knowledge management of
scholarly products is an emerging research area in the Semantic Web eld known
as Semantic Publishing [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Semantic Publishing aims at providing access to
semantic enhanced scholarly products with the aim of enabling a variety of
semantically oriented tasks, such as knowledge discovery, knowledge exploration
and data integration. The Semantic Publishing challenge [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] is a breakthrough in
this direction. Its objective is assessing the quality of systems that extract
meaningful metadata from scholarly articles and represent them as RDF. Similarly
to the Semantic Publishing challenge, the Jailbreaking the PDF initiative [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] is
aimed at creating a formal exible infrastructure to extract semantic
information from PDF documents by combining existing solutions and tools in order to
extract data from raw PDFs and convert data to domain-speci c annotations.
      </p>
      <p>
        Despite these continuous e orts, it has been argued that lots of
information about academic conferences is still missing or spread across several sources
in a largely chaotic and non-structured way and a viable solution is a strong
cooperation between researchers and publishers [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
3
      </p>
      <p>
        The cLODg tool - publishing Conference Semantic
Data
Conference metadata collected from di erent unstructured and semi-structured
resources must be expressed with appropriate vocabularies to be exposed as
linked data. cLODg currently implements two data representations: one that
7 ontologydesignpatterns.org
8 http://lod.springer.com/wiki/bin/view/Linked+Open+Data/About
9 http://data.elsevier.com/documentation/index.html
strictly follows the Semantic Web Conference ontology [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and one which
enriches the representations with the SPAR ontologies10 and novel ad-hoc
ontology patterns to also capture social data11. Nevertheless, cLODg architecture
allows the addition of other representations in the future. The Semantic Web
Conference ontology [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] is one of the vocabulary of choice to describe academic
conferences. The SWC ontology extends and combines existing widely accepted
vocabularies (i.e., FOAF [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], SIOC [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], Dublin Core [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]) to provide a reference
model to describe typical actors in an academic conference, such as accepted
papers, authors, their a liations, organizing committee and all other roles
involved. Choosing the SWC ontology as reference vocabulary makes sure that
data is homogeneous with the SWDF corpus. The additional concepts,
properties and axioms that we introduce are further described in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and the resulting
ontology is available on-line12.
      </p>
      <p>The cLODg work ow consists of four main steps:
1. Data acquisition
2. Data conversion and integration
3. Data augmentation and veri cation
4. Linked Data Publication</p>
      <p>Data acquisition is currently supported from (i) proprietary XML data
obtained through the easychair conference management system13, (ii) html based
input (iii) csv les (iv) ics les for calendar events.</p>
      <p>
        Data conversion is implemented via XSLT transformations 14 and integrated
with existing LOD (e.g. we check if some of the conference participants are
already present in the SWDF corpus). Each person in the graph is identi ed
by a URI. For each person we generate a transparent URI of the form http://
data.semanticweb.org/person/&lt;firstname&gt;-&lt;surname&gt;. The convention to
generate a URIs for a person is to use http://data.semanticweb.org/person/
as pre x and concatenate any rstname, middle names and surnames, separeted
by the dash character. This procedure should make sure that if a person is already
present in the SWDF corpus, the same URI is reused. This can in practice cause
problems due to misspelling, noise and ambiguity, for which we implemented a
naive disambiguation procedure, described in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>Produced data is used to pre-populate on-line forms15, which are
submitted to conference participants in order to (ii) verify correctness and (ii) collect
additional information such as their photos and twitter accounts.</p>
      <p>Resulting LOD data about the conference is sent as dump le to be integrated
in the SWDF corpus. Currently this step is performed manually by the SWDF
corpus administrator.
10 http://sempublishing.sourceforge.net/
11 ontologydesignpatterns.org/ont/eswc/ontology.owl
12 The ontology is available at ontologydesignpatterns.org/ont/eswc/ontology.owl
13 http://www.easychair.org/
14 http://www.w3.org/TR/xslt
15 Example of form: http://wit.istc.cnr.it/conference-live/data</p>
    </sec>
    <sec id="sec-3">
      <title>Conclusions and future work</title>
      <p>This paper describes cLODg, a set of tools to collect, re ne and produce Linked
Data to describe scienti c conferences and their publications, participants and
events. cLODg is an answer to the need of open tools for metadata generation
(in the spedi c case in the domain of scienti c conferences). It has the ambition
to foster a synergy between publishers and researches and to provide a possible
way forward to combine the e orts between the two. The main contribution
of this work is an open source tool to support the production of metadata for
conferences and scholarly data.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>D.</given-names>
            <surname>Beckett</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and D.</given-names>
            <surname>Brickey</surname>
          </string-name>
          .
          <article-title>Expressing simple dublin core in rdf/xml</article-title>
          . retrievable on line at http://dublincore.org/documents/dcmes-xml/,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>D.</given-names>
            <surname>Berrueta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Brickley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Decker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Fernandez</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>Gorn, A</article-title>
          . Harth,
          <string-name>
            <given-names>T.</given-names>
            <surname>Heath</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Idehen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kjernsmo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Miles</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Passant</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Polleres</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Polo</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Sintek. SIOC Core</surname>
          </string-name>
          <article-title>Ontology Speci cation</article-title>
          .
          <source>W3c member submission, W3C</source>
          ,
          <year>June 2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>D.</given-names>
            <surname>Brickley</surname>
          </string-name>
          and
          <string-name>
            <given-names>L.</given-names>
            <surname>Miller. FOAF Vocabulary</surname>
          </string-name>
          <article-title>Speci cation</article-title>
          .
          <source>Technical report</source>
          , FOAF project, May
          <year>2007</year>
          .
          <source>Published online on May 24th</source>
          ,
          <year>2007</year>
          at http://xmlns. com/foaf/spec/20070524.html.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>V.</given-names>
            <surname>Bryl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Birukou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Eckert</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Kessler</surname>
          </string-name>
          .
          <article-title>What is in the proceedings? combining publishers and researchers perspectives</article-title>
          .
          <source>In 4th Workshop on Semantic Publishing (SePublica</source>
          <year>2014</year>
          ), Anissaras, Greece, May
          <year>25th</year>
          ,
          <year>2014</year>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>A.</given-names>
            <surname>Garcia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Murray-Rust</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Burns</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Stevens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Tkaczyk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>McLaughlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Belin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Di</surname>
          </string-name>
          <string-name>
            <given-names>Iorio11</given-names>
            , L. Garc a,
            <surname>C.</surname>
          </string-name>
          Gruson-Daniel, et al.
          <article-title>Pdfjailbreak{a communal architecture for making biomedical pdfs semantic</article-title>
          .
          <source>Proceedings of BioLINK SIG</source>
          <year>2013</year>
          , page
          <volume>13</volume>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>A. L.</given-names>
            <surname>Gentile</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Acosta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Costabello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. G.</given-names>
            <surname>Nuzzolese</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Presutti</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D. Reforgiato</given-names>
            <surname>Recupero</surname>
          </string-name>
          .
          <article-title>Conference live: Accessible and sociable conference semantic data</article-title>
          .
          <source>In Proceedings of the 24th International Conference on World Wide Web Companion</source>
          , pages
          <volume>1007</volume>
          {
          <fpage>1012</fpage>
          . International World Wide Web Conferences Steering Committee,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>W.</given-names>
            <surname>Harrison</surname>
          </string-name>
          .
          <article-title>Eating your own dog food</article-title>
          .
          <source>Industrial and Organizational Psychology</source>
          , (June):
          <volume>5</volume>
          {
          <issue>7</issue>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>C.</given-names>
            <surname>Lange</surname>
          </string-name>
          and
          <string-name>
            <given-names>A. Di</given-names>
            <surname>Iorio</surname>
          </string-name>
          .
          <article-title>Semantic publishing challenge - assessing the quality of scienti c output</article-title>
          .
          <source>In Semantic Web Evaluation Challenge</source>
          , volume
          <volume>475</volume>
          of Communications in Computer and Information Science, pages
          <volume>61</volume>
          {
          <fpage>76</fpage>
          . Springer International Publishing,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9. K. Moller, S. Bechofer, and
          <string-name>
            <given-names>T.</given-names>
            <surname>Heath</surname>
          </string-name>
          . Semantic web conference ontology. retrievable on line at http://data.semanticweb.org/ns/swc/swc_2009-05-09.html,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10. K. Moller, T. Heath,
          <string-name>
            <given-names>S.</given-names>
            <surname>Handschuh</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Domingue</surname>
          </string-name>
          .
          <article-title>Recipes for semantic web dog food: The eswc and iswc metadata projects</article-title>
          .
          <source>In Proc. of ISWC'07/ASWC'07</source>
          , pages
          <fpage>802</fpage>
          {
          <fpage>815</fpage>
          , Berlin, Heidelberg,
          <year>2007</year>
          . Springer-Verlag.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>D.</given-names>
            <surname>Shotton</surname>
          </string-name>
          .
          <article-title>Semantic publishing: the coming revolution in scienti c journal publishing</article-title>
          .
          <source>Learned Publishing</source>
          ,
          <volume>22</volume>
          (
          <issue>2</issue>
          ):
          <volume>85</volume>
          {
          <fpage>94</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>