<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Linked Data Utilization along the Content Value Chain - Observations and Implications</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Georg Neubauer</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Applied Sciences St. Poelten Matthias Corvinus Str.</institution>
          <addr-line>15, 3100 St. Poelten</addr-line>
          ,
          <country country="AT">Austria</country>
        </aff>
      </contrib-group>
      <fpage>8</fpage>
      <lpage>11</lpage>
      <abstract>
        <p>The authors present the results of a longitudinal investigation in the utilization of Linked Data technologies along the content value chain. The authors analyzed 71 papers in the period from 2006 to 2014 that used Linked data technologies in editorial workflows. By coding the primary and secondary research topics addressed in the paper the authors draw a conclusion of the maturity of Linked Data technologies as support systems along the content value chain. The survey indicates that Linked Data technologies are constantly maturing as a support infrastructure for editorial processes. The validity of the survey results for application domains not related to editorial tasks is open to discussion.</p>
      </abstract>
      <kwd-group>
        <kwd>Linked Data</kwd>
        <kwd>Content Value Chain</kwd>
        <kwd>Semantic Metadata</kwd>
        <kwd>Semantic Web</kwd>
        <kwd>Data Journalism</kwd>
        <kwd>News Production</kwd>
        <kwd>Editorial Workflows</kwd>
        <kwd>Media Economics</kwd>
        <kwd>IPR</kwd>
        <kwd>Data Licensing</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        The growing recognition of Linked Data among the research
community as “Semantic Web done right” [
        <xref ref-type="bibr" rid="ref15">14</xref>
        ] motivates to take a
closer look if and how Linked Data research has evolved over the
recent years. Such an investigation allows to gain insights into
research trends and interdependencies thereof, and it allows to
draw conclusions whether the research field has reached a
significant degree of maturity in terms of technology diffusion and
application areas.
      </p>
      <p>As illustrated in Figure 1 a survey about the occurrence of the
phrase “Linked Data” in research publications of the ACM digital
library from the period 2006 to 2014 reveals the growing
popularity of this technological concept in the computer sciences
till 2013 with a decline in 2014. Linked Data as a generic
technology for data management is being applied across various
application areas and industries, making it very hard to come to a
general statement concerning its level of maturity and industry
adoption. So is this distribution from figure 1 an indicator for the
growing maturity of a research field? And if yes, how can this
maturity be operationalized empirically?
100
0</p>
      <p>Tassilo Pellegrini</p>
      <p>447
368
366
33
47
47
88</p>
      <p>
        208
To tackle these questions the authors chose to analyze a subset of
research papers from the ACM database that address the
application of Linked Data within editorial workflows. This subset
allowed us to apply a unified classification scheme – known as the
content value chain [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] – to the various application areas of
Linked Data. The content value chain can be described as a
process model that is comprised of several sequential steps
contributing to the content production process. By looking at the
application area of Linked Data in editorial workflows it was
possible to identify primary and secondary areas of utilization,
thus allowing us to draw conclusions towards the diffusion and
appropriability of Linked Data for the production of media
content.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. CLASSIFICATION SCHEME &amp;</title>
    </sec>
    <sec id="sec-3">
      <title>RELATED WORK</title>
      <p>
        The original concept of the value chain as developed by Michael
Porter in 1979 is used as an analytical framework for the analysis
of value creation processes at the firm level or the industry level
[
        <xref ref-type="bibr" rid="ref16">15</xref>
        ]. Over recent years the concept of the value chain has also
gained popularity in the context of open data in general [4; 6; 16]
and Linked Data in special [3; 5]. Especially research that
investigated the organizational and economic impact of Linked
Data refers to the concept of the value chain [
        <xref ref-type="bibr" rid="ref14">13</xref>
        ].
      </p>
      <p>
        In this paper we refer to a generic abstraction of the content value
chain consisting of five steps: 1) content acquisition, 2) content
editing, 3) content bundling, 4) content distribution and 5) content
consumption. As illustrated by [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] Linked Data can contribute to
each step by supporting its associated intrinsic production
function. These are in detail:
Content acquisition is mainly concerned with the collection,
storage and integration of relevant information necessary to
produce a news item. In the course of this process information and
facts are being pooled from internal or external sources for further
processing.
      </p>
      <p>Content editing entails all necessary steps that deal with the
semantic adaptation, interlinking and enrichment of data.
Adaptation can be understood as a process in which acquired data
is provided in a way that it can be used in the editorial process.
Interlinking and enrichment are often performed via processes like
tagging and/or referencing to enrich media documents either by
disambiguating existing concepts or by providing background
knowledge for deeper insights.</p>
      <p>Content bundling is mainly concerned with the contextualization
and personalization of information products. It can be used to
provide customized access to media files i.e. by using metadata for
the device-sensitive delivery of content, or to compile thematically
relevant material into Landing Pages or Dossiers thus improving
the navigability, findability and reuse of information.</p>
      <p>In a Linked Data environment the process of content distribution
mainly deals with the provision of machine-readable and
semantically interoperable (meta)data via Application
Programming Interfaces (APIs) or SPARQL Endpoints. These can
be designed either to serve internal purposes so that data can be
reused within controlled environments (i.e. within or between
units) or for external purposes so that data can be shared between
unknown users (i.e. as open SPARQL Endpoints on the Web).
Content consumption entails any means that enable a human user
to search for and interact with content items in a pleasant und
purposeful way. So according to this view this level mainly deals
with end user applications that make use of Linked Data to
provide access to content i.e. by providing reasonable retrieval
tools and/or visualizations.</p>
      <p>The five steps of the content value chain comprise the
classification scheme.</p>
    </sec>
    <sec id="sec-4">
      <title>3. METHODOLOGY</title>
      <p>We selected a sample of 71 papers (out of 1921) dealing with the
utilization of Linked Data in editorial workflows in the period
from 2006 to 2014 from the ACM Digital Library (DL). The
selected papers had to comply with the following criteria: 1) the
work must analyse the utilization of Linked Data with reference to
some sort of editorial workflow; and 2) the work must not be
purely theoretical but provide at least a proof of concept. The
relevant papers have then been analysed and clustered according
to the five classes acquisition, editing, bundling, distribution,
consumption. As most papers treated more than one of these
topics we weighted each paper according to the primary and
secondary topic discussed, thus also gaining a better
understanding how the research topics relate to each other.</p>
      <p>The weighted greyscale values have been calculated as follows.
Given that black is 100%. 50% divided by the amount of papers
with main classification (black) multiplied with the amount of the
related classifications for the secondary classification. Figure 4
illustrates the results of our survey.</p>
    </sec>
    <sec id="sec-5">
      <title>4. RESULTS</title>
    </sec>
    <sec id="sec-6">
      <title>4.1 General Findings</title>
      <p>The main application areas of Linked Data in editorial workflows
fall into the areas editing (23 papers), bundling (18 papers) and
consumption (21 papers).</p>
      <p>
        Crawling and leveraging processes could be subsumed as
acquisition process [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] using special indexing methods for several
entities found and aggregated through queries. The indexing
methods built a fundament for further scientific processing called
content editing.
      </p>
      <p>Scientific editing using algorithmic methods to classify data into
separated, semantically enriched lists or ontologies were treated in
23 papers as main topic. All of these editing methods were part of
a recognition process used for video-, text- or graphic- analysis in
terms of media-analysis and enrichment of metadata.
18 papers concerned content bundling as main topic. Bundling can
easily be defined as fine-grained representations of resource parts
used for personalization and contextualization of the content.
Just 4 papers described distributions for example in case of
improved accessibility of information. The main difference to the
content bundling process and the content consumption process
explained later on, therefore was, that only APIs can access this
data which in case of content bundling wasn't put to visualized
graphs of the content. This low number of distributions is not
significant for further conclusions.
21 papers applied Linked Data through a framework visualizing
graph-based relations of links. This sort of standard for framework
developers was to visualize links of Linked Data for purposes like
content recommendation.</p>
    </sec>
    <sec id="sec-7">
      <title>4.2 Longitudinal Perspective</title>
      <p>2006: We found just one paper in 2006 with relation to our
research focus. This paper addressed content acquisition as main
topic and editing issues as secondary topic.
2007: In 2007 one paper was classified treating content bundling
as main topic and content acquisition as secondary topic. Two
papers addressing content consumption as primary topic and
acquisition, editing and bundling in treating only content
consumption.
2008: In 2008 we determine one paper addressing content
distribution and one paper addressing content consumption both
referring to content editing.
2009: We have three papers classified as content editing, content
bundling and content consumption. The subrelations in case of
content bundling is editing and in case of content consumption the
subrelations equally refer to content bundling and content
distribution.
2010: In 2010 the authors detected one paper treating content
acquisition, one paper treating content distribution and another
one content consumption. Two papers treated content editing
frameworks. All of the five papers treated content acquisition as
their secondary topic.
2011: In 2011 one paper was about content acquisition, editing,
distribution and content consumption. The relations begin in the
content editing class including a single subrelation to content
acquisition and content consumption. Four papers have all an
equal amount of subrelations to content acquisition and editing.
Additionally one paper described a framework for content
consumption.
2012: In 2012 the authors found one paper addressing content
acquisition as main topic and content editing as secondary topic.
Two papers demonstrated the opposite pattern, discussing editing
as main topic and acquisition as secondary topic. Four papers refer
to content bundling with subrelations to content acquisition and
content editing, while one of them also mentioned content
distribution or content consumption as tertiary topic. Four papers
address content consumption as main topic showing subrelations
to content acquisition in all of their descriptions and one paper
including further treatment of editing.
2013: All papers that describe content editing frameworks in the
year of 2013 also have acquisitional processes as topic. One of
three papers addressing content editing have a subrelation to
content bundling. Two papers are subrelated to content
distribution and one to content consumption. Only one paper
related to content bundling subrelated to content acquisition and
content editing. Four papers give reason to content consumption.
Their relation to subclasses are three addressing content editing,
two addressing content bundling and four addressing content
consumption frameworks as main topic.
2014: In 2014 the classification scheme of the content value chain
seems applicable to a huge amount of papers. We analysed 25
papers and came to the conclusion that scientific content editing
utilizing combinations of vocabularies for the preparation of
linked data is high of note, i.e. automatic extraction RDF-Triples
from web sources for purposes of content enrichment. So 11
papers are classified as content editing in nearly all cases within
acquisitional preprocessing. Content bundling with 5 papers and
content consumption with 6 papers as main classification seem
very similar spreaded in relation to the former years.</p>
    </sec>
    <sec id="sec-8">
      <title>5. DISCUSSION, LIMITATIONS &amp; FUTURE</title>
    </sec>
    <sec id="sec-9">
      <title>WORK</title>
      <p>
        The results show a trend in the utilization of Linked Data
technologies towards content editing, content bundling and
content consumption. Especially the increasing amount of papers
addressing consumption purposes after 2009 is taken as an
indicator for the increasing maturity of Linked Data technologies
in editorial workflows. We also made out a reason of the
increasing usage of content acquisition processes beginning in
2008, assuming that the data infrastructure achieved reclaimable
integrity. Concerning the main result the intertwinedness of
research topics have seamless integration of distinct steps in the
content value chain. Metadata acquisition systems can minimize
the human burden in recording data [
        <xref ref-type="bibr" rid="ref13">12</xref>
        ]. Normally the content
acquisition process is the premier step to process data. We also
claim that there exists a structural relation between content
distribution and acquisition given the fact that these two processes
are technologically intertwined in interlinked data ecosystems.
Content distribution could be treated as a main goal of data
storage and supply [
        <xref ref-type="bibr" rid="ref14">13</xref>
        ]. The authors assume that well established
Linked Data stores are a precondition to content acquisition
allowing further processing like content bundling, content
distribution and content consumption. By taking this appropriate
amount of papers in 2014 we came to the conclusion that content
editing takes root, but the consistency of the result should also be
considered in a normalized way to the former years.
      </p>
      <p>To gain further insights the authors plan to extend the sample size
of their survey in their future work. The current amount of 71
papers is simply too small to draw precise conclusions on the state
of the art and future direction of Linked Data utilization in
editorial workflows. But apart from these limitations the insights
generated by the survey indicate that Linked Data technologies are
constantly maturing as a support infrastructure for editorial
processes. The validity of the survey results for application
domains not related to editorial tasks is open to discussion.</p>
    </sec>
    <sec id="sec-10">
      <title>6. REFERENCES</title>
      <p>Value of
Report,
Semantic
Semantic</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Pellegrini</surname>
          </string-name>
          , Tassilo. “
          <article-title>Integrating Linked Data into the Content Value Chain: A Review of News-Related Standards, Methodologies</article-title>
          and
          <string-name>
            <given-names>Licensing</given-names>
            <surname>Requirements</surname>
          </string-name>
          .”
          <source>In Proceedings of the 8th International Conference on Semantic Systems</source>
          ,
          <volume>94</volume>
          -
          <fpage>102</fpage>
          . ACM,
          <year>2012</year>
          . http://dl.acm.org/citation.cfm?id=
          <fpage>2362513</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Auer</surname>
            , Sören, Theodore Dalamagas, Helen Parkinson, François Bancilhon, Giorgos Flouris, Dimitris Sacharidis,
            <given-names>Peter</given-names>
          </string-name>
          <string-name>
            <surname>Buneman</surname>
          </string-name>
          , et al. “
          <article-title>Diachronic Linked Data: Towards LongTerm Preservation of Structured Interrelated Information</article-title>
          .”
          <source>In Proceedings of the First International Workshop on Open Data</source>
          ,
          <fpage>31</fpage>
          -
          <lpage>39</lpage>
          . WOD '
          <fpage>12</fpage>
          . New York, NY, USA: ACM,
          <year>2012</year>
          . http://doi.acm.
          <source>org/10</source>
          .1145/2422604.2422610.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Auer</surname>
          </string-name>
          , Sören, Jens Lehmann,
          <string-name>
            <surname>Axel-Cyrille Ngonga Ngomo</surname>
          </string-name>
          , and Amrapali Zaveri. “
          <article-title>Introduction to Linked Data and Its Lifecycle on the Web.” In Reasoning Web</article-title>
          .
          <source>Semantic Technologies for Intelligent Data Access</source>
          ,
          <fpage>1</fpage>
          -
          <lpage>90</lpage>
          . Springer,
          <year>2013</year>
          . http://link.springer.com/chapter/10.1007/978-3-
          <fpage>642</fpage>
          - 39784-
          <issue>4</issue>
          _
          <fpage>1</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Davis</surname>
          </string-name>
          , Mills. Technologies.
          <article-title>” “The Business Presentation</article-title>
          and
          <string-name>
            <surname>Technologies for E-Government</surname>
          </string-name>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          http://project10x.com/bio_downloads/business_value_of_sem anti c_technologies_
          <year>2005</year>
          .pdf,
          <source>accessed May 9</source>
          ,
          <fpage>2015</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Latif</surname>
          </string-name>
          , Atif, Anwar Us Saeed, Patrick Hoefler, Alexander Stocker, and Claudia Wagner. “
          <article-title>The Linked Data Value Chain: A Lightweight Model for Business Engineers</article-title>
          .”
          <string-name>
            <surname>In</surname>
            <given-names>ISEMANTICS</given-names>
          </string-name>
          ,
          <fpage>568</fpage>
          -
          <lpage>75</lpage>
          . Citeseer,
          <year>2009</year>
          . http://citeseerx.ist.psu.edu/viewdoc/download?doi
          <source>=10.1.1.181. 950 &amp;rep=rep1&amp;type=pdf.</source>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Pepe</surname>
          </string-name>
          , Alberto, Matthew Mayernik, Christine L.
          <string-name>
            <surname>Borgman</surname>
          </string-name>
          , and Herbert Van de Sompel. “
          <article-title>Technology to Represent Scientific Practice: Data, Life Cycles</article-title>
          , and Value Chains.”
          <source>World Wide Web Internet And Web Information Systems</source>
          ,
          <year>2009</year>
          ,
          <fpage>1</fpage>
          -
          <lpage>22</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Robak</surname>
            , Silva,
            <given-names>Bogdan</given-names>
          </string-name>
          <string-name>
            <surname>Franczyk</surname>
          </string-name>
          , and Marcin Robak. “
          <article-title>Research Problems Associated with Big Data Utilization in Logistics and Supply Chains Design</article-title>
          and Management,” n.d. https://fedcsis.org/proceedings/2014/pliks/472.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Solanki</surname>
          </string-name>
          , Monika, and Christopher Brewster. “
          <article-title>Consuming Linked Data in Supply Chains: Enabling Data Visibility via Linked Pedigrees</article-title>
          .”
          <string-name>
            <surname>In</surname>
            <given-names>COLD</given-names>
          </string-name>
          ,
          <year>2013</year>
          . http://windermere.aston.ac.uk/~monika/papers/SolankiCOLD2 013.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Taskar</surname>
            , Benjamin,
            <given-names>Eran</given-names>
          </string-name>
          <string-name>
            <surname>Segal</surname>
          </string-name>
          , and Daphne Koller. “
          <article-title>Probabilistic Classification and Clustering in Relational Data</article-title>
          .”
          <source>In International Joint Conference on Artificial Intelligence</source>
          ,
          <volume>17</volume>
          :
          <fpage>870</fpage>
          -
          <lpage>78</lpage>
          . LAWRENCE ERLBAUM ASSOCIATES LTD,
          <year>2001</year>
          . http://ai.stanford.edu/users/koller/Papers/Taskar+al:
          <fpage>IJCAI01</fpage>
          .p df.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Van</surname>
            <given-names>Erp</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marieke</surname>
          </string-name>
          , Willem Robert van Hage,
          <string-name>
            <surname>Laura Hollink</surname>
          </string-name>
          , Anthony Jameson, and Raphaël Troncy. “Detection, Representation, and
          <article-title>Exploitation of Events in the Semantic Web</article-title>
          ,”
          <year>2013</year>
          . http://ceur-ws.org/Vol1123/proceedingsderive2013.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Villazón-Terrazas</surname>
          </string-name>
          , Boris, and Oscar Corcho.
          <article-title>“Methodological Guidelines for Publishing Linked Data</article-title>
          .” Una Profesión, Un Futuro:
          <string-name>
            <surname>Actas de Las XII Jornadas Españolas de Documentación</surname>
          </string-name>
          : Málaga 25, no.
          <volume>26</volume>
          (
          <year>2011</year>
          ):
          <fpage>20</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Labrinidis</surname>
            , Alexandros, and
            <given-names>H. V.</given-names>
          </string-name>
          <string-name>
            <surname>Jagadish</surname>
          </string-name>
          . “
          <article-title>Challenges and Opportunities with Big Data</article-title>
          .
          <source>” Proc. VLDB Endow</source>
          .
          <volume>5</volume>
          , no.
          <volume>12</volume>
          (
          <year>August 2012</year>
          ):
          <fpage>2032</fpage>
          -
          <lpage>33</lpage>
          . doi:
          <volume>10</volume>
          .14778/2367502.2367572.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Edward</surname>
          </string-name>
          , Curry et al.
          <article-title>"Big Data</article-title>
          .
          <source>Technical Working Groups White Paper</source>
          ,"
          <year>2014</year>
          . http://bigproject.eu/sites/default/files/BIG_D2_
          <article-title>2_2</article-title>
          .pdf
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Berners-Lee</surname>
          </string-name>
          ,
          <source>Tim</source>
          (
          <year>2008</year>
          ).
          <article-title>Linked open Data</article-title>
          . See also: http://www.w3.org/2008/Talks/0617-lod-tbl/#%281%29, accessed May 9,
          <fpage>2015</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Porter</surname>
          </string-name>
          ,
          <string-name>
            <surname>Michael</surname>
          </string-name>
          (
          <year>1985</year>
          ). Competitive Advantage. New York: Free Press
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Archer</surname>
          </string-name>
          , Phil; Dekkers, Max; Goedertier, Stijn; Loutas,
          <string-name>
            <surname>Nikolaos</surname>
          </string-name>
          (
          <year>2013</year>
          ).
          <article-title>Study on business models for Linked Open Government Data (BM4LOGD - SC6DI06692)</article-title>
          . Services See also: http://ec.europa.eu/isa/documents/study-onbusiness
          <article-title>-modelsopen-government_en</article-title>
          .pdf,
          <source>accessed May 10</source>
          ,
          <year>2015</year>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>