<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>April</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Thomas Steiner</string-name>
          <email>tsteiner@liris.cnrs.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pierre-Antoine Champin</string-name>
          <email>pachampin@liris.cnrs.fr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hannes Mühleisen</string-name>
          <email>hannes@cwi.nl</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Benoît Encelle</string-name>
          <email>bencelle@liris.cnrs.fr</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ruben Verborgh</string-name>
          <email>ruben.verborgh@ugent.be</email>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yannick Prié</string-name>
          <email>yannick.prie@univ-nantes.fr</email>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CNRS, Université de Lyon</institution>
          ,
          <addr-line>LIRIS, UMR5205, Université Lyon 1</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>CNRS, Université de Lyon</institution>
          ,
          <addr-line>LIRIS, UMR5205, Université Lyon 1</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>CNRS, Université de Lyon</institution>
          ,
          <addr-line>LIRIS, UMR5205, Université Lyon 1</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Database Architectures Group, CWI</institution>
          ,
          <addr-line>Science Park 123, 1098 XG Amsterdam, NL</addr-line>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>LINA - UMR 6241 CNRS, Université de Nantes</institution>
          ,
          <addr-line>44322 Nantes Cedex 3</addr-line>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>Multimedia Lab, Ghent University - iMinds</institution>
          ,
          <addr-line>B-9050 Gent</addr-line>
          ,
          <country country="BE">Belgium</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2014</year>
      </pub-date>
      <volume>8</volume>
      <issue>2014</issue>
      <abstract>
        <p>Video has become a rst class citizen on the Web with broad support in all common Web browsers. Where with structured mark-up on webpages we have made the vision of the Web of Data a reality, in this paper, we propose a new vision that we name the Web(VTT) of Data, alongside with concrete steps to realize this vision. It is based on the evolving standards WebVTT for adding timed text tracks to videos and JSON-LD, a JSON-based format to serialize Linked Data. Just like the Web of Data that is based on the relationships among structured data, the Web(VTT) of Data is based on relationships among videos based on WebVTT les, which we use as Web-native spatiotemporal Linked Data containers with JSON-LD payloads. In a rst step, we provide necessary background information on the technologies we use. In a second step, we perform a largescale analysis of the 148 terabyte size Common Crawl corpus in order to get a better understanding of the status quo of Web video deployment and address the challenge of integrating the detected videos in the Common Crawl corpus into the Web(VTT) of Data. In a third step, we open-source an online video annotation creation and consumption tool, targeted at videos not contained in the Common Crawl corpus and for integrating future video creations, allowing for weaving the Web(VTT) of Data tighter, video by video.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Categories and Subject Descriptors</title>
      <p>H.5.1 [Multimedia Information Systems]: Video</p>
    </sec>
    <sec id="sec-2">
      <title>1. INTRODUCTION 1.1</title>
      <p>From &lt;OBJECT&gt; to &lt;video&gt;</p>
      <p>
        In the \ancient" times of HTML 4.01 [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ], the &lt;OBJECT&gt;
tag1 was intended for allowing authors to make use of
multimedia features like including images, applets (programs
that were automatically downloaded and ran on the user's
machine), video clips, and other HTML documents in their
pages. The tag was seen as a future-proof all-purpose
solution to generic object inclusion. In an &lt;OBJECT&gt; tag,
HTML authors can specify everything required by an
object for its presentation by a user agent: source code,
initial values, and run-time data. While most user agents
have \built-in mechanisms for rendering common data types
such as text, GIF images, colors, fonts, and a handful of
graphic elements", to render data types they did not support
natively|namely videos|user agents generally ran external
applications and depended on plugins like Adobe Flash.2.
      </p>
      <p>
        While the above paragraph is provocatively written in
past tense and while the &lt;object&gt; tag is still part of both
the current World Wide Web Consortium (W3C) HTML5
speci cation [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and the Web Hypertext Application
Technology Working Group (WHATWG) \Living Standard",3
more and more Web video is now powered by the native
and well-standardized &lt;video&gt; tag that no longer depends
on plugins. What currently still hinders the full adoption
of &lt;video&gt;, besides some licensing challenges around video
codecs, is its lack of Digital Rights Management (DRM)
support and the erce debate around it, albeit the Director of
the W3C has con rmed4 that work in form of the Encrypted
Media Extensions [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] on \playback of protected content " was
in the scope of the HTML Working Group. However, it can
well be said that HTML5 video has nally become a rst
class Web citizen that all modern browsers fully support.
1HTML 4.01 &lt;OBJECT&gt; tag (uppercased in the spirit of
the epoch): http://www.w3.org/TR/REC-html40/struct/
objects.html#edef-OBJECT
2Adobe Flash: http://get.adobe.com/flashplayer/
3HTML5 &lt;object&gt; tag in the \Living Standard" (now
lowercased): http://www.whatwg.org/specs/web-apps/
current-work/#the-object-element
4New Charter for the HTML Working Group:
http://lists.w3.org/Archives/Public/public-htmladmin/2013Sep/0129.html
1.2
      </p>
    </sec>
    <sec id="sec-3">
      <title>Contributions and Paper Structure</title>
      <p>We are motivated by the vision of a Web(VTT) of Data,
a global network of videos and connected content that is
based on relationships among videos based on WebVTT
les, which we use as Web-native spatiotemporal containers
of Linked Data with JSON-LD payloads. The paper makes
four contributions, including transparent code and data.
i) Large-Scale Common Crawl study of the state
of Web video: we have examined the 148 terabyte
size Common Crawl corpus and determined statistics
on the usage of the &lt;video&gt;, &lt;track&gt;, and &lt;source&gt;
tags and their implications for Linked Data.
ii) WebVTT conversion to RDF-based Linked Data:
we propose a general conversion process for
\triplifying" existing WebVTT, i.e., for turning WebVTT into
a specialized concrete syntax of RDF. This process is
implemented in form of an online conversion tool.
iii) Online video annotation format and editor: we
have created an online video annotation format and an
editor prototype implementing it that serves for the
creation and consumption of semantic spatiotemporal
video annotations turning videos into Linked Data.
iv) Data and code: source code and data are available.</p>
      <p>The remainder of the paper is structured as follows.
Section 2 provides an overview of the enabling technologies that
we require for our approach. Section 3 describes a
largescale study of the state of Web video deployment based on
the Common Crawl corpus. Section 4 deals with the
integration of existing videos into the Web(VTT) of Data through
a tool called LinkedVTT. Section 5 presents an online video
annotation format and an editor that implements this
format. We look at related work in Section 6 and close with
conclusions and an outlook on future work in Section 7.
2.</p>
    </sec>
    <sec id="sec-4">
      <title>TECHNOLOGIES OVERVIEW</title>
      <p>In this section, we lay the foundations of the set of
technologies that enable our vision of the Web(VTT) of Data.
The &lt;track&gt; tag allows authors to specify explicit
external timed text tracks for videos. With the &lt;source&gt; tag,
authors can specify multiple alternative media resources for
a video. Both do not represent anything on their own and
are only meaningful as direct child nodes of a &lt;video&gt; tag.</p>
      <sec id="sec-4-1">
        <title>Web Video Text Tracks format (WebVTT).</title>
        <p>
          The Web Video Text Tracks format (WebVTT, [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ]) is
intended for marking up external text track resources mainly
for the purpose of captioning video content. The
recommended le extension is vtt, the MIME type is text/vtt.
WebVTT les are encoded in UTF-8 and start with the
required string WEBVTT. Each le consists of items called cues
that are separated by an empty line. Each cue has a start
time and an end time in hh:mm:ss.milliseconds
format, separated by a stylized ASCII arrow --&gt;. The cue
payload follows in the line after the cue timings part and
can span multiple lines. Typically, the cue payload contains
plain text, but can also contain textual data serialization
formats like JSON, which later on in the paper we will show
is essential for our proposed approach to semantic video
annotation. Cues optionally can have unique WebVTT
identi ers. WebVTT-compliant Web browsers [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] support ve
di erent kinds of WebVTT tracks: subtitles, captions,
descriptions, chapters, and metadata, detailed in
Table 1 and speci ed in HTML5 [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. In this paper, we are
especially interested in text tracks of kind metadata that
are meant to be used from a scripting context and that are
not displayed by user agents. For scripting purposes, the
video element has a property called textTracks that
returns a TextTrackList of TextTrack members, each of
which correspond to track elements. A TextTrack has
a cues property that returns a TextTrackCueList of
individual TextTrackCue items. Important for us, both
TextTrack and TextTrackCue elements can be
dynamically generated. Listing 1 shows a sample WebVTT le.
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>JSON-LD.</title>
        <p>
          The JavaScript Object Notation5 (JSON) is a (despite the
name) language-independent textual syntax for serializing
objects, arrays, numbers, strings, booleans, and null. Linked
Data [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] describes a method of publishing structured data
so that it can be interlinked and become more useful, which
builds upon standard Web technologies such as HTTP, RDF
and URIs. Based on top of JSON, the JavaScript Object
Notation for Linked Data (JSON-LD, [
          <xref ref-type="bibr" rid="ref27">27</xref>
          ]) is a method for
transporting Linked Data with a smooth upgrade path from
JSON to JSON-LD. JSON-LD properties like title can be
mapped to taxonomic concepts (like dc:title from Dublin
Core6) via so-called data contexts.
5JavaScript Object Notation: http://json.org/
6Dublin Core: http://dublincore.org/documents/dces/
WEBVTT
00:01.000 --&gt; 00:04.000
Never drink liquid nitrogen.
00:05.000 --&gt; 00:09.000
It will perforate your stomach.
        </p>
        <p>Listing 1: Example WebVTT le with two cues</p>
        <sec id="sec-4-2-1">
          <title>WebVTT Kind</title>
          <p>subtitles</p>
          <p>captions
descriptions
chapters
metadata</p>
        </sec>
        <sec id="sec-4-2-2">
          <title>Description and Default Behavior</title>
          <p>Transcription or translation of speech,
suitable for when sound is available but
not understood. Overlaid on the video.
Transcription or translation of the
dialogue, sound e ects, and other relevant
audio information, suitable for when
sound is unavailable or not clearly
audible. Overlaid on the video; labeled as
appropriate for the hard-of-hearing.
Textual descriptions of the video
component of the media resource, intended
for audio synthesis when the visual
component is obscured, unavailable, or
unusable. Synthesized as audio.
Chapter titles, intended to be used for
navigating the media resource.
Displayed as an interactive (potentially
nested) list in the user agent's interface.
Metadata intended for use from script
context. Not displayed by user agent.</p>
        </sec>
      </sec>
      <sec id="sec-4-3">
        <title>Media Fragments URI.</title>
        <p>
          Media Fragments URI [
          <xref ref-type="bibr" rid="ref30">30</xref>
          ] speci es a syntax for
constructing URIs of media fragments and explains how to handle
them over the HTTP protocol. The syntax is based on the
speci cation of name-value pairs that can be used in URI
query strings and URI fragment identi ers to restrict a
media resource to a certain fragment. Media Fragments URI
supports temporal and spatial media fragments. The
temporal dimension is denoted by the parameter name t and
speci ed as an interval with begin time and end time, with
the begin time defaulting to 0 seconds and the end time
defaulting to the media item's duration. The spatial
dimension selects a rectangular area of pixels from media items.
Rectangles can be speci ed as pixel coordinates or
percentages. Rectangle selection is denoted by the parameter name
xywh. The value is either pixel: or percent: followed
by four comma-separated integers. The integers denote x, y,
width, and height respectively, with x = 0 and y = 0 being
the top left corner of the media item. If percent: is used,
x and width are interpreted as a percentage of the width of
the original media item, y and height of the original height.
        </p>
      </sec>
      <sec id="sec-4-4">
        <title>Ontology for Media Resources.</title>
        <p>
          The Ontology for Media Resources [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] serves to bridge
di erent description methods of media resources and to
provide a core set of descriptive properties. It also de nes
mappings to common metadata formats. Combined with
Media Fragments URI, this allows for making ontologically
anchored statements about media items and fragments thereof.
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>LARGE-SCALE COMMON CRAWL</title>
    </sec>
    <sec id="sec-6">
      <title>STUDY OF THE STATE OF WEB VIDEO</title>
      <p>Part of the objectives behind the Web(VTT) of Data is to
create a truly interconnected global network of and between
videos containing Linked Data pointers to related content of
all sorts, where diverse views are not ltered by the network
bubble, but where serendipitously new views can be
discovered by taking untrodden Linked Data paths. In order to get
there, we have conducted a large-scale study based on the
Common Crawl corpus to get a better understanding of the
status quo of Web video and timed text track deployment.
3.1</p>
    </sec>
    <sec id="sec-7">
      <title>Common Crawl</title>
      <p>
        The Common Crawl Foundation7 is a non-pro t
organization founded in 2008 by Gil Elbaz. Its objective is to
democratize access to Web information by producing and
maintaining an open repository of Web crawl data that is
universally accessible and analyzable. All Common Crawl data
is stored on Amazon Simple Storage Service (Amazon S3)8
and accessible to anyone via Amazon Elastic Compute Cloud
(Amazon EC2),9 allowing the data to be downloaded in
bulk, as well as directly be accessed for map-reduce
processing in EC2. The, at time of writing, latest dataset was
collected at the end of 2013, contains approximately 2.3 billion
webpages and is 148 terabyte in size [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Crawl raw data is
stored in the Web ARChive format (WARC, [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]), an
evolution of the previously used Archive File Format (ARC, [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]),
which was developed at the Internet Archive.10 Each crawl
7Common Crawl: http://commoncrawl.org/
8Amazon S3: http://aws.amazon.com/s3/
9Amazon EC2: http://aws.amazon.com/ec2/
10Internet Archive: https://archive.org/
run is hierarchically organized in segments directories that
contain the WARC les with the HTTP requests and
responses for each fetch, and individual Web Archive
Metadata (WAT, [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]) les, which describe the metadata of each
request and response. While the Common Crawl corpus
gets bigger with each crawl run, it obviously does not
represent the \whole Web", which is an illusive concept anyway,
given that a simple calendar Web application can produce
an in nite number of pages. Common Crawl decides on
the to-be-included pages based on an implementation11 of
the PageRank [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] algorithm, albeit the inclusion strategy
is unknown|despite the foundation's focus on transparency.
3.2
      </p>
    </sec>
    <sec id="sec-8">
      <title>On the Quest for WebVTT</title>
      <p>
        We have analyzed the entire 148 terabytes of crawl data
using an Elastic Compute Cloud job whose code was made
available as open-source.12 Rather than parse each
document as HTML, we have tested them for the regular
expression &lt;video[ˆ&gt;]*&gt;(.*?)&lt;/video&gt;, an approach that also
in previous experiments proved very e cient [
        <xref ref-type="bibr" rid="ref22 ref3">3, 22</xref>
        ]. We
tested exactly 2,247,615,323 webpages that had returned
a successful HTTP response to the Common Crawl bot,
and had to skip exactly 46,524,336 non-HTML documents.
On these webpages, we detected exactly 2,963,766 &lt;video&gt;
tags, resulting in a 1.37 gigabyte raw text le that we have
made available publicly.13 This means that on average only
0:132% of all webpages contain HTML5 video. The whole
job took ve hours on 80 c1.xlarge machines and costed $555,
consisting of $468 for Amazon EC2, plus an additional $87
for Amazon Elastic MapReduce (Amazon EMR).14
3.3
      </p>
    </sec>
    <sec id="sec-9">
      <title>Text Track Statistics</title>
      <p>From all 2,963,766 &lt;video&gt; tags, only 1,456 ( 0:049%)
had a &lt;track&gt; child node. Upon closer examination of
the kinds of these 1,456 &lt;track&gt; nodes (see Table 1 for
an explanation of the various kinds), we saw that the
overwhelming majority are unsurprisingly used for subtitles
or captions. Almost no chapter usage was detected and
neither metadata nor description usage at all. The full
details can be seen in Table 2. Looking at the languages
used in the captions and subtitles, these were almost
exclusively English and French, as can be seen in Table 3. The
track labels listed in Table 4 indeed con rm this
observation. In case of multiple tracks for one video, one track
can be marked as the default track. This happens through
a boolean attribute,15 whose value either needs to be the
empty string or the attribute's name, which is \default" in
the concrete case. Table 5 shows that this was used
correctly in almost all cases. When we tried to determine the
MIME type of the actual text tracks, we relied on the le
extension of the values given in the &lt;track src&gt; attributes.
As a signi cant amount of text tracks seems to be
dynam11Common Crawl PageRank code: https://github.com/
commoncrawl/commoncrawl-crawler/tree/master/src/
org/commoncrawl/service/pagerank
12EC2 job: https://github.com/tomayac/postdoc/blob/
master/demos/warczenschwein/
132,963,766 &lt;video&gt; tags: https://drive.google.com/
file/d/0B9LlSNwL2H8YdWVIQmJDaE81UEk
14Amazon EMR: http://aws.amazon.com/
elasticmapreduce/
15HTML boolean attributes: http://www.whatwg.org/
specs/web-apps/current-work/#boolean-attributes
ically generated on-the- y|and thus had no le extension
but a video identi er in the URL instead|we used an
approximation to check if some part of the URL matched the
regular expression /\bvtt\b/gi. Based on this
approximation, a little over half of all text tracks are in WebVTT
format with the extension .vtt or rarely .webvtt. The
predecessor SubRip le format16 can still be encountered in
about a quarter of all text tracks. In between SubRip and
WebVTT, a format originally called WebSRT (Web Subtitle
Resource Tracks) existed that shared the .srt le
extension. The full distribution details are available in Table 6.
Looking at the number of text tracks per video, almost all
videos had only exactly one text track rather than multiple,
as detailed in Table 7, meaning that the broad majority of
all videos are subtitled or captioned in only one language.
16SubRip le format: http://www.matroska.org/
technical/specs/subtitles/srt.html</p>
    </sec>
    <sec id="sec-10">
      <title>3.4 Video Statistics</title>
      <p>
        As in Section 5 we will report on ways to make
semantic statements about videos on the Web, we have
additionally compiled some video statistics. Unlike with images on
the Web, where semantic statements in Resource
Description Framework (RDF) can be made based on the image's
URL [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], with Web video, the situation is another. Due
to di erent Web browsers supporting di erent video codecs,
it is a common practice to provide videos in di erent
encodings. The user's Web browser then dynamically selects
a version it can play. This is realized through the &lt;source&gt;
tag. Table 8 shows the observed numbers of &lt;source&gt;
tag child nodes per &lt;video&gt; tag with &lt;track&gt; tag, with
the result that up to four sources are given for essentially
the \same" video. Table 9 con rms this observation for
the entire collection of all &lt;video&gt; tags with or without
&lt;track&gt; tag. Table 10 shows the distribution of values
for the &lt;source type&gt; attribute of &lt;video&gt; tags with
&lt;track&gt; tag, the clear leaders being the MP4 format
followed by WebM, a trend that again is also re ected in
Table 11 within the entire collection of all &lt;video&gt; tags with
or without &lt;track&gt; tag.
3.5
      </p>
    </sec>
    <sec id="sec-11">
      <title>Implications on Linked Data for Videos</title>
      <p>
        The biggest issue with this practice of putting multiple
sources is that rather than having one unique identi er (URL)
per video, there can be multiple identi ers. Listing 2 shows
a minimal example. Unless one repeats all statements for
each source, there will always remain unclear sources
without structured data. We note that a video in encoding A
and the \same" video in encoding B may not be marked
as &lt;owl:sameAs&gt;, because statements about the encoding
format of one video do not apply to the other, the
identity symmetry condition would thus be violated. In
practice, a solution similar to specifying canonical URLs in Web
search [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] seems feasible. Another approach is to require
a unique identi er in the &lt;video id&gt; attribute, which
allows for addressing the video with fragment identi ers. More
advanced approaches to the problem stemming from the
bibliographic universe like FRBR [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ] are possible, but for the
concrete use case seem quite complex.
&lt;div about="kitten.jpg"&gt;
&lt;img src="kitten.jpg" alt="Cute kitten" /&gt;
&lt;a rel="license" href="http://creativecommons.
      </p>
      <p>org/licenses/by-sa/3.0/"&gt;</p>
      <p>Creative Commons Attribution Share-Alike 3.0
&lt;/a&gt;
&lt;/div&gt;
&lt;div about="kitten.mp4"&gt;
&lt;video&gt;
&lt;source src="kitten.mp4"/&gt;
&lt;source src="kitten.webm"/&gt;
&lt;/video&gt;
&lt;a rel="license" href="http://creativecommons.</p>
      <p>org/licenses/by-sa/3.0/"&gt;</p>
      <p>Creative Commons Attribution Share-Alike 3.0
&lt;/a&gt;
&lt;/div&gt;
Listing 2: Specifying a license for an image and
attempt to do the same for a video with two sources
(the license of kitten.webm stays unclear)</p>
    </sec>
    <sec id="sec-12">
      <title>WEBVTT CONVERSION TO</title>
    </sec>
    <sec id="sec-13">
      <title>RDF-BASED LINKED DATA</title>
      <p>
        The WebVTT speci cation [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] de nes a syntax for
conveying timed video text tracks, and a semantics for this
syntax in terms of how Web browsers should process such
tracks. It achieves this by specifying an underlying data
model for those tracks. The aim of this section is to show
how this data model can easily be mapped to RDF-based
Linked Data, and thus allowing for many other usage
scenarios for this data. For this purpose, we propose an
RDFSchema ontology17 conveying the WebVTT data model. In
the rest of the paper, terms from this ontology will be
preceded by the vtt: pre x. An online implementation of this
interpretation process that we have titled LinkedVTT is
likewise available online.18 It takes the URL of any WebVTT
le, the contents of a raw WebVTT le, or a YouTube URL
of any video with closed captions as an input, and applies
the conversion from WebVTT to Linked Data on-the- y.
4.1
      </p>
    </sec>
    <sec id="sec-14">
      <title>Basic Interpretation</title>
      <p>
        A WebVTT le de nes a set of cues, which are described
by a pair of timestamps and a payload. In other words, each
cue is an annotation of the video, associating a temporal
video fragment to the payload, delimited by the two
timestamps. As there is a standard way of identifying temporal
and spatial video fragments with a URI [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ] it is
straightforward to represent this annotation as an RDF triple. We
therefore propose a property vtt:annotatedBy to serve
as predicate for those triples. To keep the context of each
annotation, we use the notion of RDF dataset [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Each
vtt:annotatedBy triple is enclosed in a named graph,
whose name is either a URI, based on the cue identi er if it
has one, or a blank node if the cue has no identi er. The
default graph of the dataset describes its overall structure,
linking the dataset URI to all the URIs and blank nodes
identifying its cues with the vtt:hasCue property. In the default
graph, each cue is also linked to the Media Fragments URI
it describes, with the vtt:describesFragment property.
As the notion of dataset is a recent addition to the RDF
core concepts (previously, it was speci c to the SPARQL
query language), we envision that some consumers will not
be able to deal with it. Hence, we propose an alternate
interpretation of WebVTT as RDF. In this at interpretation,
the contents of all named graphs is merged into the default
graph, at the expense of contextual information.
4.2
      </p>
    </sec>
    <sec id="sec-15">
      <title>Advanced Interpretation</title>
      <p>
        WebVTT is not limited to textual timed text tracks. As
Table 1 details, the HTML5 &lt;track&gt; tag supports di
erent kinds of tracks, one of them being metadata, a track
designed for machine rather than human consumption.
Although it was shown in Subsection 3.3 that there is no
measurable evidence of use for this kind of track yet|which
is understandable given that the technology is still under
development|we propose that JSON data is a good
candidate for cues of such tracks. JSON has a textual syntax that
is easy to author and easy to process in a Web browser and
elsewhere. Furthermore, JSON-LD [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ] provides a standard
way to interpret JSON data as Linked Data, which ts nicely
with our approach. More precisely, whenever the payload of
a cue successfully parses as a JSON object, we consider that
this object is meant to represent the annotated media
fragment itself, and interpret it as JSON-LD. In consequence,
all properties of the JSON object are applied directly to the
fragment, and embedded structures can be used to describe
other resources related to that fragment, e.g., depicted
persons, locations, topics, related videos or video fragments, or
17RDF-Schema ontology:
linkedvtt/onto#
18LinkedVTT: http://champin.net/2014/linkedvtt/
http://champin.net/2014/
spatiotemporal video tags. In this case, all the triples
generated from parsing the payload as JSON-LD replace the
vtt:annotatedBy triple in the cue's named graph.
Listing 3 gives an example of such JSON-LD payload. We note
that it includes the JSON-LD speci c @context key, to allow
its interpretation as Linked Data. This context can be
speci ed in each cue, but below we also provide an alternative
way to declare it once for the entire WebVTT le.
4.3
      </p>
    </sec>
    <sec id="sec-16">
      <title>Linked Data Related Metadata</title>
      <p>In addition to the cues, WebVTT les can contain
metadata headers described as key-value pairs. While the WebVTT
speci cation de nes a number of metadata headers, it leaves
it open for extensions. We propose three extended
metadata headers listed below. Most WebVTT currently does
not contain these metadata headers, but we argue that they
allow for an easy transition from plain WebVTT to Linked
Data WebVTT, just like JSON-LD makes it easy to turn
plain JSON into Linked Data by adding a @context
property. Further more, other metadata headers will be
evaluated against the JSON-LD context, and can produce
additional triples with the WebVTT le as its subject.
@base Sets the base URI used for resolving relative URIs.
This applies to any relative URIs that would be found in
the JSON-LD descriptions, but also to generate URIs for
cues based on their identi ers. It defaults to the URI of
the WebVTT le.
@context This key can be used multiple times; each value
is the URI of a JSON-LD context that should be used to
interpret the JSON payloads in the WebVTT le.
@video Sets the URI for the video for generating media
fragment URIs. If not present, the video URI must be
provided externally, e.g., the &lt;video src&gt; attribute of
the video containing the WebVTT track. This metadata
header is a direct response to an issue that we have outlined
in Subsection 3.5.
4.4</p>
    </sec>
    <sec id="sec-17">
      <title>Integrating Existing Videos Into the</title>
    </sec>
    <sec id="sec-18">
      <title>Web(VTT) of Data</title>
      <p>
        Given the currently rather manageable amount of videos
with captions or subtitles as outlined in Subsection 3.3,
approaches for the automated semantic lifting based on timed
text track data are feasible. These approaches extract the
transcribed text snippets from cues and either convert them
into one consistent block of text or treat each text snippet in
isolation before applying named entity extraction on them.
Representative examples based on this idea are [
        <xref ref-type="bibr" rid="ref18 ref19 ref20">18, 19, 20</xref>
        ]
by Li et al. or also [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ] by us. In combination with Media
Fragments URI, spatiotemporal annotations can be created
with good precision and reasonable time e ort both
on-they or in bulk for static storage in a triple store.
5.
      </p>
    </sec>
    <sec id="sec-19">
      <title>ONLINE VIDEO ANNOTATION</title>
    </sec>
    <sec id="sec-20">
      <title>FORMAT AND EDITOR</title>
      <p>Complementary to the conversion process presented in
Section 4, in this section we focus on facilitating the online
creation and consumption of metadata tracks for future
video creations and videos not contained in the Common
Crawl corpus. We begin with the annotation model.
5.1</p>
    </sec>
    <sec id="sec-21">
      <title>Annotation Model</title>
      <p>Our annotation model is the same as the one produced
by the interpretation process presented above. Annotations
take the form of RDF statements (subject-predicate-object),
where the subject is any temporal or spatiotemporal
fragment of the video, identi ed by the corresponding Media
Fragments URI. They are encoded as TextTrackCues with
JSON-LD payloads such as the one shown in Listing 3.
A dedicated data context de nes their semantics.
5.2</p>
    </sec>
    <sec id="sec-22">
      <title>WebVTT Editor</title>
      <p>We have implemented this annotation model in form of
an online demonstrator prototype. The demonstrator
interprets the existing metadata track for a video and reacts on
annotations when the currentTime of the media resource
matches the startTime or endTime of a cue. We call
existing annotations Read annotations. Users can add Write
annotations by creating new TextTrackCues at the desired
start and end times and by providing their JSON-LD
payloads. The editor facilitates this task through a graphical
user interface, abstracting the underlying details. Figure 1
shows a screenshot of the WebVTT editor. Newly generated
annotations get directly interpreted and can be persistently
stored locally or in the future remotely for collaborative
editing. We have developed a WebVTT to JSON-LD converter,
capable of transforming WebVTT metadata tracks following
our annotation model into JSON-LD for the Web of Data.
This allows for straight-forward local annotation creation
with Semantic Web compliance upon global publication.
5.2.1</p>
      <sec id="sec-22-1">
        <title>Semantic Annotation Types</title>
        <p>
          Our JSON-LD context eases common annotation tasks
by de ning the semantics of a few useful JSON properties
described below. According to this context, Listing 3 is
interpreted as in Listing 4 (RDF in JSON-LD syntax) and
Listing 5 (RDF in N-Triples syntax). More advanced
annotation tasks can be supported by extending the data context.
Plain Text Tags Annotations of type tags allow for
adding plain text tags to a media fragment. They are
interpreted as Common Tag [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] format ctag:label.
Semantic Tags Annotations of type semanticTags
allow for adding semantic tags to a media fragment. Unlike
plain text tags, semantic tags are references to well-de ned
concepts complete with their own URIs. They are
interpreted as Common Tag [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] format ctag:means.
Spatiotemporal semantic tags allow for interesting Linked Data
experiences if the tags point to well-connected concepts.
{
}
"@context": "http://champin.net/2014/linkedvtt/
context.json",
"@id": "http://ex.org/metadata.vtt",
"@type": "VideoMetadataDataset",
"video": "http://ex.org/video",
"cues": [{
"@id": "#id=cue1",
"fragment": {
"@context": "http://champin.net/2014/
        </p>
        <p>linkedvtt/demonstrator-context.json",
"@id": "http://ex.org/video#t</p>
        <p>
          =0:0.0,0:12.0",
"tags": ["wind scene", "opening credits"],
"contributors": ["http://ex.org/sintel"]
Contributors The contributors annotation type
allows for denoting the contributors in a media fragment,
like its actors. They are interpreted as Ontology for Media
Resources [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] format ma:hasContributor.
        </p>
        <p>
          Summary The summary annotation type allows for
summarizing a media fragment (note, not the whole video like
kind description tracks) with plain text. They are
interpreted as ma:description [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ].
        </p>
      </sec>
      <sec id="sec-22-2">
        <title>5.2.2 Presentation-Oriented Annotation Types</title>
        <p>Presentation-oriented annotations|similar to temporal
style sheets|do not generate RDF data, but only impact
the way videos get presented.</p>
        <p>Visual E ect Annotations of type visualEffect allow
for applying visual e ects in the syntax of Cascading Style
Sheets19 (CSS) to a media fragment, e.g., lters, zoom,
transparency, and 2D/3D transformations and animations.
Audial E ect The audialEffect annotation type allows
for applying audial e ects to a media fragment. Currently,
we support modifying the volume from 0 to 1.</p>
        <p>Playback Rate The playbackRate annotation type
allows for specifying the e ective playback rate of a media
fragment. The playback rate is expressed as a oating point
multiple or fraction of the intrinsic video speed.</p>
        <p>HTML Overlay Via the htmlOverlay annotation type,
overlays in freeform HTML code can be added to a media
fragment. Examples are graphical, textual, or combined
overlays that can contain links to (temporal fragments of)
other videos or within the current video.
19Cascading Style Sheets: http://www.w3.org/Style/CSS/
&lt;http://ex.org/metadata.vtt&gt; &lt;http://www.w3.org
/1999/02/22-rdf-syntax-ns#type&gt; &lt;http://ex.
org/VideoMetadataDataset&gt; .
&lt;http://ex.org/metadata.vtt&gt; &lt;http://champin.net
/2014/linkedvtt/onto#hasCue&gt; &lt;http://ex.org/
metadata.vtt#id=cue1&gt; .
&lt;http://ex.org/metadata.vtt#id=cue1&gt; &lt;http://
champin.net/2014/linkedvtt/onto#
describesFragment&gt; &lt;http://ex.org/video#t
=0:0.0,0:12.0&gt; .
&lt;http://ex.org/video#t=0:0.0,0:12.0&gt; &lt;http://
commontag.org/ns#label&gt; "wind scene" .
&lt;http://ex.org/video#t=0:0.0,0:12.0&gt; &lt;http://
commontag.org/ns#label&gt; "opening credits" .
&lt;http://ex.org/video#t=0:0.0,0:12.0&gt; &lt;http://www.
w3.org/ns/ma-ont#hasContributor&gt; &lt;http://ex.
org/sintel&gt; .</p>
        <p>Listing 5: RDF triples based on the JSON-LD code
from Listing 4</p>
      </sec>
    </sec>
    <sec id="sec-23">
      <title>5.3 Interpretation Layer</title>
      <p>In our WebVTT editor, we propose an interpretation layer
capable of dealing with the herein de ned annotation types.
We thus make an open world assumption by supporting
a set of pre-de ned values for predicate and object listed
below, and ignoring unknown ones. This permits others
to extend|or even completely replace|our interpretation
layer. If a TextTrackCue has a WebVTT identi er, we
use it to address its annotations via the metadata track's
URI and corresponding cue fragment identi er, allowing for
meta annotations of annotations, e.g., to attach provenance
or license information to them.</p>
    </sec>
    <sec id="sec-24">
      <title>5.4 Evaluation</title>
      <p>
        We evaluate or annotation model and related technology
stack based on a state-of-the-art hypervideo model by
Sadallah et al. [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] that builds on a careful study of prior art.
      </p>
      <sec id="sec-24-1">
        <title>The CHM Hypervideo Model.</title>
        <p>Sadallah et al. de ne hypervideo as \interactive
video-centric hypermedia document built upon audiovisual content".
The authors identify three common hypervideo
characteristics, namely (i) interactivity, which, e.g., can enable richer
navigational possibilities, (ii) non-linearity, which allows for
features like video montages, and nally (iii) enrichments
that include all sorts of supplementary material besides and
on top of hypervideos. The authors have examined
hypervideo systems of recent years and found recurring patterns,
summarized and compared to our approach in the following.
Video player and controls Hypervideo systems by de
nition provide one or multiple video players, however, the
corresponding video controls are not necessarily exposed.
X Our approach uses the (optionally customizable) default
HTML5 player that includes hidable controls (Figure 1).
Timeline A timeline is the spatial representation of
temporally situated metadata in a video. The most common
timeline pattern shows the time along the x-axis and
corresponding metadata along the y-axis.</p>
        <p>X Our approach supports temporal metadata.
Customizable timeline visualizations exist20 and can be added.
Textual or graphical overlay Additional textual or
graphical information can be displayed in form of overlays on
the video. Overlays can also serve as external or
videointernal hyperlinks, referred to as hotspots.</p>
        <p>X We realize overlays and links with htmlOverlay types.
Figure 1 shows both a graphical (yellow box) and two
textual overlays (red and green texts).</p>
        <p>Textual or graphical table of contents If a video is
logically separated into di erent parts, a table of contents lists
these in textual or graphical form, makes them navigable,
or visually summarizes them, referred to as video map.
X Textual tables of contents are directly supported via
WebVTT text tracks of type chapters. Graphical tables
of contents can be created based thereon.</p>
        <p>Transcript The textual document of the transcribed
audiovisual content of a video allows for following along the
video by reading and also serves for in-video navigation.
X Subtitles and captions are natively supported by
WebVTT tracks of the types subtitles and captions.
Figure 1 shows active subtitles (white text).</p>
      </sec>
    </sec>
    <sec id="sec-25">
      <title>RELATED WORK</title>
      <p>
        With our annotation approach, we leverage WebVTT
metadata tracks as a means for tying semantic JSON-LD
annotations to temporal or spatiotemporal video fragments. As
each &lt;track&gt; tag by pure de nition is bound to exactly one
&lt;video&gt; tag, and as modern search engines parse and
interpret JSON-LD annotations, a unique relation of annotations
to video content is made. In consequence, related work can
be regarded under the angles of online annotation creation
and large-scale Linked Data e orts for video. Many have
combined Linked Data and video, typical examples are [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]
by Lambert et al. and [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] by Hausenblas et al. We have
already described the text track enriching approaches [
        <xref ref-type="bibr" rid="ref18 ref19 ref20 ref28">18,
19, 20, 28</xref>
        ] in Subsection 4.4, [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] being closest to our idea
of a Web(VTT) of Data, albeit their approach is centered
around their application Synote. The online video hosting
platform YouTube lets video publishers add video
annotations in a closed proprietary format. From 2009 to 2010,
YouTube had a feature called Collaborative Annotations [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]
that allowed video consumers to collaboratively create video
annotations. Unlike the format of YouTube, our format
is open and standards-based. In [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ], Van Deursen et al.
present a system that combines Media Fragments URI and
the Ontology for Media Resources in an HTML5 Web
application to convert rich media fragment annotations into
a WebVTT le that can be used by HTML5-enabled
players to show the annotations in a synchronized way. Building
on their work, we additionally allow for writing annotations
by letting annotators create WebVTT cues with an editor.
The Component-based Hypervideo Model Popcorn.js21 is an
HTML5 JavaScript media framework for the creation of
media mixes by adding interactivity and context to online video
by letting users link social media, feeds, visualizations, and
20D3 timeline implementation:
jiahuang/d3-timeline
21Popcorn.js: http://popcornjs.org/
https://github.com/
other content directly to moving images. PopcornMaker22
is an interactive Web authoring environment that allows for
videos to be annotated on a video timeline. While Popcorn
media annotations are essentially JavaScript programs, our
approach is based on directly indexable WebVTT les.
7.
      </p>
    </sec>
    <sec id="sec-26">
      <title>CONCLUSIONS AND FUTURE WORK</title>
      <p>In this paper, we have introduced our vision of the
Web(VTT) of Data, a global network of videos and
connected content that is based on relationships among videos
based on WebVTT les, which we use as Web-native
spatiotemporal containers of Linked Data with JSON-LD
payloads. With the recent graduation of the JSON-LD syntax
as an o cial W3C Recommendation and a major search
engine company23 supporting embedded JSON-LD documents
in HTML documents,24 JSON-LD de nitely is here to stay.
Likewise for WebVTT, which in the more recent past has
been natively implemented by all major Web browser
vendors, the future is bright. We combine both technologies in
a fruitful way that is focused both at common Web search
engines as well as at the entire Linked Data stack of
technologies. Using WebVTT as a container for JSON-LD is
both innovative and natural. Making commonly understood
semantic statements about video fragments on the Web has
become feasible thanks to Media Fragments URI, a
standard that allows for applying Linked Data approaches to
moving images on a temporal and spatiotemporal axis. We
have organized this paper in three major steps. (i) in
order to get a better understanding of the status quo of Web
video deployment, we have performed a large-scale
analysis of the 148 terabyte size Common Crawl corpus, (ii) we
have addressed the challenge of integrating existing videos
in the Common Crawl corpus into the Web(VTT) of Data
by proposing a WebVTT conversion to RDF-based Linked
Data, and (iii) we have open-sourced an online video
annotation creation and consumption tool, targeted at videos not
contained in the Common Crawl corpus and for integrating
future video creations. In this paper, we have combined Big
Data and Small Data. On the Big Data side, we have learned
from the Common Crawl corpus which kind of timed text
tracks are out there, which allowed us to propose a realistic
approach to integrating it into the Web(VTT) of Data. On
the Small Data side, we have implemented an online editor
for the creation of semantic video annotations that can be
applied video by video, so that the Web(VTT) of Data gets
woven tighter and tighter with each new addition.</p>
      <p>Future work has several dimensions. Beginning from video
annotation, a rst concrete research task is to work on our
editor prototype. While a lot of e orts can be put in the
editor itself, far more added value is created by
proposing an extension to the most well-known online video
annotation stack, the Popcorn.js and PopcornMaker projects.
A minimal Popcorn.js example annotation can be seen in
Listing 6. Rather than storing the annotations as steps of
a JavaScript program that \arti cially" need to be aligned
to the corresponding parts of the video, an extension to
22PopcornMaker: https://popcorn.webmaker.org/
23JSON-LD in Gmail: https://developers.google.com/
gmail/actions/reference/formats/json-ld
24Embedding JSON-LD in HTML Documents:
http://www.w3.org/TR/json-ld/#embedding-jsonld-in-html-documents
Popcorn.js could use our approach of leveraging naturally
temporally aligned WebVTT cues with JSON-LD payloads
for the annotations. We have been able to play video in
Web browsers plugin-free for a couple of years now, the
next step is adding resources to videos to make them more
accessible and provide more options to the viewer.
Straightforward things to do are to pro t from recent advances in
machine-translation and speech recognition to evaluate the
usefulness of automatically transcribed and translated
captions combined with language-independent metadata
annotations based on named entity extraction for providing
Linked Data paths between videos no matter their original
language. We have learned that with regard to HTML5
video with timed text track support, it is still early days, so
at least in the short-term it will be inevitable to deal with
legacy plugin-dependent content and ways to integrate it in
the Web(VTT) of Data through adaptors or converters etc.</p>
      <p>Concluding, the vision of the Web(VTT) of Data is a
realistic one and all building blocks are in place. We are
optimistic that by leveraging the popularity of existing tools like
Popcorn.js, we can push the state of Web video forward
toward an interconnected, semantic, and wall-free experience.</p>
    </sec>
    <sec id="sec-27">
      <title>Acknowledgments</title>
      <p>The research presented in this paper was partially supported
by the French National Agency for Research project
Spectacle En Ligne(s), project reference ANR-12-CORP-0015.
&lt;video id="video" src="http://ex.org/video.mp4"&gt;
&lt;/video&gt;
&lt;div id="footnote-container"&gt;&lt;/div&gt;
&lt;div id="wikipedia-container"&gt;&lt;/div&gt;
&lt;script&gt;
// get a reference to the video
var pop = Popcorn("#video");
// add a footnote from second 2 to second 6
pop.footnote({
start: 2,
end: 6,
text: "In this scene: George Clooney",
target: "footnote-container"
});
// add a reference to a Wikipedia article
// from second 2 to 6
pop.wikipedia({
start: 2,
end: 6,
src: "http://en.wikipedia.org/wiki/George
_Clooney",
title: "George Clooney",
target: "wikipedia-container"
});
&lt;/script&gt;</p>
      <p>Listing 6: Popcorn.js example</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.</given-names>
            <surname>Bar</surname>
          </string-name>
          et al.
          <article-title>YouTube's Collaborative Annotations</article-title>
          . In Webcentives '
          <volume>09</volume>
          , 1st International Workshop on Motivation and Incentives, pages
          <volume>18</volume>
          {
          <fpage>19</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>R.</given-names>
            <surname>Berjon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Faulkner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Leithead</surname>
          </string-name>
          , et al.
          <article-title>HTML5, A Vocabulary and Associated APIs for HTML and XHTML</article-title>
          . Candidate Recommendation,
          <year>W3C</year>
          ,
          <year>2013</year>
          . http://www.w3.org/TR/html5/.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>C.</given-names>
            <surname>Bizer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Eckert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Meusel</surname>
          </string-name>
          , H. Muhleisen, M. Schuhmacher, and
          <string-name>
            <given-names>J.</given-names>
            <surname>Vo</surname>
          </string-name>
          <article-title>lker. Deployment of RDFa, Microdata, and Microformats on the Web { A Quantitative Analysis</article-title>
          . In H. Alani,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kagal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Fokoue</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Groth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Biemann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Parreira</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Aroyo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Noy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Welty</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          K. Janowicz, editors,
          <source>The Semantic Web { ISWC</source>
          <year>2013</year>
          , volume
          <volume>8219</volume>
          of Lecture Notes in Computer Science, pages
          <volume>17</volume>
          {
          <fpage>32</fpage>
          . Springer Berlin Heidelberg,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>C.</given-names>
            <surname>Bizer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Heath</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Berners-Lee</surname>
          </string-name>
          .
          <article-title>Linked Data|The Story So Far</article-title>
          .
          <source>Int. J. Semantic Web Inf. Syst.</source>
          ,
          <volume>5</volume>
          (
          <issue>3</issue>
          ):1{
          <fpage>22</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>C.</given-names>
            <surname>Bizer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Heath</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Berners-Lee</surname>
          </string-name>
          , and M. Hausenblas, editors.
          <source>WWW2012 Workshop on Linked Data on the Web</source>
          , Lyon, France, 16 April,
          <year>2012</year>
          , volume
          <volume>937</volume>
          <source>of CEUR Workshop Proceedings. CEUR-WS.org</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Burner</surname>
          </string-name>
          and
          <string-name>
            <given-names>B.</given-names>
            <surname>Kahle</surname>
          </string-name>
          . Arc File Format.
          <source>Technical report, Jan</source>
          .
          <year>1996</year>
          . http://archive.org/web/ researcher/ArcFileFormat.php.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>R.</given-names>
            <surname>Cyganiak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Wood</surname>
          </string-name>
          , and M.
          <source>Lanthaler. RDF 1.1 Concepts</source>
          and
          <string-name>
            <given-names>Abstract</given-names>
            <surname>Syntax</surname>
          </string-name>
          . Proposed Recommendation, W3C, Jan.
          <year>2014</year>
          . http://www.w3.org/TR/rdf11-concepts/.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>D.</given-names>
            <surname>Dorwin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bateman</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Watson</surname>
          </string-name>
          . Encrypted Media Extensions. Working Draft, W3C, Oct.
          <year>2013</year>
          . http://www.w3.org/TR/encrypted-media/.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Dutton</surname>
          </string-name>
          .
          <article-title>Getting Started With the Track Element</article-title>
          , Feb.
          <year>2012</year>
          . http://www.html5rocks.com/en/ tutorials/track/basics/.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>V.</given-names>
            <surname>Goel</surname>
          </string-name>
          .
          <article-title>Web Archive Metadata File Speci cation</article-title>
          .
          <source>Technical report, Apr</source>
          .
          <year>2011</year>
          . https: //webarchive.jira.com/wiki/display/Iresearch/ Web+Archive+Metadata+File+Specification.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>L.</given-names>
            <surname>Green</surname>
          </string-name>
          .
          <article-title>Winter 2013 Crawl Data Now Available</article-title>
          , Jan.
          <year>2014</year>
          . http://commoncrawl.org/winter-2013
          <string-name>
            <surname>-</surname>
          </string-name>
          crawl
          <article-title>-data-now-available/.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hausenblas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Troncy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Raimond</surname>
          </string-name>
          , and T. Burger. Interlinking Multimedia:
          <article-title>How to Apply Linked Data Principles to Multimedia Fragments</article-title>
          .
          <source>In Linked Data on the Web Workshop (LDOW 09)</source>
          ,
          <source>in conjunction with the 18th International World Wide Web Conference (WWW 09)</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>A.</given-names>
            <surname>Iskold</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Passant</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Milicic</surname>
          </string-name>
          , et al.
          <article-title>Common Tag Speci cation</article-title>
          ,
          <year>June 2009</year>
          . http://commontag.org/Specification.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <source>[14] ISO 28500</source>
          .
          <article-title>Information and documentation { The WARC File Format</article-title>
          .
          <source>International Standard</source>
          ,
          <year>2008</year>
          . http://bibnum.bnf.fr/WARC/WARC_ISO_
          <volume>28500</volume>
          _ version1_latestdraft.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>J.</given-names>
            <surname>Kupke</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Ohye</surname>
          </string-name>
          .
          <article-title>Specify your canonical</article-title>
          ,
          <source>Feb</source>
          .
          <year>2009</year>
          . http://googlewebmastercentral.blogspot. de/2009/02/specify-your-canonical.html.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>D.</given-names>
            <surname>Lambert</surname>
          </string-name>
          and
          <string-name>
            <given-names>H. Q.</given-names>
            <surname>Yu</surname>
          </string-name>
          .
          <article-title>Linked Data based Video Annotation and Browsing for Distance Learning</article-title>
          . In SemHE '
          <volume>10</volume>
          : The Second International Workshop on Semantic Web Applications in Higher Education,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>W.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Bailer</surname>
          </string-name>
          , T. Burger, et al.
          <source>Ontology for Media Resources</source>
          <volume>1</volume>
          .0. Recommendation, W3C, Feb.
          <year>2012</year>
          . http://www.w3.org/TR/mediaont-10/.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Rizzo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. L. Redondo</given-names>
            <surname>Garc</surname>
          </string-name>
          <string-name>
            <surname>a</surname>
          </string-name>
          , R. Troncy,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wald</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Wills. Enriching Media</surname>
          </string-name>
          <article-title>Fragments with Named Entities for Video Classi cation</article-title>
          .
          <source>In Proceedings of the 22nd International Conference on World Wide Web Companion, WWW '13 Companion</source>
          , pages
          <volume>469</volume>
          {
          <fpage>476</fpage>
          , Republic and Canton of Geneva, Switzerland,
          <year>2013</year>
          . International World Wide Web Conferences Steering Committee.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Rizzo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Troncy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wald</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Wills. Creating Enriched YouTube Media</surname>
          </string-name>
          <article-title>Fragments with NERD Using Timed-Text</article-title>
          .
          <source>In 11th International Semantic Web Conference (ISWC2012)</source>
          ,
          <year>November 2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wald</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Omitola</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shadbolt</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Wills</surname>
          </string-name>
          . Synote:
          <article-title>Weaving Media Fragments and Linked Data</article-title>
          . In Bizer et al. [
          <volume>5</volume>
          ].
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>P.</given-names>
            <surname>Linsley</surname>
          </string-name>
          .
          <article-title>Specifying an image's license using RDFa, Aug</article-title>
          .
          <year>2009</year>
          . http://googlewebmastercentral. blogspot.com/
          <year>2009</year>
          /08/specifying-imageslicense
          <article-title>-using-rdfa.html.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>H.</given-names>
            <surname>Mu</surname>
          </string-name>
          <article-title>hleisen and C. Bizer. Web Data Commons { Extracting Structured Data from Two Large Web Corpora</article-title>
          . In Bizer et al. [
          <volume>5</volume>
          ].
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>L.</given-names>
            <surname>Page</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Brin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Motwani</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Winograd</surname>
          </string-name>
          .
          <article-title>The PageRank Citation Ranking: Bringing Order to the Web</article-title>
          .
          <source>Technical report</source>
          , Stanford InfoLab, Nov.
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>S.</surname>
          </string-name>
          <article-title>Pfei er and I. Hickson. WebVTT: The Web Video Text Tracks Format</article-title>
          . Draft Community Group Speci cation,
          <issue>W3C</issue>
          , Nov.
          <year>2013</year>
          . http://dev.w3.org/html5/webvtt/.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>D.</given-names>
            <surname>Raggett</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Le Hors</surname>
          </string-name>
          ,
          <string-name>
            <surname>and I. Jacobs.</surname>
          </string-name>
          <article-title>HTML 4.01 Speci cation</article-title>
          . Recommendation, W3C, Dec.
          <year>1999</year>
          . http://www.w3.org/TR/html401.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>M.</given-names>
            <surname>Sadallah</surname>
          </string-name>
          et al.
          <article-title>CHM: An Annotation-</article-title>
          and
          <article-title>Component-based Hypervideo Model for the Web</article-title>
          .
          <source>Multimedia Tools and Applications</source>
          , pages
          <volume>1</volume>
          {
          <fpage>35</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>M.</given-names>
            <surname>Sporny</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Longley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Kellogg</surname>
          </string-name>
          , et al.
          <article-title>JSON-LD 1.0, A JSON-based Serialization for Linked Data</article-title>
          . Proposed Recommendation, W3C, Nov.
          <year>2013</year>
          . http://www.w3.org/TR/json-ld/.
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>T.</given-names>
            <surname>Steiner</surname>
          </string-name>
          .
          <article-title>SemWebVid { Making Video a First Class Semantic Web Citizen and a First Class Web Bourgeois</article-title>
          . In A. Polleres and H. Chen, editors,
          <source>Proceedings of the ISWC 2010 Posters &amp; Demonstrations Track: Collected Abstracts</source>
          , Shanghai, China, November 9,
          <year>2010</year>
          , volume
          <volume>658</volume>
          <source>of CEUR Workshop Proceedings ISSN 1613-0073</source>
          , pages
          <fpage>97</fpage>
          {
          <fpage>100</fpage>
          ,
          <string-name>
            <surname>Nov</surname>
          </string-name>
          .
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>B.</given-names>
            <surname>Tillett</surname>
          </string-name>
          .
          <article-title>FRBR: A Conceptual Model for the Bibliographic Universe</article-title>
          .
          <source>Technical report</source>
          ,
          <year>2004</year>
          . http://www.loc.gov/cds/downloads/FRBR.PDF.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>R.</given-names>
            <surname>Troncy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Mannens</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          <article-title>Pfei er</article-title>
          , et al.
          <source>Media Fragments URI 1</source>
          .0 (
          <issue>basic</issue>
          ).
          <source>Recommendation, W3C</source>
          , Sept.
          <year>2012</year>
          . http://www.w3.org/TR/media-frags/.
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <surname>D. Van Deursen</surname>
            ,
            <given-names>W. Van</given-names>
          </string-name>
          <string-name>
            <surname>Lancker</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Mannens</surname>
          </string-name>
          , et al.
          <source>Experiencing Standardized Media Fragment Annotations Within HTML5. Multimedia Tools and Applications</source>
          , pages
          <volume>1</volume>
          {
          <fpage>20</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>