<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Wikidata: A platform for data integration and dissemination for the life sciences and beyond</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Elvira Mitraka</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andra Waagmeester</string-name>
          <email>andra@micelio.be</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sebastian Burgstaller-Muehlbacher</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lynn M. Schriml</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrew I. Su</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Benjamin M. Good</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Maryland School of Medicine</institution>
          ,
          <addr-line>Baltimore, USA Micelio, Antwerp</addr-line>
          ,
          <institution>Belgium Department of Molecular and Experimental Medicine, Scripps Research Institute</institution>
          ,
          <addr-line>La Jolla</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Wikidata is an open, Semantic Web-compatible database that anyone can edit. This 'data commons' provides structured data for Wikipedia articles and other applications. Every article on Wikipedia has a hyperlink to an editable item in this database. This unique connection to the world's largest community of volunteer knowledge editors could help make Wikidata a key hub within the greater Semantic Web. The life sciences, as ever, faces crucial challenges in disseminating and integrating knowledge. Our group is addressing these issues by populating Wikidata with the seeds of a foundational semantic network linking genes, drugs and diseases. Using this content, we are enhancing Wikipedia articles to both increase their quality and recruit human editors to expand and improve the underlying data. We encourage the community to join us as we collaboratively create what can become the most used and most central semantic data resource for the life sciences and beyond.</p>
      </abstract>
      <kwd-group>
        <kwd>Wikidata</kwd>
        <kwd>Wikipedia</kwd>
        <kwd>Linked Data</kwd>
        <kwd>Semantic Web</kwd>
        <kwd>Crowdsourcing</kwd>
        <kwd>Knowledge Management</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        In the Stone Soup folktale [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], a group of hungry travelers arrive in a village
with its inhabitants unwilling to share their food. With a kettle of water and a stone
the travelers manage to touch the curiosity of the villagers. The curiosity finally
spawns a collaborative effort to make a great soup. This story is nowadays used to
express the power of crowdsourcing and collaborative projects [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], such as Wikipedia,
where many individuals each make small contributions but collectively produce
something larger than the sum of its parts. Wikidata extends this collaborative model
to the Web of data [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. In this article we will describe Wikidata and the ways that this
open public platform can take a central role in data sharing and management for the
life science community.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Wikidata and Wikipedia</title>
      <p>
        Wikipedia is among the most visited sites on the Internet. Articles about medical
topics were viewed more than 4.88 billion times in 2013, a number on par with
http://nih.gov and significantly greater than WebMD [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. This incredibly important
resource, created through volunteer labor, is now tightly coupled to Wikidata - an
open, Semantic Web-compatible database that anyone can edit [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Wikipedia
infoboxes - the tables of data often appearing on the right side of articles - can now
render content stored in Wikidata and each Wikipedia article now has a direct link to
the corresponding Wikidata item, thus encouraging the collaborative editing of the
data (Fig. 1).
      </p>
      <p>
        Infoboxes provide the bridge between machine-readable structured data and
the unstructured text that forms the main body of each article. Since 2008, the Gene
Wiki project has automatically created and maintained the infoboxes for around
10000 articles about human genes [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Now, this initiative is focused on generating a
foundation of biomedical knowledge in Wikidata that will be used to improve infobox
content on Wikipedia and help drive new applications. To date, we have loaded
Wikidata with items about: 56451 human and 73086 mouse genes from NCBI Gene
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], 6562 concepts in the Disease Ontology [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], and 1830 FDA-approved drugs. This
initial data load generated Wikidata items for these key biomedical concepts, mapped
them to Wikipedia articles and linked them to the corresponding identifiers in
authoritative public databases. The identifier-level connections to the source databases
ensure that Wikidata content can be easily integrated into the existing Web of
biomedical data. Moreover, the provenance of all Wikidata claims can be assessed through
inspection of the supporting references. The data is kept up to date by periodically
running ‘bots’ that propagate changes from authoritative sources to Wikidata. When
conflicts arise from human edits to Wikidata items, these are flagged for manual
review. The next phase of the project will stitch these concepts into a richly
interconnected semantic network.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Taking a sip of the data soup – Wikidata and the Semantic Web</title>
      <p>
        The first application to use Wikidata extensively is Wikipedia but this could
be the tip of the iceberg. To give a preview of what Wikidata could become, it’s
useful to briefly examine its closest ancestor, DBpedia. The DBpedia project mines
content from Wikipedia by parsing infoboxes, maps this content to their own ontology,
and provides access to this data in the form of a large RDF database available both for
bulk download and SPARQL query. While enabling interesting queries on its own,
its most important function is as a global linking hub for the Semantic Web [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. In
comparison to DBpedia, Wikidata has a number of advantages. First, it can be edited
directly and changes are reflected in real time. Second, it does not require any parsing
because all data is managed in a database from the outset. Third, it contains large
amounts of content that is not present in Wikipedia, such as items for every mouse
gene. Finally, its query API supports not only queries along its asserted knowledge
graph, but also along references, qualifiers and even edit histories. These additional
capabilities, viewed in light of the success of the DBpedia project, portend a vital
future for Wikidata in the context of the Semantic Web.
      </p>
      <p>
        Within the biomedical domain, useful queries are already possible as a result
of the ‘single-pot’ nature of Wikidata. For example, it is possible to use Wikidata’s
SPARQL endpoint (https://query.wikidata.org/) to answer questions such as “what
clinically relevant drug-drug interactions are known for the drug methadone
(CHEMBL651)” [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Importantly, the data used to answer this query came from two
groups working completely independently. Our ‘drug_bot’ bot added the CHEMBL
identifiers (as well as many other identifiers) while another bot developed by a team
at the Medical University of Vienna added the drug-drug interactions [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. This
happened without any direct coordination between our groups.
      </p>
      <p>This kind of serendipitous, automatic, cross-continental data integration is
the primary goal of the Semantic Web, but is not yet commonplace. The key beauty
and main challenge of the Semantic Web is its distributed nature. In order for this
kind of integration to happen in the absence of a centralized resource like Wikidata,
several major hurdles would need to be leaped. First, both teams would need to know
enough about the fairly complex stack of semantic technologies to provide their data
as RDF through a stable, public SPARQL endpoint. Second, they would have to
work with overlapping identifier systems. Third, the would-be consumer of their data
would need to discover both of their endpoints and be sophisticated enough with
SPARQL to identify and issue the appropriate distributed query. All of this is
possible and can work, but it is not easy.</p>
      <p>
        By integrating data in a centralized, single community pot, Wikidata
provides a platform that addresses each of these problems. Data providers do not have to
set up and maintain their own SPARQL endpoint – a challenge that very few teams
have succeeded at doing for any length of time [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. By virtue of working in the same
database, it is far less likely - though not impossible - for independent teams to
generate and publish different identifiers, as the first step in working with Wikidata is to
query it to see what is already there. Finally, the challenge of finding a relevant
endpoint is negated when there is only one. Note that Wikidata can be queried using
SPARQL or the Wikidata Query Language [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
4
      </p>
      <p>The fact that Wikidata is one centralized, community resource immediately
surfaces the challenges incurred in any collaborative ontology development
process. In Wikidata, the ‘ontology’ corresponds to its collection of linking properties
used to describe items. A new property in Wikidata has to be proposed for
community discussion and is only created after a consensus regarding the value of the property
and its relation to existing properties has been established. For those used to
controlling their own data and data models, this process can feel tedious. But this same
fundamental process must be undertaken in any attempt at data integration. The fact that
it happens up front, when data is first being loaded, should help to keep the data
consistent and reduce the downstream identifier and ontological mapping problems that
continue to plague bioinformatics.</p>
      <p>Imagine the power of combining the structured data in Wikidata, the high
accessibility and dedicated community of Wikipedia and the knowledge of the
scientific community. Contemplate further that all of this data is freely available and
accessible through a stable query interface and robust, read/write API. This makes
important, high-quality information easily accessible by anyone and opens up scientific
knowledge for public scrutiny. Further, the built-in provenance tracking can provide
detailed chains of evidence to support or refute each claim and all of this can be
discussed using the many social tools, such as ‘talk pages’ for every data item, baked
into the MediaWiki infrastructure.</p>
      <p>Aside from creating useful ways to disseminate data, this sociotechnical
structure provides a framework for the broad community to broadcast feedback back
to the original data owners. Even at this early stage of this project, this process has
already led to improvements in source data. For example, in the Disease Ontology the
term ‘Ollier disease’ had the synonym ‘Maffucci syndrome’. Upon importing the
Disease Ontology into Wikidata, members of the Wikidata community pointed out
that the two terms, though putative synonyms, linked to two different extant Wikidata
items. Upon closer review it was determined that these two terms represent two
different, albeit closely related, diseases, leading to the creation of a new term in the
Disease Ontology. As Wikidata expands it is to be expected that additional
differences in representation between it and other knowledge resources will surface. These
will first be triaged by the Wikidata community to check for errors and, if consensus
is achieved that there is an error in the original source, this will be relayed for
consideration. In this way, the Wikidata community can become the ‘many eyes’ that make
all ontology bugs shallow.
...Can Make a Delicious Soup</p>
      <p>We can create a powerful commons of biomedical knowledge by building on
established resources and the dedicated community to connect genes, proteins, drugs,
diseases, phenotypes and symptoms. Wikipedia will be the first application to use the
content in Wikidata, but certainly not the last. The fire is ready and the pot is starting
to heat up. Some villagers are already peeking out of their windows ready to join us
around the pot, but it will take the effort of the whole community to make a delicious
biomedical data soup. We invite you to join us in this effort.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <article-title>1. History of the Stone Soup Story from 1720 to now</article-title>
          . Available from: http://www.stonesoup.
          <article-title>com/history-of-the-stone-soup-story-</article-title>
          <string-name>
            <surname>from-</surname>
          </string-name>
          1720
          <string-name>
            <surname>-</surname>
          </string-name>
          to-now/.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Taylor</surname>
          </string-name>
          . J.
          <source>The Stone Soup of Data</source>
          .
          <year>2007</year>
          8 May; Available from: https://km.aifb.kit.edu/ws/ckc2007/StoneSoup-www2007.
          <fpage>pdf</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Vrandečić</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Krötzsch</surname>
          </string-name>
          ,
          <article-title>Wikidata: A Free Collaborative Knowledgebase</article-title>
          ,
          <source>in Communications of the ACM</source>
          .
          <year>2014</year>
          , ACM. p.
          <fpage>78</fpage>
          -
          <lpage>85</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Heilman</surname>
            ,
            <given-names>J.M.</given-names>
          </string-name>
          and
          <string-name>
            <given-names>A.G.</given-names>
            <surname>West</surname>
          </string-name>
          , Wikipedia and medicine: quantifying readership, editors, and
          <article-title>the significance of natural language</article-title>
          .
          <source>J Med Internet Res</source>
          ,
          <year>2015</year>
          .
          <volume>17</volume>
          (
          <issue>3</issue>
          ): p.
          <fpage>e62</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Huss</surname>
            ,
            <given-names>J.W.</given-names>
          </string-name>
          ,
          <year>3rd</year>
          , et al.,
          <article-title>A gene wiki for community annotation of gene function</article-title>
          .
          <source>PLoS Biol</source>
          ,
          <year>2008</year>
          .
          <volume>6</volume>
          (
          <issue>7</issue>
          ): p.
          <fpage>e175</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Brown</surname>
            ,
            <given-names>G.R.</given-names>
          </string-name>
          , et al.,
          <article-title>Gene: a gene-centered information resource at NCBI</article-title>
          .
          <source>Nucleic Acids Res</source>
          ,
          <year>2015</year>
          .
          <volume>43</volume>
          (Database issue): p.
          <fpage>D36</fpage>
          -
          <lpage>42</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Kibbe</surname>
            ,
            <given-names>W.A.</given-names>
          </string-name>
          , et al.,
          <source>Disease Ontology</source>
          <year>2015</year>
          update
          <article-title>: an expanded and updated database of human diseases for linking biomedical knowledge through disease data</article-title>
          .
          <source>Nucleic Acids Res</source>
          ,
          <year>2015</year>
          .
          <volume>43</volume>
          (Database issue): p.
          <fpage>D1071</fpage>
          -
          <lpage>8</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          , et al.,
          <article-title>DBpedia - A crystallization point for the Web of Data</article-title>
          .
          <source>Web Semantics: Science, Services and Agents on the World Wide Web</source>
          ,
          <year>2009</year>
          .
          <volume>7</volume>
          (
          <issue>3</issue>
          ): p.
          <fpage>154</fpage>
          -
          <lpage>165</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <article-title>Get all the drug-drug interactions for Methadone based on its CHEMBL id CHEMBL651</article-title>
          .
          <year>2015</year>
          [cited 2015 Sep. 14]; Available from: https://bitbucket.org/sulab/wikidatasparqlexamples/overview#markdown
          <article-title>-header-get-allthe-drug-drug-interactions-for-methadone-based-on-its-chembl-id-chembl651.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Pfundner</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , et al.,
          <article-title>Utilizing the Wikidata system to improve the quality of medical content in Wikipedia in diverse languages: a pilot study</article-title>
          .
          <source>J Med Internet Res</source>
          ,
          <year>2015</year>
          .
          <volume>17</volume>
          (
          <issue>5</issue>
          ): p.
          <fpage>e110</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Buil-Arand</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          , et al.
          <source>SPARQL Web-Querying Infrastructure: Ready for Action? in 12th International Semantic Web Conference</source>
          .
          <year>2013</year>
          . Sydney, Australia.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12. Wikidata Query Editor.
          <source>[cited</source>
          <year>2015</year>
          ; Available from: https://wdq.wmflabs.org/wdq/.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>