<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>H. Turki);</journal-title>
      </journal-title-group>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Automating the use of Shape Expressions for the validation of semantic knowledge in Wikidata</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Houcemeddine Turki</string-name>
          <email>turkiabdelwaheb@hotmail.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mohamed Ali Hadj Taieb</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Khalil Chebil</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mohamed Ben</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aouicha</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lane Rasberry</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Daniel Mietchen</string-name>
          <email>daniel.mietchen@ronininstitute.org</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Data Engineering and Semantics Research Unit, Faculty of Sciences of Sfax, University of Sfax</institution>
          ,
          <addr-line>Sfax</addr-line>
          ,
          <country country="TN">Tunisia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>FIZ Karlsruhe - Leibniz Institute for Information Infrastructure</institution>
          ,
          <addr-line>Berlin</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Ronin Institute for Independent Scholarship</institution>
          ,
          <addr-line>Montclair, New Jersey</addr-line>
          ,
          <country country="US">United States of America</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>School of Data Science, University of Virginia</institution>
          ,
          <addr-line>Charlottesville, VA</addr-line>
          ,
          <country country="US">United States of America</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Wikidata, Knowledge Graph Validation</institution>
          ,
          <addr-line>Shape Expressions, SPARQL, Semantic Alignment, Data Quality</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2014</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>matter. In this position paper, we discuss the semantic alignment-based approach for automating the ShEx-based validation of Wikidata items as proposed by Wikimedia Deutschland in July 2023, and we propose an alternative method that automates the shape-based validation of Wikidata entities and statements based on the conversion of ShEx EntitySchemas into SPARQL queries that identify relevant entities. We explain the advantages and drawbacks of both methods to provide the community with a useful overview of the Wikidata'23: Wikidata Workshop at ISWC 2023 ∗Corresponding author.</p>
      </abstract>
      <kwd-group>
        <kwd>Wikidata</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        With the rise and growth of open knowledge graphs, ensuring their quality becomes increasingly
challenging [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Particularly, quality assessment has been a vital component of the development
of Wikidata as an open and collaborative knowledge graph, as it follows the same principles as
Wikipedia, including its quest for consistency and completeness [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Subsequently, Wikidata
has supported the creation of EntitySchemas implemented in Shape Expressions (ShEx) to
ensure the shape-based validation of the Wikidata items [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], following preliminary experiments
conducted in 2019 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Currently, there are multiple EntitySchemas to validate a number of
Wikidata classes1, particularly the ones related to the molecular biology field [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Several eforts
have also been directed toward creating tools for the automatic generation of EntitySchemas,
like sheXer [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. There are even several adaptations of the Shape Expressions (ShEx) language to
CEUR
Workshop
Proceedings
make it more intuitive for Wikidata users to write EntitySchemas, like ShExStatements [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and
WShEx [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Furthermore, there are several initiatives to combine ShEx with SPARQL to enhance
its scope beyond shape-based validation [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Yet, there is limited progress toward automating
the use of ShEx EntitySchemas for the validation of Wikidata entities. A new initiative of
Wikimedia Deutschland aims to solve this problem based on creating semantic alignments
between ShEx EntitySchemas and Wikidata Classes.
      </p>
      <p>In this position paper, we describe the current situation of the ShEx-based validation of
semantic knowledge in Wikidata (Section 2). Then, we outline the approach of Wikimedia
Deutschland for automating the ShEx-based validation of Wikidata (Section 3). After that,
we propose an alternative approach for the ShEx-based validation of Wikidata driven by the
conversion of ShEx EntitySchemas into corresponding SPARQL queries (Section 4). Later,
we discuss the strengths and limitations of both approaches to give the Wikidata community
diferent perspectives on how to handle this issue (Section 5). Finally, we draw conclusions and
propose future directions for this research work (Section 6).</p>
    </sec>
    <sec id="sec-3">
      <title>2. ShEx-based validation of semantic knowledge in Wikidata</title>
      <p>
        In this section, we delve into the utilization of Shape Expressions (ShEx) for the validation of
semantic knowledge in Wikidata. ShEx provides a structured approach to assess the conformity
of Wikidata entities to predefined schemas. These schemas, known as EntitySchemas, define
the expected structure and constraints for diferent classes of entities. The integration of ShEx
into Wikidata’s data quality control framework introduces a standardized method for ensuring
data accuracy and consistency. The main preconditions for using ShEx to validate Wikidata
entries are as follows:
• EntitySchemas: EntitySchemas are at the core of ShEx-based validation. They define
the structural expectations and constraints for specific classes of entities within Wikidata.
Currently, Wikidata has approximately 100 million items spanning several million classes.
However, there are only around 400 EntitySchemas available. These schemas cover a
wide spectrum of classes, from highly populated ones like ”human” (E10) and ”city” (E100)
to less-represented ones like ”Nazca lines” (E148) and ”ethics committees” (E396). The
level of detail and complexity of these schemas varies significantly. Some employ direct
constraints, which are single straightforward one-line constraints, while others use open
constraints, where the object is a variable, and indirect constraints, which are sets of
chained constraints where the object of one constraint is the subject of another constraint.
There are also closed constraints, which define constraints where the object is a predefined
set of items (See the EntitySchema for a deceased person [E105] for examples of every
type of condition, as per Figure 1). Some EntitySchemas even rely on other schemas
through the use of IMPORT clauses to define the entities corresponding to them and to
provide further constraints on how entities should be defined (e.g., E192 for a virus taxon
refers to E69 when describing diseases that the virus might be involved in). This diversity
reflects the heterogeneous nature of Wikidata’s data structure.
• Identifying Entities: To apply a specific schema to a group of entities, Wikidata relies
on SPARQL queries. Schema documentation often includes sample SPARQL queries that
can be executed using tools like the Wikidata Query Service or RDF dump-based external
SPARQL endpoints [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. These queries identify the entities subject to validation based on
specified criteria, allowing for a focused validation process.
• Validation Tools: The validation process itself is facilitated by dedicated tools. On each
EntitySchema’s Wikidata page, there is a ”check entities against this Schema” hyperlink
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. This link directs users to the Simple Online Validator2, a tool hosted on the Wikimedia
Toolforge infrastructure. The validator consumes the EntitySchema, executes SPARQL
queries to retrieve relevant entities, applies the schema’s constraints, and reports
compliance or non-compliance, providing detailed feedback when necessary. Additionally,
other validation tools and methods, such as those for identifying subsets of Wikidata [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ],
are available to users.
• Signaling Compliance: Presently, the signaling of ShEx schema compliance to Wikidata
users is somewhat limited. While dedicated validation tools report compliance or
noncompliance, this information is not prominently integrated into the SPARQL endpoint
or the graphical user interface (GUI). Nonetheless, Wikidata employs various non-ShEx
quality control mechanisms [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], such as constraint statements and bot-curated pages, to
identify and report data quality issues [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Adapting these workflows to incorporate
ShEx-based compliance reporting remains a possibility.
      </p>
      <p>ShEx schemas provide a structured framework for assessing data quality in Wikidata, yet
several challenges and considerations persist. One fundamental concern revolves around the
coverage of EntitySchemas in relation to the vast number of entities within Wikidata. It is crucial
to ascertain the extent to which the 400 schemas encompass Wikidata’s diverse entity classes.
It is conceivable that these schemas predominantly address well-represented classes, potentially
leaving a long tail of classes with limited coverage. Additionally, the subjective nature of schema
definitions, particularly for highly specific or niche classes, may have contributed to the slow
adoption of ShEx in Wikidata.</p>
      <p>Another impediment to the widespread adoption of ShEx in Wikidata has been the absence
of automated tools for ShEx-based validation. In the subsequent sections, we will introduce two
solutions to address this issue. The first solution, proposed by Wikimedia Deutschland,
leverages semantic alignments between Wikidata classes and EntitySchemas. The second solution,
presented in this position paper, revolves around the transformation of ShEx EntitySchemas
into corresponding SPARQL queries through a rule-based approach.</p>
    </sec>
    <sec id="sec-4">
      <title>3. The Wikimedia Deutschland solution: Semantic alignment between Wikidata classes and EntitySchemas</title>
      <p>In July 2023, the development team of Wikimedia Deutschland deployed a new property in
test Wikidata3, the sandbox for Wikidata-related experiments, to assign EntitySchemas to their
corresponding Wikidata classes. This property is called EntitySchema for this class4 and has
2https://shex-simple.toolforge.org/wikidata/
3https://test.wikidata.org.
4https://test.wikidata.org/wiki/Property:P97725.
already been discussed by the community since 28 May 20195. The Proposal implies the creation
of a new datatype for EntitySchemas and proposes to align the Wikidata classes as subjects
to EntitySchemas as objects of the EntitySchema for this class relations, as shown in Figure
2. This will allow the development of tools that process the taxonomic relations of a given
item in Wikidata (i.e., instance of [P31], subclass of [P279], and part of [P361]) to identify the
EntitySchemas that are relevant to use for checking the consistency of the considered entity.</p>
    </sec>
    <sec id="sec-5">
      <title>4. Our solution: SPARQL-based identification of relevant items</title>
      <p>What we propose is to analyze the EntitySchema itself to identify the Wikidata items that
are relevant to it. This should be enabled by converting the ShEx statements into a SPARQL
query that can be used to retrieve the Wikidata items that should be considered. This solution
has been proposed since the early days of ShEx in 20176. The principle is based on using
closed constraints to create a SPARQL query to find the Wikidata items corresponding to the
considered EntitySchema. By a closed constraint, we mean the chain of ShEx statements having
a definite set of Wikidata items (beginning with wd:Q) as a final object, as shown in Figure 1.</p>
      <p>
        As SPARQL serves as the query language for RDF knowledge graphs, and ShEx functions
as the semantic web language for shape-based validation of knowledge graphs, it’s important
to note that these two languages do not share identical syntax or structure [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Therefore,
when dealing with closed constraints, it becomes necessary to undergo a conversion process,
transforming them into SPARQL statements prior to their utilization in retrieving specific items
from Wikidata. This conversion process is illustrated in Figure 3. In Figure 3, we can observe
this conversion process in action, employing the example of E50, which represents the Wikidata
EntitySchema for national flags. Additionally, the figure demonstrates the efectiveness of this
approach using the schema for genes or variants with references (E390), thereby proving the
versatility of this method across EntitySchemas, regardless of their complexity or scope.
      </p>
      <p>
        A preliminary edition of the source code that can convert EntitySchemas into SPARQL queries
is currently available at https://github.com/csisc/WikidataShExSPARQL. The processing of
closed constraints involves the parsing of curly brackets using the parentheses level count
algorithm to assign opening curly brackets to their corresponding closing curly brackets [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ],
the elimination of several non-alphanumerical characters (e.g., *, #, and +), and the removal of
EXTRA properties if directly following a variable name. Later, a set of rules will be used to
transform the remaining skeleton of the ShEx EntitySchema into a SPARQL query, as shown in
Table 1. The obtained SPARQL query will be run through the SPARQL endpoint of Wikidata
(https://query.wikidata.org) to identify all the Wikidata items that can be validated using the
considered EntitySchema.
      </p>
    </sec>
    <sec id="sec-6">
      <title>5. Discussion</title>
      <p>The method proposed by Wikimedia Deutschland introduces a valuable approach to aligning
Wikidata classes with EntitySchemas. However, it is essential to acknowledge its limitations,
5https://www.wikidata.org/wiki/Wikidata:Property_proposal/Shape_Expression_for_class.
6https://github.com/shexSpec/shex/issues/75.
&lt; 1&gt; { 1 EXTRA  2  1}
PREFIX  : &lt; &gt;</p>
      <p>SPARQL
? 1  1  1;  2  2.
? 1 1 1 [1 2  1].</p>
      <p>VALUES ?o { 1  2}
? 1  1 ?o.
? 1
SELECT ?id WHERE {
{?id Statements for  1}
UNION
{?id Statements for  2}
}
? 1  1 [ 2  1].</p>
      <p>
        PREFIX  : &lt; &gt;
particularly when EntitySchemas are not directly related to Wikidata classes. For instance,
consider an EntitySchema designed for Tunisian scientists. In many cases, Wikidata items within
this category may not explicitly assign Tunisian scientist as the object of an instance of [P31]
or facet of [P1269] statement. Instead, such items are often represented as a combination of
properties like {”occupation”, ”scientist”} (P106, Q901) and {”country of citizenship”, ”Tunisia”}
(P27, Q948). Adding items like ”Tunisian scientist” as objects of ”facet of” [P1269] relations could
result in an influx of redundant data without enriching the underlying semantic knowledge,
which is inadvisable given the growing size of Wikidata and its data storage challenges [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
Another potential solution is the use of reification to specify requirements for an item to be
considered beyond class membership [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], as illustrated in Figure 4. The same solution can
be applied by using the ”category contains” [P4224] property to specify how class members
should be defined in Wikidata, as in https://www.wikidata.org/wiki/Q6471216#P4224 [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
However, this approach may introduce additional complexity when developing tools based
on ”EntitySchema for this class” statements to identify corresponding EntitySchemas for a
Wikidata item [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>Regarding our proposed solution, which aims to generate SPARQL queries to identify
Wikidata items corresponding to ShEx EntitySchemas, it ofers flexibility by implicitly inferring the
conditions for item inclusion without the need to store them as RDF triples [14]. The
formulation of constraints as SPARQL queries provides adaptability, making it possible to identify
corresponding items regardless of the complexity of the constraints [15]. While the generated
SPARQL queries are primarily intended for retrieving the set of Wikidata items corresponding
to an EntitySchema (i.e., subsetting) [14], a slight adaptation of the query can eficiently verify
whether an item meets the specified criteria (as depicted in the red query in Figure 5). If an
item does not conform to the query, it will return an empty result; if it does, it will return the
Wikidata ID of the item as a result.</p>
      <p>Our method may have the drawback of increasing the computational stress on the Wikidata
Query Service (https://query.wikidata.org). The SPARQL endpoint of Wikidata has faced
scalability challenges, leading to several outages since 20207. Any additional workload imposed
on the Wikidata Query Service may exacerbate its current situation. In contrast, the approach
proposed by Wikimedia Deutschland does not rely on advanced computations, making it an
attractive option.</p>
      <p>In terms of recall, both approaches have inherent limitations. The proposed methods rely
on the availability and accuracy of data within Wikidata. If a schema imposes highly specific
constraints that are not fully represented in Wikidata entries, there may be entities that go
unidentified. For instance, identifying a ”Tunisian scientist” would be challenging if the
occupation and nationality are not explicitly documented in Wikidata entries. While these limitations
are challenging to completely mitigate, they should be considered when implementing
ShExbased validation methods in Wikidata and when interpreting the results. Further research and
refinement of these methods may help enhance their recall capabilities in the future.</p>
    </sec>
    <sec id="sec-7">
      <title>6. Conclusion</title>
      <p>In this position paper, we explain the current eforts for automating the shape-based validation of
Wikidata items based on ShEx EntitySchemas and we propose our preliminary solution for this
matter that is driven by converting ShEx EntitySchemas into SPARQL queries to identify relevant
items. We have shown that our solution is more eficient than the one proposed by Wikimedia
Deutschland, as it requires less data storage and supports complex constraints that cannot
be dealt with using semantic alignments between Wikidata items and ShEx EntitySchemas.
Although our solution is more practical, it requires an upgrade to overcome its limitations
caused by the performance limits of the Wikidata Query Service. As a future direction of
this work, we propose to optimize our source code for eficiency by adjusting its layout and
substituting the use of the Wikidata Query Service with other methods for mining Wikidata,
like CirrusSearch8.</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgments</title>
      <p>This research is funded by the Wikimedia Research Fund of Wikimedia Foundation (San
Francisco, California, United States of America) through the Adapting Wikidata to support clinical
practice using Data Science, Semantic Web and Machine Learning Project.9 Source code is made
available under the MIT License at https://github.com/csisc/WikidataShExSPARQL.
7See https://www.wikidata.org/wiki/Wikidata:SPARQL_query_service/WDQS_backend_update for further details.
8https://www.mediawiki.org/wiki/Extension:CirrusSearch
9https://meta.wikimedia.org/wiki/Research:Adapting_Wikidata_to_support_clinical_practice_using_Data_Science,
_Semantic_Web_and_Machine_Learning
(Wikidata@ISWC 2021), 2021, p. 8. URL: https://ceur-ws.org/Vol-2982/paper-8.pdf.
[14] J. E. Labra-Gayo, A. G. Hevia, D. F. Álvarez, A. Ammar, D. Brickley, A. J. G. Gray,
E. Prud'hommeaux, D. Slenter, H. Solbrig, S. A. H. Beghaeiraveri, B. Fünkfstük, A.
Waagmeester, E. Willighagen, L. Ovchinnikova, G. Benjaminsen, R. G. González, L. J. Castro,
D. Mietchen, Knowledge graphs and Wikidata subsetting, BioHackathon Europe 2020
(2021). doi:10.37044/osf.io/wu9et.
[15] V. Nguyen, O. Bodenreider, A. Sheth, Don't like RDF reification?, in: Proceedings of the
23rd international conference on World wide web, ACM, 2014. URL: https://doi.org/10.
1145/2566486.2567973. doi:10.1145/2566486.2567973.
c
S e</p>
      <p>9
t c 3
x
E
m t o
u n tt</p>
      <p>E o</p>
      <p>b
n /
a i
h
k ,</p>
      <p>g
c
ro f
P co (3 re
:
3
R s
PA lse</p>
      <p>d
S se i</p>
      <p>u i
g
idn fo .
n n
w la</p>
      <p>F
g l
a
n
o
t i
a t</p>
      <p>a
k (N
w 05
w :E
w a</p>
      <p>w em
s
e
r
r
o
c li
o E
t ,
n )
i
s (
t s
n t</p>
      <p>/
a / h
n :
i s c</p>
      <p>S
m tp</p>
      <p>e E
2 c /
t ty
h i</p>
      <p>t
: n
r i</p>
      <p>k
ou i
e i
m ra )[
e t
t
n S /
w
g
4 r</p>
      <p>o
s (
n
a .
t a
s co ry t
x e a
hE epn u id
q k</p>
      <p>i
S o
fo fo</p>
      <p>L
Q .w</p>
      <p>R w</p>
      <p>1
r (
E l</p>
      <p>t
t h
a ,</p>
      <p>)
u p
m to
roF ,se</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>K.</given-names>
            <surname>Shenoy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ilievski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Garijo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Schwabe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Szekely</surname>
          </string-name>
          ,
          <article-title>A study of the quality of Wikidata</article-title>
          ,
          <source>Journal of Web Semantics</source>
          <volume>72</volume>
          (
          <year>2022</year>
          )
          <article-title>100679</article-title>
          . doi:
          <volume>10</volume>
          .1016/j.websem.
          <year>2021</year>
          .
          <volume>100679</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Piscopo</surname>
          </string-name>
          , E. Simperl,
          <article-title>What we talk about when we talk about Wikidata quality: a literature survey</article-title>
          ,
          <source>in: Proceedings of the 15th International Symposium on Open Collaboration, ACM</source>
          ,
          <year>2019</year>
          . doi:
          <volume>10</volume>
          .1145/3306446.3340822.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Waagmeester</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. L.</given-names>
            <surname>Willighagen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. I.</given-names>
            <surname>Su</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kutmon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. E. L.</given-names>
            <surname>Gayo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>FernándezÁlvarez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Groom</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. J.</given-names>
            <surname>Schaap</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. M.</given-names>
            <surname>Verhagen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. J.</given-names>
            <surname>Koehorst</surname>
          </string-name>
          ,
          <article-title>A protocol for adding knowledge to Wikidata: aligning resources on human coronaviruses</article-title>
          ,
          <source>BMC Biology 19</source>
          (
          <year>2021</year>
          ). doi:
          <volume>10</volume>
          .1186/s12915-020-00940-y.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>K.</given-names>
            <surname>Thornton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Solbrig</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. S.</given-names>
            <surname>Stupp</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. E. L.</given-names>
            <surname>Gayo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Mietchen</surname>
          </string-name>
          ,
          <string-name>
            <surname>E.</surname>
          </string-name>
          <article-title>Prud'hommeaux, A. Waagmeester, Using Shape Expressions (ShEx) to Share RDF Data Models and to Guide Curation with Rigorous Validation</article-title>
          , in: The Semantic Web, Springer International Publishing,
          <year>2019</year>
          , pp.
          <fpage>606</fpage>
          -
          <lpage>620</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -21348-0_
          <fpage>39</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Samuel</surname>
          </string-name>
          , ShExStatements: Simplifying Shape Expressions for Wikidata,
          <source>in: Companion Proceedings of the Web Conference</source>
          <year>2021</year>
          , ACM,
          <year>2021</year>
          . doi:
          <volume>10</volume>
          .1145/3442442.3452349.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J. E.</given-names>
            <surname>Labra-Gayo</surname>
          </string-name>
          ,
          <article-title>Wshex: A language to describe and validate wikibase entities</article-title>
          ,
          <source>in: Proceedings of the 3rd Wikidata Workshop</source>
          ,
          <year>2022</year>
          , p.
          <fpage>3</fpage>
          . URL: https://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>3262</volume>
          / paper3.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>H.</given-names>
            <surname>Turki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A. H.</given-names>
            <surname>Taieb</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Shafee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lubiana</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Jemielniak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. B.</given-names>
            <surname>Aouicha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. E. L.</given-names>
            <surname>Gayo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. A.</given-names>
            <surname>Youngstrom</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Banat</surname>
          </string-name>
          ,
          <string-name>
            <surname>D. Das</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Mietchen</surname>
          </string-name>
          ,
          <article-title>on behalf of WikiProject COVID</article-title>
          ,
          <string-name>
            <surname>Representing</surname>
            <given-names>COVID</given-names>
          </string-name>
          -
          <article-title>19 information in collaborative knowledge graphs: The case of Wikidata, Semantic Web 13 (</article-title>
          <year>2022</year>
          )
          <fpage>233</fpage>
          -
          <lpage>264</lpage>
          . doi:
          <volume>10</volume>
          .3233/sw-210444.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>D.</given-names>
            <surname>Hernández</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hogan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Riveros</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Rojas</surname>
          </string-name>
          , E. Zerega,
          <article-title>Querying wikidata: Comparing SPARQL, relational and graph databases</article-title>
          ,
          <source>in: Lecture Notes in Computer Science</source>
          , Springer International Publishing,
          <year>2016</year>
          , pp.
          <fpage>88</fpage>
          -
          <lpage>103</lpage>
          . URL: https://doi.org/10.1007/ 978-3-
          <fpage>319</fpage>
          -46547-0_
          <fpage>10</fpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>319</fpage>
          -46547-0_
          <fpage>10</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J. E.</given-names>
            <surname>Labra-Gayo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. C.</given-names>
            <surname>González Cavazos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Waagmeester</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Hofmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Hosseini Beghaeiraveri</surname>
          </string-name>
          , E. Prud'hommeaux,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ul-Hasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Willighagen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ammar</surname>
          </string-name>
          ,
          <article-title>Enhancement and reusage of biomedical knowledge graph subsets</article-title>
          ,
          <source>BioHackrXiv</source>
          (
          <year>2022</year>
          ). URL: https://doi.org/10.37044/osf.io/n7qku. doi:
          <volume>10</volume>
          .37044/osf.io/n7qku.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>H.</given-names>
            <surname>Turki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Jemielniak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A. H.</given-names>
            <surname>Taieb</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. E. L.</given-names>
            <surname>Gayo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. B.</given-names>
            <surname>Aouicha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Banat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Shafee</surname>
          </string-name>
          , E. Prud'hommeaux, T. Lubiana,
          <string-name>
            <surname>D. Das</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Mietchen</surname>
          </string-name>
          ,
          <article-title>Using logical constraints to validate statistical information about disease outbreaks in collaborative knowledge graphs: the case of COVID-19 epidemiology in wikidata</article-title>
          ,
          <source>PeerJ Computer Science</source>
          <volume>8</volume>
          (
          <year>2022</year>
          )
          <article-title>e1085</article-title>
          . URL: https://doi.org/10.7717/peerj-cs.
          <volume>1085</volume>
          . doi:
          <volume>10</volume>
          .7717/peerj-cs.
          <volume>1085</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>H.</given-names>
            <surname>Turki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A. H.</given-names>
            <surname>Taieb</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. B.</given-names>
            <surname>Aouicha</surname>
          </string-name>
          ,
          <article-title>Enhancing filter-based parenthetic abbreviation extraction methods</article-title>
          ,
          <source>Journal of the American Medical Informatics Association</source>
          <volume>28</volume>
          (
          <year>2021</year>
          )
          <fpage>668</fpage>
          -
          <lpage>669</lpage>
          . URL: https://doi.org/10.1093/jamia/ocaa314. doi:
          <volume>10</volume>
          .1093/jamia/ocaa314.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>O.</given-names>
            <surname>Pelgrin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Galárraga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Hose</surname>
          </string-name>
          ,
          <article-title>Towards fully-fledged archiving for RDF datasets</article-title>
          ,
          <source>Semantic Web</source>
          <volume>12</volume>
          (
          <year>2021</year>
          )
          <fpage>903</fpage>
          -
          <lpage>925</lpage>
          . doi:
          <volume>10</volume>
          .3233/sw-210434.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>H.</given-names>
            <surname>Turki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Hadj Taieb</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Ben Aouicha, Coupling wikipedia categories with wikidata statements for better semantics</article-title>
          ,
          <source>in: Proceedings of the 2nd Wikidata Workshop</source>
          a
          <article-title>a rp i r a a E f r : o L on r o Q n . a r e v i n o i c l it S /w an fo :/ s n p m io t</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>