<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Schema.org: How is It Used?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Minh-Hoang Dang</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alban Gaignard</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hala Skaf-Molli</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pascal Molli</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Nantes Université, CNRS, INSERM, l'institut du thorax</institution>
          ,
          <addr-line>F-44000 Nantes</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Nantes Université</institution>
          ,
          <addr-line>LS2N, Nantes</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Schema.org defines a shared vocabulary for semantically annotating web pages. Due to the vast and diverse nature of the contributed annotations, it is not easy to understand the widespread use of Schema.org. In this poster, we rely on the characteristic sets computed from the web data commons datasets to provide insights into property combinations on various websites. Thanks to in-depth experiments, this poster establishes a comprehensive observatory for schema.org annotations, visually presenting the most frequently used classes, commonly used combinations of properties per class, the average number of filled properties per class, and the classes with the greatest property coverage. These findings are valuable for both the communities involved in defining Schema.org vocabularies and the users of these vocabularies.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>(a) Cakes entities at four diferent websites
Entity name image description totalTime
Cake1 X X X
Cake2 X X
Cake3 X X</p>
      <p>
        Cake4 X X X
By relying on the Product class specification defined by Schema.org, it is impossible to know
which combination of properties is the most used on the web. As a webmaster, am I following
good practices? Finding the most used combinations of properties can help discover latent soft
schemas [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] in semantic annotations on the web.
      </p>
      <p>This poster establishes an observatory for schema.org annotations, providing comprehensive
insights into the commonly used combinations of class-specific properties and the quality of
class descriptions. This information is crucial for both communities specifying Schema.org
vocabularies and Schema.org profiles ( e.g. Bioschema.org) and for the users of these specifications.
The remainder of the paper is organized as follows: Section 2 presents our approach. Section 3
details our experimental study. The last section concludes the paper and points out perspective
works.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Approach</title>
      <p>
        We rely on characteristic sets [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] to build our observatory. Characteristic sets describe
semantically similar entities by grouping them according to the set of properties the
entities share. For an entity  of in an RDF dataset , the characteristic set is defined as:
 () = {|∃ : (, , ) ∈ }. To illustrate, in Table 1a, each row represents a diferent
website, and the X indicates the presence of a property in that particular website. Table 1b
presents the cardinality of each characteristic set. We consider a class well described if many
entities of that class share a combination of properties with many properties. We computed
characteristic sets (CSets) for the JSON-LD dataset (most used format) of WebDataCommons
(October 2021) 3. The CSets are available at (https://doi.org/10.5281/zenodo.8167689) and used
as a basis to answer diferent questions detailed in the next section. All results are available
at (https://schema-obs-demo.onrender.com).
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Experimental study</title>
      <p>The experimental study answers the following questions: i) Which classes are the most
commonly used? ii) What are the common combinations of properties per class? iii) Which classes
are accurately described?
3http://webdatacommons.org/structureddata/2021-12/stats/how_to_get_the_data.html (623GB)
(b) Top-10 most used Schema.org classes
125000
iittszcoeen17050000000
rse 50000
In 25000</p>
      <p>0
&lt;schema.org/clinicalPharmacology&gt;
&lt;schema.org/isProprietary&gt;
&lt;schema.org/alcoholWarning&gt;
&lt;schema.org/breastfeedingWarning&gt;
&lt;schema.org/mechanismOfAction&gt;
&lt;schema.org/mainEntityOfPage&gt;
&lt;schema.org/pregnancyWarning&gt;
&lt;schema.org/pregnancyCategory&gt;
&lt;schema.org/drugUnit&gt;
&lt;schema.org/legalStatus&gt;
&lt;schema.org/availableStrength&gt;</p>
      <p>&lt;schema.org/manufacturer&gt;
&lt;schema.org/prescriptionStatus&gt;
&lt;schema.org/dosageForm&gt;
&lt;schema.org/image&gt;</p>
      <p>&lt;schema.org/url&gt;
&lt;schema.org/description&gt;
&lt;schema.org/activeIngredient&gt;
&lt;schema.org/nonProprietaryName&gt;
&lt;schema.org/name&gt;
isa:&lt;schema.org/Drug&gt;</p>
      <p>
        Data corpus. We used the JSON-LD (most common formats) dataset from the
WebDataCommons [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] released in October 2021. This dataset is derived from crawling 35 million websites, of
which 42% utilized Web Entities. It comprises 82 billion RDF quads (16 terabytes uncompressed)
and 6.7 billion Schema.org entities.
      </p>
      <p>Distributed computing infrastructure. To analyze the schema.org dataset composed of 6.7B
web entities, we used an 8 nodes HPC cluster (8 CPU threads, 32 GB of RAM, 20 GB of local
storage per node). The code is written in Apache Spark. We computed in total 4, 638, 824 CSets,
which took around 30 hours.</p>
      <p>Most used classes. The sunburst diagram of Figure 2a presents the class hierarchy and
the corresponding instance count per class. Table 2b shows the most common classes, with
(a) Top-10 classes ranked by coverage
(b) Classes ranked by average properties
schema:Person ranking 8th with around 306 million entities. It also reveals that the classes are
not uniformly employed, indicating varying degrees of usage across the web.
Most commonly used combination of properties per class. Figures 3a and 3b illustrate,
respectively, the Upset plot for the top-10 characteristic sets of the classes Recipe and Drug.
The 10 columns represent the top-10 combination of properties, ordered by the number of
combined properties. In the case of Drugs, the vast majority of instances (top left histograms)
are annotated with only name and nonProprietaryName properties. The last column of this plot
shows very few instances annotated with the largest combination of properties. Conversely, in
the case of Recipes, starting from column 8, we observe that a large number of instances are
better annotated. More generally, the characteristic sets and their visual representation through
an upset plot provide interesting insights into the class’s latent soft schema and the poorly used
properties.</p>
      <p>
        Classes coverage. To compare class descriptions, we use the coverage metric defined in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. A
high coverage (near 1) indicates that class entities use most properties defined in the Schema.org
type specification. For readability, we computed the coverage for the properties contained in
the top-10 CSets per class. As shown in Table 4a, Recipe has a high coverage of 0.72, whereas
we computed a low coverage of 0.14 for Drug (104 defined properties), which confirms the
results shown in Figures 3a and 3b.
      </p>
      <p>Classes average properties (AvP). We computed the average number of used properties (AvP)
per class as indicated in Table 4b. On schema.org specification, a class definition has an AvP of
70.47, but on class instances, we observe an AvP of 5. This means that most of the properties
defined in the schema.org are not used when instantiating the classes. We observed that Recipe
obtained a good ranking with an AvP of 14.08 on 144 defined properties, whereas Product has
a lower AvP of 7.3 on 68 defined properties. This indicates that the cooking community may
better populate the available Recipe properties than the e-commerce community with the class
Product.</p>
      <p>More generally, we observed (i) no correlation between the rate of type usage and its coverage,
(ii) classes with the best coverage are those with fewer properties, (iii) the rankings with AvP
and coverage metrics return diferent top-10. The coverage metric seems to be biased towards
classes specified with few properties.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusion and future works</title>
      <p>We analyzed how webmasters efectively use Schema.org types and properties to annotate web
pages extracted from the WebDataCommons dataset. For each of the 776 analyzed types, we
computed the number of instances, the characteristic sets, the average number of properties
per type, and their coverage.</p>
      <p>Thanks to the characteristic sets, we could graphically display the instantiated schema with
Upset plots. Compared to the Schema.org specifications, we observed that very few properties
are efectively instantiated, and there is a great diversity in the combination of used properties.
The Upset plots allow webmasters and Schema.org maintainers to know which properties are
efectively and commonly used.</p>
      <p>In future works, we aim to define new quality metrics and leverage this Schema.org
observatory to study: (i) the per-class and per-web domain use of properties such as sameAs, (ii) the
temporal evolution of Schema.org by analyzing the yearly published WebDataCommons datasets
and (iii) the adoption of emerging community-specific profiles (e.g., Bioschemas) promoted in
the context of FAIR, reproducible, and open sciences.</p>
      <p>Acknowledgments
This work is supported by the French ANR project DeKaloG (Decentralized Knowledge Graphs),
ANR-19-CE23-0014, CE23 - Intelligence artificielle, and the French CominLabs project MikroLog
(The Microdata Knowledge Graph) grant no. 2019-05655. We are especially grateful to the
Nantes University master students Houda Bourefis, Julien Brochard, Farah Farouh, Hajar Lazrak
Senhaji, Mohammed Ali Ahmed Ragaa, Mohammed-Amine Bouzid, Yacine-Fadl Bestaoui, and
Oswaldo Andres Jimenez Hidalgo, who contributed to the early stages of this study.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>R. V.</given-names>
            <surname>Guha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Brickley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Macbeth</surname>
          </string-name>
          , Schema.org:
          <article-title>Evolution of structured data on the web</article-title>
          ,
          <source>Queue</source>
          <volume>13</volume>
          (
          <year>2015</year>
          )
          <fpage>10</fpage>
          -
          <lpage>37</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Brinkmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Primpeli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bizer</surname>
          </string-name>
          ,
          <article-title>The web data commons schema.org data set series</article-title>
          ,
          <source>in: Companion Proceedings of the ACM Web Conference</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>136</fpage>
          -
          <lpage>139</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>H.</given-names>
            <surname>Mühleisen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bizer</surname>
          </string-name>
          ,
          <article-title>Web data commons - extracting structured data from two large web corpora</article-title>
          ,
          <source>in: LDOW</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>R.</given-names>
            <surname>Meusel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Petrovski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bizer</surname>
          </string-name>
          ,
          <article-title>The webdatacommons microdata, rdfa and microformat dataset series</article-title>
          ,
          <source>in: The Semantic Web - ISWC 2014</source>
          , Springer International Publishing, Cham,
          <year>2014</year>
          , pp.
          <fpage>277</fpage>
          -
          <lpage>292</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>R.</given-names>
            <surname>Meusel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bizer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Paulheim</surname>
          </string-name>
          ,
          <article-title>A web-scale study of the adoption and evolution of the schema.org vocabulary over time</article-title>
          ,
          <source>in: 5th International Conference on Web Intelligence</source>
          , Mining, and
          <string-name>
            <surname>Semantics</surname>
          </string-name>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>T.</given-names>
            <surname>Neumann</surname>
          </string-name>
          , G. Moerkotte,
          <article-title>Characteristic sets: Accurate cardinality estimation for rdf queries with multiple joins</article-title>
          ,
          <source>27th International Conference on Data Engineering</source>
          (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>S.</given-names>
            <surname>Duan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kementsietsidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Srinivas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Udrea</surname>
          </string-name>
          ,
          <article-title>Apples and oranges: a comparison of RDF benchmarks and real RDF datasets</article-title>
          , in: SIGMOD,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>