<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>BiodivTab: A Table Annotation Benchmark based on Biodiversity Research Data</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Friedrich Schiller University Jena</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Heinz Nixdorf Chair for Distributed Information Systems</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Michael Stifel Center Jena</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <abstract>
        <p>Semantic Table Annotation (STA) denotes the process of annotating a tabular dataset with concepts and relations from a given Knowledge Graph. The objective is to map individual table elements to their counterparts from the Knowledge Graph. The Semantic Web Challenge on Tabular Data to Knowledge Graph Matching (SemTab) aims to establish a common framework for systems that tackle the process of STA. Since 2019, it has provided a set of benchmarks each year for evaluation. However, most of the provided datasets in the rst two incarnations of the challenge are Automatically Generated (AG) and general domain datasets. This leaves the question open whether the developed systems can similarly be applied to real-world datasets that provide a di erent set of challenges. In this paper, we try to address this gap by introducing a domain-speci c benchmark named BiodivTab. It consists of 50 datasets based on real-world biodiversity research data that have further been augmented. BiodivTab was made available to SemTab participants during Round 3 in the 2021 edition of the challenge.</p>
      </abstract>
      <kwd-group>
        <kwd>Cell Entity Annotation</kwd>
        <kwd>Column Type Annotation</kwd>
        <kwd>Table Annotation</kwd>
        <kwd>Benchmark</kwd>
        <kwd>Biodiversity</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The Semantic Web Challenge on Tabular Data to Knowledge Graph Matching (SemTab)
has worked on establishing a community for Semantic Table Annotation (STA) tasks
over the course of so far three editions: 2019 [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], 2020 [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], and 20214. The challenge
formulated three tasks for the STA that are illustrated by Figure 1. Each task matches
a table component to its counterpart within a target Knowledge Graph (KG):
{ Cell Entity Annotation (CEA) matches individual cells to entities.
{ Column Type Annotation (CTA) assigns a semantic column type.
{ Column Property Annotation (CPA) links column pairs using a semantic property.
? Copyright 2021 for this paper by its authors. Use permitted under Creative
Commons License Attribution 4.0 International (CC BY 4.0).
4 https://www.cs.ox.ac.uk/isg/challenges/sem-tab/2021/index.html
      </p>
      <p>Country</p>
      <p>Area</p>
      <p>Capital</p>
      <p>Country</p>
      <p>Area</p>
      <p>Capital</p>
      <p>Country</p>
      <p>Area</p>
      <p>Capital
Egypt</p>
      <p>1,010,408
Germany
357,386</p>
      <p>Cairo
Berlin</p>
      <p>Egypt
1,010,408</p>
      <p>Cairo</p>
      <p>Egypt
1,010,408</p>
      <p>Cairo
Germany
357,386</p>
      <p>Berlin</p>
      <p>Germany
357,386</p>
      <p>Berlin
https://www.wikidata.org/wiki/Q79
https://www.wikidata.org/wiki/Q183 https://www.wikidata.org/wiki/Q6256 https://www.wikidata.org/wiki/Q5119
(a) CEA
(b) CTA
(c) CPA</p>
      <p>
        The challenge establishes common standards for systems that tackle the problem of
STA. Among the best-performing participants from the 2020 are MTab4Wikidata [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ],
LinkingPark [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], bbw [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], DAGOBAH [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], and JenTab [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        The ultimate goal is systems that can annotate real-world datasets. However, the
datasets introduced in the rst two years of the challenge are Automatically
Generated (AG) derived from di erent KGs [
        <xref ref-type="bibr" rid="ref10 ref9">9, 10</xref>
        ]. The ToughTables Dataset (2T) of 2020
is manually curated and focuses on the disambiguation of possible solutions [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. The
datasets employed, so far, adhere to no particular domain but represented a sample
from a wide range of general-purpose data. On the other hand, domain-speci c datasets
pose speci c challenges as witnessed, e.g., by evaluation campaigns in other domains
like semantic web services evaluations [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. So, ensuring that those challenges are
covered, there is a demand for other domain-speci c datasets based on real-world data.
Furthermore, these benchmarks have to comply with the standards already in use by
the community to easily highlight current shortcomings and encourage further e orts
on these challenges.
      </p>
      <p>
        In this paper, we introduce a domain-speci c tabular benchmark named BiodivTab.
We have collected real tables from the biodiversity domain and manually annotated
them using the live edition of Wikidata [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] during September 2021, resulting in a
human-level generated ground truth data for CEA and CTA tasks. Inspired by the
challenges witnessed in the domain, we introduced arti cially created variations to
increase the number of included tables.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Benchmark Description</title>
      <p>In this section, we explain the creation steps of BiodivTab, the data sources we have
used, and the biodiversity-speci c challenges encountered. Moreover, we describe the
annotation phase, the data augmentation step, and the nal assembly of the
benchmark.</p>
      <p>
        The included datasets were collected from three public repositories of biodiversity
data: data.world5, BEFChina6 [
        <xref ref-type="bibr" rid="ref16 ref20 ref22 ref5">5, 16, 20, 22</xref>
        ], and BExIS7 [4, 8, 13{15, 18]. We gathered
      </p>
      <sec id="sec-2-1">
        <title>5 https://data.world/ 6 https://data.botanik.uni-halle.de/bef-china/ 7 https://www.bexis.uni-jena.de/</title>
        <p>a large collection of biodiversity-related datasets and veri ed their licenses8 to ensure
that they allow for the use in such a benchmark. Subsequently, we manually checked
each of them concerning their suitability to the semantic table annotation tasks. In the
process, we discarded datasets that predominantly contained, e.g., internal database
\ID" columns, generic headers (e.g., \BEX 12"), or numerical columns without any
further explanation or context. We consider those datasets next to impossible to
annotate automatically and of little bene t to the community. Consequently, we decided
to include only datasets containing a substantial amount of categorical information.</p>
        <p>The datasets collected this way feature unique characteristics that can be
summarized as follows:
{ Specimen Data: The collected datasets contain observations of a particular
specimen, e.g., a speci c individual of a given species. This includes a multitude of
properties of the specimen and their particular environment. The assembled data
can only rarely be attributed to the general species.
{ Numerical Data: Most of the collected datasets describe the specimen by various
measurements in numerical form.
{ Abbreviations: Species names may be given in an abbreviated format. For example,
\Canna glauca", a particular kind of ower, is often referred to as \C.glauca" or
\Ca.glauce".
{ Special Format : Species names are identi ed by a combination of species and
subspecies. For example, \species:Atrichum sub:subserratum" may be used instead of
\Atrichum subserratum".</p>
        <p>The identi ed challenges increase the di culty during the semantic annotation
process. Commonly, data is characterized by a single subject column (usually the leftmost
one) with a group of other columns representing properties to the respective subject.
Given that most of the data represent specimen data, those numerical/object elds do
not necessarily relate to the general properties of species. As of the time of writing,
Wikidata, the target KG of BiodivTab, contains no direct equivalent to the specimen
data contained in the selected datasets. Thus, we could not annotate column-relations
in the fashion of a CPA-task. As a consequence, BiodivTab does not include a
CPAtask as of now. However, it might change in the future, if other KGs are supported, or
Wikidata is extended accordingly. Instead, we provide ground truths only for CEA and
CTA tasks. The special format found in the species names impedes the direct matching
of cell values to labels of individual entities in the KG. One possible approach might be
to handle these cases with a particular variety of misspellings. In fact, part of the cell
values might be considered additional noise that has to be removed before matching.</p>
        <p>
          After the data collection phase, we picked 13 tables to use in the annotation
process. Our naming convention follows the schema of \dataSource id", e.g., \befchina 1"
represents the rst dataset collected from BEFChina. The annotation itself is the most
time-consuming part of the benchmark creation. To ensure the quality of mappings, we
manually annotated the selected tables with entities assembled from Wikidata during
September 2021, resulting in ground truth data for both CEA and CTA tasks.
However, another set of annotations and checking the Inter-annotator Agreement [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] would
be needed. Concerning CEA, we have marked possible candidate columns to annotate
their cells. For each cell value, we assembled possible matches via Wikidata's built-in
search. If more than one match has been found, we manually selected the most suitable
8 https://github.com/fusion-jena/BiodivTab/blob/main/Read_Datasets_
        </p>
        <p>Licenses.md
ones to disambiguate the cell's semantics. If this still leaves more than one candidate,
we keep them all and consider them true matches. Consequently, the provided ground
truth contains all possible candidates that di erent systems could generate. We
followed the same procedure for CTA. We maintain separate ground truth les to ease
manual inspection, revision, and quality assurance for each table. E.g., \befchina 1"
is annotated by two such les: \befchina 1 CEA" and \befchina 1 CTA". Biodiversity
experts have partially revised these annotations. The structure of the ground truth les
follows the format of SemTab. In particular, the solution les for CEA use a format
of lename, column id, row id, and ground truth, whereas the ones for CTA employ a
structure of lename, column id, and ground truth.</p>
        <p>To increase the number of tables in our benchmark and reduce the human e ort
needed, we further resorted to data augmentation. It is a technique to increase the
amount of data by adding slightly modi ed copies of already existing entries. In our
context, we introduced challenges to the existing dataset based on our ndings during
the data collection and analysis phase (see the rst paragraph). Since abbreviations
are a common issue in biodiversity datasets, a column containing full species names
was thus abbreviated accordingly. This strategy allowed us to (i) increase the number
of included tables to 50 (almost 4 the number of just real tables), (ii) reduce the
required human e ort during annotation, and (iii) produce a benchmark that relies on
real challenges instead of arti cial ones.</p>
        <p>The benchmark dataset consists of a set with 13 real tables and 37 augmented
ones. We anonymized the le names of tables to use unique identi ers using Python's
uuid functionalities in the process. Subsequently, we aggregated the individual
solutions of CEA and CTA into one le per task resulting in CEA biodivtab 2021 gt.csv
and CTA biodivtab 2021 gt.csv, respectively. From this, we generated corresponding
\target- les" by removing the ground truth columns from these solution les. The
benchmark tables and targets can now be published during the challenge for
participants to solve. This follows the general approach of SemTab that hides the ground
truth of STA tasks from participants during the challenge.</p>
        <p>In 2021, the organizers published a call for domain-speci c benchmarks to be used
in that year's challenge. BiodivTab was submitted and accepted as one of these
benchmarks. As a result, BiodivTab represented one of three benchmarks posed during the
third round of 2021's SemTab challenge.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Conclusions &amp; Future Work</title>
      <p>
        We have introduced a tabular benchmark derived from biodiversity research data
named BiodivTab. It consists of a collection of 50 tables. We have created BiodivTab
by manually annotating 13 tables from real-world biodiversity datasets and adding 37
more tables by augmenting them with noise based on previously observed challenges.
BiodivTab was submitted to and subsequently used in Round 3 of the 2021 SemTab
challenge. Our benchmark is publicly available [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]9.
      </p>
      <p>Future Work We see multiple directions to continue this work. We plan to include
more biodiversity tables from other projects to cover a broader spectrum of the domain.
In addition, ground truth data from other KGs, in particular domain-speci c ones, can
be provided.</p>
      <sec id="sec-3-1">
        <title>9 https://github.com/fusion-jena/BiodivTab</title>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgment</title>
      <p>The authors thank the Carl Zeiss Foundation for the nancial support of the project
\A Virtual Werkstatt for Digitization in the Sciences (P5)" within the scope of the
program line \Breakthroughs: Exploring Intelligent Systems" for \Digitization - explore
the basics, use applications". We would like to especially thank our Biodiversity experts
Cornelia Furstenau and Andreas Ostrowski for feedback and validation of the created
annotations. Last but not least, we would like to thank Samira Babalou for the fruitful
discussions during the work.</p>
      <p>The tables provided in this challenge are based on real-world biodiversity research
datasets, but have been adapted for the challenge. In the form provided here, they may
be used for the challenge, only. Any publication on challenge results needs to contain
citations of the underlying datasets. The list of original datasets is available within our
GitHub repository.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Abdelmageed</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schindler</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>JenTab: Matching Tabular Data to Knowledge Graphs</article-title>
          . In: SemTab@ ISWC. pp.
          <volume>40</volume>
          {
          <issue>49</issue>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Abdelmageed</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schindler</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Knig-Ries</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          : fusion-jena/BiodivTab (Oct
          <year>2021</year>
          ). https://doi.org/10.5281/zenodo.5584180
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Artstein</surname>
          </string-name>
          , R.:
          <source>Inter-annotator Agreement</source>
          , pp.
          <volume>297</volume>
          {
          <fpage>313</fpage>
          . Springer Netherlands (
          <year>2017</year>
          ), https://doi.org/10.1007/
          <fpage>978</fpage>
          -94-024-0881-2_
          <fpage>11</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Boeddinghaus</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marhan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berner</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boch</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fischer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kattge</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klaus</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kleinebecker</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oelmann</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prati</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , Schafer,
          <string-name>
            <given-names>D.</given-names>
            , Schoning, I.,
            <surname>Schrumpf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Sorkau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Kandeler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Manning</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Kandeler</surname>
          </string-name>
          ,
          <string-name>
            <surname>E.</surname>
          </string-name>
          :
          <article-title>Plant functional trait shifts explain concurrent changes in the structure and function of grassland soil microbial communities (</article-title>
          <year>2017</year>
          ). https://doi.org/10.25829/bexis.24867-
          <issue>1</issue>
          .1.
          <fpage>23</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Bruelheide</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eichenberg</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , Krober, W., Bohnke,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Ristok</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          : Main Experiment:
          <article-title>Leaf traits and chemicals from individual trees in the Main Experiment (Site A &amp; B</article-title>
          ) (
          <year>2012</year>
          ), https://china.befdata.biow.uni-leipzig.de/datasets/323
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Karaoglu</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Negreanu</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ma</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yao</surname>
            ,
            <given-names>J.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gordon</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lin</surname>
          </string-name>
          , C.Y.:
          <article-title>LinkingPark: An Integrated Approach for Semantic Table Interpretation</article-title>
          . In: SemTab@ ISWC. pp.
          <volume>65</volume>
          {
          <issue>74</issue>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Cutrona</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bianchi</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jimnez-Ruiz</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Palmonari</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          : Tough Tables:
          <article-title>Carefully Evaluating Entity Linking for Tabular Data (Nov</article-title>
          <year>2020</year>
          ). https://doi.org/10.5281/zenodo.4246370
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Fischer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nauss</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tschapka</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weisser</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          , Muller, J.:
          <article-title>Aggregated species richness and habitat heterogeneity variables for testing the habitat-heterogeneity hypothesis,</article-title>
          <year>2006</year>
          -
          <fpage>2018</fpage>
          (
          <year>2020</year>
          ). https://doi.org/10.25829/bexis.25126-
          <fpage>1</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Hassanzadeh</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Efthymiou</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jimnez-Ruiz</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Srinivas</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>SemTab2019: Semantic Web Challenge on Tabular Data to Knowledge Graph Matching -</article-title>
          2019
          <source>Data Sets (Oct</source>
          <year>2019</year>
          ). https://doi.org/10.5281/zenodo.3518539
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Hassanzadeh</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Efthymiou</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jimnez-Ruiz</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Srinivas</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>SemTab 2020: Semantic Web Challenge on Tabular Data to Knowledge Graph Matching Data Sets (Nov</article-title>
          <year>2020</year>
          ). https://doi.org/10.5281/zenodo.4282879
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Huynh</surname>
            ,
            <given-names>V.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chabot</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Labbe</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Monnin</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Troncy</surname>
          </string-name>
          , R.: DAGOBAH:
          <article-title>Enhanced Scoring Algorithms for Scalable Annotations of Tabular Data</article-title>
          . In: SemTab@ ISWC. pp.
          <volume>27</volume>
          {
          <issue>39</issue>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12. Kuster, U.,
          <article-title>Konig-</article-title>
          <string-name>
            <surname>Ries</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Towards standard test collections for the empirical evaluation of semantic web service approaches</article-title>
          .
          <source>International Journal of Semantic Computing</source>
          <volume>2</volume>
          (
          <issue>03</issue>
          ),
          <volume>381</volume>
          {
          <fpage>402</fpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Leonhardt</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peters</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , Keller, A.:
          <article-title>Amino acids in pollen of Osmia bicornis larval provisions 2017-2018 (</article-title>
          <year>2020</year>
          ). https://doi.org/10.25829/bexis.27228-
          <fpage>4</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Leonhardt</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peters</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , Keller, A.:
          <article-title>Fatty acids in pollen of Osmia bicornis larval provisions 2017-2018 (</article-title>
          <year>2020</year>
          ). https://doi.org/10.25829/bexis.27227-
          <fpage>2</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Leonhardt</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peters</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , Keller, A.:
          <article-title>Trap nesting solitary bee species measured on all grassland VIPs 2017-</article-title>
          <year>2018</year>
          (
          <year>2020</year>
          ). https://doi.org/10.25829/bexis.27226-
          <fpage>4</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Nadrowski</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Deviations from stem breaking probabilities at species level (</article-title>
          <year>2013</year>
          ), http://china.befdata.biow.uni-leipzig.de/datasets/327
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Nguyen</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yamada</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kertkeidkachorn</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ichise</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Takeda</surname>
          </string-name>
          , H.: MTab4Wikidata at SemTab 2020:
          <article-title>Tabular Data Annotation with Wikidata</article-title>
          .
          <source>In: SemTab@ ISWC</source>
          . pp.
          <volume>86</volume>
          {
          <issue>95</issue>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Seibold</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , Gosner,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Simons</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            , Bluthgen, N., Muller, J.,
            <surname>Ambarli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Ammer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Bauhus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Fischer</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          , Furstenau,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Habel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.C.</given-names>
            ,
            <surname>Linsenmair</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.E.</given-names>
            ,
            <surname>Nauss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Ostrowski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Penone</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Prati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Schall</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Schulze</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.D.</given-names>
            ,
            <surname>Vogt</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          , Wollauer,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Weisser</surname>
          </string-name>
          ,
          <string-name>
            <surname>W.</surname>
          </string-name>
          :
          <article-title>Arthropod data from 150 grassland plots</article-title>
          ,
          <year>2008</year>
          -
          <fpage>2017</fpage>
          , and 140 forest plots, 2008
          <article-title>-2016, used in "Arthropod decline in grasslands and forests is associated with drivers at landscape level"</article-title>
          ,
          <source>Nature</source>
          (
          <year>2019</year>
          ). https://doi.org/10.25829/bexis.25786-
          <issue>1</issue>
          .3.
          <fpage>11</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Shigapov</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zumstein</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kamlah</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , Oberlander, L.,
          <string-name>
            <surname>Mechnich</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schumm</surname>
          </string-name>
          , I.:
          <article-title>bbw: Matching csv to wikidata via meta-lookup</article-title>
          .
          <source>In: CEUR Workshop Proceedings</source>
          . vol.
          <volume>2775</volume>
          , pp.
          <volume>17</volume>
          {
          <fpage>26</fpage>
          .
          <string-name>
            <surname>RWTH</surname>
          </string-name>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Staab</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schuldt</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Assmann</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bruelheide</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klein</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Ant community structure during forest succession in a subtropical forest in South-East China pp</article-title>
          .
          <volume>32</volume>
          {
          <issue>40</issue>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Vrandecic</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , Krotzsch, M.:
          <article-title>Wikidata: a free collaborative knowledgebase</article-title>
          .
          <source>Communications of the ACM</source>
          <volume>57</volume>
          (
          <issue>10</issue>
          ),
          <volume>78</volume>
          {85 (sep
          <year>2014</year>
          ). https://doi.org/10.1145/2629489
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Wubet</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buscot</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Soil Fungal metagenome from 12 CSPs based on the fungal ITS rDNA pyrotags (</article-title>
          <year>2013</year>
          ), http://china.befdata.biow.uni-leipzig. de/datasets/397
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>