<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Exploring Naive Bayes Classifiers for Tabular Data to Knowledge Graph Matching</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Brice Foko</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Azanzi Jiomekong</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hippolyte TAPAMO</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jérémy Buisson</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sanju Tiwari</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, University of Yaounde I</institution>
          ,
          <addr-line>Yaounde</addr-line>
          ,
          <country country="CM">Cameroon</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Universidad Autonoma de Tamaulipas</institution>
          ,
          <addr-line>Mexico</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The present research investigates the use of Naive Bayes classifiers to match knowledge graphs and tabular data, with particular emphasis on Column Type Annotation, Cell Entity Annotation, Column Property Annotation and Table Topic Detection. Using feature extraction techniques such as number of co-occurrences and term frequency, the study evaluates the efectiveness and performance of Naive Bayes classifiers on a variety of datasets. The proposed method is straightforward and generic, making a contribution to the field of knowledge graph matching and demonstrating the potential of Naive Bayes classifiers for the integration and interoperability of tabular data and knowledge graphs.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Tabular Data</kwd>
        <kwd>Knowledge Graph</kwd>
        <kwd>Tabular data to Knowledge Graph Matching</kwd>
        <kwd>Naive bayes</kwd>
        <kwd>TSOTSATable system</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The exponential growth of digital information has made both structured and unstructured
data increasingly prevalent. Among structured data, tabular datasets play a crucial role in
organizing and presenting information in a structured format across various domains such as
digital libraries [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], food science and nutrition [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], etc. On the other hand, Knowledge Graphs
(KGs) such as Wikidata1 and DBpedia2 provide comprehensive representations of real-world
entities and their interconnections. Therefore, matching tabular data with knowledge graphs
enables the enrichment of tabular datasets with semantic annotations and links to external
knowledge sources, resulting in enhanced data integration, interpretation, and interoperability.
However, achieving accurate and eficient matching poses significant challenges due to the
heterogeneity and complexity inherent in both tabular data and knowledge graphs.
      </p>
      <p>
        Naive Bayes classifiers are proven to be efective in various classification tasks due to their
simplicity, eficiency, and robustness [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. This study investigates the potential of Naive Bayes
classifiers in addressing four primary tasks proposed by the SemTab 2023 challenge: Column
Type Annotation (CTA), Cell Entity Annotation (CEA), Column Property Annotation (CPA),
and Table Topic Detection (TTD). The aim is to uncover new insights and practical techniques
to enhance the alignment and integration of tabular data with KGs. The source code used in
this work is available under open source license on GitHub3. We also provided a document4
demonstrating how to use the system proposed to solve the SemTab tasks.
      </p>
      <p>The rest of this paper is organized as follow: Section 2 presents some related work on table
annotations, Section 3 presents Naive Bayes classifier, Section 4 gives an overview of the research
methodology, Section 5 presents how we processed to solve the diferent tasks of the challenge,
Section 6 presents the results and finally, Section 7 conclude the paper.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related work</title>
      <p>
        The SemTab Challenge is an annual competition that evaluates table annotation systems. It
requires understanding the semantics of tabular data and knowledge graphs. Its previous
editions introduced three tasks: CTA Task, which assigns a semantic type from a KG to a
table column; CEA Task, which matches a cell of a given table to a KG entity; and CPA Task,
which assigns a KG property to the relationship between two columns. These tasks have been
addressed by diferent systems using various approaches, including:
• bbw (boosted by wiki) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. It uses Wikidata KG to annotate CSV tables using Meta-lookup
on a locally-deployed SearX metasearch engine and contextual matching. For contextual
matching, exact matching is used, followed by case-insensitive matching if no results are
found, and string matching with edit distance.
• MTAB [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. It handles CTA, CEA and CPA tasks well, using a probabilistic graph model.
      </p>
      <p>
        It improves matching by using multiple services like DBpedia Lookup, DBpedia
endpoint, Wikipedia, Wikidata, and a cross-lingual matching strategy, enhancing the overall
eficiency.
• DAGOBAH [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. It assumes that the entities are closed in the embedding space, and
employs an embedding strategy to cluster and score them in a column. For entity
clustering, it employs pre-trained Wikidata embeddings.
• TSOTSATable system [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Introduced by us in the previous challenge, it aims to match
tabular datasets to Knowledge Graphs or ontologies, specifically applying it to
TSOTSATable datasets [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. It proposes a KG refinement approach to address the matching problem
between tabular data and KGs, focusing on error correction and completing tabular data
with missing entities and relations.
      </p>
      <p>Despite the large number of annotation systems available, they are often complicated to
implement. This year, we focus on implementing a highly flexible machine learning method,
exploring Naive Bayes classification on table to KG matching problems.
3https://github.com/fokobrice3/STProbClass/tree/main/MNB_2023
4https://github.com/fokobrice3/STProbClass/blob/main/GUIDE.pdf</p>
    </sec>
    <sec id="sec-3">
      <title>3. Naive Bayes classifier</title>
      <p>A Naive Bayes classifier is a simple probabilistic classification method based on Bayes’
theorem, which calculates the probability of a specific class based on observed features. For any
occurrences of A and B, the Bayes’ theorem asserts the rule given by the equation 1.
 (|) =
 (|) *  ()
 ()
(1)
• P(A|B): this is the probability of event A given that event B has occurred.
• P(B|A): this is the likelihood of observing evidence B if the event A is true.
• P(A): this is the initial belief or knowledge about the probability of A before considering
any evidence.
• P(B): represents the overall probability of observing B, regardless of the occurrence of
event A.</p>
      <p>
        Bayes’ theorem is useful for inferring causes from their efects, as it simplifies determining
the likelihood of an efect based on its presence or absence [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. We used this theorem as
our background in multiclass-classification of tokens on labels, where the predicted class is
determined by multiplying the prior probability of each class with the conditional probabilities
of each feature.
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Research methodology</title>
      <p>
        This section describes the research methodology employed in this work. We developed the
research methodology by relying on what we know in empirical research in software engineering
[
        <xref ref-type="bibr" rid="ref10 ref9">9, 10</xref>
        ]. The methodology is adapted to the specific tasks and datasets at hand, including CTA,
CEA, CPA, and the recently introduced Table Topic Detection (TTD). In the following paragraphs,
we present the research question and the empirical research method used in this research.
      </p>
      <sec id="sec-4-1">
        <title>4.1. Research question</title>
        <p>The research question: "How to use Naive Bayes classifiers to match tabular datasets to
knowledge graph?" was used as the guideline of this work. To reply to this question, one should
provide a system that accepts a tabular dataset and a set of labels (classes or properties from a
knowledge graph) as inputs, and produces the annotated dataset with these labels. To this end,
the following questions should be replied:
• How can a column of tabular data be classified using a knowledge graph class? This task
is known as CTA and the fundamental query is "Which features must be used to classify
a column with a label?"
• How can data from a tabular data cell be classified using knowledge graph entities? The
CEA is presented here. The fundamental query is "Which features must we use to classify
a cell with a label?"
• How can a relationship between two columns of tabular data be classified using knowledge
graph property? This task is known as CPA. The fundamental query is "Which features
must we use to classify a relationship with a label?"
• How can knowledge graph class be used to classify the topic of tabular data ? This is the
TTD. The fundamental query is "Which features must we use to identify a topic with a
label?"</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Empirical methods</title>
        <p>
          The research methodology combines case study research, action research, and experimental
research, three empirical research techniques used in software engineering [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. This involves
investigating, testing, evaluating potential solutions, and proposing a solid solution that can be
applied to annotate any tabular data with a KG entity, class or property. Actually, the SemTab
organizers gave us three case studies to use in order to solve the tabular data to KG matching
including:
• Annotation of WikidataTables5 using Wikidata,
• Annotation of tFood6 using Wikidata [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ],
• Annotation of SOTAB7 using Schema.org and DBpedia.
        </p>
        <p>To enable the proposed solution to be applied in any situation, it is important to gain a deeper
understanding of the tabular data used in the knowledge graph matching problem through the
study of these case studies. We used the Scrum process, we ran, and improved our system using
the Sprint, iterative and incremental practices.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Solving the SemTab challenge</title>
      <p>The exploration of the use of Naive Bayes classifiers to solve the SemTab challenges tasks
allowed us to come up with a generic pipeline presented by Fig 1. This pipeline involves data
pre-processing, feature extraction, classification of known labels and prediction of new labels
components. In the following paragraphs, we present the diferent components of this pipeline
and the implementation of the system.</p>
      <sec id="sec-5-1">
        <title>5.1. Data Preprocessing</title>
        <p>The first step in our approach is data pre-processing, which prepares the raw datasets for
training and matching. For each dataset (test and train), the pre-processing steps consist of:</p>
        <sec id="sec-5-1-1">
          <title>1. Remove special characters,</title>
          <p>2. Set characters to lowercase,
3. Remove stopping word,</p>
        </sec>
        <sec id="sec-5-1-2">
          <title>5https://zenodo.org/record/7829583 6https://zenodo.org/record/7828163 7http://webdatacommons.org/structureddata/sotab/</title>
          <p>Once processed, the dataset contains cleaned data that can be used for feature extraction. The
training datasets were provided by the SemTab organizers and are presented in the table 1:</p>
        </sec>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Feature Extraction</title>
        <p>5.2.1. Column Type Annotation
For the CTA task, features are extracted from the table columns, including column headers when
taken into account, and column descriptions. During the training phase, these features will be
connected to the annotation (class/label) that was taken from the training data as presented by
Fig. 2.</p>
        <p>• Column Headers: names of the columns,
• Column Descriptions: a bag of words/tokens related to the column content.
5.2.2. Cell Entity Annotation
Concerning the CEA task, features are extracted from both the tabular data and the knowledge
graph to capture the relevant information for aligning each cell with the appropriate entity.
These features include cell contents, entity labels and contextual information that provide
additional context for matching as presented by Fig. 3.
5.2.3. Column Property Annotation
For the CPA task, features are extracted to capture the relationships or properties between two
columns in the tabular data. These features provide insights into the associations or connections
that exist between the columns. Examples of features for CPA (presented by Fig. 4) may include:
• Co-occurrence count: which is the number of times specific values or combinations of
values appear together in the two columns. This gives us information about the likelihood
or frequency of values appearing in both columns at the same time. Only the tokens with
a co-occurrence &gt;= 1 are taken into account.
• Relationship: relationships are obtained by linking pairs of tokens together. This
consists of combining tokens in the diferent cells of the target columns in the training
dataset for the CPA task.
5.2.4. Table Topic Detection
Concerning the TTD task, features related to the overall content and structure of the table are
extracted. They help determine the primary topic or subject matter of the table. Examples of
features for TTD (presented by Fig. 5) may include:
• Table Headers: the names of the columns provide valuable information about the
table’s content. These headers can be extracted as textual features that contribute to the
classification of the table topic when they are available.
• Key Terms: relevant tokens from the table content can serve as features for the TTD
task. These terms are extracted using Term Frequency (TF) algorithms 8.</p>
        <sec id="sec-5-2-1">
          <title>8https://www.capitalone.com/tech/machine-learning/understanding-tf-idf/</title>
        </sec>
      </sec>
      <sec id="sec-5-3">
        <title>5.3. Implementation of Naive Bayes Classifiers</title>
        <p>This Section presents how we implemented the Naives Bayes classification for each task using
the label from the KG inside the train datasets and the extracted features.
5.3.1. Goal
The objective is to build the learning function  ()−
&gt; 
•  is one of the classes/labels in the training dataset(e.g Number, Boolean, Food, Person,</p>
        <p>Hotel, P18, P651, Q625, Q31, Q54389 etc.).
•  = (1, 2, ..., ) is the representation of the data that comes in a bag of
tokens/words.</p>
        <p>Given that we have a set of classes/labels 1, 2, ..., , and we want to determine the
probability of an instance  to belong to each class/label. The Bayesian probability of a class/label is
obtained using the equation 2.</p>
        <p>(|) =
 (|) *  ()
 ()
(2)
•  (|): this is the conditional probability to be calculated.
•  (|): this is the likelihood of the features of  under the assumption that it
belongs to class/label ,  (|) =  (1|) *  (2|) * ... *  (|)
•  (): represents the initial belief or knowledge about the probability of an instance
belonging to class/label  before considering any evidence.
•  (): represents the overall probability of observing instance , regardless of its
class/label.</p>
        <p>To classify an instance , we compute  (|) for each class  and select the class with the
highest probability. The implementation is tailored to the specific requirements of each dataset.
The following paragraphs give an overview.
5.3.2. WikidataTables Dataset
This dataset consists of tables in CSV format. The CTA, CEA, and CPA targets have to be
classified using Wikidata’s classes and properties. Three Naive Bayes classifiers were trained
for each task, using the labeled dataset on train. The labels, provided by this dataset, represent
the ground truth annotations obtained from the Wikidata KG. The entity, semantic type, and
relationship predictions for new, unseen tabular data are then predicted using the three trained
classifiers.
5.3.3. tFood Dataset
The tFood dataset consists of tables in CSV format. There, the Wikidata class and properties
have to be used to classify CTA, CEA, CPA, and TTD targets. Using the same process on
WikidataTables dataset, four Naive Bayes classifiers are trained using the labeled dataset of the
tFood training data. The first column of a table is ignored as it lacks relevant information for
the TTD task (e.g., prop0, prop1, prop2, etc.). Since no corpus is provided for each table, we also
adjust the TF-IDF to term frequency. Due to time-consuming training, a limit of 50 cells from
each CSV file are randomly extracted per table during training.
5.3.4. SOTAB Dataset
The SOTAB dataset tables are provided in GZ-compressed JSON files. The task was to classify
CTA and CPA with schema.org and DBpedia classes. Using the labeled dataset of the SOTAB
training data, we trained six classifiers, including two in round 1 and four in round 2. We also
limited the number of data elements that could be retrieved from each JSON file to 20 by file at
random for each task because many JSON files contained a lot of data, which made the training
too time-consuming.
5.3.5. Development environment
The development environment was composed of VSCode as code editor, Node.js as the JavaScript
runtime, and npm as the package manager. For feature extraction and data preprocessing, we
implement the diferent utilities from scratch and we consider the Natural 9 npm package for
Porter Stemming and building classifiers. We used a desktop with a Ryzen 1700 8 core processor
and 16Gb RAM. We also consider multi-instance activity as performance optimization techniques
to enhance the eficiency of the implementation.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Results</title>
      <p>The SemTab 2023 challenge consisted of two (02) rounds lasting from April 14 to June 22, 2023.
The pipeline presented by Fig. 1 was applied to each dataset. The following paragraphs present
the results provided by the SemTab organizers for the diferent datasets and a short discussion.</p>
      <sec id="sec-6-1">
        <title>9https://naturalnode.github.io/natural</title>
        <sec id="sec-6-1-1">
          <title>6.1. WikidataTables Dataset</title>
          <p>WikidataTables dataset was provided in Round 1 of the challenge. The test data are presented
in Table 2 and the challenge’s CTA, CEA and CPA results on this dataset are shown in Fig. 6.</p>
        </sec>
        <sec id="sec-6-1-2">
          <title>6.2. tFood Dataset</title>
          <p>The tFood dataset was provided in Round 1 of the challenge. This dataset was divided into two
main parts: tFood-horizontal and tFood-entity. The test data are presented in Table 3.</p>
          <p>The tFood datasets revealed challenges in training and inference time, as well as the lack
of a corpus to accurately extract the key terms for TTD task. In addition, we discovered that
the table’s redundancies and unclear data were not appropriate to the approach presented in
this paper. Unfortunately, to reduce training time issues, we took a subset in each table with a
maximum of 50 rows per table. Typically, 6-8 hours were spent on training and 9-10 hours for
inference per task. Fig. 7 presents the results of the challenge on these datasets.</p>
        </sec>
        <sec id="sec-6-1-3">
          <title>6.3. SOTAB Dataset</title>
          <p>The SOTAB dataset was provided in Rounds 1 and 2. An overview of the test data is presented
in Table 4.</p>
          <p>Due to the time required for training and inference, and the limited performance of our
training environment, we just submitted to round 2 with this dataset.</p>
          <p>The SOTAB dataset has a larger number of tables and rows in JSON format compared to
other datasets. To reduce training and inference time, we significantly decreased the number
of training tables , and the rows were reduced to 20 per table. We also point out that the
redundancy and incorrect data in the tables were not helpful for the proposed approach. Each
task took 10-12 hours for training and 8-9 hours for inference.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusion</title>
      <p>This paper outlines the method we propose for annotating tabular data with knowledge graph
classes, entities, and properties using Naives Bayes multiclass classifiers. Data pre-processing
aims to ensure the accuracy and integrity of the tabular data as well as to make the subsequent
matching process easier. Feature extraction techniques such as the number of co-occurrences
and the frequency of terms provide useful information to capture the semantic relationships
between tabular data and knowledge graphs. Due to redundant and misleading data in the
training dataset, the computation time, the approach was severely limited but we are exploring
further solutions to improve our feature extraction and computation times.</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgment</title>
      <p>We are grateful to SemTab organizers for having given us the opportunity to share this work
with the community.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Oelen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Stocker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Auer</surname>
          </string-name>
          ,
          <article-title>Creating a scholarly knowledge graph from survey article tables</article-title>
          ,
          <source>in: Digital Libraries at Times of Massive Societal Transition</source>
          , Springer International Publishing,
          <year>2020</year>
          , pp.
          <fpage>373</fpage>
          -
          <lpage>389</lpage>
          . URL: https://doi.org/10.1007%
          <fpage>2F978</fpage>
          -
          <fpage>3</fpage>
          -
          <fpage>030</fpage>
          -64452-9_
          <fpage>35</fpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -64452-9_
          <fpage>35</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Jiomekong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Etoga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Foko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Tsague</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Folefac</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kana</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. M. Sow</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Camara</surname>
          </string-name>
          ,
          <article-title>A large scale corpus of food composition tables</article-title>
          ,
          <year>2022</year>
          , pp.
          <fpage>34</fpage>
          -
          <lpage>36</lpage>
          . URL: https://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>3320</volume>
          /paper4.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Ting</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Ip</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Tsang</surname>
          </string-name>
          ,
          <article-title>Is naïve bayes a good classifier for document classification?</article-title>
          ,
          <source>International Journal of Software Engineering and its Applications</source>
          <volume>5</volume>
          (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>R.</given-names>
            <surname>Shigapov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Zumstein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kamlah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Oberländer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mechnich</surname>
          </string-name>
          ,
          <string-name>
            <surname>I.</surname>
          </string-name>
          <article-title>Schumm, bbw: Matching csv to wikidata via meta-lookup</article-title>
          , in: SemTab@ISWC,
          <year>2020</year>
          . URL: https: //api.semanticscholar.org/CorpusID:229242235.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>P.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          , I. Yamada,
          <string-name>
            <given-names>N.</given-names>
            <surname>Kertkeidkachorn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ichise</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Takeda</surname>
          </string-name>
          ,
          <year>Semtab 2021</year>
          :
          <article-title>Tabular data annotation with mtab tool</article-title>
          , in: SemTab@ISWC,
          <year>2021</year>
          . URL: https://api.semanticscholar. org/CorpusID:247363605.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>V.-P.</given-names>
            <surname>Huynh</surname>
          </string-name>
          , J. Liu,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chabot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Deuzé</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Labbé</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Monnin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Troncy</surname>
          </string-name>
          ,
          <article-title>Dagobah: Table and graph contexts for eficient semantic annotation of tabular data</article-title>
          , in: SemTab@ISWC,
          <year>2021</year>
          . URL: https://api.semanticscholar.org/CorpusID:247363666.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Jiomekong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Foko</surname>
          </string-name>
          ,
          <article-title>Towards an approach based on knowledge graph refinement for tabular data to knowledge graph matching</article-title>
          ,
          <year>2022</year>
          , pp.
          <fpage>111</fpage>
          -
          <lpage>122</lpage>
          . URL: https://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>3320</volume>
          /paper12.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>B.</given-names>
            <surname>Alsafy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Mosad</surname>
          </string-name>
          , W. Mutlag,
          <source>Multiclass classification methods: A review</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Jiomekong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Tapamo</surname>
          </string-name>
          , G. Camara,
          <article-title>Combining Scrum and Model Driven Architecture for the development of the EPICAM platform</article-title>
          ,
          <source>in: CARI</source>
          <year>2022</year>
          , Yaounde, Cameroon,
          <year>2022</year>
          . URL: https://hal.archives-ouvertes.fr/hal-03712484.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J.</given-names>
            <surname>Azanzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Tapamo</surname>
          </string-name>
          , G. Camara,
          <article-title>Combining Scrum and Model Driven Architecture for the development of an epidemiological surveillance software</article-title>
          ,
          <source>Revue Africaine de Recherche en Informatique et Mathématiques Appliquées</source>
          Volume
          <volume>39</volume>
          -
          <fpage>2023</fpage>
          (
          <year>2023</year>
          ). URL: https://arima.episciences.org/11537. doi:
          <volume>10</volume>
          .46298/arima.9873.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>P.</given-names>
            <surname>Ralph</surname>
          </string-name>
          , et al.,
          <source>Empirical standards for software engineering research</source>
          ,
          <year>2021</year>
          . URL: https: //arxiv.org/abs/
          <year>2010</year>
          .03525. arXiv:
          <year>2010</year>
          .03525.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>N.</given-names>
            <surname>Abdelmageed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Jimènez-Ruiz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Hassanzadeh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>König-Ries</surname>
          </string-name>
          ,
          <source>tFood: Semantic Table Annotations Benchmark for Food Domain</source>
          ,
          <year>2023</year>
          . URL: https://doi.org/10.5281/zenodo. 7828163. doi:
          <volume>10</volume>
          .5281/zenodo.7828163.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>