<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Pythagoras: Semantic Type Detection of Numerical Data Using Graph Neural Networks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sven Langenecker</string-name>
          <email>sven.langenecker@mosbach.dhbw.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christoph Sturm</string-name>
          <email>christoph.sturm@mosbach.dhbw.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christian Schalles</string-name>
          <email>christian.schalles@mosbach.dhbw.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Carsten Binnig</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Semantic Type Detection, Data Discovery in Data Lakes, Tabular data, Work in Progress</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>DFKI Darmstadt</institution>
          ,
          <addr-line>Hochschulstrasse 10, 64289 Darmstadt</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>DHBW Mosbach</institution>
          ,
          <addr-line>Lohrtalweg 10, 74821 Mosbach</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>LWDA'23: Lernen</institution>
          ,
          <addr-line>Wissen, Daten, Analysen</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>TU Darmstadt</institution>
          ,
          <addr-line>Karolinenplatz 5, 64289 Darmstadt</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Detecting semantic types of table columns is a crucial task to enable dataset discovery in data lakes. However, prior semantic type detection approaches have primarily focused on non-numeric data despite the fact that numeric data play an essential role in many enterprise data lakes. Therefore, typically, existing models are rather inadequate when applied to data lakes that contain a high proportion of numerical data. In this paper, we introduce Pythagoras, our new learned semantic type detection approach specially designed to support numerical data along with non-numerical data. Pythagoras uses a graph neural network based on a new graph representation of tables to predict the semantic types for numerical data with high accuracy. In our initial experiments, we thus achieve F1-Scores of 0.829 (support-weighted) and 0.790 (macro), respectively, exceeding the state-of-the-art performance significantly.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        Dataset discovery of numerical data is important in enterprise data lakes. Enterprise
data lakes serve as invaluable repositories of diverse data types, enabling organizations to
store and manage vast amounts of information [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. In these data lakes, numerical data plays a
dominant role, making up a much larger proportion compared to non-numerical data [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and
providing insights into various business domains, including finance, manufacturing, healthcare,
and marketing. Such data often contain critical information such as sales figures, production
metrics, customer demographics, and financial records. Therefore, it is essential to automatically
detect the correct semantic type of table columns with numerical data enabling data scientists
to find required data for downstream analysis and thus address the dataset discovery problem
in data lakes [
        <xref ref-type="bibr" rid="ref3 ref4 ref5">3, 4, 5</xref>
        ].
      </p>
      <p>
        Existing approaches are mainly designed for non-numerical data. In order to provide
the task of semantic type detection, many solutions using deep learning techniques have been
proposed in the past [
        <xref ref-type="bibr" rid="ref10 ref6 ref7 ref8 ref9">6, 7, 8, 9, 10</xref>
        ]. Unfortunately, all these existing approaches have primarily
focused on detecting the semantic type of non-numerical data table columns, leaving a critical
nEvelop-O
CEUR
Workshop
Proceedings
Serializations
      </p>
      <p>Tablename Columns
[CLS] tablename [SEP] [CLS] Val 1 ... Val n [SEP]</p>
      <sec id="sec-2-1">
        <title>BERT</title>
        <p>Feature Specific
Subnetwork</p>
        <p>Input Features
(192 Units)</p>
        <p>ReLu
(512 Units)
s
de iton
iiltIanoN trsenepeaRTabNleondaeme ColTuemxtuNaoldes
GNN
Numerical
Colum Nodes
Numerical Values
Statistic Node</p>
      </sec>
      <sec id="sec-2-2">
        <title>Graph Convolution Layer</title>
      </sec>
      <sec id="sec-2-3">
        <title>ReLU</title>
      </sec>
      <sec id="sec-2-4">
        <title>Graph Convolution Layer</title>
        <p>Textual Colum Nodes
Numerical Colum Nodes</p>
      </sec>
      <sec id="sec-2-5">
        <title>Final Classification Layer</title>
        <p>(b) Model architecture
Content</p>
        <p>CNNToaaalmubmmleeesn PLJNelaaaBmymreeoersn PSoFBFisea/itPlsdioFknetbaPlloGPi3nal1atms.y3eeprerStAatsisGsit7asic.mt5sseper Rpeebr8oG.u2anmdes</p>
        <p>TMuyrnleesr PF/C 15.4 2.1 9.8
CTeoxlutumanl NCuomluemricnal</p>
        <p>
          Graph
Representation
need for innovative approaches that efectively handle the detection of semantic types for
numerical table columns [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ].
        </p>
        <p>
          Towards a new learned semantic type detection model for numerical data. In this paper,
we introduce our new vision of a semantic type detection model called Pythagoras, which can
not only predict the semantic type of non-numerical table columns with high accuracy but
also of numerical table columns. To achieve this, the main idea of the new model architecture
is to use graph neural networks (GNNs) together with a novel graph representation of tables
and their columns. This graph representation includes directed edges to provide necessary
context information (e.g. neighboring non-numerical columns) for predicting the semantic type
of numerical columns using GNNs message passing mechanism. The graph representation and
the new model architecture are the main contributions of this paper. Moreover, as a second
contribution, we show initial highly promising results comparing Pythagoras against five
existing state-of-the-art models on the SportsTables corpus [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. The results of this experiment
demonstrate that we outperform all existing semantic type detection models on numerical data.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>2. Overview of Pythagoras</title>
      <p>In the following, we will introduce our new semantic type detection model Pythagoras and
discuss the main design aspects that will lead to better predictions on numerical table columns.</p>
      <p>
        Figure 1a demonstrates how we convert a table and its columns into a graph representation
using an example. We can see that the table is transformed into a graph containing four
diferent node types. The green node represents the table name. The orange and blue nodes are
responsible for the representation of the textual and numerical columns. In addition, there is
another node in the graph for each numerical column, which contains 192 selected statistical
features (see Table 2 in the Appendix) of the numerical column values (red node). Because
detecting semantic types of numerical columns is generally harder than for textual columns,
using only the numerical column values to specify the type is too limited [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Hence, we
designed the graph structure with directed edges to inject necessary context information into
the numerical column representation and thus enrich it for better predictions. Looking at a
Numerical Column Node we can see that three directed edge types go towards the node. With
that, the node will embed information from its connected neighbors into its own representation
during a GNN layer iteration based on the message passing paradigm [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Specifically, the
green edge provides information about the table name, the yellow edges convey information
from each textual column within the table and the red edge facilitates the transmission of the
additional statistical features. As a consequence, the GNN layers transforms the representation
of the Numerical Column Nodes, leading to enhanced information content for accurate semantic
type prediction. For instance, when faced with a numerical column with values in the range of
60-100, where the semantic type could be ambiguous (e.g., basketball.player.weight or humidity),
the embedding of information from a neighboring textual column containing basketball player
names allows for a more precise identification of the semantic type as basketball.player.weight.
      </p>
      <p>
        In Figure 1b we can see the whole model architecture of Pythagoras which encodes the table
structure as a graph. The upper part of the architecture illustrates how we generate the initial
embedding vector representations of the nodes in the graph. For encoding table names as
well as cell values (textual and numerical), we use the pre-trained transformer-based language
model BERT1[
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. In addition, to embed the features of the Numerical Values Statistic Nodes, we
train a feature specific subnetwork similar to the approach in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. This subnetwork embeds the
extracted features from the numerical column to an output of fixed length using one hidden
layer with a rectifier linear unit (ReLu) activation function. For each numerical column, we
extract 192 features2, including for example mean and median of all values, as well as statistical
metrics regarding the occurrence of individual digits. The initial node representations, along
with the discussed graph structure, serve as the input for the GNN model. As GNN, we use
a graph convolutional neural network [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. After traversing the GNN layers, we extract the
hidden states of Textual as well as Numerical Column Nodes from the last convolutional layer.
These hidden states are then passed as inputs to a final classification layer to perform the
semantic type classification task.
      </p>
    </sec>
    <sec id="sec-4">
      <title>3. Initial Experimental Results</title>
      <p>
        In this section, we present initial experimental results applying Pythagoras on the SportsTables
corpus. We compare the performance of our approach against five state-of-the-art models.
Baseline models. As state-of-the-art models we consider Sherlock [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], Sato [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], Dosolo [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ],
Doduo [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and GPT-3 [
        <xref ref-type="bibr" rid="ref13 ref14">13, 14</xref>
        ]. While Sherlock and Dosolo are models that utilize only the values
of a single column for the prediction, Sato and Doduo are successors of them that adopt a
1Note that Pythagoras is independent of how to generate these initial embeddings, and there may exist alternative
language models or embedding methods that could potentially yield even better results in this context.
2The complete feature list can be found in Table 2 in the Appendix
context-based approach similar to our model. Despite their similarities to our model, Sato and
Doduo do not specifically address the prediction of semantic types for numerical-based columns
and do not ofer a well-defined approach for injecting contextual information into the prediction
process. Furthermore, to have another benchmark, we developed in our experiments a
finetuned GPT-3 model for the task of semantic type detection. We chose fine-tuning over prompt
designs for higher model quality and the capacity to train on a larger number of examples3.
Experiment setup. In our experiments, we use the SportsTables dataset, due to its high
proportion of numerical-based columns. To perform the experiments, we split the corpus into
60/20/20 for train, validation, and test set. After training the models, the checkpoint with the
best accuracy on the validation set is used for evaluation on the test set. We report end results
as an average of five runs with diferent random seeds using the evaluation metrics
supportweighted and macro F1-Score as in previous studies [
        <xref ref-type="bibr" rid="ref15 ref6 ref7 ref8">6, 7, 8, 15</xref>
        ]. To implement Pythagoras we
used Python together with PyTorch [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], DGL [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] and the Transformers library [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ].
Results of study. The experimental results are shown in Table 1. For each model, we list the
F1-Scores overall data types to show the total performance, but also the separate average
F1Scores for only numerical and non-numerical data types, respectively. The results show that our
model Pythagoras outperforms all existing models in detecting the semantic type of numerical
columns. To the best performing existing model Sato with F1-Scores of 0.703/0.650
(supportweighted/macro F1-Score) we can achieve an improvement of +0.126/+0.140. Furthermore,
a notable observation across all existing models is the substantial performance discrepancy
between the prediction on non-numerical and numerical columns. On non-numerical data, the
accuracy of the models is generally high, whereas their performance on numerical data tends
to be poorer. In contrast, our model exhibits a diferent behavior, as we are able to achieve a
more balanced accuracy for both numerical and non-numerical data. In summary, the results
demonstrate that our model, in conjunction with the graph representation of tables, leads to a
significantly improved performance.
      </p>
      <p>The road ahead. To establish the generalizability of our approach, additional experiments on
diverse datasets are crucial. These experiments will validate the efectiveness of our approach
across diferent data domains and assess its robustness. Furthermore, conducting an ablation
study is essential to examine the impact of various design choices in our architecture.
3For fine-tuning the GPT-3 model, we use OpenAIs API described in https://platform.openai.com/docs/guides/
fine-tuning (visited on 09/04/2023)
Acknowledgements. This research and development project was funded by DHBW Mosbach.
We also want to thank the NHR Program, the BMBF project KompAKI (grant number 02L19C150),
the HMWK cluster project 3AI, hessian.AI, and DFKI Darmstadt for their support.
#
150</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J.</given-names>
            <surname>Dixon</surname>
          </string-name>
          , Data Lakes Revisited, https://jamesdixon.wordpress.com/
          <year>2014</year>
          /09/25/ data-lakes-revisited/,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Langenecker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Sturm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Schalles</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Binnig</surname>
          </string-name>
          ,
          <article-title>Sportstables: A new corpus for semantic type detection</article-title>
          , in: B.
          <string-name>
            <surname>König-Ries</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Scherzinger</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          <string-name>
            <surname>Lehner</surname>
          </string-name>
          , G. Vossen (Eds.),
          <source>Datenbanksysteme für Business</source>
          ,
          <source>Technologie und Web (BTW</source>
          <year>2023</year>
          ),
          <fpage>20</fpage>
          .
          <article-title>Fachtagung des GI-Fachbereichs „Datenbanken und Informationssysteme” (DBIS</article-title>
          ),
          <volume>06</volume>
          .-
          <fpage>10</fpage>
          ,
          <year>März 2023</year>
          , Dresden, Germany, Proceedings, volume P-331
          <string-name>
            <surname>of</surname>
            <given-names>LNI</given-names>
          </string-name>
          , Gesellschaft für Informatik e.V.,
          <year>2023</year>
          , pp.
          <fpage>995</fpage>
          -
          <lpage>1008</lpage>
          . URL: https://doi.org/10.18420/BTW2023-68. doi:
          <volume>10</volume>
          .18420/BTW2023- 68.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>G.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <article-title>Table discovery in data lakes: State-of-the-art and future directions</article-title>
          ,
          <source>in: Companion of the 2023 International Conference on Management of Data, SIGMOD '23</source>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2023</year>
          , p.
          <fpage>69</fpage>
          -
          <lpage>75</lpage>
          . URL: https://doi.org/10.1145/3555041.3589409. doi:
          <volume>10</volume>
          .1145/3555041.3589409.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>F.</given-names>
            <surname>Nargesian</surname>
          </string-name>
          , E. Zhu,
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. Q.</given-names>
            <surname>Pu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. C.</given-names>
            <surname>Arocena</surname>
          </string-name>
          ,
          <article-title>Data lake management: Challenges and opportunities</article-title>
          ,
          <source>Proc. VLDB Endow</source>
          .
          <volume>12</volume>
          (
          <year>2019</year>
          )
          <fpage>1986</fpage>
          -
          <lpage>1989</lpage>
          . URL: https://doi.org/10.14778/ 3352063.3352116. doi:
          <volume>10</volume>
          .14778/3352063.3352116.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Khatiwada</surname>
          </string-name>
          , G. Fan,
          <string-name>
            <given-names>R.</given-names>
            <surname>Shraga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Gatterbauer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Riedewald</surname>
          </string-name>
          , Santos:
          <article-title>Relationship-based semantic table union search</article-title>
          ,
          <source>Proc. ACM Manag. Data</source>
          <volume>1</volume>
          (
          <year>2023</year>
          ). URL: https://doi.org/10.1145/3588689. doi:
          <volume>10</volume>
          .1145/3588689.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hulsebos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bakker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Zgraggen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Satyanarayan</surname>
          </string-name>
          , T. Kraska, c. Demiralp,
          <string-name>
            <given-names>C.</given-names>
            <surname>Hidalgo</surname>
          </string-name>
          ,
          <article-title>Sherlock: A deep learning approach to semantic data type detection</article-title>
          ,
          <source>in: SIGKDD, KDD '19</source>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA,
          <year>2019</year>
          , p.
          <fpage>1500</fpage>
          -
          <lpage>1508</lpage>
          . URL: https://doi.org/10. 1145/3292500.3330993. doi:
          <volume>10</volume>
          .1145/3292500.3330993.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hulsebos</surname>
          </string-name>
          , Y. Suhara, c. Demiralp,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.-C.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <article-title>Sato: Contextual semantic type detection in tables</article-title>
          ,
          <source>in: VLDB</source>
          , volume
          <volume>13</volume>
          ,
          <string-name>
            <given-names>VLDB</given-names>
            <surname>Endowment</surname>
          </string-name>
          ,
          <year>2020</year>
          , p.
          <fpage>1835</fpage>
          -
          <lpage>1848</lpage>
          . URL: https://doi.org/10.14778/3407790.3407793. doi:
          <volume>10</volume>
          .14778/3407790.3407793.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Suhara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , c. Demiralp,
          <string-name>
            <given-names>C.</given-names>
            <surname>Chen</surname>
          </string-name>
          , W.-C. Tan,
          <article-title>Annotating columns with pre-trained language models</article-title>
          , in: SIGMOD,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA,
          <year>2022</year>
          , pp.
          <fpage>1493</fpage>
          -
          <lpage>1503</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>X.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lees</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>Yu, TURL: Table Understanding through Representation Learning</article-title>
          , in: VLDB, volume
          <volume>14</volume>
          ,
          <string-name>
            <given-names>VLDB</given-names>
            <surname>Endowment</surname>
          </string-name>
          ,
          <year>2021</year>
          , pp.
          <fpage>307</fpage>
          -
          <lpage>319</lpage>
          . URL: https://github. com/sunlab-osu/TURL. doi:
          <volume>10</volume>
          .14778/3430915.3430921. arXiv:
          <year>2006</year>
          .14806v2.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S.</given-names>
            <surname>Langenecker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Sturm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Schalles</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Binnig</surname>
          </string-name>
          ,
          <article-title>Steered training data generation for learned semantic type detection</article-title>
          ,
          <source>Proc. ACM Manag. Data</source>
          <volume>1</volume>
          (
          <year>2023</year>
          )
          <volume>201</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>201</lpage>
          :
          <fpage>25</fpage>
          . URL: https://doi.org/10.1145/3589786. doi:
          <volume>10</volume>
          .1145/3589786.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>T. N.</given-names>
            <surname>Kipf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Welling</surname>
          </string-name>
          ,
          <article-title>Semi-supervised classification with graph convolutional networks</article-title>
          ,
          <source>in: 5th International Conference on Learning Representations, ICLR</source>
          <year>2017</year>
          , Toulon, France,
          <source>April 24-26</source>
          ,
          <year>2017</year>
          , Conference Track Proceedings, OpenReview.net,
          <year>2017</year>
          . URL: https: //openreview.net/forum?id=
          <fpage>SJU4ayYgl</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , BERT:
          <article-title>Pre-training of Deep Bidirectional Transformers for Language Understanding, in: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Association for Computational Linguistics</article-title>
          , Minneapolis, Minnesota,
          <year>2019</year>
          , pp.
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          . URL: https://aclanthology.org/ N19-1423. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>N19</fpage>
          - 1423.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>T. B. Brown</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Mann</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Ryder</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Subbiah</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Kaplan</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Dhariwal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Neelakantan</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Shyam</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Sastry</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Askell</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Agarwal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Herbert-Voss</surname>
            , G. Krueger,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Henighan</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Child</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Ramesh</surname>
            ,
            <given-names>D. M.</given-names>
          </string-name>
          <string-name>
            <surname>Ziegler</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Winter</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Hesse</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Chen</surname>
            , E. Sigler,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Litwin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Gray</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Chess</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Berner</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>McCandlish</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Radford</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Amodei</surname>
          </string-name>
          ,
          <article-title>Language models are few-shot learners</article-title>
          ,
          <source>in: Proceedings of the 34th International Conference on Neural Information Processing Systems</source>
          , NIPS'20, Curran Associates Inc.,
          <string-name>
            <surname>Red</surname>
            <given-names>Hook</given-names>
          </string-name>
          ,
          <string-name>
            <surname>NY</surname>
          </string-name>
          , USA,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>L.</given-names>
            <surname>Ouyang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Almeida</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. L.</given-names>
            <surname>Wainwright</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mishkin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Agarwal,
          <string-name>
            <given-names>K.</given-names>
            <surname>Slama</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ray</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Schulman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hilton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Kelton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. E.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Simens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Askell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Welinder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. F.</given-names>
            <surname>Christiano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Leike</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Lowe</surname>
          </string-name>
          ,
          <article-title>Training language models to follow instructions with human feedback</article-title>
          ,
          <source>ArXiv abs/2203</source>
          .02155 (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>X.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lees</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <article-title>Turl: Table understanding through representation learning</article-title>
          ,
          <source>Proc. VLDB Endow</source>
          .
          <volume>14</volume>
          (
          <year>2020</year>
          )
          <fpage>307</fpage>
          -
          <lpage>319</lpage>
          . URL: https://doi.org/10.14778/3430915. 3430921. doi:
          <volume>10</volume>
          .14778/3430915.3430921.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>A.</given-names>
            <surname>Paszke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gross</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Massa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lerer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bradbury</surname>
          </string-name>
          , G. Chanan,
          <string-name>
            <given-names>T.</given-names>
            <surname>Killeen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Gimelshein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Antiga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Desmaison</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Köpf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>DeVito</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Raison</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Tejani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Chilamkurthy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Steiner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Fang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Chintala</surname>
          </string-name>
          ,
          <string-name>
            <surname>Pytorch:</surname>
          </string-name>
          <article-title>An imperative style, high-performance deep learning library</article-title>
          ., in: H.
          <string-name>
            <surname>M. Wallach</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Larochelle</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Beygelzimer</surname>
          </string-name>
          , F.
          <string-name>
            <surname>d'Alché Buc</surname>
            ,
            <given-names>E. B.</given-names>
          </string-name>
          <string-name>
            <surname>Fox</surname>
          </string-name>
          , R. Garnett (Eds.), NeurIPS,
          <year>2019</year>
          , pp.
          <fpage>8024</fpage>
          -
          <lpage>8035</lpage>
          . URL: http://dblp.uni-trier.de/db/conf/nips/nips2019.html#PaszkeGMLBCKLGA19.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>M.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Ye</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Gan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Ma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Gai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Xiao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Karypis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <surname>Z. Zhang,</surname>
          </string-name>
          <article-title>Deep graph library: A graph-centric, highly-performant package for graph neural networks</article-title>
          , arXiv preprint arXiv:
          <year>1909</year>
          .
          <volume>01315</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Jia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Fu</surname>
          </string-name>
          , S. Han,
          <string-name>
            <surname>D</surname>
          </string-name>
          . Zhang, Tuta:
          <article-title>Tree-based transformers for generally structured table pre-training</article-title>
          ,
          <source>Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery &amp; Data Mining</source>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>