<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Graphs and Language Models to Answer Questions over Tables</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Judith Knoblach</string-name>
          <email>judith.knoblach@iais.fraunhofer.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nikhil Acharya</string-name>
          <email>nikhil.acharya@iais.fraunhofer.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bhavya Koranemkattil</string-name>
          <email>bhavya.koranemkattil@iais.fraunhofer.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andreas Both</string-name>
          <email>andreas.both@datev.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Diego Collarana</string-name>
          <email>diegocollarana@upb.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Knowledge Graphs, Transformers, Language Models</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>DATEV eG</institution>
          ,
          <addr-line>Nuremberg</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Fraunhofer Institute for Intelligent Analysis and Information Systems</institution>
          ,
          <addr-line>Dresden</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Leipzig University of Applied Sciences</institution>
          ,
          <addr-line>Leipzig</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Universidad Privada Boliviana</institution>
          ,
          <addr-line>Cochabamba</addr-line>
          ,
          <country country="BO">Bolivia</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2022</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>Tables remain a primary modality for organizing and presenting information to people. We interact every day with Excel sheets, CSV files, tables in PDF documents, and web tables. Providing a natural language interface to query table information is paramount for several use cases. This demo shows a solution to query semantically described tables using natural-language questions. Our solution employs knowledge graphs as a medium to integrate tables coming from heterogeneous sources. Then, a transformer-based language model analyzes a user's question and finds the answer in the semantically represented tables. During the demo session, we will show a use case developed in collaboration with DATEV eG, where tax consultants can eficiently query information from financial tables. Attendees will experience how a natural-language interface speeds up the information retrieval process from tables. They will also be allowed to ask their questions to a prepared dataset, showing the scalability of our solution. The video demo is available at https://owncloud.fraunhofer.de/index.php/s/uXFmUfzCta70rqN.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Companies still use tables as the main modality to present information to employees. For
example, DATEV eG is a German company mainly providing large-scale business software (e.g.,
accounting). These services are widely used by tax consultants, lawyers, auditors, small and
medium-sized enterprises, municipalities, start-ups, and many more. More than two million
German companies use financial accounting programs from DATEV, interacting with hundreds
of tables daily. To continue this success story and remain competitive in the market, DATEV
relies on employees who are experts in their field and on state-of-the-art software solutions to
accelerate internal processes even in environments that become increasingly data-dependent.
DATEV uses AI solutions that facilitate the development and relieve the employees’ workload.
One of the most relevant aspects is data management and information retrieval. DATEV
employees handle a wide range of information on diferent domains, e.g., taxes in European
countries, specific software versions for wire transfers, or (strict) deadlines for storing personal,
sensitive data. Often the data is presented as tables, either as web tables or CSV files coming
from internal, external, or oficial data providers. This data-intensive environment makes it
challenging for users to find information eficiently.</p>
      <p>In the scope of the SPEAKER1 project, Fraunhofer IAIS and DATEV eG have teamed up to
provide an application to query tables with natural language, i.e., Question Answering (QA)
over tables. Our approach allows non-technical users to express what they want in the table
more naturally in text form. Moreover, our solution integrates tables represented in various
formats, e.g., CSV, Web Tables, or even tables encoded in XML, as in DATEV’s use case. Figure 1
depicts the structure of the overall application. It consists of two proprietary components –
the Fraunhofer Smart Data Connector (SDC) and the Fraunhofer QA Component. The SDC
is a solution to create knowledge graphs by transforming heterogeneous enterprise data into
actionable knowledge. The QA component allows answering questions expressed in natural
language over the knowledge graph. The following section presents an overview of our approach
that combines the SDC and QA components to provide a solution to the problem of answering
questions over tables. The last section describes the demonstration of the use case in detail.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Architecture</title>
      <sec id="sec-2-1">
        <title>2.1. Smart Data Connector</title>
        <p>
          The SDC is a generic component to transform and store enterprise data sources into a knowledge
graph. Figure 2a depicts a general overview of the SDC components and their interactions.
Following a mapping approach [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], the SDC uses mapping rules to transform tables into RDF
triples following the CSVW vocabulary. In DATEV’s use case, we define a mapping rule to
transform tables encoded in XML files to RDF. However, our component is generic enough to
handle other scenarios, e.g., transforming CSV tables into RDF. At runtime, the SDC Engine
1cf. https://www.speaker.fraunhofer.de/en
(a) Components of the Smart Data Connector
(b) The CSVW vocabulary
provides a service to upload XML files. Then, the Mapping Engine takes the file and the mapping
rules and creates the semantic representation of the table using CSVW vocabulary. Figure 2b
shows the main concepts from CSVW that we are using in this application. Once the tables
are transformed into RDF, the SDC ofers diferent services to the QA component. The SDC
Entity-Relation Index provides a search service for entities and relations (including their labels,
types, and synonyms), and it can bed used for Entity Linking tasks by the QA component.
The SDC Engine provides a SPARQL endpoint used by the QA component to query the tables.
Finally, the SDC Embeddings Generator ofers advanced services, e.g., entity similarity based
on embeddings. The SDC Embeddings Generator uses PyKEEN [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] to generate embeddings of
the entities and relations of the graph with diferent models.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. QA component</title>
        <p>
          Our QA component is inspired by the TaPas model [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. In our approach, the model is extended
to answer questions from tables semantically described in knowledge graphs. TaPas is a deep
learning model based on BERT’s encoder architecture [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] and is specifically designed for question
answering over tabular data. The TaPas model is built through two stages, pre-training, and
ifne-tuning. The pre-training has been done over millions of tables and related text segments
crawled from Wikipedia and this is a crucial reason behind the performance of the model [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
        </p>
        <p>
          The fine-tuning process of TaPas has been done in a supervised fashion in multiple public
datasets such as WIKISQL, WIKITQ, and SQA. Figure 3 depicts the architecture of the TaPas
model. In addition to the BERT’s encoder embeddings [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], the table data and structure are
also encoded as inputs for the TaPas model. A table is flattened to a string format where
the column headers and cells are concatenated as string tokens. Then question tokens are
appended to the sequence. TaPas, trained in the weakly supervised setting, achieves close to
state-of-the-art performance for WIKISQL (accuracy: 83.6). For SQA, TaPas leads to substantial
improvements on all metrics: improving all metrics by at least 11 points. For WIKITQ [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] the
model trained only from the original training data reaches an accuracy of 42.6, which surpasses
similar approaches [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Demonstration</title>
      <p>This demo shows the two-fold approach of our solution:</p>
      <p>Building a knowledge graph from domain- and format-independent tables. The
biggest challenge for the DATEV’s use case is to map tables from diferent domains and formats
without the additional efort of creating a complex ontology, keeping the mapping rules to a
minimum. We select to reuse the CSV on the Web2 (CSVW) vocabulary for having a generic
semantic representation of the tables. Figure 2b shows the main concepts we use, i.e., “Table”
and connects it with the concepts “Row”, “Cell” and “Column”. In the demo video, XML files
from DATEV (document storage period, VAT rate in European countries, and payment initiation
versions for the SEPA formats) are mapped to the CSVW with the rule “transform-tables”. Using
the query service in the SDC dashboard, the structure of the populated knowledge graph can be
explored in more detail. The knowledge graph contains triples that connect the table URI with
the corresponding columns and rows. There are also triples that link the content of a cell to the
expected row via the row URI.</p>
      <p>
        Answering questions over knowledge graph. The table URI is the key element for
connecting with the “QA component”. The user sends both the URI of the table and a question in
natural language, e.g., “What is the intermediate VAT rate in Belgium?”. Thanks to the use of a
shared vocabulary for all tables, i.e., CSVW, the QA component can send a standard SPARQL to
retrieve the table’s columns and rows based on the URI. Then the Cells and Columns instances
are pre-processed as embeddings used to represent the table in TaPas, i.e., column embeddings
which indicate the column the token belongs to, row embedding, which indicates the row the
token belongs to, rank embeddings which indicate the rank of a particular cell according to the
column it belongs. Thus, our QA component can combine tables stored in knowledge graphs
and the TaPas model with this information. It is also possible to answer questions requiring
cell aggregation, e.g., “How many European countries have a standard VAT rate of 20 percent?”
(COUNT), “What is the average standard VAT rate?” (AVERAGE), “Can you tell me the total
years I need to keep all invoices?” (SUM). As a weakly supervised deep learning model, TaPas
can represent the relationships between columns and values in tables and has an excellent
semantic understanding of natural language queries. We can deploy TaPas for multiple use
cases in multiple domains since the model has the ability for cross-domain [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusions</title>
      <p>The application to the DATEV use case is just one of many possible applications of our
question answering over tables solution. This demo emphasizes the combination of RDF
knowledge graphs with Language Models to solve the problem of answering questions
over heterogeneous tables. In future work, we will explore improving the TaPas model by
taking advantage of the semantically described tables. Moreover, we want to extend the
TaPas model to the German language. Also, we plan to add more aggregation operators,
like MINIMUM and MAXIMUM, into the TaPas architecture. Finally, the QA component
can be improved to automatically identify the right table and answer without explicitly
specifying the table URI. This way, it would be possible to answer questions from a pool of tables.
Acknowledgements: We acknowledge the support of the EU H2020 Projects Opertus Mundi
(GA 870228), and the Federal Ministry for Economic Afairs and Energy (BMWi) project
SPEAKER (FKZ 01MK20011A).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Dimou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. V.</given-names>
            <surname>Sande</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Colpaert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Verborgh</surname>
          </string-name>
          , E. Mannens, R. V. de Walle,
          <article-title>RML: A generic language for integrated RDF mappings of heterogeneous data</article-title>
          ,
          <source>in: WWW</source>
          , Seoul, Korea, volume
          <volume>1184</volume>
          <source>of CEUR Workshop Proceedings</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Ali</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Berrendorf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. T.</given-names>
            <surname>Hoyt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Vermue</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sharifzadeh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Tresp</surname>
          </string-name>
          , J. Lehmann, PyKEEN
          <volume>1</volume>
          .
          <article-title>0: A Python library for training and evaluating knowledge graph embeddings</article-title>
          ,
          <source>J. Mach. Learn. Res</source>
          .
          <volume>22</volume>
          (
          <year>2021</year>
          )
          <volume>82</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>82</lpage>
          :
          <fpage>6</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Herzig</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. K.</given-names>
            <surname>Nowak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Piccinno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Eisenschlos</surname>
          </string-name>
          ,
          <article-title>TaPas: Weakly supervised table parsing via pre-training</article-title>
          , in: ACL, Online,
          <year>2020</year>
          , pp.
          <fpage>4320</fpage>
          -
          <lpage>4333</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          ,
          <article-title>BERT: pre-training of deep bidirectional transformers for language understanding</article-title>
          ,
          <source>in: NAACL-HLT</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Lin</surname>
          </string-name>
          , J.-g. Lou,
          <article-title>TaPEx: Table pre-training via learning a neural SQL executor</article-title>
          ,
          <source>arXiv preprint arXiv:2107.07653</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Neelakantan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q. V.</given-names>
            <surname>Le</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Abadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>McCallum</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Amodei</surname>
          </string-name>
          ,
          <article-title>Learning a natural language interface with neural programmer</article-title>
          , in: ICLR, Toulon, France,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>