<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Padua, Italy
$ benjamin.haettasch@cs.tu-darmstadt.de (B. Hättasch)
 https://benjaminhaettasch.de/ (B. Hättasch)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>WannaDB: Ad-hoc Structured Exploration of Text Collections Using Queries</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Benjamin Hättasch</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sunjet Aviation</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Hendrick Motorsports, Inc</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Southwest Airlines, Inc</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Technical University of Darmstadt (TU Darmstadt)</institution>
          ,
          <addr-line>Karolinenplatz 5, 64289 Darmstadt</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>airline Big Island Air Business Jet Services Ltd. city Aberdeen Volcano Burbank Stuart</p>
      </abstract>
      <kwd-group>
        <kwd>Textual collection containing information</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>date
Houston
information before they can compute an answer to their
query or to build extraction pipelines (e.g., when using
[1]) which however require substantial eforts.</p>
      <p>Hence, we advocate for a diferent route where users
can extract structured data relevant to satisfy an
information need from a collection of text documents without
the need to program, train or specify extraction systems.</p>
      <p>Instead, the aim is to provide a system that allows
users to explore new (unseen) text collections by simply
issuing a query to receive structured information from
the corpus. In contrast to [2] this should not require data
already in tabular form, rather the idea is to automatically
identify the relevant target structure and then, again
automatically, fill it from unstructured text.</p>
    </sec>
    <sec id="sec-2">
      <title>Contributions Therefore, we propose WannaDB, a</title>
      <p>system for ad-hoc structured text exploration. The main
Figure 1: Aim: Query a text collection and receive an approx- idea of WannaDB is that a user specifies their
informaimate structured result without manual extraction tion need by composing SQL-style queries over the text
collection. For example, in Figure 1, the user issues a
query to extract information about dates, airlines, and
Motivation In many domains, users face the problem cities of incidents. WannaDB then takes the query and
of needing to quickly extract insights from large col- evaluates it over the given document collection by
autolections of textual documents. For example, imagine a matically populating the table(s) required to answer the
journalist who wants to write an article about airline query with information nuggets from the documents.
security that was triggered by some recent incidents of To do this, WannaDB uses a novel pipeline as shown
a well-known US airline. For this reason, the journalist in Figure 2 which first extracts a superset of
informamight decide to explore a collection of textual accident tion nuggets from texts (e.g., all named entities), then
reports from the National Transportation Safety Board in
order to answer questions like ’What incident types are
the most frequent ones?’ or ’Which airlines are involved aCpopmropxuitmea&amp;tereretusurnlt
most often in incidents?’. And clearly, there are many
more domains where end users want to explore textual User
doYcuetm, eexnitsctionlglecatpiponrosaicnhaessitmoialanrs wfaesrhisounc.h queries over qisusueerys ApproKxBimate
new text collections force users to either read through SELECT...
vast amounts of text and manually extract the relevant</p>
      <p>Extract &amp;
organize
(additional)
relevant information
determines the information need from the query, and
ifnally matches nuggets to the relevant attributes of the
user’s query. As a result, WannaDB allows to answer the
queries, even if the information is not explicitly stated in
the corpus but has to be calculated (e.g., when the query
contains aggregation functions like AVG or SUM). A main
observation here is that in many cases a sample of
extractions (i.e., a table with partially missing or incorrect
values) is suficient to produce approximate results to
answer the user’s query.</p>
      <p>Acknowledgments</p>
    </sec>
    <sec id="sec-3">
      <title>This work has been supported by the German Federal</title>
      <p>Ministry of Education and Research as part of the Project
Software Campus 2.0 (TUDA), Microproject INTEXPLORE,
under grant ZN 01IS17050, by the German Research
Foundation as part of the Research Training Group Adaptive
Preparation of Information from Heterogeneous Sources
(AIPHES) under grant No. GRK 1994/1, as well as the
German Federal Ministry of Education and Research and the
State of Hesse through the National High-Performance
Computing Program.</p>
      <p>Architecture &amp; Initial Evaluation WannaDB
contains components to determine the information need
from queries, aggregate the relevant information, and References
compute the actual query result. For the extraction of
possibly relevant information nuggets, WannaDB relies on [1] C. De Sa, A. Ratner, C. Ré, J. Shin, F. Wang,
of-the-shelf extractors like Stanza [ 3]. The key contribu- S. Wu, C. Zhang, Deepdive: Declarative
knowltion of WannaDB, however, is a new matching approach edge base construction, SIGMOD Rec. 45 (2016)
that uses a novel embedding space exploration algorithm 60–67. URL: https://doi.org/10.1145/2949741.2949756.
incorporating interactive user feedback: The matching doi:10.1145/2949741.2949756.
process is done separately for each relevant attribute. [2] S. Zhang, K. Balog, Ad hoc table retrieval using
It starts by selecting information embeddings close to semantic similarity, in: Proceedings of the 2018
the attribute embedding. Afterwards, other embeddings World Wide Web Conference, WWW ’18,
Internathat might be matches are searched by applying several tional World Wide Web Conferences Steering
Comselection rules based on the closeness of embeddings to mittee, Republic and Canton of Geneva, CHE, 2018,
known matches. Each candidate is presented for feedback p. 1553–1562. URL: https://doi.org/10.1145/3178876.
(yes/no) to the user. The algorithm balances between ex- 3186067. doi:10.1145/3178876.3186067.
ploration and exploitation to select those information [3] P. Qi, Y. Zhang, Y. Zhang, J. Bolton, C. D.
Mannuggets for feedback that quickly allow identifying the ning, Stanza: A python natural language
processareas in the embedding space relevant for the attribute ing toolkit for many human languages, ArXiv
with as little feedback as possible. This area can then be abs/2003.07082 (2020).
used to populate the remaining rows in the target table. [4] B. Hättasch, J.-M. Bodensohn, C. Binnig, Aset:
AdPreviously extracted information (and user feedback) can hoc structured exploration of text collections, in: 3rd
be reused for follow-up queries. A detailed description International Workshop on Applied AI for Database
of the matching process can be found in [4]. Systems and Applications, Copenhagen, Denmark,</p>
      <p>Our experiments on diferent text-collections each fo- 2021. URL: https://sites.google.com/view/aidb2021/
cused around certain topics lead to promising results: home/accepted-papers.
10 − 25 quick iterations of feedback for each attribute
(i.e., confirming whether an information nuggets belongs
in a certain column of a table) suficed for high matching
scores for both textual and numeric attributes.</p>
      <p>The Road Ahead In the next steps, we want to
enlarge the scope, i.e., support more general corpora. We
also want to leverage certainties from the extraction and
matching process for the computation of the approximate
result and provide useful interfaces for the end users, not
only through a standalone application but also e.g., in
form of a Jupyter Notebook extension.</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>