<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Discovery from Scientific Documents</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Vijay S. Kumar</string-name>
          <email>v.kumar1@ge.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Varish Mulwad</string-name>
          <email>varish.mulwad@ge.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jenny Weisenberg Williams</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tim Finin</string-name>
          <email>finin@umbc.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sharad Dixit</string-name>
          <email>sharad.dixit@ge.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anupam Joshi</string-name>
          <email>joshi@umbc.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Bases (VLDBW'23) - TaDA'23: Tabular Data Analysis Workshop</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>GE Research</institution>
          ,
          <addr-line>1 Research Circle, Niskayuna, NY</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>GE Research, John F. Welch Technology Center</institution>
          ,
          <addr-line>Whitefield, Bengaluru</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Joint Workshops at 49th International Conference on Very Large Data</institution>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>University of Maryland Baltimore County</institution>
          ,
          <addr-line>1000 Hilltop Circle, Baltimore, MD</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>Workshop Proce dings</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Synthesizing information from collections of tables embedded within scientific and technical documents is increasingly critical to emerging knowledge-driven applications. Given their structural heterogeneity, highly domain-specific content, and difuse context, inferring a precise semantic understanding of such tables is traditionally better accomplished through linking tabular content to concepts and entities in reference knowledge graphs. However, existing tabular data discovery systems are not designed to adequately exploit these explicit, human-interpretable semantic linkages. Moreover, given the prevalence of misinformation, the level of confidence in the reliability of tabular information has become an important, often overlooked, factor in discovery over open datasets. We describe a preliminary implementation of a discovery engine that enables table-based semantic search and retrieval of tabular information from a linked knowledge graph of scientific tables. We discuss the viability of semantics-guided tabular data analysis operations, including on-the-fly table generation under reliability constraints, within discovery scenarios motivated by intelligence production from documents.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Scientific tables, semantic tabular data discovery, on-the-fly table generation, data fusion
© 2023 Copyright for this paper by its authors. Use permitted under Creative Commons License proach and preliminary implementation of a tabular data
discovery system driven by knowledge graph technology.</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>Tables are ubiquitous across domains in organizing and
concisely communicating information in a structured
form. Increasing democratization of generative AI
techniques and the popularity of their application in the
comprehension, creation, and refinement of complex
multimodal digital content such as technical documents is
triggering a fresh revisit of tables—often, a key
structured data component embedded within various kinds
of published documents ranging from scientific papers
and preprint articles to patents and contractual
agreements to intelligence reports and impact assessment
statements. We use the term “scientific tables” to denote this
sub-category of tabular data objects whose primary role
(alongside other structured artifacts like charts) is one
of supplementing a document’s textual content with
vital visual cues, and whose access and interpretation is
nEvelop-O
(A. Joshi)
(A. Joshi)</p>
    </sec>
    <sec id="sec-3">
      <title>2. Scientific Tables</title>
      <p>Tables in scientific documents often capture a
point-intime summary view or meta-analysis over a more
comprehensive set of information—–including over large
structured datasets hosted across diverse locations on the web
(e.g., open data portals, data markets, research data
repositories, etc.). While these latter datasets typically feed data
preparation tasks aimed at automating data science and
analysis pipelines, we focus more on the former kind of
tables to automate technical content generation pipelines.
Specifically, we seek to enable technical experts and
analysts to (i) eficiently discover useful information in
existing scientific tables, (ii) analyze knowledge extracted
from across these tables in the wider context of their
visual appearance and reliability, and (iii) present their
learnings and valuable information (again, in the form of
scientific tables) for possible inclusion in new documents
or reports. Our work is motivated by how intelligence
community standards instruct analysts to “incorporate
efective visual presentations” of information (including
via tables) to enhance the overall usefulness of
intelligence reports [10]. Additionally, we derive inspiration
from a recent trend of scientists resorting to AI and
conversational search engines as an evolving modern-day
lazyweb personification for generating tabular content
in lieu of conducting the research themselves.</p>
      <sec id="sec-3-1">
        <title>2.1. Characteristics and Challenges</title>
        <p>The very practices seeking to ease human comprehension
of scientific tables also introduce challenges for
machinedriven table understanding. Scientific tables exhibit
certain distinctive characteristics borne out of the general
circumstances of their creation:
1. High structural heterogeneity: Constrained by
‘publication real estate’ and desire to place tables in
close proximity to any accompanying text, collections
of scientific tables display high structural variability,
even more so than web tables. Data discovery and
integration systems with data and schema matching
techniques that assume ‘well-structured’ or relational
tables do not adequately address this complexity.
semantics of its content. Based on how tables are
visually formatted to optimize informational content
for human consumption, this context may include
inferred semantics of other cells in a row or column,
table captions or other descriptive text (from within
the body of the containing documents) that refers to
these tables.</p>
        <sec id="sec-3-1-1">
          <title>4. In this era of scientific misinformation and non-peer</title>
          <p>reviewed preprints, there is a dearth of approaches to
tackle the lack of information reliability of
scientific tables , as is the case with web tables and open
datasets.</p>
          <p>In [8], we describe in detail our solution to address
some of these challenges via extensive structural
characterization of over 120,000 tables drawn from scientific
publications, and by matching tabular content to
reference knowledge graphs with high precision. We also
developed a practical entity linker [9], adaptable to
different domains, to eficiently match COVID-19-related
scientific tables to Wikidata [ 12].</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>2.2. Data Discovery from Scientific Tables</title>
        <sec id="sec-3-2-1">
          <title>As an illustrative example, consider an analyst seeking</title>
          <p>meta-analyses information about phase I clinical trials for
COVID-19 vaccines developed around the world. Unless
such information is centrally curated, it will be more
expeditiously available directly within scientific tables
in published documents (e.g., figure 1).</p>
          <p>While some publisher sea.rch services [13] can
specif2. Domain-specific entities: Like open datasets, sci- ically return tables from relevant papers, such matches
entific tables typically contain more numerical cell are not based on tabular data. In reality, most tabular
content than text. Where they do contain text, it is content in scientific documents is not directly accessible
usually in the form of literals or idiomatic strings and even in instances where they are internally maintained
entities specific to a scientific domain. Data semantics as machine-readable (e.g., HTML-formatted) objects.
As[11] play a key role in disambiguating such content. suming that analyst queries are best served by
synthesizing tabular responses and that the primary source of
3. As with web tables, scientific tables exhibit difuse information for doing so are existing scientific tables, one
context wherein one must draw upon additional con- could potentially adapt current semantic schema
matchtextual information that lies outside an individual ing techniques to help discover ranked lists of tables with
table cell (or, even an entire table body) to infer the matching content. However, as with open datasets, it
is unlikely that individual scientific tables will contain
all information requested by queries. Instead, relevant
“scientific views” must be composed on the fly by suitably
merging content from multiple scientific tables. We
introduce technical challenges with on-the-fly generation
of relational scientific tables in response to search
requests under contextual constraints and briefly describe
our heuristic approach to address them in section 3.3.
3. Technical Approach
parse an input table-based search request into an
intermediate query plan comprising a set of abstract foundational
primitives: SELECT, FILTER, RANK (and FUSE). Akin to
relational algebra operators, SELECT logically returns a list
of identifiers for all tables that semantically match the
query table. FILTER prunes this list by applying one or
more temporal, cell coverage-based, or reliability-based
constraints on matching tables. RANK orthogonally
reorders the list of tables based on some ranking criteria.</p>
          <p>Any table-based search request can be expressed as a
query plan comprising these primitives.</p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>3.2. Tabular Data Discovery Engine</title>
        <sec id="sec-3-3-1">
          <title>Our high-level strategy for tabular data discovery is one</title>
          <p>of knowledge-based analysis to generate new tables on
the fly using information integrated from multiple
scientific tables. Our goal is to populate a result table with
new knowledge, potentially even one cell at a time.</p>
          <p>Intermediate query plans are automatically translated
into subgraph triple pattern-matching queries for
execution against our KG of scientific tables. We built a
• We first constructed a KG from scientific tables tabular data discovery engine to incrementally construct
[8]—Specifically, we perform column type annota- SPARQL queries by adding or modifying ad hoc graph
tion (CTA) and cell entity annotation (CEA) [14]—i.e., patterns corresponding to each primitive instance in a
header and body cells are automatically annotated with query plan. As depicted in figure 2, the discovery of
relWikidata concepts and entities respectively (e.g., the evant existing scientific tables without on-the-fly table
cell with content ‘Platform’ from the table in figure 1 is generation requires a single pass over the engine’s
comlinked to entity with QID: Q108028785: “Vaccine Plat- ponents. This engine is also responsible for packaging
form”). Each table’s structural assessments, inferred the results of SPARQL query execution in the form of
semantics, and provenance-based estimates of reliabil- relational result tables for easy consumption.
ity are all encoded as RDF triples in accordance with By translating query plans into SPARQL as shown,
linked data principles. our engine is efectively emulating “database-like”
analysis against tables in our KG—where each primitive’s
• We then designed and implemented a prototype system implementation is driven by explicit semantic linkages
for ultimately discovering tabular data from this KG automatically inferred for each table. SELECT compares
of scientific tables. As an initial step towards discov- the set of header-cell QIDs for the query table against
ery and on-the-fly generation [ 15], we first developed those for each KG table. If the two sets of QIDs
overfoundational capabilities—including a query model and lap completely, or if set overlap exceeds some
threshdiscovery engine—to enable table-based search over old specified in the request, then the candidate table is
our KG under rich contextual constraints. Finally, we deemed semantically similar and included in the results.
extended these foundational capabilities with a pre- By default, RANK sorts a list of result tables based on this
liminary, heuristic approach to identify and fuse the header-cell coverage metric (i.e., tables with maximum
content of semantically compatible scientific tables on QID set overlap are ranked higher). Besides exact
comthe fly via union operations. Our overall approach parison of header-cell QIDs, the engine also optionally
is analogous to the ‘reference architecture’ approach supports QID similarity based on pre-trained knowledge
described in [16] to discover project-join views. graph embeddings from Wembedder [17].
While a detailed algorithmic description of query
trans3.1. Table-based Semantic Search lation is beyond the scope of this paper, in general, given
a search request, a Query Parser identifies and assembles
Search requests against our KG can take the form of a key- (in a specific order) all information needed to formulate
word list or a (potentially partially-specified query-by- a SPARQL SELECT query—including: return variables
example) [16]) input table—–which can then be semanti- and limits, (subject, predicate and object) for base triple
cally resolved to reduce ambiguity—along with any asso- patterns, entire subgraph patterns (if applicable),
variciated contextual constraints. In response to a request, able bindings and clauses for the query constraints, and
we match the semantics of the query table with inferred ranking preferences. A SPARQL Formulator then acts on
semantics of tables in our KG. Since KG queries operate at these inputs one by one in the prescribed order,
expandthe granularity of low-level subgraph triple patterns, we ing them into actual triple patterns, nested sub-queries,
elevate scientific tables and relational-style analysis oper- and clauses (FILTER, HAVING, OPTIONAL, etc.) as
reations on these tables to first-class citizens in our KG. We quired, to produce a functional SPARQL query.</p>
        </sec>
      </sec>
      <sec id="sec-3-4">
        <title>3.3. On-the-fly Table Generation</title>
        <p>On-the-fly generation of tables or views by fusing the
relevant portions of a list of existing tables brings additional
complex challenges to knowledge-based analysis, e.g., de- 4. Conclusions
termining if a pair of candidate tables are compatible for
merge operations based on their inferred semantics, de- Supporting tabular data discovery over collections of
tatermining the order in which a list of candidate tables bles published in scientific documents presents unique
must be merged to produce an optimal tabular response, challenges that require use of the explicit meaning of
assessing the order in which diferent kinds of merge scientific tables as inferred by linking them to reference
operations (e.g., union, join, cell-expand) need to be ap- knowledge graphs. We described our approach and
implied on the tables, etc. Moreover, the context needed to plementation of a novel engine that can discover tabular
efectively merge content across tables may come from data from knowledge graphs of scientific tables with early
other sources like reliability assessments captured in our support for automatically merging content from multiple
KG. Knowledge derived from other modalities such as tables on the fly via union operations. While preliminary,
text may act as a ‘bridge’ between a pair of tables where we believe our approach is foundational and can be highly
compatibility may not be established based on semantics efective when expanded to cover on-the-fly cell-based
of table content alone. table expansion techniques. We believe this work will</p>
        <p>To demonstrate on-the-fly table generation capability, motivate new research directions in knowledge-guided
we implemented a preliminary dynamic approach to fuse scientific table generation and analysis at large scales.
content of relevant tables via union operations. Ours is a
heuristic-based approach that greedily seeks to maximize
the amount of populated cells in the result table via new Acknowledgments
rows while satisfying provenance-based reliability
constraints. Our engine breaks up a list of initial candidate
tables into groups where tables in each group have
identical header-cell set overlap with the query table. It then
performs a union of tables within each group to create
super-tables. If there is partial set overlap across
supertables, it performs a union across groups in decreasing
order of their set overlap sizes. Finally, these merged</p>
        <sec id="sec-3-4-1">
          <title>This research is based on work supported in part by the</title>
          <p>Ofice of the Director of National Intelligence (ODNI),
Intelligence Advanced Research Projects Activity (IARPA),
via [2021-21022600004]. The views and conclusions
contained herein are those of the authors and should not be
interpreted as necessarily representing the oficial
policies, either expressed or implied, of ODNI, IARPA, or the
U.S. Government.
[14] N. Abdelmageed, J. Chen, V. Cutrona, V. Efthymiou,</p>
          <p>O. Hassanzadeh, M. Hulsebos, E. Jiménez-Ruiz, J.
Se[1] M. Cafarella, A. Halevy, H. Lee, J. Madhavan, C. Yu, queda, K. Srinivas, Results of SemTab 2022,
volD. Z. Wang, E. Wu, Ten Years of Webtables, Proc. ume 3320 of CEUR Workshop Proceedings, 2022. URL:
VLDB Endow. 11 (2018) 2140–2149. doi:10.14778/ https://ceur-ws.org/Vol-3320/paper0.pdf.
3229863.3240492. [15] S. Zhang, K. Balog, On-the-fly Table Generation, in:
[2] S. Zhang, K. Balog, Web Table Extraction, Retrieval, The 41st International ACM SIGIR Conference on
and Augmentation: A Survey, ACM Trans. Intell. Research &amp; Development in Information Retrieval,
Syst. Technol. 11 (2020). doi:10.1145/3372117. SIGIR ’18, Association for Computing Machinery,
[3] A. Chapman, E. Simperl, L. Koesten, G. Konstan- New York, NY, USA, 2018, p. 595–604. doi:10.1145/
tinidis, L.-D. Ibáñez, E. Kacprzak, P. Groth, Dataset 3209978.3209988.</p>
          <p>Search: A Survey, The VLDB Journal 29 (2019) [16] Y. Gong, Z. Zhu, S. Galhotra, R. C.
Fernan251–272. doi:10.1007/s00778-019-00564-x. dez, Ver: View Discovery in the Wild, 2022.
[4] D. Brickley, M. Burgess, N. Noy, Google Dataset arXiv:2106.01543.</p>
          <p>Search: Building a Search Engine for Datasets in [17] F. Årup Nielsen, Wembedder: Wikidata entity
eman Open Web Ecosystem, in: The World Wide Web bedding web service, 2017. arXiv:1710.04099.
Conference, WWW ’19, Association for Computing
Machinery, New York, NY, USA, 2019, p. 1365–1375.</p>
          <p>doi:10.1145/3308558.3313685.
[5] Elicit: The AI Research Assistant, Ought, 2022.</p>
          <p>https:elicit.org.
[6] REASON - Rapid Explanation, Analysis and
Sourcing ONline, IARPA, 2023. https://www.iarpa.gov/
research-programs/reason.
[7] PMC Open Access Subset, National Library of</p>
          <p>Medicine, Bethesda, MD, 2003. https://www.ncbi.</p>
          <p>nlm.nih.gov/pmc/tools/openftlist/.
[8] V. Mulwad, V. S. Kumar, J. Weisenberg Williams,</p>
          <p>T. Finin, S. Dixit, A. Joshi, Towards Semantic
Exploration of Tables in Scientific Documents, in:
ESWC 2023 Workshops and Tutorials Joint
Proceedings, volume 3443 of CEUR Workshop Proceedings,
2023. URL: https://ceur-ws.org/Vol-3443/ESWC_
2023_SemTech4STLD_paper_2.pdf.
[9] V. Mulwad, T. Finin, V. S. Kumar, J.
Weisenberg Williams, S. Dixit, A. Joshi, A Practical
Entity Linking System for Tables in Scientific
Literature, in: Proceedings of the Workshop on
Scientific Document Understanding (SDU 2023),
colocated with 37th AAAI Conference on Artificial</p>
          <p>Inteligence (AAAI), 2023.
[10] Army Techniques Publication (ATP) 2-33.4.
Intelligence Analysis, 2020. URL: https://irp.fas.org/
doddir/army/atp2-33-4.pdf.
[11] U. Khurana, K. Srinivas, S. Galhotra, H. Samulowitz,</p>
          <p>A Vision for Semantically Enriched Data Science,
2023. arXiv:2303.01378.
[12] D. Vrandečić, M. Krötzsch, Wikidata: A Free
Collaborative Knowledgebase, Commun. ACM 57 (2014)
78–85. doi:10.1145/2629489.
[13] The New England Journal of Medicine - Advanced</p>
          <p>Search, NEJM, 2023. https://www.nejm.org/
search?searchType=advancedSearch&amp;allWords=
covid19&amp;searchWithin=fullText&amp;objectType=
nejm-media&amp;mediaType=Table.</p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>