<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Interactive Enrichment of Tabular Data with SemTUI⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Flavio De Paoli</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roberto Avogadro</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marco Ripamonti</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Matteo Palmonari</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>SINTEF AS</institution>
          ,
          <addr-line>Oslo</addr-line>
          ,
          <country country="NO">Norway</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Milano-Bicocca</institution>
          ,
          <addr-line>Milan</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this work, we demonstrate the usage of SemTUI, an open-source tool to explore and define semanticbased data enrichment operations. The tool is designed to define sequences of data enrichment operations generalizing a link &amp; extend paradigm by integrating external services for end-to-end tabular data annotation, data reconciliation, and data extension. These services are operated through a graphical user interface that helps users explore alternatives and revise the results.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        In this paper, we present a demonstration of SemTUI [
        <xref ref-type="bibr" rid="ref10 ref9">9, 10</xref>
        ], a modular framework for
interactive enrichment of tabular data that exploits semantics under the three perspectives
discussed above. It comes with a graphical user interface (GUI), where users can upload, visualize,
enrich, and finally export a table. SemTUI interoperates with multiple services, especially data
reconciliation and extension services, including end-to-end table annotation services designed
for STI. Some services were developed by us, e.g., Alligator [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]2 for reconciliation and
end-toend annotation, while others are based on third-party APIs integrated into the framework, e.g.,
GeoNames 3 and Here 4 APIs for geocoding.
      </p>
      <p>An example of enrichment with the SemTUI GUI is described in Figure 1. SemTUI is designed
as an open system, where more services can be integrated with limited efort and is released as
open source under Apache 2.0 License. Links to GitHub repositories and two demonstration
videos are available on the tool documentation page 5.</p>
    </sec>
    <sec id="sec-2">
      <title>2. State of the Art and Novelty</title>
      <p>
        The main novelty in SemTUI consists in combining five main features, eventually enabling
sequences of data enrichment operations according to the link and extend paradigm: 1)
specification of data enrichment operations and inspection/editing of the results through a GUI
(with a special focus on reconciliation services); 2) support for end-to-end STI; 3) support for
an open set of reconciliation and extension services; 4) support for KG-based enrichment 5)
generalization of semantic-based enrichment to consider third-party APIs.
2The paper presents an improved method for cell entity annotation; for all the other STI tasks, Alligator reuses the
algorithms of s-elBat [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]
3https://www.geonames.org/
4https://www.here.com/docs/
5https://i2tunimib.github.io/I2T-docs/resources/
      </p>
      <p>
        The primary competitors of SemTUI that ofer semantic table enrichment functionalities do
not encompass all five features. OpenRefine 6 served as an inspiration for our work, sharing
many similarities in approach despite its primary focus on data cleaning operations.
Kgextension [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] is a library designed to facilitate KG-based enrichment within scikit-learn. However,
its interaction is mediated through a notebook interface, which is less intuitive compared to a
graphical user interface. MAGIC [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] provides support for KG-based enrichment, but it is tailored
to a specific knowledge graph (KG). DAGOBAH-UI [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] is a tool focused on STI annotations
and leverages links for enrichment from Wikidata. ASIA, a previous tool developed by us and
integrated into a broader data transformation framework [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], can be considered a precursor to
SemTUI. ASIA had limitations in visual exploration and extensibility for integrating external
enrichment services, these issues have been addressed in SemTUI through a redesign of the
interaction paradigm and graphical user interface (GUI). While SemTUI supports automatic
STI annotation services, it focuses on enrichment features and does not support arbitrary data
transformations like ASIA7. Additionally, few other tools such as MantisTable [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] exist for
visualizing or editing STI annotations; some of these tools are not maintained anymore and
some other have been developed in the context of the SemTab challenge [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], an initiative to
evaluate STI approaches; however, all these tools ofer limited support for data enrichment
scenarios similar to link and extend processes.
      </p>
      <p>For space constraints, Table 1 provides a comparative summary of SemTUI and other
frameworks that deliver semantic enrichment functionalities. Features exclusive to SemTUI, which
are not supported by the other tools, are indicated by no. The features labeled partially denote
those ofered by the respective tool but enhanced to some extent in SemTUI. Features marked
yes indicate functionalities that are comparable to those provided by SemTUI.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Main Features and Demonstration</title>
      <p>
        For more details, we refer to [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], where we also report some preliminary user studies; we
summarize the main features and some insights into pilot applications here below.
      </p>
      <p>Annotation and reconciliation services. End-to-end STI annotations (computed by
clicking on the Automatic Annotation button shown on the top-right corner in Figure 1) are currently
served by a selected service (Alligator). Other reconciliation services interoperate following the</p>
      <sec id="sec-3-1">
        <title>6https://openrefine.org/</title>
        <p>
          7Another limitation of the current version of SemTUI w.r.t. ASIA is the lack of support for manual editing of schema
annotations; this feature is planned for future integration
specification of W3C Reconciliation Service API v0.2 8 and include, among others: geocoding
services based on GeoNames and Here APIs, linking services based on GeoNames APIs and
Wikidata (via Alligator [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] and OpenRefine APIs 9 ).
        </p>
        <p>Extension services. Extension services are accessed by clicking on a reconciled column that
forms the input for the extension service and by specifying in a widget other input conditions
(more optional columns) and the features to add in new columns. Extension services currently
supported include, among others: Wikidata properties, geographical data (GeoNames attributes,
and route information from Here APIs), and meteorological data (from OpenMeteo10).</p>
        <p>Visual exploration and edit functionalities. Entities identified by shared systems of
identifiers can be inspected by the users. The users can look at the best candidates for each cell,
modify the link, or enter a new identifier. Visual codes (colors and shapes) are introduced to
communicate the status of reconciliation for each cell (confidence, manual revision, etc.).</p>
        <p>
          Pilot applications. We summarize various pilot applications of SemTUI for data enrichment
tested so far. 1) Enrichment from annotations computed by STI approaches, where additional
data are collected from an annotated table using found links (see the example in Figure 1 or a
short online demo. 2) Enrichment of data from digital marketing campaigns, where historical
performance data is enriched using API services to obtain coordinates and weather data; this
approach streamlines an enrichment use case discussed in prior work [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. 3) Enrichment of
procurement data with information from Wikidata, where organization names are linked to
Wikidata for additional information to support classification algorithms (see also a short online
demo). 4) Enrichment of urban data with further insights from human interaction [
          <xref ref-type="bibr" rid="ref10 ref16">10, 16</xref>
          ].
5) Enrichment of company data with corporate knowledge graphs, integrating Atoka’s entity
reconciliation and data extension services to streamline client data enrichment without needing
specialized data scientists.
        </p>
        <p>Demonstration. In the demonstration, we plan to use data from the first and the second
scenarios. In the first scenario, we use tables from the SemTab challenges and show an example
of end-to-end annotation process with the Alligator STI framework and show examples of
entity links found with other reconciliation, e.g., the Wikidata lookup service; we will show
how to explore the results and change links that are not correct; finally, we will show examples
of enrichment using Wikidata-based extension services. In the second scenario, we use digital
marketing data reconciled and enriched with third-party services to demonstrate the
servicebased generalization of the link &amp; extend paradigm using APIs. In both scenarios, we will walk
the audience through diferent enrichment steps, discussing the features of SemTUI.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Acknowledgements</title>
      <p>This work has been partially funded by the European innovation action enRichMyData (HE
101070284) and the Italian PRIN project Discount Quality for Responsible Data Science:
Human-inthe-Loop for Quality Data (202248FWFS) funded by the European Community - Next Generation
EU.</p>
      <sec id="sec-4-1">
        <title>8https://www.w3.org/community/reports/reconciliation/CG-FINAL-specs-0.2-20230410/ 9https://wikidata.reconci.link/ 10https://open-meteo.com/</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hameed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Naumann</surname>
          </string-name>
          ,
          <article-title>Data preparation: A survey of commercial tools</article-title>
          ,
          <source>SIGMOD Rec</source>
          .
          <volume>49</volume>
          (
          <year>2020</year>
          )
          <fpage>18</fpage>
          -
          <lpage>29</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Ciavotta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Cutrona</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>De Paoli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Nikolov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Palmonari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Roman</surname>
          </string-name>
          ,
          <article-title>Supporting semantic data enrichment at scale</article-title>
          ,
          <source>in: Technologies and Applications for Big Data Value</source>
          , Springer,
          <year>2022</year>
          , pp.
          <fpage>19</fpage>
          -
          <lpage>39</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>T.-C. Bucher</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Jiang</surname>
            , O. Meyer,
            <given-names>S.</given-names>
            Waitz, S.
          </string-name>
          <string-name>
            <surname>Hertling</surname>
          </string-name>
          , H. Paulheim, scikit
          <article-title>-learn pipelines meet knowledge graphs: The python kgextension package</article-title>
          ,
          <source>in: ESWC 2021 Satellite Events: Revised Selected Papers</source>
          , Springer,
          <year>2021</year>
          , pp.
          <fpage>9</fpage>
          -
          <lpage>14</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Harari</surname>
          </string-name>
          , G. Katz,
          <article-title>Automatic features generation and selection from external sources: a DBpedia use case</article-title>
          ,
          <source>Information Sciences 582</source>
          (
          <year>2022</year>
          )
          <fpage>398</fpage>
          -
          <lpage>414</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>V.</given-names>
            <surname>Cutrona</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>De Paoli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Košmerlj</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Nikolov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Palmonari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Perales</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Roman</surname>
          </string-name>
          ,
          <article-title>Semantically-enabled optimization of digital marketing campaigns</article-title>
          ,
          <source>in: ISWC</source>
          , Springer,
          <year>2019</year>
          , pp.
          <fpage>345</fpage>
          -
          <lpage>362</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>C.</given-names>
            <surname>Sarthou-Camy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Jourdain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chabot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Monnin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Deuzé</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.-P.</given-names>
            <surname>Huynh</surname>
          </string-name>
          , J. Liu,
          <string-name>
            <given-names>T.</given-names>
            <surname>Labbé</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Troncy</surname>
          </string-name>
          ,
          <article-title>DAGOBAH-UI: a new hope for semantic table interpretation</article-title>
          ,
          <source>in: ESWC Demo Papers</source>
          , Springer,
          <year>2022</year>
          , pp.
          <fpage>107</fpage>
          -
          <lpage>111</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>B.</given-names>
            <surname>Steenwinckel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>De Turck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ongenae</surname>
          </string-name>
          , MAGIC:
          <article-title>Mining an augmented graph using INK, starting from a CSV.</article-title>
          , in: SemTab@ ISWC,
          <year>2021</year>
          , pp.
          <fpage>68</fpage>
          -
          <lpage>78</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chabot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Troncy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.-P.</given-names>
            <surname>Huynh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Labbé</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Monnin</surname>
          </string-name>
          ,
          <article-title>From tabular data to knowledge graphs: A survey of semantic table interpretation tasks and methods</article-title>
          ,
          <source>Journal of Web Semantics</source>
          <volume>76</volume>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Ripamonti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>De Paoli</surname>
          </string-name>
          , M. Palmonari,
          <article-title>SemTUI: a framework for the interactive semantic enrichment of tabular data</article-title>
          ,
          <source>arXiv preprint arXiv:2203.09521</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>F. De Paoli</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Ciavotta</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Avogadro</surname>
            , E. Hristov,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Borukova</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Petrova-Antonova</surname>
            ,
            <given-names>I. Krasteva</given-names>
          </string-name>
          ,
          <article-title>An interactive approach to semantic enrichment with geospatial data, Data Knowledge Engineering (</article-title>
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>R.</given-names>
            <surname>Avogadro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ciavotta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>De Paoli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Palmonari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Roman</surname>
          </string-name>
          ,
          <article-title>Estimating link confidence for human-in-the-loop table annotation</article-title>
          , in: IEEE/WIC International Joint Conference on Web Intelligence and
          <article-title>Intelligent Agent Technology (WI-IAT)</article-title>
          , IEEE,
          <year>2023</year>
          , pp.
          <fpage>142</fpage>
          -
          <lpage>149</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M.</given-names>
            <surname>Cremaschi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Avogadro</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          <article-title>Chieregato, s-elBat: a semantic interpretation approach for messy table-s, in: SemTab @ ISWC, CEUR-WS</article-title>
          . org,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>V.</given-names>
            <surname>Cutrona</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ciavotta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>De Paoli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Palmonari</surname>
          </string-name>
          , et al.,
          <article-title>ASIA: a tool for assisted semantic interpretation and annotation of tabular data</article-title>
          ,
          <source>in: ISWC Demo Papers</source>
          , volume
          <volume>2456</volume>
          , CEUR-WS.org,
          <year>2019</year>
          , pp.
          <fpage>209</fpage>
          -
          <lpage>212</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>M.</given-names>
            <surname>Cremaschi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rula</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Siano</surname>
          </string-name>
          , F. De Paoli,
          <article-title>MantisTable: a tool for creating semantic annotations on tabular data</article-title>
          ,
          <source>in: ESWC Demo Papers</source>
          , Springer,
          <year>2019</year>
          , pp.
          <fpage>18</fpage>
          -
          <lpage>23</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>O.</given-names>
            <surname>Hassanzadeh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Abdelmageed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Efthymiou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Cutrona</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hulsebos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Jiménez-Ruiz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Khatiwada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Korini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Kruit</surname>
          </string-name>
          , et al.,
          <source>Results of SemTab</source>
          <year>2023</year>
          , in: SemTab @ ISWC, volume
          <volume>3557</volume>
          , CEUR-WS.org,
          <year>2023</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>14</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>I.</given-names>
            <surname>Krasteva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Petrova-Antonova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>De Paoli</surname>
          </string-name>
          , E. Hristov,
          <string-name>
            <given-names>M.</given-names>
            <surname>Borukova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ciavotta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Avogadro</surname>
          </string-name>
          ,
          <article-title>Geospatial enrichment of urban data for advanced city planning: a pilot study</article-title>
          , in: BigData, IEEE,
          <year>2023</year>
          , pp.
          <fpage>3139</fpage>
          -
          <lpage>3143</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>