<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Web Tool for the Semantic Integration of Heterogeneous and Complex Spreadsheet Tables</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sara Bonfitto</string-name>
          <email>sara.bonfitto@unimi.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luca Cappelletti</string-name>
          <email>luca.cappelletti@unimi.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Elena Casiraghi</string-name>
          <email>elena.casiraghi@unimi.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paolo Perlasca</string-name>
          <email>paolo.perlasca@unimi.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fabrizio Trovato</string-name>
          <email>fabrizio.trovato@team1994.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giorgio Valentini</string-name>
          <email>giorgio.valentini@unimi.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marco Mesiti</string-name>
          <email>marco.mesiti@unimi.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Area s.r.l.</institution>
          ,
          <addr-line>Via Torino 10/B, 12084 Mondovì</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Dep. of Computer Science, Università di Milano</institution>
          ,
          <addr-line>Via Celoria 18, 20133 Milano</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>The acquisition and integration of data contained in spreadsheet tables is a complex task because they do not impose any regular structure on the organization of the data, or constraints on valid values. Moreover, mistakes can occur due to the passage from a format to another one or misspelled words in the original sources. The automatic extraction of their content, interpretation and integration is thus a complex task. In this paper, we outline the characteristics of a semi-automatic, interactive tool conceived for creating a knowledge base by extracting semantic information from heterogeneous spreadsheets.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Heterogeneous Spreadsheet Tables</kwd>
        <kwd>Semantic Table Interpretation</kwd>
        <kwd>User Interfaces</kwd>
        <kwd>Machine Learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The table understanding problem consists in the meaningful extraction of information from
tabular data [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] contained in diferent kinds of files (HTML, pdf, csv, spreadsheets). As reported
in a recent survey [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], many approaches have been proposed for the localization of the table(s)
within a file, the segmentation of its cells, each with an eventually diferent size, the automatic
identification of their functional role, in terms of either data cell or access key (i.e. the column
headers or later stubs) and of their structural role (i.e. the identification of hierarchies on the
access keys), and the interpretation of the table content. The interpretation of a table usually
regards the transformation of its content in terms of well-known data models (like the relational
model or RDF). Nowadays, diferent approaches for table understanding have been proposed
that exploit machine learning (ML) techniques for inferring the table metadata by exploiting
domain Ontologies and (automatically or manually) tagged knowledge bases (e.g. [
        <xref ref-type="bibr" rid="ref3 ref4 ref5 ref6">3, 4, 5, 6</xref>
        ]).
      </p>
      <p>We have faced the table understanding problem in the context of a research project with an
Italian debt collection agency. The agency daily receives spreadsheets in diferent formats ( csv,
xls, xlsx) from local authorities (e.g. municipalities, tax agency), each containing batches
with thousand of invoices to be rescued. These spreadsheets are large, heterogeneous, and do
not follow any standard format or notation. Moreover, the labels used for the column headers,
if not missing, do not belong to any pre-defined dataset and sometimes are not informative
of the column content. Finally, spreadsheets often contain mistakes that need to be fixed. In
this context, usual approaches for table understanding are not successful due to the following
two reasons. First, they mainly rely on the homogeneity of the types of data contained in the
columns of the identified table (i.e. columns containing only person names, or only date of
births), and therefore exploit rules maximizing coherence in their cell segmentation methods, as
well as in their structural and functional analysis algorithms; this is impossible in our scenario,
where heterogeneity often characterizes both rows, columns, and even cells. In practice, it often
happens that both rows and columns contain diferent types of information (e.g. the name
of a person and the name of a company in the same column) and the information of a single
invoice is contained in contiguous table rows. Moreover, each cell often contains complex string
contents, which express multiple information (e.g. street number, street name, ZIP code and
city can be extracted from the string “N. 3425 Stone Street, BN12HB - London”). Finally, current
approaches rarely consider the presence of syntactic mistakes (e.g. the SSN number that does
not fit the length constraint, or a misspelled city name) and semantic mistakes (e.g. a city in the
wrong region, or the amount of a debt that does not match its detailed voices).</p>
      <p>We believe that the development of an automatic technique for addressing all the mentioned
issues, even if exploiting sophisticated ML techniques, would be unpractical. Therefore, we
introduce a semi-automatic approach that combines data management approaches with ML
techniques and intelligent user-interfaces to reduce the human efort in the extraction, cleaning
and semantic characterization of the information contained in the invoices and their integration
in a single and common knowledge base. Our human-in-the-loop approach heavily relies
on the user-interaction to tune the prediction system depending on the feedback obtained
while processing new spreadsheets. Easy-to-use graphical interfaces are thus fundamental
for correcting mistakes and improving the overall performance of the system. To reach this
goal, we propose the adoption of a three-phase approach. Phase I (Section 2) is responsible
for the spreadsheet cleaning, the identification of the column types and the syntactic error
correction. Phase II (Section 3) aims at creating a semantic characterization of the table content
to be extracted from the spreadsheet and relies on the use of a domain Ontology (DO) and,
when possible, a Knowledge Base (KB). Phase III (Section 4) relies on the construction of a KB
containing the extracted information.</p>
      <p>In the paper we use as running example the spreadsheet in Fig. 1 reporting trafic tickets which
is organized in three parts: title lines, a table of data, and footer notes. Below the heading rows,
a row contains the access keys to the columns (column schema). Column SSN/VAT contains
the code identifying either individual or companies (it is an example of heterogeneous column
containing values of two types). On the other side, the address column contains complex (string)
contents expressing multiple information; a proper string parser must be defined to correctly
extract all such information. Blanks and semi-blank rows can occur in the main table.
Semiblank rows usually contain totals or aggregated data. While some blank rows are sometimes
used for aesthetic reasons, some others for delimiting correlated rows. This is the case of the two
rows marked with a green border that represent an invoice with its the legal representative that
must pay the invoice. Last but not least, diferent kinds of errors can occur in the spreadsheet.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Phase I: Extraction, Cleaning and Typing</title>
      <p>The main purpose of this phase is to correct syntax errors occurring in the data, identifying
the existing relations among table rows, and identifying basic types of each column and cell as
detailed in the remainder of the section.</p>
      <p>
        Table extraction. The table is extracted by using the approach described in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Essentially, for
the localization of the table, we consider the density of the information contained in the table
w.r.t. external data. For the functional analysis of the cells, we use their position within the table
and adopt dictionaries of already processed spreadsheets for identifying the column schema. For
removing blank and semi-blank rows, predefined thresholds are used on the minimal number
of non-empty cells in each row w.r.t. the number of columns.
      </p>
      <p>Correlation between rows. For identifying correlation among table rows, we adopt a
declarative pattern-based language, allowing the user to define rules, which express its knowledge about
the existence of a relationship between consecutive rows. The rules can be specified through
a GUI, and the user can decide the rules to be applied in processing a specific spreadsheet.
More precisely, each rule is identified by a unique name , and is applied to consecutive
rows, the current row, , and the next row, +1. A rule is composed by a conjunction of
basic conditions (that check for the existence of a relationship) and an action that expresses the
way the information from the two rows should be joined when the condition is verified. The
following two basic  can be specified: 1) []   (named basic condition) requiring
that the -th cell in the row  ∈ {current, next} is compared according to the operator 
with a value , where  ∈ {=, ̸=, startwith, endwith}; 2) current[] = next[′] (named
equijoin condition) imposing that the -th cell of current row is equal to the ′-th cell of
next row.  is specified by a tuple (, ) determining the way in which  and +1
should be concatenated ( ∈ {natural, inverse}) and the kind of relationship that exists
among the two rows. The relationship can be: extra, representing further information about
the invoice; LR, when the row contains the legal representative of the invoice; heir when the
invoice is titled to a subject that is dead and one of his/her heirs should be contacted.
Example 1. In our scenario, a company is represented by a Legal Representative whose data are
reported in the row above the one containing the invoice. To identify this situation, the rule in
Fig. 2 presents two conditions: a basic one that identifies a cell whose value starts with Legal
Representative, and an equijoin to verify that the two addresses are equal. Whenever a pair
of rows satisfy the condition, the first row is concatenated after the second one (inverse order).</p>
      <p>Rules can be specified in advance, through a graphical interface (see Fig. 2), or on the fly, by
clicking on the operations on tuples loaded in the interface in Fig. 3 (first column). In both cases,
the table schema is duplicated and correlated rows are concatenated to form a single table line.
Type Inference. By applying the aforementioned techniques, a table  = ⟨,  ⟩ is extracted
from the spreadsheet, where  = [1, . . . ,  , . . . , ] is the list of column names, and
 = {1, . . . , } is the set of table rows. Each row , 1 ≤  ≤ , is a list of values, and
 = [,1, . . . , , , . . . , ,], one for each column identified in the column schema.</p>
      <p>
        At this stage, to discriminate the role of each cell within the table and to correctly assign
a meaning to its content, a preliminary step is the identification of the column and cell types
among those in our type system  composed of ,  and  types (details in
[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]).  types are basic domains and specific domains for our application (e.g. SSN, VAT,
emails, zip codes). A  type (e.g. address column) is a record-type associated with
diferent patterns used for extracting the record (in the case of address column, it is rec(city,
streetName, streetNumber, zip)) from a string.   types represent the occurrence
of instances of diferent types in the same column and are denoted (1, . . . , ℎ) where
1, . . . , ℎ are simple or mixed types.
      </p>
      <p>For type-inference we exploit decision trees, which are trained on synthetised table corpora,
which have been automatically tagged with the types of our type system. The decision of using
synthetised tables is firstly due to the fact that our tables use Italian language and are from
a financial context; to the best of our knowledge, no such corpora exist yet. Secondly, our
approach must be robust against mistakes, so that we have specifically developed the synthetizer
in order to produce common mistakes we experienced in the first spreadsheets we received
(and tagged as error). The corpora consists of 10.000.000 synthetic tables that are randomly
split 10 times in 8.000.000 documents for training and 2.000.000 for the validation set.</p>
      <p>The decision tree works on a value , represented by a vector  consisting of two parts:
1) (,,1, . . . , ,,) are Boolean values obtained by the application of basic-type recognizers
on the value , (several kinds of recognizers have been realized for each of the considered
basic and mixed types); 2) (,1, . . . , ,) are average numbers of recognition for each
type recognizer applied on the values occurring in the column . The generated vectors are
used for training a decision tree that predicts the type of each value contained in the table  .</p>
      <p>In order to determine the type of each column, we cannot exploit a majority vote approach,
because in our context we admit the presence of values belonging to a union type (as the case of
the column SSN/VAT). For this reason, we have adopted an approach based on the frequency of
the types in a column, where the types included in the union types are those that exceed a basic
cutof threshold. All the values whose type is not included in the union type are considered
error for the current column.</p>
      <p>Example 2. Fig. 3 shows the result of the automatic type inference approach on our running
example. White columns denote values of the same single type, whereas yellow cells denote empty
values and red cells denote type mistakes w.r.t. the type identified for the column. Cells with
mistakes need to be fixed by the user (also exploiting facilities made available in the platform for
reducing the typing eforts). When more than one colour is used for the same column, it means that
a union type was detected for such a column. Values of a mixed type are denoted by using diferent
text colours (as in the address column where each sub-component has a diferent text colour).</p>
      <p>In the address column, there are also values that the ML algorithm was not able to extract
because a non-supported pattern was used. In this case, the user exploits the interface in Fig. 4 for
extracting the pattern. Note that, the identification of a correlated row led to the introduction of
further columns w.r.t. those initially contained in the spreadsheet for representing all the
information related to a single invoice in the same row. Further, a new column named correlation has
been introduced for representing the correlation typed identified by the rule.
Interfaces for the manual typing and correction. The automatically identified types,
however, may contain errors due to the occurrence of mistakes in the data or in the classification
algorithm. Therefore, specific interfaces have been developed to make ease the correction. A
key point of this manual modification is to consider the concept and properties of the Domain
Ontology (details will be given in Section 3) in this way the human operator can give further
information for the generation of the semantic description of the data contained in the table.</p>
      <p>Fig. 4 shows the interface for altering the type associated with a single value (a similar
interface is used for altering one of the types specified for a column). In the left panel, the
Domain Ontology concepts are reported. Once a concept has been chosen, the list of its
properties is shown in the top part of the interface, so that the user can use the selected property
for labelling a portion of the value or the entire string. By assigning diferent properties to
difering parts of the string, the pattern of a mixed type is identified. The identified pattern can
thus be checked to other values of the same column (those sharing the same initial type). In
this way, a single type specification can be applied to many values. Moreover, the identified
pattern is stored in the system and made available for processing subsequent spreadsheets.
Preliminary experiments. To validate the type prediction algorithm we considered 40
spreadsheets ofered by the debt collection agency. Each spreadsheet may contain errors and has
around 100 rows and a variable number of columns (between 7 and 25). Starting from this
original dataset, the ground truth dataset has been created by manual correction of the errors
and subsequent type labelling. After processing the original documents as described in this
section, the precision, recall, and F1 measures have been computed by taking into account the
occurrences of mixed, union types, errors and empty values. We remark that the presence
of union types requires to consider as true positives (TP) those cells whose predicted type
corresponds to the manually assigned type, as false positives (FP) those cells whose predicted
type does not match the manually assigned type (among those specified in the union type), as
false negatives (FN) those cells for which an error has been predicted and the expected type is
one of the union type, as true negatives (TN) those cells whose predictive type does not match
those occurring in the union type.</p>
      <p>The results are shown in Figure 5. For more than a half of all the documents we obtained a
precision of 100%, while the lowest result is above the 60%; this result highlights that the system
rarely identifies an error as a correct value. The recall of two-thirds of the documents is over
the 85%, which means that most error cases are identified by the ML algorithm. The F1 measure
is above 80% for 32 documents. In most cases, the mistakes of the ML algorithm (FN) are due to
cells containing mixed types. As future work, we plan to exploit the user-defined patterns in
the prediction system in order to incrementally reducing the user eforts in manual typing.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Phase II: Semantic Model Generation</title>
      <p>
        The aim of this phase is to provide a semantic description of the spreadsheet tables by means
of annotations w.r.t. a Domain Ontology. Even if many approaches have been proposed for
this problem ( [
        <xref ref-type="bibr" rid="ref3 ref4 ref5 ref6">3, 4, 5, 6</xref>
        ]), in our research we wish to face the presence of union types for table
columns, which is generally neglected. Our semantic description is inspired to the one used in
Karma [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] but difers because it is created starting from the types identified in Phase I and allows
the extraction of several data from a column (while Karma only admits a 1:1 correspondence).
The Domain Ontology. Starting from the analysis of the information system currently used
by the debt collection agency for the management of the invoices, we have developed a Domain
Ontology containing the concepts and relationships useful for the considered scenario. Among
them, we have identified the invoice concept, representing the document assigned by the local
authority to the debt collection agency. The invoice contains many properties (e.g. the date in
which it has been issued, the total amount) and is associated with a subject. Subjects can be
classified in individuals and companys and can be associated with one of more addresses.
Addresses can specify the residence/domicile of an individual or the legal/administrative ofice
of a company. The invoice is associated with the detailed credit/debit voices (concept details)
that the total amount to be rescued is composed of. Each invoice can be also associated with
other subjects that are co-obligated in its payment (e.g. a legal representative or an heir).
Formally, a Domain Ontology  consists of a set of Concepts  = {1, . . . , }. Concepts
can be organized in a hierarchy of concepts: 1 ⊑ 2 denotes that 1 is subclass of 2.
Each Concept can have associated basic properties taken from a set  = {1, . . . , }; each
property is associated with basic types of our type system  ; the properties of a concept  are
denoted  (). Among the properties belonging to a class, we identified a subset of mandatory
properties. These properties need to be specified for properly characterizing an individual of the
concept and it is relevant for identifying the presence of instances of real world concepts in the
KB (and thus avoiding the introduction of duplicates). Relationships can be identified among
concepts and denoted by means of a set of roles  that identifies the kind of relationships that
bind the individuals of two concepts. Some relationships can be considered mandatory to be
identified in the semantic description (e.g. we cannot have an invoice without its subject and
the corresponding address).
      </p>
      <p>The Semantic Description of a table. The semantic description of a table  = ⟨,  ⟩ is
a graph ( ,  , ,  ,  ,  ), where  is a multiset of concepts belonging to the
Ontology,  ⊆ , (we consider a multiset because a concept can appear more than once),
 is a set of nodes corresponding to the columns in  ( ⊆ ),  ⊆  ×  are edges
representing the relationships existing among concepts in  , and  are edges representing
the extraction of values from the columns of  that are used for feeding properties of concepts
in  . Moreover labeling functions  ,  are introduced. The first one is used for assign
to edges in  the name of the relationship that they represent (i.e.  :  → , where
 is the set of relation names), whereas the second one is used for determining the type of
values that can be extracted from a column of  , i.e.  ×  →  , where  () is the set of
properties. This function assigns to an edge ( ,  ) and a type  ∈  a property among those
specified for the concept associated with the vertex  . We remark that in our context more
than one type of information can be extracted from the same column because of the presence
of union and mixed types. In the case of union type, the extraction assumes the meaning of
identifying diferent kinds of values in diferent rows. For mixed type, the extraction assumes
the meaning of identifying a record from a string value of a table row.</p>
      <p>Example 3. The top part of Fig. 6 reports the semantic description obtained for our running
example. In our web tool, this description is editable and the user can fix it or include other properties
and relationships depending on the needs. The blue nodes represent the column of the table from
which data are extracted. The green nodes represent classes. Edges among green nodes represent
relationships available in the Domain Ontology, whereas, edges between green and blue nodes
represent extraction rules. Whenever more than one edge arrives at the same blue nodes, it means that
diferent alternative information (i.e. values of a union type) or a mixed value can be extracted.
Semi-automatic Construction of Semantic Descriptions. Even if our graphical interfaces
allow the manual creation of the semantic description starting from the types obtained at the
end of phase I, this approach is tedious and does not take into account the semantic descriptions
1, . . . ,  already developed for other spreadsheets. Two approaches have been adopted for
constructing a semantic description by relying on previous experience. First, each  (1 ≤  ≤ )
is associated with an hash value obtained by applying an hash function on the ordered list of
types obtained at the end of phase I. This value is used as an index for identifying a possible
semantic description for the novel spreadsheet. This approach is adequate when the spreadsheets
always adhere to a pre-defined structure. Another approach is to keep track of the types of
the single columns and fragments of the semantic descriptions already associated with them.
Whenever a column is labeled with a type, the fragments associated with it are collected and
integrated into a single description. Since many combination of fragments could be possible, we
adopt a scoring function that try to maximize the adoption of fragments extracted from types
used in the same table. However, there is no guarantee that the final semantic description of
all the the columns of the spreadsheet is correct. For this reason, the semantic description is
drawn as a graph, and graphical facilities support the user for updating its representation.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Phase III: Construction of the Knowledge Base</title>
      <p>At the end of the previous phase, we have constructed a semantic description of the information
contained in the table extracted from the initial spreadsheet. The description characterizes the
information contained in the table, but can also be used for transforming the table rows in
instances of a KB that adhere to the constraints imposed by the Domain Ontology. In this way,
the knowledge base becomes a data-store of the invoices treated by the debt agency and can be
exploited for predicting the capacity of the invoice subjects in paying the due imports. Indeed,
by combining the invoice information with the subjects’ previous payments/not-payments, it is
possible to predict when a given invoice will be positively rescued by a given subject.</p>
      <p>In the transformation process, the following issues should be taken into account. First, the
presence of union types and missed properties can lead to the generation of diferent semantic
descriptions for each row of the table and special attention should be kept in generating
Algorithm 1 GenerateInstances</p>
      <p>Input: A knowledge base , A semantic description , A table  = ⟨,  ⟩</p>
      <p>
        Output:  enhanced with the content extracted from 
1: for each row of  do
2: create a copy ¯ = ( ,  , , ,  ,  ) of 
3: apply the extraction rule associated with 
4: assign the extracted value to the corresponding vertex in 
5: for each class  in  do
6: if all the properties in  have been removed and
7:  is not involved in mandatory relationships then
8: remove the class  and all the relationships in which it is involved
9: end if
10: if  still exists and mandatory properties in  are missing then
11: include the mandatory properties with fake values
12: end if
13: end for
14: Consider the updated structure ¯, for creating the triples to include in 
15: end for
consistent values according to our Domain Ontology. Second, when a mandatory property is
missing in the considered row, this property should be included in any case in the semantic
description and proper interfaces should be developed for supporting the user in fixing these
properties. Finally, the data extracted from the spreadsheet should be included in the KB.
Therefore, we need to identify overlapping with the instances of the KB that need to be integrated.
The integration process should guarantee that coherent and correct information are maintained
in the KB by exploiting well known properties of data quality [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
      <p>
        The last problem is a well-known problem in the context of data quality (object identification
problem [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]) and we plan to introduce facilities in this direction in the next release of our
platform. For what concerns the first two issues, Algorithm 1 has been developed with the aim
of generating the proper instances of the Domain Ontology that describe the single invoices
contained in the data table. The algorithm works row by row on the table identified at the end
of phase 1 (see Fig. 3) and allows the extraction of values according to the properties and classes
identified in the semantic description. The algorithm takes into account the possibility that
classes of the semantic description are not instantiated because alternative classes have been
fed (due to the presence of union types). Moreover, the algorithm imposes the introduction of
properties when they are mandatory in the Domain Ontology.
      </p>
      <p>Example 4. Fig. 7 reports the generated graph. Brown nodes represent invoices. An invoice
is sometimes titled to individuals, other times to companies. In the case of the correlated tuples
discussed in Example 1 for invoice 4 (the area highlighted on the top part of the figure), the invoice
has been titled to a company and presents a legal representative (note that in this specific case,
the two subjects share the same address). Finally, the subject Jane Doe (the area highlighted on
the bottom part of the figure) is represented only once in the KB because the two instances of the
subject share the same SSN, and so they are considered the representation of the same individual.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Concluding Remarks</title>
      <p>
        Both approaches for inferring simple data types and concepts of a knowledge base have been
proposed in the literature. In the first category falls many studies, including wrangling tools
[
        <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
        ], software libraries (e.g. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]) and probabilistic approaches (e.g. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]) that exploit
validation functions and regular expressions on data samples for inferring a single type. However,
these approaches are able to infer a very limited number of types and usually do not work very
well when missing and anomalous data occur in the sample. Many approaches are nowadays
proposed for inferring concepts of a knowledge base by exploiting diferent pre-annotated
corpora (e.g. [
        <xref ref-type="bibr" rid="ref14 ref4 ref8">8, 14, 4</xref>
        ]). However, the tolerance to mistakes is quite limited and in many cases
they require that table values occur in the KB.
      </p>
      <p>In this paper, we have proposed an approach that is not completely automatic but combines
ML techniques, data management facilities and user-friendly interfaces for dealing with
heterogeneous tables that can contain diferent kinds of mistakes. The approach has been tested on a
collection of documents made available from the debt collection agency that needs to handle
every day this kind of documents that are highly heterogeneous and with many mistakes. Our
initial experiments prove the feasibility of the approach and the utility of our interfaces for
easily fixing the mistakes and generating a consistent KB. The tool is not yet publicly available
because it contains sensitive information, however interested readers can get in touch with us
for a demo.</p>
      <p>
        The work discussed in this paper can be extended in several directions. By means of the
graphical interfaces described in Fig. 4 it is possible to define new patterns and include them in
the ML process described in Section 2. This is an interesting research direction in the spirit of
incremental learning and thus being able to adapt the model without the entire re-training. The
extension of the learning facilities discussed in Section 3 for learning a semantic description
from previous generated descriptions is another interesting research direction that can take
advantage of the work discussed in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Moreover, a systematic evaluation of the performances of
the entire web application should be conducted. Furthermore, we wish to combine our approach
for the construction of the semantic description of a table with the approaches proposed in
[
        <xref ref-type="bibr" rid="ref3 ref4 ref5 ref6">3, 4, 5, 6</xref>
        ] by taking into account the peculiarity of our domain. Finally, we also plan to apply the
concepts discussed in this paper for dealing with other kinds of semi-structured data in order to
semantically characterize data produced in heterogeneous sources (e.g. IoT data [
        <xref ref-type="bibr" rid="ref15 ref16">15, 16</xref>
        ], graphs
[
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], and tree-structured data) and represented in diferent models for their integration.
Acknowledgements. The authors wish to thank Giovanni Mancinelli for working on the
implementation of the tool.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hurst</surname>
          </string-name>
          ,
          <source>The Interpretation of Tables in Texts, Ph.D. thesis, Uni. of Edinburgh</source>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Bonfitto</surname>
          </string-name>
          , E. Casiraghi,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mesiti</surname>
          </string-name>
          ,
          <article-title>Table understanding approaches for extracting knowledge from heterogeneous tables, WIREs Data Mining and Knowledge Discovery (</article-title>
          <year>2021</year>
          ). doi:https: //doi.org/10.1002/widm.1407.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <article-title>Efective and eficient semantic table interpretation using tableminer +</article-title>
          ,
          <source>Semantic Web</source>
          <volume>8</volume>
          (
          <year>2017</year>
          )
          <fpage>921</fpage>
          -
          <lpage>957</lpage>
          . doi:
          <volume>10</volume>
          .3233/SW-160242.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Jiménez-Ruiz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Horrocks</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Sutton</surname>
          </string-name>
          ,
          <article-title>Colnet: Embedding the semantics of web tables for column type prediction</article-title>
          ,
          <source>in: AAAI Conf. on Artificial Intelligence</source>
          , volume
          <volume>33</volume>
          ,
          <year>2019</year>
          , pp.
          <fpage>29</fpage>
          -
          <lpage>36</lpage>
          . doi:
          <volume>10</volume>
          .1609/aaai.v33i01.
          <fpage>330129</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , E. Meij,
          <string-name>
            <given-names>K.</given-names>
            <surname>Balog</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Reinanda</surname>
          </string-name>
          ,
          <article-title>Novel entity discovery from web tables</article-title>
          ,
          <source>in: Proc. of The Web Conf., ACM</source>
          ,
          <year>2020</year>
          , p.
          <fpage>1298</fpage>
          -
          <lpage>1308</lpage>
          . doi:
          <volume>10</volume>
          .1145/3366423.3380205.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Cremaschi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. D.</given-names>
            <surname>Paoli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rula</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Spahiu</surname>
          </string-name>
          ,
          <article-title>A fully automated approach to a complete semantic table interpretation</article-title>
          ,
          <source>Future Gener. Comput. Syst</source>
          .
          <volume>112</volume>
          (
          <year>2020</year>
          )
          <fpage>478</fpage>
          -
          <lpage>500</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.future.
          <year>2020</year>
          .
          <volume>05</volume>
          .019.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>S.</given-names>
            <surname>Bonfitto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Cappelletti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Trovato</surname>
          </string-name>
          , G. Valentini,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mesiti</surname>
          </string-name>
          ,
          <article-title>Semi-automatic column type inference for csv table understanding</article-title>
          ,
          <source>in: 47th Int. Conf. on Current Trends in Theory and Practice of Computer Science</source>
          ., Springer,
          <year>2021</year>
          , pp.
          <fpage>535</fpage>
          -
          <lpage>549</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -67731-2.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Taheriyan</surname>
          </string-name>
          , et al.
          <article-title>Learning the semantics of structured data sources</article-title>
          ,
          <source>J. of Web Semantics</source>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>C.</given-names>
            <surname>Batini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Scannapieco</surname>
          </string-name>
          ,
          <source>Data and Information Quality - Dimensions, Principles and Techniques, Data-Centric Systems and Applications</source>
          , Springer,
          <year>2016</year>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>319</fpage>
          -24106-7.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Trifacta</surname>
          </string-name>
          , Trifacta wrangler,
          <year>2020</year>
          . URL: https://www.trifacta.com/.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Google</surname>
          </string-name>
          ,
          <article-title>Openrefine: A free, open source, powerful tool for working with messy data, 2020</article-title>
          . URL: https://openrefine.org/.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>T.</given-names>
            <surname>Petricek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Guerra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Syme</surname>
          </string-name>
          ,
          <article-title>Types from data: Making structured data first-class citizens in f#</article-title>
          ,
          <source>in: Proc. of 37th ACM SIGPLAN Conf. on Programming Language Design and Implementation</source>
          , ACM,
          <year>2016</year>
          , p.
          <fpage>477</fpage>
          -
          <lpage>490</lpage>
          . doi:
          <volume>10</volume>
          .1145/2908080.2908115.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>I.</given-names>
            <surname>Valera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Ghahramani</surname>
          </string-name>
          ,
          <article-title>Automatic discovery of the statistical types of variables in a dataset</article-title>
          ,
          <source>in: Proc. of Machine Learning Research</source>
          , volume
          <volume>70</volume>
          ,
          <year>2017</year>
          , pp.
          <fpage>3521</fpage>
          -
          <lpage>3529</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hulsebos</surname>
          </string-name>
          , et al.
          <article-title>Sherlock: A deep learning approach to semantic data type detection</article-title>
          ,
          <source>in: SIGKDD Int'l Conf. on Knowledge Discovery and Data Mining, ACM</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>S.</given-names>
            <surname>Bonfitto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Hachem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. G.</given-names>
            <surname>Belay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Valtolina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mesiti</surname>
          </string-name>
          ,
          <article-title>On the bulk ingestion of iot devices from heterogeneous iot brokers</article-title>
          ,
          <source>in: 2019 IEEE International Congress on Internet of Things (ICIOT)</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>189</fpage>
          -
          <lpage>195</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICIOT.
          <year>2019</year>
          .
          <volume>00039</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>S.</given-names>
            <surname>Valtolina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Ferrari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mesiti</surname>
          </string-name>
          ,
          <article-title>Ontology-based consistent specification of sensor data acquisition plans in cross-domain iot platforms</article-title>
          ,
          <source>IEEE Access 7</source>
          (
          <year>2019</year>
          )
          <fpage>176141</fpage>
          -
          <lpage>176169</lpage>
          . doi:
          <volume>10</volume>
          .1109/ ACCESS.
          <year>2019</year>
          .
          <volume>2957855</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>M.</given-names>
            <surname>Mesiti</surname>
          </string-name>
          ,
          <article-title>Mergegraphs: a web-based system for merging heterogeneous big graphs</article-title>
          ,
          <source>in: Proc. of the 17th Int'l Conf. on Information Integration and Web-based Applications &amp; Services</source>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          ,
          <year>2015</year>
          , pp.
          <volume>1</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>1</lpage>
          :
          <fpage>10</fpage>
          . doi:
          <volume>10</volume>
          .1145/2837185.2837211.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>