<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Learning Regular Expressions for the Extraction of Product Attributes from E-commerce Microdata</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Petar Petrovski</string-name>
          <email>petar@informatik.uni-mannheim.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Volha Bryl</string-name>
          <email>volha@informatik.uni-mannheim.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christian Bizer</string-name>
          <email>chris@informatik.uni-mannheim.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Mannheim, Germany Research Group Data and Web Science</institution>
        </aff>
      </contrib-group>
      <fpage>45</fpage>
      <lpage>54</lpage>
      <abstract>
        <p>A large number of e-commerce websites have started to markup their products using standards such as Microdata, Microformats, and RDFa. However, the markup is mostly not as fine-grained as desirable for applications and mostly consists of free text properties. This paper discusses the challenges that arise in the task of matching descriptions of electronic products from several thousand e-shops that o↵ er Microdata markup. Specifically, our goal is to extract product attributes from product o↵ ers, by means of regular expressions, in order to build well structured product specifications. For this purpose we present a technique for learning regular expressions. We evaluate our attribute extraction approach using 1.9 million product o↵ ers from 9,240 e-shops which we extracted from the Common Crawl 2012, a large public Web corpus. Our results show that with our approach we are able to reach a similar matching quality as with manually defined regular expressions.</p>
      </abstract>
      <kwd-group>
        <kwd>Feature Extraction</kwd>
        <kwd>Entity Linking</kwd>
        <kwd>Microdata</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Recently more and more websites have started to embed structured data
describing various items into their HTML pages using markup standards, such as
Microdata1 [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], Microformats2 and RDFa [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. This results in millions of records
from thousands of data sources becoming publicly available on the Web. Being
able to integrate the data, e.g. in the e-commerce domain, would enable the
creation of powerful applications.
      </p>
      <p>
        In this paper our use case is the Common Crawl corpus3, the largest and most
up-to-date publicly available web corpus, which, among others, contains product
data from 9,000 e-shops. Bizer et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] have successfully extracted the
structured data from the 2012 Common Crawl into the WebDataCommons (WDC)
1 Microdata is a standardized HTML extension for marking up the structured data
within web pages so that it can be easily parsed by computer programs.
2 http://microformats.org/
3 Common Crawl – http://commoncrawl.org/
data set4, finding that almost 15% percent of the pages contain structured data.
Among the top topical domains of the data is the e-commerce domain with more
than 9,000 e-shops using structured data. Bringing the information from these
disparate data sets into a common integrated dataset, e.g. an open e-commerce
product catalog, would substantially increase the value of the information
collected.
      </p>
      <p>
        Figure 1a shows two product o↵ ers from the WDC dataset coming from two
di↵ erent e-shops describing the same product. Matching these product o↵ ers is
a non-trivial task [
        <xref ref-type="bibr" rid="ref10 ref11 ref16 ref9">9–11, 16</xref>
        ] since (1) product descriptions often follow di↵ erent
patterns and/or di↵ erent levels of detail (the top product description contains
less technical description than the bottom one); (2) numeric values are often
imprecise, e.g. due to rounding (the top product description contains an 11.6
inch laptop versus the 11 inch laptop in the bottom one); (3) abbreviations are
used di↵ erently (the bottom product description uses SSD as abbreviation, while
the top refers to the full name Solid State Drive).
      </p>
      <p>
        In this paper we present a feature extraction method, in order to get more
fine-grained structured data as an input for entity linking tools such as e.g.
Silk [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], and thus improve the matching precision.
      </p>
      <p>
        In [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] we have presented an approach covering the process of integration of
product data, where we proposed feature extraction methods as a key
preprocessing step for entity linking (matching).
      </p>
      <p>4 http://www.webdatacommons.org/</p>
      <p>
        The feature extraction method we proposed in [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] either require manual
configuration, and thus, good understanding of the input data, or are not able
to extract values that are not present in the training data. Therefore, in this
paper we propose a method that relies on learning regular expressions to extract
product attributes from product o↵ ers, in order to build well structured product
specifications for product matching. We explore a genetic programming approach
for learning regular expressions.
      </p>
      <p>The rest of this paper is structured as follows: In Section 2 we discuss the
state of the art in the area of product matching and learning regular expressions
for feature extraction. Section 3 gives the problem description and introduces
the automated technique for learning regular expressions. In Section 4 we report
on the evaluation of the proposed approach measuring the accuracy of the
extracted attributes as well as the product matching performance. Finally, Section
5 concludes the paper.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related</title>
    </sec>
    <sec id="sec-3">
      <title>Work</title>
      <p>
        The problem of feature extraction has been studied extensively under the topic of
entity disambiguation including product matching [
        <xref ref-type="bibr" rid="ref11 ref16 ref9">9,11,16</xref>
        ]. Specifically, K¨opcke
et al. [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] perform property mapping on product o↵ ers. While the domain is
the same as ours, only free-text properties were used for entity resolution in
the study. Product features were extracted from the title by manually defining
regular expressions. Similarly, record linkage between free-text product o↵ ers
and structured product specifications has been studied in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Structured product
specifications were used to learn lists of property values as a model for extracting
new products features from the product o↵ ers and labeling. Even though the
approach shows promising results, it lacks the ability to extract feature values
which have not been present in the training set. Petrovski et al. [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] propose a
combination of the previous two, by allowing di↵ erent product properties to be
extracted by di↵ erent extraction methods.
      </p>
      <p>
        Di↵ erently from [
        <xref ref-type="bibr" rid="ref11 ref16 ref9">9, 11, 16</xref>
        ] in this paper we propose an approach that
follows automatic induction of deterministic regular expressions [
        <xref ref-type="bibr" rid="ref14 ref2">2, 14</xref>
        ], i.e. using
genetic programming (GP) to learn regular expressions from examples in order
to perform feature extraction. The problem of inducing regular expressions from
positive and negative examples has been studied in the past, even outside the
context of feature extraction [
        <xref ref-type="bibr" rid="ref13 ref17 ref3">3, 13, 17</xref>
        ]. Most of the studies assume very strong
pattern in the examples, and thus the problem reduces to learning simple regular
expressions. For instance, applications motivated by DNA and RNA [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] view
input sequences as multiple atomic events, where each atomic event is a simple
regular expression. In a similar manner, in DTD inference [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] documents are
described using simple DTDs, thus again simple regular expressions are often
enough to capture a DTD definition. However, regular expressions for
information extraction rely on more complex constructs. Li et al. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] introduces a novel
evolutionary approach to learn regular expressions for information extraction,
starting from handful of seeds. The study presents experiments mainly on
properties with strong patterns like telephone numbers, e-mails and software names.
Similarly, to our approach Bartoli et al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] use GP to learn regular expressions
from examples. However it di↵ ers from our approach in that it is using di↵ erent
fitness function and does not mention any specific breeding techniques which are
commonly known to boost performance of GP algorithms.
3
3.1
      </p>
    </sec>
    <sec id="sec-4">
      <title>Methodology</title>
      <sec id="sec-4-1">
        <title>Problem Description</title>
        <p>We have a set D of product o↵ ers, represented as RDF statements. Every product
o↵ er d 2 D consists of set of properties (property-name, value). The properties
most frequently found are: title, with 86% usage, and description with 64% usage.
Both these properties often have an unstructured free text as values. An example
of such product description can be seen in Figure 2a. D represents a subset of the
WDC dataset containing more than 1.9 million product descriptions originating
from 9,240 e-shops.</p>
        <p>On the other hand, we have a set S of product specifications from the Amazon
product catalog, where every specification consists of well defined properties. The
properties of a product specification can be of numerical or categorical nature,
as can be seen in the product specification example in Figure 1b. In addition, the
specifications contain a textual description similar to the one found in D. Our
objective is to extract new product specifications from d, shown in Figure 2b, in
order to match them against the already existing specifications in S.</p>
        <p>
          As in [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], our key observation is that product descriptions frequently contain
product properties such as: product brand, product model, height, weight etc.
Since most of these properties follow a pattern we approach the extraction
problem by learning regular expressions for specific properties from s and applying
the same regular expressions to the free text in d.
3.2
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>Learning Regular Expressions from Examples</title>
        <p>
          Similarly to [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], we represent each valid regular expression as a tree by defining,
for each operator of a regular expression, sub-trees suitable for that operator.
The function set consists of the following regular expressions operators:
– concatenate node - a binary node that concatenates other nodes or leaves;
– possessive quantifiers - quantifier is possessive by placing an extra + after
it, making the quantifier greedy;
• ”*+” - a greedy zero or more repetitions of the preceding element;
• ”++” - a greedy one or more repetitions of the preceding element;
• ”?+” - a greedy zero or one repetitions of the preceding element;
• ”{m,n}+” - a greedy matching of the preceding element at least m and
not more than n times;
– the group operator ”()”;
– the character class ”[]”.
        </p>
        <p>The terminal set used for the leaves of the tree consists of:
– constants - a single character, a number or a string,
– ranges - ”a-z”, ”A-Z”, ”0-9” or ”a-z0-9”,
– character classes - ”\w” or ”\d”,
– white spaces -”\s”,
– the wildcard - the ”.” character.</p>
        <p>
          As input the algorithm takes a set of examples. Each example is composed of
a pair of strings: (text, the string we want to extract). For instance in Figure 1b
the pair for the Display property would be (”[the whole textual description]”,
”11.6-inch”). An example is considered negative in the case the string we want
to extract is empty. As an initial population the algorithm takes 2 times the size
of the training set, or 2 ⇤ |T |. Half of the population is generated from the
examples themselves, by changing every character sequence by \w and each number
sequence by \d. The other half of the initial population is generated randomly by
the ramped half-and-half method [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. This method uses two methods to create
trees: (a) the full method producing full/bushy trees and (b) the grow method
producing diverse tree structures with some branches longer than others. The
maximum depth of the trees generated is ramped, so that individuals are created
in a range of sizes. Using this method allows creating a diverse initial population
in terms of structure.
        </p>
        <p>The quality of a learned regular expression is assessed by the fitness function
based on user-provided training data. The prediction of the regular expression
is compared with the positive examples while counting true positives (TP) and
false negatives (FN), and the negative examples while counting false positives
(FP) and true negatives (TN). Based on these counts, a fitness value between
-1 and 1 is assigned to an individual regular expression by calculating Matthews
correlation coe cient (MCC):</p>
        <p>M CC =</p>
        <p>T P ⇥ T N F P ⇥ F N
p(T P + F P )(T P + F N )(T N + F P )(T N + F N )</p>
        <p>In contrast to many other popular fitness measures such as the F-measure
(i.e. the harmonic mean of precision and recall), Matthews correlation coe cient
yields good results even for heavily unbalanced training data.</p>
        <p>
          To improve the population, our approach makes use of two of the most
common genetic operations: crossover and mutation. A crossover operator is used to
learn more complex trees by selecting a random path of nodes in two individuals.
It then combines both paths by executing a two point crossover. A two point
crossover, shown in Figure 3, is executed by selecting two nodes on the parents
and swapping the sub-trees between these nodes, rendering two children.
Selection of the individuals is done by the tournament selection method [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ], which
involves splitting the population into groups and running several tournaments
among the individuals in the groups. The winners from the tournaments are
selected for crossover. The mutation operator is implemented similarly by selecting
the crossover operator and executing a headless chicken crossover [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] i.e. crossing
an individual from the population with a randomly generated one.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Evaluation</title>
      <p>In this section, we present performance results from two experiments using the
approach presented in the previous section. The first experiment, presented in
Section 4.1, involves the use of our approach to extract specific properties from a
set of product o↵ ers. In addition the second experiment, presented in Section 4.2,
involves product matching where we use the output from the first experiment
and match it to a subset of the already existing product specifications. The
data used in the experiments and the implementation of our approach can be
found at http://www.webdatacommons.org/structureddata/2012-08/data/
product/howto.html.
4.1</p>
      <sec id="sec-5-1">
        <title>Property extraction</title>
        <p>
          We use a set of 5,000 electronics product o↵ ers from the WDC product dataset;
the same dataset as in [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. The training set T consists of 500 product
specifications, as shown in Figure 1b, from the Amazon product catalog. The electronics
ll
a
c
e
R
ll
a
c
e
R
Precision
0.8
1.0
0.0
0.2
0.8
        </p>
        <p>1.0
0.4</p>
        <p>0.6
Precision
(c) Storage property</p>
        <p>(d) Processor property
product o↵ ers are selected to closely match the Amazon product catalog. This
was done by performing a simple pair wise matching (see the baseline method
in Section 4.2) and manually annotating o↵ ers that are rightly matched. The
Amazon product specifications are selected from the first 500 featured
electronics products on their website. The e↵ ort to create the input examples consisted
of pairing the textual description of the product specification with the property
value for the property that is being learned. The experiment involved
learning regular expressions for 5 di↵ erent properties from T : Model, Display size (in
inches), Processor, Storage size, and Dimension. The algorithm was set to run for
not more than 100 iterations and stop if the best fitness is reached (M CC = 1).
Subsequently the learned regular expressions are applied to the whole text (title
and description) of the product o↵ ers.</p>
        <p>In the following we list the learned regular expressions:
– Model - (?:[^\d]+\s[a-z0-9]+)*+
– Storage - (?:\d+[^B]+[B]+)++
– Display - \d+.[^nc]*+nc[^o]*+
– Processor - \d+\s?[^z]++z
– Dimension - \d[^.]x?[\d]++</p>
        <p>Figure 4 shows the F-measure for each of the 7 of the most popular
product categories (Smart Phones, Tablets, Laptops, TVs, Digital Cameras, HDDs,
MP3s) and for each property. This experiment indicates that this approach yields
good results, when the properties are of numeric or semi-numeric nature.
Generally, if the property is numeric or a simple combination of numbers and letters,
our approach performs with 89.4% F-measure (as can be seen from Figures 4b
to 4e) on average. The Dimension property has the highest 94.2% F-measure,
which is expected since dimension follows a strong pattern in ”length x height
x thickness”. On the other hand, in Figure 4e, the F-measure of the Model
property, which does not show a strong pattern (model can be only character
based, e.g. iPpod Nano), has an F-measure of 77.2% on average for all product
categories.
4.2</p>
      </sec>
      <sec id="sec-5-2">
        <title>Product matching</title>
        <p>
          The second experiment showcases the application of our property extraction
approach as a preprocessing step for product matching. We use the same set
of 5,000 electronics product o↵ ers from the WDC dataset as in the previous
section, and we match them against 20 products from the Amazon product
catalog (see Figure 1b for an example). We use the Silk link discovery framework5
for generating linkage rules and matching the products [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. As a baseline for the
matching task we execute pairwise matching using just the title and description
from both WDC and Amazon datasets with Jacard similarity as our similarity
measure. In order to get a more precise comparison we extract patterns from the
5 Silk is an open source tool for generation and learning linkage rules –
http://wifo503.informatik.uni-mannheim.de/bizer/silk/
title and description. From the title we extract the product brand and model by
matching a regular expression:
^.*(\w+_[a_zA-Z0-9]+)_\d.*(gb|hd|p[x]|inche?s?|m).*\$. From the description we find a number/unit of measurement
pairs, which usually correspond to numeric attributes like 5 m, 3.5 inches, 256
MB, etc.
        </p>
        <p>The second setting involves applying the learned regular expressions on the
5,000 electronics product o↵ ers in order to get new product specifications each
containing at most 5 properties (the output from the previous experiment). In
most cases the regular expression matches to one or none values, however in the
case of multiple matches we use a simple approach of choosing the first match.</p>
        <p>
          Table 1 shows the precision, recall and F-measure for the two configurations.
As can be seen, there is a big improvement in precision and F-measure when
using our automated technique for feature extraction. The precision and F-measure
are comparable to the numbers we report in [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] for the case of feature extraction
with manually created regular expressions: 82% precision and 80.9% F-measure.
Therefore, we can conclude that our approach reaches a similar matching
quality, without the need of manually assigning regular expressions, which requires
knowledge about the data as well as regular expression syntax.
5
        </p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>This paper presents an approach for extracting product features by learning
regular expressions. The evaluation indicates that this approach yields good
results when the properties are of numeric or semi-numeric nature, even though
the approach also proved competent when learning a regular expressions for more
complex properties. Moreover, the study shows that learning regular expressions
for feature extraction reaches a similar matching quality compared to the case
of feature extraction with manually created regular expressions, which requires
knowledge about the data and regular expression syntax.</p>
      <p>
        There are a number of potential future research directions. We currently do
a selection of the first match in case there are several matches when it comes
to the extraction. Ideally, we would like to rank all matches in order to improve
the extraction. One possibility of attaining this would be to perform semantic
parsing [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], where each match is compared to tagged values in a knowledge base.
Another direction is studying the e↵ ect of the elitist strategy [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], i.e. keeping
the top 1% of the population in the next iteration when it comes to breeding.
Finally, it would be interesting to apply the proposed approach to other topical
domains, such as local businesses, postal addresses, etc.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Ben</given-names>
            <surname>Adida</surname>
          </string-name>
          and
          <string-name>
            <given-names>Mark</given-names>
            <surname>Birbeck</surname>
          </string-name>
          .
          <article-title>RDFa primer - bridging the human and data webs - W3C recommendation</article-title>
          . http://www.w3.org/TR/xhtml-rdfa-primer/,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Alberto</given-names>
            <surname>Bartoli</surname>
          </string-name>
          , Giorgio Davanzo, Andrea De Lorenzo, Marco Mauri, Eric Medvet, and
          <string-name>
            <given-names>Enrico</given-names>
            <surname>Sorio</surname>
          </string-name>
          .
          <article-title>Automatic generation of regular expressions from examples with genetic programming</article-title>
          .
          <source>In Proceedings of the 14th Annual Conference Companion on Genetic and Evolutionary Computation</source>
          ,
          <source>GECCO '12</source>
          , pages
          <fpage>1477</fpage>
          -
          <lpage>1478</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Geert</given-names>
            <surname>Jan</surname>
          </string-name>
          <string-name>
            <surname>Bex</surname>
          </string-name>
          , Wouter Gelade, Frank Neven, and
          <string-name>
            <given-names>Stijn</given-names>
            <surname>Vansummeren</surname>
          </string-name>
          .
          <article-title>Learning deterministic regular expressions for the inference of schemas from XML data</article-title>
          .
          <source>ACM Trans. Web</source>
          ,
          <volume>4</volume>
          (
          <issue>4</issue>
          ):
          <volume>14</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>14</lpage>
          :
          <fpage>32</fpage>
          ,
          <year>September 2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>C.</given-names>
            <surname>Bizer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Eckert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Meusel</surname>
          </string-name>
          , H. Mu¨hleisen, M. Schuhmacher, and
          <string-name>
            <given-names>J.</given-names>
            <surname>Vo</surname>
          </string-name>
          <article-title>¨lker. Deployment of RDFa Microdata and Microformats on the Web - A Quantitative Analysis</article-title>
          .
          <source>In 12th International Semantic Web Conference In-Use track</source>
          , pages
          <fpage>17</fpage>
          -
          <lpage>32</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Moises G. de Carvalho</surname>
          </string-name>
          ,
          <string-name>
            <surname>Alberto H. F. Laender</surname>
          </string-name>
          , Marcos Andre Goncalves, and
          <article-title>Altigran Soares da Silva. A genetic programming approach to record deduplication</article-title>
          .
          <source>IEEE Trans. Knowl</source>
          . Data Eng.,
          <volume>24</volume>
          (
          <issue>3</issue>
          ):
          <fpage>399</fpage>
          -
          <lpage>412</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>I. Hickson. HTML</given-names>
            <surname>Microdata</surname>
          </string-name>
          . http://www.w3.org/TR/microdata/,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>R.</given-names>
            <surname>Isele</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Bizer</surname>
          </string-name>
          .
          <article-title>Learning Linkage Rules using Genetic Programming</article-title>
          .
          <source>In 6th International Workshop on Ontology Matching</source>
          , pages
          <fpage>1638</fpage>
          -
          <lpage>1649</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>Terry</given-names>
            <surname>Jones</surname>
          </string-name>
          . Crossover, macromutation, and
          <article-title>population-based search</article-title>
          .
          <source>In Proceedings of the Sixth International Conference on Genetic Algorithms</source>
          , pages
          <fpage>73</fpage>
          -
          <lpage>80</lpage>
          . Morgan Kaufmann,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>A.</given-names>
            <surname>Kannan</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Givoni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Agrawal</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Fuxman</surname>
          </string-name>
          .
          <article-title>Matching unstructured o↵ ers to structured product descriptions</article-title>
          .
          <source>In International Conference on Knowledge Discovery and Data Mining (KDD)</source>
          , pages
          <fpage>404</fpage>
          -
          <lpage>412</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10. H. Ko¨pcke,
          <string-name>
            <given-names>A.</given-names>
            <surname>Thor</surname>
          </string-name>
          , and
          <string-name>
            <given-names>E.</given-names>
            <surname>Rahm</surname>
          </string-name>
          .
          <article-title>Evaluation of entity resolution approaches on real-world match problems</article-title>
          .
          <source>In Proc. 36th Intl. Conference on VLDB / Proceedings of the VLDB Endowment</source>
          <volume>3</volume>
          (
          <issue>1</issue>
          ), pages
          <fpage>484</fpage>
          -
          <lpage>493</lpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11. H. Ko¨pcke,
          <string-name>
            <given-names>A.</given-names>
            <surname>Thor</surname>
          </string-name>
          , and
          <string-name>
            <given-names>E.</given-names>
            <surname>Rahm</surname>
          </string-name>
          .
          <article-title>Tailoring entity resolution for matching product o↵ ers</article-title>
          .
          <source>In Proc. 15th Intl. Conference on Extending Database Technology (EDBT)</source>
          , pages
          <fpage>545</fpage>
          -
          <lpage>550</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>John R. Koza</surname>
          </string-name>
          .
          <article-title>Genetic Programming: On the Programming of Computers by Means of Natural Selection</article-title>
          . MIT Press, Cambridge, MA, USA,
          <year>1992</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13. W. B.
          <string-name>
            <surname>Langdon</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Rowsell</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A. P.</given-names>
            <surname>Harrison</surname>
          </string-name>
          .
          <article-title>Creating regular expressions as mRNA motifs with GP to predict human exon splitting</article-title>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Yunyao</surname>
            <given-names>Li</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Rajasekar</given-names>
            <surname>Krishnamurthy</surname>
          </string-name>
          , Sriram Raghavan, Shivakumar Vaithyanathan, and
          <string-name>
            <given-names>H. V.</given-names>
            <surname>Jagadish</surname>
          </string-name>
          .
          <article-title>Regular expression learning for information extraction</article-title>
          .
          <source>In Proceedings of the Conference on Empirical Methods in Natural Language Processing, EMNLP '08</source>
          , pages
          <fpage>21</fpage>
          -
          <lpage>30</lpage>
          , Stroudsburg, PA, USA,
          <year>2008</year>
          .
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Brad L. Miller</surname>
          </string-name>
          , Brad L.
          <string-name>
            <surname>Miller</surname>
            ,
            <given-names>David E.</given-names>
          </string-name>
          <string-name>
            <surname>Goldberg</surname>
            , and
            <given-names>David E.</given-names>
          </string-name>
          <string-name>
            <surname>Goldberg</surname>
          </string-name>
          .
          <article-title>Genetic algorithms, tournament selection, and the e↵ ects of noise</article-title>
          .
          <source>Complex Systems</source>
          ,
          <volume>9</volume>
          :
          <fpage>193</fpage>
          -
          <lpage>212</lpage>
          ,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Petar</surname>
            <given-names>Petrovski</given-names>
          </string-name>
          , Volha Bryl, and
          <string-name>
            <given-names>Christian</given-names>
            <surname>Bizer</surname>
          </string-name>
          .
          <article-title>Integrating product data from websites o↵ ering Microdata markup</article-title>
          .
          <source>In Proceedings of the 4th Workshop on Data Extraction and Object Search (DEOS)</source>
          ,
          <source>WWW</source>
          <year>2014</year>
          , pages
          <fpage>1299</fpage>
          -
          <lpage>1304</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <given-names>Borge</given-names>
            <surname>Svingen</surname>
          </string-name>
          .
          <article-title>Learning regular languages using genetic programming</article-title>
          .
          <source>In Genetic Programming 1998: Proceedings of the Third Annual Conference</source>
          , pages
          <fpage>374</fpage>
          -
          <lpage>376</lpage>
          . Morgan Kaufmann,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>