<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Interactions between Data Mining and Natural Language Processing</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>1st International Workshop</institution>
          ,
          <addr-line>DMNLP 2014 Nancy</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Marie-Francine Moens Department of Computer Science</institution>
          ,
          <addr-line>KU Leuven Celestijnenlaan 200A, B-3001 Heverlee</addr-line>
          ,
          <country country="BE">Belgium</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Stan Matwin Faculty of Computer Science, Dalhousie University 6050 University Ave.</institution>
          ,
          <addr-line>PO BOX 15000, Halifax, NS B3H 4R2</addr-line>
          ,
          <country country="CA">Canada</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2014</year>
      </pub-date>
      <fpage>106</fpage>
      <lpage>147</lpage>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Volume Editors</title>
    </sec>
    <sec id="sec-2">
      <title>Peggy Cellier INSA Rennes, IRISA Campus Beaulieu, 35042 Rennes cedex, France E-mail: peggy.cellier@irisa.fr</title>
      <p>Copyright c 2014 for the individual papers by the papers' authors. Copying permitted
only for private and academic purposes. This volume is published and copyrighted by
its editors.
Recently, a new eld has emerged taking bene t of both domains: Data Mining (DM)
and Natural Language Processing (NLP). Indeed, statistical and machine learning
methods hold a predominant position in NLP research1, advanced methods such as
recurrent neural networks, Bayesian networks and kernel based methods are
extensively researched, and "may have been too successful (. . . ) as there is no longer much
room for anything else"2. They have proved their e ectiveness for some tasks but one
major drawback is that they do not provide human readable models. By contrast,
symbolic machine learning methods are known to provide more human-readable model that
could be an end in itself (e.g., for stylistics) or improve, by combination, further
methods including numerical ones. Research in Data Mining has progressed signi cantly in
the last decades, through the development of advanced algorithms and techniques to
extract knowledge from data in di erent forms. In particular, for two decades Pattern
Mining has been one of the most active eld in Knowledge Discovery.</p>
      <p>This volume contains the papers presented at the ECML/PKDD 2014 workshop:
DMNLP'14, held on September 15, 2014 in Nancy. DMNLP'14 (Workshop on
Interactions between Data Mining and Natural Language Processing) is the rst
workshop dedicated to Data Mining and Natural Language Processing cross-fertilization,
i.e a workshop where NLP brings new challenges to DM, and where DM gives future
prospects to NLP. It is well-known that texts provide a very challenging context to both
NLP and DM with a huge volume of low-structured, complex, domain-dependent and
task-dependent data. The objective of DMNLP is thus to provide a forum to discuss
how Data Mining can be interesting for NLP tasks, providing symbolic knowledge, but
also how NLP can enhance data mining approaches by providing richer and/or more
complex information to mine and by integrating linguistic knowledge directly in the
mining process.</p>
      <p>Out of 23 submitted papers, 9 were accepted as regular papers amounting to an
acceptance rate of 39%. In addition to regular contributions, two less mature works,
which were still considered valuable for discussion, were accepted as posters and appear
as extended abstract in this volume.</p>
      <p>The high quality of the program of the workshop was ensured by the
muchappreciate work of the authors and the Program Committee members. Finally, we
wish to thank the local organization team of ECML/PKDD 2014, and more speci
cally Amedeo Napoli and Chedy Rassy, and the ECML/PKDD 2014 workshop chairs
Bettina Berendt and Patrick Gallinari.</p>
      <p>September 2014
Peggy Cellier, Thierry Charnois</p>
      <p>Andreas Hotho, Stan Matwin
Marie-Francine Moens, Yannick Toussaint</p>
      <sec id="sec-2-1">
        <title>1 D. Hall, D. Jurafsky, and C. M. Manning. Studying the History of Ideas Using Topic</title>
        <p>Models. In Proceedings of the 2008 Conference on Empirical Methods in Natural
Language Processing, pp. 363{371, 2008</p>
      </sec>
      <sec id="sec-2-2">
        <title>2 K. Church. A Pendulum Swung Too Far. Linguistic Issues in Language Technology,</title>
        <p>Vol. 6, CSLI publications, 2011.</p>
        <sec id="sec-2-2-1">
          <title>Organization</title>
          <p>Program Chairs
Peggy Cellier
Thierry Charnois
Andreas Hotho
Stan Matwin
Marie-Francine Moens
Yannick Toussaint
Program Commitee</p>
          <p>INSA Rennes, IRISA, France
Universite Paris 13, Sorbonne Paris cite, LIPN, France
University of Kassel, Germany
Dalhousie University, Canada
Katholieke Universiteit Leuven, Belgium</p>
          <p>INRIA Nancy Grand-Est, LORIA, France
Martin Atzmueller
Delphine Battistelli
Yves Bestgen
Philipp Cimiano
Bruno Cremilleux
Beatrice Daille
Francois Jacquenet
Jiri Klema
Yves Lepage
Amedeo Napoli
Adeline Nazarenko
Claire Nedellec
Maria Teresa Pazienza
Pascal Poncelet
Stephen Poteet
Solen Quiniou
Mathieu Roche
Arnaud Soulet
Ste en Staab
Koichi Takeuchi
Isabelle Tellier
Johanna Volker
Xifeng Yan
Pierre Zweigenbaum</p>
          <p>University of Kassel, Germany
MoDyCo-Universite Paris Ouest, France
F.R.S-FNRS, Universite catholique de Louvain, Belgium
University of Bielefeld, Germany
Universit de Caen, France
Laboratoire d'Informatique de Nantes Atlantique, France
Laboratoire Hubert Curien, France
Czech Technical University, Prague, Czech Republic
Waseda University, Japan
LORIA Nancy, France
Universite de Paris 13, LIPN, France
Institut National de Recherche Agronomique, France
University of Roma "Tor Vergata", Italy
LIRMM Montpellier, France
Boeing, USA
Laboratoire d'Informatique de Nantes Atlantique, France
Cirad, TETIS, Montpellier, France
Universite Francois Rabelais Tours, France
University of Koblenz-Landau, Germany
Okayama University, Japan
Lattice, Paris, France
University of Mannheim, Germany
University of California at Santa Barbara, USA</p>
          <p>LIMSI-CNRS, Paris, France
Additional Reviewers
Eric Kergosien</p>
          <p>LIRMM, Montpellier, France
Author Index : : : : : : : : : : : : : : : : : : : : : : : : : : : : : : : : : : : : : : : : : : : : : : : : : : : : : : : : : VI</p>
        </sec>
        <sec id="sec-2-2-2">
          <title>Automatically Detecting and Rating Product</title>
        </sec>
        <sec id="sec-2-2-3">
          <title>Aspects from Textual Customer Reviews</title>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Wouter Bancken, Daniele Alfarone and Jesse Davis</title>
      <p>Department of Computer Science, KU Leuven
Celestijnenlaan 200A - box 2402, 3001 Leuven, Belgium
wouter.bancken@student.kuleuven.be
daniele.alfarone@cs.kuleuven.be</p>
      <p>jesse.davis@cs.kuleuven.be
Abstract. This paper proposes a new approach to aspect-based
sentiment analysis. The goal of our algorithm is to obtain a summary of
the most positive and the most negative aspects of a specific product,
given a collection of free-text customer reviews. Our approach starts by
matching handcrafted dependency paths in individual sentences to find
opinions expressed towards candidate aspects. Then, it clusters together
different mentions of the same aspect by using a WordNet-based
similarity measure. Finally, it computes a sentiment score for each aspect,
which represents the overall emerging opinion of a group of customers
towards a specific aspect of the product. Our approach does not require
any seed word or domain-specific knowledge, as it only employs an
offthe-shelf sentiment lexicon. We discuss encouraging preliminary results
in detecting and rating aspects from on-line reviews of movies and MP3
players.</p>
      <p>Keywords: aspect-based sentiment analysis, opinion mining, syntactic
dependency paths, text mining
1</p>
      <p>Introduction
Sentiment analysis is the task of detecting subjectivity in natural language.
Approaches to this task mainly draw from the areas of natural language processing,
data mining, and machine learning. In the last decade, the exponential growth
of opinionated data on the Web fostered a strong interest in the insights that
sentiment analysis could reveal. For example, companies can analyze user
reviews on the Web to obtain a good picture of the general public opinion on their
products at very little cost.</p>
      <p>While the first efforts in sentiment analysis were directed towards
determining the general polarity (positive or negative) of a certain sentence or document,
the interest has recently shifted towards a more qualitative analysis, that aims
to detect the different aspects of a topic towards which an opinion is expressed.
For example, we may be interested in analyzing a movie review to capture the
opinions of the reviewer towards aspects such as the plot, the cinematography,
or the performance of a specific actor. The most challenging part in aspect-based
sentiment analysis is that a system needs to detect the relevant aspects before
these can be associated with a polarity.</p>
      <p>
        In this paper we introduce Aspectator, a new algorithm for automatically
detecting and rating product aspects from customer reviews. Aspectator can
discover candidate aspects by simply matching few syntactic dependency paths,
while other approaches [
        <xref ref-type="bibr" rid="ref14 ref16 ref21 ref6">6, 14, 16, 21</xref>
        ] require seed words in input and use
syntactic dependencies or some bootstrapping technique to discover new words and
the relations between them. Additionally, it does not require any domain-specific
knowledge in input, but only few handcrafted syntactic dependency paths and
an off-the-shelf sentiment lexicon. Consequently, the proposed system can detect
and rate aspects of products in any domain, while many existing approaches [
        <xref ref-type="bibr" rid="ref16 ref18 ref21">16,
21, 18</xref>
        ] focus on domains for which machine-readable knowledge is available.
Concretely, Aspectator combines a first high-recall step where candidate aspects
are extracted from individual sentences through syntactic dependency paths,
with a second and third high-precision steps, where aspect mentions are clustered
and their sentiment scores are aggregated by leveraging an external sentiment
lexicon.
      </p>
      <p>In our opinion, the considered setting represents an ideal testbed for
investigating interactions between natural language processing and data mining.
Indeed, our focus is not on extracting the aspects discussed in a single sentence
or document, which could be seen as a problem of deep text understanding, but
on crunching hundreds of reviews of a specific product to capture the aspects
that best summarize the opinions of a group of customers, which requires
linguistic knowledge to extract information from single sentences, along with data
mining expertise to make sense of large amounts of data.
2</p>
      <p>
        Related Work
Historically, sentiment analysis has been concerned with assigning a binary
classification to sentences or entire documents, that represents the polarity (i.e., the
orientation) of the writer towards the discussed contents [
        <xref ref-type="bibr" rid="ref13 ref19 ref51 ref59">13, 19</xref>
        ]. Nevertheless,
the overall polarity gives no indication about which aspects the opinions refer to.
For this reason, in 2004 Hu and Liu [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] introduced the more interesting problem
of aspect-based sentiment analysis, where polarity is not assigned to sentences
or documents, but to single aspects discussed in them. In their approach, given
a large number of reviews for a specific product, they first attempt to identify
interesting product aspects by using association mining, and then attach a
sentiment score to each aspect by exploiting a small seed set of opinion words, along
with their synonyms and antonyms present in WordNet. Next, they use newly
detected opinion words to extract additional infrequent product aspects. Instead
of using association mining, our work will detect aspects through dependency
paths, and will use an external sentiment lexicon to rate them. However, their
work remains the most similar to ours, as in both cases the goal is to summarize a
collection of reviews for a specific product by detecting the most interesting and
discussed aspects, while most approaches focus on analyzing individual reviews.
      </p>
      <p>
        Qiu et al. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] continued to pursue the idea that opinion words can be used
to detect product aspects and vice versa, focusing on single reviews. In their
approach, a seed set of opinion words is combined with syntactic dependencies
to identify product aspects and new opinion words. To detect the polarity of
the newly identified opinion words, they consider the given polarities of the
seed words and make the assumption that opinion words expressing a sentiment
towards the same aspect in the same review share the same polarity. While Qiu
et al. use syntactic dependencies solely to capture word sequences that contain
aspects or opinion words already observed, our approach uses dependency paths
to detect new product aspects, with the potential advantage of achieving higher
coverage.
      </p>
      <p>
        A different line of work on aspect-based sentiment analysis is based on topic
models. Brody and Elhadad [
        <xref ref-type="bibr" rid="ref3 ref66">3</xref>
        ] have tried to use Latent Dirichlet Allocation
(LDA) [
        <xref ref-type="bibr" rid="ref2 ref65">2</xref>
        ] to extract topics as product aspects. To determine the polarity
towards each topic/aspect, they start from a set of seed opinion words and
propagate their polarities to other adjectives by using a label propagation algorithm.
Instead of treating aspect detection and sentiment classification as two separate
problems, Lin and He [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] and Jo and Oh [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] directly integrate the sentiment
classification in the LDA model, so that it natively captures the sentiment towards
the topic/aspect. While these LDA-based approaches provide an elegant model
of the problem, they produce topics that are often not directly interpretable as
aspects, and thus require manual labelling to achieve a readable output.
      </p>
      <p>
        The work discussed so far proposes domain-independent solutions for
aspectbased sentiment analysis, where also our approach is positioned. However, several
works make use of domain-specific knowledge to improve their results. For
instance, Thet et al. [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] focus on aspect-based classification of movie reviews, and
include as input for their algorithm movie-specific terms such as the name of the
movie, the cast and the director. Additionally, they include some domain-specific
opinion words as input for their algorithm. As expected, including
domainspecific knowledge yields a more accurate sentiment classification. To make an
example, the word “unpredictable” has a negative polarity in general English, but
in the movie domain it is often used to praise the unpredictability of a storyline.
Since all relevant aspects are given as input, they exclusively focus on detecting
opinions towards the given aspects by (1) capturing new opinion words through
syntactic dependencies, and (2) rating the product aspects based on an external
sentiment lexicon and some given domain-specific opinion words.
      </p>
      <p>
        Similarly, Zhu et al. [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] use product aspects and some aspect-related terms
as input for their algorithm, but then attempt to discover new aspect-related
terms by applying a bootstrapping algorithm based on co-occurrence between
seed terms and new candidate terms. A sentiment score is again obtained by
accessing an external sentiment lexicon. While our approach retains from these
works the usage of an external lexicon, it requires neither labelled examples nor
domain-specific knowledge, thus it has wider applicability.
      </p>
      <p>Aspectator: a New Approach
Aspectator takes as input a collection of textual customer reviews for one
specific product, and automatically extracts the most positive and the most
negative aspects of the product, together with all the sentences that contribute
to the sentiment polarity of each aspect. More precisely:
Given: a collection of textual reviews of one specific product
Extract:
– The n most positive product aspects, along with a list of all sentences
containing positive and negative mentions of each aspect.
– The n most negative product aspects, along with a list of all sentences
containing positive and negative mentions of each aspect.</p>
      <p>Aspectator works in three steps, depicted in Fig. 1. First, it detects
mentions of aspects and their associated opinion by matching handcrafted paths
in dependency trees. Second, it clusters the different mentions of an aspect
extracted in the first step by means of a WordNet-based similarity measure. Third,
it attaches a sentiment score to each mention, and aggregates the scores from
all mentions belonging to the same cluster in order to obtain a final sentiment
score for each aspect.</p>
      <p>Aspectator does not require labelled examples and it is domain-independent,
thus it can run on any collection of reviews for a specific product. The only
required external knowledge is in the form of ten handcrafted dependency paths
and an English lexicon with a sentiment score for every word.
3.1</p>
      <p>Detecting Product Aspects
The objective of the first step is to extract from customer reviews mentions of
a product aspect and the words that express the opinion of the writer towards
that aspect. For instance, given the sentence:</p>
      <p>“The action music used in the movie wasn’t too good.”
Aspectator extracts the following pair:
&lt; not too good ; action music &gt;</p>
      <p>| Smenot{dizmifieenrt } | mAes{npzteicotn }</p>
      <p>We call this an opinion pair, as the first part is the opinion of a reviewer
towards the second part. The first part can optionally be negated, as in the
above example, causing an inversion of the polarity expressed by the sentiment
modifier.
Input: a set of customer reviews for one product, e.g. the movie Batman &amp; Robin
Step 1: detection of product aspects</p>
      <p>Step 2: clustering of product aspects
cheesy film
top-notch acting
almost bad uma thurman
extremely attractive uma thurman
very cheesy acting
not bad movie
[…]
[…]
[…]
top-notch acting
very cheesy acting
…
extremely attractive uma thurman
almost bad uma thurman
…
cheesy film
not bad movie
…</p>
      <p>Step 3: rating of product aspects
top-notch acting: +0.63
very cheesy acting: -0.77
…
extremely attractive uma thurman: +0.74
almost bad uma thurman: -0.57</p>
      <p>…
-9.76
Output
+7.03
cheesy film: -0.75
not bad movie: +0.57
…
-25.6
Positive aspects:
• Uma Thurman
23 positive mentions, e.g.: “Batman and Robin has plenty of big name actors, Uma</p>
      <p>Thurman is extremely attractive as Poison Ivy and …”
9 negative mentions, e.g.: “The great Uma Thurman (Pulp Fiction, The Avengers) who
plays Poison Ivy, is almost as bad as Schwarzenegger.”
• […]
Negative aspects:
• Acting
5 positive mentions, e.g.: “The acting, storyline and visual effects were top-notch.”
22 negative mentions, e.g.: “The acting was very cheesy and predictable, but there
is some parts that boggles my mind...george clooney as batman?!”
• […]
Fig. 1. Aspectator’s full pipeline, with example extractions from reviews for the
movie Batman &amp; Robin. Scores greater or lower than zero represent positive or negative
sentiment polarity, respectively.</p>
      <p>
        Aspectator extracts opinion pairs by using ten simple handcrafted
dependency paths, in three steps:
1. For each sentence, Aspectator extracts a syntactic dependency tree by
using the Stanford dependency parser [
        <xref ref-type="bibr" rid="ref10 ref4 ref48 ref67">4, 10</xref>
        ]. Fig. 2 shows the dependencies
for the example sentence above.
2. Given a dependency tree, it attempts to extract a basic opinion pair
composed by a single-word sentiment modifier and a single-word aspect mention
by matching one of the five dependency paths shown in Table 1. For the
example sentence, this step extracts the opinion pair &lt; good ; music &gt; through
the dependency path A ←n−s−u−b−j M −c−o→p ∗.
3. Given a matched opinion pair, it attempts to extend the match to
neighbouring words by applying the additional dependency paths shown in Table 2.
This allows to (1) capture multi-word expressions, such as “action music”
and “too good ” in the running example, and (2) capture negations, such as
“wasn’t ” in the example. The final opinion pair for the running example
becomes &lt; not too good ; action music &gt;.
      </p>
      <p>
        Note that our approach leverages syntactic dependency paths for two
purposes: (1) detecting aspect mentions and sentiment modifiers, and (2)
discovering relations between them. This is a significant difference with other approaches
that are based on syntactic dependencies. For example, Qiu et al. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] only use
syntactic dependencies to identify relations between word sequences that contain
an aspect or an opinion word that has been detected before.
      </p>
      <p>While our technique for extracting aspect mentions and sentiment modifiers
yields high recall, its precision is low, since several irrelevant word sequences
are captured. Nevertheless, the following steps allow our system to assign lower
confidence to incorrect extractions, thus ultimately yielding accurate top-ranked
extractions.
∗ ←a−u−x</p>
      <p>cop
∗ ←au−x− M −−→ ∗
3.2</p>
      <p>Clustering Product Aspects
The goal of this step is to cluster the previously-extracted opinion pairs by
searching for all semantically similar aspect mentions, independently from their
sentiment modifier. For example, in the context of movie reviews, we would like to
cluster together the opinion pairs &lt; very bad ; music &gt; and &lt; awesome ; soundtrack &gt;,
as they both express opinions towards the same aspect of a movie.</p>
      <p>
        To identify semantically similar aspect mentions, Aspectator uses a
WordNetbased similarity metric called J cn [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Zhai et al. [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] experimented with several
WordNet-based similarity metrics in the context of clustering for aspect-based
sentiment analysis, concluding that J cn delivers the best results.
      </p>
      <p>J cn is based on the principle that two terms are similar if their least common
subsumer (LCS) in the WordNet taxonomy has high information content (IC).
For instance, the terms (car, bicycle), having LCS “vehicle”, are more similar
than (car, fork), having LCS “artifact ”, because “vehicle” is a more informative
term than “artifact ”. Formally, the J cn similarity between two terms t1 and t2
is defined as:</p>
      <p>J cn(t1, t2) =</p>
      <p>1</p>
      <p>IC(t1) + IC(t2) − 2 · IC(LCS(t1, t2))
where LCS(t1, t2) is the least common subsumer of t1 and t2 in WordNet, and
the information content of a term is equivalent to:</p>
      <p>IC(t) = −log P (t)
where P (t) is the probability of observing, in a large English corpus, the term t
or any term subsumed by t in the WordNet hierarchy. The higher the probability
of observing a term t or any of its subsumed terms, the lower the information
content of t.</p>
      <p>
        Concretely, in order to cluster opinion pairs, Aspectator first computes
the J cn similarity for every possible pair of aspect mentions, by using an
implementation available in the WS4J library [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Next, it normalizes all mentions
by stemming them, in order to increase data density. When two terms map
to the same root, for instance “act ” and “acting ”, a comparison with another
term is made by picking the stem that maximizes the J cn similarity. Finally,
Aspectator uses the pairwise similarity values as input for the K-Medoids
clustering algorithm [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], which will return clusters of opinion pairs, with each
cluster representing a collection of opinions towards a single aspect. K-Medoids is
preferred over K-Means because it can compute the centroid of a cluster without
the need of defining a mean.
3.3
      </p>
      <p>Rating Product Aspects
In the final stage of our approach, each cluster receives a sentiment score, which
represents the overall emerging opinion of a group of customers towards a specific
aspect of a product. Concretely, Aspectator undertakes three sub-steps for
each cluster:
(1)
(2)
1. For each opinion pair in the cluster, it assigns an individual sentiment score
to each word that composes the sentiment modifier. For instance, given the
opinion pair &lt; just plain stupid ; action music &gt;, it attaches three individual
scores to “just ”, “plain” and “stupid ”.
2. It combines the scores for the individual words into a single score for the
entire sentiment modifier, e.g., “just plain stupid ”.
3. It extracts a final sentiment score for the entire cluster by aggregating the
scores of all sentiment modifiers.</p>
      <p>
        Step 1. In order to obtain a sentiment score for individual words, Aspectator
uses the external sentiment lexicon SentiWordNet [
        <xref ref-type="bibr" rid="ref1 ref64">1</xref>
        ]. SentiWordNet extends
WordNet by attaching three scores to each synset :1 a positive sentiment score, a
negative sentiment score and a neutrality score. These three scores always sum
to 1. For example, the word “mediocre”, in the sense of “lacking exceptional
quality or ability” has the scores 0.25, 0.125 and 0.625 as its positive, neutral
and negative score, respectively.
      </p>
      <p>
        For simplicity, our approach does not use three different sentiment scores,
but combines them in one score in the range [
        <xref ref-type="bibr" rid="ref1 ref64">-1,1</xref>
        ] by subtracting the negative
score from the positive score. The neutrality score is thus ignored, as “almost
neutral” opinions will have a score close to zero, and consequently will have no
significant impact in the following aggregation steps. Instead of performing word
sense disambiguation, Aspectator simply aggregates the sentiment scores of
all the synsets in which a word w appears, as follows:
n
      </p>
      <p>P score(synseti)/i
score(w) = i=1
n
P 1/i
i=1
(3)
where i ∈ N is the rank of a synset in WordNet based on the synset’s frequency
in general English, and synseti is the ith synset of w in the ranking. Intuitively,
dividing a synset’s score by i allows our approach to give higher weight to synsets
that are more likely to represent the right sense of the word w in a certain context,
given their overall higher popularity in English.</p>
      <p>
        Step 2. The word-level scores obtained in the previous step are then combined
into a single score for the entire sentiment modifier by adopting an approach
based on the work of Thet et al. [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. Specifically, Aspectator takes as
initial score the sentiment score of the rightmost (i.e., most specific) word in the
sentiment modifier. Then, it iteratively uses the score of each preceding word
to either intensify or attenuate the current score depending on the polarity of
the considered words, remaining in the range [
        <xref ref-type="bibr" rid="ref1 ref64">-1,1</xref>
        ]. Concretely, the score for a
sentiment modifier composed by words wn wn−1 . . . w1 w0 is computed as:
      </p>
      <sec id="sec-3-1">
        <title>1 A synset is a group of synonymous words, corresponding to a node in the WordNet</title>
        <p>hierarchy.
(4a)
(4b)
score(wi . . . w0) = score(wi−1 . . . w0) − (score(wi−1 . . . w0) · |score(wi)|)
if score(wi−1 . . . w0) &gt; 0 and score(wi) &lt; 0
score(wi . . . w0) = σ · |score(wi−1 . . . w0)| + (1 − |score(wi−1 . . . w0)|) · |score(wi)|
with σ = sign(score(wi−1 . . . w0))
otherwise</p>
        <p>In case the sentiment modifier is negated, the resulting sentiment score is
multiplied by −1 to obtain the opposite polarity.</p>
        <p>Equation (4b) models the general case, where the next word wi in the iterative
procedure intensifies the current score functionally to |score(wi)|. This follows
Thet et al.’s observation that (1) words with the same polarity tend to intensify
each other (e.g., “super nice”, “terribly stupid ”), and (2) a negative current
score becomes more negative when the next word has positive score (e.g., “super
bad ”). Equation (4a) is introduced to handle the particular case in which the
current score is positive and the next word to be processed is negative (e.g.,
“hardly interesting ”). In this case, applying (4b) would make the final score
more positive, while a negative modifier should make the score less positive.</p>
        <p>As a full example, we show how our iterative procedure computes a sentiment
score for the opinion pair &lt; just plain stupid ; action music &gt;:</p>
        <p>Example opinion pair : &lt; just plain stupid ; action music &gt;</p>
        <p>w2 w1 w0
Individual scores: 0.07 0.12 −0.51
score(plain stupid ) = (−1) · 0.51 + (1 − 0.51) · 0.12
= −0.57
score(just plain stupid ) = (−1) · 0.57 + (1 − 0.57) · 0.07
= −0.60
Thus, the resulting sentiment score for the aspect mention “action music” in
this example is −0.60.</p>
        <p>Step 3. Lastly, Aspectator computes a final sentiment score for each aspect,
by summing the scores computed in the previous step for all sentiment modifiers
belonging to the aspect’s cluster. A simple algebraic summation supports the
intuition that few strongly positive/negative opinions should result in a
sentiment score comparable to the one of many weakly positive/negative opinions.
We refer back to Fig. 1 for a complete example.</p>
        <p>In order to produce the final output, Aspectator ranks the aspects by their
sentiment score, and returns only the n most positive and the n most negative
aspects, where n is specified by the user. This ranking places at the top the most
interesting aspects, i.e., the ones that (1) are frequently mentioned in the reviews,
and (2) are subjected to strong positive or negative opinions of the reviewers.
This has also the advantage that many incorrect opinion pairs extracted in the
first step of the pipeline (Sect. 3.1) will be excluded from the final output, as they
typically have very few mentions and are not associated with strong opinions.
4</p>
        <p>Experiments
In this section, we present a preliminary evaluation of Aspectator. The
objective of our experiments is to address the following questions:
1. Can our approach detect interesting and relevant product aspects?
2. Can our approach provide meaningful evidence that supports the sentiment
score assigned to each aspect?
Additionally, we discuss the main sources of error of our approach.
Aspectator’s output was manually evaluated on a portion of two public datasets
from different domains by two annotators, out of which only one was a co-author.</p>
        <p>
          The first dataset is a collection of movie reviews taken from Amazon,2
published by McAuley and Leskovec [
          <xref ref-type="bibr" rid="ref12 ref54 ref58">12</xref>
          ]. Since manual evaluation is required, we
sampled ten movies to create a validation set and a test set, in the following way.
First, we considered only the 50 movies with the highest number of reviews, as
we want to the test the ability of our algorithm to summarize a large amount of
data for a single movie. Since most of these movies have a majority of positive
reviews, in order to obtain a more balanced dataset we first took the five movies
with the highest number of negative reviews, and then randomly sampled five
other movies from the remaining set. This resulted in a collection of 700 to 850
reviews for each movie.
        </p>
        <p>
          The second dataset consists of reviews of MP3 players taken from Amazon,3
published by Wang, Lu and Zhai [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. From this dataset we selected the five
products with the highest number of reviews in the dataset, obtaining a collection
of 500 to 770 reviews for each MP3 player.
        </p>
        <p>From these samples, we used eight movies and three MP3 players as our
validation set, and the remaining two movies and two MP3 players as our test
set. We used the validation set to determine the optimal k for the K-Medoids
clustering applied in Sect. 3.2, which should ideally be equal to the total number
of unique aspects appearing in a set of reviews. We found that the optimal k is
0.9 times the number of aspect mentions to be clustered. For instance, if 1700
aspect mentions have been identified for a certain product, we set k = 1530.</p>
        <p>We used the test set consisting of two movies and two MP3 players to
manually evaluate our algorithm. For each product, two annotators were given a form
with the ten most positive and the ten most negative product aspects, along
2 http://snap.stanford.edu/data/web-Movies.html
3 http://sifaka.cs.uiuc.edu/~wang296/Data/LARA/Amazon/mp3/
with the six sentences containing the three most positive and three most
negative mentions of each aspect. The annotators were asked to mark each product
aspect and each sentence mentioning the aspect as either correct or incorrect. For
simplicity, in the form given to the annotators each aspect was only represented
by the aspect mention appearing most frequently in the reviews. An aspect is
considered correct if it is an interesting and relevant aspect for the considered
product, such as “battery life” and “display ” for an MP3 player, or “storyline”
and the name of an actor for a movie. A sentence listed by our algorithm for a
certain aspect is considered correct only if (1) the sentence mentions the
considered aspect, and (2) the sentence expresses an opinion towards the considered
aspect that matches the polarity extracted by our algorithm for that specific
opinion pair.
4.2</p>
        <p>Results
The accuracy of the top-n aspects is shown in Fig. 3. On average, the 10 most
positive and the 10 most negative aspects were considered to be correct in 72.5%
of the cases. The sentences mentioning an aspect were only considered correct in
59.8% of the cases. However, this last result can be studied more closely. Table 3
shows the accuracy of these sentences in function of the polarity of both aspects
and sentences. Clearly, the detected sentences are generally more accurate when
the aspect and the corresponding sentence have the same polarity. This is due
to the fact that for an aspect there are typically many more sentences with a
matching polarity than sentences with the opposite polarity, so when the top-3
sentences are drawn from a larger number of sentences, these tend to have higher
accuracy.</p>
        <p>1
2
3
4
5
6
7
8
9</p>
        <p>10
top-n aspects
Fig. 3. Percentage of top-1, top-3, top-5, top-10 aspects marked as correct by two
annotators.</p>
        <p>100%
90%
80%
ts 70%
c
sep 60%
tca 50%
e
rro 40%
c%30%
20%
10%
0%
Annotator #1
Annotator #2
creases the quality of clusters but at the same time it also increases the number
of context based lists tagged as “NOTVALIST” which decreases the word
coverage of clusters. (2) Decreasing threshold values of M inCoverage for parameters
Coverage although decreases the quality of clusters but at the same time it
increases the word coverage of clusters by decreasing the number of context based
lists tagged as “NOTVALIST”. (3) By varying the threshold value of Minprobdif
from 5% to 30% for parameter TagProbDif we found that increasing the
threshold value increases the precision values of POS tags but slightly decreases their
recall because the number of words tagged as “NOTAG” increases. Practical
advantage of this parameter is that it ensures that tagging of ambiguous and
non-confident cases is avoided. (4) The number of POS tag clusters obtained in
the classifier model is almost independent of the selected threshold values of the
parameters. For the datasets given in Table 1 and for the range of threshold
values M inConf idence = 60% to 90% and M inCoverage = 0% to 75%, number
of POS tag clusters found for English was 100 to 101, for Hindi was 29 to 31, for
Tamil was 22 to 26, for Bengali was 25 and for Telugu was 23. We noted that
the POS tags missing from the set of clusters were the rare POS tags having
very low frequencies.</p>
        <p>Conclusions and Future Work
In this work we developed TagMiner, a semi-supervised associative classification
method for POS tagging. We used the concept of context based list and context
based association rule mining. We developed a method to find interestingness
TagMiner
measures required to find the association rules in a semi-supervised manner
from a training set of tagged and raw untagged data combined. We showed
that TagMiner gives good performance for resource rich as well as resource poor
languages without using extensive linguistic knowledge. It works well even with
less tagged training data and less untagged training data. It can also tag unknown
words. To some extent, it handles class imbalance and data sparsity problems
using the untagged data and a special method to find interestingness measures.
It handles phrase boundary problem using a set of parameters. These advantages
make it very suitable for resource poor languages and can be used as an initial
POS tagger while developing linguistic resources for them.</p>
        <p>Future work includes (1) using other contexts instead of trigram, (2) finding
methods to include linguistic features in the current approach, (3) mining tagging
patterns from the clusters to find tag of a test word and (4) using this approach
for other lexical item classification tasks.</p>
        <p>Sequential Patterns of POS Labels
Help to Characterize Language Acquisition</p>
        <p>Isabelle Tellier1,2, Zineb Makhlouf1, Yoann Dupont1
(1) Lattice, CNRS - UMR 8094, (2) University Paris 3 - Sorbonne Nouvelle
Abstract. In this paper, we try to characterize various steps of the syntax
acquisition of their native language by children with emerging sequential patterns of
Part Of Speech (POS) labels. To achieve this goal, we first build a set of corpora
from the French part of the CHILDES database. Then, we study the
linguistic utterances of the children of various ages with tools coming from Natural
Language Processing (morpho-syntactic labels obtained by supervised machine
learning) and sequential Data Mining (emerging patterns among the sequences
of morpho-syntactc labels). This work thus illustrates the interest of combining
both approaches. We show that the distinct ages can be characterized by
variations of proportions of morpho-syntactic labels, which are also clearly visible
inside the emerging patterns.
1</p>
        <p>Introduction
The acquisition of their native language by children, especially how grammatical
constructions are gradually mastered, is a process which largely remains
mysterious. Some researches address this issue within a Natural Language Processing
framework, for example by implementing programs trying to mimic the learning
process [CM06,Ali10]. Our approach in this paper is different: we do not target
to reproduce, but to mine children productions, from a morphosyntactic point
of view. More precisely, we study the linguistic utterances of children of
various ages, seen as sequences of part-of-speech (POS) labels, with sequential data
mining tools.</p>
        <p>Sequential data mining can be applied to any kind of data following an order
relation. This relation is often related to time; for texts, it is only the linear
order of words in sentences. Sequential data mining allows to extract sequential
patterns, that is sequences or sub-sequences of itemsets that repeatedly occur in
the data. This domain has given rise to many works [AS95,SA96,Zak01,NR07].
If the extracted sequences are contiguous portions of texts, patterns coincides
with the older notion of repeated segments [Sal86].</p>
        <p>When data are composed of natural language texts, the itemsets are not
necessarily reduced to words: lemmas and POS labels can also be taken into account.
The use of sequential data mining technics in such a linguistic context has
recently been tested for the extraction of Named Entities [NAFS13], the discovery
of relations between entities in the biological field [CPRC09,CCP10,BCCC12] or
the study of stylistic differences between textual genres [QCCL12]. As we look
at the emergence of grammatical constructions in children, we are mainly
interested here in patterns of morpho-syntactic labels. As a matter of fact, they are
more general than words or lemmas and provide more abstract characterizations
of a given age. We seek in particular to exhibit specific emerging patterns for
different age groups.</p>
        <p>The remaining of the article is as follows. First, we present the way our
corpora of children’s productions of different age groups have been collected.
Then, we explain how we processed their morpho-syntactic analysis. Observing
that usual POS taggers available for French made many mistakes on our data, we
have built a new one, by training a machine learning device (a CRF model) on a
reduced set of manually corrected data. We show that, despite this reduced set of
manual corrections, the new tagger obtained behaves far better than the previous
one on our data. Finally, the last part of the paper describes the technique used
for the extraction of n-grams of morpho-syntactic labels of each specific age
group and provides quantitative and qualitative analyses of the corresponding
emerging patterns.
2
Several resources collecting children’s productions exist online, as those
available in the CNRTL1. But the best known and most widely used database is
CHILDES2 [Elm01], a multilingual corpus of transcriptions of recorded
interactions between adults and children. In this article, we are only interested in
the French part of these data. The recordings of a child cover several months or
years, the age of the children may therefore vary from one record to another.
Relying on the transcription manual3 which explicits the meta-data associated
with the corpus, we created six different sub-corpora corresponding to six age
groups: from the "1-2 years" to the "6-7 years".
In this corpus, children and parents communicate by speech turns. Each speech
turn is transcribed and delimited by a period. In the following, we consider that
each line corresponds to a "sentence". The transcriptions are annotated and are
often followed by additional information in a (semi-)standard format allowing
to describe elements of the situation (e.g. objects which are in the scene). We
performed a preprocessing step to focus only on linguistic productions. We have
1 Centre National des Ressources Textuelles et Linguistiques (http://www.cnrtl.fr for
children’s production): see Traitement de Corpus Oraux en Français (TCOF) corpus
2 http://childes.psy.cmu.edu/
3 http://childes.psy.cmu.edu/manuals/CHAT.pdf
removed all special characters related to standards of transcription, as well as
all information of phonetic nature, which are not relevant for the analysis of
syntactic constructions and prevent the use of a tagger. We have also eliminated
from our data all adult utterances.</p>
        <p>The characteristics of each of our initial sub-corpora are presented in the
table of Figure 1. There are differences between them: the corpus for the age of
"6-7 years" is the smallest one. To balance the corpora of the different age groups,
we have sampled them according to the number of words : this feature is more
reliable than the number of sentences, because the length of the sentences is a
key factor which significantly varies from one age to another (see the following).
To have comparable sub-corpora, the number of words is thus more reliable than
the number of sentences.</p>
        <p>1-2 years
2-3 years
3-4 years
4-5 years
5-6 years
6-7 years</p>
        <p>Fig. 1. Characteristics of the initial sub-corpora
The smallest corpus in terms of words (the one of "6-7 years") is the reference
sample for the other age groups. So, we chose to take 20,000 words per corpus,
with a rate of 0.01% tolerance. To build our new corpora from the initial ones, we
sampled sentences randomly until the sum of all words in all sentences reaches
this size. After the sampling, we have six new corpora, whose properties are
given in the table of Figure 2.</p>
        <p>The corpora now have comparable size in terms of words. The number of
sentences in each corpus have of course decreased, but we note that the average
lengths of the sentences follow the same evolution than in the initial corpora. This
is crucial because, as long as the children grow up, they tend to produce longer
sentences. This is a well-known key feature of language acquisition [Bro73,MC81].
To go further in our exploration, we will now label the productions of the children
with morpho-syntactic labels.
1-2 years
2-3 years
3-4 years
4-5 years
5-6 years
6-7 years</p>
        <p>POS labeling</p>
        <p>Use of an existing tagger
As we want to characterize the acquisition of syntactic constructions, we need
more information than simple transcriptions of words. Our experiments in this
article rely on a morpho-syntactic tagging of children’s productions: we must
thus assign to each word in the sub-corpora a label corresponding to its
grammatical category. Several tools are available to annotate plain text in French
with "Part of Speech" (POS) labels, such as TreeTagger [Sch94]. In our work,
we have used SEM4 [TDE+12], which was obtained by training a linear CRF
(Conditional Random Fields) model on the French Treebank [ACT03]. The set
of labels adopted in SEM, similar to the one of [CC08], includes 30 different
categories among which the main important ones for the following are: NC (for
common nouns), V (for verbs), DET (for determiners), P (for prepositions), I
(for interjections) and CLS (for subject clitic). SEM also integrates the external
lexical resource Lefff [CSL04] to help achieve a better labeling.</p>
        <p>SEM has been learned with labeled sentences extracted from the French
newspaper "Le Monde". Our texts of children productions have very different
properties, and we therefore expect many annotation errors. Indeed, the
corpus CHILDES is composed of oral transcription, whose conventions differ from
those of writing (especially concerning punctuations). Furthermore, children
utterances are often far from standard French. It has already been observed that,
even if SEM is supposed to reach 97% accuracy on texts similar to those on
which it has been learned, it reaches 95.6% accuracy on more casual written
texts from blogs, and only 81.6% on oral productions of adults.</p>
        <p>To assess the quality of SEM on our data, we have randomly selected 200
sentences from each of our six corpora, tagged them with SEM and manually
corrected the labeling errors, following the annotation conventions of the French
Treebank. The accuracy of SEM on these samples (see table of Figure 4) ranges
from 70% (2-3 years) to 87% (6-7 years). The detailed F-measures of the main
categories for each age group can also be seen in the table of Figure 3: the
label interjection (I), very rare in the French Treebank but very frequent in our
4 http://www.lattice.cnrs.fr/sites/itellier/SEM.html
corpora, are particularly not well recognized by SEM (the F-measures goes from
33.33 for the "1-2 years" age group to 0 for the the "6-7 years" one).
3.2</p>
        <p>Learning a New tagger
As we want to perform statistical measures on the morpho-syntactic labels,
labeling errors must be reduced as much as possible. In [TDEW13], it has been
shown that to learn a good tagger by supervised machine learning, it is more
efficient to have a small annotated corpus similar to the target data than to have
a large too different training set. So, we decided to use the labelled sentences
which have been manually corrected for the evaluation of SEM as training data
to learn a new tagger adapted to our corpora.</p>
        <p>For this, we have used the same tools as those used to learn SEM, that is
CRFs (Conditional Random Fields), introduced by [LMP01] and implemented in
the software Wapiti [LCY10]. CRFs are graphical models that have proven their
effectiveness in the field of automatic annotation by supervised machine learning
[TTA09,TDE+12]. They allow to assign the best sequence of annotations y to
an observable sequence x. For us, the elements of x are words enriched with
endogenous attributes (presence of caps, digits, etc.) or exogenous ones (e.g.
associated properties in Lefff), while y is the corresponding sequence of
morphosyntactic labels.</p>
        <p>We trained our new tagger thanks to 200 ∗ 6 = 1200 annotated and manually
corrected sentences (which is a very small number to learn a POS tagger), and
we tested it on 50 ∗ 6 = 300 other independent sentences, equally sampled from
the 6 distinct sub-corpora. The table of Figure 3 gives the F-measures of the
main labels obtained by SEM and by the re-learned tagger for each age group,
while the accuracy of both taggers are provided in the table of Figure 4.
corpus CLS DET I NC P V
1-2 years 100/100 80/100 33.33/57.14 76.92/84.21 0/0 80/100
2-3 years 71.43/93.33 66.67/54.55 12.5/90.91 71.43/80 40/33.33 71.43/63.64
3-4 years 77.42/100 80/78.26 13.33/88.89 88.89/94.74 71.43/71.43 83.87/94.74
4-5 years 89.8/94.55 80.95/89.36 8.7/97.78 75.76/93.15 90.91/80 88.89/95.89
5-6 years 81.08/97.56 91.18/93.15 0/94.74 86.32/96.08 78.05/88.89 92.96/90.14
6-7 years 96.55/100 87.88/97.14 0/80 90/92.13 89.47/87.8 93.88/89.36
Fig. 3. F-measures of the main distinct labels before (with SEM) /after the re-learning</p>
        <p>We observe that the relearning leads to a significant improvement of the
accuracy of about 10% in average. SEM is better for only 4 cells out of 36 in the
table of Figure 3, probably thanks to its better vocabulary exposure: the French
Treebank on which SEM was learned was about ten times larger than our training
corpus. The improvement brought by relearning is larger for oral-specific labels
such as I. It is therefore very beneficial, despite a very small training corpus. This
corpus SEM
1-2 years 82%
2-3 years 70%
3-4 years 73%
4-5 years 75%
5-6 years 80%
6-7 years 87%
average 77.83%</p>
        <p>Fig. 4. Impact of the re-learning on the accuracy of the distinct age groups
can be explained by the fact that the vocabulary used in our texts is relatively
limited and redundant: few data are therefore sufficient to obtain a tagger which
is effective on our corpus, even if it is not uniformly better than SEM on every
label (it would obviously be much less effective on other types of data). In the
following, we systematically use the new version of the tagger.</p>
        <p>Analysis of POS labels
Figure 5 shows the distribution of the main morpho-syntactic categories in the
different age groups. For example, we see that the curve of the label I
(interjection) is decreasing (except for the 4-5 years age group): it seems that children
use fewer and fewer interjections in their productions as long as grow up. In
contrast, the label P (preposition) is strictly increasing, which is consistent with
an acquisition of increasingly sophisticated syntactic constructions. Curves for
the labels CLS (subject clitic) and V (verb) follow very similar variations,
probably because they are often used together: they increase till the age of 4, then
decrease from 4 to 6, and finally stabilize at the age of 6. Observing labels DET
(determiner) and NC (common nouns), we notice that until the age of 4 years,
NC is the most common label, but not yet being systematically associated with a
DET. It is only at the age of 4 that both curves become parallel (most probably
when most NC is preceded by a DET). We finally note that from the age of 5
years, the proportions of different labels stabilize.</p>
        <p>The residual errors of the tagger (there is more than 10% remaining labeling
errors) lead us to be prudent with these observations. But it is clear that some of
the phenomena observed here would not have been possible without re-learning:
interjections, for example, were the words most poorly recognized by the original
SEM, because they are very rare in newspaper articles. However, their production
appears to be an important indicator of the child’s age group. Example sentences
like "ah maman" ("ah mom") or "heu voilà " ("uh there") were respectively
labeled as "ADJ NC" and "ADV V" with the original SEM tagger. After the
re-learning, the labels became "I NC " and "I V", which is at least more correct.</p>
        <p>Although we can already draw some interesting conclusions from these curves,
we cannot characterize the syntactic acquisition of children from single isolated
categories. We thus decided to use sequential data mining techniques on our data
to explore them further.
4
Many studies have focused on the analysis of texts seen as sequential data.
For example, the notion of repeated segment is used in textometrics [Sal86] to
characterize a contiguous sequence of items appearing several times in a text.
Sequential data mining [AS95] generalizes such concept, with notions like
sequential patterns of itemsets. In our case, itemsets can be composed of words
and POS labels. A sequence of itemsets is an ordered list of itemsets. An order
relation can be defined on such sequences: a sequence S1 = hI1, I2, ..., Ini is
included into a sequence S2 = hI10, I20, ..., Im0i, which is noted S1 ⊆ S2, if there exist
integers 1 ≤ j1 ≤ j2 ≤ ... ≤ jn ≤ m such that I1 ⊆ Ij01 , I2 ⊆ Ij02 , ..., In ⊆ I0
jn (in
the classical sense of itemset inclusion). The table of Figure 6 provides examples
of sequences of itemsets found in our corpus labelled with the re-trained tagger.</p>
        <p>The support of a sequence S, denoted sup(S), is equal to the number of
sentences of the corpus containing S. For example, in the table of Figure 6,
sup(h(ADJ) (NC)i) = 2. The relative support of a sequence S is the proportion
of sequences containing S in the base of initial sequences. It is worth 12 for
the sequence in our example, because this sequence is present in 2 out of the
4 sequences of the database. Algorithms mining sequential patterns are based
on a minimum threshold for extracting frequent patterns. A frequent pattern is
thus a sequence for which the support is greater than or equal to this threshold.
Other concepts are also useful to limit the number of extracted patterns.
3</p>
        <p>sequence
h(le, DET) (petit, ADJ) (chat, NC)i</p>
        <p>("the little cat")
h(le, DET) (grand, ADJ) (arbre, NC)i</p>
        <p>("the big tree")
h(le, DET) (chat, NC)i</p>
        <p>("the cat")
h(tombé, VPP) (et, CC) (cassé, VPP)i</p>
        <p>("fallen and broken")</p>
        <p>Extraction of Sequential Patterns under constraints
In [YHA03], was introduced the notion of closed patterns that allows to eliminate
redundancies without loss of information. A frequent pattern S is closed, if there
is no other frequent pattern S0 such S ⊆ S0 and sup(S) = sup(S0). In our
example, if we fix minsup=2, the frequent pattern h(DET) (NC)i, extracted from
Figure 6, is not closed because it is included in the pattern h(le, DET) (NC)i
and they both have a support equals to 3. But the pattern h(DET) (small, ADJ)
(NC)i is closed. A length constraint can also be used. It defines the minimum
and maximum number of items contained in a pattern [BCCC12].
There are several available tools for extracting sequential patterns such as GSP
[SA96] and SPADE [Zak01]. CloSpan [YHA03] and BIDE [WH04] are able to
extract frequent closed sequential patterns. SDMC5, used here, is a tool based
on the method proposed in [PHMA+01]. It extracts several types of
sequential patterns, where items can correspond to simple words, lemma and/or their
morpho-syntactic category (the tagger is parameterized, which allowed us to use
our tagger). In this work, we wanted to characterize grammatical constructions,
and we thus focused only on sequences of POS labels. The algorithm of SDMC
implements the pattern growth technic; it is briefly discussed in [BCCC12]. It
allows to extract sequential patterns under several constraints.
[DL99] introduced the concept of emerging pattern. A frequent sequential
pattern is called emerging if its relative support in a set of data set is significantly
higher than in another set of data. Formally, a sequential pattern P of a set of
data D1 is emerging relatively to another set of data D2 if GrowthRate(P ) ≥ ρ,
5 https://sdmc.greyc.fr, login and password to be asked
with ρ &gt; 1. The growth rate function is defined by:
∞
ssuuppppDD12 ((PP )) otherwise</p>
        <p>if supportD2 (P ) = 0
where suppD1 (P ) (respectively suppD2 (P )) is the relative support of the
pattern P in D1 (respectively D2). Any pattern P whose support is zero in a set is
neglected.
5
The corpora used in our experiments are those described in section 2.3. We are
interested here in sequences of itemsets restricted to POS labels without any gap
(thus corresponding to n-grams, or repeated segments of labels), under some
constraints (such as having a support strictly greater than a given threshold
or pruning non-closed patterns), to limit their number. To set the lengths of
sequences, we took account of the average size of sentences. So, we have decided
to select patterns of length between 1 and 10. The minsup threshold is set to 2
and ρ = 1.001. To find the emerging patterns of a certain age group, we do as
[QCCL12] did for literary genres: each age group (D1) is compared to the set of
every other age groups (D2).
Figure 7 shows the number of frequent and emerging patterns obtained under
our constraints for each age group. For example, for the age of 4-5 years, there are
1933 frequent patterns but only 842 emerging ones (42.6%). A serious reduction
has occurred, which will make the observation easier. The number of emerging
patterns is relatively stable across ages from 3-4 years and is important in each
age group. As these emerging patterns are defined relatively to every other age
group, this suggests the existence for each age group of characteristic phases of
grammatical acquisitions.</p>
        <p>Figure 8 shows the average size of the frequent and emerging patterns for each
age group. The curves are very similar, suggesting that emerging patterns have
properties which are similar to frequent patterns. In both cases, the length is
increasing and reaches its maximum at the age of 5-6 years old. This parameter
seems very correlated to the one of sentence length (see Figure 1): not only
utterances become longer as the children grow up, but also the grammatical
patterns they instantiate.</p>
        <p>Figures 9 and 10 show the distributions of the main morpho-syntactic
labels in frequent and emerging patterns respectively for each age group. These
results are consistent with those obtained on the entire corpus (cf. Figure 5).</p>
        <p>Fig. 8. Average length of frequent versus emerging patterns for each age group
The proportion of interjections still regularly decreases, while the one of
prepositions increases, which is consistent with syntactic constructions of increasing
complexity. We also note that the CLS and V curves are parallel and that, before
the age of 4 years, the NC label is very frequent without being associated with
the label DET. These curves show that the proportions of labels in the frequent
and emerging patterns of each age group are similar to those of the corpus. In
this sense, these patterns seem to be representative of the different age groups.
5.3
The table of Figure 11 provides examples of emerging patterns of each age group,
and some corresponding sentences. These examples show that a single pattern
can correspond to various sentences, and that they have increasing complexity.
We note that even before the age of 2, children can produce sentences with a
CN preceded by a DET. We also note, for example, that the patterns "(DET)
(NC)" and "(DET) (NC) (CLS) (V) (VINF)" respectively extracted of the age
"1-2 years" and "4-5 years are included in "(P) (DET) (NC) "and" (DET) (NC)
(CLS) (V) (VINF) (DET) (NC)" respectively, of the following age group. This
is consistent with a gradual acquisition of complex syntactic constructions.
6</p>
        <p>Conclusion
In this article, we have applied techniques from Natural Language Processing,
machine learning and sequential Data Mining to study the evolution of children’s
utterances of different ages. The phase of morpho-syntactic labeling required the
learning of a specific tagger, adapted to our data. It was a necessity, considering
that current available taggers do not properly handle oral transcriptions, and
even less those of children: interjections, for example, which are very specific of
1-2 years (P) (NC) - à maman ("to mom")</p>
        <p>- sac à dos ("backpack")
(DET) (NC) - le ballon ("the ball")</p>
        <p>- des abeilles ("some bees")
2-3 years (P) (DET) (NC) - de la tarte
("some pie")
- poissons dans l’eau
("fishes in the water")
(ADVWH) (CLS) (V) - où il est ?
("where it is ?")
- comment il marche ?
("how it works ?")
3-4 years (ADV) (CLS) (V) - non il est par terre
(" no it is on the floor")
- ici il pourra passer
("here it will be able to pass")
4-5 years (ADV) (CLS) (CLO) (V) - alors tu m’as vue ?
("so you saw me ?")
- oui j’en fais souvent
("yes I do some often")
(DET) (NC) (CLS) (V) (VINF) - les lapins ils vont rentrer
("the rabbits they will come in")
- le chat il veut attraper l’oiseau
("the cat it wants to catch the bird")
5-6 years (DET) (NC) (CLS) (V) (VINF) - l’enfant il va chercher le chat
(DET) (NC) ("the child he goes and fetch the cat")
- le monsieur il va chercher les cerises
("the man he goes and catch the cherries")
(CC) (DET) (NC) (CLS) (V) - la maman et le papa ils regardaient le garçon
(DET) (NC) ("the mommy and the daddy they watched the boy")
- et le chat il mange les cerises
("and the cat it eats the cherries")
6-7 years (P) (VINF) (DET) (NC) - les oiseaux les aident à ramasser les cerises
("the birds help them to pick up the cherries")
- il y a un chat qui essaie de chasser des oiseaux
("there is a cat trying to catch birds")
(DET) (NC) (PROPEL) (V) - il y a un chat qui suit la fille avec son panier
(DET) (NC) (P) (DET) (NC) ("there is a cat which follows the girl with a basket")
- et aussi un monsieur qui ramasse des cerises
dans un arbre
("and a man picking up cherries in a tree")</p>
        <p>Fig. 11. Examples of emerging patterns in each age group
oral productions, would have been poorly recognized without re-learning. This
is crucial, as the curves of label proportions show that their frequency appears
as an important way to characterize a child’s age group.</p>
        <p>We currently restricted our research to n-grams of POS labels but further
work could use richer itemsets of the type (word, lemma, POS tag). Our
exploration seems to confirm that the extracted emerging patterns are representative
of the age group in which they arise. The provided examples further confirm
the intuition that (at least some of) the patterns of increasing age groups are
included into each other, going in the direction of a grammatical sophistication.</p>
        <p>As far as we know, these kinds of analyses had never been performed before.
Of course, a detailed analysis of the patterns obtained remains to be done by
specialists of language acquisition. They could for example allow to characterize
typical evolutions of grammatical knowledge, or help to diagnose pathological
evolution of a child’s productions. We hope that they will provide valuable tools
for the study of language acquisition phases.</p>
        <p>Aknowlegment
This work is supported by a public grant overseen by the French National
Research Agency (ANR) as part of the "Investissements d’Avenir" program
(reference: ANR-10-LABX-0083).</p>
        <p>The authors acknowledge Christophe Parisse, for his advice.</p>
        <p>References
[ACT03]
[Ali10]
[AS95]
[BCCC12]
[Bro73]
[CC08]
[CCP10]
[CM06]
[CPRC09]
[CSL04]</p>
        <sec id="sec-3-1-1">
          <title>RegExpMiner: Automatically discovering frequently matching regular expressions</title>
          <p>Julien Rabatel1, J´erˆome Az´e1, Pascal Poncelet1, and Mathieu Roche1,2</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>1 LIRMM, CNRS UMR 5506, Univ. Montpellier 2, 34095 Montpellier, France</title>
      </sec>
      <sec id="sec-3-3">
        <title>2 UMR TETIS - Cirad, Irstea, AgroParisTech, 34093 Montpellier, France</title>
        <p>Regular expressions (REs) are a very powerful and popular tool to
manipulate string data in a variety of applications. They are used as search templates
to look for the occurrences of a given piece of text in a document, or to define
how a given piece of text should be formatted in order to be valid (e.g., to check
that the value entered in an email field of a Web form is correctly formatted),
or even to help solving more complex NLP tasks [NS07]. Their popularity in
those various application domains arises from several reasons. First, they are
easy to understand and manipulate for common usages, despite their wide
expressiveness and power of abstraction. Second, they are natively usable within
a large variety of programming languages, hence making them suitable to be
integrated into every project addressing text processing tasks. Their usage
often relies on a very limited amount of hand-crafted REs. It is indeed difficult
to automatically obtain the REs matching with a given set of strings for which
no a priori knowledge about their underlying formatting rules is given. Such
an automatic discovery of REs would nonetheless offer some very interesting
prospects. Regular expressions indeed have an interesting abstraction power as
they are able to provide information about how textual content is formatted,
rather than focusing on the actual sequences of characters. Having a more
abstract description space for describing textual content then offers new insights.
For instance, an application scenario consists in data cleaning problems. Given a
database containing some textual content about entities (e.g., addresses, names,
phone numbers, etc.), one may be interested in finding values contained in the
database that are mistakes from the people who entered them. Such typos and
formatting mistakes can easily be highlighted if they result in strings that do
not match the same regular expressions as the majority of the other strings.</p>
        <p>While regular expressions can be seen as interesting descriptors of textual
data for various NLP and machine learning tasks, they are hard to obtain. The
literature does not offer fully relevant solutions when one wishes to
enumerate some REs to describe a given set of strings. Regular Expression learning
[Fer05], for instance, consists in building a single regular expression matching
with a given set of positive string examples. Such approaches typically do not
allow exceptions w.r.t. the set of strings to be matched, hence losing their
interest as soon as input data are noisy. Additionally, only one RE is learned
while one can expect to obtain several REs reflecting the different templates
that co-exist in the data. E.g., one cannot expect all the values of a list of
international ZIP codes to respond to only one template, as each country may use
a different one. Constructing one single RE matching with all of them will
of</p>
        <p>J. Rabatel, J. Aze, P. Poncelet and M. Roche
ten lead to an over-generalization of the underlying templates that would make
the obtained RE irrelevant in practical applications. On the other hand, the
sequence mining literature, when applied to string data, offers the possibility
to discover more various templates via frequent patterns, i.e., data fragments
occurring in a sufficient amount of strings. While this general principle answers
the problems above-mentioned for RE learning approaches, the type of extracted
patterns (e.g., sequential patterns [AS95], episodes [MTV97]) is typically much
less expressive than REs. Some efforts have however been put in allowing the
generalization of sequence elements [PLL+10] but extracted sequential patterns
have little commonality with REs, as they only aim at discovering sequence
elements that are frequently found in the same order.</p>
        <p>
          We propose an approach for extracting regular expressions under the form of
frequent patterns in textual data. To this end, we define a relevant pattern
language that offers some interesting algorithmic properties. While we do not aim at
exploiting all the characteristics and expressiveness of the RE language, we focus
on providing a preliminary approach by keeping some of its main features. In
particular, we fully consider the problem of allowing the generalization of
characters via the use of predefined character classes, commonly used in REs3. Another
aspect that this approach takes into account is the repetition of some
characters in strings. For instance, we assume that the strings “012 ” and “9876543 ”,
should both be generalizable to the RE /[
          <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4 ref47 ref49 ref5 ref52 ref53 ref6 ref60 ref61 ref62 ref63 ref64 ref65 ref66 ref67 ref7 ref8 ref9">0−9</xref>
          ]+/, i.e., a list of consecutive digit
characters, even if they do not contain the same digits nor the same amount
of digits. We define the frequent regular expression pattern mining problem by
providing a theoretical framework linking together the RE and sequence mining
worlds, and highlight some properties that, while inspired from known properties
in sequence mining, are specific to the problem we consider study and employs
them to design the RegExpMiner algorithm to mine such patterns.
References
[AS95] Rakesh Agrawal and Ramakrishnan Srikant. Mining sequential patterns. In
Data Engineering, 1995. Proceedings of the Eleventh International
Conference on, pages 3–14. IEEE, 1995.
[Fer05] Henning Fernau. Algorithms for learning regular expressions. In Algorithmic
        </p>
        <p>Learning Theory, pages 297–311. Springer, 2005.
[MTV97] Heikki Mannila, Hannu Toivonen, and A Inkeri Verkamo. Discovery of
frequent episodes in event sequences. Data Mining and Knowledge Discovery,
1(3):259–289, 1997.
[NS07] David Nadeau and Satoshi Sekine. A survey of named entity recognition and
classification. Lingvisticae Investigationes, 30(1):3–26, 2007.
[PLL+10] Marc Plantevit, Anne Laurent, Dominique Laurent, Maguelonne Teisseire,
and Yeow Wei Choong. Mining multidimensional and multilevel sequential
patterns. ACM Transactions on Knowledge Discovery from Data, 4(1), 2010.</p>
      </sec>
      <sec id="sec-3-4">
        <title>3 Character classes are sets of characters. When tested against a string, a character</title>
        <p>
          class matches with any of the characters it contains. For instance, the character
class [
          <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4 ref47 ref49 ref5 ref52 ref53 ref6 ref60 ref61 ref62 ref63 ref64 ref65 ref66 ref67 ref7 ref8 ref9">0−9</xref>
          ] contains all the digit characters 0, 1, · · · , 9, which allows it to match with
strings such as “3” or “8”, but not with “A”.
        </p>
        <p>NLP-based Feature Extraction for Automated Tweet</p>
        <p>Classification
Anna Stavrianou, Caroline Brun, Tomi Silander, Claude Roux</p>
        <p>Xerox Research Centre Europe, Meylan, France</p>
        <p>Introduction
Traditional NLP techniques cannot alone deal with twitter text that often does not
follow basic syntactic rules. We show that hybrid methods could result in a more
efficient analysis of twitter posts. Tweets regarding politicians have been annotated
with two categories: the opinion polarity and the topic (10 predefined topics). Our
contributions are on automated tweet classification of political tweets.
2</p>
        <p>
          Combination of NLP and Machine Learning Techniques
Initially we used our syntactic parser [
          <xref ref-type="bibr" rid="ref1 ref64">1</xref>
          ] which has given high results on opinion
mining when applied to product reviews [
          <xref ref-type="bibr" rid="ref2 ref65">2</xref>
          ] or the Semeval 2014 Sentiment Analysis
Task [
          <xref ref-type="bibr" rid="ref3 ref66">3</xref>
          ]. However, when applied to Twitter posts, results were not satisfactory. Thus,
we use a hybrid method and combine knowledge given by our parser with learning.
        </p>
        <p>Linguistic information has been extracted from every annotated tweet. We have
used features such as bag of words, bigrams, decomposed hashtags, negation,
opinions, etc. The“liblinear” library (http://www.csie.ntu.edu.tw/~cjlin/liblinear/) was
used to classify tweets. We used logistic regression classifier (with L2-regularization),
where each class c has a separate vector of weights for all the input features. More
formally, , where is the th feature and the is its weight
in class c. When learning the model, we try to find the vectors of weight that
maximize the product of the class probabilities in the training data.</p>
        <p>
          Our objective has been to identify the optimal combination of features that yields
good prediction results, while avoiding overfitting. Some features used are: Snippets:
during annotation, we kept track of the snippets that explained why the annotator
tagged the post with a specific topic or polarity, Hashtags: decomposition techniques
have been applied to hashtags, and they are analyzed by an opinion detection system
that extracts the semantic information they carry [
          <xref ref-type="bibr" rid="ref4 ref67">4</xref>
          ].
        </p>
        <p>We have selected the models using a 10-fold cross validation in the training data
and evaluated them by their accuracy in the test data. For the topic-category task,
(6,142 tweets, 80% used for training), the annotation had &lt;0.4 inter-annotator
agreement, which shows the difficulty of the task. Table 1. shows the results when
NLP features are used, as well as when some semantic merging of classes takes place.</p>
        <p>A. Stavrianou, C. Brun, T. Silander and C. Roux</p>
        <p>NLP features 44.38 29.37</p>
        <p>NLP features + merging 48.91 34.17</p>
        <p>Binary classification was applied to improve the results. We selected the class with
the highest distribution and annotated the dataset with CLASS1 and NOT_CLASS1
tags. We created a model for the prediction of CLASS1, the prediction of CLASS2
and a model for the prediction of the rest of the 8 classes. Merging these models gave
an accuracy of 40.03%, higher than the max accuracy of Table 1.</p>
        <p>For the opinion polarity task (5,754 tweets, 80% used for training), the
interannotator agreement was higher (~ 0.8). As Table 3. shows, we have used not only
NLP features from the tweet but also from the ‘snippet’. The “syntactic analysis” is
the opinion tag given from our opinion analyser.
NLP features (syntactic analysis of opinion)
NLP features of snippet (syntactic analysis)</p>
        <p>As a conclusion, in this paper we provide a model that predicts opinions and topics
for a tweet in the political context. More research around feature analysis will be
carried out. We also plan to add more features yielded by our syntactic analyzer such
as POS tags, or tense. We should also consider a multiple-class labelling.</p>
        <p>This work was partially funded by the project ImagiWeb
ANR-2012-CORD-0023
01.
4</p>
        <p>Acknowledgements</p>
        <p>References
Bancken, Wouter
Bethard, Steven
Boella, Guido
Brun, Caroline
Davis, Jesse
Di Caro, Luigi
Do, Quynh Ngoc Thi
Dupont, Yoann
Gabriel, Alexander
Janssen, Frederik
Paulheim, Heiko
Poncelet, Pascal
Pudi, Vikram
Rabatel, Julien
Rani, Pratibha
Roche, Mathieu
Roux, Claude
1
143</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Agrawal</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , Imielin´ski, T.,
          <string-name>
            <surname>Swami</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Mining Association Rules Between Sets of Items in Large Databases</article-title>
          .
          <source>In: Proc. of SIGMOD</source>
          . pp.
          <fpage>207</fpage>
          -
          <lpage>216</lpage>
          (
          <year>1993</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Banko</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moore</surname>
            ,
            <given-names>R.C.</given-names>
          </string-name>
          :
          <article-title>Part-of-Speech Tagging in Context</article-title>
          .
          <source>In: Proc. of COLING</source>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bharati</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Misra Sharma</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bai</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sangal</surname>
          </string-name>
          , R.:
          <article-title>AnnCorra : Annotating Corpora Guidelines For POS And Chunk Annotation For Indian Languages</article-title>
          .
          <source>Tech. Rep. TRLTRC-31</source>
          , Language Technologies Research Centre,
          <string-name>
            <given-names>IIIT</given-names>
            ,
            <surname>Hyderabad</surname>
          </string-name>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Bhatt</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Narasimhan</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Palmer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rambow</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sharma</surname>
            ,
            <given-names>D.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xia</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>A multi-representational and multi-layered treebank for Hindi/Urdu</article-title>
          . In
          <source>: Proc. of the Third Linguistic Annotation Workshop</source>
          . pp.
          <fpage>186</fpage>
          -
          <lpage>189</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Biemann</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Unsupervised Part-of-Speech Tagging Employing Efficient Graph Clustering</article-title>
          .
          <source>In: Proc. of ACL</source>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Brants</surname>
          </string-name>
          , T.:
          <article-title>TnT: a statistical part-of-speech tagger</article-title>
          .
          <source>In: Proc. of ANLP</source>
          . pp.
          <fpage>224</fpage>
          -
          <lpage>231</lpage>
          (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Brill</surname>
          </string-name>
          , E.:
          <article-title>A Simple Rule-Based Part of Speech Tagger</article-title>
          .
          <source>In: Proc. of ANLP</source>
          . pp.
          <fpage>152</fpage>
          -
          <lpage>155</lpage>
          (
          <year>1992</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Brill</surname>
          </string-name>
          , E.:
          <article-title>Transformation-Based Error-Driven Learning and Natural Language Processing: A Case Study in Part-of-Speech Tagging</article-title>
          .
          <source>Comput. Linguist</source>
          .
          <volume>21</volume>
          (
          <issue>4</issue>
          ),
          <fpage>543</fpage>
          -
          <lpage>565</lpage>
          (
          <year>1995</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Brin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Motwani</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ullman</surname>
            ,
            <given-names>J.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tsur</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Dynamic Itemset Counting and Implication Rules for Market Basket Data</article-title>
          .
          <source>In: Proc. of SIGMOD</source>
          . pp.
          <fpage>255</fpage>
          -
          <lpage>264</lpage>
          (
          <year>1997</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Cutting</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kupiec</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pedersen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sibun</surname>
            ,
            <given-names>P.:</given-names>
          </string-name>
          <article-title>A practical part-of-speech tagger</article-title>
          .
          <source>In: Proc. of the third conference on ANLP</source>
          . pp.
          <fpage>133</fpage>
          -
          <lpage>140</lpage>
          (
          <year>1992</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Dandapat</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sarkar</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Basu</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Automatic Part-of-speech Tagging for Bengali: An Approach for Morphologically Rich Languages in a Poor Resource Scenario</article-title>
          .
          <source>In: Proc. of the 45th Annual Meeting of the ACL on Interactive Poster and Demonstration Sessions</source>
          . pp.
          <fpage>221</fpage>
          -
          <lpage>224</lpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Dubey</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pudi</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Class Based Weighted K-Nearest Neighbor over Imbalance Dataset</article-title>
          .
          <source>In: Proc. of PAKDD (2)</source>
          . pp.
          <fpage>305</fpage>
          -
          <lpage>316</lpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Ekbal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hasanuzzaman</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bandyopadhyay</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Voted Approach for Part of Speech Tagging in Bengali</article-title>
          .
          <source>In: Proc. of PACLIC</source>
          . pp.
          <fpage>120</fpage>
          -
          <lpage>129</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Gadde</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yeleti</surname>
            ,
            <given-names>M.V.</given-names>
          </string-name>
          :
          <article-title>Improving statistical POS tagging using Linguistic feature for Hindi and Telugu</article-title>
          .
          <source>In: Proc. of ICON</source>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Gimenez</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marquez</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Svmtool: A general pos tagger generator based on support vector machines</article-title>
          .
          <source>In: Proc. of LREC</source>
          . pp.
          <fpage>43</fpage>
          -
          <lpage>46</lpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Goldwater</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Griffiths</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>A fully Bayesian approach to unsupervised part-ofspeech tagging</article-title>
          .
          <source>In: Proc. of the 45th Annual Meeting of the ACL</source>
          . pp.
          <fpage>744</fpage>
          -
          <lpage>751</lpage>
          (
          <year>June 2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Ide</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suderman</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>The American National Corpus First Release</article-title>
          .
          <source>In: Proc. of LREC</source>
          . pp.
          <fpage>1681</fpage>
          -
          <lpage>1684</lpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Kamruzzaman</surname>
            ,
            <given-names>S.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haider</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hasan</surname>
            ,
            <given-names>A.R.</given-names>
          </string-name>
          :
          <article-title>Text Classification using Association Rule with a Hybrid Concept of Naive Bayes Classifier and Genetic Algorithm</article-title>
          .
          <source>CoRR abs/1009</source>
          .4976 (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Lafferty</surname>
            ,
            <given-names>J.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McCallum</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pereira</surname>
            ,
            <given-names>F.C.N.</given-names>
          </string-name>
          :
          <article-title>Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data</article-title>
          .
          <source>In: Proc. of ICML</source>
          . pp.
          <fpage>282</fpage>
          -
          <lpage>289</lpage>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          , Han,
          <string-name>
            <given-names>J</given-names>
            .,
            <surname>Pei</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.:</surname>
          </string-name>
          <article-title>CMAR: Accurate and Efficient Classification Based on Multiple Class-Association Rules</article-title>
          .
          <source>In: Proc. of ICDM</source>
          . pp.
          <fpage>369</fpage>
          -
          <lpage>376</lpage>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hsu</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          , Ma, Y.:
          <article-title>Integrating Classification and Association Rule Mining</article-title>
          .
          <source>In: Proc. of KDD</source>
          . pp.
          <fpage>80</fpage>
          -
          <lpage>86</lpage>
          (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>P.V.S.</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.</surname>
          </string-name>
          :
          <article-title>Part Of Speech Tagging Using Conditional Random Fields and Transformation Based Learning</article-title>
          .
          <source>In: Proc. of IJCAI Workshop SPSAL</source>
          . pp.
          <fpage>21</fpage>
          -
          <lpage>24</lpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Shaohong</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guidan</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <source>Research of POS Tagging Rules Mining Algorithm. Applied Mechanics and Materials 347-350</source>
          ,
          <fpage>2836</fpage>
          -
          <lpage>2840</lpage>
          (
          <year>August 2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Shrivastava</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bhattacharyya</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <string-name>
            <surname>Hindi POS Tagger Using Naive Stemming: Harnessing Morphological Information Without Extensive Linguistic</surname>
          </string-name>
          <article-title>Knowledge</article-title>
          .
          <source>In: Proc. of ICON</source>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Søgaard</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Semisupervised condensed nearest neighbor for part-of-speech tagging</article-title>
          .
          <source>In: Proc. of ACL HLT: short papers - Volume</source>
          <volume>2</volume>
          . pp.
          <fpage>48</fpage>
          -
          <lpage>52</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Soni</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vyas</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Using Associative Classifiers For Predictive Analysis In Health Care Data Mining</article-title>
          .
          <source>Int. Journal Of Computer Application</source>
          <volume>4</volume>
          (
          <issue>5</issue>
          ),
          <fpage>33</fpage>
          -
          <lpage>37</lpage>
          (
          <year>July 2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Subramanya</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Petrov</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pereira</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Efficient graph-based semi-supervised learning of structured tagging models</article-title>
          .
          <source>In: Proc. of EMNLP</source>
          . pp.
          <fpage>167</fpage>
          -
          <lpage>176</lpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Thabtah</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>A Review of Associative Classification Mining</article-title>
          .
          <source>The Knowledge Engineering Review</source>
          <volume>22</volume>
          (
          <issue>1</issue>
          ),
          <fpage>37</fpage>
          -
          <lpage>65</lpage>
          (
          <year>Mar 2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Thonangi</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pudi</surname>
          </string-name>
          , V.:
          <article-title>ACME: An Associative Classifier Based on Maximum Entropy Principle</article-title>
          .
          <source>In: Proc. of ALT</source>
          . pp.
          <fpage>122</fpage>
          -
          <lpage>134</lpage>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Toutanova</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klein</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manning</surname>
            ,
            <given-names>C.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Singer</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Feature-rich part-of-speech tagging with a cyclic dependency network</article-title>
          .
          <source>In: Proc. of NAACL HLT'03 - Volume 1</source>
          . pp.
          <fpage>173</fpage>
          -
          <lpage>180</lpage>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31. V.,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          , G.,
          <string-name>
            <surname>S.</surname>
          </string-name>
          , P., S.K.,
          <string-name>
            <surname>S.</surname>
          </string-name>
          , R.:
          <article-title>Tamil POS Tagging using Linear Programming</article-title>
          .
          <source>Int. Journal of Recent Trends in Engineering</source>
          <volume>1</volume>
          (
          <issue>2</issue>
          ),
          <fpage>166</fpage>
          -
          <lpage>169</lpage>
          (May
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          32.
          <string-name>
            <surname>Yarowsky</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>One sense per collocation</article-title>
          .
          <source>In: Proc. of the workshop on Human Language Technology</source>
          . pp.
          <fpage>266</fpage>
          -
          <lpage>271</lpage>
          (
          <year>1993</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          33.
          <string-name>
            <surname>Yin</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          , Han,
          <string-name>
            <surname>J</surname>
          </string-name>
          .: CPAR:
          <article-title>Classification based on Predictive Association Rules</article-title>
          .
          <source>In: Proc. of SDM</source>
          . pp.
          <fpage>331</fpage>
          -
          <lpage>335</lpage>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          34. Za¨ıane,
          <string-name>
            <given-names>O.R.</given-names>
            ,
            <surname>Antonie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.L.</given-names>
            ,
            <surname>Coman</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          :
          <article-title>Mammography Classification By an Association Rule-based Classifier</article-title>
          .
          <source>In: Proc. of MDM/KDD</source>
          . pp.
          <fpage>62</fpage>
          -
          <lpage>69</lpage>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          <string-name>
            <surname>In A</surname>
          </string-name>
          . Abeillé, editor, Treebanks. Kluwer, Dordrecht,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Alishahi</surname>
          </string-name>
          .
          <article-title>Computational modeling of human language acquisition (Synthesis lectures on human language technologies)</article-title>
          . San Rafael: Morgan and
          <string-name>
            <given-names>Claypool</given-names>
            <surname>Publisher</surname>
          </string-name>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          <string-name>
            <given-names>Data</given-names>
            <surname>Engineering</surname>
          </string-name>
          : IEEE,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          <string-name>
            <given-names>N.</given-names>
            <surname>Béchet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Cellier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Charnois</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Crémilleux</surname>
          </string-name>
          .
          <article-title>Discovering linguistic patterns using sequence mining</article-title>
          .
          <source>In proceedings of CICLing'2012</source>
          , pages
          <fpage>154</fpage>
          -
          <lpage>165</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          <string-name>
            <given-names>R. W.</given-names>
            <surname>Brown</surname>
          </string-name>
          .
          <article-title>A first language: the early stages</article-title>
          . Cambridge, Mass. Harvard University Press, Cambridge, Massashusetts,
          <year>1973</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          <string-name>
            <given-names>B.</given-names>
            <surname>Crabbé</surname>
          </string-name>
          and
          <string-name>
            <given-names>M. H.</given-names>
            <surname>Candito</surname>
          </string-name>
          .
          <article-title>Expériences d'analyse syntaxique statistique du français</article-title>
          . In Actes de TALN'
          <volume>08</volume>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          <string-name>
            <given-names>P.</given-names>
            <surname>Cellier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Charnois</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Plantevit</surname>
          </string-name>
          .
          <article-title>Sequential patterns to discover and characterise biological relations</article-title>
          . In A. Gelbukh, editor,
          <source>CICLing</source>
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          <string-name>
            <surname>LNCS</surname>
          </string-name>
          , vol.
          <volume>6008</volume>
          , pages
          <fpage>537</fpage>
          -
          <lpage>548</lpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          <string-name>
            <given-names>N.</given-names>
            <surname>Chater</surname>
          </string-name>
          and
          <string-name>
            <given-names>C. D.</given-names>
            <surname>Manning</surname>
          </string-name>
          .
          <article-title>Probabilistic models of language processing and acquisition</article-title>
          .
          <source>In Trends in Cognitive Science</source>
          ,
          <volume>10</volume>
          (
          <issue>7</issue>
          ), pages
          <fpage>335</fpage>
          -
          <lpage>344</lpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          <string-name>
            <given-names>T.</given-names>
            <surname>Charnois</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Plantevit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Rigotti</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Crémilleux</surname>
          </string-name>
          .
          <article-title>Fouille de données séquentielles pour l'extraction d'information</article-title>
          .
          <source>In Traitement Automatique des Langues</source>
          ,
          <volume>50</volume>
          (
          <issue>3</issue>
          ),
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref45">
        <mixed-citation>
          <string-name>
            <given-names>L.</given-names>
            <surname>Clément</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Sagot</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Lang</surname>
          </string-name>
          .
          <article-title>Morphology based automatic acquisition of large-coverage lexica</article-title>
          .
          <source>In LREC</source>
          <year>2004</year>
          , Lisbonne,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref46">
        <mixed-citation>
          [DL99]
          <string-name>
            <given-names>G.</given-names>
            <surname>Dong</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          .
          <article-title>Efficient mining of emerging patterns: Discovering trends and differences</article-title>
          .
          <source>In Proc. of SIGKDD'99</source>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref47">
        <mixed-citation>
          [Elm01]
          <string-name>
            <given-names>J.</given-names>
            <surname>Elman</surname>
          </string-name>
          .
          <article-title>Connectionism and language acquisition. In Essential readings in language acquisition</article-title>
          .
          <source>In Oxford : Blackwell</source>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref48">
        <mixed-citation>
          [LCY10]
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Lavergne</surname>
          </string-name>
          , Olivier Cappé, and
          <string-name>
            <given-names>François</given-names>
            <surname>Yvon</surname>
          </string-name>
          .
          <article-title>Practical very large scale CRFs</article-title>
          .
          <source>In Proceedings of ACL'2010</source>
          , pages
          <fpage>504</fpage>
          -
          <lpage>513</lpage>
          . Association for Computational Linguistics,
          <year>July 2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref49">
        <mixed-citation>
          <string-name>
            <surname>[LMP01] John D. Lafferty</surname>
          </string-name>
          ,
          <string-name>
            <surname>Andrew McCallum</surname>
          </string-name>
          , and
          <string-name>
            <surname>Fernando</surname>
            <given-names>C. N.</given-names>
          </string-name>
          <string-name>
            <surname>Pereira</surname>
          </string-name>
          .
          <article-title>Conditional random fields: Probabilistic models for segmenting and labeling sequence data</article-title>
          .
          <source>In Proceedings of the Eighteenth International Conference on Machine Learning (ICML)</source>
          , pages
          <fpage>282</fpage>
          -
          <lpage>289</lpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref50">
        <mixed-citation>
          [MC81]
          <string-name>
            <given-names>J. F.</given-names>
            <surname>Miller</surname>
          </string-name>
          and
          <string-name>
            <given-names>R. S.</given-names>
            <surname>Chapman</surname>
          </string-name>
          .
          <article-title>The relation between age and mean length of utterance in morphemes</article-title>
          .
          <source>In Journal of Speech and Hearing Research</source>
          ,
          <volume>24</volume>
          , pages
          <fpage>154</fpage>
          -
          <lpage>161</lpage>
          ,
          <year>1981</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref51">
        <mixed-citation>
          [NAFS13]
          <string-name>
            <given-names>D.</given-names>
            <surname>Nouvel</surname>
          </string-name>
          ,
          <string-name>
            <surname>J-Y. Antoine</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Friburger</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Soulet</surname>
          </string-name>
          . Fouille de rè
          <article-title>- gles d'annotation partielles pour la reconnaissance d'entités nommées</article-title>
          .
          <source>In TALN'13</source>
          , pages
          <fpage>421</fpage>
          -
          <lpage>434</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref52">
        <mixed-citation>
          [NR07]
          <string-name>
            <given-names>M.</given-names>
            <surname>Nanni</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Rigotti</surname>
          </string-name>
          .
          <article-title>Extracting trees of quantitative serial episodes</article-title>
          .
          <source>In Proc. of KDID'07</source>
          , pages
          <fpage>170</fpage>
          -
          <lpage>188</lpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref53">
        <mixed-citation>
          [PHMA+01]
          <string-name>
            <given-names>J.</given-names>
            <surname>Pei</surname>
          </string-name>
          , J. Han,
          <string-name>
            <given-names>B.</given-names>
            <surname>Mortazavi-Asl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Pinto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.</given-names>
            <surname>Dayal</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Hsu</surname>
          </string-name>
          . Prefixspan:
          <article-title>Mining sequential patterns by prefix-projected growth</article-title>
          .
          <source>In ICDE, IEEE Computer Society</source>
          , pages
          <fpage>215</fpage>
          -
          <lpage>224</lpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref54">
        <mixed-citation>
          [QCCL12]
          <string-name>
            <given-names>S.</given-names>
            <surname>Quiniou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Cellier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Charnois</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Legallois</surname>
          </string-name>
          .
          <article-title>Fouille de données pour la stylistique : cas des motifs séquentiels émergents</article-title>
          .
          <source>In Proceedings of the 11th International Conference on the Statistical Analysis of Textual Data, Liege</source>
          , pages
          <fpage>821</fpage>
          -
          <lpage>833</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref55">
        <mixed-citation>
          [SA96]
          <string-name>
            <given-names>R.</given-names>
            <surname>Srikant</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Agrawal</surname>
          </string-name>
          .
          <article-title>Mining sequential patterns: Generalizations and performance improvements</article-title>
          .
          <source>In EDBT 1996. LNCS</source>
          , vol.
          <volume>1057</volume>
          , pages
          <fpage>3</fpage>
          -
          <lpage>17</lpage>
          ,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref56">
        <mixed-citation>
          [Sal86]
          <string-name>
            <given-names>A.</given-names>
            <surname>Salem</surname>
          </string-name>
          .
          <article-title>Segments répétés et analyse statistique des données textuelles</article-title>
          .
          <source>In Histoire &amp; Mesure volume 1 - numéro 2</source>
          , pages
          <fpage>5</fpage>
          -
          <lpage>28</lpage>
          ,
          <year>1986</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref57">
        <mixed-citation>
          [Sch94]
          <string-name>
            <given-names>Helmut</given-names>
            <surname>Schmid</surname>
          </string-name>
          .
          <article-title>Probabilistic part-of-speech tagging using decision trees</article-title>
          .
          <source>In Proceedings of International Conference on New Methods in Language Processing</source>
          , pages
          <fpage>44</fpage>
          -
          <lpage>49</lpage>
          ,
          <year>1994</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref58">
        <mixed-citation>
          [TDE+12]
          <string-name>
            <given-names>I.</given-names>
            <surname>Tellier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Duchier</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Eshkol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Courmet</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Martinet</surname>
          </string-name>
          .
          <article-title>Apprentissage automatique d'un chunker pour le français</article-title>
          . In Actes de TALN'
          <volume>12</volume>
          , papier court (poster),
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref59">
        <mixed-citation>
          [TDEW13]
          <string-name>
            <given-names>I.</given-names>
            <surname>Tellier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Dupont</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Eshkol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and I.</given-names>
            <surname>Wang</surname>
          </string-name>
          .
          <article-title>Adapt a text-oriented chunker for oral data: How much manual effort is necessary?</article-title>
          <source>In The 14th International Conference on Intelligent Data Engineering and Automated Learning (IDEAL'</source>
          <year>2013</year>
          ),
          <article-title>Special Session on Text Data Learning</article-title>
          ,
          <string-name>
            <surname>LNAI</surname>
          </string-name>
          ,
          <source>Hefei (Chine)</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref60">
        <mixed-citation>
          [TTA09]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tsuruoka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tsujii</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Ananiadou</surname>
          </string-name>
          .
          <article-title>Fast full parsing by linear-chain conditional random fields</article-title>
          .
          <source>In Proceedings of EACL 2009</source>
          , pages
          <fpage>790</fpage>
          -
          <lpage>798</lpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref61">
        <mixed-citation>
          [WH04]
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Han</surname>
          </string-name>
          . Bide:
          <article-title>Efficient mining of frequent closed sequences</article-title>
          .
          <source>In ICDE, IEEE Computer Society</source>
          , pages
          <fpage>79</fpage>
          -
          <lpage>90</lpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref62">
        <mixed-citation>
          [YHA03]
          <string-name>
            <given-names>X.</given-names>
            <surname>Yan</surname>
          </string-name>
          , J. Han, and
          <string-name>
            <given-names>R.</given-names>
            <surname>Afshar</surname>
          </string-name>
          .
          <article-title>Mining closed sequential patterns in large databases</article-title>
          .
          <source>In SDM SIAM</source>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref63">
        <mixed-citation>
          <string-name>
            <surname>[Zak01] M. J. Zaki</surname>
          </string-name>
          . Spade:
          <article-title>An efficient algorithm for mining frequent sequences</article-title>
          .
          <source>In Machine Learning Journal</source>
          <volume>42</volume>
          (
          <issue>1</issue>
          /2), pages
          <fpage>31</fpage>
          -
          <lpage>60</lpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref64">
        <mixed-citation>
          1.
          <string-name>
            <surname>Ait-Mokthar</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chanod</surname>
            ,
            <given-names>J.P.</given-names>
          </string-name>
          :
          <article-title>Robustness beyond Shallowness: Incremental Dependency Parsing</article-title>
          .
          <source>NLE Journal</source>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref65">
        <mixed-citation>
          2.
          <string-name>
            <surname>Brun</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Learning opinionated patterns for contextual opinion detection</article-title>
          .
          <source>COLING</source>
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref66">
        <mixed-citation>
          3.
          <string-name>
            <surname>Brun</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Popa</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roux</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>XRCE: Hybrid Classification for Aspect-based Sentiment Analysis</article-title>
          .
          <source>In International Workshop on Semantic Evaluation (SemEval)</source>
          ,
          <year>2014</year>
          (to appear).
        </mixed-citation>
      </ref>
      <ref id="ref67">
        <mixed-citation>
          4.
          <string-name>
            <surname>Brun</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roux</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Décomposition des « hash tags » pour l'amélioration de la classification en polarité des « tweets »</article-title>
          .
          <source>In TALN, July</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>