<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Grammatical Feature Engineering for fine-grained IR tasks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Danilo Croce</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roberto Basili</string-name>
          <email>basilig@info.uniroma2.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Enterprise Engineering University of Roma</institution>
          ,
          <addr-line>Tor Vergata</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>Information Retrieval tasks include nowadays more and more complex information in order to face contemporary challenges such as Opinion Mining (OM) or Question Answering (QA). These are examples of tasks where complex linguistic information is required for reasonable performances on realistic data sets. As natural language learning is usually applied to these tasks, rich structures, such as parse trees, are critical as they require complex resources and accurate pre-processing. In this paper, we show how good quality language learning methods can be applied to the above tasks by using grammatical representations simpler than parse trees. These features are here shown to achieve the state-of-art accuracy in different IR tasks, such as OM and QA.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>hints for modeling different semantic inferences, such as in document topical
classification, predicate and role recognition in sentences as well as question classification in
Question Answering. Lexical features here include lemmas, multiword expressions or
Named Entities that can be directly observed in the texts. Features are then
generalized into predictive components in the final model, induced from the training examples.
Obviously, lexical information usually implies different words to provide different
contributions but usually neglect other crucial linguistic properties, such as word ordering.</p>
      <p>
        The information about the sentence syntactic structure can be thus exploited and
symbolic expressions derived from the parse trees of training examples are used as
features for language learning systems. These features denote the position and the
relationship between words that can be seemingly realized by different trees independently
from irrelevant differences. For example, in a declarative sentence (such as in a S NP
VP structure), the relationship between a verbal predicate (VP) and its immediately
preceding grammatical subject (NP) is literally translated in the feature VP"VP"S#NP,
where arrows indicate upward or downward movements through the tree. Linear
kernels over the resulting Parse Tree Path features are employed in NLP tasks such as for
Semantic Role Labeling [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] or Opinion Mining [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. This idea is further expanded in
tree kernels, introduced by [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. These model similarity between training examples as a
function of the shared subtrees in their corresponding parses. Tree kernels have been
successfully applied to different tasks ranging from parsing [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] to semantic role
labeling [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. Tree kernels are known to determine a better grammatical representation for
the targeted examples and provide an implicit method for robust feature engineering.
      </p>
      <p>
        However, the adoption of grammatical features and tree kernels is still affected by
significant drawbacks. First, strict requirements exist in terms of the size of the
training data set as high dimensionality spaces are generated, whose data sparseness can be
prohibitive. Usually, the application of exact learning algorithms gives rise to complex
training processes whose convergence is quite slow. Although specific forms of
optimization have been proposed to limit their inherent complexity (e.g. [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]), tree kernels
do not scale well over very large training data sets. Finally it must be noticed that most
of the methods extracting grammatical features from parse trees, are strongly biased by
parsing errors.
      </p>
      <p>
        We want to explore here a possible solution to the above problems through the
adoption of shallow but more consistent grammatical features that avoid the use of a full
parser in semantic tasks. Parsing accuracy is highly varying across corpora, and it is
often poorly effective for some natural languages or application domains where limited
resources are available or the syntactic structure of the test instances is very different
with respect to the training material. In particular [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] investigates the accuracy loss of
well known syntactic parsers applied to micro-blogging datasets. In particular they
observed a drastic drop in performance moving from the in-domain test set to the new
Twitter dataset. Avoiding the adoption of full parsing obviously increases the number
and nature of possible uses of language technologies in a variety of complex NLP
applications. In IR, part of speech information has been generally used for stemming,
generating stop-word lists, and identifying pertinent terms or phrases in documents and/or in
queries. Generally, the state of the art in IR systems tend to benefit from the adoption
of parts of speech to index or retrieve information [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ].
      </p>
      <p>The open research questions are: which shallow grammatical representation is
suitable to support the learning of fine-grained semantic models? Which grammatical
generalizations can be usefully achieved over shallow syntactic representations for
sentencebased inferences?</p>
      <p>In the rest of this work, we show how embedding shallow grammatical information
in a sentence representation, as a special case of enriched lexical information, produces
useful generalizations in standard machine learning settings. Empirical findings in
support to this thesis are discussed against two complex sentence-based semantic tasks, i.e.
question classification and sentiment analysis in micro-blogging.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Shallow Parsing and Grammatical Feature engineering</title>
      <p>Grammatical feature engineering is required as lexical information alone is, in general,
not sufficient to characterize linguistic generalizations useful for fine-grained semantic
inferences. For example, sentence (3) is the appropriate answer for the question (1),
although both sentences (2) and (3) are reasonable candidates.</p>
      <p>The grapes which produce the Cognac grow in the province and the French government ... (2)</p>
      <p>What French province is Cognac produced in? (1)</p>
      <p>Cognac is a brandy produced in Poitou-Charentes. (3)</p>
      <p>
        Suppose we use a lexical overlap rule for a Question Answering (QA) task: given
the overlapping terms outlined in bold1, it would result in the wrong answer (2). A
simple lexical overlap model is too simplistic, as syntactic information characterizing
the individual sentences (1) and (3) is here necessary. Syntactic features provide more
information to estimate the similarity between the question and the candidate answers,
as in general explored by tree kernels in Answer Classification/Re-ranking [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. The
parse tree in Figure 1 corresponds to sentence (3) and represents:
– lexical information through its terminal nodes (e.g., words as Cognac; is; : : : )
– Coarse-grained grammatical information through the POS tag characterizing
preterminal nodes (e.g. N N P or V BZ)
– Fine-grained grammatical information as subtrees correspond to the production
rules of the underlying context free grammar (CFG).
      </p>
      <p>
        Examples of the CFG rules involved in Figure 1 are: S ! N P V P , V P !
V BZ N P , N P ! N P P or N P ! DT N N . Stochastic context free grammars
(e.g. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]), are generative models for parse trees, seen as complex joint events, whose
overall probability depends on the individual CFG rules (i.e., subtrees), and lexical
information as well. Our aim here is to acquire these rules implicitly, as a side effect of
the learning for semantic inference process. Specific features can in fact be designed
to surrogate the syntactic structures of the parse tree, implicitly. Observable POS tag
sequences correspond to subtrees and can be considered their shallow counterpart.
1 Sentence (2) shares five terms with the sentence (1), while (3) shares only four terms.
      </p>
      <p>
        They express linearly special properties, in analogy with the Parse Tree Paths in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
In other words, subtrees can be artificially replaced introducing POS tag sequences (or
POS n-grams), instead of parse tree fragments. The idea is that the syntactic structure
of a sentence could be surrogated as the POS n-grams, instead of the set of possible
syntactic tree fragments, as used by tree kernels. For example, the partial tree expressed
by VP!VBN PP in Fig. 1 can be represented through the pseudo token given by
VBNIN-NNP.
      </p>
      <p>NP
NNP
Cognac</p>
      <p>S
VBZ
is</p>
      <p>VP
NP</p>
      <p>NP
DT
a</p>
      <p>NN
brandy</p>
      <p>VBN
produced</p>
      <p>VP</p>
      <p>IN
in</p>
      <p>PP</p>
      <p>NP</p>
      <p>NNP
Poitou-Charentes</p>
      <p>Lexicalized features (i.e., true words) as well as shallow syntactic information (i.e.,
the POS n-grams) are thus made available as flat features, thus constraining the
capacity of the underlying learning machine. A sentence s of length jsj is thus represented
as a set of words (in a bag-of-word fashion), extended by the pseudo tokens
defining the corresponding POS tag sequences whose length is smaller that n (n-POS tag
grams). Given the word sequence s = fw1; : : : ; wjsjg whose corresponding
part-ofspeeches are fpos1; : : : ; posjsjg, the representation of the pseudo tokens is the set of
pairs f(w1:pos1); : : : ; (wjsj:posjsjg, where each lemmatized word is coupled with its
POS tag.</p>
      <p>Moreover, in order to capture syntactic structures of interest, POS tags are also
mapped into pseudo-tokens expressing their sequences (i.e., POS n-grams). Given n
as the maximal size of the extracted sequences, every subsequence of length at most n
is mapped into a pseudo-token. These novel grammatical tokens of length are
expressed as fpj ; : : : ; pj+ g where = 1; :::; n. In these patterns the representation of
prepositions (POS tag IN) is made explicit. Every position k 2 [j; j + ] for which
posk=IN is represented through wk itself, so that at-NP or of -DT-NN are obtained as
pseudo-tokens for fragments such as “at Whitlock” or “of the vineyard”. The
representation of sentence (3) is shown in Table 2, where words (wi:posi) and n-gram tokens
are shown.
unigrams cognac.NNP be.VBZ a.DT brandy.NN produce.VBN in.IN</p>
      <p>poitou-charentes.NNP</p>
      <p>DT-NN-VBN</p>
      <p>NN-VBN-in
2.1</p>
      <sec id="sec-2-1">
        <title>Shallow Syntactic Features for Question Classification</title>
        <p>
          In Question Answering three main processing stages are foreseen: question processing,
document retrieval and answer extraction [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. Question processing is usually centered
around the so called question classification (QC) task that maps a question into one
of k predefined answer classes [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. Typical examples of classes characterize
different answer strategies and range from questions regarding persons or organizations (e.g.
Who killed JFK?) to definition questions (e.g. What is a perceptron?) or modalities (e.g.
How fast does boiling water cool?). Highly accurate QC systems apply supervised
machine learning techniques, e.g. Support Vector Machines (SVMs) [
          <xref ref-type="bibr" rid="ref20 ref23">20, 23</xref>
          ] or the SNoW
model [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ], where questions are encoded using a variety of lexical, syntactic and
semantic features. In [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ], it has been shown that the questions’ syntactic structure contributes
remarkably to the classification accuracy. This task is thus strictly syntax-dependent,
especially because individual sentences are targeted.
        </p>
        <p>As questions can be regarded as individual sentences, we will adopt the feature
extraction scheme proposed in Table 2 for our QC models. These features represent both
lexical and grammatical information that can be efficiently feed a statistical classifier
based on linear kernels. Section 3.1 will discuss comparative experiments with previous
works on Question Classification.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Shallow Syntactic Features for Sentiment Analysis over micro-blogging</title>
        <p>
          Microblogging has been already established as a significant form of electronic
wordof-mouth for sharing opinions, suggestions and consumer reviews concerning ideas,
products or brands. Microblogging is also referred to as micro-sharing or Twittering
(from Twitter2 by far the most popular microblogging application). While opinion
mining over traditional text sources (e.g. movie reviews or forums) has been significantly
studied [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ], sentiment analysis over tweets has a more recent history, [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] or [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. It has
been usually addressed on the basis of only lexical information whereas the syntactic
structure of tweets is often neglected [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ]. In [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ] the linguistic redundancy in Twitter
is investigated and several types of linguistic features are tested in a supervised setting,
showing that tweet syntactic structure does not provide alone a statistically significant
contribution with respect to lexical typed features. The main problem of syntax-driven
2 http://www.twitter.com
approaches over tweets is the quality of the available grammatical information as tweets
are sentences lacking of a proper grammatical structure.
        </p>
        <p>
          Here the modeling through POS n-grams is suitable to overcome these problems,
as it provides a simpler representation of the tweets’ syntax and, on the other hand,
it should be more robust as for tagging accuracy. However even POS taggers, trained
over standard texts, may be inadequate, as the linguistic form of tweets is rather non
standard with a large use of jargon and shortcuts. An interesting finding in [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] was that
one of the main cause of the syntactic parsing errors over the Twitter dataset is due
to the propagation of part-of-speech tagging errors. In line with other works (see for
example [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] or [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]), we propose to pre-process tweets before a standard POS tagger
is applied. This avoids the noise in applying traditional POS tagging to odd symbols
(e.g. re-tweets or emoticons) or jargon expressions and also reduces data sparseness, as
canonical forms are adopted. The following set of actions is applied before training:
– fully capitalized words are first converted in their lowercase counterpart, i.e. ”DOG”
into ”dog”, before applying POS tagging
– reply marks (i.e. @user name) are replaced with the pseudo-token USER whose
        </p>
        <p>POS tag is set back to PUSER after POS tagging
– hyperlinks are replaced by the token LINK whose POS is PLINK
– hash tags (i.e. #thread name) are replaced by the pseudo-token THREAD whose</p>
        <p>POS is imposed to PTHREAD
– repeated letters and punctuation characters (e.g. looove, loooove or !!!) are cleansed
as they cause high levels of lexical data sparseness. Characters occurring more than
twice are all replaced with a double occurrence expression, so that looove or !!! are
mapped into loove or !!, respectively
– all emoticons, e.g. :-) or :P, are used as sentence separators although they are
systematically misinterpreted by a standard POS tagger. Accordingly, they are first
replaced with a full stop ”.” and then recovered at their original form after POS
tagging. Their POS is always set to SMILE.</p>
        <p>After the above pre-processing phase, a tweet like @jdoe I looove Twitter! :-)
http://twitpic.com/2y2e0 can be represented according to the model proposed in
Section 2. Here the lists of lexical unigrams and grammatical n-grams are reported:
USER.PUSER i.PRP loove.VBP twitter.NNP!.PUNC :-).SMILE LINK.PLINK</p>
        <p>PUSER PRP PRP VBP VBP NNP NNP PUNC PUNC SMILE SMILE PLINK USER PRP VBP ...</p>
        <p>As it is clear from the example, the resulting POS sequences are able to better
capture the intended syntax and act as good models of relevant grammatical relations:
the sequence USER.PUSER i.PRP loove.VBP :::, for example, is a good hint for
the positive bias introduced by loove as a verb.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Performance Evaluation</title>
      <p>
        In this section we evaluate the use of POS n-grams in two applications previously
discussed as standard example of different semantic inferences useful for IR. In all the
experiments POS tagging is carried out by the tagger available in the LTH parser [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
The performance achievable by POS n-grams is thus compared with the one derived by
richer grammatical representations based on parse trees.
3.1
      </p>
      <sec id="sec-3-1">
        <title>Question Classification Results</title>
        <p>
          This first experiment studies the impact of combining lexical and shallow syntactic
information (i.e. POS n-grams), on question classification. The targeted dataset is the
UIUC corpus, largely adopted for benchmarking [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. UIUC contains a training set of
5,452 questions and a test set of 500 questions, both extracted from TREC. Question
classes are organized in two levels of granularity. At the first level, 6 coarse-grained
classes are defined, like ABBREVIATION, ENTITY, DESCRIPTION. A second level
explodes the first level classes into a set of 50 fine-grained sub-classes, e.g., Plant and
Food are subclasses of the ENTITY category.
        </p>
        <p>SVM learning is applied over the feature vectors discussed in Section 2.1 and
multiclassification is modeled through a one-vs-all scheme. The quality of classification is
measured through accuracy, i.e. the percentage of questions associated with the
correct class. A development set is derived from the 20% of the training material. In the
experiments two sentence models are compared:
– POS tagged Unigrams (PU): a question is mapped into a bag of POS tagged
lemmas, i.e. into pairs of (lemma:pos). This model is based only on lexical
information.
– POS n-grams (PnG): each question is modeled by augmenting the P U model
through the shallow syntactic information provided by the sequence of n-grams of
POS tags, with n &lt; 4. The POS of W h-determiners and prepositions are replaced
in the individual POS n-grams by the corresponding lemmas.</p>
        <p>
          In this evaluation the voted perceptron [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] and SM V light [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] have been both
applied3 . Results, compared with the results achieved by the system discussed in [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ] on
the same UIUC dataset, are shown in Table 2. The authors combine a kernel classifier
based on BOW with two semantic kernels: one (i.e. K(LS)) is based on Latent Semantic
Indexing applied to Wikipedia, and the other (i.e. K(semRel)) uses semantic
information acquired through manually constructed lists of words, i.e., a task-specific lexicon
related to the answer types.
        </p>
        <p>
          In the coarse-grained test, i.e. the question classification with respect to the 6 coarse
grained classes, Table 2 shows how the syntactic generalization supported by the P nG
model achieves the best known results on the UIUC dataset, i.e., 91.8% that correspond
to the accuracy reported by a tree kernel approach [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ], without any semantic extension.
This improves the best results of [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ] (i.e., the K(BOW ) + K(LS) + K(semRel))
that refer to a task-dependent use of manually annotated resources. Note how the
kernel K(LS) that uses only lexical information, gathered by an external corpus like
Wikipedia [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ] is also weaker than the P nG model, that makes no use of trees or other
3 In the experiments a polynomial kernel of degree 2 has been applied with SM V light, as it
achieved the best result on the development set
resources. The results in Table 2 are also remarkable from a computational point of
view: the P nG method only requires POS tagged sentences and no parsing. Moreover,
the training time of tree kernel based SVMs on benchmarking data sets are in the order
of hours or days for large training collections (e.g., Prop Bank, as reported in [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]).
        </p>
        <p>
          In [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] an extension to the tree kernel formulation has been proposed, i.e. the
semantic Smoothed Partial Tree Kernel that enriches the similarity among syntactic tree
structures with lexical information gathered by en external corpus, in line with the
K(LS) described in [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ]. State-of-the art results of 94.8% have been obtained in the
coarse-grained test. However it is still a complex approach that need explicit syntactic
parsing of the sentences and an external corpus that provides lexical knowledge. This
is beyond the scope of this work, that aims at providing an efficient and practical
engineering method for natural language learning systems. The training complexity of the
proposed models is very low. Consider that for a short sentence (i.e. a question or a
micro-blogging message) the number of feature is reduced. For example a sentence of
10 words, will generate 10 lexical, 9 bi-gram, 8 three-gram and 7 four-gram features,
i.e. a feature vector of 34 features. It generates a hi-dimensional but very sparse space,
where both SVM and the vote perceptron algorithms can very effectively find a
solution. The efficiency of the proposed method in the QC task is thus proved, as the PnG
model has been trained over 5,452 examples in less than 2 minutes and 40 seconds, with
SM V light and the voted perceptron, respectively.
3.2
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Sentiment Analysis Results</title>
        <p>
          The POS n-grams model has been also applied in the task of Sentiment Analysis over
tweets, as introduced in Section 2.2. The goal here is to classify a tweet according to
its sentiment polarity. The adopted dataset is Twitter Sentiment, released by [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]4, as
other studies (e.g. [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]) do not allow a full comparative analysis. It provides a training
4 http://www.stanford.edu/ alecmgo/cs224n/twitterdata.2009.05.25.c.zip
set automatically generated by selecting the positive (or negative) examples from the
tweets containing positive (or negative) emoticons, e.g. :-) (or :-( ). The test set,
also made available by [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ], includes 183 tweets, manually annotated according to their
binary sentiment polarity, i.e. 1. Each tweet is modeled as a feature vector, including
words as well as the pseudo-tokens generated in the pre-processing phase, including the
resulting POS n-grams (see Section 2.2). SM V light has been applied, with a 50-50%
train-development splitting: in this setting a linear kernel provided the best results.
        </p>
        <p>
          As Table 3 suggests, the results improve on [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ], as the adopted grammatical
information is helpful. The test set employed in our experiments is slightly more complex,
as the Unigrams model achieves a significantly worse result than in [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. Moreover,
without pre-processing, POS tags are inaccurate and this reflects in the lower
performances of the Noisy POS 4-grams model. Our approach achieves a new state-of-art
(i.e. 83.61%) on the dataset. This results due to the grammatical information provided
by the POS n-grams and the contribution of the proposed pre-processing method is
crucial. When no pre-processing is applied, the noise introduce by the POS-tagger would
produce a consistent performance reduction, i.e. 77.59% vs 82.51%. Error analysis
suggests that mistakes (e.g. the positive polarity given to the tweet ”Kobe is the best
in the world not Lebron”) are due to lack of information. If LeBron James (and not
Kobe) is the focus then the polarity is negative. But the alternative decision would have
been perfectly acceptable, otherwise. Figure 2 reports the learning curve for the system
with and without POS n-grams: POS n-grams are responsible of a faster convergence
to higher accuracy levels.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusions</title>
      <p>
        In this paper shallow grammatical features as sequences of POS tags (i.e. POS n-grams)
are proposed as a robust and effective model of grammatical information in
different semantic tasks. Every experiment shows that state-of-the-art results are achieved
or closely approximated by our modeling. Although standard training algorithms are
here adopted, simple kernels over POS n-grams are quite effective, as for example the
sentiment analysis tests demonstrate. Surprisingly, in Question Classification our model
equals the accuracy of a performant tree kernel. The training complexity of the proposed
models is very low. Although several optimization methods for tree kernel learning have
been proposed (e.g. [
        <xref ref-type="bibr" rid="ref18 ref6">6, 18</xref>
        ]), our simpler approach is more applicable by posing much
weaker requirements in terms of quality and size of the annotated datasets. This makes
the proposed technology quite appealing for complex NLP and IR applications, such as
the treatment of noisy sources that current micro-blogging trends require. This is also
shown by the performances observed in the tweet sentiment analysis task, for which
state-of-the-art results are obtained.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Baeza-Yates</surname>
            ,
            <given-names>R.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ribeiro-Neto</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Modern Information Retrieval</article-title>
          .
          <string-name>
            <surname>Addison-Wesley Longman</surname>
          </string-name>
          Publishing Co., Inc., Boston, MA, USA (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Barbosa</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Feng</surname>
          </string-name>
          , J.:
          <article-title>Robust sentiment detection on twitter from biased and noisy data</article-title>
          .
          <source>In: Coling</source>
          <year>2010</year>
          : Posters. pp.
          <fpage>36</fpage>
          -
          <lpage>44</lpage>
          .
          <article-title>Coling 2010 Organizing Committee</article-title>
          , Beijing, China (
          <year>August 2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bilotti</surname>
            ,
            <given-names>M.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elsas</surname>
            ,
            <given-names>J.L.</given-names>
          </string-name>
          , Carbonell, J.,
          <string-name>
            <surname>Nyberg</surname>
          </string-name>
          , E.:
          <article-title>Rank learning for factoid question answering with linguistic and semantic constraints</article-title>
          .
          <source>In: Proceedings of ACM CIKM</source>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4. Collins,
          <string-name>
            <surname>M.</surname>
          </string-name>
          :
          <article-title>Three generative, lexicalised models for statistical parsing</article-title>
          .
          <source>In: Proceedings of ACL 1997</source>
          . pp.
          <fpage>16</fpage>
          -
          <lpage>23</lpage>
          (
          <year>1997</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5. Collins,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Duffy</surname>
          </string-name>
          , N.:
          <article-title>Convolution kernels for natural language</article-title>
          .
          <source>In: Proceedings of Neural Information Processing Systems (NIPS)</source>
          . pp.
          <fpage>625</fpage>
          -
          <lpage>632</lpage>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Croce</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moschitti</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Basili</surname>
          </string-name>
          , R.:
          <article-title>Structured lexical similarity via convolution kernels on dependency trees</article-title>
          .
          <source>In: Proceedings of EMNLP. Edinburgh</source>
          , Scotland, UK. (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Foster</surname>
          </string-name>
          , J., O¨ zlem C¸
          <article-title>etinog˘lu,</article-title>
          <string-name>
            <surname>Wagner</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roux</surname>
            ,
            <given-names>J.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hogan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nivre</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hogan</surname>
          </string-name>
          , D., van Genabith, J.:
          <article-title>#hardtoparse: Pos tagging and parsing the twitterverse</article-title>
          .
          <source>In: Prooceedings of AAAI-11 Workshop on Analysing Microtext</source>
          . San Francisco, CA (
          <year>August 2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Freund</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schapire</surname>
            ,
            <given-names>R.E.</given-names>
          </string-name>
          :
          <article-title>Large margin classification using the perceptron algorithm</article-title>
          .
          <source>Machine Learning Journal</source>
          <volume>37</volume>
          (
          <issue>3</issue>
          ),
          <fpage>277</fpage>
          -
          <lpage>296</lpage>
          (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Gildea</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jurafsky</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Automatic Labeling of Semantic Roles</article-title>
          .
          <source>Computational Linguistics</source>
          <volume>28</volume>
          (
          <issue>3</issue>
          ),
          <fpage>245</fpage>
          -
          <lpage>288</lpage>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Go</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bhayani</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Twitter Sentiment Classification using Distant Supervision</article-title>
          .
          <source>In: CS224N Project Report</source>
          , Stanford (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Joachims</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Text categorization with support vector machines: Learning with many relevant features</article-title>
          .
          <source>In: In Proceedings of the European Conference on Machine Learning</source>
          (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Johansson</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moschitti</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Extracting opinion expressions and their polarities - exploration of pipelines and joint models</article-title>
          .
          <source>In: Proceedings of ACL-HLT. Portland</source>
          , Oregon, USA (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Johansson</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nugues</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Dependency-based syntactic-semantic analysis with propbank and nombank</article-title>
          .
          <source>In: Proceedings of CoNLL-2008</source>
          . Manchester,
          <source>UK (August</source>
          <volume>16</volume>
          -17
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Johansson</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nugues</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>The effect of syntactic representation on semantic role labeling</article-title>
          .
          <source>In: Proceedings of COLING. Manchester, UK (August</source>
          <volume>18</volume>
          -22
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Kaufmann</surname>
          </string-name>
          , J.,
          <string-name>
            <surname>Kalita</surname>
          </string-name>
          , J.:
          <article-title>Syntactic normalization of twitter messages</article-title>
          .
          <source>In: International Conference on Natural Language Processing</source>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Kwok</surname>
            ,
            <given-names>C.C.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Etzioni</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weld</surname>
            ,
            <given-names>D.S.</given-names>
          </string-name>
          :
          <article-title>Scaling question answering to the web</article-title>
          .
          <source>In: WWW</source>
          . pp.
          <fpage>150</fpage>
          -
          <lpage>161</lpage>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roth</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Learning question classifiers</article-title>
          .
          <source>In: Proceedings of ACL'02</source>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Moschitti</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Efficient convolution kernels for dependency and constituent syntactic trees</article-title>
          .
          <source>In: ECML</source>
          . pp.
          <fpage>318</fpage>
          -
          <lpage>329</lpage>
          .
          <source>Machine Learning: ECML 2006, 17th European Conference on Machine Learning, Proceedings</source>
          , Berlin, Germany (
          <year>September 2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Moschitti</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pighin</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Basili</surname>
          </string-name>
          , R.:
          <article-title>Tree kernels for semantic role labeling</article-title>
          .
          <source>Computational Linguistics</source>
          <volume>34</volume>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Moschitti</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Quarteroni</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Basili</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manandhar</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Exploiting syntactic and shallow semantic kernels for question answer classification</article-title>
          .
          <source>In: In Proc. of ACL-07</source>
          . pp.
          <fpage>776</fpage>
          -
          <lpage>783</lpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Moschitti</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Quarteroni</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Basili</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manandhar</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Exploiting syntactic and shallow semantic kernels for question/answer classification</article-title>
          .
          <source>In: Proceedings of ACL'07</source>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Pang</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Opinion mining and sentiment analysis</article-title>
          .
          <source>Foundations and Trends in Information Retrieval</source>
          <volume>2</volume>
          (
          <issue>1-2</issue>
          ),
          <fpage>1</fpage>
          -
          <lpage>135</lpage>
          (
          <year>Jan 2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Tomas</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giuliano</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>A semi-supervised approach to question classification</article-title>
          .
          <source>In: Proceedings of the 17th European Symposium on Artificial Neural Networks</source>
          , Bruges, Belgium (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Voorhees</surname>
            ,
            <given-names>E.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harman</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Overview of the seventh text retrieval conference trec-7</article-title>
          .
          <source>In: Proceedings of the Seventh Text REtrieval Conference (TREC-7</source>
          . pp.
          <fpage>1</fpage>
          -
          <lpage>24</lpage>
          (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Zanzotto</surname>
            ,
            <given-names>F.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pennacchiotti</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tsioutsiouliklis</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Linguistic redundancy in twitter</article-title>
          .
          <source>In: Proceedings of 2011 Conference on Empirical Methods on Natural Language Processing (EmNLP)</source>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>