<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Question Answering using Semantic Annotation</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Lili Aunimo and Reeta Kuuskoski Department of Computer Science University of Helsinki</institution>
          ,
          <addr-line>P.O. Box 68 FIN-00014 UNIVERSITY OF HELSINKI</addr-line>
          ,
          <country country="FI">Finland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Question Answering</institution>
          ,
          <addr-line>Evaluation, Experimentation, Multilingual Information Access</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents a question answering (QA) system called Tikka. Tikka's approach to QA relies heavily on the semantic annotation of text documents and on the usage of answer extraction patterns. In this way, Tikka applies to QA pattern-based techniques traditionally used in named entity recognition and information extraction. In the experiments presented in this paper, Tikka's performance is evaluated in the following tasks: monolingual Finnish and French and bilingual Finnish-English QA. Its performance in the monolingual tasks is near the average when it is compared with the QA systems' performance that participated in the monolingual French task. In the monolingual Finnish task, Tikka was the only participating system.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <sec id="sec-1-1">
        <title>QUESTION ANALYSIS</title>
        <p>GAZETTEER
CLASSIFI−
CATION</p>
        <p>RULES
EXTRACTION
PATTERNS
BILINGUAL
DICTIONARY
SYNTACTIC
PARSER</p>
        <p>SEMANTIC
ANNOTATOR
QUESTION</p>
        <p>CLASSIFIER
TOPIC AND TAR−
GET EXTRACTOR
TRANSLATOR
QUERY
TERMS AND
QUESTION
CLASS,
TOPIC AND
TARGET</p>
      </sec>
      <sec id="sec-1-2">
        <title>ANSWER EXTRACTION</title>
        <p>TEXT
DOCUMENTS
GAZETTEER
PATTERN
PROTOTYPES
DOCUMENT
RETRIEVER
PARAGRAPH
SELECTOR
SEMANTIC
ANNOTATOR</p>
        <p>PATTERN
INSTANTIATOR
AND MATCHER
ANSWER
SELECTOR</p>
      </sec>
      <sec id="sec-1-3">
        <title>ANSWER</title>
        <p>The question analysis component of the QA system consists of five software modules: 1) the
syntactic parser for Finnish, 2) the semantic annotator, which is detailed in Section 4, 3) the
question classifier, 4) the topic and target extractor and 5) the translator, which is described in
the system description of the previous version of Tikka, that participated in QA@CLEF 2004 [1].
All these modules, along with the databases that they use, are illustrated in Figure 1.</p>
        <p>Table 1 shows through an example how question analysis is performed. First, a natural
language question is given as input to the system, for example: D FI EN Mika¨ on WWF? 1. Next, the
Finnish question is parsed syntactically and the French question is annotated semantically. Then
both questions are classified according to the expected answer type, and the topic and target words
are extracted from them. The expected answer types are determined by the multinine corpus, and
they are: LOCATION, MEASURE, ORGANIZATION, OTHER, PERSON and TIME. The
target words are extracted or inferred from the question and they further restrict the answer type,
e.g. age, kilometers and capital city[2].The topic words are words extracted from the question that
in a sentence containing the answer to the question carry old information. For example, in the
question What is WWF?, WWF is the topic because in the answer sentence WWF is the World
Wide Fund for Nature., WWF is the old information and the World Wide Fund for Nature is the
new information. The old and new information of a sentence are contextually established [9]. In
our case, the question is the context. In Tikka, topic words are useful query terms along with the
target words, and they are also used to fill slots in the answer pattern prototypes.</p>
        <p>
          1D stands for a definition question and FI EN means that the source language is Finnish and the target language
is English. In English, the question means What is WWF?
(
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) Parser
(
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) Semantic
        </p>
        <sec id="sec-1-3-1">
          <title>Annotator</title>
          <p>
            (
            <xref ref-type="bibr" rid="ref3">3</xref>
            ) Classifier
(
            <xref ref-type="bibr" rid="ref4">4</xref>
            ) T &amp; T
          </p>
        </sec>
        <sec id="sec-1-3-2">
          <title>Extractor</title>
          <p>
            (
            <xref ref-type="bibr" rid="ref5">5</xref>
            ) FI → EN
WWF
          </p>
        </sec>
        <sec id="sec-1-3-3">
          <title>Example</title>
        </sec>
        <sec id="sec-1-3-4">
          <title>English Finnish</title>
          <p>D FI EN Mik¨a D FI FI Mik¨a on WWF?
on WWF ?
1 Mik¨a mik¨a subj:&gt;2 &amp;NH PRON SG NOM
2 on olla main:&gt;0 &amp;+MV V ACT IND PRES SG3
3 WWF wwf &amp;NH N</p>
          <p>N/A</p>
        </sec>
        <sec id="sec-1-3-5">
          <title>French</title>
          <p>D FR FR Qu’est-ce que
la WWF?
N/A
Qu’est-ce que &lt;organization&gt;
la WWF&lt;/organization&gt;?
Organization
Topic: WWF
Target: N/A</p>
          <p>N/A
The answer extraction component consists of five software modules: 1) the document retriever, 2)
the paragraph selector, 3) the semantic annotator, 4) the pattern instantiator and matcher and
5) the answer selector. All these modules, along with the databases that they use, are illustrated
in Figure 1. The dotted arrows that go from the document retriever back to itself as well as from
the answer selector back to the document retriever illustrate that if no documents or answers are
found, answer extraction starts all over.
3.1</p>
          <p>An Example</p>
        </sec>
        <sec id="sec-1-3-6">
          <title>Module</title>
          <p>
            (
            <xref ref-type="bibr" rid="ref1">1</xref>
            ) Document
retriever
(
            <xref ref-type="bibr" rid="ref2">2</xref>
            ) Paragraph
selector
(
            <xref ref-type="bibr" rid="ref3">3</xref>
            ) Semantic
          </p>
        </sec>
        <sec id="sec-1-3-7">
          <title>Annotator</title>
          <p>
            (
            <xref ref-type="bibr" rid="ref4">4</xref>
            ) Pattern
I &amp; M
(
            <xref ref-type="bibr" rid="ref5">5</xref>
            ) Answer
selector
          </p>
        </sec>
        <sec id="sec-1-3-8">
          <title>English</title>
          <p>22 docs retrieved
22 docs inspected
70 paragraphs
selected</p>
        </sec>
        <sec id="sec-1-3-9">
          <title>Finnish French</title>
          <p>Query terms: WWF, Topic: WWF, Target: N/A
76 docs retrieved, 313 docs retrieved,
30 docs inspected 10 docs inspected
99 paragraphs 39 paragraphs
selected selected</p>
        </sec>
        <sec id="sec-1-3-10">
          <title>Example</title>
          <p>See Table 4
0 patterns
0 matches
nothing to
choose from
0 NIL
12 instantiated patterns
match 7 different answers
chooses the answer
with the highest score, 18
0.25 AAMU19950818-000016
Maailman Luonnon S¨a¨ati¨o
18 instantiated patterns
match 4 different answers
chooses the answer
with the highest score, 8
0.75 ATS.940527.0086 le
Fonds mondial pour la nature
((&lt;[a-z]+&gt;[^&lt;&gt;]+&lt;\/[a-z]+&gt; )+)\( (&lt;[a-z]+&gt;)?TOPIC(&lt;\/[a-z]+&gt;)? \)Score:9
((&lt;[a-z]+&gt;[^&lt;&gt;]+&lt;\/[a-z]+&gt; )+)\( (&lt;[a-z]+&gt;)?Wwf(&lt;\/[a-z]+&gt;)? \)Score:9</p>
          <p>The text snippet that matched the above pattern is in Table 4. (The patterns are case
insensitive.) Only at least partly semantically annotated candidates can be extracted. The score of a
unique answer candidate is the sum of the scores of the patterns that extracted the similar answer
instances, or more formally:
score(answer) =</p>
          <p>
            X patternScore(i),
i in A
(
            <xref ref-type="bibr" rid="ref1">1</xref>
            )
where A is the set of similar answers and patternScore(i) is the score of the pattern that has
matched i in text. The confidence value of a non-NIL answer candidate is determined by the
candidate’s score and by the total number of candidates. This is illustrated in Figure 2. For
example, if the total number of candidates is between 1 and 5, and the score of the candidate is
17 or greater, confidence is 1, but if the score of the candidate is between 1 and 16, confidence
is 0.75. If the confidence score 1 is reached, the answer is selected and no further answers are
searched. Otherwise, all paragraphs are searched for answers, and the one with the highest score
is selected. Since the confidence of the answer for Finnish in Table 2 is 0.25, and the score of the
answer is 18, we can deduce that the number of answer candidates is at least 11.
          </p>
          <p>If document retrieval does not return any documents, or no answer is extracted from the
paragraphs, Tikka has several alternative ways in which to proceed, depending on which task
it is performing and how many times document retrieval has been tried. This is illustrated in
Table 5. As can be seen from the figure, if no documents are retrieved in the first iteration, the
parameter settings of the retrieval engine are altered and document retrieval is performed again in
the monolingual Finnish and bilingual Finnish-English tasks. However, in the monolingual French
task the system halts and returns NIL with a confidence of 1 as an answer. In the monolingual
Finnish task, the system halts after the second try, but in the bilingual English-Finnish task,
document retrieval is performed for a third time if either no documents are retrieved or no answer
is found. Alternatively, in the monolingual Finnish and English tasks, if documents are retrieved,
but no answers are found after the first try, document retrieval is tried once more. In all tasks,
the system returns a confidence value of 1 for the NIL answer if no documents are found and a
confidence value of 0 for the NIL answer if documents are found but no answer can be extracted.
3.2</p>
          <p>Document Retrieval
The document retrieval module of Tikka consists of the vector space model [10] based search engine
Lucene 2 and of the document indices for English, Finnish and French newspaper text built using
it. Tikka has one index for the English document collection, two indices for the Finnish document
collection and two for the French document collection. In each of the indices, one newspaper
article forms one document. The English index (enstem) is a stemmed one. It is stemmed using
2http://lucene.apache.org/java/docs/index.html
17
10
5
0</p>
          <p>1
the implementation of Porter’s stemming algorithm [7] included in Lucene. One index (filemma) to
the Finnish collection is created using the lemmatized word forms as index terms. The Connexor’s
parser is used for the lemmatization. The other Finnish index (fistem) consists of stemmed word
forms. The stemming is done by Snowball [8] project’s 3 stemming algorithm for Finnish. A
Snowlball stemmer is also used to create one of the indices for French (frstem). The other French
index (frbase) is built using the words of the documents as such. This index is case-insensitive.</p>
          <p>In the document retrieval phase, Lucene determines the similarity between the query (q) and
the document (d) using the formula presented in Equation 2 [4].</p>
          <p>similarity(q, d) =</p>
          <p>
            X tf (t in d) · idf (t) · boost(t.f ield in d) · lengthN orm(t.f ield in d), (
            <xref ref-type="bibr" rid="ref2">2</xref>
            )
t in q
where tf is the term frequency factor for the term t in the document d, and idf (t) is the inverse
document frequency of the term. The factor boost adds more weight to the terms appearing in a
given field, and it can be set at indexing time. The last factor is a coefficient that normalizes the
score according to the length of the field. After all the scores regarding a single query have been
calculated, they are normalized from the highest score if that score is greater than 1. Since we do
not use the field specific term weighting, the two last terms of the formula can be discarded, and
the formula is equal to calculating the dot product between a query with binary term weights and
a document with tf idf [5] term weights.
          </p>
          <p>Lucene does not use the pure boolean information retrieval (IR) model, but we model the
conjunctive boolean query by requiring all of the query terms to appear in each of the documents
in the result set. This differs from the pure boolean IR model in that the relevance score for each
document is calculated according to Equation 2, and the documents are ordered according to it.
The ordering is important in Tikka. This is what the term boolean means in Figure 5. In the
same figure, the term ranked means a normal Lucene query where all of the query words are not
required to appear in the retrieved documents.
4</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Semantic Annotation</title>
      <p>Semantic annotation is in many ways a similar task to named entity recognition (NER). NER
is commonly done based on preset names lists and patterns [11] or using machine learning
techniques [3]. Our method relies on the first method. The main difference between NER and semantic
annotation is that the first one aims at recognizing proper names whereas the second aims at
recognizing both proper names and common nouns.</p>
      <p>In Tikka, French questions and selected paragraphs from the search results are annotated
semantically. We have 14 semantic classes that are listed in table 3. To the classes consisting
mainly of proper nouns, some common nouns are added in order to be able to analyze the questions
correctly. For instance, in the gazetteers of the class organization, there are proper names denoting
companies (IBM, Toyota), but also some common nouns referring to organizations in each language
(school, bank, union). As can be seen from Table 3, the organization gazetteer in English is
significantly shorter than those in other two languages. This is due to the NER from Connexor
that is used in addition to our own semantic annotator.</p>
      <p>Class
person
country
language
nationality
capital
location
organization</p>
      <p>The semantic annotator uses a window of two words for identifying the items to be annotated.
In that way we can only find the entities consisting of one or two words. The external NER that
is used in the English annotation is able to identify person names, organizations and locations.
Hence, there are no limitations on the length of entities on these three classes in English. For
Finnish, we exploit Connexor’s syntactic parser for part of speech recognition to eliminate the
words that are not nouns, adjectives or numerals. For French, the semantic annotator builds
solely on the text as it is without any linguistic analysis.</p>
      <p>In the text to be annotated, persons are identified based on a list of first names and the
subsequent capital word. The subsequent capital words are added to the list of known names in
the document. In this way the family names appearing alone later in the document can also be
identified to be names of a person. The class location consists of names of large cities that are
not capitals and of the names of states and other larger geographical items. To the class measure
belong numerals and numeric expressions, for instance dozen. Unit consists of terms such as
percent, kilometer. The event class is quite heterogeneous, since to it belong terms like Olympics,
Christmas, war and hurricane. Time gazetteer lists time related terms, the names of the months,
week days etc. Example annotations for each of the languages can be seen in Table 4.
5</p>
    </sec>
    <sec id="sec-3">
      <title>Analysis of Results</title>
      <p>Tikka was evaluated by participating in the monolingual Finnish and French tasks and in the
bilingual Finnish-English task. The evaluation results are described in detail in Section 4 (Results)
of the QA track overview paper [6]. In each of the tasks, two different parameter settings (run 1
and run 2) for the document retrieval component were tested. These settings are listed in Table 5.
The results of the runs are shown in Figure 3. We can observe that the difference between runs
is not very big for the French monolingual run. The accuracy of the artificial combination run 4
(C) is not much higher than that of the the French monolingual run 1, which means that almost
4In this case, the artificial combination run represents a run where the system is somehow able to choose for
each question the better answer from the two answers provided by the runs 1 and 2. For more information on the
combination runs, see the track overview paper [6].</p>
      <sec id="sec-3-1">
        <title>Lang.</title>
        <p>English
Finnish
French</p>
      </sec>
      <sec id="sec-3-2">
        <title>Example</title>
        <p>The &lt;organization&gt;World Wide Fund for Nature&lt;/organization&gt; (
&lt;organization&gt;WWF&lt;/organization&gt; ) reported that only &lt;measure&gt;35,000&lt;/measure&gt;
to &lt;measure&gt;50,000&lt;/measure&gt; of the species remained in mainly isolated pockets .
&lt;organization&gt;Maailman Luonnon S¨a¨ati¨o&lt;/organization&gt; ( &lt;ne&gt;WWF&lt;/ne&gt; )
vetoaa kaikkiin &lt;country&gt;Suomen&lt;/country&gt; metsstjiin , ett ei
&lt;person&gt;Toivoa&lt;/person&gt; ja sen perhett ammuttaisi niiden
&lt;unit&gt;matkalla&lt;/unit&gt; toistaiseksi tuntemattomille talvehtimisalueille .
&lt;ne&gt;La&lt;/ne&gt; dcision de &lt;organization&gt;la Commission&lt;/organization&gt; baleini`ere
internationale ( &lt;ne&gt;CBI&lt;/ne&gt; ) de cr´eer un sanctuaire pour les c´etac´es est ”une victoire
historique” , a comment´e &lt;time&gt;vendredi&lt;/time&gt; &lt;ne&gt;le Fonds&lt;/ne&gt;
mondial pour la nature ( &lt;organization&gt;WWF&lt;/organization&gt; )
all answers given by the runs are equal. On the contrary, there is a difference of 4 points between
the accuracies of the two monolingual Finnish runs, and in addition, as the difference between
run 1 and the combination run for Finnish is 3.5 points, we can conclude that some of the correct
answers returned by run 2 are not included in the set of correct answers given by run 1. This
means that the different parameter settings of the runs produced an effect on Tikka’s overall
performance. Between the runs, both the parameters for type of index and the maximum number
of documents were altered. In the bilingual Finnish-English task, some difference between the
runs can be observed, but the difference is not as big as in he monolingual Finnish task.</p>
        <p>
          Accuracy
20%
10%
26.5
C
23.0
1
Tikka is a QA system that uses pattern-based techniques to extract answers from text. In the
experiments presented in this paper, its performance is evaluated in the following tasks: monolingual
Finnish and French and bilingual Finnish-English QA. Its performance in the monolingual tasks is
near the average when it is compared with the QA systems’ that participated in the monolingual
French task. In the monolingual Finnish task, Tikka was the only participating system.
(
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) →
(
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) if no documents →
(
          <xref ref-type="bibr" rid="ref3">3</xref>
          ) if no answers →
Iterations
(
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) →
(
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) if no documents →
(
          <xref ref-type="bibr" rid="ref3">3</xref>
          ) if no answers →
        </p>
        <p>In the future, as the document databases are not very big (about 1.6 GB), the documents could
be annotated semantically before indexing. This would speed up the interactive processing time
and the semantic classes could be used as fields in index creation. In addition, indexing based on
paragraph level instead of document level might raise the ranking of the essential results and it
would speed up the processing time of the interactive phase.
7</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgements</title>
      <p>The authors thank Connexor Ltd 5 for providing the Finnish and English parsers and NE
reconizers, Kielikone Ltd 6 for providing the bilingual dictionary and Juha Makkonen for providing the
question classifier for Finnish.</p>
      <p>5http://www.connexor.com
6http://www.kielikone.fi/en</p>
      <p>Available at</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Lili</given-names>
            <surname>Aunimo</surname>
          </string-name>
          , Reeta Kuuskoski, and
          <string-name>
            <given-names>Juha</given-names>
            <surname>Makkonen</surname>
          </string-name>
          .
          <article-title>Finnish as Source Language in Bilingual Question Answering</article-title>
          . In C. Peters,
          <string-name>
            <given-names>P. D.</given-names>
            <surname>Clough</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. J. F.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kluck</surname>
          </string-name>
          , and B. Magnini, editors,
          <source>Multilingual Information Access for Text, Speech and Images: 5th Workshop of the Cross-Language Evaluation Forum</source>
          ,
          <string-name>
            <surname>CLEF</surname>
          </string-name>
          <year>2004</year>
          ,
          <article-title>Bath</article-title>
          , UK,
          <source>September 15-17</source>
          ,
          <year>2004</year>
          , Revised Selected Papers, volume
          <volume>3491</volume>
          of Lecture Notes in Computer Science. Springer Verlag,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Lili</given-names>
            <surname>Aunimo</surname>
          </string-name>
          , Juha Makkonen, and
          <string-name>
            <given-names>Reeta</given-names>
            <surname>Kuuskoski</surname>
          </string-name>
          .
          <article-title>Cross-language Question Answering for Finnish</article-title>
          . In Eero Hyvo¨nen, Tomi Kauppinen, Mirva Salminen, Kim Viljanen, and
          <string-name>
            <surname>Pekka</surname>
          </string-name>
          Ala-Siuru, editors,
          <source>Proceedings of the 11th Finnish Artificial Intelligence Conference STeP</source>
          <year>2004</year>
          ,
          <article-title>September 1-3</article-title>
          , Vantaa, Finland, volume
          <volume>2</volume>
          of Conference Series - No 20, pages
          <fpage>35</fpage>
          -
          <lpage>49</lpage>
          .
          <source>Finnish Artificial Intelligence Society</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Daniel</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Bikel</surname>
          </string-name>
          , Richard Schwartz, and
          <string-name>
            <surname>Ralph</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Weischedel</surname>
          </string-name>
          .
          <article-title>An algorithm that learns what's in a name</article-title>
          .
          <source>Machine Learning</source>
          ,
          <volume>34</volume>
          (
          <issue>1-3</issue>
          ):
          <fpage>211</fpage>
          -
          <lpage>231</lpage>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Erik</given-names>
            <surname>Hatcher</surname>
          </string-name>
          and Otis Gospodneti´c. Lucene in Action. Manning Publications Co.,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Karen</given-names>
            <surname>Sparck Jones</surname>
          </string-name>
          .
          <article-title>A statistical interpretation os term specificity and its application in retrieval</article-title>
          .
          <source>Journal of Documentation</source>
          ,
          <volume>28</volume>
          (
          <issue>1</issue>
          ):
          <fpage>11</fpage>
          -
          <lpage>21</lpage>
          ,
          <year>1972</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>B.</given-names>
            <surname>Magnini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Vallin</surname>
          </string-name>
          , ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Aunimo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Ayache</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Erbach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Penas</surname>
          </string-name>
          , M. de Rijke,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rocha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Simov</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Sutcliffe</surname>
          </string-name>
          .
          <article-title>Overview of the CLEF 2005 Multilingual Question Answering Track</article-title>
          . In Carol Peters and Francesca Borri, editors,
          <source>Proceedins of the CLEF 2005 Workshop</source>
          , Vienna, Austria, sept
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M. F.</given-names>
            <surname>Porter</surname>
          </string-name>
          .
          <article-title>An algorithm for suffix stripping</article-title>
          .
          <source>Program</source>
          ,
          <volume>14</volume>
          (
          <issue>3</issue>
          ):
          <fpage>130</fpage>
          -
          <lpage>137</lpage>
          ,
          <year>1980</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.F.</given-names>
            <surname>Porter</surname>
          </string-name>
          .
          <article-title>Snowball: A language for stemming algorithms</article-title>
          ,
          <year>2001</year>
          . http://snowball.tartarus.org/texts/introduction.
          <source>html[22.8</source>
          .
          <year>2005</year>
          ].
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>R.</given-names>
            <surname>Quirk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Greenbaum</surname>
          </string-name>
          , G. Leech, and
          <string-name>
            <given-names>J.</given-names>
            <surname>Svartvik</surname>
          </string-name>
          .
          <article-title>A Comprehensive Grammar of the English Language</article-title>
          . Longman,
          <year>1985</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Gerard</given-names>
            <surname>Salton</surname>
          </string-name>
          .
          <source>The SMART Retrieval System: Experiments in Automatic Document Processing. Prentice Hall</source>
          ,
          <year>1971</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Martin</given-names>
            <surname>Volk</surname>
          </string-name>
          and
          <string-name>
            <given-names>Simon</given-names>
            <surname>Clematide. Learn - Filter -</surname>
          </string-name>
          Apply - Forget.
          <article-title>Mixed Approaches to Named Entity Recognition</article-title>
          .
          <source>In Proceedings of the 6th International Workshop of Natural Language for Information Systems</source>
          , Madrid, Spain,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>