<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Information Extraction from the Weather Reports in Serbian</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Staša Vujičić Stanković</string-name>
          <email>stasa@matf.bg.ac.rs</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>BCI'12, September 16-20, 2012, Novi Sad, Serbia.</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vesna Pajić</string-name>
          <email>svesna@agrif.bg.ac.rs</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Copyright © 2012 by the paper's authors. Copying permitted only for private and</institution>
          ,
          <addr-line>academic purposes. This volume is published and copyrighted by its editors., Local Proceedings also appeared in ISBN 978-86-7031-200-5</addr-line>
          ,
          <institution>Faculty of Sciences, University of Novi Sad.</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Belgrade, Faculty of Agriculture</institution>
          ,
          <addr-line>Nemanjina 6, Zemun, Serbia, +381 64 2977630</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Belgrade, Faculty of Mathematics</institution>
          ,
          <addr-line>Studentski trg 16, Belgrade, Serbia, +381 11 202 78 01</addr-line>
        </aff>
      </contrib-group>
      <fpage>105</fpage>
      <lpage>108</lpage>
      <abstract>
        <p>In this paper, we describe a process of extracting information from meteorological texts in Serbian. The text corpus consists of almost 46000 sentences. Having in mind the specifics of Serbian and characteristics of meteorological sublanguage, we develop a classification schema for structuring extracted information and transducers for annotating pieces of information in the text corpus. We describe the transducer for extracting information about daily temperatures and give some evaluation parameters for all other transducers used in the information extraction process.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Information extraction</kwd>
        <kwd>transducers</kwd>
        <kwd>Serbian language</kwd>
        <kwd>language resources</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        Weather forecast reports are interesting for natural language
processing because of their properties and the possibility of
various uses of extracted data. These texts have been studied over
the years in the areas such as information extraction, text mining
or text understanding, and the obtained data were used for
machine translation from one language to another
(TAUMMETEO system developed in Canada for machine translation
from English to French and vice versa [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]), data
visualization described in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], web information extraction using
extraction ontologies represented in [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], creating dialogue
manager system as in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], summarization of data from multiple
sources ([
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]), etc.
      </p>
      <p>In this paper we present the process of extracting information
about weather conditions from meteorological texts in Serbian,
which can be used for different purposes (for example, for
automatic creation of lexicon or annotation of texts). The main
goal of this research was to provide foundations for developing
electronic resources in Serbian, construction of sublanguages,
ontologies, machine translation system from Serbian to English,
and vice versa, and different kinds of linguistic researches in the
domain of weather forecast. Some specifics of Serbian that are
important for this research are presented in Section 2. The corpus
of meteorological texts in Serbian, collected during 2010, 2011,
and 2012 years from several sources is described in Section 3.
The main goal of the extraction process was to annotate
information contained in a text description. Three types of
information were of interest: location, time, and meteorological
phenomena. Semantic classes of information used to structure the
data are described in Section 4.</p>
      <p>
        The process of information extraction is presented in Section 5.
The extraction rules are defined by finite state transducers (FST)
([
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]) and recursive transition networks (RTN) with
output ([
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]), both referred to as transducers in this paper.
We used the UNITEX software system [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] for the creation and
application of the transducers.
      </p>
      <p>Finally, we evaluate the information extraction process and give
the directions for the future research.</p>
    </sec>
    <sec id="sec-2">
      <title>2. THE SPECIFICS OF SERBIAN</title>
      <p>The specific features of Serbian determine, to a great extent,
approach and method that will be used for the information
extraction from texts written in Serbian.</p>
      <p>
        Serbian is a language with rich morphology. For example, the
most adjectives in Serbian may take more than 40 different forms.
There are algorithms for different Natural Language Processing
(NLP) tasks that have excellent results when applied to texts in
English, but very bad when it comes to texts in a language with
rich morphology, such as Serbian. The rich morphological system
of Serbian requires the use of additional linguistic resources, such
as electronic dictionaries and grammars, for text processing. Thus,
it is possible to develop systems for the information extraction
that would be efficient when applied to texts in Serbian.
This paper describes a process of extracting information from
texts in Serbian, in which the electronic dictionary for Serbian
([
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]) was used. This dictionary was written in the DELA
format [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. It contains 125269 lemmas of simple words and
4378245 simple word forms, as well as 5251 lemmas of
compounds and 106731 forms of compounds [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. THE CHARACTERISTICS OF THE</title>
    </sec>
    <sec id="sec-4">
      <title>TEXT CORPUS</title>
      <p>Meteorological texts have been collected during 2010, 2011, and
2012 years from several sources (Republic Hydrometeorological
Service of Serbia,1 the Meteos agency,2 the Politika daily news,3
B92,4 SMedia5 and Internet portal Krstarica6). The created text
corpus contains 13705 text descriptions, which consist of a total
of 45862 sentences.</p>
    </sec>
    <sec id="sec-5">
      <title>3.1 Weather Forecast Sublanguage</title>
      <p>The language used for describing weather conditions in textual
reports is very specific and easily recognizable. A limited set of
words from natural language, which is used to describe the
meteorological phenomenon, can be treated as a sublanguage,
along with its characteristics:
–
–
–
limited vocabulary – the same words are used to describe a
meteorological phenomenon in almost every weather report;
irregular syntax – sentences in meteorological reports
typically do not contain auxiliary verb, and often do not have
a predicate (“Vetar slab, jugoistočni.” – “Wind weak,
southeast.”) or adverbs;
text structure – it is not possible to distinguish different
statements based only on punctuation, since a sentence often
contains multiple statements, and a few sentences sometimes
merges into one separated with commas.</p>
      <p>On the one hand, the existence of such sublanguage facilitates the
text processing, since many syntactic rules are simplified in
comparison to natural language. On the other hand, it is contempt
of natural language syntax rules that prevents the use of existing
electronic grammars, developed and available for a given natural
language.</p>
    </sec>
    <sec id="sec-6">
      <title>3.2 The Structure of Textual Meteorological</title>
    </sec>
    <sec id="sec-7">
      <title>Descriptions</title>
      <p>The descriptions of weather conditions consist of smaller
fragments (sentences and parts of sentences), which carry three
types of information (meteorological phenomenon, location and
time), combined together in a statement. Therefore, every
semantic unit of the text structure (particular statement) can be
treated as a triple &lt;location, time, phenomenon&gt;. The ideal
information extraction process from the following description in
Serbian “Ujutru i pre podne u nižim delovima grada magla ili
sumaglica.“ (“In the morning and before the noon in the lower
parts of the city fog or haze.”) would extracts the following
triples:
&lt;“niži delovi grada”, “ujutru”, “magla ili sumaglica”&gt;
&lt;“niži delovi grada”, “pre podne”, “magla ili sumaglica”&gt;
(&lt;“the lower parts of the city”, “In the morning”, “ fog or haze”&gt;
&lt;“the lower parts of the city”, “before the noon”, “ fog or haze”&gt;)
The statements mutually overlap in the textual descriptions,
usually with no clear boundary between two different statements.
This semantic structure requires a special approach, semantically
oriented, in order to resolve coreferences between different parts.
However, the first steps in this process are the detection and
isolation of the values of individual features. This paper describes
exactly this process, while merging isolated pieces of information
and their values into the statements will be the subject of a future
research.
1 http://www.hidmet.gov.rs
2 http://www.meteos.rs
3 http://www.politika.rs
4 http://www.b92.net
5 http://www.smedia.rs
6 http://www.krstarica.com</p>
    </sec>
    <sec id="sec-8">
      <title>4. SEMANTIC CLASSES FOR</title>
    </sec>
    <sec id="sec-9">
      <title>INFORMATION STRUCTURING</title>
      <p>The information contained in the textual descriptions of weather
conditions, which were of interest in the research, are grouped
into semantic classes of different levels. A semantic class,
together with possible additional classification, should be
assigned to each separate fragment of the text. Hierarchical
classes are shown in Table 1.
(TTeemmppeerraattuurrea) (CKTaaettmeTgpeoemrryap)tuerrea-ture (nmaajvxiišmau,mju,tmaorrnnjiang..….)
(DOTepesmicsrpTiepertmaiotupnr)eer-ature (hcloalddn,wo,atrompelirj,er,ispiongra…s)t ...</p>
      <p>Pojava TipPojave magla, oluja ...
(Phenomenon) (PhenomenonType) (fog, storm...)</p>
      <p>ImeTeritorije Srbija, Evropa, Beograd
n Teritorija (TeritoryName) ...
ito (Teritory) DeoTeritorije severoistok, južni delovi
aco (TeritoryPart) (northeast, southern parts)
L Lokalitet Lokalitet na planinama, lokalno...
(Locality) (Locality) (in the mountains, localy)
Datum 15. januar
(Date) (January 15th)
Dan ImeDana ponedeljak, utorak ..
(Day) (DayName) (Monday, Tuesday...)
ieTm (DDeaoyDPaarnta) (auifjntuettrhrneuo,mopnoo)rnsliengp, oindnthee
sledeće nedelje, tokom
Period Period februara
(Period) (Period) (next week, during
February…)
The names of the features, given in Table 1, are used for
annotating pieces of information in the text.</p>
      <p>The annotations had the following syntax:
&lt;Feature&gt;text segment&lt;/Feature&gt;</p>
      <p>Hence, the example sentence “U većem delu zemlje
promenljivo oblačno, mestimično slaba kiša, pljuskovi i
grmljavina.” (“In most of the country variable cloudiness, with
areas of light rain, showers, and thunder.”), should be annotated
as follows:
&lt;lokalitet&gt;U većem delu zemlje&lt;/ lokalitet&gt;
&lt;obimOblacnosti&gt;promenljivo&lt;/obimOblacnosti&gt;
&lt;prisustvoOblaka&gt;oblačno&lt;/prisustvoOblaka&gt;,
&lt;lokalitet&gt;mestimično&lt;/lokalitet&gt;
&lt;obimPadavina&gt;slaba&lt;/obimPadavina&gt;
&lt;tipPadavina&gt;kiša &lt;/tipPadavina&gt;,
&lt;tipPojave&gt;pljuskovi&lt;/tipPojave&gt; i
&lt;tipPojave&gt;grmljavina&lt;/tipPojave&gt;.
(&lt;Locality&gt;In most of the country &lt;/Locality&gt;
&lt;CloudAmount&gt;variable&lt;/CloudAmount&gt;
&lt;CloudPresence&gt;cloudiness&lt;/CloudPresence&gt;,
&lt;Locality&gt;with areas&lt;/Locality&gt; of
&lt;PrecipitationAmount&gt;light &lt;/PrecipitationAmount&gt;
&lt;PrecipitationType&gt;rain&lt;/PrecipitationType&gt;,
&lt;PhenomenonType&gt;showers&lt;/PhenomenonType&gt; and
&lt;PhenomenonType&gt;thunder&lt;/PhenomenonType&gt;.)</p>
    </sec>
    <sec id="sec-10">
      <title>5. INFORMATION EXTRACTION</title>
    </sec>
    <sec id="sec-11">
      <title>PROCESS</title>
      <p>
        We used transducers (FST and RTN) as extraction rules. The
transducer that describes the rule for extracting particular piece of
information was created for each feature given in Table 1. The
rules were applied through the software system UNITEX, where
the structuring of data was done by annotating text segments that
carry information. The application of transducers was performed
sequentially, one by one. The application order was not important
for the majority of created transducers, although it is possible to
organize the information extraction process so that the successive
application of transducers improves the efficiency of the process
(a cascade of transducers, one operating after the other using the
results of previously applied transducers [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]). In this section, we
will present one of the transducers that extracts information
related to the temperature.
      </p>
      <p>Temperature data have been presented in the texts as values (12
stepeni – 12 degrees, 12°C, 12 C, dva stepena – two degrees,
ispod nule – below zero, minus 5 ...) or descriptive (hladno - cold,
hladnije - colder, toplo - warm, toplije – warmer, pad
temperature – the temperature drop, temperatura u porastu - the
temperature rising ...). For each way of representing temperature,
a special extraction rule has been created. Figure 1 shows the
main transducer (temperatura.grf) in the RTN for extracting
information related to the temperature.</p>
      <p>Subgraph calls are marked with gray colour. Subgraph
vrednost.grf recognizes different expressions for the specific
value (number of degrees) of the temperature. This subgraph is
shown in Figure 2.</p>
      <p>The lexical mask &lt;NB&gt; recognizes successive digits. The lexical
mask &lt;NUM&gt; recognizes all the words in the dictionary that are
marked with a code NUM (jedan, dva, tri - one, two, three, ...).
Thus, this subgraph recognizes, among others, the following
expressions: 10, minus dva – minus two, +5 ili jedanaest – eleven.
The main transducer temperatura.grf (Figure 1) contains a
subgraph call stepen.grf. This graph is intended to recognize
expressions that describe the degrees on the Celsius scale, as the
common unit of temperature measure, in the texts in Serbian
language. Subgraph stepen.grf is shown in Figure 3.</p>
      <p>The lexical mask which refers to a dictionary word (&lt;stepen&gt;)
recognizes any form of the word stepen – degree (stepena,
stepeni, stepenima etc.). Graph temperatura.grf recognizes the
following phrases: oko +8 °C (approximately +8 ° C), - 1C, - 30
° C,- 4 stepena (- 4 degrees), od -1 C do 1 C (from -1 C to 1 C), -1
do +3 stepena (-1 to +3 degrees), -12 do -8 (-12 to -8), od 11 do
15 stepeni (from 11 to 15 degrees), 11 stepeni (11 degrees), od pet
do devet stepeni (from five to nine degrees), oko četiri (about
four), ispod 0 (below 0) etc.</p>
      <p>Similarly, for each feature in the Table 1 a rule extraction is
created for annotation of the text segments that carry specific
information.</p>
    </sec>
    <sec id="sec-12">
      <title>5.1 Analysis of Extracted Information and</title>
    </sec>
    <sec id="sec-13">
      <title>Process Efficiency</title>
      <p>The process of information extraction from the meteorological
texts is in the initial phase. During this phase, the analysis of the
texts from the described corpus was performed and the
transducers for extracting simple features were created. Since the
extraction rules are still evolving, and the text corpus over which
the extraction is carried out is fairly large (45862 sentences with
more than one million tokens), a comprehensive evaluation of the
system’s efficiency, which would accurately assess the precision
and recall, is not currently possible. However, an initial analysis
of the created transducers, which would determine the directions
for further development, is possible.</p>
      <p>Table 2 lists the transducers which were used to extract
information, in order of their implementation. The number of
extracted text segments is shown in the third column of the table,
while the evaluation of precision is presented in the fourth.</p>
      <sec id="sec-13-1">
        <title>Transducer</title>
      </sec>
      <sec id="sec-13-2">
        <title>Features</title>
        <p>opisTemp OpisTemperature
(Temperature</p>
        <p>Description)
temperature Temperatura</p>
        <p>(Temperature)
katTemp KatTemperature
(Temperature-Category)
vetarPre JacinaVetra
(WindAmount) and
PravacVetra
(WindDirection)
vetarPost JacinaVetra
(WindAmount) and
PravacVetra
(WindDirection)
padavine TipPadavina
(PrecipitationType) and
ObimPadavina
(Precipitation-Amount)
oblacnost ObimOblacnosti
(CloudAmount) and
PrisustvoOblaka
(CloudPresence)
deoTeritorije DeoTeritorije
(TeritoryPart)
imeTeritorije ImeTeritorije</p>
        <p>(TeritoryName)
lokalitet Lokalitet (Locality)
pojava Pojava (Phenomenon)
Number
of Evaluation
extracted of
text precision
segments
11518 100%
25618
14817
7720
99.6%
100%
100%
1559</p>
        <p>100%
18878</p>
        <p>100%
18875</p>
        <p>98%
4918
6036
7623
3737
99.8%
95%
98%
100%</p>
      </sec>
    </sec>
    <sec id="sec-14">
      <title>6. CONCLUSION</title>
      <p>The high precision of the transducers is expected, given that this
is an early stage of the system design and the extraction rules
creation process. Further development of the process, in order to
extract a larger number of individual pieces of information (i.e. to
increase recall), will surely reduce the precision. However, it is
expected the transducers will still maintain high efficiency.
We would like to emphasize that the next step in the process, after
the extraction of simple features, is merging the extracted data
into classes of higher semantic level. During that process, it will
be possible to further improve efficiency, by resolving
ambiguities or correcting wrongly interpreted text segments.</p>
    </sec>
    <sec id="sec-15">
      <title>7. ACKNOWLEDGMENTS</title>
      <p>This research was conducted through the projects 178006 and III
47003, financed by the Serbian Ministry of Science.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Brkić</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Matetić</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <year>2007</year>
          .
          <article-title>Modeling Natural Language Dialogue for Croatian Weather Forecast System</article-title>
          .
          <source>In Proceedings of the 18th International Conference on Information and Intelligent Systems (Varaždin, Croatia</source>
          ,
          <year>2003</year>
          ),
          <fpage>391</fpage>
          -
          <lpage>396</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Chevalier</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dansereau</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Poulin</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <year>1978</year>
          .
          <article-title>TAUMMETEO: Description du Système</article-title>
          . Universite de Montreal, Canada.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Friburger</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Maurel</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <year>2004</year>
          .
          <article-title>Finite-state transducer cascades to extract named entities in texts</article-title>
          .
          <source>Theoretical Computer Science</source>
          <volume>313</volume>
          ,
          <issue>1</issue>
          (
          <year>2004</year>
          ),
          <fpage>93</fpage>
          -
          <lpage>104</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Jurafsky</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Martin</surname>
            ,
            <given-names>J. H.</given-names>
          </string-name>
          <year>2008</year>
          .
          <article-title>Speech and language processing, 2nd edition</article-title>
          . Prentice-Hall Inc.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Kerpedjiev</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Noncheva</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <year>1990</year>
          .
          <article-title>Intelligent Handling of Weather Forecasts</article-title>
          .
          <source>In Proceedings of the 13th International Conference on Computational Linguistics COLING-90</source>
          ,
          <issue>3</issue>
          (Helsinki, Finland,
          <source>August 20-25</source>
          ,
          <year>1990</year>
          ),
          <fpage>379</fpage>
          -
          <lpage>381</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Kononenko</surname>
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Popov</surname>
            <given-names>I.</given-names>
          </string-name>
          , and
          <string-name>
            <given-names>Zagorulko</given-names>
            <surname>Yu</surname>
          </string-name>
          .
          <year>1999</year>
          .
          <article-title>Approach to Understanding Weather Forecast Telegrams with AgentBased Technique</article-title>
          . In Perspectives of System Informatics, Third International Andrei Ershov Memorial Conference,
          <source>PSI'99 (Novosibirsk, Russia, July 6-9</source>
          ,
          <year>1999</year>
          ),
          <fpage>511</fpage>
          -
          <lpage>516</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Kononenko</surname>
            <given-names>I.</given-names>
          </string-name>
          , Kononenko S.,
          <string-name>
            <surname>Popov</surname>
            <given-names>I.</given-names>
          </string-name>
          , and
          <string-name>
            <given-names>Zagorulko</given-names>
            <surname>Yu</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>Information extraction from non-segmented text (on the material of weather forecast telegrams)</article-title>
          .
          <source>In Proceedings of the 6th International Conference</source>
          , RIAO 2000 (College de France, France,
          <source>April 12-14</source>
          ,
          <year>2000</year>
          ),
          <fpage>1069</fpage>
          -
          <lpage>1088</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Krstev</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <year>2008</year>
          .
          <article-title>Processing of Serbian Automata, Texts and Electronic dictionaries</article-title>
          .
          <source>Faculty of Philology</source>
          , University of Belgrade, Belgrade, Serbia.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Krstev</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Vitas</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <year>2005</year>
          .
          <article-title>Corpus and Lexicon - Mutual Incompleteness</article-title>
          .
          <source>In Proceedings from the Corpus Linguistics Conference Series</source>
          ,
          <volume>1</volume>
          , 1, ISSN 1747-
          <fpage>939</fpage>
          (Birmingham University, UK,
          <source>July 14-17</source>
          ,
          <year>2005</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Krstev</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vitas</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Obradović</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Utvić</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <year>2011</year>
          .
          <article-title>EDictionaries and Finite-State Automata for the Recognition of Named Entities</article-title>
          .
          <source>In Proceedings of the 9th International Workshop on Finite State Methods and Natural Language Processing (Blois, France, July 12-15</source>
          ,
          <year>2011</year>
          ),
          <fpage>48</fpage>
          -
          <lpage>56</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Labsky</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nekvasil</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Svatek</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <year>2007</year>
          .
          <article-title>Towards web information extraction using extraction ontologies and (indirectly) domain ontologies</article-title>
          .
          <source>In Proceedings of the 4th international conference on Knowledge capture K-CAP '07 (Whistler</source>
          ,
          <string-name>
            <surname>BC</surname>
          </string-name>
          , Canada,
          <source>October 28-31</source>
          ,
          <year>2007</year>
          ), ACM New York, NY, USA,
          <fpage>201</fpage>
          -
          <lpage>202</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Paumier</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <year>2008</year>
          .
          <article-title>Unitex 2.1 User Manual</article-title>
          . http://www-igm.
          <source>univmlv.fr/~unitex/UnitexManual2</source>
          .1.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Silberztein</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <year>1993</year>
          .
          <article-title>Dictionnaires électroniques et analyse automatique de textes: le système INTEX</article-title>
          . Edition Masson, Paris.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Slocum</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <year>1985</year>
          .
          <article-title>A Survey of Machine Translation: its History, Current Status, and Future Prospects</article-title>
          .
          <source>In: Computational Linguistics</source>
          <volume>11</volume>
          ,
          <issue>1</issue>
          (
          <year>1985</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>17</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Vitas</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <year>2006</year>
          .
          <article-title>Prevodioci i interpretatori: Uvod u teoriju i metode kompilacije programskih jezika</article-title>
          .
          <source>Faculty of Mathematics</source>
          , University of Belgrade, Belgrade, Serbia.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Woods</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          <year>1970</year>
          .
          <article-title>Transition network grammars for natural language analysis</article-title>
          ,
          <source>In Communications of the ACM</source>
          <volume>13</volume>
          ,
          <issue>10</issue>
          (
          <year>1970</year>
          ),
          <fpage>591</fpage>
          -
          <lpage>606</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>