<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Hybrid Approach for Entity Extraction in Code-Mixed Social Media Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Shweta Comp. Sc.</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Engg. Deptt. IIT Patna</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>India shweta.pcs</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>@iitp.ac.in</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Asif Ekbal Comp. Sc.</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Engg. Deptt. IIT Patna</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>India asif@iitp.ac.in</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>CCS Concepts</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Computing methodologies</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Deepak Gupta Comp. Sc. &amp; Engg. Deptt. IIT Patna</institution>
          ,
          <country country="IN">India</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Pushpak Bhattacharyya Comp. Sc. &amp; Engg. Deptt. IIT Patna</institution>
          ,
          <country country="IN">India</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Shubham Tripathi Electrical Engg. Deptt. MNIT Jaipur</institution>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Entity extraction is one of the important tasks in various natural language processing (NLP) application areas. There has been a signi cant amount of works related to entity extraction, but mostly for a few languages (such as English, some European languages and few Asian languages) and doamins such as newswire. Nowadays social media have become a convenient and powerful way to express one's opinion and sentiment. India is a diverse country with a lot of linguistic and cultural variations. Texts written in social media are informal in nature, and perople often use more than one script while writing. User generated content such as tweets, blogs and personal websites of people are written using Roman script or sometimes users may use both Roman as well as indigenous scripts. Entity extraction is, in general, a more challenging task for such an informal text, and mixing of codes further complicates the process. In this paper, we propose a hybrid approah for enity extraction from code mixed language pairs such as English-Hindi and EnglishTamil. We use a rich linguistic feature set to train Conditional Random Field (CRF) classi er. The output of classi er is post-processed with a carefully hand-crafted feature set. The proposed system achieve the F-scores of 62:17% and 44:12% for English-Hindi and English-Tamil language pairs, respectively. Our system attains the best F-score among all the systems submitted in Fire 2016 shared task for the English-Tamil language pairs.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        Code-mixing refers to the mixing of two or more languages
or language varieties. Code switching and code mixing are
interchangeably used by the peoples. With the
availability of easy internet access to people, social media
involvement has been increased a lot. Over the past decade,
Indian language content on various media types such as blogs,
email, website, chats has increased signi cantly. And it is
observed that with the advent of smart phones more people
are using social media such as whatsapp, twitter, facebook
to share their opinion on people, products, services,
organizations, governments. The abundant of social media data
created many new opportunities for information access, but
also many new challenges. To deals with these challenges
many of the research is going on and it have become one of
the prime present-day research areas. Non-English speakers,
especially Indians, do not always use Unicode to write
something in social media in Indian languages. Instead, they use
their roman script or transliteration and frequently use
English words or phrases through code-mixing and also often
mix multiple languages in addition to anglicisms to express
their thoughts. Although English is the principal language
for social media communications, there is a necessity to
develop mechanism for other languages, including Indian
languages. According to the Indian constitution there are 22
o cial language in India. However Census of India of 2001,
reported India has 122 major languages and 1599 other
languages. The 2001 Census recorded 30 languages which were
spoken by more than a million native speakers and 122 which
were spoken by more than 10; 000 people. Language
diversity and dialect changes instigate frequent code-mixing in
India. Hence, Indians are multi-lingual by adaptation and
necessity, and frequently change and mix languages in social
media contexts, which poses additional di culties for
automatic social media text processing on Indian language. The
growth of Indian language content is expected to increase
by more than 70% every year. Hence there is a great need
to process this huge data automatically. Named Entity
Recognition (NER) is one of the key information extraction
tasks, which is concerned with identifying names of
entities such as people, location, organization and product. It
can be divided into two main phases: entity detection and
entity typing (also called classi cation)[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Recently,
Information extraction over micro-blogs have become an active
research topic [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], following early experiments which showed
this genre to be extremely challenging for state-of-the-art
algorithms[
        <xref ref-type="bibr" rid="ref2 ref5">5, 2</xref>
        ]. For instance, named entity recognition
methods typically have 85-90% accuracy on longer texts,
but 30-50% on tweets[
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. First, the shortness of micro-blogs
(maximum 140 characters for tweets) makes them hard to
interpret. Consequently, ambiguity is a major problem since
semantic annotation methods cannot easily make use of
coreference information. Unlike longer news articles, there is
a low amount of discourse information per microblog
document, and threaded structure is fragmented across multiple
documents, owing in multiple directions. Second,
microtexts exhibit much more language variation, tend to be less
grammatical than longer posts, contain unorthodox
capitalization, and make frequent use of emoticons,
abbreviations and hashtags, which can form an important part of the
meaning. To combat these problems, research has focused
on microblog-speci c information extraction algorithms (e.g.
named entity recognition for Twitter using CRFs[
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] or
hybrid methods[
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. Particular attention is given to micro-text
normalization[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], as a way of removing some of the
linguistic noise prior to part-of-speech tagging and entity
recognition. In literature primarily machine learning and rule
based approach has been used for named entity recognition
(NER). Machine learning (ML) based techniques for NER
make use of a large amount of NE annotated training data
to acquire higher level knowledge by extracting relevant
features from the labeled data. Several ML techniques have
already been applied for the NER tasks such as Support
vector vector classi er[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], Maximum Entropy[
        <xref ref-type="bibr" rid="ref10 ref3">3, 10</xref>
        ], Markov
Model(HMM)[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], Conditional Random Field (CRF)[
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] etc.
The rule based techniques have also been explored in the
task by[
        <xref ref-type="bibr" rid="ref13 ref19 ref6">6, 13, 19</xref>
        ]. The hybrid approaches that combines
di erent Machine learning based approaches are also used
by Rohini et al.[
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] by combining Maximum entropy,
Hidden Markov Model and handcrafted rules to build an NER
system. Entity extraction has been actively researched for
over 20 years. Most of the research has, however been
focused on resource rich languages, such as English, French
and Spanish. However entity extraction and recognition
from social media text on for Indian language have been
introduced on FIRE-15 workshop[
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. The code-mixing entity
extraction from social media text on for Indian mix language
introduced in the FIRE-2016. Entity extraction from
codemixing social media text poses some key challenge which are
as follows:
1. The released data set contains code mixing as well as
uni-language utterance.
2. Set of entity are not limited to only traditional set of
entity e.g. Person Name, Location Name,
Organization Name etc. There are 22 di erent types of entities
are there to extract from text.
3. There are lack of resources/tools for Indian languages.
code-mixing makes problems more di cult for
preprocessing tasks required for NER such as sentence
splitter, tokenization, part-of-speech tagging and
chunking etc.
      </p>
    </sec>
    <sec id="sec-2">
      <title>PROBLEM DEFINITION</title>
      <p>The problem de nition of code-mixing entity extraction
comprises two sub-problem entity extraction and entity
classi cation. Mathematically the problem of code-mixing
entity extraction can be described as follows: Lets S is
codemixing sentence having n tokens t1; t2 : : : tn. E is the set of
k pre-de ned entity E = fE1; E2; : : : Ekg.</p>
      <p>1. Entity Extraction step: Extract set of tokens TE =
fti; tj : : : tkg from S whose characteristics is similar to
any of the entity from entity set E.</p>
      <sec id="sec-2-1">
        <title>2. Entity classi cation step: Classify each of the to</title>
        <p>kens of set TE into one of the entity type from entity
set E.
3.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>DATASET</title>
      <p>There are two language pair data set available to
evaluate the system performance. It was crawled from tweeter,
mainly the crawled tweet are in English-Hindi and
EnglishTamil language mix. There are 22 types of entities present
in the training data set in which the majority of entities
are from `Entertainment', `Person' `Location' and
`Organization'. The statistics of the training data set is shown in the
Table-1. We have also shown some of the sample tweets from
both Language pair in Table-2. English-Tamil language pair
tweets contains some of the tweets from only Tamil language
only. English-Hindi tweet data set contains total 2700 tweets
from 2699 tweeter users. Similarly English-Tamil tweet data
set contains total 2183 tweets from 1866 tweeter users.</p>
      <sec id="sec-3-1">
        <title>Entities</title>
        <sec id="sec-3-1-1">
          <title>COUNT</title>
          <p>PLANTS</p>
          <p>PERIOD</p>
          <p>LOCOMOTIVE
ENTERTAINMENT
MONEY</p>
          <p>TIME
LIVTHINGS</p>
          <p>DISEASE
ARTIFACT</p>
          <p>MONTH
FACILITIES</p>
          <p>PERSON
MATERIALS
LOCATION</p>
          <p>YEAR</p>
          <p>DATE
ORGANIZATION
QUANTITY</p>
          <p>DAY</p>
          <p>SDAY
DISTANCE</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>Total</title>
        <p>English-Hindi
# Entity
132
1
44
13
810
25
22
7
7
25
10
10
712
24
194
143
33
109
2
67
23
0
2413</p>
        <p>English-Tamil
# Entity
94
3
53
5
260
66
18
16
5
18
25
23
661
28
188
54
14
68
0
15
6
4
1624
j tj (yi 1; yi; x; i) + X
k
ksk(yi; x; i))
(1)
where tj (yi 1; yi; x; i) is a transition feature function of the
entire observation sequence and the labels at positions i and
i 1 in the label sequence; sk(yi; x; i) is a state feature
function of the label at position i and the observation sequence;
and j and k are parameters to be estimated from training
data. When de ning feature functions, we construct a set
of real-valued features b(x; i) of the observation to expresses
some characteristic of the empirical distribution of the
training data that should also hold of the model distribution. An
example of such a feature is
b(x; i) =
1 if the observation word at position i is `religion'
0 otherwise
Each feature function takes on the value of one of these
realvalued observation features b(x; i) if the current state (in the
case of a state function) or previous and current states (in
the case of a transition function) take on particular values.
All feature functions are therefore real-valued. For example,
consider the following transition function:
tj (yi 1; yi; x; i) =
b(x; i) if yi 1 = B-ORG and yi = I-ORG</p>
        <p>0 otherwise
This allows the probability of a label sequence y given an
observation sequence x to be written as</p>
        <p>P (yjx; ) =</p>
        <p>1 exp(X
Z(x)
j
j Fj (y; x))
where Fj (y; x) can be written as follows:</p>
        <p>Fj (y; x) =
n
X fj (yi 1; yi; x; i)
i=1
where each fj (yi 1; yi; x; i) is either a state function
s(yi 1; yi; x; i) or a transition function t(yi 1; yi; x; i).</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. FEATURE EXTRACTION</title>
      <p>The proposed system uses an exhaustive set of features
for NE recognition. These features are described below.
1. Context word: Local contextual information is
useful to determine the type of the current word. We use
the contexts of previous two and next two words as
features.
2. Character n-gram: Character n-gram is a
contiguous sequence of n characters extracted from a given
word. The set of n-grams that can be generated for a
(2)
(3)
given token is basically the result of moving a window
of n characters along the text. We extracted
character n-grams of length one (unigram), two(bigram) and
three (trigram), and use these as features of the
classi ers.
3. Word normalization : Words are normalized in
order to capture the similarity between two di erent
words that share some common properties. Each
uppercase letter is replaced by `A', lowercase by 'a' and
number by '0'.</p>
      <sec id="sec-4-1">
        <title>Words</title>
        <p>NH10
Maine
NCR</p>
      </sec>
      <sec id="sec-4-2">
        <title>Normalization</title>
        <p>AA00
Aaaaa
AAA
4. Pre x and Su x: Pre x and su x of xed length
character sequences (here, 3) are stripped from each
token and used as features of the classi er.
5. Word Class Feature: This feature was de ned to
ensure that the words having similar structures belong
to the same class. In the rst step we normalize all
the words following the process as mentioned above.
Thereafter, consecutive same characters are squeezed
into a single character. For example, the normalized
word AAAaaa is converted to Aa. We found this
feature to be e ective for the biomedical domain, and we
directly adapted this without any modi cation.
6. Word Position: In order to capture the word
context in the sentence, we have used a numeric value to
indicate the position of word in the sentence. The
normalized position of word in the sentence is used as a
features. The feature values lies in the ranges between
0 and 1.
7. Number of Upper case Characters: This features
takes into account the number of uppercase alphabets
in the word. The feature is relative in nature and
ranges between 0 and 1.
8. Test Word Probability: This feature nds the
probability of the test word to be labeled the same as in
training data. The length of this vector feature is the
total number of labels or output tags, where each bit
represents an output tag, It is initialized with 0. If the
test word does not appear in training, every bits retain
their initially marked value 0. Based on the probability
value, we have two features:
(a) Top@1-Probability: For the output tag with
highest probability, its corresponding bit in the
feature vector is set to 1. All other bits remain as
0.
(b) Top@2-Probability: For the output tag with
highest and second highest probability, their
corresponding bit are set to 1 in the feature vector.</p>
        <p>All other bits remain as 0.
9. Binary Features: These binary features are
identied after the through analysis of training data.
(a) isSu cientLength: Since most of the entity
from training data have a signi cant length.
Therefore we set a binary feature to re when the length
of token is greater than a speci c threshold value.
The threshold value 4 is used to extract the
binary features.
(b) isAllCapital: This value of this feature is set
when all the character of current token is in
uppercase.
(c) isFirstCharacterUpper: This value of this
feature is set when the rst character of current
token is in uppercase..
(d) isInitCap: This feature checks whether the
current token starts with a capital letter or not. This
provides an evidence for the target word to be of
NE type for the English language.
(e) isInitPunDigit: We de ne a binary-valued
feature that checks whether the current token starts
with a punctuation or a digit. It indicates that
the respective word does not belong to any
language. Few such examples are 4u,:D, :P etc.
(f) isDigit: This feature is red when the current
token is numeric.
(g) isDigitAlpha: We de ne this feature in such a
way that checks whether the current token is
alphanumeric. The word for which this feature has
a true value has a tendency of not being labeled
as any named entity type.
(h) isHashTag: Since we are dealing with tweeter
data , therefore we encounter a lot of hashtag is
tweets. We de ne the binary feature that checks
whether the current token starts with # or not.</p>
        <sec id="sec-4-2-1">
          <title>Input : Disease Name list as DN</title>
          <p>Living things list as LT ;
Special Days list as SD</p>
          <p>List of pair obtained from CRF as L(W,C)
Output: Post-processed list of (token,label) pair
obtained after post-processing as PL(W,C')
PL(W,C')=L(W,C)
while L(W,C) is non-empty do
if DN contains Wi then</p>
          <p>C=DISEASE;</p>
          <p>C'=C;
else if LT contains Wi then</p>
          <p>C=LIVINGTHINGS;</p>
          <p>C'=C;
else if SD contains Wi then</p>
          <p>C=SPECIALDAYS;</p>
          <p>C'=C;
else</p>
          <p>C'=C;
end
return PL(W,C');
Algorithm 1: Post-processing algorithm for code mixed
data set
5.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>EXPERIMENTAL SETUP</title>
      <p>To extract the entity from code mixed data we have
followed three step approach, which are described in this
section. Fig-1 shows a architecture diagram of our proposed
approach.
5.1</p>
    </sec>
    <sec id="sec-6">
      <title>Pre-processing</title>
      <p>Pre-processing stage is an important task before applying
any classi er. The release data set was in raw text sentence
having the list of entity. There are two step was performed
as part of pre-processing.</p>
      <p>
        1. Tokenization: Since the data set are crawled from
Twitter therefore a suitable tokenizer which can deals
with social media data need to be used. We used the
CMU PoS tagger[
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] for tokenization and PoS tagging
of tweets.
2. Token Encoding: We used the IOB encoding for
tagging token. The IOB format (Inside, Outside,
Beginning) is a common tagging format for tagging tokens
in a chunking task in computational linguistics (ex.
Named Entity Recognition). The B- pre x before a
tag indicates that the tag is the beginning of a chunk,
and an I- pre x before a tag indicates that the tag
is inside a chunk. An O tag indicates that a token
belongs to no chunk.
5.2
      </p>
    </sec>
    <sec id="sec-7">
      <title>Sequence Labeling</title>
      <p>
        In literature primarily HMM, MEMM and CRF has been
used for sequence labeling task. Here we used the CRF
classi er for label the sequence of token. The features set
described in section-4 were nally formulated to provide as
input to our CRF classi er[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. We used the CRF++1
implementation of CRF. The default parameter setting was used
to carry out the experiment.
5.3
      </p>
    </sec>
    <sec id="sec-8">
      <title>Post-processing</title>
      <p>The rule and dictionary based post-processing was
performed followed by labeling obtained from CRF classi er.
The detailed explanation are as follows:</p>
    </sec>
    <sec id="sec-9">
      <title>English-Hindi</title>
      <p>For English-Hindi code mixed data we used dictionary of
Disease Name, Living Things and Special Days. A list
consist of 250 disease name was obtained from wiki page2. A
manual list are created of 668 living things from di erent
web page source. Similarly for list of special days, a
manual list of 92 special days was obtained from this website3.
The dictionary element was red in the order as mentioned
in Algorithm-1. At last regular expressions are formed to
correct PERIOD, MONEY and TIME on post-processed
output.</p>
    </sec>
    <sec id="sec-10">
      <title>English-Tamil</title>
      <p>Since none of our team member are native Tamil speaker so
we could not do deep error analysis of CRF predicted output.
We used the same dictionary which was used in
EnglishHindi data set, because those dictionary contains only
English word. Finally regular expressions are formed to correct
1https://taku910.github.io/crfpp/
2https://simple.wikipedia.org/wiki/List of diseases
3www.drikpanchang.com/calendars/indian/indiancalendar.html
Run-2</p>
      <p>R
NA
9.93
19.59
12.21
NA</p>
      <p>Run-3
F P R F</p>
      <p>NA
17.53 79.51 21.88 34.32
31.44 NA
20.22 58.94 11.94 19.86</p>
      <p>NA
the PERIOD, MONEY and TIME on post-processed
output.</p>
    </sec>
    <sec id="sec-11">
      <title>RESULT &amp; ANALYSIS</title>
      <p>An entity extraction model for English-Hindi &amp;
EnglishTamil language pair are trained by using CRF as base
classi er. We test our system on the test data for the concerned
language pair. The proposed approach was used to extract
the entity from both the language pair data.</p>
      <p>The developed entity extraction and identi cation system
has been evaluated using the precision(P), recall (R) and
F-measure (F). The organizers of the CMEE-IL task FIRE
2016, released the data in two phases: in the rst phase,
training data is released along with the corresponding NE
annotation le. In the second phase, the test data is released
and no NE annotation le is provided. The extracted NE
annotation le for test data was nally sent to the
organizers for evaluation. The organizers evaluate the di erent runs
submitted by the various teams and send the o cial results
to the participating teams.</p>
      <p>The o cial results for English-Hindi language pair are shown
in Table-3. Our system(Deepak-IITPatna) performance are
shown in bold font. Our system got the highest Precision
of 81:15% among all the submitted system. The proposed
approach achieved F-score of 62:17% on English-Hindi
language pair. Table-4 shows the o cial result for
EnglishTamil language pair data set. Our system(Deepak-IITPatna)
performance are shown in bold font. Our system is the best
performing system among all the submitted system. We
achieved the 79:92%, 30:47% and 44:12% precision(p),
recall(r) and F-score(f) respectively. The reason for lower
Fscore on Tamil-English could be the lack of good features
which can help to recognize a Tamil word as named entity.
7.</p>
    </sec>
    <sec id="sec-12">
      <title>CONCLUSION &amp; FUTURE WORK</title>
      <p>This paper describes a code mixed named entity
recognition from Social media text in English-Hindi and
EnglishTamil language pair data. Our proposed approach is a
hybrid model of machine learning and rule based system. The
experimental results show that our system is the best
performer among the systems participated in the CMEE-IL task
for code mixed English-Tamil language pair. For
EnglishHindi language pair system achieved the highest precision
value 81:15% among all the submitted system. In future we
would like to build a more robust code mixed NER system
by using deep learning system.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>D. M.</given-names>
            <surname>Bikel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Schwartz</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Weischedel</surname>
          </string-name>
          .
          <article-title>Nymble: a high-performance learning name- nder</article-title>
          .
          <source>In Proceedings of the fth conference on Applied natural language processing</source>
          , pages
          <volume>194</volume>
          {
          <fpage>201</fpage>
          . Association for Computational Linguistics,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>K.</given-names>
            <surname>Bontcheva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Derczynski</surname>
          </string-name>
          ,
          <string-name>
            <surname>and I. Roberts.</surname>
          </string-name>
          <article-title>Crowdsourcing named entity recognition and entity linking corpora</article-title>
          .
          <source>The Handbook of Linguistic Annotation</source>
          (to appear),
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Borthwick</surname>
          </string-name>
          .
          <article-title>A maximum entropy approach to named entity recognition</article-title>
          .
          <source>PhD thesis</source>
          , Citeseer,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A. E. Cano</given-names>
            <surname>Basave</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Varga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rowe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Stankovic</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.-S.</given-names>
            <surname>Dadzie</surname>
          </string-name>
          .
          <article-title>Making sense of microposts (# msm2013) concept extraction challenge</article-title>
          .
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>L.</given-names>
            <surname>Derczynski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Maynard</surname>
          </string-name>
          , G. Rizzo, M. van Erp,
          <string-name>
            <given-names>G.</given-names>
            <surname>Gorrell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Troncy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Petrak</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Bontcheva</surname>
          </string-name>
          .
          <article-title>Analysis of named entity recognition and linking for tweets</article-title>
          .
          <source>Information Processing &amp; Management</source>
          ,
          <volume>51</volume>
          (
          <issue>2</issue>
          ):
          <volume>32</volume>
          {
          <fpage>49</fpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>R.</given-names>
            <surname>Grishman</surname>
          </string-name>
          .
          <article-title>The nyu system for muc-6 or where's the syntax</article-title>
          ?
          <source>In Proceedings of the 6th conference on Message understanding</source>
          , pages
          <volume>167</volume>
          {
          <fpage>175</fpage>
          . Association for Computational Linguistics,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>R.</given-names>
            <surname>Grishman</surname>
          </string-name>
          and
          <string-name>
            <given-names>B.</given-names>
            <surname>Sundheim</surname>
          </string-name>
          . Message understanding conference-6
          <article-title>: A brief history</article-title>
          .
          <source>In Proceedings of the 16th Conference on Computational Linguistics - Volume 1, COLING '96</source>
          , pages
          <fpage>466</fpage>
          {
          <fpage>471</fpage>
          ,
          <string-name>
            <surname>Stroudsburg</surname>
          </string-name>
          , PA, USA,
          <year>1996</year>
          .
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>B.</given-names>
            <surname>Han</surname>
          </string-name>
          and
          <string-name>
            <given-names>T.</given-names>
            <surname>Baldwin</surname>
          </string-name>
          .
          <article-title>Lexical normalisation of short text messages: Makn sens a# twitter</article-title>
          .
          <source>In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies-Volume</source>
          <volume>1</volume>
          , pages
          <fpage>368</fpage>
          {
          <fpage>378</fpage>
          . Association for Computational Linguistics,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>H.</given-names>
            <surname>Isozaki</surname>
          </string-name>
          and
          <string-name>
            <given-names>H.</given-names>
            <surname>Kazawa</surname>
          </string-name>
          .
          <article-title>E cient support vector classi ers for named entity recognition</article-title>
          .
          <source>In Proceedings of the 19th international conference on Computational linguistics-Volume</source>
          <volume>1</volume>
          , pages
          <fpage>1</fpage>
          <lpage>{</lpage>
          7. Association for Computational Linguistics,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>N.</given-names>
            <surname>Kumar</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Bhattacharyya</surname>
          </string-name>
          .
          <article-title>Named entity recognition in hindi using memm</article-title>
          .
          <source>Techbical Report</source>
          , IIT Mumbai,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J. La erty</given-names>
            , A.
            <surname>McCallum</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F. C.</given-names>
            <surname>Pereira</surname>
          </string-name>
          .
          <article-title>Conditional random elds: Probabilistic models for segmenting and labeling sequence data</article-title>
          .
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>W.</given-names>
            <surname>Li</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>McCallum</surname>
          </string-name>
          .
          <article-title>Rapid development of hindi named entity recognition using conditional random elds and feature induction</article-title>
          .
          <source>ACM Transactions on Asian Language Information Processing (TALIP)</source>
          ,
          <volume>2</volume>
          (
          <issue>3</issue>
          ):
          <volume>290</volume>
          {
          <fpage>294</fpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>D.</given-names>
            <surname>McDonald</surname>
          </string-name>
          .
          <article-title>Internal and external evidence in the identi cation and semantic categorization of proper names</article-title>
          .
          <source>Corpus processing for lexical acquisition</source>
          , pages
          <volume>21</volume>
          {
          <fpage>39</fpage>
          ,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>O.</given-names>
            <surname>Owoputi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. O</given-names>
            <surname>'Connor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Dyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Gimpel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Schneider</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N. A.</given-names>
            <surname>Smith.</surname>
          </string-name>
          <article-title>Improved part-of-speech tagging for online conversational text with word clusters</article-title>
          .
          <source>Association for Computational Linguistics</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>P. R.</given-names>
            <surname>Rao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Malarkodi</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S. L.</given-names>
            <surname>Devi.</surname>
          </string-name>
          Esm-il:
          <article-title>Entity extraction from social media text for indian languages@ re 2015{an overview</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ritter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Clark</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Etzioni</surname>
          </string-name>
          , et al.
          <article-title>Named entity recognition in tweets: an experimental study</article-title>
          .
          <source>In Proceedings of the Conference on Empirical Methods in Natural Language Processing</source>
          , pages
          <volume>1524</volume>
          {
          <fpage>1534</fpage>
          . Association for Computational Linguistics,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>R.</given-names>
            <surname>Srihari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Niu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W.</given-names>
            <surname>Li</surname>
          </string-name>
          .
          <article-title>A hybrid approach for named entity and sub-type tagging</article-title>
          .
          <source>In Proceedings of the sixth conference on Applied natural language processing</source>
          , pages
          <volume>247</volume>
          {
          <fpage>254</fpage>
          . Association for Computational Linguistics,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>M. Van Erp</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Rizzo</surname>
            , and
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Troncy</surname>
          </string-name>
          .
          <article-title>Learning with the web: Spotting named entities on the intersection of nerd and machine learning</article-title>
          .
          <source>In # MSM</source>
          , pages
          <volume>27</volume>
          {
          <fpage>30</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>T.</given-names>
            <surname>Wakao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Gaizauskas</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wilks</surname>
          </string-name>
          .
          <article-title>Evaluation of an algorithm for the recognition and classi cation of proper names</article-title>
          .
          <source>In Proceedings of the 16th conference on Computational linguistics-Volume</source>
          <volume>1</volume>
          , pages
          <fpage>418</fpage>
          {
          <fpage>423</fpage>
          . Association for Computational Linguistics,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>