<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Reconocimiento y Clasicacin de Entidades Nombradas independiente de la lengua y el dominio mediante perles</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Language</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Domain Independent Named Entity Recognition</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Classication through Proles</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Isabel Moreno Departamento de Lenguajes y Sistemas InformÆticos Universidad de Alicante Apdo. de correos</institution>
          ,
          <addr-line>99 E-03080 Alicante</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>Named Entity Recognition and Classication (NERC) is a prerequisite to many natural language processing applications. Nevertheless, the adaptation of NERC systems is usually expensive given that most of them only work appropriately on the scenario for which they were created. Therefore, the main purpose of this thesis is to research, analyse and develop an adaptable system, named CARMEN, for NERC through proles and supervised machine learning. Attention would be focused on CARMEN being domain and language independent so as to achieve similar results, using the same method, regardless of the training corpus utilised.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Desde hace tiempo estamos en la era de
la informacin digital y, aunque esta crece
sin descanso, nuestra habilidad para
explotarla y procesarla continoea constante
        <xref ref-type="bibr" rid="ref3 ref37">(Bendov y Feldman, 2010)</xref>
        . Desde esta
perspectiva, el Procesamiento del Lenguaje
Natural (PLN) investiga y formula mecanismos
computacionales para facilitar la interrelacin
hombre-mÆquina por medio del lenguaje
natural, en lugar de otros lenguajes mÆs
formales y restrictivos, sin perder efectividad
        <xref ref-type="bibr" rid="ref18 ref29">(Manaris, 1998; Moreno et al., 1999)</xref>
        . MÆs
concretamente, las tØcnicas de Extraccin de
Informacin (EI) procesan texto para detectar la
informacin textual explcita de interØs y
convertirla a un formato fÆcilmente comprensible
por las mÆquinas, tambiØn conocido como
estructurado
        <xref ref-type="bibr" rid="ref3 ref37">(Ben-dov y Feldman, 2010)</xref>
        .
      </p>
      <p>
        Una de las tareas que lleva a cabo la EI es
el Reconocimiento y la Clasicacin de
Entidades Nombradas (RCEN)
        <xref ref-type="bibr" rid="ref16 ref19 ref30">(Nadeau y Sekine,
2007; Marrero et al., 2013)</xref>
        , que tiene dos
objetivos diferenciados. Primero, identicar las
menciones de nombres propios en un texto,
lo que se conoce como la fase de
reconocimiento (REN). Segundo, asignar una
categora, de entre un conjunto predeterminado, a
cada una de las entidades previamente
reconocidas, llamada fase de clasicacin (CEN).
Ambos objetivos pueden abordarse de
manera conjunta o separada.
      </p>
      <p>
        Los sistemas RCEN juegan un papel
importante en muchas aplicaciones que procesan
informacin textual. La razn es que el RCEN
es un prerrequisito para diversas tareas como:
la minera de opiniones
        <xref ref-type="bibr" rid="ref13 ref5 ref5">(Ding, Liu, y Zhang,
2009; Jin, Hay Ho, y Srihari, 2009)</xref>
        , la
generacin automÆtica de resoemenes
        <xref ref-type="bibr" rid="ref12 ref2 ref22 ref24 ref7">(Fuentes y
Rodrguez, 2002; Alcn y Lloret, 2015)</xref>
        , la
generacin de lenguaje natural
        <xref ref-type="bibr" rid="ref1 ref21 ref38">(Vicente y
Lloret, 2016)</xref>
        , los sistemas de boesqueda de
respuestas
        <xref ref-type="bibr" rid="ref16 ref17 ref23 ref30 ref31 ref36">(Peregrino, TomÆs, y Pascual, 2012;
Lee, Hwang, y Jang, 2007; Lee et al., 2006)</xref>
        o los sistemas de recuperacin de
informacin
        <xref ref-type="bibr" rid="ref4 ref9">(Guo et al., 2009; Chen, Ding, y Tsai,
1998)</xref>
        , entre otras aplicaciones.
      </p>
      <p>
        A pesar de que los sistemas RCEN son de
uso comoen, su utilizacin no siempre es
directa. La mayora de sistemas RCEN fueron
desarrollados ad-hoc para un dominio1
concreto, con requisitos especcos y, a su vez,
un conjunto reducido de tipos de entidades de
interØs en ese dominio. Como resultado,
cuando se quiere portar una herramienta RCEN a
otro dominio, con otros requisitos y un
conjunto diferente de entidades, se requiere un
esfuerzo considerable para que funcione
adecuadamente
        <xref ref-type="bibr" rid="ref19">(Marrero et al., 2013)</xref>
        .
      </p>
      <p>
        AdemÆs, el RCEN estÆ condicionado por
la lengua para la que se desarrollan los
sistemas. La mayora de herramientas se
construyen para un corpus especco y, como
consecuencia, existe una dependencia de la
lengua de dicho corpus. La adaptacin de un
RCEN a un nuevo idioma no siempre es
posible por tres razones principales: (i) estos
sistemas RCEN suelen necesitar de herramientas
de anÆlisis lingstico que no siempre estÆn
disponibles para todos los idiomas
        <xref ref-type="bibr" rid="ref10">(Indurkhya, 2014)</xref>
        ; (ii) el RCEN depende comoenmente
de otros recursos (como diccionarios) que
varan entre lenguas
        <xref ref-type="bibr" rid="ref19">(Marrero et al., 2013)</xref>
        , si es
que existen; y (iii) cada idioma supone retos
diferentes que pueden afectar al rendimiento
del RCEN, como se observ en
        <xref ref-type="bibr" rid="ref33 ref35">(Tjong Kim
Sang, 2002; Sang y De Meulder, 2003)</xref>
        .
      </p>
      <p>Por ello, el presente proyecto de tesis se
centrarÆ en analizar, proponer y desarrollar
un sistema RCEN, llamado CARMEN,
basado en perles que emplearÆ aprendizaje
automÆtico supervisado. Se buscarÆ que dicho
sistema sea independiente del dominio y de la
lengua para, con el mismo mØtodo, conseguir
resultados similares sin importar el corpus de
entrenamiento utilizado.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Trabajo relacionado</title>
      <p>
        Hace mÆs de dos dØcadas que fue acuaeado el
tØrmino Entidad Nombrada (EN) en la
sex1En esta tesis se entiende por dominio al tpico
o Ærea de interØs de un corpus, como pueden ser el
dominio mØdico o el educativo.
ta conferencia Message Understanding
Conference (MUC). En ella el objetivo era la
RCEN de personas, organizaciones, lugares
as como expresiones numØricas de tiempo y
cantidad
        <xref ref-type="bibr" rid="ref8">(Grishman y Sundheim, 1996)</xref>
        .
Desde entonces, diversos foros de PLN han
seguido sus pasos, promocionando tareas
para evaluar sistemas RCEN
        <xref ref-type="bibr" rid="ref11 ref12 ref2 ref22 ref24 ref3 ref32 ref33 ref34 ref35 ref37 ref6">(Tjong Kim Sang,
2002; Sang y De Meulder, 2003; Uzuner, Solti,
y Cadag, 2010; Segura-Bedmar, Martnez, y
Herrero-Zazo, 2013; Ji, Nothman, y Hachey,
2014; Pradhan et al., 2014; Elhadad et al.,
2015; Ji, Nothman, y Hachey, 2015)</xref>
        .
      </p>
      <p>
        Se observan dos patrones en el foco de
las mismas: (i) un dominio y moeltiples
idiomas
        <xref ref-type="bibr" rid="ref11 ref12 ref2 ref22 ref24 ref33 ref35">(Tjong Kim Sang, 2002; Sang y De
Meulder, 2003; Ji, Nothman, y Hachey, 2014; Ji,
Nothman, y Hachey, 2015)</xref>
        ; o (ii) un dominio
restringido y un solo idioma
        <xref ref-type="bibr" rid="ref3 ref32 ref34 ref37 ref6">(Uzuner, Solti,
y Cadag, 2010; Segura-Bedmar, Martnez, y
Herrero-Zazo, 2013; Pradhan et al., 2014;
Elhadad et al., 2015)</xref>
        .
      </p>
      <p>
        Un ejemplo del primer caso lo
encontramos en la conferencia CoNLL, donde se
organizaron dos competiciones
        <xref ref-type="bibr" rid="ref33 ref35">(Tjong Kim Sang,
2002; Sang y De Meulder, 2003)</xref>
        para tratar el
RCEN en noticias de peridicos en inglØs,
holandØs, castellano y alemÆn. En ambas
ediciones, los sistemas obtuvieron diferentes
resultados en cada idioma. Por ejemplo, el mejor
sistema de cada edicin tuvo una diferencia
de al menos 15 puntos en la F1 global, por
lo que es discutible que sean
completamente independientes del idioma. MÆs
recientemente se han investigado otras
aproximaciones RCEN multilinges
        <xref ref-type="bibr" rid="ref1 ref15 ref21 ref38">(Konkol et al., 2015;
Agerri y Rigau, 2016)</xref>
        , donde tambiØn se
observan diferentes resultados en cada uno de
los idiomas (aproximadamente 15 puntos de
F1 global).
      </p>
      <p>
        Respecto al oeltimo caso, un ejemplo de
competicin centrada en un dominio
restringido y un idioma es el DDIExtraction
2013
        <xref ref-type="bibr" rid="ref34">(Segura-Bedmar, Martnez, y
HerreroZazo, 2013)</xref>
        , organizado dentro del taller
internacional SemEval. Uno de sus objetivos
principales es el RCEN en dos fuentes
mØdicas de informacin textual (DrugBank y
MedLine). TambiØn en este caso los
participantes obtuvieron resultados diferentes segoen la
fuente (al menos 20 puntos en la F1 global).
      </p>
      <p>
        Fuera de estos marcos de evaluacin y
la multilingualidad, Tkachenko y Simanovsky
(2012) diseaean un RCEN y experimentan con
varios gØneros textuales presentes en el corpus
OntoNotes.
        <xref ref-type="bibr" rid="ref14">Kitoogo y Baryamureeba (2008</xref>
        )
denen un RCEN que se prob en dos
dominios (general y legislativo): entrenando en el
dominio general
        <xref ref-type="bibr" rid="ref33">(Sang y De Meulder, 2003)</xref>
        y
evaluando en el legislativo, y viceversa.
Ambos trabajos obtienen una diferencia de al
menos 20 puntos en la F1 global cuando cambian
de dominio o gØnero.
      </p>
      <p>Aunque vemos que se han hecho progresos
considerables en el RCEN, los resultados de
las investigaciones ponen de maniesto que
los sistemas no han mostrado un rendimiento
ptimo cuando cambia el idioma o el dominio,
as como la fuente o el gØnero textual.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Propuesta de investigacin</title>
      <p>Dado este panorama general, esta tesis
doctoral plantea como objetivo la investigacin,
anÆlisis y desarrollo de un sistema adaptable
para el RCEN, llamado CARMEN. La
atencin se centrarÆ, sobre todo, en que
CARMEN proporcione salidas consistentes aun
cuando cambie el dominio o la fuente o el
gØnero o el idioma del corpus de entrenamiento.</p>
      <p>Por tanto, la hiptesis de partida es que
el desarrollo de un sistema RCEN basado en
perles con aprendizaje automÆtico,
redunda en herramientas con mnima adaptacin,
evitando las diferencias observadas en los
resultados actuales en relacin a dependencias
del dominio o de la lengua. En cuanto a la
dependencia del dominio, concretamente, nos
planteamos el estudio de nuestra
aproximacin en al menos dos dominios: (i) general,
que representa necesidades de informacin
comunes; y (ii) farmacoterapØutico, que
representa necesidades de informacin especcas
durante la atencin sanitaria. Ambos
dominios son altamente representativos en cuanto
a gØneros textuales, idiomas y entidades
nombradas. Por tanto, estos dominios permiten
denir un escenario de evaluacin apropiado
para conrmar nuestra hiptesis y denir
objetivos especcos:
O1 Realizar un estado de la cuestin,
sistemÆtico y exhaustivo, para detectar las
limitaciones tanto de las aproximaciones
para RCEN como de los corpus
existentes, al menos, en dos dominios: general y
farmacoterapØutico.</p>
      <p>O2 Analizar las entidades nombradas
relevantes en el dominio
farmacoterapØutico y crear un corpus en espaaeol para el</p>
      <sec id="sec-3-1">
        <title>RCEN en este dominio, as como evaluar</title>
        <p>la calidad del recurso generado.</p>
        <p>O3 Diseaear e implementar nuevas tØcnicas
de RCEN que permitan solventar alguna
de las limitaciones de las aproximaciones
encontradas en el estado de la cuestin:</p>
      </sec>
      <sec id="sec-3-2">
        <title>O3.1 diseaear nuevas tØcnicas de REN,</title>
        <p>O3.2 diseaear nuevas tØcnicas de CEN, y
O3.3 diseaear nuevas tØcnicas de
desambiguacin de entidades con respecto a
bases de conocimiento y las
relaciones entre las mismas.</p>
        <p>O4 Diseaear y analizar experimentos para
evaluar el resultado del objetivo O3, en al
menos dos dominios y dos lenguas,
realizando para ello una evaluacin intrnseca
y extrnseca, basada en mØtodos
cuantitativos y cualitativos.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Metodologa y experimentos</title>
      <p>Con el n de demostrar la hiptesis y los
objetivos presentados en la seccin anterior, hasta
ahora hemos llevado a cabo tres grandes
tareas:</p>
      <p>
        Primero, la creacin del corpus
DrugSemantics para el RCEN en el dominio
farmacoterapØutico ha concluido con la creacin de un
gold standard, siguiendo la metodologa
descrita en
        <xref ref-type="bibr" rid="ref20 ref25 ref26 ref27 ref28">(Moreno et al., 2017)</xref>
        .
      </p>
      <p>
        Segundo, la implementacin de un RCEN,
llamado MaNER, basado en lexicones
especcos del dominio mØdico
        <xref ref-type="bibr" rid="ref12 ref12 ref2 ref2 ref20 ref22 ref22 ref23 ref24 ref24 ref25 ref26 ref27 ref28 ref36">(Moreno, Moreda, y
RomÆ-Ferri, 2012; Moreno, Moreda, y
RomÆFerri, 2015; Moreno, Moreda, y RomÆ-Ferri,
2015; Moreno et al., 2017)</xref>
        , que sirvi de
apoyo en la construccin de un gold standard de
calidad.
      </p>
      <p>
        Tercero, el desarrollo del mdulo CEN
basado en perles del sistema CARMEN, que
emplea aprendizaje automÆtico supervisado
y cuyas caractersticas incluyen informacin
local a la entidad (ajos, longitud y la
propia entidad) as como informacin de
contexto en una ventana. El contexto se consigue
mediante perles generados para cada una de
las entidades. Los resultados de diversos
experimentos nos han permitido estudiar
diferentes parÆmetros en corpus de diferentes
dominios
        <xref ref-type="bibr" rid="ref20 ref25 ref26 ref27 ref28">(Moreno, RomÆ-Ferri, y Moreda, 2017d)</xref>
        ,
as como su rendimiento
        <xref ref-type="bibr" rid="ref1 ref20 ref20 ref21 ref25 ref25 ref26 ref26 ref27 ref27 ref28 ref28 ref38">(Moreno, Moreda,
y RomÆ-Ferri, 2016; Moreno, RomÆ-Ferri, y
Moreda, 2017b; Moreno, RomÆ-Ferri, y
Moreda, 2017c)</xref>
        . AdemÆs, estamos
experimentando en varios idiomas
        <xref ref-type="bibr" rid="ref20 ref25 ref26 ref27 ref28">(Moreno, RomÆ-Ferri, y
Moreda, 2017a)</xref>
        .
5
tir:
      </p>
    </sec>
    <sec id="sec-5">
      <title>Elementos especcos para discusin</title>
      <p>Siendo el RCEN un tema de gran interØs en
el PLN, queremos intercambiar experiencias
para orientar nuestra investigacin.</p>
      <p>En concreto, son dos los intereses a
debanuestra prxima tarea, la REN, as
como las diferentes tØcnicas, caractersticas
y herramientas que permitiran construir
este mdulo independiente del dominio y
la lengua.
posibles escenarios y experimentos que
nos permitan reforzar nuestra hiptesis.</p>
    </sec>
    <sec id="sec-6">
      <title>Agradecimientos</title>
      <p>
        Esta investigacin ha sido nanciada
parcialmente por el Gobierno Espaaeol
        <xref ref-type="bibr" rid="ref12 ref2 ref22 ref24">(TIN2015-65100-R y
TIN2015-65136-C22-R)</xref>
        , la Generalitat Valenciana
(PROMETEOII/2014/001), la Universidad de
Alicante (GRE16-01: Plataforma inteligente
para recuperacin, anÆlisis y representacin
de la informacin generada por usuarios en
Internet) y las Ayudas Fundacin BBVA
a equipos de investigacin cientca 2016
(ASAP - AnÆlisis de Sentimientos Aplicado
a la Prevencin del Suicidio en las Redes
Sociales).
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Agerri</surname>
            , R. y
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Rigau</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Robust multilingual Named Entity Recognition with shallow semi-supervised features</article-title>
          .
          <source>Articial Intelligence</source>
          ,
          <volume>238</volume>
          :
          <fpage>6382</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Alcn</surname>
            , . y
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Lloret</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Estudio de la inuencia de incorporar conocimiento lØxicosemÆntico a la</article-title>
          <string-name>
            <surname>tØcnica de AnÆlisis de Componentes</surname>
          </string-name>
          <article-title>Principales para la generacin de resoemenes multilinges</article-title>
          .
          <source>LinguamÆtica</source>
          ,
          <volume>7</volume>
          (
          <issue>1</issue>
          ):
          <fpage>53</fpage>
          <lpage>63</lpage>
          , Julio.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Ben-dov</surname>
            , M. y
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Feldman</surname>
          </string-name>
          .
          <year>2010</year>
          . Text Mining and
          <string-name>
            <given-names>Information</given-names>
            <surname>Extraction. En O. Maimon</surname>
          </string-name>
          y L.
          <article-title>Rokach, editores, Data Mining and Knowledge Discovery Handbook</article-title>
          . Springer US, Boston, MA, 2nd edicin, captulo
          <volume>42</volume>
          , pÆginas 809835.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Chen</surname>
            , H., Y. Ding, y
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Tsai</surname>
          </string-name>
          .
          <year>1998</year>
          .
          <article-title>Named Entity Extraction for Information Retrieval</article-title>
          .
          <source>En COMPUTER PROCESSING OF ORIENTAL LANGUAGES, volumen 11.</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Ding</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          , B. Liu, y
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhang</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Entity Discovery and Assignment for Opinion Mining Applications</article-title>
          .
          <source>En Proceedings of the 15th ACM SIGKDD international conference on Knowledge discovery and data mining , pÆginas 11251134.</source>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Elhadad</surname>
            , N.,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Pradhan</surname>
            ,
            <given-names>S. L.</given-names>
          </string-name>
          <string-name>
            <surname>Gorman</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Manandhar</surname>
          </string-name>
          , W. W. Chapman, y G. Savova.
          <year>2015</year>
          . SemEval-2015
          <source>Task 14 : Analysis of Clinical Text. Proceedings of the 9th International Workshop on Semantic Evaluation</source>
          , pÆginas
          <volume>303310</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Fuentes</surname>
            , M. y
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Rodrguez</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>Using cohesive properties of text for automatic summarization</article-title>
          . En Actas de las Jornadas de tratamiento y recuperacin de la informacin (Jotri'
          <year>2002</year>
          ) .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Grishman</surname>
            , R. y
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Sundheim</surname>
          </string-name>
          .
          <year>1996</year>
          .
          <article-title>Message understanding conference-6: A brief history</article-title>
          .
          <source>En Proceedings of the 16th Conference on Computational Linguistics - Volume</source>
          <volume>1</volume>
          , COLING '
          <volume>96</volume>
          , pÆginas 466471,
          <string-name>
            <surname>Stroudsburg</surname>
          </string-name>
          , PA, USA. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Guo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Xu</surname>
          </string-name>
          , X. Cheng, y
          <string-name>
            <given-names>H.</given-names>
            <surname>Li</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Named Entity Recognition in Query</article-title>
          .
          <source>En Proceedings of the 32nd international ACM SIGIR conference on Research and development in information retrieval, pÆginas 267274</source>
          , Boston, Massachusetts, USA.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Indurkhya</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <year>2014</year>
          .
          <article-title>Natural Language Processing</article-title>
          . En T. Gonzalez J. Daz-Herrera, y A. Tucker, editores, Computing Handbook, Third Edition: Computer Science and Software Engineering. CRC Press, captulo
          <volume>40</volume>
          , pÆginas
          <volume>40</volume>
          :
          <fpage>117</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Ji</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          , J. Nothman, y
          <string-name>
            <given-names>B.</given-names>
            <surname>Hachey</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Overview of TAC-KBP2014 Entity Discovery</article-title>
          and
          <string-name>
            <given-names>Linking</given-names>
            <surname>Tasks</surname>
          </string-name>
          .
          <source>En Proceedings of Text Analysis Conference.</source>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Ji</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          , J. Nothman, y
          <string-name>
            <given-names>B.</given-names>
            <surname>Hachey</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Overview of TAC-KBP2015 Entity Discovery</article-title>
          and
          <string-name>
            <given-names>Linking</given-names>
            <surname>Tasks</surname>
          </string-name>
          .
          <source>En Proceedings of Text Analysis Conference</source>
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Jin</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          , H. Hay Ho, y
          <string-name>
            <given-names>R. K.</given-names>
            <surname>Srihari</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>OpinionMiner: A Novel Machine Learning System for Web Opinion Mining and Extraction</article-title>
          .
          <source>En Proceedings of the 15th ACM SIGKDD international conference on Knowledge discovery and data mining, pÆginas 11951204</source>
          , Paris, France.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Kitoogo</surname>
            ,
            <given-names>F. y V.</given-names>
          </string-name>
          <string-name>
            <surname>Baryamureeba</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Towards domain independent named entity recognition. En Strengthening the Role of ICT in Development, volumen IV</article-title>
          .
          <source>Fountain publishers, pÆginas 84 95.</source>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Konkol</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Brychcn</surname>
          </string-name>
          , Konop, y
          <string-name>
            <surname>M. K.</surname>
          </string-name>
          <year>2015</year>
          .
          <article-title>Latent semantics in Named Entity Recognition</article-title>
          .
          <source>Expert Systems with Applications</source>
          ,
          <volume>42</volume>
          (
          <issue>7</issue>
          ):
          <fpage>34703479</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          , Y.-G. Hwang, y M.
          <article-title>-</article-title>
          G. Jang.
          <year>2007</year>
          .
          <article-title>Fine-Grained Named Entity Recognition and Relation Extraction for Question Answering</article-title>
          .
          <source>En Proceedings of the 30th annual international ACM SIGIR conference on Research and development in information retrieval , pÆginas 799800</source>
          , Amsterdam, The Netherlands.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          -H.,
          <string-name>
            <given-names>Y. G.</given-names>
            <surname>Hwang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. J.</given-names>
            <surname>Oh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Heo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. H.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. J.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. H.</given-names>
            <surname>Wang</surname>
          </string-name>
          , y
          <string-name>
            <given-names>M. G.</given-names>
            <surname>Jang</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Fine-grained Named Entity Recognition using Conditional Random Fields for Question Answering</article-title>
          .
          <source>En Information Retrieval Technololgy, Proceedings , volumen 4182</source>
          . Springer, Berlin, Heidelberg, pÆginas
          <volume>581587</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>Manaris</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <year>1998</year>
          .
          <article-title>Natural Language Processing: A Human-Computer Interaction Perspective</article-title>
          .
          <source>En Advances in Computers , volumen 47. pÆ- ginas 166.</source>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <surname>Marrero</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Urbano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>SÆnchez-Cuadrado</surname>
          </string-name>
          , J. Morato,
          <string-name>
            <given-names>y J. M.</given-names>
            <surname>Gmez-Berbs</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Named Entity Recognition: Fallacies, challenges and opportunities</article-title>
          .
          <source>Computer Standards and Interfaces</source>
          ,
          <volume>35</volume>
          (
          <issue>5</issue>
          ):
          <fpage>482489</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <surname>Moreno</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          , E. Boldrini, P. Moreda, y M.
          <string-name>
            <surname>T.</surname>
          </string-name>
          RomÆ-Ferri.
          <year>2017</year>
          .
          <article-title>Drugsemantics: A corpus for named entity recognition in spanish summaries of product characteristics</article-title>
          .
          <source>Journal of Biomedical Informatics</source>
          ,
          <volume>72</volume>
          :
          <fpage>8</fpage>
          22.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <surname>Moreno</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          , P. Moreda, y M.
          <string-name>
            <surname>T.</surname>
          </string-name>
          RomÆ-Ferri.
          <year>2016</year>
          .
          <article-title>An active ingredients entity recogniser system based on proles</article-title>
          .
          <source>En 21st International Conference on Applications of Natural Language to Information Systems</source>
          , volumen 9612 de LNCS, pÆginas
          <volume>276284</volume>
          , Salford. Springer.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <surname>Moreno</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          , P. Moreda, y M.
          <string-name>
            <surname>T.</surname>
          </string-name>
          RomÆ-Ferri.
          <year>2015</year>
          . Estudio de abilidad y viabilidad de
          <source>la Web 2</source>
          .
          <article-title>0 y la Web semÆntica para enriquecer lexicones en el dominio farmacolgico</article-title>
          .
          <source>Procesamiento del Lenguaje Natural</source>
          ,
          <volume>55</volume>
          :
          <fpage>6572</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <surname>Moreno</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          , P. Moreda, y M.
          <string-name>
            <surname>RomÆ-Ferri</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Reconocimiento de entidades nombradas en dominios restringidos</article-title>
          .
          <source>En Actas del III Workshop en Tecnologas de la InformÆtica . pÆginas 4157.</source>
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <surname>Moreno</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          , P. Moreda, y M.
          <string-name>
            <surname>RomÆ-Ferri</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>MaNER: a MedicAl Named Entity Recogniser for Spanish</article-title>
          .
          <source>En 20th International Conference on Applications of Natural Language to Information Systems</source>
          , volumen 9103 de LNCS, pÆginas
          <volume>418423</volume>
          , Passau. Springer.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <surname>Moreno</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>M. T.</surname>
            RomÆ-Ferri,
            <given-names>y P.</given-names>
          </string-name>
          <string-name>
            <surname>Moreda</surname>
          </string-name>
          . 2017a.
          <article-title>Language independent proposal to prole-based named entity classication</article-title>
          .
          <source>En The First Workshop on Multi-Language Processing in a Globalising World , pÆginas 2130</source>
          , Dublin.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <surname>Moreno</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>M. T.</surname>
            RomÆ-Ferri,
            <given-names>y P.</given-names>
          </string-name>
          <string-name>
            <surname>Moreda</surname>
          </string-name>
          . 2017b.
          <article-title>Named entity classication based on proles: A domain independent approach</article-title>
          .
          <source>En 22nd International Conference on Applications of Natural Language to Information Systems</source>
          , volumen 10260 de LNCS, pÆginas
          <volume>142</volume>
          146, Lieja. Springer.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <surname>Moreno</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>M. T.</surname>
            RomÆ-Ferri,
            <given-names>y P.</given-names>
          </string-name>
          <string-name>
            <surname>Moreda</surname>
          </string-name>
          . 2017c. Propuesta de un sistema de clasicacin de
          <article-title>entidades basado en perles e independiente del dominio</article-title>
          .
          <source>Procesamiento del Lenguaje Natural</source>
          ,
          <volume>59</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <string-name>
            <surname>Moreno</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          , M. RomÆ-Ferri,
          <string-name>
            <given-names>y P.</given-names>
            <surname>Moreda</surname>
          </string-name>
          .
          <year>2017d</year>
          .
          <article-title>A domain and language independent named entity classication approach based on proles and local information</article-title>
          .
          <source>En Recent Advances in Natural Language Processing , pÆginas 510 518</source>
          ,
          <string-name>
            <surname>Varna</surname>
          </string-name>
          (To appear).
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <surname>Moreno</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Palomar</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Molina, y A.
          <source>FernÆndez</source>
          .
          <year>1999</year>
          .
          <article-title>Introduccin al procesamiento del lenguaje natural</article-title>
          . Publicaciones Universidad de Alicante.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <surname>Nadeau</surname>
            , D. y
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Sekine</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>A survey of named entity recognition and classication</article-title>
          .
          <source>Lingvisticae Investigationes</source>
          ,
          <volume>30</volume>
          (
          <issue>1</issue>
          ):326, jan.
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          <string-name>
            <surname>Peregrino</surname>
            ,
            <given-names>F. S.</given-names>
          </string-name>
          , D. TomÆs, y
          <string-name>
            <given-names>F. L.</given-names>
            <surname>Pascual</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Question Answering and Multi-search Engines in Geo-Temporal Information Retrieval</article-title>
          . En A. Gelbukh, editor,
          <source>Computational Linguistics and Intelligent Text Processing: 13th International Conference, CICLing</source>
          <year>2012</year>
          , New Delhi, India, March
          <volume>11</volume>
          -17,
          <year>2012</year>
          , Proceedings, Part II. Springer, Berlin, Heidelberg, pÆginas 342
          <fpage>352</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          <string-name>
            <surname>Pradhan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Elhadad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. W.</given-names>
            <surname>Chapman</surname>
          </string-name>
          , S. Manandhar, y G. Savova.
          <year>2014</year>
          .
          <article-title>SemEval2014 Task 7: Analysis of Clinical Text</article-title>
          . pÆginas 5462.
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          <string-name>
            <surname>Sang</surname>
            ,
            <given-names>E. F.</given-names>
          </string-name>
          <string-name>
            <surname>T. K. y F. De Meulder</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>Introduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition</article-title>
          .
          <source>Proceedings of the 7th Conference on Natural Language Learning</source>
          , pÆginas
          <volume>142147</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          <string-name>
            <surname>Segura-Bedmar</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          , P. Martnez, y M.
          <year>HerreroZazo</year>
          .
          <year>2013</year>
          . SemEval
          <article-title>-2013 Task 9: Extraction of Drug-Drug Interactions from Biomedical Texts</article-title>
          (DDIExtraction
          <year>2013</year>
          ).
          <source>En Proceedings of the 7th International Workshop on Semantic Evaluation</source>
          , pÆginas
          <volume>341350</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          <string-name>
            <given-names>Tjong</given-names>
            <surname>Kim Sang</surname>
          </string-name>
          ,
          <string-name>
            <surname>E. F.</surname>
          </string-name>
          <year>2002</year>
          .
          <article-title>Introduction to the CoNLL-2002 shared task</article-title>
          .
          <source>En Proceeding of the 6th Conference on Natural Language Learning.</source>
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          <string-name>
            <surname>Tkachenko</surname>
            , M. y
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Simanovsky</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Selecting Features for Domain-Independent Named Entity Recognition</article-title>
          .
          <source>En Proceedings of KONVENS</source>
          <year>2012</year>
          , pÆginas 248253.
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          <string-name>
            <surname>Uzuner</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          , I. Solti,
          <string-name>
            <given-names>y E.</given-names>
            <surname>Cadag</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Extracting medication information from clinical text</article-title>
          .
          <source>Journal of the American Medical Informatics Association</source>
          ,
          <volume>17</volume>
          (
          <issue>5</issue>
          ):
          <fpage>5148</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          <string-name>
            <surname>Vicente</surname>
            , M. y
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Lloret</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Exploring Flexibility in Natural Language Generation throughout Discursive Analysis of New Textual Genres</article-title>
          .
          <source>Proceedings of the 2nd International Workshop Future and Emerging Trends in Language Technologies, Machine Learning and Big Data (FETLT) .</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>