<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Open Domain Named Entity Discovery and Linking Task</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name> Yeqiang Xu</string-name>
          <email>yeqiang@summba.com</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zhongmin Shi</string-name>
          <email>shi@summba.com</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Peipeng Luo</string-name>
          <email>peipeng@summba.com</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yunbiao Wu</string-name>
          <email>yunbiao@summba.com</email>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Summba Inc.</institution>
          ,
          <addr-line>Guangzhou</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <fpage>2</fpage>
      <lpage>7</lpage>
      <abstract>
        <p>This paper describes a named entity discovery and linking system, which compete the CCKS2017 question named entity discovery and linking task. We are facing challenges including short-text, small training samples and open domain, making the existing solutions unfeasible. In this paper, we propose a CRF + rules method to recognize the corresponding named entity, which employs several features, such as bag-of-word features, POS features, parsing features etc. As for entity linking, context information, popularity, word embeddings, and online public corpus are used. The experiment results show that, the F1 score of named entity discovery is 0.815, while the accuracy of the entity linking is 0.736. The overall F1 score is 0.600, which proves the effectiveness of our system.</p>
      </abstract>
      <kwd-group>
        <kwd>Named Entity Discovery</kwd>
        <kwd>Entity Linking</kwd>
        <kwd>Context Information</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        With the development of Internet techniques, unstructured text has become one of the
most popular information carriers. As the basis of text analysis, Named Entity
Discovery (NED) and Entity Linking (EL) are widely studied. However, the
traditional NED techniques can only be applied to the situation with very few entity
types (e.g. persons, locations, organizations etc.), and the existing EL methods need
rich context information. In this task, we mainly encounter three difficulties. Firstly,
the boundary of Named Entity (NE) definition is fuzzy. For example, “苹果手机”
(Apple iPhone) is not a named entity, but “苹果” (Apple) is a named entity as a
company name; The second challenge is that the questions in training set are
short-texts with lots of noises; The last challenge is that the corpus is from open
domain and entity types are relatively much more.
Our work can be divided into two parts: NED and EL. The existing approaches of NED
are mainly based on statistical models, e.g. HMM (Hidden Markov Model), CRF
(Conditional Random Field) and DNN (Deep Neural Network) etc[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ][
        <xref ref-type="bibr" rid="ref2">2</xref>
        ][
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Most
systems of EL adopt supervised methods to disambiguate, including binary
classification[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ][
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and Machine-Learned Ranking-based EL (MLR-based EL)[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ][
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
Some studies suppose NED and EL are related sub-task, and achieved a high
performance by optimized the joint model[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>However, the statistical models of NED principally focus on very few entity types,
it is unsuitable in open domain. As for EL, supervised approach is difficult to apply,
with short-text and limited information about each sense of ambiguous named entity.</p>
    </sec>
    <sec id="sec-2">
      <title>Named Entity Discovery &amp; Entity Linking</title>
      <sec id="sec-2-1">
        <title>Named Entity Discovery</title>
        <p>
          Our strategy is CRF + rules in NED. Firstly, bag-of-word features and POS features
from the training corpus are extracted with the pyltp tool[
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. The training corpus is then
transformed into character-level corpus, and divided into training set and test set. Also,
it is necessary to further filter, split and re-identify them based on rules.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Rule 1: Filter Rules</title>
        <p>The filter rules are divided into the following categories (more than 180 cases are
included, while the size of training set is 1395):
a) Version number filtering: the product number should be included as part of the
named entities, while the version number should not.
b) Suffix filtering: some suffixes affecting NED must be filtered. For instance, for
the phrase "戴尔笔记本" (Dell Laptop), the real NE is "戴尔" (Dell), while "笔
记本" ( Laptop) is a generic entity cannot be included.
c) Verb-prefix filtering: in the prediction results, there are very few named entities
which has structure of "verb + noun", in which case the verb prefix needs to be
filtered out.</p>
      </sec>
      <sec id="sec-2-3">
        <title>Rule 2: Split Rules</title>
        <p>Among the predicted named entities, if conjunctions exist, such as "and" etc., the
results need to be split into more parts.</p>
      </sec>
      <sec id="sec-2-4">
        <title>Rule 3: Re-identify Rules</title>
        <p>The CRF model-based result can be revised by re-identify rules. In the beginning, a
named entity dictionary in the training set is constructed. By using the dictionary, we
re-identify a new result based on all-matching rule. If the re-identified result does not
overlap with the model-based result, it can be added to the final result set. In the case
of overlap, when the model-based result is contained in re-identified result, the
re-identified result should be added to the final result set. In other cases, the result is
based merely on the model-based result. Just as one example, assuming the
model-based result is “百度知道” (Baidu Zhidao), while the re-identified result is “百
度知道企业平台” (Baidu Zhidao Enterprise Platform), the final result can be revised
as “百度知道企业平台” (Baidu Zhidao Enterprise Platform).</p>
      </sec>
      <sec id="sec-2-5">
        <title>Entity Linking</title>
        <p>Several useful features can be used for EL, such as entity context information in
original sentence, popularity of each entity’s sense, word co-occurrence in Baidu
Zhidao1 for each entity’s sense etc.</p>
      </sec>
      <sec id="sec-2-6">
        <title>Information Extension</title>
        <p>In this task, information extension is necessary for the reason that each ambiguous
named entity has several senses with less information. We collect each sense’s
relative questions as candidate question sets by searching Baidu Zhidao and picking
out the questions in top 5 pages. For example, a sense from knowledgeworks2 calling
“小米（小米公司）” (Xiaomi Inc.) is extended to “小米公司的经营理念是什么?”
(What is the business philosophy of Xiaomi Inc.) as one candidate question by
searching Badidu Zhidao.</p>
      </sec>
      <sec id="sec-2-7">
        <title>Information Filter</title>
        <p>Firstly, all questions are segmented. Then stop words are removed. At last, the
similarity of original question with each candidate question is calculated by using
Jaccard Similarity, and those candidate questions with low similarity are removed.</p>
      </sec>
      <sec id="sec-2-8">
        <title>Context Similarity</title>
        <p>
          The original and candidate questions need to be represented as vectors. In a question,
each word’s vector representation is gained by pre-trained word2vec[
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] model. After
filtration, a sentence has only a few words left, therefore, we simply average word
vectors in a sentence. Let W = {wi} be a word vector set whose size is N. Then the
question’s sentence vector v is presented as formula (1).
        </p>
        <p>v = N1 ∙∑1N wi (1)</p>
        <p>The context similarity score of an entity’s sense can be calculated by the average
cosine similarity of the original question with each one in the candidate set of this
sense. Let o be the original question’s sentence vector, m be one of entity’s sense,
V = {vi} be the sentence vectors in candidate set of the sense whose size is K. The
context similarity score of o and mis shown as formula (2).</p>
        <p>scorecontext(o, m) = 1 ∙ ∑lK=1 similaritcyosine(vo, vl)</p>
        <p>K
(2)</p>
      </sec>
      <sec id="sec-2-9">
        <title>Popularity</title>
        <p>We collect every sense’s visits of a named entity from Baidu Baike3 as one index of
popularity. But, some of entity’s senses in knowledgeworks are different from those
in Baidu Baike, in which case edit distance is needed. For those entity’s senses never
show up in Baidu Baike, we use the minimum score from other entity’s senses to
ensure score’s smoothness. Let o be the original question, M be the entity’s senses
1 https://zhidao.baidu.com/
2 http://knowledgeworks.cn:30001/
3 https://baike.baidu.com/
in knowledgeworks, m ∈ M be one of entity’s sense, k ∈ M, k ≠ m be another
entity’s sense different from m, S be the size of senses in Baidu Baike, “editDist”
be the edit distance of m and the processing sense in Baidu Baike. The score for an
entity’s senseis shown as formula (3).
scorevisit(o, m) = max((1 ∙ ∑iS=1 visitsi ∙ editDist) , (</p>
        <p>S</p>
        <p>min scorevisit(o, k))) (3)
k∈M,k≠m</p>
        <p>For each entity, knowledgeworks provides a frequently-used sense calling primary
sense. Let t be 1 or 0 presenting whether m is primary, g be the weight when m
is not primary for score’s smoothness. The primary score is shown as formula (4).</p>
        <p>scoreprimary(o, m) = g1−t</p>
      </sec>
      <sec id="sec-2-10">
        <title>Word Co-occurrence</title>
        <p>After processing irrelevant information filtration and Chinese segmentation for
extended questions, word co-occurrence frequencies are calculated by counting each
word shown up together in the two extension sets of original questions and candidate
ones in the entity’s senses, and then normalized. Let o be the original question, m
be one of entity’s sense whose size is N, counti, i ∈ N be the word co-occurrence of
arbitrarily one of the entity’s senses. The score is shown as formula (5).
(4)
(5)
(6)
scoreco−occ(o, m) = ∑c1Nocuonutnit , i ∈ N</p>
      </sec>
      <sec id="sec-2-11">
        <title>EL Scoring Method</title>
        <p>According to the four scores described above, the weighting method is shown as
formula (6).</p>
        <p>scorefinal(o, m) = α ∙ scorecontext + β ∙ scorevisits + γ ∙ scoreprimary + μ
∙ scoreco−occ</p>
        <p>s. t. α + β + γ + μ = 1</p>
      </sec>
      <sec id="sec-2-12">
        <title>EL Parameters</title>
        <p>The following Table 1 lists EL parameters for formulas described above.
All parameters in Table 1 are adjusted by repeated testing. The editDist means the
edit distance of senses. the g presents the weight of non-primary parameter. As for
the combination parameters, the sentence context similarity α contributes most to
EL; the visit β and primary γ presenting popularity are assigned the same value; the
word co-occurrence μ also act as an important role for EL.</p>
        <p>Method
CRF
CRF+Rule1
CRF+Rule1+Rule2
CRF+Rule1+Rule2+Rule3</p>
      </sec>
      <sec id="sec-2-13">
        <title>EL Result</title>
        <p>The Table 3 shows EL results.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Evaluation</title>
      <sec id="sec-3-1">
        <title>Dataset and Pre-training of Word Embedding</title>
        <p>
          In this experiment, we use the dataset provided by CCKS2017 Task 1. Besides we
crawl tens of millions of questions from Baidu Zhidao and train a word2vec[
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] model.
4.2
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Experimental Result</title>
      </sec>
      <sec id="sec-3-3">
        <title>Named Entity Discovery Result</title>
        <p>The following Table 2 lists NED results.
According to the result, the best score of NED in training set is based on CRF with
rules, which F1 is 0.815. The top precision of EL comes from the combination of four
features, which is 0.736. So as to the overall F1 is 0.600.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusion and Future Work</title>
      <p>
        In this paper, we propose some strategies for NED and EL to deal with open domain
short-text issues. Experiments show that our method has effective performance. In the
future, we are trying to apply NED by using Heuristic and DNN[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] method. As for
EL, CNN[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] can be considered as research direction.
      </p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <p>This work was supported in the Research on People's Heterogeneous Information
Aggregation Technology, and Research and Development of Intelligent Question
Answering System Based on Text Automatic Abstract Technology, and the Research
on Key Techniques of Knowledge Map in Intelligent Home.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Morwal</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jahan</surname>
            <given-names>N</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chopra</surname>
            <given-names>D.</given-names>
          </string-name>
          <article-title>Named entity recognition using hidden Markov model (HMM)[J]</article-title>
          .
          <source>International Journal on Natural Language Computing (IJNLC)</source>
          ,
          <year>2012</year>
          ,
          <volume>1</volume>
          (
          <issue>4</issue>
          ):
          <fpage>15</fpage>
          -
          <lpage>23</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dai</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Automatic recognition of Chinese organization name based on cascaded conditional random fields</article-title>
          .
          <source>Acta Electronica Sinica</source>
          <volume>34</volume>
          (
          <issue>5</issue>
          ),
          <volume>804</volume>
          (
          <year>2006</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Ma</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hovy</surname>
          </string-name>
          , E.:
          <article-title>End-to-end Sequence Labeling via Bi-directional LSTM-CNNs-CRF</article-title>
          .
          <source>In: Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics</source>
          , pp.
          <fpage>1064</fpage>
          -
          <lpage>1074</lpage>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Su</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tan</surname>
            ,
            <given-names>C-L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
          </string-name>
          , W-T.:
          <article-title>Entity Linking leveraging automatically generated annotation</article-title>
          .
          <source>In: International Conference on Computational Linguistics (COLING</source>
          <year>2010</year>
          ), pp.
          <fpage>1290</fpage>
          -
          <lpage>1298</lpage>
          (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Francis-Landau</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Durrett</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Klein</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Capturing semantic similarity for Entity Linking with Convolutional Neural Networks</article-title>
          .
          <source>In: Proceedings of NAACL-HLT</source>
          , pp.
          <fpage>1256</fpage>
          -
          <lpage>1261</lpage>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Ratinov</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roth</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Downey</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Anderson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Local and global algorithms for disambiguation to Wikipedia</article-title>
          . In:
          <article-title>Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies</article-title>
          -Volume, pp.
          <fpage>1375</fpage>
          -
          <lpage>1384</lpage>
          (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Shen</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Luo</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Linden: linking named entities with knowledge base via semantic knowledge</article-title>
          .
          <source>In: Proceedings of the 21st international conference on World Wide Web</source>
          , pp.
          <fpage>449</fpage>
          -
          <lpage>458</lpage>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Che</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          , Liu,
          <string-name>
            <surname>T.</surname>
          </string-name>
          :
          <article-title>LTP: A chinese language technology platform</article-title>
          .
          <source>In: Proceedings of the 23rd International Conference on Computational Linguistics</source>
          , pp.
          <fpage>13</fpage>
          -
          <lpage>16</lpage>
          (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          .
          <source>Advances in neural information processing systems</source>
          , pp.
          <fpage>3111</fpage>
          -
          <lpage>3119</lpage>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>