<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Keyword Extraction and Technology Entity Extraction for Disruptive Technology Policy Texts</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Aofei Chang †</string-name>
          <email>chaf@pku.edu.cn</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bolin Hua †</string-name>
          <email>huabolin@pku.edu.cn</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dahai Yu</string-name>
          <email>1900016644@pku.edu.cn</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Information, Management, Peking University</institution>
          ,
          <addr-line>Beijing</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <fpage>36</fpage>
      <lpage>40</lpage>
      <abstract>
        <p>The rapid development of disruptive technologies has attracted the attention of major countries in the world in recent years, and the mining and research on the texts of disruptive technology policies of these countries can reveal the key layout, focus areas, and development pattern of each countries disruptive technology. This article first crawls the texts of disruptive technologies from the science and technology policy websites of major countries. Then, the text is segmented by Spacy, the segment result is filtered by a word list to construct an applicable TF*IDF matrix, and finally the matrix weights are optimized with manually collected domain core words and important words. After these, extraction and statistics of technical entity are performed according to a specified word list. Through comprehensive analysis, it can be found that the keyword hotspots of the experimental texts are focused on artificial intelligence, information security, new energy, etc. The key areas of specific disruptive technologies are artificial intelligence, air and space, and new generation communication technologies. The result reflects the current situation and policy focus of disruptive technology development in these countries.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>CCS CONCEPTS</title>
      <p>• Computing methodologies • Artificial intelligence • Natural
language processing • Information extraction
disruptive technology, science and technology policy, keywords
extraction, technology entity extraction
Science and technology field are changing rapidly, especially under
the wave of big data and artificial intelligence. Some countries have
introduced many S&amp;T policies to promote the development of S&amp;T.
2. Unlike academic papers, S&amp;T policy texts do not carry keywords,
so keywords need to be extracted from a large number of long texts
of S&amp;T policies. The text of S&amp;T policy will involve both S&amp;T
projects, innovation mechanisms, transformation of results,
industrial support, S&amp;T rewards, S&amp;T management and
configuration, but also other specific S&amp;T content, which are the
research frontier, high-tech, industry common technology,
disruptive technologies, and so on.</p>
      <p>
        Disruptive technology was proposed by Harvard University
Professor Christensen in 1997 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and have become a hot topic of
interest for international institutions and researchers in recent years.
      </p>
      <p>It is generally believed that disruptive technologies are strategic
innovative technologies that open up new technological tracks
based on new principles, combinations and applications of S&amp;T,
and produce an overall or fundamental replacement for traditional
or mainstream technologies. Disruptive technologies have strong
application capabilities, can enhance the scientific and
technological competitiveness of enterprises and even countries,
promote the renewal of scientific and technological products,
improve social production efficiency, and are expected to have
farreaching impact in many fields. Disruptive technology policies can
stimulate technological innovation and provide corresponding
support and guarantee, so it is necessary to study the mining of
disruptive technology policies text.</p>
      <p>Named entity extraction (NER) is a hot topic in the field of
natural language processing (NLP), which specifically includes
unsupervised or supervised recognition of specialized domain
vocabulary such as names of people, places, time, products, and
organization names, and textual keyword extraction on this basis is
also an important application. Named entity extraction is usually
the first step in intelligent information retrieval, relationship
extraction, and even knowledge graph construction.
Domainspecific entity extraction is also the key to explore the trends and
hotspots in the field and to build dynamic knowledge graphs.</p>
      <p>Keyword extraction is pivotal in NLP, and a few keywords can
effectively reveal the main contents and topics of long documents.</p>
      <p>
        The study of S&amp;T policy texts has also been paid attention by
some scholars in recent years, Alan Porter proposed that Tech
Copyright 2021 for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0).
Mining makes exploitation of text databases meaningful to those
who can gain from derived knowledge about emerging
technologies. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] And he came up with “science overlay maps” as
a new tool for research policy and library management. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] Scholars
such as Wen Zeng argue that China's S&amp;T policies play an
important role in promoting economic and social development，so
they make use of semantic technologies to extract and analyze the
relatively important information from massive S&amp;T policies in
China, but they mainly focus on exploratory terms and sentences
extraction. [
        <xref ref-type="bibr" rid="ref4 ref5">4</xref>
        ] T Dmitrievna focused on the construction of a
corpus in the field of S&amp;T, he suggests that corpus linguistics tools
can be applied to solve the problem of aerospace terminology
which presents a good idea to better study the scientific and
technical policy text. [5] Michael T. Gorczyca and other scholars
stress value in leveraging language representation models (LRMs)
on domain-specific text corpora for domain-specific tasks, which
helps solve the problem that algorithms for developing text mining
models require a large amount of training data. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] SV Podolkova
considers classification of scientific and technical texts based on
the criterium of text communicative purport, he focused on
structure and composition peculiarities, exact definitions and clear
organization of text representation. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]
      </p>
      <p>At present, although there has been a lot of research on S&amp;T
policy, there is no specialist algorithm and thesaurus for disruptive
technology policies, much less extracting keywords from these
policies documents.</p>
      <p>
        Keyword extraction algorithms mainly include unsupervised
methods and supervised methods. Unsupervised algorithms are
based on statistical features of the text, such as the TF*IDF
algorithm based on the word frequency of the document collection,
and Campos, Ricardo et al. had proposed the YAKE algorithm,
which makes full use of the basic features of the text, it uses 5
features: “Position of Word in the text”,” Word frequency”,” Term
Relatedness to Context” and “Term Different Sentence”. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]
Based on the algorithmic idea of PageRank, the graph-based
keyword extraction algorithm TextRank emerged, and later
Xiaojun Wan and Jianguo Xiao proposed the Expand Rank
algorithm, which extended the TextRank algorithm from
independent documents to a collection of similar documents. Based
on TextRank.[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] Corina Florescu assign larger weights to words
that are found early in a document, which is the core thought of the
Position Rank. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] The algorithm based on text features and
document graph structure does not require a large amount of data
annotation, but it does not completely consider the semantic
relationship between words and documents, and it is difficult to
make a breakthrough in accuracy.
      </p>
      <p>
        Deep learning, as an emerging supervised approach, provides
new ideas in keyword extraction. The basic approach is to vectorize
the candidate words by pre-training the model with word
embedding and then calculate the similarity to the document. Based
on this, the Embed Rank algorithm was proposed by
BennaniSmires et al. it uses both Sen2Vec and Doc2Vec, then a score will
be calculated by the Maximal Marginal Relevance (MMR) formula
which take into account similarity between text and phrase as well
as diversity of the keywords set.[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] After the powerful BERT
model is proposed, Grootendorst proposed KeyBERT's keyword
extraction algorithm in 2020，it extracts phrases that have better
cosine similarity to the document vector which is produced by
using pre-trained domain-specific BERT model.[
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]
      </p>
      <p>Although deep learning helps improve extraction precision,
training such models often requires large amounts of annotated data,
which is expensive to gather. In the case of keyword research in
policy texts, especially in S&amp;T policy texts, the data are mostly
semi-structured long texts, and large-scale annotation training
would be very difficult. In addition, for the traditional unsupervised
algorithm, it will not have a high accuracy rate in a specific field,
especially in our focus on S&amp;T, and the extracted keywords are
often not close to the topic of S&amp;T.</p>
      <p>To address these issues, we combine automatic word separation
and technology domain word lists to achieve simple and fast entity
extraction in the technology domain, including the extraction of
named entities, and based on this, we achieve keyword extraction
by the weighted optimized TF*IDF algorithm, and also perform
simple statistical analysis of disruptive technology entities. Our
approach combines supervised and unsupervised algorithms, and
since it focuses only on the technology domain, it is relatively easy
to build our corpus, and the keyword extraction based on the
technology domain word list itself focuses on the thematic and
semantic relationships between keywords and documents.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Method design</title>
      <p>By researching major global S&amp;T policy websites, we determined
the search strategy and key websites, followed by automatic
crawling of S&amp;T policies using python to form a database of
disruptive technology policy texts. Afterwards, we judged the
initial data for disruptive technology relevance and screened out
1005 policy texts with high relevance. For these experimental texts,
we performed automatic keyword extraction and specific disruptive
technology extraction, respectively. For keyword extraction, we
performed automatic word segmentation by self- designed
automatic word segmentation based on spacy, then determined the
TF*IDF matrix of candidate words and the overall text collection
by screening the domain keyword list, and finally optimized the
weights based on the self- designed core words and technology.
The final results and rankings are calculated based on the
selfdesigned core words and technology related word list for weight
optimization.</p>
    </sec>
    <sec id="sec-3">
      <title>2.1 Text collection and relevance judgement</title>
      <p>The main large websites crawled are the UK government website,
the EU publications website, the US Center for Strategic and
International Studies, etc.</p>
      <p>In terms of relevance evaluation, considering that the
occurrence of keywords is not limited to fixed pairings, for example,
disruptive and technology may not appear next to each other, but if
the keywords are split and then matched statistically, the accuracy
rate will be reduced, so we took a compromise approach, that is, to
detect whether keyword pairs appear in a sentence.</p>
      <p>In general, the keywords in the title and abstract of an article are
relatively more important, so the content of the article is divided
into three parts according to the position: title, abstract and body,
and the corresponding weights are 5,3,1. In addition, only
considering the word frequency will ignore the influence of the
length of the article on the score, so a balancing factor of the
average article length is added to the formula.</p>
      <p>Here is the formula for calculating the correlation score:
The meaning of each of these symbols is as follows.</p>
      <p>N： Total number of keywords in a document
b：A free constant that specifies how much the document
length affects the score
k：Free constants to specify the upper limit of the impact of a
single word on the rating
ld：Length of the document
lavg： Average document length</p>
      <p>The total number of texts crawled was over 10,000, and 1005 of
them were identified as experimental texts after relevance
screening.
2.2 The specific process of keyword extraction
2.2.1 Words segmentation and entity extraction. Specific
process: document reading and sentence slicing, using Spacy
natural language processing tools, specifically using Spacy's
dependent syntactic analysis to determine the predicate of each
sentence, that is, the root of a sentence, as one of the basis for
slicing, using Spacy's deactivation table to mark the deactivation
words, and then mark some specific lexical words, specifically
including'ADV','AUX','CONJ','INTJ','NUM','PRON','SYM','SCO
NJ','PREP'. The above words are used as cut nodes to split the
words and get the preliminary splitting results.</p>
      <p>2.2.2 Disruptive technology lexicon construction and weighted
TF*IDF calculation. The first step is the identification of core
terms, key terms and domain word lists. Considering that the topics
of the texts we collected are all about technology policy and
disruptive technologies, when extracting keywords we need to
prioritize the words that directly match this topic, so we first
defined 27 core terms as follows: disruptive innovation, radical
innovation, innovation, disruptive technology, incremental
innovation, open innovation, new product development, business
model, absorptive capacity, technological innovation, developing
technology, advanced technology, integrated technology, future
technology, promising technology, next generation technology,
evolving technology, radical technology, Next Big Thing, radical
technology, breakthrough technology, game changer, gaming
changing technology, emerging technology, revolutionary
technology, transformative technology.</p>
      <p>A crawler program on Web of Science was used to obtain the
search results of the advanced search formula “TS = (disruptive
technology OR disruptive innovation)”, then we extracted all the
keyword fields and performed a simple word separation process
and lemmatization. Then the keyword frequency statistics were
arranged in descending order. The keyword statistics were then
used as the base keyword list, followed by manual identification to
construct a keyword list of specific disruptive technologies. Finally,
7400 base keywords and more than 300 non-repetitive
technologyspecific words was collected.</p>
      <p>Since Web of Science integrates academic journals, invention
patents, academic conferences, academic websites and various
other high-quality information resources to provide academic
information in multiple fields, it is accurate and reasonable to grasp
the dynamics of disruptive technologies and keywords in academia
through Web of Science.</p>
      <p>Finally, we constructed a TF*IDF matrix of 1005*5246 based
on the results of word separation to determine the words and word
frequencies for each text. For the core words, we multiplied the
TF*IDF results by 20 as weights, and the words appearing in the
specific technical word list were multiplied by 10 as weights.
Finally, the top ten words in weight for each document were filtered
as candidate keywords.</p>
    </sec>
    <sec id="sec-4">
      <title>2.3 The simple process of technology entity extraction</title>
      <p>Through the previous collection of more than 300 disruptive
technology entity words, we used python's FastText tool to match
these entity words to get the results of each text technology entity
extraction. The process of the entity extraction part does not have
much detail and focuses on the analysis of the results</p>
    </sec>
    <sec id="sec-5">
      <title>3 Analysis and measurement of results</title>
    </sec>
    <sec id="sec-6">
      <title>3.1 Keyword extraction algorithm measurement</title>
      <p>In order to effectively measure keyword extraction algorithms, we
will compare them with some mainstream unsupervised keyword
extraction algorithms. These algorithms include: Yake, TextRank,
KeyBert.</p>
      <p>3.1.1 P/R/F-Score comparison without considering keyword
order.
they are domain-independent and more applicative. However, since
our proposed algorithm is supervised algorithm based on domain
word lists, it has some degree of advantage in terms of accuracy
and recall. If these unsupervised algorithms are optimized in
combination with domain word lists, they will be more effective.</p>
    </sec>
    <sec id="sec-7">
      <title>3.2 Keyword and technical entity extraction results</title>
      <p>We conducted statistical analysis on the keyword extraction results,
selected the top 120 keywords in terms of word frequency, and
generated the word cloud map by python program as follows</p>
      <p>Figure 3 Classification statistics of hot technical words
Observing the word cloud map, it can be seen that the
experimental text has numerous disruptive technology hotspots,
involving information security, artificial intelligence, Internet of
Things, new energy and even military technology fields.</p>
      <p>We performed a simple count of high frequency words from the
results of disruptive technology entity extraction and categorized
them by domain.</p>
      <p>As Figure 5 shows, the hot high frequency words in the chart
mainly belong to the fields of artificial intelligence, air and space
technology and new generation communication technology, such
as AI and Machine Learning in the field of artificial intelligence,
aircraft and satellite in the field of air and space technology and
internet, ICT and 5G in the new generation communication
technology, which all reflect the technology hotspots in these fields .</p>
    </sec>
    <sec id="sec-8">
      <title>4 Conclusion and Discussion</title>
      <p>As for the disruptive technology entity extraction part, the method
is mainly depend on a manually designed words list, so it has great
scope for improvement. There are existing excellent named entity
recognition (NER) algorithms that can be utilized, for example the
currently popular named entity recognition model of
BILSTMCRF via BERT pre-training. And Emma Strubell et al. first used
IDCNN for entity recognition, reducing the time complexity of the</p>
      <p>In the keyword extraction part, considering the special
characteristics of scientific and technical texts, i.e., the focus topic
of the text is often the current development of a scientific and
technical field or technology, we enlarge the weight of scientific
and technical terms so that the keyword extraction procedure can
be more sensitive to these field terms and often extract the less
frequent but important scientific and technical field terms. The
results of the automatic word screening are intuitive and measured
well because of the use of a manually developed word list and the
increased weight of technical terms in the calculation of TF*IDF.</p>
      <p>However, considering the subjective nature of the manually
developed word list, the fact that the content of the word list limits
the determination of candidate keywords and phrases, and the fact
that the field of S&amp;T is developing rapidly and the nouns of S&amp;T
are changing rapidly, it is not possible to extract technical noun
entities and candidate keywords by relying on the manual word list
alone.</p>
      <p>The keyword algorithm is insufficient: it does not have the
ability to extract new words, and the extraction results are overly
dependent on the word list, so it needs a great degree of
optimization. Optimization direction: Later on, neural networks
and deep learning algorithms can be used to extract a wider range
of noun entities in the field of S&amp;T in combination with the word
list, and the ability to extract words outside the word list is
enhanced.</p>
      <p>In specific disruptive technology extraction part, the extraction
in this paper relies entirely on the manually developed word list,
and later we can also rely on deep learning algorithms for
optimization. Through stronger autonomous extraction, statistical
analysis, we will optimize discovery of hotspots in S&amp;T fields and
make preparation for domain knowledge extraction.</p>
    </sec>
    <sec id="sec-9">
      <title>ACKNOWLEDGMENTS</title>
      <p>This work was supported in part by The National Social Science
Fundation of China (Number: 17BTQ066).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Christensen</surname>
            ,
            <given-names>Clayton. "M.</given-names>
          </string-name>
          (
          <year>1997</year>
          )
          <article-title>The Innovators Dilemma: When New Technologies Cause Great Firms to Fail." Harvard Business Review Press (</article-title>
          <year>1997</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Porter</surname>
            , Alan L.,
            <given-names>and Scott W.</given-names>
          </string-name>
          <string-name>
            <surname>Cunningham</surname>
          </string-name>
          .
          <article-title>Tech mining: exploiting new technologies for competitive advantage</article-title>
          . Vol.
          <volume>29</volume>
          . John Wiley &amp; Sons,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Rafols</surname>
          </string-name>
          , Ismael, Alan L.
          <string-name>
            <surname>Porter</surname>
            , and
            <given-names>Loet</given-names>
          </string-name>
          <string-name>
            <surname>Leydesdorff</surname>
          </string-name>
          .
          <article-title>"Science overlay maps: A new tool for research policy and library management</article-title>
          .
          <source>" Journal of the American Society for information Science and Technology 61.9</source>
          (
          <year>2010</year>
          ):
          <fpage>1871</fpage>
          -
          <lpage>1887</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Zeng</surname>
            , Wen,
            <given-names>Changqing</given-names>
          </string-name>
          <string-name>
            <surname>Yao</surname>
            , and
            <given-names>Hui</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
          </string-name>
          .
          <article-title>"The exploration of information extraction and analysis about science and technology policy in China." The Electronic Library (</article-title>
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Butenko</surname>
            ,
            <given-names>Iuliia</given-names>
          </string-name>
          <string-name>
            <surname>Ivanovna</surname>
          </string-name>
          , Tatiana Dmitrievna Margaryan, and Elizaveta Evgeneevna Bolotova.
          <article-title>"Scientific and Technical Text Corpus as the Basis for Aerospace Terminology Standardization."</article-title>
          <source>Applied Linguistics Research Journal</source>
          <volume>5</volume>
          .3:
          <fpage>113</fpage>
          -
          <lpage>119</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Gorczyca</surname>
          </string-name>
          , Michael T., et al.
          <article-title>"A comparison of language representation models on small text corpora of scientific and technical documents." Artificial Intelligence and Machine Learning for Multi-Domain Operations Applications II</article-title>
          . Vol.
          <volume>11413</volume>
          .
          <source>International Society for Optics and Photonics</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Podolkova</surname>
            ,
            <given-names>S. V.</given-names>
          </string-name>
          "
          <source>Principles of Scientific and Technical Text Analysis." Філол огічні трактати 10</source>
          ,№
          <volume>2</volume>
          (
          <year>2018</year>
          ):
          <fpage>90</fpage>
          -
          <lpage>100</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Campos</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ricardo</surname>
          </string-name>
          , et al.
          <article-title>"YAKE! Keyword extraction from single documents using multiple local features."</article-title>
          <source>Information Sciences</source>
          <volume>509</volume>
          (
          <year>2020</year>
          ):
          <fpage>257</fpage>
          -
          <lpage>289</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Wan</surname>
            , Xiaojun, and
            <given-names>Jianguo</given-names>
          </string-name>
          <string-name>
            <surname>Xiao</surname>
          </string-name>
          .
          <article-title>"Single Document Keyphrase Extraction Using Neighborhood Knowledge."</article-title>
          <source>AAAI</source>
          . Vol.
          <volume>8</volume>
          .
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Florescu</surname>
            , Corina, and
            <given-names>Cornelia</given-names>
          </string-name>
          <string-name>
            <surname>Caragea</surname>
          </string-name>
          .
          <article-title>"Positionrank: An unsupervised approach to keyphrase extraction from scholarly documents." Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics</article-title>
          (Volume
          <volume>1</volume>
          :
          <string-name>
            <surname>Long</surname>
            <given-names>Papers).</given-names>
          </string-name>
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Bennani-Smires</surname>
          </string-name>
          , Kamil, et al.
          <article-title>"Simple unsupervised keyphrase extraction using sentence embeddings." arXiv preprint arXiv:</article-title>
          <year>1801</year>
          .
          <volume>04470</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Grootendorst</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>"KeyBERT: minimal keyword extraction with BERT</article-title>
          ,
          <year>v0</year>
          .
          <fpage>1</fpage>
          .3.
          <string-name>
            <surname>"</surname>
          </string-name>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Strubell</surname>
          </string-name>
          ,
          <string-name>
            <surname>Emma</surname>
          </string-name>
          , et al.
          <article-title>"Fast and accurate entity recognition with iterated dilated convolutions</article-title>
          .
          <source>" arXiv preprint arXiv:1702</source>
          .
          <year>02098</year>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>