<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>June</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>An Examination of the Validity of General Word Embedding Models for Processing Japanese Legal Texts</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Linyuan Tang</string-name>
          <email>linyuan-tang@g.ecc.u-tokyo.ac.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kyo Kageura</string-name>
          <email>kyo@p.u-tokyo.ac.jp</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Graduate School of Interdisciplinary Information Studies, The University of Tokyo</institution>
          ,
          <addr-line>Tokyo</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Interfaculty Initiative in Information Studies, The University of Tokyo</institution>
          ,
          <addr-line>Tokyo</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <volume>21</volume>
      <issue>2019</issue>
      <abstract>
        <p>Thanks to the recent developments in distributed representation learning and the large amounts of published and digitized legal texts, computational linguistic analysis of legal language becomes possible and eficient. However, most of these open language resources and shared tasks are in English. For the languages that have little open legal texts like Japanese, a word embedding model trained on the specific language usages is accompanied by the concern of less accuracy and representativeness. Based on the observation that legal language shares a modest common vocabulary with general language, we examined the validity of using the pretrained general word embedding model for processing legal texts by an intrinsic evaluation constructed on pairs of synonyms and related terms which were extracted from a legal term dictionary. We first investigated the settings of hyperparameters of the embedding models trained on legal texts. Then we compared the performances of our domain-specific models with general models. The pre-trained Wikipedia model conducted a better performance than domain-specific models on detecting semantic relations. This model also showed a higher compatibility with legal texts than the general model trained on newspaper articles. Although researchers tend to indicate the importance of domain-specific representation models, a general model can still be an alternative solution when there is little language resource.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        Owing to the emergence of Word2Vec [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ] and the following
explosive improvements in distributed representation learning, use
of distributed representation models as features becomes a
paradigm in automated semantic analysis. In general, for the
construction of such models, trainings on large-scale balanced corpora are
ideal and necessary, and for evaluation, shared downstream tasks
and robust evaluation measures are required.
      </p>
      <p>
        Resources of general language usages are abundant in major
languages. However, when processing texts in specialised domains,
vocabularies of these domains can be very diferent from general
language. Besides those so-called “technical terms” appeared in
every specialised domain, there are also words called “sub-technical
terms” that “activate a specialised meaning in the legal field,
being frequently used as general words in everyday language” [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
The high ratio of sub-technical terms in legal English vocabulary
diferentiates it “from the lexicon of other LSP (Language for
Specific Purposes) varieties” [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. The existence of technical terms and
sub-technical terms indicates that words in the specialised domain
are not only diferent from general language in the aspect of
vocabulary, but also in the aspect of semantics. That can make the
application of general word embedding models to specialised
domains ineficient and unreasonable due to the inconsistency of the
semantic spaces. Thus, to let the compositions of the processing
texts stay along with the embedding models in the same semantic
space, the domain-specific models are preferable.
      </p>
      <p>
        Unfortunately, in comparison with English, there are less open
source data of legal texts in Japanese. Either the documents are
not in machine-readable data format, or they are not even open to
public. As estimated in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] that “both using more data and higher
dimensional word vectors will improve the accuracy”, less data will
cause lower accuracy conversely. Nevertheless, that legal language
has a high ratio of overlapping of the vocabulary with general
language provides a possibility for us to apply general embedding
models.
      </p>
      <p>Therefore, in this paper, we examine whether general
embedding models, specifically, a Japanese word embedding model
pretrained on Japanese Wikipedia and a model trained on newspaper
articles, can be used when processing legal texts. We start with
constructing a similarity and relatedness task as an intrinsic
evaluation of trained embedding models. Pairs of synonyms and
related terms are extracted from a Japanese legal term dictionary. We
train domain-specific embedding models on two legal text datasets
with diferent settings of hyperparameters and investigate the best
configurations. The comparison of the general embedding models
and the domain-specific models are then conducted by the intrinsic
evaluation.</p>
      <p>Although the performance of a model mostly depends on
downstream tasks, we believe that it is also important for researchers
to have an awareness of the distributed representations inside of
the embedding models when trying to use them to achieve better
scores in specific tasks and to solve the real world problems.
2</p>
    </sec>
    <sec id="sec-2">
      <title>RELATED WORK</title>
      <p>NLP tasks related to legal issues, including legal information
retrieval, document classification, question answering methods and
so on, have been increasingly attracting attention from both
computational linguists and legal professionals.</p>
      <p>
        To improve the performances of theses tasks with the
assistance of semantic analysis, there were two word embedding
models specifically trained on legal texts. One was the pre-trained model
built in a Python library called LexNLP [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. LexNLP focused on
natural language processing and machine learning for legal and
regulatory text. The pre-trained models were based on thousands
of real documents and various judicial and regulatory proceedings.
The other one was Law2Vec 1 provided by LIST2. This model
“oriented to legal text trained on large corpora comprised of legislation
from UK, EU, Canada, Australia, USA, and Japan among other
legal documents.” Although legal texts of Japan seemed to be used
for achieving semantic representations of words in the legal
domain, the used texts were English-translated and the models were
for legal English.
      </p>
      <p>
        COLIEE (the legal question answering Competition on Legal
Information Extraction/Entailment) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] is the only competition about
Japanese legal texts and providing law articles both in Japanese and
English as a knowledge resource. COLIEE 2017 focused on
extraction and entailment identification aspects of legal information
processing related to answering yes/no questions from Japanese legal
bar exams. Carvalho et el. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and Nanda et el. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] both tested the
Google News dataset pre-trained vectors3 in information retrieval,
and the former team also found that the “pure common text
embedding” resulted in poor performance, “most probably due to the
absence of legal vocabulary and corresponding semantics.”
      </p>
      <p>
        The evaluation of the word embeddings trained from diferent
textual resources has been conducting in the biomedical domain.
Roberts [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] revealed that combinations of corpora led to a better
performance. Wang et al. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] concluded that the word embeddings
trained on the biomedical domain did not necessarily have better
performance than those trained on the general domain. While they
both agreed that the eficiency of a word embedding model was
task-dependent, Gu et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] argued that even smaller
domainspecific corpora may be preferable to pre-trained word
embeddings built on a general corpus if the diversity of vocabulary was
low.
      </p>
      <p>In general, related work tends to indicate the importance of
domain-specific distributed representation models for processing
specialised texts.
3</p>
    </sec>
    <sec id="sec-3">
      <title>DATA</title>
      <p>Our dataset consists of three corpus, a dictionary of legal terms
(hereinafter, referred to as dictionary), “the fact of the crime” parts
of the judgements obtained from Westlaw Japan 4 judicial
precedent corpus (referred to as judgements), and newspaper articles
contained in Mainichi Newspaper Corpus (referred to as
newspaper). Basic statistics of our corpus are given in Table 1. Detailed
descriptions of each corpus are given below.</p>
      <p>Dictionary. The technical term dictionary adopted in this work
was Yuhikaku Legal Term Dictionary (4th edition). The dictionary
consists of 13,812 entry words with the definitions written by
experts and carefully edited. We simply referred the word “legal term”
(or “term”) to the entry words that were recorded in the dictionary
instead of getting involved in the sophisticated discussion about
the meaning of the word. In the dictionary, a synonym of term t is
1https://archive.org/details/Law2Vec.
2http://www.luxli.lu/university-of-athens/.
3https://code.google.com/archive/p/word2vec.
4https://www.westlawjapan.com/.
given when t has no definition and is labeled with a See tag, while
a related term of t is labeled with a See Also tag when the experts
thought more information was needed</p>
      <p>Judgements. We obtained 2,306 judgements passed on criminal
cases in district courts nationwide from 2008 to 2017. Legal English
is known as legalese because of its tedious and puzzling language
usage. Legal Japanese also shared these problems. Therefore, in
order to conduct a moderate comparison with newspaper articles
in the aspect of contents and document lengths, we extracted “the
fact of the crime” part from each judgement.</p>
      <p>Newspaper. When a case happened, it is often reported as an
article in the social section of the newpaper. Additionally, the
language usage in a newspaper article can be considered as a general
usage, or at least less specialized than legalese used in legal texts.
We obtained all the articles from a one-year (2015) corpus. 1748
legal terms were observed in these articles.
4</p>
    </sec>
    <sec id="sec-4">
      <title>METHODS</title>
      <p>Before processing, we applied Japanese morphological analyzer
Chasen5 to split the sentences into words and remove signals and
numbers. The similarity measured between two vectors in this
paper were all cosine similarity.</p>
      <p>The examining procedure was in two steps. First, we built a term
pair inventory for performance evaluation. Term pairs were
separated into synonym pairs and related pairs. Domain-specific
models were then trained with hyper parameter tuning on this
inventory. Second, we focused on the common term pairs existed in both
general models and domain-specific models. The performances of
the models were examined on both synonym detection and related
term detection.
4.1</p>
    </sec>
    <sec id="sec-5">
      <title>Task Design</title>
      <p>We extracted 1440 pairs of synonyms and 6641 pairs of related
terms by exploiting the indicative tags provided in the dictionary.
These pairs constructed the gold standards of synonym detection
task and related term detection task for evaluating each model’s
ability of catching semantic relations between terms.</p>
      <p>We evaluated the performances of models by counting how many
semantic relations were correctly caught by each model.
Specifically, we first obtained top n most similar words of term t from the
model. n was set to {1, 5, 10}. If the synonym or the related term
was in these most similar words, we treated the trial as a correct
one. The performance was represented by accuracy as the ratio of
correctly predicted pairs to all synonym or related term pairs.
5version: 0.996, neologd 102.</p>
      <p>The Validity of General Embedding Models for Processing Japanese Legal Texts</p>
    </sec>
    <sec id="sec-6">
      <title>Model Training</title>
      <p>We applied pre-trained Wikipedia Entity Vectors as our general word
embedding model 6. It is a 300-dimension Skip-Gram Negative
Sampling (SGNS) model. With the same training configuration of it,
we trained another general model on newspaper articles for the
comparison within general models. We then trained our
domainspecific models on the dictionary and the judgements, respectively
and together. The size of the source data and vocabularies are given
in Table 2.</p>
      <p>
        The performance of word embedding models can be improved
by hyperparameter tuning. Since the efects of diferent
configurations can be diverse, we investigated hyperparameter settings as
in [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. We exploited gensim [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] for model training. Examined
parameters and values are shown in Table 3. Each model had five
chances on each task.
5
5.1
      </p>
    </sec>
    <sec id="sec-7">
      <title>RESULTS</title>
    </sec>
    <sec id="sec-8">
      <title>Model Tuning</title>
      <p>The best accuracy scores of models on the two tasks under
diferent configurations are shown in Table 4, 5. The models trained on
judgements failed in detecting both synonym and relatedness
relations. The best accuracy of those judgement models was 0 (0.0%)
6https://github.com/singletongue/WikiEntVec. Wikipedia data until 2018.10.01.
for synonym pairs, and 19 (0.3%) for related term pairs. In both
tasks, additional legal texts (i.e., judgements) did not improve the
performance of our domain-specific models, which indicated that
our legal text dataset is biased to the dictionary dataset and the
more data does not always lead to the better performance.</p>
      <p>
        The default training configuration of gensim is {dimension =
100, window size = 5, min.count = 5, negative sample = 5}. The
selected configuration after a hyperparameter tuning on an English
domain-specific model training [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] was {dimension = 400,
window size = 5, min.count = 5, negative sample = 5}. However, we
found that window size or negative sample that was lower than 10
would led to worse performances in all circumstances. Due to the
relatively tiny data size, min.count that larger than 3 also had a
negative efect on the performances.
      </p>
      <p>The most suitable configuration for the models trained on the
dictionary across the variation of top_n was {dimension = 300,
window size = 15, min.count = 3, negative sample = 10}. It is similar
to the configuration of the Wikipedia model which is {dimension
= 300, window size = 10, min.count = 3, negative sample = 10}.</p>
      <p>We selected the same values of parameters as the Wikipedia
model as the training configuration of our domain-specific model
with which the general models would be compared on the next
stage.</p>
    </sec>
    <sec id="sec-9">
      <title>Intrinsic Evaluation</title>
      <p>As shown in Table 4, 5, the Wikipedia model achieved higher
performances on detecting semantic relations of legal terms, even those
relations were obtained from the legal domain. This result can be
due to the absence of low frequency terms in the dictionary corpus.
Therefore, we further conducted two detection tasks on the
common pairs among the domain-specific model, the Wikipedia model
and the newspaper model. There were 18 common synonym pairs
and 465 common related term pairs. Results of the experiment are
shown in Table 6, 7.</p>
      <p>The Wikipedia model achieved the best accuracy score among
three models, while the same general embedding model, the
newspaper model, was the worst. The performance diference between
the Wikipedia model and the newspaper model also confirmed that
the performance of general models are efected by the diversity of
general language resources. The similar results of the examination
on the common term pairs to the examination on all term pairs
indicated that the Wikipedia model is superior to the domain-specific
dictionary model for catching the intrinsic semantic relations of
legal terms.
6</p>
    </sec>
    <sec id="sec-10">
      <title>CONCLUSION</title>
      <p>Since the usefulness of an embedding model mostly depends on the
downstream tasks, we don’t argue that which embedding model is
better or worse for legal NLP tasks. The purpose of this research
is to investigate whether a general corpus could be used when the
training on the specific domain is not practicable. The word
embedding model built on Wikipedia showed a considerable
performance on the intrinsic evaluation. The legal domain is diferent
from other specialised domains in the aspect of the ratio of
overlapping words with general language. This characteristic is helpful
when there are not enough domain-specific language resources. In
this paper, we provided some evidence that domain-specific word
embedding models are not always outperform general models and
not all the domain-specific texts are useful when constructing the
semantic relations among technical terms. The using of general
word embedding models, especially the models trained on a
balanced large-scale corpus, therefore can be considered as an
alternative way to processing those domain-specific texts.</p>
    </sec>
    <sec id="sec-11">
      <title>ACKNOWLEDGMENTS</title>
      <p>The authors would like to thank YUHIKAKU Publishing Co., Ltd.
for providing the legal dictionary dataset. We are also grateful to
the reviewers for their valuable comments and suggestions.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Michael</given-names>
            <surname>James</surname>
          </string-name>
          <string-name>
            <surname>Bommarito</surname>
          </string-name>
          , Daniel Martin Katz, and
          <string-name>
            <given-names>Eric</given-names>
            <surname>Detterman</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <source>LexNLP: Natural Language Processing and Information Extraction For Legal and Regulatory Texts. SSRN Electronic Journal</source>
          (
          <year>2018</year>
          ). https://doi.org/10.2139/ssrn.3192101
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Danilo</surname>
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Carvalho</surname>
          </string-name>
          , Vu Tran, Khanh Van Tran, and Nguyen Le Minh.
          <year>2017</year>
          .
          <article-title>Improving Legal Information Retrieval by Distributional Composition with Term Order Probabilities</article-title>
          .
          <source>In COLIEE 2017. 4th Competition on Legal Information Extraction and Entailment (EPiC Series in Computing)</source>
          , Ken Satoh,
          <string-name>
            <surname>Mi-Young</surname>
            <given-names>Kim</given-names>
          </string-name>
          , Yoshinobu Kano,
          <source>Randy Goebel, and Tiago Oliveira (Eds.)</source>
          , Vol.
          <volume>47</volume>
          . EasyChair,
          <fpage>43</fpage>
          -
          <lpage>56</lpage>
          . https://doi.org/10.29007/2xzw
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Yang</surname>
            <given-names>Gu</given-names>
          </string-name>
          , Gondy Leroy, Sydney Pettygrove, Maureen Kelly Galindo, and
          <string-name>
            <surname>Margaret</surname>
          </string-name>
          Kurzius-Spencer.
          <year>2018</year>
          .
          <article-title>Optimizing Corpus Creation for Training Word Embedding in Low Resource Domains: A Case Study in Autism Spectrum Disorder (ASD)</article-title>
          .
          <source>AMIA Annual Symposium proceedings (2018)</source>
          ,
          <fpage>508</fpage>
          -
          <lpage>517</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Yoshinobu</given-names>
            <surname>Kano</surname>
          </string-name>
          ,
          <string-name>
            <surname>Mi-Young</surname>
            <given-names>Kim</given-names>
          </string-name>
          , Randy Goebel, and
          <string-name>
            <given-names>Ken</given-names>
            <surname>Satoh</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Overview of COLIEE 2017</article-title>
          .
          <source>In COLIEE 2017. 4th Competition on Legal Information Extraction and Entailment (EPiC Series in Computing)</source>
          , Ken Satoh,
          <string-name>
            <surname>Mi-Young</surname>
            <given-names>Kim</given-names>
          </string-name>
          , Yoshinobu Kano,
          <source>Randy Goebel, and Tiago Oliveira (Eds.)</source>
          , Vol.
          <volume>47</volume>
          . EasyChair, 1-
          <fpage>8</fpage>
          . https://doi.org/10.29007/fm8f
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>María</given-names>
            <surname>José</surname>
          </string-name>
          Marín and
          <string-name>
            <given-names>Camino</given-names>
            <surname>Rea</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Researching Legal Terminology: A Corpus-based Proposal for the Analysis of Sub-technical Legal Terms</article-title>
          .
          <source>ASp</source>
          <volume>66</volume>
          (nov
          <year>2014</year>
          ),
          <fpage>61</fpage>
          -
          <lpage>82</lpage>
          . https://doi.org/10.4000/asp.4572
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>María</given-names>
            <surname>José Marín Pérez</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Measuring the Degree of Specialisation of Sub-technical Legal Terms through Corpus Comparison: A Domain-independent Method</article-title>
          .
          <source>Terminology</source>
          <volume>22</volume>
          ,
          <issue>1</issue>
          (
          <year>2016</year>
          ),
          <fpage>80</fpage>
          -
          <lpage>102</lpage>
          . https://doi.org/10.1075/term.22.1.04mar
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          , Kai Chen, Greg Corrado, and
          <string-name>
            <given-names>Jefrey</given-names>
            <surname>Dean</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Efifcient Estimation of Word Representations in Vector Space</article-title>
          . (
          <year>2013</year>
          ).
          <source>arXiv:cs.CL/1301.3781v3</source>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          , Ilya Sutskever, Kai Chen, Greg S Corrado, and
          <string-name>
            <given-names>Jef</given-names>
            <surname>Dean</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Distributed Representations of Words and Phrases and their Compositionality</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          .
          <volume>3111</volume>
          -
          <fpage>3119</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Rohan</given-names>
            <surname>Nanda</surname>
          </string-name>
          , Adebayo Kolawole John, Luigi Di Caro, Guido Boella, and
          <string-name>
            <given-names>Livio</given-names>
            <surname>Robaldo</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Legal Information Retrieval Using Topic Clustering and Neural Networks</article-title>
          .
          <source>In COLIEE 2017. 4th Competition on Legal Information Extraction and Entailment (EPiC Series in Computing)</source>
          , Ken Satoh,
          <string-name>
            <surname>Mi-Young</surname>
            <given-names>Kim</given-names>
          </string-name>
          , Yoshinobu Kano,
          <source>Randy Goebel, and Tiago Oliveira (Eds.)</source>
          , Vol.
          <volume>47</volume>
          . EasyChair,
          <fpage>68</fpage>
          -
          <lpage>78</lpage>
          . https://doi.org/10.29007/psgx
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Farhad</surname>
            <given-names>Nooralahzadeh</given-names>
          </string-name>
          , Lilja Øvrelid, and Jan Tore Lønning.
          <year>2018</year>
          .
          <article-title>Evaluation of Domain-specific Word Embeddings using Knowledge Resources</article-title>
          .
          <source>In Proceedings of the 11th Language Resources and Evaluation Conference. European Language Resource Association</source>
          , Miyazaki, Japan. https://www.aclweb.org/anthology/L18-1228
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Radim</given-names>
            <surname>Řehůřek</surname>
          </string-name>
          and
          <string-name>
            <given-names>Petr</given-names>
            <surname>Sojka</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Software Framework for Topic Modelling with Large Corpora</article-title>
          .
          <source>In Proceedings of the LREC 2010 Workshop on New Challenges for NLP Frameworks. ELRA</source>
          , Valletta, Malta,
          <fpage>45</fpage>
          -
          <lpage>50</lpage>
          . http://is.muni.cz/publication/884893/en.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Kirk</given-names>
            <surname>Roberts</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Assessing the Corpus Size vs. Similarity Trade-of for Word Embeddings in Clinical NLP</article-title>
          .
          <source>In Proceedings of the Clinical Natural Language Processing Workshop</source>
          (ClinicalNLP).
          <source>The COLING 2016 Organizing Committee</source>
          , Osaka, Japan,
          <fpage>54</fpage>
          -
          <lpage>63</lpage>
          . https://www.aclweb.org/anthology/W16-4208
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Yanshan</surname>
            <given-names>Wang</given-names>
          </string-name>
          , Sijia Liu, Naveed Afzal, Majid Rastegar-Mojarad,
          <string-name>
            <given-names>Liwei</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Feichen</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Paul</given-names>
            <surname>Kingsbury</surname>
          </string-name>
          , and Hongfang Liu.
          <year>2018</year>
          .
          <article-title>A Comparison of Word Embeddings for the Biomedical Natural Language Processing</article-title>
          .
          <source>Journal of Biomedical Informatics</source>
          <volume>87</volume>
          (nov
          <year>2018</year>
          ),
          <fpage>12</fpage>
          -
          <lpage>20</lpage>
          . https://doi.org/10.1016/j.jbi.
          <year>2018</year>
          .
          <volume>09</volume>
          .008
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>