<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Deep Ensemble Learning for Legal Query Understanding</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Arunprasath Shankar</string-name>
          <email>arunprasath.shankar@lexisnexis.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Venkata Nagaraju Buddarapu</string-name>
          <email>venkatanagaraju.buddarapu@lexisnexis.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>LexisNexis</institution>
          ,
          <addr-line>Raleigh</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Legal query understanding is a complex problem that involves two natural language processing (NLP) tasks that needs to be solved together: (i) identifying intent of the user and (ii) recognizing entities within the queries. The problem equates to decomposing a legal query into its individual components and deciphering the underlying diferences that can occur due to pragmatics. Identifying the desired intent and recognizing correct entities helps us return back relevant results to the user. Deep Neural Networks (DNNs) have recently achieved great success surpassing traditional statistical approaches. In this work, we experiment with several DNN architectures towards legal query intent classification and entity recognition. Deep Neural architectures like Recurrent Neural Networks (RNNs), Long Short Term Memory (LSTM), Convolutional Neural Networks (CNNs) and Gated Recurrent Units (GRU) were applied and compared against one another both individually and as combinations. The models were also compared against machine learning (ML) and rule-based approaches. In this paper, we describe a methodology that integrates posterior probabilities produced by the best DNN models and create a stacked framework for combining the diferent predictors to improve prediction accuracy and F-measure for legal intent classification and entity recognition.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Copyright © CIKM 2018 for the individual papers by the papers'
authors. Copyright © CIKM 2018 for the volume as a collection
by its editors. This volume and its papers are published under
the Creative Commons License Attribution 4.0 International (CC
BY 4.0).
1.</p>
    </sec>
    <sec id="sec-2">
      <title>Introduction</title>
      <p>
        With US legal services market approaching $500 bn
in 2018 that is 2% of US GDP [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and substantial
incentives to increase market share, maximizing user
satisfaction with search results continues to be the
primary focus for legal search industry. Recently, there is
an upward trend of embracing vertical search engines
by legal companies since they provide a specialized
type of information service to satisfy a user’s intent.
Using a specialized user interface, a vertical search
engine can return more relevant results than a general
search engine for in-domain legal queries.
      </p>
      <p>In practice, a user has to decide his choice of
vertical search engine beforehand to satisfy his query
intent. It would be convenient if a query intent identifier
could be provided in a general search engine that could
precisely predict whether a query should trigger a
vertical search in a certain domain. Moreover, since a
user’s query may implicitly express more than one
intent, it would be very helpful if a general search engine
could detect all query intents, distribute it to
appropriate vertical search engines and efectively organize
the results from the diferent vertical search engines to
satisfy a user’s need. Consequently, understanding a
query intent is crucial for providing better search
results and thus improving the overall satisfaction of the
user.</p>
      <p>Understanding intent and identifying legal entities
from a user’s query can help legal search
automatically route the query to corresponding vertical search
engines and obtain relevant contents, thus, greatly
improving user satisfaction meaning better assistance to
law researches to support legal arguments and
decisions. Law researchers have to cope with a tremendous
load of legal content since sources of law can originate
from diversified sources like judicial branch,
legislative branch (statutes), legal reference books, journals,
news etc. This makes legal search and understanding
queries in legal search such an important aspect to
retrieve relevant support documents for supporting one’s
argument.</p>
      <p>Legal search is hard as it demands writing complex
queries to retrieve desired content from information
retrieval (IR) systems. Classifying legal queries and
identifying domain specific legal entities from queries
is even harder. For e.g., in the query: “who is supreme
court magistrate John Roberts and abortion law?”, the
word “magistrate” can be resolved to a &lt;judge title&gt;
when observed along with the context phrase “supreme
court”. Similarly, the phrase “abortion law” can be
identified as a &lt;practice area&gt; when seen alongside
a supporting context. However, since we also observe
the interrogative phrase “who is”, we can safely assume
that the intent of this query is &lt;judge&gt; or &lt;person&gt;
search.</p>
      <p>Similar to google queries, legal queries can be
classified into one of three categories. (i) Do: The users
wants to do something, like buy a product or
service. E.g., Buy Law Books, Research Guides,
Police/Personal Reports, Real Estate Property Records
etc. (ii) Know: An informational query, where the
user wants to learn about a subject. E.g., law, statute,
doctrine etc. Very often single word queries are
classified at least partially as “Know” queries. (iii) Go:
Also known as a navigational query, the user wants to
go to a specific site scoped to a particular legal entity.
E.g., a query where a user wants to know analytics of
a specific legal entity like a judge or an expert
witness. Our research in this work focuses only on “Go”
type of queries, e.g., an user wanting to understand a
particular judge profile with the query “ judge John D.
Roberts”.</p>
      <p>The process of finding named entities in a text and
classifying them to a semantic type is called named
entity recognition (NER). Legal NER is nearly always
used in conjunction with intent classification systems.
Given a query, the goal of NER system described in
this paper is two-fold: (i) segment the input into
semantic chunks, and (ii) classify each chunk into a
predefined set of semantic classes. For e.g., given a
query “judge john d roberts”, the desired output would
be: “judge” = &lt;OTHER&gt; and “john d roberts” =
&lt;JUDGE&gt;. Here the class &lt;JUDGE&gt; represents a
person or a judge entity and the class &lt;OTHER&gt;
represents any non judge specific term.</p>
      <p>The aim of this research is to improve the
understanding of queries involved in legal research. In this
paper, we explore three diferent approaches: (a)
machine learning (ML) with feature engineering, (b) deep
learning (DL) without any feature engineering and (c)
ensemble of deep neural networks. We then perform
quantitative evaluations in comparison to an already
existing baseline model which is a rule-based
classification system in production. Finally, we show from
our experimental results that a deep ensemble model
significantly outperforms other approaches for both
intent classification and legal NER.
2.</p>
      <p>
        Background and Related Work
DL systems have dramatically improved the state of
several domains like NLP, computer vision, image
processing etc. Various deep architectures and learning
methods have been developed with distinct strengths
and weaknesses in recent years. Deep ensemble
learning is a learning paradigm where ensembles of several
neural networks show improved generalization
capabilities that outperform those of single networks. For
deep learning of multi-layer neural networks,
ensemble learning is still applicable [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ]. However, there is
not much work done towards legal domain involving
neither deep learning nor ensemble learning. How can
ensemble learning be applied to various DNN
architectures to achieve better results for legal tasks is the
primary focus of this paper.
      </p>
      <p>
        Most ML approaches to text understanding consists
of tokenizing a string of characters into structures such
as words, phrases, sentences or paragraphs, and then
apply some statistical classification algorithm onto the
statistics of such structures [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. These techniques work
well when applied to a narrowly defined domain.
      </p>
      <p>
        Typical queries submitted to legal search engines
contain very short keyword phrases, which are
generally insuficient to fully describe a user’s information
need. Thus, it is a challenging problem to classify
millions of queries into some predefined categories. A
variety of related topical query classification problems
have been investigated in the past [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Most of
them seek to use statistical machine learning methods
to train a classifier to predict the category of an input
query. From the statistical learning perspective, in
order to obtain a classifier that has good generalization
ability in predicting future unseen data, two
conditions should be satisfied: discriminative feature
representation and suficient training samples. However, for
the problem of query intent classification, even though
there are huge volumes of legal queries, both
conditions are hardly to met due to the sparseness of query
features coupled with the sparseness of labeled
training data. In [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], Beitzel et al. attempted to solve this
problem through augmenting the query with more
features using external knowledge, such as search engine
results and achieved fair results.
      </p>
      <p>DNNs have revolutionized the field of NLP. RNNs
and CNNs, the two main types of DNN architectures
are widely explored to handle various NLP tasks. CNN
is supposed to be good at extracting positional
invariant features and RNN at modeling units in sequence.
The state-of-the-art on many NLP tasks often switches
due to the battle of CNNs and RNNs. While CNNs
take advantage of local coherence in the input to cut
down on the number of weights, RNNs are used to
process sequential data (often with LSTM cells). RNNs
are also good at representing very large implicit
internal structures that are dificult even to think about.</p>
      <p>In summary, the conventional wisdom is that RNNs
should be used when the context is richer and there
is more state information that needs to be captured.
This proposition has been challenged by CNNs
recently with the claim that finite state information of
limited scope can be more eficiently handled by
multiple convolution layers. We think both are true, and
one should not go for RNNs just for the sake of it,
instead more eficient deep CNNs should be tried in
limited context situations. But for more complex
implicit mappings where context and state information
spans are much bigger, RNNs are the best and at this
point almost the only tool.</p>
      <p>There is not a lot of previous research work
involving DL and legal domain for the problems of intent
classification and NER. However, there are a handful
of papers that talk about DNN approaches for
nonlegal intent classification and general NER. Also
recently, RNNs and CNNs have been applied on a variety
of NLP tasks with various degree of success. Below,
we talk about the evolution of various DNNs and their
applications towards NLP tasks.</p>
      <p>
        In [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], Zhai et al. adopted RNNs as building blocks
to learn desired representations from massive user click
logs. The authors proposed a novel attention network
that learns to assign attention scores to words within a
sequence (query or ad). In [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], Hochreiter and
Schmidhuber introduced LSTM which solved the most
complex, artificial long-time-lag tasks that have never been
solved by previous recurrent network algorithms. In
[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], Chung et al. compared diferent types of
recurrent units in RNNs especially focusing on units that
implement a gating mechanism, such as LSTM and
GRU units. They evaluated these units on the tasks
of polyphonic music modeling and speech signal
modeling and proved LSTMs and GRUs work better than
traditional recurrent units.
      </p>
      <p>
        In [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], Hinton et al. introduced CNN and used
it to for image classification on the ImageNet dataset
and established new state-of-the-art results. In [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ],
LeCun et al. demonstrated that we can apply DL
to text understanding from character-level inputs all
the way up to abstract text concepts, using
temporal CNNs. They applied CNNs to various large-scale
datasets, including ontology classification, sentiment
analysis, text categorization etc. and showed that
temporal CNNs can achieve astonishing performance
without any prior knowledge of syntactic or semantic
structures with regards to a human language. With respect
to intent classification, Hu et al. devised a
methodology to identify query intent by mapping the query to
a representation space backed by Wikipedia [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. In
[
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], the authors applied CNNs for intent classification
and achieved results in par with state-of-the-results.
      </p>
      <p>
        There are also many recent works that combine
RNNs with CNNs for diferent NLP tasks. For
example, in [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], Chiu et al. presented a novel neural
network architecture that automatically detects word
and character-level features using a hybrid
bidirectional LSTM and CNN architecture, eliminating the
need for most of feature engineering and established
a new state-of-the-art performance with an F1 score
of 91.62% on CoNLL-2003 data set. In [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ],
Limsopatham et al. proposed a bidirectional LSTM
(BiLSTM) to automatically learn orthographic features
from tweets.
      </p>
      <p>
        There is a handful of papers that delve into legal
applications using DNNs. In [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], Sugathadasa et al.
proposed a system that includes a page ranking graph
network with TF-IDF to build document embeddings
by creating a vector space for the legal domain, which
can be trained using a doc2vec neural network model
supporting incremental and extensive training for
scalability of the system. In [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], Nanda et al. proposed
a hybrid model using LSTM and CNN which utilizes
word embeddings trained on the Google News vectors
and evaluated the results on COLIEE 2017 dataset.
They demonstrated that the performance of LSTM +
CNN model was competitive with other textual
entailment systems. Similarly, in [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], the authors proposed
a methodology to employ DNNs and word2vec for
retrieval of civil articles.
      </p>
      <p>
        For NER, Huang et al. proposed a variety of LSTM
and LSTM variants like LSTM + CRF for NER and
achieved state-of-the-art results [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. Recently this
year, in [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], Peters et al. introduced a new type of
deep contextualized word representation that models
both semantics and polysemy and improved the state
of the art across six challenging NLP problems,
including question answering, textual entailment, sentiment
analysis and NER.
      </p>
      <p>In this paper, we scope our research limited to judge
queries meaning queries that are meant to be a search
for judge profile. Being this the scope of the problem,
the intent classification task needs to classify a given
query as a judge query or not. Once the query is
identified as a judge query, the NER system that follows,
recognizes if the query contains any person entities.
The entities are then routed to vertical searches that
follow. The structure of the paper is as follows.
Section 3. discusses the various experiments that was
carried out to solve legal query understanding. That is
followed by Section 4. that presents and analyses
overall results. The paper ends with Section 5. which gives
the conclusion and discussion on future work.</p>
      <p>MODELS AND EXPERIMENTS
The adoption of artificial Intelligence (AI) technology
is undoubtedly transforming the practice of law. Many
in the legal profession are aware that using AI can
greatly reduce time and costs while increasing
prediction measures. In the following subsections, we talk
about data collection, augmentation and training of
the diferent DNN models for legal tasks we
experimented for: (i) intent classification and (ii) legal entity
recognition.
3.1</p>
      <sec id="sec-2-1">
        <title>Data Collection</title>
        <p>For the purpose of our experiments, we created data
sets comprising of diferent types of legal queries such
as judge search, case law search, statutes/elements
search, etc. For intent classification, since the
classification task is binary, we labeled all judge queries
to be positive and the rest as negative. For NER, we
used three labels: (i) &lt;OTHER&gt; used to tag any
tokens that are not part of a judge name (ii) &lt;B-PER&gt;
denotes the beginning of a person (judge) name and
(iii) &lt;I-PER&gt; denotes the inside of a name. Table 1
portrays the diferent query types by volume we used
for our experiments. All of the collected data labeled
by our subject matter experts (SMEs). All of the data
discussed here are proprietary to LexisNexis.
3.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Data Augmentation and Balancing</title>
        <p>The labeled data was good enough to get started but
was highly imbalanced with respect to NER labels. In
order to balance out the data set, we augmented the
data by (i) expanding queries by pattern using judge
names from our proprietary judge master database, (ii)
oversampling the under represented patterns and (iii)
under sampled the over exemplified patterns. The top
10 patterns by frequency are shown in table 3. These
10 patterns alone contributed to 78% of the labeled
data. The balancing act was carried out by following
two strategies: For oversampling, we pick a pattern
and create synthetic data points (queries) by randomly
substituting judge names/titles around the patterns to
be fitted. For the process of under sampling, random
data points are selected from buckets of pattern and
removed. The sampling process is stopped when there
are equal number of queries in all the buckets of
patterns. Once the data is augmented and balanced, we
split the data into train, dev and test sets. Table 2
shows the ratio of data split for both intent
classification and NER.
3.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Word Embedding</title>
        <p>Learning a high-dimensional dense representation for
vocabulary terms, also known as a word embedding,
Type Example
Judge judge John D. Roberts
Word Wheel Ohio Municipal Court,</p>
        <p>Bellefontaine
Query Log sexual harassment
Statute statute /s limitations /s
ac</p>
        <p>tual /s fraud
Elements Law contract defense
uncon</p>
        <p>
          scionability elements
Case Search Powers v. USAA
has recently attracted much attention in NLP and IR
tasks. The embedding vectors are typically learned
based on term proximity in a large corpus and is used
to accurately predict adjacent word(s) for a given word
or context [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ]. For the purpose of training our NER
DNN models, we created word vectors of size 580,614
to be used as word embeddings to the embedding layer
of our DNN models. The word vectors were created by
training a word2vec continous bag of words (CBOW)
model on AWS ml.p3.2xlarge instance with a single
NVIDIA Tesla V100 GPU. As an input to the model,
10 Million queries were obtained from user session logs
were used. These queries were ordered by frequency.
The word2vec model was trained in 112.2 minutes for
100 epochs.
        </p>
        <p>Pattern
B-PER I-PER I-PER</p>
        <p>B-PER I-PER</p>
        <p>O B-PER I-PER I-PER
B-PER I-PER I-PER I-PER</p>
        <p>O B-PER I-PER
O B-PER I-PER I-PER I-PER I-PER</p>
        <p>B-PER I-PER I-PER I-PER
B-PER I-PER I-PER I-PER I-PER</p>
        <p>O B-PER
O O O B-PER I-PER I-PER I-PER O
FP
57
137
1789
137
25
139
107</p>
        <p>FN
209
141
20
256
647
142
160</p>
        <p>F</p>
      </sec>
      <sec id="sec-2-4">
        <title>Intent Classification</title>
      </sec>
      <sec id="sec-2-5">
        <title>Baseline vs ML Approaches:</title>
        <p>For intent classification, we tried a few diferent ML
classifiers firsthand before delving into the deep
learning arena. The ML models were compared against
a rule-based query recognition system highlighted
that was already established as our baseline system
to improve. Classifiers were selected from both the
linear and non-linear category. The list of classifiers
tried and their corresponding results (F1 score) against
baseline are shown in table 4 above. Our main
requirement while picking and implementing the classifiers
was to minimize the overall number of false positives.
This is because false positives with respect to intent
classification can intercept queries that are meant to
be routed to a diferent vertical search engine
afecting the overall customer satisfaction towards the
product. Amongst the linear classifiers, linear SVM seemed
to perform slightly better than the logistic regression
classifier. Amongst the non-linear category, the
multilayer perceptron highlighted seemed to perform the
best beating decision trees, adaboost and naive bayes
approaches.
3.4.2</p>
      </sec>
      <sec id="sec-2-6">
        <title>Feature Engineering for ML Classifiers:</title>
        <p>There are diferent ways one can address a judge in
legal taxonomy such as “chief justice”, “associate
justice”, “magistrate” etc. All of the diferent phrases
used for addressing a judge are compiled into a bag
of words representation. In addition to bag of words,
all the ML classifiers use POS, gazetteer, word shape
and orthographic features to represent semantic and
linguistic meaning of a query.
3.4.3</p>
      </sec>
      <sec id="sec-2-7">
        <title>LSTM based Intent Classifier:</title>
        <p>
          The LSTM, first described in [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], attempts to
circumvent the vanishing gradient problem by separating the
memory and output representation, and having each
dimension of the current memory unit depending
linearly on the memory unit of the previous time step.
The DNN architecture we tried using LSTM is shown
in figure 1. The embedding layer that feeds the LSTM
layer is composed of a vocabulary of 49,701 words with
an output dimension of 200. 100 hidden units are used
for the LSTM layer. The flatten layer uses 200 hidden
units. The flatten layer takes a tensor of any shape and
transform it into a one dimensional tensor. Both uni
and bi-directional LSTMs were trained for the task.
3.4.4
        </p>
      </sec>
      <sec id="sec-2-8">
        <title>CNN based Intent Classifier:</title>
        <p>CNNs are built out of many layers of pattern
recognizers stacked on top of each other. Convolutional is
a way of saying that the machine looks at small parts
of a query first rather than trying to account for the
whole thing. Each successive layer combines
information from these small parts to fill in the bigger picture
and assemble complex patterns of meaning. After
trying few variants of LSTM architectures, we started
experimenting with CNNs for intent classification. The
general neural architecture we used for CNN is show
in figure 2. The major diference here compared to the
LSTM models are: CNNs use two dense layers after
the flatten layer whereas LSTMs use only one layer.
The 1D convolutional layer uses 32 filters with a
kernel size = 8 and the max pooling layer uses 2 strides
with a pool size = 2.
3.4.5
As a next step, the best performing models from both
the LSTM and CNN pools were picked and combined
into a hybrid model for classification. We built two</p>
        <p>Conv1D [32, 8, relu]
MaxPooling1D [2, 2, valid]
models for this experiment, LSTM + CNN and
BiLSTM + CNN. The architectural diagram for
BiLSTM + CNN model is shown in figure 3.
3.4.6</p>
      </sec>
      <sec id="sec-2-9">
        <title>Deep Ensemble for Intent Classification:</title>
        <p>Ensemble learning is a ML paradigm where multiple
learners are trained to solve the same problem. In
contrast to ordinary ML approaches which try to learn one
hypothesis from training data, ensemble methods try
to construct a set of hypotheses and combine them to
use. For intent classification, we use stacking which
applies several models to original data. In stacking
we don’t have just an empirical formula for our weight
function, rather use a logistic regression model to
estimate the input together with outputs of every model
to estimate the weights or, in other words, to
determine what models perform well and what badly given
these input data.</p>
        <p>To simplify the stacking procedure for ensemble
LSTM
Bi-LSTM</p>
        <p>CNN</p>
        <p>LSTM + CNN
Bi-LSTM + CNN
learning, we perform a linear combination of the
original class posterior probabilities produced by the best
DNN models at the word level (see table 8). A set of
parameters in the form of full matrices are associated
with the linear combination, which are learned using
the training data consisting of the word-level
posterior probabilities of the diferent models and its
corresponding word-level target values (0 or 1). Figure 4
depicts the model architecture of the ensemble
classiifer for the task of intent classification.
3.4.7</p>
      </sec>
      <sec id="sec-2-10">
        <title>Training:</title>
        <p>Table 5 below shows the overall training statistics
for the diferent DNN models deployed for intent
classification. The models were trained on AWS
ml.p3.8xlarge instance with 4 NVIDIA Tesla V100
GPUs. Average time shown below is measured in
minutes.
3.5
3.5.1</p>
      </sec>
      <sec id="sec-2-11">
        <title>Baseline:</title>
      </sec>
      <sec id="sec-2-12">
        <title>Named Entity Recognition</title>
        <p>Baseline model for NER is shown in table 12 . The
baseline model (highlighted in ) was previously
established and it is the same rule-based system that was
set as baseline for intent classification. The row
highlighted in shows metrics from a conditional random
Architecture
Dropout</p>
        <p>Hidden Units</p>
        <p>Embedding
LSTM I
LSTM II
LSTM III</p>
        <p>LSTM IV
Bi-LSTM I
Bi-LSTM II</p>
        <p>False
False
False
True
False
True
50
50
50
50
100
100
200
100
200
200
200
200</p>
        <p>Architecture
Filters</p>
        <p>Kernel Padding
CNN I
CNN II
CNN III
CNN IV
CNN V
32
32
64
64
64
16
8
8
16
4
same
same
same
same
same
LSTM III
Bi-LSTM I</p>
        <p>CNN II
LSTM III + CNN II
Bi-LSTM I + CNN II</p>
        <p>Ensemble (Top 5)
same
same
same
same
100
100
100
100
100
100
100</p>
        <p>FP
25
52
23
54
19
38
FN
13
13
11
15
9</p>
        <p>F</p>
        <p>FP
14
26
17
24
11
26
FN
18
18
18
15
14</p>
        <p>FN
42
27
25
21
22
23
F-score
ifeld (CRF) probabilistic model that was previously
implemented. This model was eventually discarded
since it was outperformed by most of DNN models.
3.5.2</p>
      </sec>
      <sec id="sec-2-13">
        <title>RNN based Named Entity Recognition:</title>
        <p>For NER, we started with RNNs. The neural
architecture for RNN is shown in figure 5 above. For
the embedding layer, we used pre-trained word2vec
embeddings of dimensions (577,149 x 200). For the
RNN layer, 200 hidden units were used. The time
distributed output layer has 3 units and uses a softmax
function since we have 3 classes in total (&lt;OTHER&gt;,
&lt;B-PER&gt; &amp; &lt;I-PER&gt;. A dropout value of 0.2 was
also used for the configuration.</p>
      </sec>
      <sec id="sec-2-14">
        <title>3.5.3 LSTM based Named Entity Recognition:</title>
        <p>LSTMs in general circumvent the vanishing gradient
problem faced by RNNs. For NER, we used both uni
and bi-directional LSTMs. The architecture is similar
to shown in figure 5 except the RNN layer is replaced
3.5.4</p>
      </sec>
      <sec id="sec-2-15">
        <title>CNN based Named Entity Recognition:</title>
        <p>For NER, we also experimented with CNN. The
neural architecture for CNN is shown in figure 6. The
Conv1D layer consists of 128 filters with kernel size
set to 5. In contrary to the CNN used for intent
classification, this architecture does not use a max pooling
layer.
3.5.5</p>
      </sec>
      <sec id="sec-2-16">
        <title>GRU based Named Entity Recognition:</title>
        <p>
          More recently, gated recurrent units have been
proposed [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] as a simplification of the LSTM, while
keeping the ability to retain information over long
sequences. Unlike LSTM, GRU uses only two gates,
memory units do not exist, and the linear interpolation
occurs in the hidden state. As part of our experiment,
we replaced the RNN layer in the architecture shown
in figure 5 with a GRU layer.
3.5.6
For NER, we combined the above discussed models.
Some of the hybrid models we built are RNN + CNN,
LSTM + CNN and GRU + CNN (both uni and
bidirectional). Architecture of Bi-LSTM + CNN is
shown in figure 7.
        </p>
        <p>Conv1D [128, 5, relu]
To create an ensemble for NER, the DNN models are
ranked by their F1 score. The top 5 best models are
then picked and stacked into an ensemble. Ensemble
of top 10 models was also experimented and discarded
since it underperformed compared to the ensemble of
top 5. Figure 8 shows the architecture of the chosen
ensemble model.
We utilize standard measures to evaluate the
performance of our classifiers, i.e., precision, recall and
F1measure. Precision (P) is the proportion of actual
positive class members returned by our method among
all predicted positive class members returned by our
method. Recall (R) is the proportion of predicted
positive members among all actual positive class members
in the data. F1 is the harmonic average of precision
and recall which is defined as F1 = 2PR/(P+R).
Empirical results for intent classification are shown in
tables 6, 7 &amp; 8. Results for NER are shown in tables
9, 10, 11 &amp; 12. Based on evaluation metrics, we can
clearly see that the ensemble models outperform all
other DNN models in both tasks, intent classification
as well as NER by a good margin. In case of intent
classification, the ensemble model (top 5) highlighted
in table 8 has the lowest count of false positives and
false negatives on both the dev and test data sets. It
also has the highest F1 score value = 99.91% beating
the baseline rule-based system by a margin of 1.5%.
In case of NER, the ensemble model (top 5) highlighted
in table 12 outperforms all other DNN models and
beats the baseline model by a margin of 6%.</p>
        <p>Conclusion and Future Work
Our results show an ensemble model of stacking
different DNNs of varying architectures outperforms
individual performances of DNNs for the tasks of
legal intent classification and entity recognition. RNNs,
LSTMs, GRUs and even CNNs, all compress the
necessary information of a source query into a fixed-length
vector. This makes it dificult for the DNNs to cope
with long queries, especially those that are longer than
the queries in the training corpus. In future, we plan
to use attention within queries. Attention is the idea
of freeing a DNN architecture from the fixed-length
internal representation. The DNN models we trained
are at the word level, in future we plan to expand the
size of the training data and try DNN models at the
character level. Moreover, since the diference in
performance between the DNN models were rather small,
we plan to run tests of statistical significance and error
analysis to capture performance by patterns. Lastly,
we also plan to look into the impact on our models
with respect to data and covariance shifts.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Acknowledgements</title>
      <p>This research was supported by LexisNexis, Raleigh
Technology Center, USA.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>[1] “http://www.legalexecutiveinstitute.com.”</mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>L.</given-names>
            <surname>Deng</surname>
          </string-name>
          and
          <string-name>
            <given-names>J. C.</given-names>
            <surname>Platt</surname>
          </string-name>
          , “
          <article-title>Ensemble deep learning for speech recognition,”</article-title>
          <source>in INTERSPEECH</source>
          <year>2014</year>
          ,
          <article-title>15th Annual Conference of the International Speech Communication Association</article-title>
          , Singapore,
          <source>September 14-18</source>
          ,
          <year>2014</year>
          , pp.
          <fpage>1915</fpage>
          -
          <lpage>1919</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , and
          <string-name>
            <surname>Y. Zhang,</surname>
          </string-name>
          “
          <article-title>An Ensemble of Deep Neural Networks for Object Tracking</article-title>
          ,” in
          <source>2014 IEEE International Conference on Image Processing (ICIP)</source>
          , pp.
          <fpage>843</fpage>
          -
          <lpage>847</lpage>
          ,
          <year>Oct 2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S. G.</given-names>
            <surname>Soderland</surname>
          </string-name>
          , “
          <article-title>Building a Machine Learning based Text Understanding System</article-title>
          ,”
          <volume>05</volume>
          2001.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Beitzel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. C.</given-names>
            <surname>Jensen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Frieder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. D.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Chowdhury</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Kolcz</surname>
          </string-name>
          , “
          <article-title>Improving Automatic Query Classification via Semi-supervised Learning,”</article-title>
          <source>in Fifth IEEE International Conference on Data Mining (ICDM'05)</source>
          , pp.
          <volume>8</volume>
          pp.-,
          <year>Nov 2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>D.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Pan</surname>
          </string-name>
          , J.-T. Sun,
          <string-name>
            <given-names>J. J.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yin</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Q.</given-names>
            <surname>Yang</surname>
          </string-name>
          , “
          <article-title>Q2c@ust:Our Winning Solution to Query Classification in KDDCUP 2005,” SIGKDD Explorations</article-title>
          , vol.
          <volume>7</volume>
          , pp.
          <fpage>100</fpage>
          -
          <lpage>110</lpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Z. M.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , “
          <article-title>DeepIntent: Learning Attentions for Online Advertising with Recurrent Neural Networks,”</article-title>
          <source>in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</source>
          , San Francisco, CA, USA,
          <year>August</year>
          13-
          <issue>17</issue>
          ,
          <year>2016</year>
          , pp.
          <fpage>1295</fpage>
          -
          <lpage>1304</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S.</given-names>
            <surname>Hochreiter</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Schmidhuber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>“Long</given-names>
            <surname>Short-Term</surname>
          </string-name>
          <string-name>
            <surname>Memory</surname>
          </string-name>
          ,” Neural Comput., vol.
          <volume>9</volume>
          , pp.
          <fpage>1735</fpage>
          -
          <lpage>1780</lpage>
          , Nov.
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>Chung</surname>
          </string-name>
          , Ç. Gülçehre,
          <string-name>
            <given-names>K.</given-names>
            <surname>Cho</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bengio</surname>
          </string-name>
          , “
          <article-title>Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling,” CoRR</article-title>
          , vol.
          <source>abs/1412.3555</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A.</given-names>
            <surname>Krizhevsky</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Sutskever</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G. E.</given-names>
            <surname>Hinton</surname>
          </string-name>
          , “
          <article-title>Imagenet Classification with Deep Convolutional Neural Networks,”</article-title>
          <source>in Advances in Neural Information Processing Systems 25: 26th Annual Conference on Neural Information Processing Systems 2012. Proceedings of a meeting held December 3-6</source>
          ,
          <year>2012</year>
          ,
          <string-name>
            <given-names>Lake</given-names>
            <surname>Tahoe</surname>
          </string-name>
          , Nevada, United States., pp.
          <fpage>1106</fpage>
          -
          <lpage>1114</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhang</surname>
          </string-name>
          and Y. LeCun, “
          <article-title>Text Understanding from Scratch,” CoRR</article-title>
          , vol.
          <source>abs/1502.01710</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>J.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. H.</given-names>
            <surname>Lochovsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sun</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Z.</given-names>
            <surname>Chen</surname>
          </string-name>
          , “
          <article-title>Understanding User's Query Intent with Wikipedia,”</article-title>
          <source>in Proceedings of the 18th International Conference on World Wide Web, WWW</source>
          <year>2009</year>
          , Madrid, Spain,
          <source>April 20-24</source>
          ,
          <year>2009</year>
          , pp.
          <fpage>471</fpage>
          -
          <lpage>480</lpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>H. B. Hashemi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Asiaee</surname>
            , and
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Kraft</surname>
          </string-name>
          , “
          <article-title>Query Intent Detection using Convolutional Neural Networks,”</article-title>
          <source>in WSDM QRUMS Workshop</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>J. P. C.</given-names>
            <surname>Chiu</surname>
          </string-name>
          and E. Nichols, “
          <article-title>Named Entity Recognition with Bidirectional LSTM-CNNs,” TACL</article-title>
          , vol.
          <volume>4</volume>
          , pp.
          <fpage>357</fpage>
          -
          <lpage>370</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>N.</given-names>
            <surname>Limsopatham</surname>
          </string-name>
          and
          <string-name>
            <given-names>N.</given-names>
            <surname>Collier</surname>
          </string-name>
          , “
          <article-title>Bidirectional LSTM for Named Entity Recognition in Twitter Messages,”</article-title>
          <source>in Proceedings of the 2nd Workshop on Noisy User-generated Text, NUT@COLING</source>
          <year>2016</year>
          , Osaka, Japan, December
          <volume>11</volume>
          ,
          <year>2016</year>
          , pp.
          <fpage>145</fpage>
          -
          <lpage>152</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>K.</given-names>
            <surname>Sugathadasa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Ayesha</surname>
          </string-name>
          , N. de Silva,
          <string-name>
            <given-names>A. S.</given-names>
            <surname>Perera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Jayawardana</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lakmal</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Perera</surname>
          </string-name>
          , “
          <article-title>Legal Document Retrieval using Document Vector Embeddings and Deep Learning</article-title>
          .,” CoRR, vol. abs/
          <year>1805</year>
          .10685,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>R.</given-names>
            <surname>Nanda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. J.</given-names>
            <surname>Adebayo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. D.</given-names>
            <surname>Caro</surname>
          </string-name>
          , G. Boella, and L. Robaldo, “
          <article-title>Legal Information Retrieval using Topic Clustering and Neural Networks,” in COLIEE 2017</article-title>
          .
          <article-title>4th Competition on Legal Information Extraction and Entailment, held in conjunction with the 16th</article-title>
          <source>International Conference on Artificial Intelligence and Law (ICAIL</source>
          <year>2017</year>
          )
          <article-title>in King's College London</article-title>
          , UK., pp.
          <fpage>68</fpage>
          -
          <lpage>78</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>A. H. N.</given-names>
            <surname>Tran</surname>
          </string-name>
          , “
          <article-title>Applying Deep Neural Network to Retrieve Relevant Civil Law Articles</article-title>
          ,”
          <source>in Proceedings of the Student Research Workshop Associated with RANLP</source>
          <year>2017</year>
          , (Varna), pp.
          <fpage>46</fpage>
          -
          <lpage>48</lpage>
          ,
          <string-name>
            <given-names>INCOMA</given-names>
            <surname>Ltd</surname>
          </string-name>
          .,
          <year>September 2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Xu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Yu</surname>
          </string-name>
          , “
          <article-title>Bidirectional LSTM-CRF Models for Sequence Tagging</article-title>
          .,” CoRR, vol.
          <source>abs/1508</source>
          .
          <year>01991</year>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>M. E.</given-names>
            <surname>Peters</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Neumann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Iyyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gardner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Clark</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          , “Deep Contextualized Word Representations.,” CoRR, vol. abs/
          <year>1802</year>
          .05365,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>T.</given-names>
            <surname>Mikolov</surname>
          </string-name>
          , I. Sutskever,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. S.</given-names>
            <surname>Corrado</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Dean</surname>
          </string-name>
          , “
          <article-title>Distributed Representations of Words and Phrases and their Compositionality</article-title>
          .,” in
          <string-name>
            <surname>NIPS (C. J. C. Burges</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Bottou</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Ghahramani</surname>
            , and
            <given-names>K. Q.</given-names>
          </string-name>
          <string-name>
            <surname>Weinberger</surname>
          </string-name>
          , eds.), pp.
          <fpage>3111</fpage>
          -
          <lpage>3119</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>