<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The Gap between Deep Learning and Law: Predicting Employment Notice</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jason T. Lam</string-name>
          <email>jason.lam@queensu.com</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Samuel Dahan</string-name>
          <email>samuel.dahan@queensu.ca</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>David Liang</string-name>
          <email>david.liang@queensu.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Farhana Zulkernine</string-name>
          <email>farhana@cs.queensu.com</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Faculty of Law, Queen's University</institution>
          ,
          <addr-line>Kingston, Ontario</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Faculty of Law, Queen's University</institution>
          ,
          <addr-line>Kingston, Ontario</addr-line>
          ,
          <institution>Cornell Law School, Cornell University</institution>
          ,
          <addr-line>Ithaca</addr-line>
          ,
          <country>New York</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Queen's University</institution>
          ,
          <addr-line>Kingston, Ontario</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>School of Computing, Queen's University</institution>
          ,
          <addr-line>Kingston, Ontario</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <abstract>
        <p>This study aims to determine whether Natural Language Processing with deep learning models can shed new light on the Canadian calculation system for employment notice. In particular, we investigate whether deep learning can enhance the predictability of notice period, that is, whether it is possible to predict notice period with high accuracy. A major challenge with the classification of reasonable notice is the inconsistency of the case law. As argued by the Ontario Court of Appeal, the process of determining reasonable notice is "more art than science". In a previous study, we assessed the predictability of reasonable notice periods by applying statistical machine learning to a hand-annotated dataset of 850 cases. Building on this past study, this paper utilizes state-of-the-art deep learning models on a free-text summary of cases. We further experiment with a variety of domain adaptations of state-of-the-art pretrained BERT-esque models. Our results appear to show that the domain adaptations of BERT-esque models negatively afected performance. Our best performing model was an out-of-the-box RoBERTa base model which achieved a 69% accuracy using a +/-2 prediction window.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>CCS CONCEPTS</title>
      <p>• Applied computing → Law; • Computing methodologies →</p>
      <sec id="sec-1-1">
        <title>Artificial intelligence ; Natural language processing.</title>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>INTRODUCTION</title>
      <p>
        The system of law that governs work in Canada (outside Quebec)
consists of three overlapping regimes: the common law regime, the
regulatory regime and the collective bargaining regime (also called
labour law or the law of unionized workers). This paper focuses on
the common law regime and in particular the employment law
principles that apply to notice of termination, one of the most litigated
issues in Canada. One peculiarity of the Canadian system is that
while each province has specific regulatory standards, common
law of employment, in principle, applies in the western provinces
of Canada [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Common law is usually defined as a system of
judge-made rules that uses a precedent-based approach to case law.
Earlier decisions pertaining to similar facts or legal issues guide
later decisions in an attempt to create legal predictability. In an
effort to ensure legal certainty and predictability, judges must follow
the reasoning in earlier cases that address the same legal issues and
similar facts. That said, common law rules and their interpretations
evolve with societal values, and therefore, the interpretation and
application of common law principles can sometime be
inconsistent and unpredictable, including when it comes to termination
notice [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
      </p>
      <p>
        Upon termination, if the employment relationship is governed
by an indefinite contract and if there is no termination provision
limiting the employee’s rights, the employer has the obligation to
provide either notice or pay in lieu of notice1. Should the employer
fail to comply with this obligation, courts attempt to determine
what compensation the employee would have received during that
period if adequate notice had been provided, as well as damages
for that loss, less any mitigation income. Courts typically begin
their analysis of what constitutes "reasonable notice" by looking
at the so-called "Bardal factors", described in the landmark case
Bardal v. Globe &amp; Mail Ltd: 1) age of the employee, 2) length of
service, 3) character of employment and 4) availability of similar
employment2 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
1There is no obligation to provide reasonable notice when there is a lawful termination
provision (see Machtinger v. HOJ Industries, SCC) or where there is a fixed-term
contract [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]
2Most employment contracts are of indefinite length, and the law implies a term that
employers must provide employees with reasonable notice that the relationship is
ending. See, e.g., Machtinger v. HOJ Industries Ltd, [1992] 1 SCR 986, 7 OR (3d) 480
      </p>
      <p>
        While the Bardal test has been designed as an objective
calculation system, there is no clear indication of how much weight should
be given to each factor, nor of how the factors should be utilized [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
Accordingly, the case law on employment notice has been noted
to be inherently inconsistent and subjective, and there does not
seem to be a "right" figure for reasonable notice. As Justice
Dunphy wrote of calculation of reasonable notice periods, "[it] is more
art than science but must be one that is fair in all of the
circumstances" [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. There have also been additional layers of complication
as judges have considered factors beyond those explicitly
mentioned in Bardal, such as inducement, in which an employer’s act
in bad faith results in aggravated damages (Wallace Damages) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        In this paper we investigate whether deep learning models can
enhance the predictability of termination notice. Using advances in
pretrained models such as BERT [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], we investigate the efectiveness
of domain adaptations as well as benchmark our results against
deep learning models that have shown success across multiple
domains and in law. We build on previous work in Dahan et al. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ],
in which we analyzed the prediction of reasonable notice using
statistical machine learning on a hand-annotated tabular dataset
and demonstrated a lack of consistency amongst the judgements.
      </p>
      <p>In the balance of this paper we present related deep learning
research for legal text analytics and a description of the problem,
followed by a discussion of the data, models and methods utilized
in this research. We end with our results and discussions followed
by a conclusion and recommendations for future work.
2</p>
    </sec>
    <sec id="sec-3">
      <title>RELATED WORK</title>
      <p>
        In literature, the application of NLP to legal analytics is still new and
there exist very few implementations in predicting court decisions
or classification of legal data [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Soh et al. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] explored
multiple statistical machine learning approaches, out-of-the-box deep
learning models such as BERT and a shallow convolutional neural
network to classify 6,277 Singapore Supreme Court Judgements into
their 31 diferent legal areas. Their results showed that the
statistical model performed best with a micro-F1 of 63.2 and a macro-F1
of 73.3 [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Our problem statement difers from Soh et al. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] as
they classified the area of law utilizing the entirety of a case while
we wish to predict the outcome. While in machine learning, 3-4
instances of a sample is considered to be too few, in law, 3-4 samples
of precedence are considered to be plenty. Few-shot predictors
attempt to generalize classes with few training samples. Luo et al. [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]
proposed a model for predicting criminal charges leveraging
related law articles. They used a hierarchical attention mechanism to
create a document representation and a diferent stack of attention
components was trained to select the best supporting legislative
statutes for a given case. Their model was trained on 50,000 case
documents extracted from China Judgments Online3, and only
predicted criminal charges that had at least 80 cases. For the sake of
simplicity, the authors only considered cases in which there was
a single defendant. Luo et al. [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] reported an F1 micro/macro of
90.21/80.48 and compared the performance of their model to other
baselines they had built. Hu et al. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] performed three experiments
that trained multiple baselines including the model presented by
Luo et al. [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], on three datasets comprising 61,589/153,521/306,900
3http://wenshu.court.gov.cn/
factual case summaries generated from China Judgments Online.
They reported better results than Luo et al. [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] macro-F1 scores of
64.0/67.1/73.1 on their small/medium/large datasets, respectively.
We note in Hu et al. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] they segment the maximum document
length to 500 and China Judgements Online appear to be a database
of fact descriptions not full cases. Chalkidis et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], used a dataset
of European Court of Human Rights(ECHR) cases totalling 11,500.
They reported the results of a variety of deep learning architectures
on three tasks: binary classification, multi-class classification and
case importance prediction. They further introduced a
HierarchicalBERT, which first produced fact embeddings which are used with a
self-attention mechanism to formulate the document embedding.
Their Hierarchical-BERT was their performing model in both their
binary and multi-label classification tasks with an F1 of 82 and 60.8.
In their case importance task, their majority-class classifier achieved
the lowest mean-squared error of 0.369, with Hierarchical-BERT
achieving the best Spearman’s  of .527. Although we implement
similar models to Chaldkidis et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], we note that our task of
predicting an outcome difers from the classification of an entire
document.
      </p>
      <p>We note that as of this writing, there does not appear to be
existing literature on using deep learning for the prediction of
reasonable notice from free text.
3</p>
    </sec>
    <sec id="sec-4">
      <title>PROBLEM DESCRIPTION</title>
      <p>The aim of this project is to predict how judges determine
‘reasonable notice’ in employment termination cases. As argued earlier, if
the employment relationship is governed by an indefinite contract
and the employer wishes to terminate the employment relationship
for any reason, the employer has the obligation to provide notice
— usually denoted in months — or pay in lieu of notice, calculated
according to the Bardal factors4.</p>
      <p>While the primary goal of this research is to assess the
predictive power of deep learning models when it comes to notice of
termination, it is worth noting that this research is drawn from
a larger project aiming at developing an open-source system for
small-claims disputes, including employment disputes:
MyOpenCourt. This system aims to promote access to small-claims justice
and provide legal help to self-represented litigants by democratizing
legal analytics technology. Like many AI legal tools, MyOpenCourt
provides legal information that requires users to fill out a
multiplechoice questionnaire and outputs a single numerical prediction
along with a list of relevant precedents.</p>
      <p>
        Instead, in this research we propose a system that outputs a
prediction of reasonable notice based on a free-text summary of
the case law. We used deep learning in conjunction with Natural
Language Processing (NLP) to calculate the period of reasonable
notice from manually typed summaries of adjudicated cases. The
summaries are unstructured text data written in plain English (i.e.
not "legalese"), collected from WestLaw’s Quantum service5. These
summaries contain suficient information for a trained legal
professional to approximately determine the notice award without
looking at the outcome of the case. We decided to use summaries
4Payment in lieu of notice’ is immediate compensation at an amount equal to that an
employee would have earned as salary or wages by working through the whole notice
period
5https://www.westlawnextcanada.com/quantums/
instead of entire cases because of the the "fussy" nature of common
law text [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. The lack of styling in Canadian cases made it dificult
to automatically parse the extraction of the legal facts. As our goal
is to predict the outcome of a case, we require the inputs to not
contain any mention of the legal analysis or the outcome. The full
legal cases are long, often more than ten pages and averaging 5,000
words, and introduce ambiguity through an abundance of
information that must be carefully filtered to extract useful information.
Furthermore, legal cases are decided by a multitude of judges, which
leads to many idiosyncrasies in writing styles and case structuring.
4
      </p>
    </sec>
    <sec id="sec-5">
      <title>DATA</title>
      <p>The WestLaw Quantum Service5 provides a brief synopsis of the
case, often including the judgement on the reasonable notice period.
Since the goal of this research is to predict reasonable notice period,
any mention of the judgement regarding the notice period was
manually removed, leaving only factual descriptions of the plaintif
and the characteristics of the case. We prepended the input with
the year of the judgement, occupation category, age, salary, job
title and duration of employment of the plaintif extracted from
the summaries. If the information could not be found, nothing was
prepended to the summary. The summaries appeared to be written
in plain English by human writers. Each case outcome is considered
to be its own class, with outcomes of 25 months or greater being
grouped together. Our classification task had a total of 25 classes.
5</p>
    </sec>
    <sec id="sec-6">
      <title>MODELS AND METHODS</title>
      <p>Our dataset totaled 1,695 cases for training and 409 for testing. We
did not utilize a development set and evaluated the final models on
the testing set. All experiments were completed on an IBM Power8
server with 512 GB of memory, 64 cores (hyper-threaded to 128),
four Nvidia K80 GPUs, Red Hat Enterprise Linux Server release 7.6
and ppc64le architecture.
5.1</p>
    </sec>
    <sec id="sec-7">
      <title>Hierarchical Attention Network</title>
      <p>
        Extracting deep semantic and contextual understanding from text
data is essential for every NLP task. Bahdanau et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] first
proposed attention mechanisms for machine translation by learning
how to align the original text with the translated words. Rather
than the traditional approach of attempting to distill an entire
document into a vector, Yang et al. [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] introduced the Hierarchical
Attention Network (HAN) to encode smaller chunks of text that are
then used to inform future encodings. The HAN first learned the
importance of each word which informed its sentence embedding,
and a separate attention mechanism learned the importance of each
sentence to inform the final document representation. In our HAN
we used SpaCy sentence bound detection and tokenization [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>We utilized pretrained 200-dimension GloVe vectors with a LSTM
containing hidden dimensions of 75 and an attention dimension
of 50. We optimized with a stochastic gradient descent optimizer
with a learning rate of 0.06 , a batch size of 32, a momentum of 0.9
and dropout of 0.5. The learning rate was reduced by a factor of
0.95 when our performance stopped improving. Each epoch took
approximately six minutes to execute.
5.2</p>
    </sec>
    <sec id="sec-8">
      <title>Few-shot Learning</title>
      <p>
        While in machine learning 3-4 instances of a sample is
considered to be too few, in law, 3-4 samples are considered to be plenty.
Few-shot models often utilize a method which only requires a few
instances of a training sample to be able to generalize. Given our
sparse dataset and the success Hu et al. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] had demonstrated in
criminal law, we implemented a modified version for the prediction
of reasonable notice. We followed the implementation details
presented in Hu et al. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] except we adopted the sentence embeddings
presented by Lin et al.[
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] and generated r number of attention
vectors for each attribute, as the original model yielded poor results.
The architecture presented by Hu et al. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] created a document
representation by combining an attribute-aware embedding with
one that is attribute-free. The attribute-aware embedding resulted
from an average pooling of the sentence embeddings from four
stacked self-attention mechanisms. For a single mechanism we
predict an attribute (e.g. a person’s age or duration) by self-attending
to multiple parts of the text simultaneously. Four labels were used
to train the attribute-aware mechanisms to predict the length of
employment, age of employee, character of employment, and
availability of similar employment. These labels were hand-annotated
by a team of Queen’s Law students. The attribute-free embedding
comprised a max-pooling of the hidden states generated by the
encoder.
      </p>
      <p>We utilized pretrained 300-dimension GloVe vectors that were
ifne-tuned, a hidden dimension of 300, and dropout of 0.5. An
attention r of 30 was used. We used an Adam optimizer with a
learning rate of 0.001 and reduced the learning rate when a metric
has stopped improving by a factor of 0.95. Alpha, the scaling on
the attribute-aware loss, had a value of 0.3 and each epoch took
approximately 8 minutes to execute.
5.3</p>
    </sec>
    <sec id="sec-9">
      <title>BERT-esque and Domain adaptation</title>
      <p>
        The Bidirectional Encoder Representations from Transformers (BERT)
from Devlin et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] has recently laid the foundation of pretraining
models by leveraging language modelling, transfer learning, and
ifne-tuning on downstream tasks. As one of the main strengths of
BERT is its generalized understanding of the English language, and
pretrained BERT models have been publicly released6, we leveraged
RoBERTa for predicting reasonable notice. We further experimented
with Robustly Optimized BERT Pretraining Approach (RoBERTa),
which held the top spot in the GLUE benchmarks [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] during the
course of our research. Architecturally, RoBERTa did not difer
from the original BERT model; instead, Liu et al. [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] utilized an
additional 160GBs of data and further fine-tuned hyperparameters.
Through experimentation of learning rates and batch sizes, the
authors determined that the next sentence prediction only ofered
marginal to no performance improvements. Using only
hyperparameter fine-tuning and additional data, RoBERTa achieved an
almost 7% improvement over the original BERT model on the GLUE
benchmarks.
5.3.1 Domain Adaptations. Following common practice for
improving pretrained model performance, we experimented by
domainadapting our BERT-esque models on full reasonable notice case
6https://github.com/google-research/bert
texts as well as the Harvard case law dataset [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. In our BERT
implementations, we utilized the  model from HuggingFace7.
We further experimented with domain adapting a  
model using Facebook AI’s implementation8. In addition, we further
pretrained both models using only the masked language modelling
(MLM) criterion, consistent with the results from Liu et al. [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ].
Five epochs were used for all MLM pretraining. For our end results,
pretrained language models were further trained for text
classiifcation with ten epochs and fine-tuned on 409 remaining cases.
Classification training was performed using a batch size of 16, using
the default losses and optimizers. The classification head of BERT
took 20 minutes to train while MLM took 2.5 hours. For BERT,
we fine-tuned on the full case text that correspond to each of the
respective 1,695 reasonable notice cases in our training set.
      </p>
      <p>
        In our RoBERTa implementation, we fine-tuned on
approximately four million cases from Harvard’s case law project. We
determined cases before 1960 to be linguistically diferent from
present-day legal documents and thus these cases were removed.
We note that the Harvard case law project only includes cases from
the United States. To create an accurate comparison with our BERT
implementation, we domain-adapted   to the same set
of full case texts. The classification head of RoBERTa took 30
minutes to fine-tune, while MLM took 3 hours. Our batch size was 256
and peak learning rate was 0.0001 in accordance with Liu et al. [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ].
6
      </p>
    </sec>
    <sec id="sec-10">
      <title>RESULTS AND DISCUSSION</title>
      <p>Approach Acc. (+/-2)
HAN 67%
Few-shot w/ Self-attention 51%
BERT+base 61%
BERT+full cases 49%
RoBERTa+full cases 63%
RoBERTa+Harvard 65%</p>
      <sec id="sec-10-1">
        <title>RoBERTa+base 69% Table 1: Summary of results for predicting the number of months awarded for reasonable notice using case summaries.</title>
        <p>The output of our system was classiefid as correct if it was within
+/-2 months of the ground truth label to account for situational
variability (i.e.Eq 1). We refer to this as the output window.
ℎ − 2 ≤  ≤ ℎ + 2
(1)</p>
        <p>Our   model had the highest accuracy of 69%. A
summary of the results can be seen in Table 1.</p>
        <p>
          Interestingly, our best-performing model was  
outof-the-box and not a domain-adapted version. Our HAN out-performed
the majority of our pretrained models and was only marginally
worse than   . The performance of HAN may be
attributable to architectural structuring, as each sentence in our case
summaries roughly contains a statement of fact, allowing our HAN
7huggingface.co
8https://github.com/pytorch/fairseq
to learn the importance of each sentence (fact) and weigh it
accordingly prior to creating a document representation. This may
be replicating the thought process of the judiciary. Furthermore,
we believe our mixed results from our few-shot model may be
attributable to the size of our dataset. In Hu et al. [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ], their smallest
training dataset contained over 61,000 training samples, compared
to our 1,695. The issue of data in our few-shot model appear to
be supported with the performance of our BERT-esque models in
which the majority of models performed better. BERT-esque models
are pretrained to have a generalized understanding of language out
of the box, making it easier to fine-tune on a specific classification
task. Furthermore, the task of predicting reasonable notice requires
additional knowledge beyond an understanding of the natural
language. In fact, it requires knowledge on judicial bias and dispute
settlement - something that seasoned employment lawyers have
built throughout their interactions with judges and colleagues and
cannot be learned from case law. This may partially explain our
mixed results insofar as the majority of employment disputes are
resolved through negotiation. Thus, considering that the case law
constitutes only a small piece of the data, it may be argued that our
model could have performed better had it be trained on settlement
agreements.
        </p>
        <p>
          In our experiments we utilized a classification approach over a
regression as we believe this best replicates the decision-making
process of a judge. In deciding the amount of reasonable notice a
plaintif should receive one could presume a judge would locate
similar past cases and adjust their ruling based on diferences of
fact. We utilize classification to anchor the number of months of
reasonable notice and broaden the output window in our metric
to account for the diferences of fact. In addition, our findings in
Dahan et al. [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], in which we used regression along with a
handlabeled tabular dataset, indicated regression to be a poor predictor
of reasonable notice.
        </p>
        <p>
          A key finding of this research is that our domain adaptations
did not yield significant improvements when compared to the
outof-the-box pretrained models, as reported in Rietzler et al. [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ]
and commonly noted as a promising avenue for other domains
(e.g. SciBERT [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] and BioBERT [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]). Instead, domain adaptations
appeared to negatively afect the performance of both our  
and  models. Despite the language of the Canadian reasonable
notice cases being more similar to our case summaries than the
Harvard Case Law dataset, our results from   ,
domainadapted on the Harvard dataset, performed slightly better than the
ones trained on full Canadian cases. This finding was in line with
one of the conclusions set out by Liu et al. [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ], which emphasized
the importance of volume for datasets. Furthermore, as Liu et al. [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]
performed extensive hyperparameter experimentation we opted to
utilize their recommended learning rates and batch sizes. It appears
current deep learning solutions may not be able to accurately grasp
the many unique characters of each legal dispute. Unfortunately,
our performance suggests that in spite of recent advances, deep
learning may not be able to accurately predict reasonable notice
from free text. That said, our mixed results may also be explained
by the fact that judges are inherently unpredictable when it comes
to the application of the Bardal test, and thus predicting notice is
an almost impossible task for deep learning models, but also for
experienced lawyers.
7
        </p>
      </sec>
    </sec>
    <sec id="sec-11">
      <title>CONCLUSION AND FUTURE WORK</title>
      <p>
        In this paper we applied multiple deep-learning solutions to
humanwritten case summaries to classify the reasonable notice period
a plaintif should be awarded. Our best performing model was
  which achieved a 69% accuracy with a +/-2 month
window. Domain adaptations negatively afected our performance
in the case of RoBERTa, and marginally improved performance in
the case of BERT. Given the significant successes of deep learning in
various domains, the relatively poor eficiency found in predicting
notice periods may come as a surprise. In fact, results from this
paper appear to be consistent with our previous findings in Dahan
et al. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], suggesting the inconsistencies in the case law make it
dificult to predict reasonable notice accurately. Thus, it can be
argued that our results have very little to do with the quality of
the model, and that it may not be possible to reduce our prediction
error, mainly because of the inherently inconsistent nature of the
dataset.
      </p>
      <p>
        At the time of our research, RoBERTa was the top-performing
model on the GLUE benchmarks [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ], where it ranks 10th as of this
writing. Utilizing the top performing model from the benchmarks
may yield better results. While our domain adaptation used 1,695
full cases, and the superior performance of domain adapting on
an adjacent domain of American law, collecting a larger dataset of
Canadian cases may be an interesting avenue to explore for further
experimentation. While deep learning and BERT-esque models
have proven successful in numerous domains, it appears a gap still
exists for deep learning applications in the legal field. That said,
other areas — such as art, which has been considered too artistic
to be impacted by AI — have shown some recent successes. In
particular, deep learning models have successfully been trained to
paint like humans [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], which leads us to be optimistic about future
applications in the legal field. We believe that further advances in
the field of NLP and deep learning will be able to perform well on
this task in the future.
      </p>
    </sec>
    <sec id="sec-12">
      <title>ACKNOWLEDGMENTS</title>
      <p>We would like to thank the team at the Conflict Analytics Lab, the
Scotiabank Centre for Customer Analytics and the Center for Law
in the Contemporary Workplace at Queen’s University ( Jonathan
Touboul, Maxime Cohen, Kevin Banks, Simon Townsend, William
Quiglietta, Zach Berg, Mackenzie Anderson, Brandon Loehle, Max
Saunders, Sean Gulrajani, Shane Liquornik, Yuri Levin, Mikhail
Nediak, Stephen Thomas) as well as our colleagues Joshua Karton
and Bill Flanagan for supporting the project as well as contributing
to the development of the idea.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1] [n.d.]. Bardal v.
          <source>Globe &amp; Mail Ltd</source>
          . (
          <year>1960</year>
          ), 24
          <string-name>
            <surname>D.L.R.</surname>
          </string-name>
          (
          <year>2d</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2] [n.d.]. Fraser v. Canerector Inc.,
          <source>2015 CarswellOnt 4796</source>
          ,
          <year>2015</year>
          ONSC 2138 at para 32.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3] [n.d.]. Wallace v. United Grain Growers Ltd.,
          <source>1997 CanLII 332 (SCC)</source>
          , (
          <year>1997</year>
          ) 3
          <string-name>
            <surname>S.C.R.</surname>
          </string-name>
          <year>701</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Dzmitry</given-names>
            <surname>Bahdanau</surname>
          </string-name>
          , Kyunghyun Cho, and
          <string-name>
            <given-names>Yoshua</given-names>
            <surname>Bengio</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Neural machine translation by jointly learning to align and translate</article-title>
          .
          <source>arXiv preprint arXiv:1409.0473</source>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Iz</given-names>
            <surname>Beltagy</surname>
          </string-name>
          , Arman Cohan, and
          <string-name>
            <given-names>Kyle</given-names>
            <surname>Lo</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Scibert: Pretrained contextualized embeddings for scientific text</article-title>
          . arXiv preprint arXiv:
          <year>1903</year>
          .
          <volume>10676</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Ilias</given-names>
            <surname>Chalkidis</surname>
          </string-name>
          and
          <string-name>
            <given-names>Dimitrios</given-names>
            <surname>Kampas</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Deep learning in law: early adaptation and legal word embeddings trained on large corpora</article-title>
          .
          <source>Artificial Intelligence and Law</source>
          <volume>27</volume>
          ,
          <issue>2</issue>
          (
          <year>2019</year>
          ),
          <fpage>171</fpage>
          -
          <lpage>198</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Bruce</given-names>
            <surname>Curran</surname>
          </string-name>
          and
          <string-name>
            <given-names>Sara</given-names>
            <surname>Slinn</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Just Notice Reform: Enhanced Statutory Termination Provisions for the 99%</article-title>
          .
          <source>Osgoode Legal Studies Research Paper</source>
          <volume>61</volume>
          (
          <year>2016</year>
          ),
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Samuel</given-names>
            <surname>Dahan</surname>
          </string-name>
          , Jonathan Touboul, Jason Lam, and
          <string-name>
            <given-names>Dan</given-names>
            <surname>Sfedj</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Predicting Employment Notice with Machine Learning: Promises and Limitation</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Jacob</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ming-Wei</surname>
            <given-names>Chang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kenton</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Kristina</given-names>
            <surname>Toutanova</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Bert: Pre-training of deep bidirectional transformers for language understanding</article-title>
          . arXiv preprint arXiv:
          <year>1810</year>
          .
          <volume>04805</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>David J Doorey</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>The Law of Work: Common Law and the Regulation of Work</article-title>
          .
          <source>Regulation</source>
          <volume>285</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>James</given-names>
            <surname>Holland and Julian Webb</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Learning legal rules: a students' guide to legal method and reasoning</article-title>
          . Oxford university press.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Matthew</given-names>
            <surname>Honnibal</surname>
          </string-name>
          and
          <string-name>
            <given-names>Mark</given-names>
            <surname>Johnson</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>An Improved Non-monotonic Transition System for Dependency Parsing</article-title>
          .
          <source>In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics</source>
          , Lisbon, Portugal,
          <fpage>1373</fpage>
          -
          <lpage>1378</lpage>
          . https://aclweb.org/anthology/D/D15/ D15-1162
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Jerrold</given-names>
            <surname>Soh Tsin Howe</surname>
          </string-name>
          , Lim How Khang, and Ian Ernst Chai.
          <year>2019</year>
          .
          <article-title>Legal Area Classification: A Comparative Study of Text Classifiers on Singapore Supreme Court Judgments</article-title>
          . arXiv preprint arXiv:
          <year>1904</year>
          .
          <volume>06470</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Zikun</surname>
            <given-names>Hu</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Xiang</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Cunchao</given-names>
            <surname>Tu</surname>
          </string-name>
          , Zhiyuan Liu, and
          <string-name>
            <given-names>Maosong</given-names>
            <surname>Sun</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Few-shot charge prediction with discriminative legal attributes</article-title>
          .
          <source>In Proceedings of the 27th International Conference on Computational Linguistics</source>
          .
          <fpage>487</fpage>
          -
          <lpage>498</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Zhewei</surname>
            <given-names>Huang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Wen</given-names>
            <surname>Heng</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Shuchang</given-names>
            <surname>Zhou</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Learning to paint with model-based deep reinforcement learning</article-title>
          .
          <source>In Proceedings of the IEEE International Conference on Computer Vision</source>
          . 8709-
          <fpage>8718</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>Grant</given-names>
            <surname>Lamond</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Precedent and analogy in legal reasoning</article-title>
          . (
          <year>2006</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>Jinhyuk</given-names>
            <surname>Lee</surname>
          </string-name>
          , Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and
          <string-name>
            <given-names>Jaewoo</given-names>
            <surname>Kang</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>BioBERT: a pre-trained biomedical language representation model for biomedical text mining</article-title>
          .
          <source>Bioinformatics</source>
          <volume>36</volume>
          ,
          <issue>4</issue>
          (
          <year>2020</year>
          ),
          <fpage>1234</fpage>
          -
          <lpage>1240</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Zhouhan</surname>
            <given-names>Lin</given-names>
          </string-name>
          , Minwei Feng, Cicero Nogueira dos Santos, Mo Yu, Bing Xiang,
          <string-name>
            <surname>Bowen Zhou</surname>
            , and
            <given-names>Yoshua</given-names>
          </string-name>
          <string-name>
            <surname>Bengio</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>A structured self-attentive sentence embedding</article-title>
          .
          <source>arXiv preprint arXiv:1703.03130</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Yinhan</surname>
            <given-names>Liu</given-names>
          </string-name>
          , Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen,
          <string-name>
            <surname>Omer Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Mike</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Luke</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Veselin</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Roberta: A robustly optimized BERT pretraining approach</article-title>
          . arXiv preprint arXiv:
          <year>1907</year>
          .
          <volume>11692</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Bingfeng</surname>
            <given-names>Luo</given-names>
          </string-name>
          , Yansong Feng, Jianbo Xu, Xiang Zhang, and
          <string-name>
            <given-names>Dongyan</given-names>
            <surname>Zhao</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Learning to Predict Charges for Criminal Cases with Legal Basis</article-title>
          .
          <source>In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing</source>
          .
          <fpage>2727</fpage>
          -
          <lpage>2736</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Alexander</surname>
            <given-names>Rietzler</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Sebastian</given-names>
            <surname>Stabinger</surname>
          </string-name>
          , Paul Opitz, and
          <string-name>
            <given-names>Stefan</given-names>
            <surname>Engl</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Adapt or get left behind: Domain adaptation through bert language model finetuning for aspect-target sentiment classification</article-title>
          . arXiv preprint arXiv:
          <year>1908</year>
          .
          <volume>11860</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Alex</surname>
            <given-names>Wang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Amanpreet</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <surname>Julian Michael</surname>
          </string-name>
          , Felix Hill,
          <string-name>
            <given-names>Omer Levy</given-names>
            , and
            <surname>Samuel R Bowman</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Glue: A multi-task benchmark and analysis platform for natural language understanding</article-title>
          . arXiv preprint arXiv:
          <year>1804</year>
          .
          <volume>07461</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Zichao</surname>
            <given-names>Yang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Diyi</given-names>
            <surname>Yang</surname>
          </string-name>
          , Chris Dyer, Xiaodong He,
          <string-name>
            <surname>Alex Smola</surname>
            , and
            <given-names>Eduard</given-names>
          </string-name>
          <string-name>
            <surname>Hovy</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Hierarchical attention networks for document classification</article-title>
          .
          <source>In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          .
          <fpage>1480</fpage>
          -
          <lpage>1489</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>