<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Unsupervised Identi cation of Relevant Cases &amp; Statutes Using Word Embeddings</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Soumil Mandal</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sourya Dipta Das</string-name>
          <email>dipta.juetceg@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Jadavpur University</institution>
          ,
          <addr-line>Kolkata</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>SRM University</institution>
          ,
          <addr-line>Chennai</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper, we have described the systems that we submitted as team JU SRM for FIRE 2019 track on Arti cial Intelligence for Legal Assistance (AILA 2019). The two tasks in this track were 1) identifying relevant prior cases and 2) identifying relevant statutes. For both of these tasks, we took an unsupervised approach using pre-trained wordembeddings for encoding texts and calculating relevance using cosinesimilarity between the query and target documents.</p>
      </abstract>
      <kwd-group>
        <kwd>Arti cial Intelligence</kwd>
        <kwd>Legal Assistance</kwd>
        <kwd>Sent2Vec</kwd>
        <kwd>Fast-</kwd>
        <kwd>Text</kwd>
        <kwd>BERT</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Similar to a lot of other practical domains, the domain of law and legality is
gradually incorporating automated methods as well, especially after the rapid
growth and development in machine learning and information retrieval models.
To encourage researchers in delving into such automated methods, FIRE 2019
included a track named Arti cial Intelligence for Legal Assistance (AILA) [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
In a lot of countries, when a lawyer is presented with a case, the nal verdict is
generally based on two things, 1) statutes (established laws) and 2) precedents
(prior cases). The statutes informs the lawyer regarding applying legal
principals based on a certain situation, while precedents informs the lawyer about
how similar cases were dealt with in the past. If this pipeline of collecting
relevant statutes and precedents can be automated as a information retrieval based
model, this will not only help the lawyer but as well as several other people
including the clients and subjects. Motivated by this, the organizers of AILA
added two tasks, 1) identifying relevant prior cases for a given situation and 2)
identifying most relevant statutes for a given situation. The goal was to given
a case description as query, rank the target documents prior to relevancy. The
datasets which the organizers provided consisted of 2914 prior cases, 197 statues
and 50 queries, which were summarized case descriptions. To test our model prior
to the nal run submissions, the organizers provided us with the task 1 and task
2 outputs of the rst 10 queries. To build our systems, we took an unsupervised
approach using text-embeddings and cosine-similarity. The primary motivation
behind this was semantic level modelling and better scaling. Before building our
systems, we performed some basic NER removal using Spacy 3 with the following
tags PERSON, ORDINAL, CARDINAL, WORK OF ART, TIME, PERCENT,
QUANTITY.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        In the legal domain, researchers have contributed in several problems like text
classi cation, text summarising and information mining Gonalves et.al [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] showed
how some linguistic techniques like lemmatization and POS identi cation can be
used to increase accuracy of the model for classi cation of legal texts with low
dimensional feature vector. As we know, legal documents follow a certain
structure of information written with formal languages and de ned terminologies,
researchers have tried to use these prior structural information to summarize
the data which can be useful for other tasks like case recommendations,
document classi cation and labeling. Saravanan et.al [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ][
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] have contributed in legal
document summarizing in their subsequent works by using graphical models and
CRFs. Conrad et.al [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] introduced a query based sentiment summarization for
legal texts which can be very useful for mining the opinions. In the legal document
labeling problem, Schweighofer et.al [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] proposed a solution by using
hierarchical self-organizing map and Mencia et.al [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] introduced a multilabel classi cation
using one vs all classi ers with e cient perceptron algorithms.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Task 1 - Identifying Relevant Prior Cases</title>
      <p>
        For this task, the goal was to identify the relevant prior case from a collection of
past cases. Here, we have used two types of word embeddings, namely pretrained
Sent2Vec [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and FastText [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] trained on the prior 2914 cases. On a whole, we
created three models. For the rst two models, we used a simple algorithm where
the queries and precedents were encoded using pretrained word-embeddings.
Instead of encoding whole cases or queries as a single long vector, we extracted
sentences of max 20 tokens, then encoded each of them, and nally calculated
the average of all these vectors to get the vector of the respective case or query
of size 20. Then, for each query-case pair, we computed the cosine-similarity
score and the ranked them accordingly. The only di erence was that for the rst
model, we used Sent2Vec, while in the second we used FastText. We tested both
of these models on the training data. The Sent2Vec model secured an BPREF
of 0.0215 while the FastText model secured an BPREF of 0.0124. Using these
values, we calculated weights of the models. For Sent2Vec, it was 0.0215/(0.0215
+ 0.0124) = 0.63 and for FastText it was 0.0124/(0.0215 + 0.0124) = 0.36. With
      </p>
      <sec id="sec-3-1">
        <title>3https://spacy.io/api/annotationnamed-entities</title>
        <p>these values, we created our third model, which was a weighted voting ensemble
model. The performance metrics 4 of all of these models on the testing data is
shown below in Table 1. 1/Ro1R denotes 1/(rank of rst relevant document).</p>
        <p>Model P@10 MAP BPREF 1/Ro1R
Sent2Vec 0.0250 0.0478 0.0284 0.131
FastTex 0.0175 0.0228 0.0163 0.065</p>
        <p>Ensemble 0.0200 0.0181 0.0060 0.044</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Task 2 - Identifying Relevant Statutes</title>
      <p>
        In this task, the goal was to identify relevant statutes given a summarized version
of a case as input. To do this, we rst extracted key-phrases from the queries
and the statutes using the rake-nltk 5 library. For statutes, we further performed
some manual augmentation as well as removal of key-phrases based on relevance.
Example, for statute S10, "equality of opportunity in matters of public
employment", the library didn't select "discriminated" as a keyword so it was manually
added while " fty per cent", which was picked up was removed. Finally, we
encoded each of these key-phrases using a pretrained BERT [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] model. Using these
encoding vectors, we created three models. For the rst model, we computed
the cosine similarity scores between each of the key-phrase pairs of every
querystatute pairs. Then, for each of the query-statute pair cosine similarity scores,
we took the max and second-max values and multiplied them to get the nal
rank determining score. For the second model, we took a similar approach as
the rst one, but this time, took an average of the key-phrase cosine similarity
scores. In the third model, we used the product of the scores calculated for the
rst and the second model to get the nal relevance score, i.e. the product of
the max, second-max and the average score. The performance metrics of all of
the models on the testing data is shown in Table 4.
      </p>
      <p>Model P@10 MAP BPREF 1/Ro1R
M*SM 0.0600 0.0767 0.0309 0.1460
Average 0.0600 0.0918 0.0402 0.2010</p>
      <p>Ensemble 0.0600 0.0831 0.0285 0.1620</p>
      <sec id="sec-4-1">
        <title>4https://trec.nist.gov/pubs/trec15/appendices/CE.MEASURES06.pdf 5https://pypi.org/project/rake-nltk/</title>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion &amp; Future Work</title>
      <p>We have demonstrated that satisfactory results in both of the tasks can be
achieved by taking a simple and fast unsupervised approach using pre-trained
embeddings and cosine-similarity scores. Our Sent2Vec and average based system
got a rank of 7 and 4 in task 1 and task 2 respectively based on the metric Ro1R.
In the future, we would like to collect more legal data and annotate them to build
supervised classi cation models.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgement</title>
      <p>The authors would like to thank Silversparro Pvt, Ltd. for providing necessary
support &amp; computational resources to complete this work. We particularly
extend our gratitude to Mr. Ankit Agarwal, CTO and Mr. Ravikant Bhargav, R&amp;D
head for their constant support and encouragement.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Schweighofer</surname>
            , Erich,
            <given-names>Andreas</given-names>
          </string-name>
          <string-name>
            <surname>Rauber</surname>
            , and
            <given-names>Michael</given-names>
          </string-name>
          <string-name>
            <surname>Dittenbach</surname>
          </string-name>
          .
          <article-title>"Automatic text representation, classi cation and labeling in European law."</article-title>
          <source>In Proceedings of the 8th international conference on Arti cial intelligence and law</source>
          , pp.
          <fpage>78</fpage>
          -
          <lpage>87</lpage>
          . ACM,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Gonalves</surname>
            , Teresa, and
            <given-names>Paulo</given-names>
          </string-name>
          <string-name>
            <surname>Quaresma</surname>
          </string-name>
          .
          <article-title>"Is linguistic information relevant for the classi cation of legal texts?."</article-title>
          <source>In Proceedings of the 10th international conference on Arti cial intelligence and law</source>
          , pp.
          <fpage>168</fpage>
          -
          <lpage>176</lpage>
          . ACM,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Saravanan</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Balaraman</given-names>
            <surname>Ravindran</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Raman</surname>
          </string-name>
          .
          <article-title>"Improving legal document summarization using graphical models."</article-title>
          <source>Frontiers in Arti cial Intelligence and Applications</source>
          <volume>152</volume>
          (
          <year>2006</year>
          ):
          <fpage>51</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Mencia</surname>
            , Eneldo Loza, and
            <given-names>Johannes</given-names>
          </string-name>
          <string-name>
            <surname>Frnkranz</surname>
          </string-name>
          .
          <article-title>"E cient pairwise multilabel classication for large-scale problems in the legal domain."</article-title>
          <source>In Joint European Conference on Machine Learning and Knowledge Discovery in Databases</source>
          , pp.
          <fpage>50</fpage>
          -
          <lpage>65</lpage>
          . Springer, Berlin, Heidelberg,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Conrad</surname>
            , Jack G., Jochen L. Leidner, Frank Schilder, and
            <given-names>Ravi</given-names>
          </string-name>
          <string-name>
            <surname>Kondadadi</surname>
          </string-name>
          .
          <article-title>"Querybased opinion summarization for legal blog entries."</article-title>
          <source>In Proceedings of the 12th International Conference on Arti cial Intelligence and Law</source>
          , pp.
          <fpage>167</fpage>
          -
          <lpage>176</lpage>
          . ACM,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Saravanan</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Balaraman</given-names>
            <surname>Ravindran</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Raman</surname>
          </string-name>
          .
          <article-title>"Automatic identi cation of rhetorical roles using conditional random elds for legal document summarization."</article-title>
          <source>In Proceedings of the Third International Joint Conference on Natural Language Processing: Volume-I</source>
          .
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Bansal</surname>
          </string-name>
          , Trapit, David Belanger, and
          <string-name>
            <surname>Andrew McCallum</surname>
          </string-name>
          .
          <article-title>"Ask the gru: Multi-task learning for deep text recommendations."</article-title>
          <source>In Proceedings of the 10th ACM Conference on Recommender Systems</source>
          , pp.
          <fpage>107</fpage>
          -
          <lpage>114</lpage>
          . ACM,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Devlin</surname>
          </string-name>
          , Jacob,
          <string-name>
            <surname>Ming-Wei</surname>
            <given-names>Chang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kenton</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Kristina</given-names>
            <surname>Toutanova</surname>
          </string-name>
          .
          <article-title>"Bert: Pretraining of deep bidirectional transformers for language understanding." arXiv preprint arXiv:</article-title>
          <year>1810</year>
          .
          <volume>04805</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Pagliardini</surname>
            , Matteo, Prakhar Gupta, and
            <given-names>Martin</given-names>
          </string-name>
          <string-name>
            <surname>Jaggi</surname>
          </string-name>
          .
          <article-title>"Unsupervised learning of sentence embeddings using compositional n-gram features</article-title>
          .
          <source>" arXiv preprint arXiv:1703.02507</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Joulin</surname>
            , Armand, Edouard Grave, Piotr Bojanowski, Matthijs Douze, Hrve Jgou, and
            <given-names>Tomas</given-names>
          </string-name>
          <string-name>
            <surname>Mikolov</surname>
          </string-name>
          .
          <article-title>"Fasttext. zip: Compressing text classi cation models</article-title>
          .
          <source>" arXiv preprint arXiv:1612.03651</source>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>P.</given-names>
            <surname>Bhattacharya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Pal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mehta</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Bhattacharya.</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Majumder</surname>
          </string-name>
          ,
          <article-title>Overview of the Fire 2019 AILA track: Arti cial Intelligence for Legal Assistance</article-title>
          .
          <source>In Proc. of FIRE 2019 - Forum for Information Retrieval Evaluation</source>
          , Kolkata, India,
          <source>December 12-15</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>