<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Best Matching Algorithm to Identify and Rank the Relevant Statutes</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Kayalvizhi S</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thenmozhi D</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chandrabose Aravindan</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>SSN College Of Engineering</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chennai</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>SSN College Of Engineering</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chennai</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>SSN College Of Engineering</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chennai</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>Automated identification of relevant documents or statutes for the given query is a time eficient process. Artificial Intelligence in Legal Assistance task (AILA) task is about using the artificial intelligence to identify the suitable prior documents or statute for the given query and semantic segmentation of the legal documents. We have identified the suitable statute for the given query using Best Matching (BM25) ranking algorithm. We have calculated the scores using BM25 algorithm and sorted the statutes according to the scores and then identified the suitable statutes with better MAP score of 0.2975 than most of the other approaches on AILA@FIRE-2020 dataset.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Artificial Intelligence</kwd>
        <kwd>Best Matching</kwd>
        <kwd>Identify</kwd>
        <kwd>Ranking</kwd>
        <kwd>Sorting</kwd>
        <kwd>Statute</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
    </sec>
    <sec id="sec-2">
      <title>2. Task Description</title>
      <p>Artificial Intelligence in Legal Assistance task is about identifying the relevant cases or statutes
for the given query and semantic segmentation. There are two tasks namely Task 1: Precedent</p>
      <sec id="sec-2-1">
        <title>Task 1</title>
      </sec>
      <sec id="sec-2-2">
        <title>No. of queries</title>
      </sec>
      <sec id="sec-2-3">
        <title>Object case_docs</title>
      </sec>
      <sec id="sec-2-4">
        <title>Object_statutes</title>
      </sec>
      <sec id="sec-2-5">
        <title>Task 2 No. of case documents</title>
        <p>&amp; Statute retrieval and Task 2: Rhetorical Role Labeling for Legal Judgements. Task 1 has
subtasks with Task 1a for identifying the precedent case document and Task 1b for identifying the
relevant statutes. Task 2 is the classification task of classifying the legal statements in one of
the seven semantic segments/rhetorical roles.</p>
        <p>The seven classes include:
Facts : the events that led to the case.</p>
        <p>Ruling by Lower Court : the decision of the lower court
Argument : arguments
Statute : relevant prior statute
Precedent : relevant prior case
Ratio of the decision : statement given for judgement
Ruling by Present Court : the decision of the Supreme court</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Dataset Description</title>
      <p>The AILA@FIRE-2020 data set has 2914 case documents and 197 case statutes from which the
relevant case documents and statutes has to be identified. The training set had 50 queries
along with their relevant case documents and statutes. 10 queries were given as a testing set
for identifying the relevant case document and statutes. For Task 2 of Rhetorical Role Labeling
for Legal Judgements, 50 case documents along with their seven labels of rhetorical roles were
given as training set and the test set is of 10 case documents which has to be classified according
to the semantic segments which is explained in Table1.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Proposed Methodology</title>
      <p>
        We have presented a ranking algorithm for the task 1b of identifying the relevant statutes
among the tasks provided by the Artificial Intelligence for Legal Assistance. BM25 [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] is the
Best Matching ranking algorithm that ranks the documents according to its relevance to the
given search query. It is a bag-of-words retrieval function where the documents are ranked on
the basis of appearance of query terms in each document.
4.1.2. Ranking
4.1.3. Sorting
      </p>
    </sec>
    <sec id="sec-5">
      <title>5. Results</title>
      <p>4.1. Task 1b: Identifying relevant statutes - BM25
In this approach, the statutes are pre-processed, vectorized and ranked according to the
relevant queries.</p>
      <sec id="sec-5-1">
        <title>4.1.1. Pre-processing</title>
        <p>All the given statutes are written in a single file by removing the characters that are not letters
or numbers, the stop words in the NLTK corpus stop words and converting the words in the
sentences to a lower case.</p>
        <p>For ranking the statutes, initially a word dictionary is built with the words of the statutes except
“for, a, of, the, and, to, in" and its frequency. Then, the document length, average document
length , inverse term frequency are also calculated to calculate the BM25 score. The default
parameters are taken which is 1.5 for k (saturation parameter) and 0.75 for b(length parameter)
to calculate the score.</p>
        <p>After calculating the score of individual statute document vs queries, the statutes are sorted
according to their scores.</p>
        <p>The results of Task 1b of identifying the relevant statutes is shown in Table 2. The performance
of our team is denoted by SSN_NLP. Our approach has attained a MAP score of 0.2975, BPREF
score of 0.2531, recip_rank of 0.4769 and P@10 score of 0.15. Our approach performed better
than other approaches and is comparable to the top ranking approaches.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>Thus, artificial intelligence can be used for the identification of precedent documents to identify
the documents efectively. We have used Best Matching ranking algorithm (BM25) to rank the
relevant statute with respect to the given query which have attained better results than most of
the other approaches with a MAP score of 0.2975, BPREF score of 0.2531, recip_rank of 0.4769
and P@10 score of 0.15. The performance can further be improved by altering the values of
hyper parameters such as saturation parameter and length parameter.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>We would like thank the Department Of Science and Technology (DST)-SERB funding scheme
and HPC laboratory for providing the resources and space for our research.</p>
      <p>SSN_NLP
SSNCSE_NLP
IMS_UNIPD
Uottawa_NLP</p>
      <p>UB</p>
      <p>LAWNICS
TUW-informatics
nlpninjas
scnu_1
fs_hit_1</p>
      <p>fs_hu
fs_hit_2
0.2975
0.3423
0.3383
0.2506
0.3134
0.2962
0.2619
0.00917
0.3851
0.2139
0.235
0.2003
0.2531
0.136
0.279
0.186
0.2633
0.2812
0.2033
0.024
0.3054
0.1587
0.198
0.1587
0.4769
0.3423
0.5349
0.3144
0.5787
0.4607
0.4855
0.1204
0.5615
0.3371
0.3581
0.3452</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>P.</given-names>
            <surname>Bhattacharya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mehta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Pal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bhattacharya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Majumder</surname>
          </string-name>
          ,
          <article-title>Overview of the FIRE 2020 AILA track: Artificial Intelligence for Legal Assistance</article-title>
          ,
          <source>in: Proceedings of FIRE 2020 - Forum for Information Retrieval Evaluation</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>P.</given-names>
            <surname>Bhattacharya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Pal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mehta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bhattacharya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Majumder</surname>
          </string-name>
          ,
          <article-title>Overview of the fire 2019 aila track: Artificial intelligence for legal assistance</article-title>
          .,
          <source>in: FIRE (Working Notes)</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>12</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Mandal</surname>
          </string-name>
          , S.
          <string-name>
            <surname>D. Das</surname>
          </string-name>
          ,
          <article-title>Unsupervised identification of relevant cases &amp; statutes using word embeddings</article-title>
          .,
          <source>in: FIRE (Working Notes)</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>31</fpage>
          -
          <lpage>35</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Kayalvizhi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Thenmozhi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Aravindan</surname>
          </string-name>
          ,
          <article-title>Legal assistance using word embeddings</article-title>
          .,
          <source>in: FIRE (Working Notes)</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>36</fpage>
          -
          <lpage>39</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Robertson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zaragoza</surname>
          </string-name>
          ,
          <article-title>The probabilistic relevance framework: BM25 and beyond</article-title>
          , Now Publishers Inc,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S.</given-names>
            <surname>Renjit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Idicula</surname>
          </string-name>
          , Cusat nlp@
          <fpage>aila</fpage>
          -fire2019:
          <article-title>Similarity in legal texts using document level embeddings</article-title>
          .,
          <source>in: FIRE (Working Notes)</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>25</fpage>
          -
          <lpage>30</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Ning</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Han</surname>
          </string-name>
          ,
          <string-name>
            <surname>L</surname>
          </string-name>
          . Kong, H. Qi,
          <article-title>Fire2019@ aila: Legal retrieval based on information retrieval model</article-title>
          .,
          <source>in: FIRE (Working Notes)</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>64</fpage>
          -
          <lpage>69</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>B.</given-names>
            <surname>Gain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Bandyopadhyay</surname>
          </string-name>
          , A. De,
          <string-name>
            <given-names>T.</given-names>
            <surname>Saikh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ekbal</surname>
          </string-name>
          , Iitp at aila 2019:
          <article-title>System report for artificial intelligence for legal assistance shared task</article-title>
          .,
          <source>in: FIRE (Working Notes)</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>19</fpage>
          -
          <lpage>24</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Shao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Ye</surname>
          </string-name>
          , Thuir@ aila
          <year>2019</year>
          :
          <article-title>Information retrieval approaches for identifying relevant precedents and statutes</article-title>
          .,
          <source>in: FIRE (Working Notes)</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>46</fpage>
          -
          <lpage>51</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>