<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Summarization of Indian Legal Judgement Documents via Ensembling of Contextual Embedding based MLP Models</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Deepali Jain</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Malaya Dutta Borah</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anupam Biswas</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science &amp; Engineering, National Institute of Technology Silchar</institution>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Automatic summarization of lengthy legal documents can provide great help to the involved legal practitioners, as well as all the other end users. In this work, an extractive summarization approach has been developed which represents legal document sentences in terms of domain specific pre-trained embeddings and performs subsequent multilayer perceptron based classification to find their summary worthiness. With this approach, we have participated in the summarization related shared tasks of AILA 2021 (Task 2(a) and 2(b)). The results on the test dataset for Task 2 show that our proposed approach is able to outperform most of the other competitors, achieving 2nd position in Task 2(a) and best ROUGE-F1 scores across all the ROUGE metrics for Task 2(b). While the proposed approach has produced impressive results for Task 2, the same approach could not do well for the rhetorical labeling task (Task 1) as per the results provided by the organizers on the test dataset.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;legal bert</kwd>
        <kwd>legal judgement documents</kwd>
        <kwd>extractive summarization</kwd>
        <kwd>contextual embeddings</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Whereas, in extractive summarization, the main idea is to extract summary worthy sentences
from the input document itself to form a summary. Several research works have attempted to
summarize such lengthy legal documents [
        <xref ref-type="bibr" rid="ref1 ref3 ref4 ref5 ref6 ref7">3, 4, 5, 1, 6, 7</xref>
        ]. Some of the works have also made
use of rhetorical roles for performing the downstream summarization task [
        <xref ref-type="bibr" rid="ref8">8, 9, 10</xref>
        ]. A detailed
discussion on the various legal document summarization techniques along with several future
research directions can be found in [11]. It is important to note here that both the subtasks of
Task 2 directly correspond to the idea of extractive summarization, which is why in this work,
we have primarily focused on performing summarization of legal judgement documents via
eficient selection of summary-worthy sentences from the input documents. More specifically,
we find domain specific vectorized representation of sentences in the input documents, followed
by their summary worthiness classification, with a Multilayer Perceptron (MLP) model.
      </p>
      <p>Following the introduction, the organization of the rest of the paper is as follows: Section 2
presents a detailed methodology for the automatic legal document summarization task. Results
and analysis are given in Section 3. Finally, Section 4 concludes this paper with a summarization
of the findings along with potential future research directions.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Data and Methods</title>
      <sec id="sec-2-1">
        <title>2.1. Datasets</title>
        <p>
          The organizers of the AILA track have provided the training dataset for Task 1 and Task 2. For
Task 1, they provide 60 documents, where each sentence of a document is labeled with one of
the seven rhetorical labels (Facts, Ruling by Lower Court, Argument, Statute, Precedent, Ratio
of the decision, Ruling by Present Court) [12, 13]. For Task 2, the organizers have provided 500
document-summary pairs of judgments by the Supreme Court of India [14]. The organizers have
provided the pre-processed and sentence tokenized versions for both judgments and summaries,
along with summary-worthiness and rhetorical roles labels for each sentence in the documents.
A description of the AILA track FIRE 2021 is given in [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] and the overview of the tasks organized
in this track is presented in [15].
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Methodology</title>
        <p>In order to perform the summarization task, we have randomly split the training dataset into five
folds consisting of 400 training samples and 100 validation samples. Now, for each fold, we have
performed the steps as shown in Fig. 1(a). All the sentences are firstly converted into contextual
embeddings using a pre-trained Legal-Bert model [16]. In this way, a 768-dimensional vector is
obtained since Legal-Bert has a fixed hidden size of 768 dimensions. These sentence embeddings
are then fed through the MLP model as shown in Fig. 1(b). For this summary-worthiness
classification problem, it is to be noted here that the last dense layer consists of single node
with sigmoid activation function and binary cross-entropy loss.</p>
        <sec id="sec-2-2-1">
          <title>2.2.1. Summary-Worthy Sentence Identification Task (Task 2(a))</title>
          <p>Specifically for Task 2(a), firstly, contextual embeddings of every sentence in a document are
found using the Legal-Bert pre-trained model. Then these embeddings are fed through the MLP
5 training folds
Training documents (400 train, 100 val.)</p>
          <p>Legal Bert
embeddings</p>
          <p>Train
MLP
model</p>
          <p>Trained model
Testing documents
Dropout(0.4)
Dense(128)
Dropout(0.4)
Dense(32)
Dropout(0.4)</p>
          <p>Dense(1)
(b) MLP model
model. Since it is a binary classification task, we have considered the dense layer of 1 node
as the last layer with sigmoid activation and binary cross-entropy. This way, training is done
for all five models (for five folds). We predict the summary-worthiness probabilities of each
sentence of a test document using all the five models and take an average of all five probabilities
to get an overall probability measure. We assign label ‘1’ to the sentence if the predicted average
probability is greater than or equal to the threshold of 0.4, otherwise the ‘0’ label is assigned.
This is the exact approach utilized for Run-1 submission of our team (nits_legal).</p>
          <p>For Run-2 submission, instead of using a simple MLP model, we use a multi-task learning
based MLP model. In this approach, we have used both rhetorical and summary-worthy labels
and fed the training dataset into a multi-task learning MLP model. We have chosen to perform
a multi-task learning based approach hoping that learning rhetorical labeling might be helpful
for appropriately predicting the summary-specific relevance labels of sentences. The exact MLP
model utilized for this is depicted in Fig. 2. The training of such a model again results in 5
diferent versions for five diferent folds, which enables averaging based model ensembling.
At inference time, we predict in the same manner as that of Run-1 using all the individual
multi-task learning models.</p>
          <p>For Run-3 submission, we take the average of all the individual summary-worthiness
probabilities resulting from each individual trained model of Run-1 and Run-2. If the average probability
value is found to be greater than or equal to 0.4, we assign label ‘1’ to that sentence, otherwise
a ‘0’ label is assigned.</p>
        </sec>
        <sec id="sec-2-2-2">
          <title>2.2.2. Legal Document Summarization Task (Task 2(b))</title>
          <p>Participating in the summarization task, we have performed a sentence classification based
extractive summarization of legal judgement documents. We have made use of those models
that have been saved for Task 2(a). We used those models corresponding to each Run during
the testing phase and got the probabilities for each sentence of a document. We took average
probabilities directly as scores for each sentence and ranked the sentences in decreasing order
according to these probability scores. Sentences are then picked up according to desired
)
8
6
7
(
e
s
n
e
D
.t()
4
0
u
o
p
o
r
D
)
8
2
1
(s
e
n
e
D
.t()
4
0
u
o
p
o
r
D
summary length (in number of words) as given by the organizers. For Task 2(b), we have
submitted three runs of our proposed approaches each of which directly utilizes the individual
models saved during the Task 2(a) submissions for Run-1, Run-2 and Run-3.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Experimental results and analysis</title>
      <p>3.1. Setup
For training an MLP model, we have chosen a batch size of 32 and adam optimizer with a
learning rate of 0.001. We run the MLP model for 500 epochs with an early-stopping of 100
epochs. The classification task is evaluated using the standard metrics such as Precision, Recall,
and F-score. In contrast, the summarization task (task 2(b)) is evaluated with the help of ROUGE
metrics, which is a prevalent metric for evaluating automatically generated summaries. All the
experiments have been performed with the help of a Linux based machine with RTX-2070 (8
GB GPU).</p>
      <sec id="sec-3-1">
        <title>3.2. Results</title>
        <p>The results on the test dataset for Task 2 has been shown in Tables 1,2,3 and 4. Our team
nits_legal, has ranked second for the Task 2(a), whereas the Run id 1 for Task 2(b) has achieved
the highest scores in terms of ROUGE-F1 scores for all variants of ROUGE as shown in Table
2. In terms of ROUGE-recall, our team has achieved the highest scores for ROUGE-3 and
ROUGE-4 metrics as shown in Table 3. Whereas, for ROUGE-precision, our team has achieved
the second-highest scores for all variants of ROUGE metrics except ROUGE-1 metric as shown
in Table 4. Please note that the best scores for each measure is in bold in the tables.</p>
        <p>One of the key observations drawn from the results on the test dataset in the case of Task
2(b), is that there is not much diference between the performances of our proposed approaches
across all the ROUGE-recall, precision, and F1-scores. All three metrics are very well-balanced,
which demonstrates that our proposed approach is able to produce very precise summaries,
while having decent recall values. Usually, when the target summary lengths are not known,
top k words are taken during the summary formation step, where k depends on the average
summary lengths in the training set. However, in this sub-task, the desired summary lengths
for each test document were already given by the organizers, and in spite of such constraints,
our model is able to achieve very well-balanced summarization.</p>
        <p>We tried to apply the same MLP based classification approach for Rhetorical role labeling
(Task 1) multi-class classification problem also. However, our approach could not perform as
eficiently for this task as it did for the summarization specific tasks, as shown in Table 5. For this
task, apart from the straightforward MLP based submission, we also submitted a minority class
oversampling based run (Run id 2) which did improve the classification performances slightly.
Interestingly, the best performing team’s prediction accuracy scores for the minority class
(Ruling by Lower Court) is 0 across all the metrics, which is much worse than our results for the
minority class. However, even with this improvement, the results were not very encouraging.</p>
        <p>Chandigard_concordia
Such reduced performance may be attributed to the fact that we have considered the rhetorical
labeling task as a sentence level multi-class classification problem and not as a sequential
sentence classification problem at the document level. In our proposed approach, although
each sentence is represented by its contextual embedding, it still lacks the information on
how these sentences contribute towards the overall document. Further explorations can be
performed where the rhetorical labeling task is considered both at the sentence as well as at the
document levels, in a hierarchical fashion. Please note that the highest overall precision, recall
and F1-score is in bold in the case of Table 5.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusion</title>
      <p>Legal document summarization task becomes very important when there is unstructured and
lengthy documents such as Indian case judgements documents. In this paper, we describe
our methodology for summarization of legal documents, as part of the AILA shared task in
FIRE 2021. We have explored the application of the Legal-Bert model to find efective sentence
embeddings and then fed these embeddings as input to an MLP model for the purpose of
extractive summarization. This kind of MLP based classification approach is found to be very
efective at generating extractive summaries of legal judgement documents with very impressive
ROUGE scores. Our proposed approach is able to obtain the best summarization scores among
all the participants for most of the ROUGE metrics under consideration in Task 2(b). Moreover,
for the sentence summary-worthiness prediction task (Task 2(a)) also, our proposed approach
was able to attain the second position among all the participants. We found that even though
this approach is very efective at summarization of legal documents, it is not as eficient at the
task of rhetorical role labeling (Task 1).</p>
      <p>In order to further improve the performances for both of the tasks, hierarchical representation
of the documents should be taken into consideration, along with the exploration of some of the
recent neural architectures such as Graph Neural Networks (GNN).
[9] A. Farzindar, G. Lapalme, Letsum, an automatic legal text summarizing system, Legal
knowledge and information systems, JURIX (2004) 11–18.
[10] C. Grover, B. Hachey, I. Hughson, C. Korycinski, Automatic summarisation of legal
documents, in: Proceedings of the 9th international conference on Artificial intelligence
and law, 2003, pp. 243–251.
[11] D. Jain, M. D. Borah, A. Biswas, Summarization of legal documents: Where are we now
and the way forward, Computer Science Review 40 (2021) 100388.
[12] P. Bhattacharya, P. Mehta, K. Ghosh, S. Ghosh, A. Pal, A. Bhattacharya, P. Majumder,
Overview of the fire 2020 aila track: Artificial intelligence for legal assistance, in: FIRE
(working notes), 2020.
[13] P. Bhattacharya, S. Paul, K. Ghosh, S. Ghosh, A. Wyner, Identification of rhetorical roles
of sentences in indian legal judgments, in: Proc. International Conference on Legal
Knowledge and Information Systems (JURIX), 2019.
[14] V. Parikh, V. Mathur, P. Mehta, N. Mittal, P. Majumder, Lawsum: A weakly supervised
approach for indian legal document summarization, arXiv preprint arXiv:2110.01188v3
(2021).
[15] V. Parikh, U. Bhattacharya, P. Mehta, A. Bandyopadhyay, P. Bhattacharya, K. Ghosh,
S. Ghosh, A. Pal, A. Bhattacharya, P. Majumder, Overview of the third shared task on
artificial intelligence for legal assistance at fire 2021, in: FIRE (Working Notes), 2021.
[16] I. Chalkidis, M. Fergadiotis, P. Malakasiotis, N. Aletras, I. Androutsopoulos, Legal-bert:
The muppets straight out of law school, arXiv preprint arXiv:2010.02559 (2020).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>P.</given-names>
            <surname>Bhattacharya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Hiware</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rajgaria</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Pochhi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          ,
          <article-title>A comparative study of summarization algorithms applied to legal case judgments</article-title>
          ,
          <source>in: European Conference on Information Retrieval</source>
          , Springer,
          <year>2019</year>
          , pp.
          <fpage>413</fpage>
          -
          <lpage>428</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>V.</given-names>
            <surname>Parikh</surname>
          </string-name>
          , U. Bhattacharya,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mehta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bandyopadhyay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bhattacharya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Pal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bhattacharya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Majumder</surname>
          </string-name>
          ,
          <article-title>Fire 2021 aila track: Artificial intelligence for legal assistance</article-title>
          ,
          <source>in: Proceedings of the 13th Forum for Information Retrieval Evaluation</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>P.</given-names>
            <surname>Bhattacharya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Poddar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Rudra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          ,
          <article-title>Incorporating domain knowledge for extractive summarization of legal case documents</article-title>
          ,
          <source>in: Proceedings of the Eighteenth International Conference on Artificial Intelligence and Law</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>22</fpage>
          -
          <lpage>31</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D.</given-names>
            <surname>Jain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. D.</given-names>
            <surname>Borah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Biswas</surname>
          </string-name>
          ,
          <article-title>Fine-tuning textrank for legal document summarization: A bayesian optimization based approach</article-title>
          ,
          <source>in: Forum for Information Retrieval Evaluation</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>41</fpage>
          -
          <lpage>48</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>D.</given-names>
            <surname>Jain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. D.</given-names>
            <surname>Borah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Biswas</surname>
          </string-name>
          ,
          <article-title>Automatic summarization of legal bills: A comparative analysis of classical extractive approaches</article-title>
          , in: 2021
          <source>International Conference on Computing, Communication, and Intelligent Systems (ICCCIS)</source>
          , IEEE,
          <year>2021</year>
          , pp.
          <fpage>394</fpage>
          -
          <lpage>400</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>D.</given-names>
            <surname>Anand</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Wagh</surname>
          </string-name>
          ,
          <article-title>Efective deep learning approaches for summarization of legal texts</article-title>
          ,
          <source>Journal of King</source>
          Saud University-Computer and Information Sciences (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Farzindar</surname>
          </string-name>
          , G. Lapalme,
          <article-title>Legal text summarization by exploration of the thematic structure and argumentative roles</article-title>
          ,
          <source>in: Text Summarization Branches Out</source>
          ,
          <year>2004</year>
          , pp.
          <fpage>27</fpage>
          -
          <lpage>34</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Saravanan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Ravindran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Raman</surname>
          </string-name>
          ,
          <article-title>Improving legal document summarization using graphical models</article-title>
          ,
          <source>Frontiers in Artificial Intelligence and Applications</source>
          <volume>152</volume>
          (
          <year>2006</year>
          )
          <fpage>51</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>