<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>An Indian Court Decision Annotated Corpus and Knowledge Graph</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Pariskhit Kamat</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shubham Kalson</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Suraj S</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pooja Harde</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nandana Mihindukulasooriya</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sarika Jain</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>IBM Research</institution>
          ,
          <addr-line>Dublin</addr-line>
          ,
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>National Institute of Technology Kurukshetra</institution>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <fpage>79</fpage>
      <lpage>90</lpage>
      <abstract>
        <p>Document collection is increasing enormously in the legal domain, which requires automatic steps to analyze the data and curate the information from the same. Many challenges are being faced by the legal stakeholders to extract the information from the lengthy and unstructured court judgment documents relating to the main concepts, topics, and named entities in the documents. It has become an essential task in the current scenario to automate the information extraction process and store the documents in a properly structured format along with the diferent named legal entities for ease in the information extraction. In this paper, we introduce an annotated Indian Court Decision Document Corpus consisting of 10 coarse-grained classes and 30 fine-grained classes as a benchmark data set for constructing the knowledge graph. We also construct the Indian Court Case Documents' knowledge graph by utilizing a rule-based approach for Named Entity Recognition (NER) and Relation Extraction (RE). The results are evaluated against the proposed benchmark based on precision, recall, and F1 score and also qualitatively using SPARQL queries. The proposed approach gives a good F1 measure, though, further work is required to improve the recall.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Entity Extraction</kwd>
        <kwd>Relation Extraction</kwd>
        <kwd>Knowledge Graph</kwd>
        <kwd>Legal Domain</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        India’s vast and complex legal system routinely creates and processes large volumes of legal
documents. The limited knowledge of the general public in the field of law along with the
complex language and legal terminologies makes it dificult for them to understand the ideas
and information conveyed by legal documents. Even legal professionals who compile court
judgment documents [
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1, 2, 3</xref>
        ] find it cumbersome to go through long documents, understand
and form opinions [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The existing portals retrieve the court decision documents either in PDF
or unstructured format using basic keywords. To overcome the constraints of keyword-based
search, semantic web and semantic search can be utilized. The semantic web[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] is a network in
which the information is stored as knowledge graphs rather than a collection of documents
linked using hyperlinks. Through semantic search, we will be able to focus on the intent and
contextual meaning of the used keywords rather than completely relying on the keywords for
information retrieval.
      </p>
      <p>
        One of the major objectives of this work is to create an annotated data set for the Indian
Court Decision Documents so that the machines can extract maximum information from the
case documents and represent them in a uniform structured format with the help of Knowledge
Graphs [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. The work also explores the rule-based approach to extract and annotate the legal
entities to construct a Knowledge Graph. The resultant Knowledge base will be a prodigious
collection of interlinked case documents which can benefit the legal stakeholders. The primary
step towards constructing the Knowledge Graph is the Information Extraction (IE) in which the
various legal entities will be identified with the help of Named Entity Recognition (NER) and
the relations among these entities will be extracted through Relation Extraction (RE).
      </p>
      <p>
        The existing legal ontologies like JuDo [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], LKIF [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], LRI-Core [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and so on are available for
creating the legal domain knowledge bases and to serve the legal reasoning in the legal domain,
these ontologies address legal entities and common sense entities at the very abstract level. Thus
the already available ontologies do not fulfill the purpose of extracting the information from
the court judgments like the Court, Jurisdiction under which the court can hear the case, type
of evidence presented in the court hearing of the case, the origin of the case, case background,
and so on. Other than ontologies there is a widely used XMLSchema defined as Akoma Ntoso
1 used for legal document structuring. The limitations with XML is they lack integration of
heterogeneous data on the web and also does not provide the inference to the data which RDFS
and OWL provides. Thus we move forward with the ontology which can provide the inference
to the legal data for legal reasoning and semantics. NyOn[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] (Nyaya Ontology) is designed
to primarily extract the relevant information from the court judgment documents and derive
the relationship between this information making it available for diferent uses cases like legal
reasoning, question-answering, legal analysis and so on. In paper [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], the JCO ontology for
Indian Court Judgments is created but it lacks the reusability and publishing of the ontology.
NyOn [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] Ontology has been chosen to be served as the metadata for Information Extraction.
The second step is to store the extracted data in a graph database in form of triples to query the
database for information retrieval.
      </p>
      <p>
        The prominent contributions put forth by this paper are as follows:
• Creation of Indian Court Decision Documents Annotated Corpus guided by NyOn [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
• Identify the legal entities and relations from the court decision documents using the
      </p>
      <p>Rule-Based Approach.
• Construction of Knowledge Graph from the extracted data.</p>
      <p>• Evaluation based upon quantitative and qualitative approaches.</p>
      <p>The remaining portion of the paper is compiled as follows. Section 2 examines the related
works carried out in data set construction and the rule-based approach for information extraction.
Section 3 focuses on Dataset Construction which includes the methodologies followed in creating
and validating the data set. Section 4 discusses the construction of a Knowledge Graph using
NER and RE through a rule-based approach. Section 5 sheds light on the evaluation results and
Section 6 interprets the conclusions drawn from this work and provides insights on areas of
improvement and ideas for future work.
Though the legal domain has been benefiting from the semantic web and ontology technologies
in the past years, dedicated work on the Indian legal domain is yet to be developed. We have
analyzed a few works done on ontology-based information retrieval and the same is discussed
below.</p>
      <p>
        Aboaoga et al. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] and R.Alfred et al. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]have proposed a Rule-Based approach for
recognizing the named entity type (person names) for Arabic and Malay articles respectively.
      </p>
      <p>
        Judith Jeyafreeda Andrew et al. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] developed a system that helps journalists to recognize
legal entities like names of people, organizations, roles, and functions. The author used 2
methods; first, Conditional Random Fields as a statistical method and another is the rule-based
technique for generating language-specific regular expressions.
      </p>
      <p>
        P. H. Luz de Araujo et al. [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] presents named entity recognition dataset for Brazilian legal
documents. Along with the open domain tags such as persons, locations, time entities, and
organizations, the dataset contains law and legal cases entities specific tags.
      </p>
      <p>
        Based on the annotated corpus Prathamesh et al. [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] created baseline models for
automatically predicting rhetorical functions in court documents. They also demonstrated the use of
rhetorical roles to increase performance on summarization and legal judgement prediction tests.
      </p>
      <p>
        Vladislav Korablinov and Pavel Braslavski [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] provided the first Russian knowledge base
question answering (KBQA) dataset known as RuBQ. The high-quality dataset included 1,500
Russian questions of varied dificulty, their English machine translations, SPARQL queries
to Wikidata, reference responses, and a Wikidata sample of triples comprising entities with
Russian labels.
      </p>
      <p>
        Elena Leitner et al. [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] developed data set for Named Entity Recognition in German legal
documents. They manually annotated approx 67,000 sentences with 2 million tokens.
      </p>
      <p>
        Riaz et al. [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] have discussed the diferences between Hindi and Urdu NER and concluded
that the NER computational models for Hindi cannot be applied to Urdu. They have also
presented a NER algorithm rule-based Urdu that outperforms the models that use statistical
learning.
      </p>
      <p>
        Thomas et al. [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] have investigated natural language texts of domains lacking generic named
entities labelled domain data sets. They created a hybrid NER system that combines rule-based
deep learning with clustering-based techniques to enable the extraction of generic entities.
      </p>
      <p>
        In a paper published by Crotti Junior et al. [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] the author discussed the dificulties they
had and the progress they achieved when creating a knowledge graph-based search engine for
Wolters Kluwer Germany’s collection of German court case data.
      </p>
      <p>Filtz et al. [22] highlights the data representation and search problems in the legal domain
data. They suggested a method for representing Austrian legal information (legal standards and
court rulings), and they demonstrated how to use such information to create a legal knowledge
graph.</p>
      <p>
        Breukers et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] presented two legal core ontologies for law. The first was the outcome of
Valente’s Ph.D. thesis [23] called FOLaw. This core ontology, known as LRI-Core, is made up of
ifve main sections (or "worlds"): roles, occurrences, physical, mental, and abstract classes.
      </p>
      <p>
        Jain et al. [24] present the similar approach for extracting the NER and RE from the legal
documents. For identifying the named entities, ontology presented in [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] is being used by
the author in the paper. Author used the JAPE rules for extracting the legal entities and legal
documents are processed using the GATE tool which are later exported in the inline XML for
RE.
      </p>
      <p>
        A generic architecture for legal knowledge systems, as described by Hoekstra et al. in their
publication [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], includes a legal core ontology (LKIF) that enables knowledge exchange between
current legal knowledge systems. LKIF had two primary roles 1) translation of legal knowledge
bases expressed in various representation formats and formalisms and 2) knowledge
representation formalism that is a component of a wider architecture for creating legal knowledge
systems.
      </p>
      <p>
        Ceci et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] introduces an OWL2 ontology library of legal knowledge that relies on the
metadata contained in judicial documents known as JudO. The ontology addresses
meaningful legal semantics at the same time retaining a strong connection to source documents (i.e.
fragments of legal texts).
      </p>
      <p>
        Thomas et al. [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] presents a legal case ontology named Judicial Case Ontology (JCO) that
incorporates the concepts and relations existent in the legal domain cases including the related
terms from a set of real-life judicial decisions. The ontology supports the extraction of taxonomic
and non-taxonomic domain-specific relationships from e-judgments.
      </p>
    </sec>
    <sec id="sec-2">
      <title>3. Dataset Construction</title>
      <sec id="sec-2-1">
        <title>3.1. Dataset Description</title>
        <p>For creating an indian legal corpus, the legal documents are collected from the ’Indian Kanoon’2
website, an online search engine provided for Indian legal documents. The Python script used
for scraping the dataset is given in the Github repository. For ease of processing, the collected
PDF documents were converted to text format. The pre-processing such as sentence splitting,
tokenization, and POS tags annotation using SPACY3 are performed on these text files data. To
restrict the scope of the data we made use of the list of the competency questions such as:
1. List all the cases of month X.
2. List all the cases filed in the year X.
3. What are the total number of cases filed under case type ’criminal’?
4. List all the cases with X is a judge.
5. What is the count of cases with ’Appeal is accepted’ as the judgment?
6. What is the date of judgement for the case X.
7. List all the cases filed under ’Appellant Jurisdiction’.
8. Petitioner Name with CASE NO.: X.</p>
        <p>9. List all the cases involving X as one of the party.</p>
        <p>10. Count of appeals ’rejected’ by the judge X.</p>
        <p>
          To address the scope of the data, the required legal terms were taken from NyOn [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] Ontology,
a modular ontology to describe court judgments, and has been published adhering to the
2https://indiankanoon.org/
3https://spacy.io/
Semantic Web best practices and FAIR principles. Two semantic classes are defined to support
domain-specific tags. That is, a coarse-grained class and a fine-grained class consisting of 10
and 30 attributes respectively. Coarse is a more general legal semantic class that includes Court,
Party, CourtDecision, Document, Jurisdiction, Location, CaseType, Author, CourtOficial, and
DateOfJudgment classes.
        </p>
        <p>The created dataset is the gold standard dataset with manually identified Named Legal Entities
from the tokens and tagged with domain-specific tags using the CoNLL-2003 format. In
CoNLL2003 data files, there contains four columns separated by a single space. Each word of the
sentence is added to a single line and each sentence is followed by an empty line. A word is an
initial item on each line, followed by a part-of-speech (POS) tag, a syntactic chunk tag, and a
named entity tag. The dataset is encoded in three diferent encodings in CoNLL-2003 format:
BILOU ((B-Beginning, I-Internal, L-Last, O-outside, U-Unit), IOB (I-Inside, O-Outside, B-Begin)
and IOBES (I-Inside, O-Outside, B-Begin, E-End, S-Single). It is to be noted that the syntactic
tags are not considered for the preparation of the data set. While a named entity is a pronoun
or noun, which usually refers to the name of a person, place, etc., legal entities are basically
the legal terms from the legal documents that might be names of parties involved, document
numbers, bench, the title of the legal document, etc. A total of four annotators have participated
in the construction of the corpus. The manually developed dataset consists of a total of 50 legal
documents with 80,733 rows of tokenized words and their corresponding annotated legal tags.
Table 1 depict the count of the particular legal tag in the whole dataset (represented by #) for
the corresponding coarse-grained and fine-grained classes.</p>
        <p>Some example attributes of the coarse-grained class and fine-grained class are as follows:
Party The coarse-grained class Party PT contains fine-grained classes Respondent RES,
Appellant APLT, Plaintif PLNF, Petitioner PETR.</p>
        <sec id="sec-2-1-1">
          <title>Ex. PETITIONER PT : B. SHANKRANAND PETR Vs.</title>
          <p>RESPONDENT PT : COMMON CAUSE &amp; ORS. RES DATE OF JUDGMENT:
11/03/1996 BENCH: RAMASWAMY, K. BENCH: RAMASWAMY, K. G.B.
PATTANAIK (J).</p>
        </sec>
        <sec id="sec-2-1-2">
          <title>CourtOficial The coarse-grained class CourtOficial</title>
          <p>Investigator INVG, Solicator SOL and Judge JD.</p>
        </sec>
        <sec id="sec-2-1-3">
          <title>CRTOF contains fine-grained classes</title>
        </sec>
        <sec id="sec-2-1-4">
          <title>Ex. DATE OF JUDGMENT: 29/04/1991 BENCH: KANIA,</title>
          <p>KANIA, M. H. JD VERMA, JAGDISH SARAN. JD</p>
        </sec>
        <sec id="sec-2-1-5">
          <title>M. H. BENCH:</title>
          <p>(J) CRTOF</p>
        </sec>
        <sec id="sec-2-1-6">
          <title>RAMASWAMI, V. JD</title>
          <p>(J) CRTOF</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>3.2. DataSet Validation and Publication</title>
        <p>The constructed dataset is validated by an expert Mr. Vaibhav Vats, Advocate, Punjab and
Haryana High Court, Chandigarh. The expert reported an annotation accuracy of 92%. It was
observed that the tag SPECIAL LEAVE PETITION was wrongly annotated as PETITION. The
expert noticed still more entities that should have been included and the list is not limited to:
Party (Complainant, Defendant, Prosecution, and Accused for the criminal cases); Jurisdiction
(Regional, Pecuniary, Writ, Special Leave Petition); Case Type (Matrimonial, Consumer, IPR);
and more. The documents list can be very exhaustive if we want. The annotation of all these
will be considered as future scope work.</p>
        <p>The data set is published using FigShare4 with CC by 4.0 licence with the DOI:https://doi.
org/10.6084/m9.figshare.19719088.v4</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>4. Knowledge Graph Construction</title>
      <p>Knowledge graphs are network representations of real-world entities consisting of nodes, edges,
and labels. Representing a copious collection of unstructured data using knowledge graphs will
ease the process of abridging the facts and information from extensive documents. Though
there are multiple approaches for constructing knowledge graphs from unstructured data, we
have used the rule-based approach as it can closely simulate human intelligence and ofers
the flexibility to incorporate cognitive processes into machines. For extracting named entities
and their relations, the rule-based approach uses regular expressions to identify various lexical
patterns and trigger words.</p>
      <sec id="sec-3-1">
        <title>4.1. Named Entity Recognition</title>
        <p>
          The entity extraction process is carried out by referring to the NyOn [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] Ontology and a
total of 10 named legal entities, namely Party, Court, Date of Judgment, Court Oficial, Author,
Location, Case Type, Court Decision, Jurisdiction and Documents were identified as given in Table1.
The scraped data from Indian Kanoon2 is passed through Python rules which contain regular
expressions that trigger target words. The major predicaments faced while coding the Python
rules are the amorphous nature of the legal documents which made it dificult to code regular
expressions that could fit the entire corpus. Despite the irregular structure and format, we were
able to come up with reasonable rules that fit a decent cut of the corpus. Each case will be mapped
to a central entity "CASE" in the knowledge graph. An entity "CASE_NAME" is formed with
the help of three other identified entities, namely "Petitioner"/"Appellant", "Respondent" and
"Date Of Judgment". And if none of the above three entities are identified, the "CASE_NAME"
will be assigned with "CASE_NO" or "APPEAL_NO" respectively, subject to their identification
in the document.
        </p>
        <p>The output from the NER phase is stored in a single text file with the extracted token and its
corresponding identified entity to pass to the Relation Extraction Phase for obtaining relations
between the entities. The code and the output files are provided in the Github repository.</p>
      </sec>
      <sec id="sec-3-2">
        <title>4.2. Relation Extraction</title>
        <p>
          The relation between the entities extracted in the NER phase are identified in this relation
extraction phase using a small python script. The NyOn [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] is referred for identifying the
various relationships between the extracted entities obtained in NER Phase. A total of 14
relations, namely hasCaseName, hasParty, hasDate, hasYear, hasMonth, hasAppealNo,
hasCaseType, hasAuthor, hasCourtOficial, hasJurisdiction, hasCourt, hasLocation and hasCourtDecision
are identified. For uniquely identifying each case, a new entity "CASE_ID" is generated by
concatenating the year and month of judgment, an abbreviation of our system name(KS for
Kanoon Sarathi) along with the serial number of the case in the current month, and the court
abbreviation to which it belongs to(SC for Supreme Court, HC for High Court and DC for
District Court). A new relation hasCaseId is also derived from the new entity "CASE_ID".
        </p>
        <p>Since the output of the NER stage does not contain sentences, we use ’if’ statements to
annotate the relationships between the extracted entities. A sample output containing a few
entities and relations from both the NER phase and RE phase corresponding to the case KEWAL
KRISHAN VS. STATE OF PUNJAB on 06/03/1962 is given in Table 3 The code used for relation
extraction along with the output file containing identified relations(predicates) along with the
corresponding entities(subject and object) is provided in the Github repository.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5. Evaluation</title>
      <sec id="sec-4-1">
        <title>5.1. Quantitative Evaluation</title>
        <p>For quantitative evaluation of the rule-based Named Entity Recognition (NER) and Relation
Extraction (RE) with respect to our data set, we use the metrics F1-Score, Recall(for measuring
the reliability of the model in correctly identifying entity tags out of actually existing entity
tags), and Precision (for measuring the reliability of the model in correctly identifying entity
tags out of total identified entity tags). Table 4 shown below represents the Evaluation metrics
of Named Entity Recognition and Table 5 depicts the Evaluation Metrics for Relation Extraction.</p>
      </sec>
      <sec id="sec-4-2">
        <title>5.2. Qualitative Evaluation</title>
        <p>For the Qualitative evaluation of the Knowledge Graph, 10 competency questions were
formulated and the knowledge graph is queried using SPARQL to retrieve the relevant information.
The sample queries based on competency questions are performed on the knowledge graph
that are shown in figure 1. The list of the competency questions with the corresponding queries
and their outputs can be found in the Github repository.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>6. Conclusion and Future Scope</title>
      <p>
        In this paper, we have presented a dataset for Knowledge Base construction in the Indian Legal
domain. We have also discussed the modus operandi for constructing Knowledge Graph from
the Indian Court Decision corpus through a rule-based approach. NyOn [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] Ontology was
used as a reference for entity extraction and relation extraction with the help of which the
triples were annotated. After triple generation, the RDF conversion process is followed using
python script, and the same is stored in Apache Jena Fuseki.
      </p>
      <p>The results derived were arguably good and comparatively better than the existing works
using a rule-based approach, albeit we have identified numerous shortcomings which can be
improved. In terms of future work, we plan to extend the dataset in two dimensions; one,
add more documents to increase the size of the dataset which will provide a good sample for
approaching the machine learning algorithms for extracting named entities, and second, to add
more entities for annotating legal norms, solicitors, evidence and so on.
6.0.1. Acknowledgements
This work is supported by the IHUB-ANUBHUTI-IIITD FOUNDATION set up under the
NMICPS scheme of the Department of Science and Technology, India. We thank Mr. Vaibhav Vats,
Advocate, Punjab and Haryana High Court, Chandigarh for providing his valuable reviews of
the data set.</p>
      <p>Supplemental Material Availability: The code and the data set are available on the GitHub
Repository with link: https://github.com/semintelligence/KING.
(a) List all the cases from the year 1996.
(b) Count of all the criminal cases.
(c) List all the cases with Union of India as the
party.
(d) List all the appeals rejected by the judge V.</p>
      <p>BOSE
[22] E. Filtz, Building and processing a knowledge-graph for legal data, 2017. doi:10.1007/
978-3-319-58451-5_13.
[23] A. Valente, Legal knowledge engineering: A modelling approach, volume 30, Penn State</p>
      <p>Press, 1995.
[24] S. Jain, P. Harde, N. Mihindukulasooriya, S. Ghosh, A. Dubey, A. Bisht, Constructing
a knowledge graph from indian legal domain corpus, in: International Workshop On
Knowledge Graph Generation From Text (Text2kg) Co-located with the Extended Semantic
Web Conference (ESWC 2022), CEUR Workshop Proceedings, volume 3184, 2022, pp. 80–93.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Elnaggar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Gebendorfer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Glaser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Matthes</surname>
          </string-name>
          <article-title>, Multi-task deep learning for legal document translation, summarization and multi-label classification</article-title>
          ,
          <source>in: Proceedings of the 2018 Artificial Intelligence and Cloud Computing Conference</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>9</fpage>
          -
          <lpage>15</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>O</given-names>
            <surname>.-M. Sulea</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zampieri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Malmasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Vela</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. P.</given-names>
            <surname>Dinu</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Van Genabith</surname>
          </string-name>
          ,
          <article-title>Exploring the use of text classification in the legal domain</article-title>
          ,
          <source>arXiv preprint arXiv:1710.09306</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>N.</given-names>
            <surname>Ramrakhiyani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pawar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. K.</given-names>
            <surname>Palshikar</surname>
          </string-name>
          ,
          <article-title>A system for classification of propositions of the indian supreme court judgements</article-title>
          ,
          <source>in: Post-Proceedings of the 4th and 5th Workshops of the Forum for Information Retrieval Evaluation</source>
          ,
          <year>2013</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>4</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>V.</given-names>
            <surname>Malik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Sanjay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. K.</given-names>
            <surname>Nigam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Guha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bhattacharya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Modi</surname>
          </string-name>
          ,
          <article-title>Ildc for cjpe: Indian legal documents corpus for court judgment prediction and explanation</article-title>
          , ????
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Sharma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Jain</surname>
          </string-name>
          ,
          <article-title>Comprehensive study of semantic annotation: Variant and praxis, Advances in Computational Intelligence, its Concepts Applications (ACI</article-title>
          <year>2021</year>
          )
          <volume>2823</volume>
          (
          <year>2021</year>
          )
          <fpage>102</fpage>
          -
          <lpage>116</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Dragoni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Villata</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Rizzi</surname>
          </string-name>
          , G. Governatori,
          <source>Combining Natural Language Processing Approaches for Rule Extraction from Legal Documents: AICOL International Workshops</source>
          <year>2015</year>
          -2017:
          <article-title>AICOL-VI@JURIX 2015</article-title>
          ,
          <article-title>AICOL-VII@EKAW 2016</article-title>
          ,
          <article-title>AICOL-VIII@JURIX 2016</article-title>
          ,
          <article-title>AICOL-IX@ICAIL 2017, and AICOL-X@JURIX 2017, Revised Selected Papers</article-title>
          , volume
          <volume>10791</volume>
          ,
          <year>2018</year>
          , pp.
          <fpage>287</fpage>
          -
          <lpage>300</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -00178-0_
          <fpage>19</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Ceci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gangemi</surname>
          </string-name>
          ,
          <article-title>An owl ontology library representing judicial interpretations</article-title>
          ,
          <source>Semantic Web</source>
          <volume>7</volume>
          (
          <year>2016</year>
          )
          <fpage>229</fpage>
          -
          <lpage>253</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>R.</given-names>
            <surname>Hoekstra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Breuker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. Di</given-names>
            <surname>Bello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Boer</surname>
          </string-name>
          , et al.,
          <article-title>The lkif core ontology of basic legal concepts</article-title>
          .
          <source>, LOAIT</source>
          <volume>321</volume>
          (
          <year>2007</year>
          )
          <fpage>43</fpage>
          -
          <lpage>63</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>Breukers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Hoekstra</surname>
          </string-name>
          ,
          <article-title>Epistemology and ontology in core ontologies: Folaw and lri-core, two</article-title>
          ,
          <source>in: Proceedings of EKAW Workshop on Core ontologies [Internet]. Northamptonshire</source>
          , UK:
          <string-name>
            <surname>Sun SITE Central Europe</surname>
          </string-name>
          , Citeseer,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S.</given-names>
            <surname>Jain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Harde</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Mihindukulasooriya</surname>
          </string-name>
          ,
          <article-title>Nyon - a multilingual modular legal ontology for representing court judgments</article-title>
          ,
          <source>in: International Semantic Intelligence Conference (ISIC 2022) held during May 17-19</source>
          ,
          <year>2022</year>
          , Georgia Southern University (Armstrong Campus), Savannah, United States,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>A.</given-names>
            <surname>Thomas</surname>
          </string-name>
          ,
          <string-name>
            <surname>S. S.,</surname>
          </string-name>
          <article-title>A legal case ontology for extracting domain-specific entity-relationships from e-judgments</article-title>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M.</given-names>
            <surname>Aboaoga</surname>
          </string-name>
          ,
          <article-title>Arabic person names recognition by using a rule based approach</article-title>
          ,
          <source>Journal of Computer Science</source>
          <volume>9</volume>
          (
          <year>2013</year>
          )
          <fpage>922</fpage>
          -
          <lpage>927</lpage>
          . doi:
          <volume>10</volume>
          .3844/jcssp.
          <year>2013</year>
          .
          <volume>922</volume>
          .927.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>R.</given-names>
            <surname>Alfred</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. M. C.</given-names>
            <surname>Leong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. K.</given-names>
            <surname>On</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Anthony</surname>
          </string-name>
          ,
          <article-title>Malay named entity recognition based on rule-based approach</article-title>
          ,
          <source>International Journal of Machine Learning and Computing</source>
          <volume>4</volume>
          (
          <year>2014</year>
          )
          <fpage>300</fpage>
          -
          <lpage>306</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>J. J.</given-names>
            <surname>Andrew</surname>
          </string-name>
          ,
          <article-title>Automatic extraction of entities and relation from legal documents</article-title>
          ,
          <source>in: Proceedings of the Seventh Named Entities Workshop</source>
          , Association for Computational Linguistics, Melbourne, Australia,
          <year>2018</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          . URL: https://aclanthology.org/W18-2401. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>W18</fpage>
          -2401.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>P. H. Luz de Araujo</surname>
            , T. de Campos,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Oliveira</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Staufer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Couto</surname>
          </string-name>
          , P. De Souza Bermejo,
          <string-name>
            <surname>LeNER-Br</surname>
          </string-name>
          :
          <article-title>A Dataset for Named Entity Recognition in Brazilian Legal Text: 13th International Conference</article-title>
          , PROPOR 2018, Canela, Brazil,
          <source>September 24-26</source>
          ,
          <year>2018</year>
          , Proceedings,
          <year>2018</year>
          , pp.
          <fpage>313</fpage>
          -
          <lpage>323</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>319</fpage>
          -99722-3_
          <fpage>32</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>P.</given-names>
            <surname>Kalamkar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Tiwari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Agarwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Karn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Raghavan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Modi</surname>
          </string-name>
          ,
          <article-title>Corpus for automatic structuring of legal documents</article-title>
          ,
          <source>CoRR abs/2201</source>
          .13125 (
          <year>2022</year>
          ). URL: https: //arxiv.org/abs/2201.13125. arXiv:
          <volume>2201</volume>
          .
          <fpage>13125</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>V.</given-names>
            <surname>Korablinov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Braslavski</surname>
          </string-name>
          ,
          <article-title>Rubq: A russian dataset for question answering over wikidata</article-title>
          , CoRR abs/
          <year>2005</year>
          .10659 (
          <year>2020</year>
          ). URL: https://arxiv.org/abs/
          <year>2005</year>
          .10659. arXiv:
          <year>2005</year>
          .10659.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>E.</given-names>
            <surname>Leitner</surname>
          </string-name>
          , G. Rehm,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Schneider</surname>
          </string-name>
          ,
          <article-title>A dataset of german legal documents for named entity recognition</article-title>
          , CoRR abs/
          <year>2003</year>
          .13016 (
          <year>2020</year>
          ). URL: https://arxiv.org/abs/
          <year>2003</year>
          .13016. arXiv:
          <year>2003</year>
          .13016.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>K.</given-names>
            <surname>Riaz</surname>
          </string-name>
          ,
          <article-title>Rule-based named entity recognition in urdu</article-title>
          ,
          <source>in: Proceedings of the 2010 Named Entities Workshop</source>
          , NEWS '10,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computational Linguistics, USA,
          <year>2010</year>
          , p.
          <fpage>126</fpage>
          -
          <lpage>135</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>A.</given-names>
            <surname>Thomas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sangeetha</surname>
          </string-name>
          ,
          <article-title>An innovative hybrid approach for extracting named entities from unstructured text data</article-title>
          ,
          <source>Computational Intelligence</source>
          <volume>35</volume>
          (
          <year>2019</year>
          )
          <fpage>799</fpage>
          -
          <lpage>826</lpage>
          . URL: https://onlinelibrary.wiley.com/doi/ abs/10.1111/coin.12214. doi:https://doi.org/10.1111/coin.12214. arXiv:https://onlinelibrary.wiley.com/doi/pdf/10.1111/coin.12214.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>A.</given-names>
            <surname>Crotti Junior</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Orlandi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Graux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hossari</surname>
          </string-name>
          ,
          <string-name>
            <surname>D. O'Sullivan</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Hartz</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Dirschl</surname>
          </string-name>
          ,
          <article-title>Knowledge graph-based legal search over german court cases</article-title>
          , in: European Semantic Web Conference, Springer,
          <year>2020</year>
          , pp.
          <fpage>293</fpage>
          -
          <lpage>297</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>