<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Menghan Wu</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zhengyu Wu</string-name>
          <email>wuzhengyu5@outlook.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Xiangyu Wang</string-name>
          <email>cn.wangxiangyu@foxmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zhongyuan Han</string-name>
          <email>hanzhongyuan@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Foshan University</institution>
          ,
          <addr-line>Foshan</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Heilongjiang Institute of Technology</institution>
          ,
          <addr-line>Harbin</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Methods of Identifying Relevant Prior Cases</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper, we introduce our methods used in the Task1 and Task2 evaluations of AILA2020 (Artificial Intelligence for Legal Assistance). For Task1a, we used BM25 and Cosine Similarity to rank the documents. And the language models with Jelinek-Mercer smoothing method and Dirichlet smoothing method are used for Task1b. For the result we submitted for task2, Logistic Regression model based on TF-IDF feature and BERT is used. Artificial Intelligence for Legal Assistance, retrieval model, classification model In the age of information, studying previous legal cases is very important for judgments. How to quickly and efficiently find similar precedents, and regulations along with the task of document semantic segmentation is particularly important. Therefore, Artificial Intelligence for Legal Assistance 2020 (AILA), treats these tasks [1] as its theme, aims to develop data sets and models to solve these problems.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>2.1.</p>
      <p>2020 Copyright for this paper by its authors.
For the task of Identifying Relevant Prior Cases, we submitted three sets of run files.</p>
      <p>First, we preprocessed the data, used Lucene to index the queried documents, and removed common
punctuation marks and stop words during indexing.</p>
      <p>
        The first submission fs_hit_1_task1a_01 uses Lucene combined with BM25Similarity [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. BM25 uses
IDF (Inverse Document Frequency) to distinguish between common words (relatively unimportant) and
important words. At the same time, it believes that the more frequently a word in a document appears,
the more relevant the document is to the word.
      </p>
      <p>BM25 has two adjustable parameters: k1 and b. The parameter k1 controls the rising speed of the
word frequency result in word frequency saturation. The smaller k1 means the faster the saturation
change, and vice versa. The default value is 1.2. The parameter b controls the role played by the field
length normalization value. The value range of b is (0,1]. When b=0, normalization is disabled, and
when b=1.0, it is completely normalized. The default value is 0.75. And in the experiment, we have
used the default value, but the optimal value of k1 and b still depends on the document collection, and
finally the first set of results submitted k=1.0, b=1.0.</p>
      <p>The second submission fs_hit_1_task1a_02 is used the BM25 model. For the selection of parameters,
we adjusted the parameters by referring to the training results of fs_hit_1_task1a_01. Prior to this, the
first 50% of the query terms were calculated by TF-IDF ((term frequency-inverse document frequency)
is selected as a new query to put into the BM25 retrieval model. The parameters k1=2.0 and b=1.0 were
modified.</p>
      <p>In addition, we also tried the cosine similarity measure. Get the word frequency vector of the
sentence through word segmentation, and calculate the similarity, but the result is unsatisfactory. The
formula for cosine similarity is as follows:
similarity(A, B) =</p>
      <p>∙ 
| | × | |
=</p>
      <p>∑ =1(  ×   )
√∑ =1
(  )2 × √∑ =1
(  )2
(1)
2.2.</p>
    </sec>
    <sec id="sec-2">
      <title>Methods of Identifying Relevant Statutes</title>
      <p>
        For the task of Identifying Relevant Statutes, we submitted three run files. The first two sets of
results use the Language model with Jelinek-Mercer smoothing method [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], and the last set uses the
Language model with Dirichlet smoothing method [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>The document length of task1b is much shorter than task1a, the average length is close to 200 words.
In the statistical word results, there must be a lot of sparseness, it is impossible to include all the
keywords that can be used to query them, and the inclusion of certain keywords, the relevance between
keyword and document is relatively weak. Therefore, in order to improve the accuracy of the query,
language models are used and smoothed.</p>
      <p>Before the calculation, we processed the documents as follows: Some common punctuation marks
were removed when indexing, and stop words were also removed. In fs_hit_1_task1b_01 and
fs_hit_1_task1b_02, we used Jelinek Mercer smoothing with different parameter λ. The value of λ is
between (0,1). The choice of λ parameter is very important for the model. The best value depends on
the specific collection and query. The closer λ is to 1, the greater the smoothness effect. During our
experiment, when λ=1-10-5(fs_hit_1_task1b_01) results the best result, we also submitted the results
when λ=1-10-6(fs_hit_1_task1b_02).</p>
      <p>The last set of results submitted uses the default value of Dirichlet smoothing.
2.3.</p>
    </sec>
    <sec id="sec-3">
      <title>Methods of Rhetorical Role Labeling for Legal Judgements</title>
      <p>For the task of Rhetorical Role Labeling for Legal Judgements, the file named fs_hit_1_task2_01
uses the Logistic Regression with the feature of TF-IDF which provided by Scikit-learn.</p>
      <p>
        The rest of two submit, fs_hit_1_task2_02 and fs_hit_1_task2_03, are generated by Bert [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]
with different random seeds. In this method, we did not do any processing on the data. When the test
set is not released, a random 20% of train set is selected as the test set. On the pre-training model, we
chose L-12_H-768_A-12 (BERT-Base). We finetuned it with the train set. When the train epoch is
greater than 5, the model is overfitting.
      </p>
    </sec>
    <sec id="sec-4">
      <title>3. Experimental Setting</title>
    </sec>
    <sec id="sec-5">
      <title>3.1. Parameter Selection</title>
      <p>For Task1A, the k1 and b we used in BM25Similarity are listed as follow.</p>
      <p>Table 1
The result of different parameters k1 and b in BM25Similarity</p>
      <p>Parameters
K1=1.2, b=0.75
K1=1.0, b=0.75
K1=1.0, b=1.0</p>
      <p>MAP
0.1082
0.1108
0.1220</p>
      <p>For Task1B, we selected Jelinek-Mercer smoothing method with different lambda.</p>
      <p>After knowing the experimental results, we used some other parameters for Language Model with
Dirichlet smoothing. The specific parameters and results are shown in the figure:</p>
      <p>Evaluation result of parameter μ
0.5
0.45
0.4
0.35
0.3
0.25
0.2
0.15
0.1
0.05
0
0.1f</p>
      <p>0.6f
MAP</p>
      <p>1000f</p>
      <p>BPREF</p>
      <p>2000f
Recip-rank
3000f</p>
      <p>For Language Model with Dirichlet smoothing, the default parameter mu of this method is 2000 and
it is also the parameter we use to submit the results. We tried experiments when the parameter was close
to 0, and the result was slightly higher than the instantiation similarity using the default mu value of
2000.
3.2.</p>
    </sec>
    <sec id="sec-6">
      <title>Experimental Results</title>
      <p>From the experimental results that BM25Similarity has achieved good result, which wins sixth
place on MAP and BPREF. When using BM25+TF-IDF the result is lower than first run file there is
not much difference in results.</p>
      <p>Among the evaluation metrics of the fs_hit_1_task1a_02 file, the metrics p@10 of 0.09 won the
second place, and the metrics recip_rank of 0.2078 won the third place.</p>
      <p>From the experimental results, the methods in task1b are not achieved good performance, and the result of
using Language Model with Dirichlet smoothing is slightly higher.</p>
    </sec>
    <sec id="sec-7">
      <title>4. Conclusions</title>
      <p>The paper presents the methods we used in the FIRE2020 AILA. We propose text retrieval methods
(BM25 and Language Model) for Task1. Bert model and Logistic Regression method are employed
for tackling the Task2. Our ranks in three tasks are in a modest position among all the participants in
the score board. In future we will continue study the retrieval model and classification model that we
used now, and find other methods or models like Bert-large model.</p>
    </sec>
    <sec id="sec-8">
      <title>5. Acknowledgements</title>
    </sec>
    <sec id="sec-9">
      <title>6. References</title>
      <p>This work is supported by National Social Science Fund of China (No.18BYY125)</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Bhattacharya</surname>
          </string-name>
          ,
          <article-title>Paheli and Mehta, Parth and Ghosh, Kripabandhu and Ghosh, Saptarshi and Pal, Arindam and Bhattacharya, Arnab and Majumder, Prasenjit, Overview of the FIRE 2020 AILA track: Artificial Intelligence for Legal Assistance</article-title>
          .
          <source>Proceedings of FIRE 2020 - Forum for Information Retrieval Evaluation</source>
          . Hyderabad, India, December,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Robertson</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Steve</surname>
            <given-names>Walker</given-names>
          </string-name>
          , and
          <string-name>
            <given-names>H.</given-names>
            <surname>BM</surname>
          </string-name>
          .
          <article-title>"GM Okapi at trec-3."</article-title>
          <source>Proceedings of the Third Text REtrieval Conference (TREC</source>
          <year>1994</year>
          ).
          <year>1994</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Zhai</surname>
            , Chengxiang,
            <given-names>and John</given-names>
          </string-name>
          <string-name>
            <surname>Lafferty</surname>
          </string-name>
          .
          <article-title>"A study of smoothing methods for language models applied to ad hoc information retrieval." ACM SIGIR Forum</article-title>
          . Vol.
          <volume>51</volume>
          . No. 2. New York, NY, USA: ACM,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Devlin</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            <given-names>M W</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            <given-names>K</given-names>
          </string-name>
          , et al.
          <article-title>Bert: Pre-training of deep bidirectional transformers for language understanding</article-title>
          .
          <source>arXiv preprint arXiv:1810.04805</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>