<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Study on Lemma vs Stem for Legal Information Retrieval Using R Tidyverse. IMS UniPD @ AILA 2020 Task 1</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Giorgio Maria Di Nunzio</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Information Engineering, University of Padova</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Mathematics, University of Padova</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <volume>2517</volume>
      <fpage>16</fpage>
      <lpage>20</lpage>
      <abstract>
        <p>In this paper, we describe the results of the participation of the Information Management Systems (IMS) group at AILA 2020 Task 1, precedents and statutes retrieval. In particular, we participated in both subtasks: precedents retrieval (task a) and statutes retrieval (task b). The goal of our work was to compare and evaluate the eficacy of a simple reproducible approach based on the use of either lemmas or stems with a tf-idf vector space model and a plain BM25 model. The results vary significantly from one subtask/evaluation measure to another. For the subtask of statutes retrieval, our approach performed well, being second only to a participant that used BERT to represent documents.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Legal IR</kwd>
        <kwd>BM25</kwd>
        <kwd>TF-IDF</kwd>
        <kwd>Text Pipelines</kwd>
        <kwd>R Tidyverse</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>The remainder of the paper will introduce the methodology and a brief summary of the
experimental settings that we used in order to create the oficial runs that we submitted for this
task.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Method</title>
      <p>In this section, we summarize the pipeline for text pre-processing which has been developed
in the last years [4]. In general, our method follows the principles described by [5] where the
idea is to mine textual information from large text collections in an eficient and efective by
means of organized workflows named pipelines. Pipelines are an efective way to manage the
sequential process of text analysis by splitting the source code into steps, where the output of
one step is the input for the subsequent step. The R programming language has an interesting
set of packages that follow this idea, named tidyverse,2 that we will use in our experiments.</p>
      <p>Apart from being a tidy way of organizing software, an important advantage in working with
pipelines is that this practice promotes shareability and reproducibility in research workflows
which is one of the main pillars in the European Open Science Cloud (EOSC).3</p>
      <sec id="sec-2-1">
        <title>2.1. Pipeline for Data Cleaning</title>
        <p>In order to produce the clean dataset, we followed the same pipeline for data ingestion and
preparation for all the experiments:
• split text into words;4
• remove stopwords;
• remove words with less than two characters;
• lemmatize/stem words;5
• compute tf-idf for each word;
• compute relevance score (BM25) for each word.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Experiments</title>
      <p>In this section, we briefly describe the setting of oficial runs that we submitted for this task
and the preliminary results sent by the organizers before the workshop.</p>
      <sec id="sec-3-1">
        <title>3.1. Dataset</title>
        <p>The datasets of the two subtasks consisted in:</p>
        <p>• Corpus: 3,257 casedocs
• Queries: 50 description of situations
• Corpus: 197 statutes
• Queries: 50 description of situations (same as subtask 1)
2https://www.tidyverse.org
3https://www.eosc-portal.eu
4https://www.tidytextmining.com
5https://cran.r-project.org/web/packages/corpus/vignettes/stemmer.html</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Run Settings</title>
        <p>For subtask 1, we split each casedocs into the following parts (we add an example for casedoc
number 1):</p>
        <p>In the runs of subtask1, we used only the last field (that we named ‘text’) for the retrieval of
precedents.</p>
        <p>The goal of our experiments is to compare the efectiveness of the diferent lexical choices
(stem or lemma) with a baseline (BM25).</p>
        <p>We submitted three runs for each subtask, and we followed the same procedure for each
subtask:
• bm25_lemma: this run uses a BM25 retrieval model with lemmas;
• tfidf_lemma: this run uses a tfidf document representation and a cosine similarity score
on lemmas;
• tfidf_stem: this run uses a tfidf document representation and a cosine similarity score on
stems;</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Results</title>
        <p>The organizers of this task provided the results (averaged across topics) achieved by many
baselines compared to the runs of each participant. In Table 1 and Table 2, we show a summary
of these results.</p>
        <p>A preliminary analysis of the results shows that, in terms of standard evaluation measures
such as MAP, BPREF, reciprocal rank, and P@10 the use of the BM25 with lemmas performed
worse compared to the other two runs.</p>
        <p>For subtask 1 (Table 1), the performance was poor compared to the median values for all the
measures. This result, compared with the performance of the other subtask, requires a failure
analysis to understand how the choice of just on field afected the performance, as well as how
additional information in the statistics of the word may be useful or not (see for example team
HLJIT2019-AILA of AILA 2019 [6]).</p>
        <p>For subtask 2, the use of a vector space model with tf-idf on stems was the second best run
overall. It is interesting to see that, despite the type of query (which is a complex description
of the situation), and the result in the previous subtask, this approach performed very well
compared to the other systems.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Final remarks and Future Work</title>
      <p>The aim of our participation to the FIRE AILA 2020 Task 1 was to test the efectiveness of
a reproducible baseline without any learning strategy. The initial results show a completely
diferent scenario according which subtask is considered. The approach seems very promising,
but a failure analysis and a topic-by-topic comparison is needed to understand when and how
the diferent combination in the retrieval pipeline are significantly better/worse than other
models.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>P.</given-names>
            <surname>Bhattacharya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mehta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Pal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bhattacharya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Majumder</surname>
          </string-name>
          ,
          <article-title>Overview of the FIRE 2020 AILA track: Artificial Intelligence for Legal Assistance</article-title>
          ,
          <source>in: Proceedings of FIRE 2020 - Forum for Information Retrieval Evaluation</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>L.</given-names>
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Suominen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          , G. Pasi,
          <string-name>
            <given-names>G. S.</given-names>
            <surname>Gonzales</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Viviani</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>Xu, Overview of the CLEF eHealth 2020 task 2: Consumer health search with ad hoc and spoken queries</article-title>
          , in: Working Notes of Conference and
          <article-title>Labs of the Evaluation (CLEF) Forum</article-title>
          , CEUR Workshop Proceedings,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>