<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>NLPalma @ CLEF 2023 SimpleText: Complexity and Simplification Task</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Victor Manuel Palma Preciado</string-name>
          <email>victorpapre@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Carolina Palma Preciado</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Grigori Sidorov</string-name>
          <email>sidorov@cic.ipn.mx</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Instituto Politécnico Nacional de México</institution>
          ,
          <addr-line>Gustavo A. Madero, Ciudad de México</addr-line>
          ,
          <country country="MX">México</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Université de Bretagne Occidentale</institution>
          ,
          <addr-line>HCTI</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The following work has the purpose of describing the participation in the SimpleText 2023 track, on the identification of the term and its identification of the terms and their difficulty as a term, among themselves, all this belonging to Task 2. To solve this task, we used an approach of language models using Bloom but opted for its BLOOMZ version for a fine-tuning more focused on human instructions or in a more understandable way with more description-style prompts given by text input on a task. To solve the handling of the difficulty between terms a very simple classifier based on BERT-multilingual was used since this was developed as a binary classification and for the term vs. term evaluation a small algorithm was taken to accommodate the internal term classifications. On the other hand, we also participated in Task 3 in which the objective was to simplify passages extracted from abstracts, using the same approach as Task 2, BLOOMZ was used for the simplification of this text since different prompts were tested in case it was necessary to make several passes with those parts that yielded poor results or null results. Given that this was the first time participating in such tasks we can say that the results obtained were quite satisfactory even though we believe that they can be substantially improved with some other approach, which would have to be further reviewed.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        For this work the main objective we seek to fulfill the tasks as completely as possible, we believe
that language models can present a good starting point to develop the tasks of this track, therefore
BLOOMZ [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] was taken as a model to use because its qualities allow us to have input more consistent
with human instruction, This means that you have to be very clear with the instructions given to the
model and therefore the prompt to be used, it is also important to provide the model with as much
context as possible to get the desired results. In this case, the simplification of the passage can alter the
results depending on the previous context given, the same happening if the prompt is altered.
      </p>
      <p>A very important part that had to be taken into account is the extraction of important terms in the
given sentence for the task, besides assigning a number according to its complexity, so this complexity
can have different parameters to measure the level of complexity whether it is the size of the word, the
linguistic context, among others, so taking any of them could perhaps impact this type of classification
based on complexity, so it was decided to opt for a binary classification.</p>
      <p>It is true that the starting point is BLOOMZ but we also have to use other types of models to do
classification, although BLOOMZ is able to do this classification with the correct prompt, we believe
that the use of other tools can speed up the process without having to rely entirely on BLOOMZ in
which the more parameters you try to take, the longer it will take to give results, obviously this in many
cases allows better results despite the delay that could represent. To classify we use a BERT-type model,
in this case the multilingual one, to foresee any eventuality of the dataset.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Approach to the task</title>
      <p>
        From the task proposed by SimpleText [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], Task 2 and Task 3 were carried out by implementing
NLP models specifically deep learning models like transformers. In Task 2 identification and
explanation of difficult concepts in the paragraph are done with the aim to define the meaning of key
concepts for better conceptualization and understanding of a paragraph. On the other hand, Task 3 aims
to create summarized text from scientific abstracts written in English, this concise version should
maintain the original meaning in a simplified form.
      </p>
      <p>
        The data [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] provided is divided into three dataset sizes (small, medium, and large), for this approach
it was decided to use the medium dataset since is the middle ground of a sample of the text and sufficient
data to observe the performance of the applied method. Text summarization is a difficult task that can
be achieved from different approaches, in this case, a deep learning model known as BLOOMZ &amp; mT0
is used to generate text summarization and determine difficult words in a paragraph. Since there are
multiple versions of this model with different parameters, different versions were used when performing
the tasks as the resources available were not sufficient to run some models.
      </p>
      <p>
        The first step to complete Task 2, was to identify up to five of the most difficult terms in a sentence
the first pass was done by using mt0-xl since this version could be loaded in a local environment without
running out of memory, the first pass resolved around half of the dataset (2,320 sentences). After having
several empty results with various prompts like “Give me up to five of the most difficult words of the
next text:”, “Give me up to five of the most difficult words of the next sentence:”, “Give me the top
five most difficult words of the next text:”, and “Suggest me at up five of the most difficult words on
the next text:” it was decided to do the second pass and third pass with the Hugging Face inference API
[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] that uses the BLOOMZ finetuned model with 176B parameters.
      </p>
      <p>To explain the terms obtained in the first step, the same model is now implemented with prompts
such as “Meaning of:”, “Definition of:”, “Give me the definition:”, “Give me the meaning:” all of this
prompt were used in succession if one of them did not produce results, the next prompt was used until
a favorable output was obtained. Since the task accepts both the definition of the word and an example
of it, no further process was done.</p>
      <p>
        The next step consists in scoring and ranking the terms, for the first part a multilingual BERT [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]
was trained with the corpus provided, the score in the training text varies from 1 to 2 so the terms were
assigned this range of scoring, 2 indicating a difficult term and 1 an easier one. The BERT model trained
for the scoring was executed with the help of the Ktrain [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] wrapper that allows to load and perform
the fine-tuning process.
      </p>
      <p>Then Term ranking is done using the score and the term length, fist the words of the sentences are
found and separated by the score, then a sub-ranking is designated from the length of the words, where
the term with the most characters is the most difficult and the term with the fewest is the easiest term.
If the terms are the same length they are arranged in order of appearance. Finally, when both lists for
scores 1 and 2 are ordered, they are united to create the final ranking.</p>
      <p>
        To elaborate Task 3 BLOOMZ is also implemented, this time instead of using the inference API a
system for inference and fine-tuning called Petals [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] is used since the Hugging Face API, this platform
joins computer sources from different servers to increase the computational power. The summarization
prompts include “Summarize the text:” and “Summarize the sentence:”. As with previous processes,
multiple runs were executed to obtain the best results.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Resources employed</title>
      <p>Among the resources used to train and evaluate the models, Google's Colab environment was used,
this platform allows Python programming and execution. It also allows an easier use of GPU, which
served to perform the described tasks faster since this resource allows performing multiple simultaneous
computations. The server used has the following specifications GPU NVIDIA-SMI 525.85.12, CUDA
v12.0, and 25 of RAM.</p>
      <p>As part of the resources used, it was decided to use a combination of Google Collaboratory with
Petals for Task 2 in complexity spotting and Task 3 for simplification of scientific text, since we do not
have the computational capacity to run models of parameter size 176B, which represent a larger amount
of computation than we can handle, so the use of Petals becomes indispensable for those jobs that do
not have enough computational capacity since it allows the benefits of crowd computing to make an
inference with this type of large models in a relatively easy/semi-efficient way and in a way having the
flexibility of an API and the power of PyTorch.</p>
      <p>
        On the other hand, for the complexity ranking, we used BERT-multilingual [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], and for the internal
complexity between terms of the same sentence we used a small algorithm that takes the length of the
sentence and the value obtained from the previous classification to make 2 groups, if necessary, between
those that have the value of 1 and 2 in complexity to generate the internal complexity ranking.
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <p>As part of the results that we will present below, we will explain from a very general point of view
the type of results obtained for each task and some of their particularities. We will explain some of the
cases that we believe could be interesting to take into account for our future participation in these tasks
for SimpleText in subsequent deliveries, taking into account that perhaps we could use some other more
ad hoc models for tasks such as simplification.</p>
      <p>Task 2: “What is unclear?” Difficult concept identification and explanation</p>
      <p>In the following example, we can see that the word that was extracted is CNN, but it did not know
how to obtain the context of the sentence, so when throwing the meaning or the explanation it
misunderstands CNN (Convolutional Neural Network) for CNN (Cable News Network) and this is
completely incorrect, since as we can infer they are completely different things, one refers to a Machine
learning model and the other to a news network, so there is no point of comparison and the result
obtained refers to the latter.</p>
      <p>•
•
•
•</p>
      <p>A conventional CNN for the end-to-end control is designed to map a single front-facing
camera image to a steering command.</p>
      <p>Term rank: 2
Word: CNN</p>
      <p>Word Meaning or explanation: CNN.com</p>
      <p>It is logical to think that cnn.com refers to the web address of the news channel and clearly not to
the type of neural network since there is no way that such a domain refers to an abstract element rather
than a company.</p>
      <p>The following example is one that we believe represents a term that could be difficult to explain,
since being acronyms these can mean different things, which will depend on the field of study that
makes reference to the context of the sentence, so it is extremely important to take this into account and
in this case, it is clear that the training data off ease in the way of how to find a difficult term and
therefore the context that has these acronyms, throw a correct explanation.</p>
      <p>•</p>
      <p>This work explores the electromagnetic (EM) side-channel analysis literature for the
purpose of assisting digital forensic investigations on IoT devices.
•
•
•
•
•
•
•
•</p>
      <p>Term Rank: 2
Word: IoT
Word Meaning or explanation: Internet of Things (IoT) is a network of physical objects or
"things" embedded with electronics</p>
      <p>It could also refer to another acronym such as EM but since this term could be complex, but is
already explained in that sentence, it would not make much sense to take it into account.</p>
      <p>The classification was another aspect that had to be applied in the sentences, this is a classification
for the term extracted given its complexity and if more than one term was extracted from the sentence
these had to be classified among themselves obtaining an ascending value given its complexity, this can
be seen in the following words extracted from the same passage:
attack Term difficulty: 2 Internal term difficulty: 1
system Term difficulty: 2 Internal term difficulty: 2
blind Term difficulty: 2 Internal term difficulty: 3
False data Term difficulty: 1 Internal term difficulty: 4
injection Term difficulty: 1 Internal term difficulty: 5</p>
      <p>It is observed that the length of the word does not play such an important role in the rank assigned
to a term under its difficulty, the explanation can only be inferred but that certain words are ranked
higher than others is due to the role they play in the context and the difficulty to explain them in the
given context.</p>
    </sec>
    <sec id="sec-5">
      <title>Task 3: Rewrite this! Rewriting scientific text</title>
      <p>The case of simplifying a passage is a very interesting one, the idea of compacting a sentence or idea
to such an extent, can be complicated, since this can lead, if it is very compact, to lose information and
not be able to fully understand the idea or lend itself to an erroneous interpretation, on the other hand,
that the idea is quite loose and does not explain things well, by summarizing incorrect ideas.</p>
      <p>As can be seen, the following sentence was summarized erroneously since, as can be seen, it loses
the focus of the information and does not mention that the data were collected by researchers, the center
of information is the process, not the people who collected it.</p>
      <p>•
•</p>
      <p>Excerpt: Permissions were taken from required authorities who made sure that minimal risk
(to pedestrians) is involved in the data collection process.</p>
      <p>Simplified excerpt: The data was collected by a team of trained researchers</p>
      <p>In this positive case of simplification, we can observe the type of simplification of the information
in terms of content while maintaining the context and even explaining it in much simpler terms, instead
of saying that the information is produced and accessed, it changes it to a source of information, which
in this context means the same thing, although the final part of the sentence could be omitted and would
not affect the result producing a more compact fragment.</p>
      <p>•
•</p>
      <p>Excerpt: The advent of WWW changed the way we can produce and access information.</p>
      <p>Simplified excerpt: The Web has become a major source of information for many people</p>
    </sec>
    <sec id="sec-6">
      <title>5. Conclusions</title>
      <p>We can conclude that our decisions in the approach taken can give us good results and observe that
they can be improved in a substantial way in which the results are presented, for example in the
simplification task it can be observed that many times the context of the original sentence is lost, perhaps
giving more information about the same sentence could solve this problem, On the other hand, for the
task of term extraction, we could observe that it is difficult to obtain an accurate way of knowing that
the term is really difficult in terms of its context or that perhaps it is a compound term or an acronym
that makes it complex, this leaves open different possibilities for improvement that we believe can be
exploited in subsequent works.</p>
    </sec>
    <sec id="sec-7">
      <title>6. References</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Liana</given-names>
            <surname>Ermakova</surname>
          </string-name>
          ,
          <string-name>
            <surname>Eric</surname>
            <given-names>SanJuan</given-names>
          </string-name>
          , Stéphane Huet, Olivier Augereau, Hosein Azarbonyad, and
          <string-name>
            <given-names>Jaap</given-names>
            <surname>Kamps</surname>
          </string-name>
          .
          <year>2023</year>
          .
          <article-title>Overview of SimpleText - CLEF-2023 track on Automatic Simplification of Scientific Texts</article-title>
          . In Avi Arampatzis, Evangelos Kanoulas, Theodora Tsikrika, Stefanos Vrochidis, Anastasia Giachanou,
          <string-name>
            <given-names>Dan</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Mohammad</given-names>
            <surname>Aliannejadi</surname>
          </string-name>
          , Michalis Vlachos, Guglielmo Faggioli, Nicola Ferro (Eds.)
          <string-name>
            <surname>Experimental IR Meets Multilinguality</surname>
          </string-name>
          , Multimodality, and
          <string-name>
            <surname>Interaction</surname>
          </string-name>
          .
          <source>Proceedings of the Fourteenth International Conference of the CLEF Association (CLEF</source>
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Muennighoff</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutawika</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roberts</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Biderman</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Scao</surname>
            ,
            <given-names>T. L.</given-names>
          </string-name>
          ,et.al. (
          <year>2022</year>
          ).
          <article-title>Cross-lingual generalization through multitask finetuning</article-title>
          .
          <source>ArXiv Preprint ArXiv:2211</source>
          . 01786.S. Cohen,
          <string-name>
            <given-names>W.</given-names>
            <surname>Nutt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Sagic</surname>
          </string-name>
          ,
          <article-title>Deciding equivalances among conjunctive aggregate queries</article-title>
          ,
          <source>J. ACM</source>
          <volume>54</volume>
          (
          <year>2007</year>
          ). doi:
          <volume>10</volume>
          .1145/1219092.1219093.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Ermakova</surname>
          </string-name>
          , Liana, Eric SanJuan, Stephane Huet, Olivier Augereau, Hosein Azarbonyad, and Jaap Kamps.
          <source>“CLEF 2023 SimpleText Track: What Happens If General Users Search Scientific Texts?” ECIR 2023 Proceedings.</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Devlin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
          </string-name>
          , M.-W.,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Toutanova</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          (
          <year>2018</year>
          ).
          <article-title>BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding</article-title>
          . CoRR, abs/
          <year>1810</year>
          .04805. Retrieved from http://arxiv.org/abs/
          <year>1810</year>
          .04805I.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Arun</surname>
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Maiya</surname>
          </string-name>
          (
          <year>2020</year>
          ).
          <article-title>ktrain: A Low-Code Library for Augmented Machine Learning</article-title>
          . arXiv preprint arXiv:
          <year>2004</year>
          .10703.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Pires</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schlinger</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Garrette</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          (
          <year>2019</year>
          ). How Multilingual is Multilingual BERT? https://doi.org/10.18653/v1/p19-
          <fpage>1493</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Inference</surname>
            <given-names>API - Hugging</given-names>
          </string-name>
          <string-name>
            <surname>Face</surname>
            . (s. f.). https://huggingface.co/inference-apiCañete,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chaperon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fuentes</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ho</surname>
            ,
            <given-names>J.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kang</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Pérez</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2020</year>
          ).
          <article-title>Spanish PreTrained BERT Model and Evaluation Data</article-title>
          .
          <source>In PML4DC at ICLR</source>
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Borzunov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baranchuk</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dettmers</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ryabinin</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Belkada</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chumachenko</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Samygin</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Raffel</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          (
          <year>2022</year>
          ).
          <article-title>Petals: Collaborative Inference and Fine-tuning of Large Models</article-title>
          . arXiv (Cornell University). https://doi.org/10.48550/arxiv.2209.01188
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>