<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>MedSimples: An Automated Simpli cation Tool for Promoting Health Literacy in Brazil?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Luis Antonio Leiva Hercules</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maria Jose Bocorny Finatto</string-name>
          <email>natto@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Automatic Data Processing</institution>
          ,
          <addr-line>ADP</addr-line>
          ,
          <country country="BR">Brazil</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Universidade Federal do Rio Grande do Sul</institution>
          ,
          <addr-line>UFRGS</addr-line>
          ,
          <country country="BR">Brazil</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Surrey</institution>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <abstract>
        <p>Functional illiteracy rates in Brazil are critical. According to a recent study (2018) conducted by the Paulo Montenegro Institute, 3 out of 10 Brazilians are considered functional illiterates. Also according to INAF4, only 12% are truly pro ccient readers. On the other hand, with the signi cant increase of Internet access in Brazil in the past years, information is available to a much larger number of people. According to the Brazilian Institute of Geography and Statistics (IBGE), in 2017, 67,00% of the Brazilian population had access to the Internet. Although a search on the Internet by a layperson cannot replace going to the doctor, the growing number of Internet users who rely on it as a source of information is a reality that cannot be overlooked. Therefore, it is important that the source be reliable, but also that the information provided on these sources be linguistically accessible and understandable by people with low levels of literacy. In this scenario, our research problem is this: How can we render health-related information available on the Web in a linguistically accessible way to people with limited education and low literacy skills? Our project combined Natural Language Processing, Corpus Linguistics and Terminology Studies to develop MedSimples, an online tool that automatically identi es complex phrases in health-related texts presented by the user and o ers suggestions of simpli cation. MedSimples is an example of how research on medical language, associated with NLP, can enrich the current scenario of Digital Humanities in its broadest scope. The prototypical user for the tool is the Health or Communication professional interested in producing accessible texts graduated according to the needs of their audience. However, it can also be used by anyone who is interested in getting simpli ed health-related information for themselves. Our Text Simpli cation approach is focused on lexical and terminological levels, and the goal is to reduce complexity while preserving meaning. Lexical simpli cation can have</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        L. Paraguassu et al.
two approaches [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]: modifying the vocabulary on a text by selecting words that
are more adequate to the reader's reading skills, or adding explanations or
definitions to the vocabulary that cannot be replaced. MedSimples combines both
approaches, using the suggestions of modi cations mostly for complex phrases,
and then using simpler de nitions for the terminology that is present in the text.
We believe that this way the information is preserved, being more accurate and
reliable, and, in addition to that, by explaining the complex vocabulary, we can
promote health literacy by educating our readers on health issues.
      </p>
      <p>The tool was initially built based on a Parkinson's Disease (PD) corpus. The
PD corpus is a representative collection of original texts available on the Internet
and published by reliable sources. These texts have been simpli ed by
Linguistic graduate and undergraduate students and validated by health specialists to
build a parallel simpli ed PD corpus. This corpus was then further analysed
and converted into an initial list of terms and a list of example-sentences with
varying degrees of complexity, that can be used for consultation by the health
professionals.</p>
      <p>
        In terms of automatic processing of texts, for identifying complex phrases
and terms, and o ering suggestions of simpli cation or simpler de nitions for
the user, MedSimples currently relies on three lexical resources and one parser.
The three lexical resources comprise a list of words that are considered simple,
a list of words that are considered complex along with simpler synonyms, and a
glossary of terms with simple de nitions. The rst step in MedSimples process is
to parse the text that was selected by the user, we used the PassPort parser [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] for
annotating morphological and lemma information on the text. This annotated
text is then processed and, when MedSimples recognizes a term that is listed in
our glossary it automatically provides a simpli ed explanation. For the complex
words that are not considered health-related terms, we used the synsets present
in the Thesaurus of Portuguese (TeP 2.0) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], which were ltered based on a list
of simple words extracted from CorPop [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], a corpus of texts considered fairly
accessible to the average Brazilian reader.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Maziero</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pardo</surname>
          </string-name>
          , T.:
          <article-title>Interface de acesso ao tep 2.0{thesaurus para o portugu^es do brasil</article-title>
          . Relatorio tecnico. University of Sao Paulo (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Pasqualini</surname>
            ,
            <given-names>B.F.</given-names>
          </string-name>
          :
          <article-title>CorPop: um corpus de refer^encia do portugu^es popular escrito do Brasil</article-title>
          .
          <source>Ph.D. thesis, Universidade Federal do Rio Grande do Sul</source>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Saggion</surname>
          </string-name>
          , H.:
          <article-title>Automatic text simpli cation: Synthesis lectures on human language technologies</article-title>
          , vol.
          <volume>10</volume>
          (
          <issue>1</issue>
          ). California, Morgan &amp; Claypool Publishers (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Zilio</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wilkens</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fairon</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Passport: A dependency parsing model for portuguese</article-title>
          .
          <source>In: International Conference on Computational Processing of the Portuguese Language</source>
          . pp.
          <volume>479</volume>
          {
          <fpage>489</fpage>
          . Springer (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>