<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Carolina's Methodology: building a large corpus with provenance and typology information</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Center for Artificial Intelligence, University of São Paulo. Av. Prof. Lúcio Martins Rodrigues</institution>
          ,
          <addr-line>370 - 05508-020 - Butantã, São Paulo</addr-line>
          ,
          <country country="BR">Brazil</country>
        </aff>
      </contrib-group>
      <fpage>0000</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>This paper presents the salient aspects of WaC-wiPT methodology, developed for the construction of the Carolina Open Corpus for Linguistics and Artificial Intelligence, a large corpus for contemporary Brazilian Portuguese. Both the corpus and the methodology are under development at the Center for Artificial Intelligence of the University of São Paulo. This paper describes the paths we took this far into the making of the Carolina Corpus, presents its current state and discloses the future agenda of the project.</p>
      </abstract>
      <kwd-group>
        <kwd>Open Corpus</kwd>
        <kwd>Brazilian Portuguese</kwd>
        <kwd>Provenance</kwd>
        <kwd>Typology</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The Carolina Open Corpus for Linguistics and Artificial Intelligence is a general corpus
of contemporary Brazilian Portuguese texts written after 1970 hosting provenance and
typology information. It is under development since September 2020 as part of the
Natural Language Processing for Portuguese (NLP2) project of the Center for Artificial
Intelligence of the University of São Paulo (C4AI-USP).</p>
      <p>
        With Carolina, we expect to build a large and reliable resource for research in both
Linguistics and Computer Sciences, with more than a billion tokens. In opposition to
other corpora built under the “Web as corpus” view, which aim to gather large amounts
of texts for language-modeling retrieving them from multiple untraceable origins,
Carolina’s intention is to curate sources in large quantities of text. In doing that, we expect
to provide further information about the texts, especially on provenance and typology
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], which benefit linguistic research.
      </p>
      <p>By not tailoring the corpus for any specific linguistic application, we aim to avoid
restraining the possibilities for future projects and posing obstacles for any researchers
interested in studying a wide range of language aspects. For instance, investigation of
1 Copyright © 2022 for this paper by its authors. Use permitted under Creative Commons License</p>
      <p>Attribution 4.0 International (CC BY 4.0).}
typological characteristics, word collocation, language detection and Historical
Linguistics. To reach this goal, we developed the WaC-wiPT (Web-as-Corpus with
Provenance and Typology) methodology, which combines the automation and large
extension of language-modeling corpora with the careful text-information curatorship of
smaller linguistic corpora.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related works</title>
      <p>
        Over the last decades, corpus-building initiatives have increasingly resorted to the Web
as their main source. As part of this endeavor, the WaCky (Web-As-Corpus Kool
Yinitiative) methodology was developed [
        <xref ref-type="bibr" rid="ref2 ref3 ref4">2, 3, 4</xref>
        ]. As it proved to be a relatively easy and
not high-resource-demanding method, this framework quickly became popularized.
There have been undertakings to apply it to the Portuguese language, such as the
Brazilian Portuguese Web as Corpus (brWaC), considered to be the “biggest Brazilian
Portuguese corpus available” at the time [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], with 2.68 billion tokens.
      </p>
      <p>
        One example of a Brazilian Portuguese corpus with a significant size is the Brazilian
Corpus [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], with approximately one billion syntactically annotated words. There are
also other important corpora out of the envisioned scope of language or size, such as
the Oscar Corpus [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], or the Corpus do Português: Web/Dialects [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Other corpora with
provenance and typology information are known, such as ReLi [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and CETENFolha
[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], albeit their smaller sizes according to their specific goals.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Methodology</title>
      <p>
        During the first stages of our research, some effort was put into investigating the
possibility of implementing pre-existing corpora-building frameworks, such as the WaCky
methodology [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. However, it does not ensure the transparency of the content that is
scraped, so a post-hoc investigation is necessary [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. This poses a challenge for
provenance tracking, quality control, and rights-of-use compliance, which are at the core of
Carolina's objectives. Therefore, we built on the knowledge provided by this
investigation to develop WaC-wiPT, a Web-as-Corpus with Provenance and Typology
methodology, which is constantly being improved.
3.1
      </p>
      <sec id="sec-3-1">
        <title>Web text prospection</title>
        <p>The fundamental steps of the method are based on broad types of domains, the Carolina
broad typology, which are not intended to reflect the textual content of our documents,
but a macro-structure that guides the corpus development. They were defined after the
surveying process described below. Currently, we are working with eight: datasets and
other corpora, Brazilian Judicial branch and Legislative branch, journalistic texts,
public domain works, social media2, university domains and wikis. The Carolina broad
2 It is important to clarify that we only incorporated open access content, in compliance with each
domain’s license.
typology opposes the Carolina narrow typology, which aims to effectively reflect the
textual types of our documents, which are not yet defined for they require further
analysis of the texts and may depend on linguistic theory; as well as the source typology, a
simple typological organization that reflects exclusively what is declared on the sources
and therefore are declared when available.</p>
        <p>We began by conducting prospective surveys, which are in-depth research of each
Web domain, prioritizing open access content available online. In these surveys, we
verified if the texts were in our scope, in addition to searching for metadata information
and mapping the basic directory structure of each domain. Therefore, this first step is
important to help us systematize metadata for future automatic annotation — so we do
not lose any important information from our sources — and organize the provenance
of the data as well. These surveys also facilitate the download and extraction stages, for
they allow the downloaded content to be mostly deliberate and not randomly crawled.</p>
        <p>At the beginning of the data collection step, we mirrored some websites to keep raw
copies of our sources, but even in those cases, we could directly extract the desired texts
because of the mapping of the directory structure previously made. However, most
websites were not obtained with this method for some had defense mechanisms against
machine download and others could only be used partially, as they contained texts that
were not in our scope or were under restrictive licenses. In all cases, special care was
taken to verify the rights of use: during the data collection step, we downloaded
exclusively open access texts and later verified if they allowed derivative works. Should the
occasion arise that any data is copyright-claimed, our methodology enables the easy
removal of any set of texts from the corpus.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Metadata and extraction</title>
        <p>As we aim to build a corpus with provenance and typology information, each text is
embedded in an XML header carrying annotated metadata — such as source URL and
license — following the TEI (Text Encoding Initiative) guidelines. To gather
information for the header, after the download stage some surveys are complemented by
opening a small sample of the raw documents and searching for any additional
metadata. However, as we have a large number of texts, information that cannot be
automatically annotated is not mandatory, thus, most categories are optional, such as
Author and Regional Origin. To make this process easier we extract the texts by batches
arranged by coincident information, usually grouped by downloaded directory
structure, which mirrors the broad typology's.</p>
        <p>This means that for the extraction of a batch, we inform the metadata that holds
for the whole set, prioritizing the fulfilling of the header’s mandatory categories.
Therefore, the metadata collection and insertion are carefully made, to prevent errors from
being repeated in the whole batch. We centralize this latter process with an extraction
module developed in Python3, which obtains some metadata by input and others
automatically, organizes them and generates the XML file with the clean text embedded in
the header3. It also verifies if the text is valid by assessing its language and size, for
example. In addition to the traditional search by words, the structured header allows
searches by tags, facilitating the metadata recovery, thus providing further query tools.</p>
        <p>In order to test our tools, we processed a portion of our raw data before our first
official extraction, with over a billion tokens total and about 24 hours of CPU time.
This first test version will not be made publicly available; however, it sheds some light
on what is to be expected of the first official publication in terms of size and typology
distribution. The texts obtained were sampled randomly (590 files in total) by Carolina
broad typology, of which only 4 broad types were included.</p>
        <p>We carefully looked for problems concerning not only the cleaning process, but also
the metadata provided, and implemented computational solutions to improve textual
quality, such as removing remaining blank lines and corrupted characters. Some
recurring issues were over or under-cleaned texts, as well as the formatting of the data
provided automatically, which sometimes did not match our chosen standards. After this
process, we established the importance of human inspection of the files, as some
problems would not be easily identified and fixed without it. Therefore, this method of
examining the samples of the extracted files will be kept in the future as well as some
machine inspections will be made to verify if all the files are well-formatted.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Current state</title>
      <p>At present, Carolina is at a prototypical stage. After the test extraction, some
considerable yet expected size-reduction took place in relation to the raw crawled content4. The
main reason for that is that files were disposed of tags and non-content elements aiming
at a text as clean as possible, and many files did not reach a significant number of
characters and were thus discarded. The table ahead illustrates these reductions.
3 An example of a generated XML header can be accessed at:
https://sites.usp.br/corpuscarolina/exemplo/
4 The raw crawled content encompasses everything downloaded from each domain, including
media files, source metadata and content out of the project’s current scope.
5 A list of the content available on the first version of the Carolina Corpus can be accessed at:
https://sites.usp.br/corpuscarolina/repositorios/</p>
      <p>Based on this initial testing of our methods' performance, we have already begun the
proceedings for the first version, to be released in March 2022, the Carolina 1.0 (Ada),
following the steps of this under development methodology.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion and future steps</title>
      <p>Both the preliminary surveys and the close analysis of the samples extracted for the test
version proved to be essential to keep track of the metadata collection and sustain such
complex text headers. Thus, our methodology helped to guarantee the provenance and
typology information we aim to preserve in the automatic processes of text extraction.</p>
      <p>As for future intentions, we wish to develop language verification tools capable of
distinguishing Brazilian Portuguese from other variants and assess the percentage of all
languages used in a text, hoping that would allow studies on language contact and
cultural influences. Additionally, there are plans to build a historical corpus for
Philological and Historical Linguistics research based on our methodology.</p>
      <p>For the future versions of the Carolina corpus, we will work to ensure better
balancing of text types, which require further effort put into new surveys. In our
understanding, this continuous development of the methodology is an essential part of the labor
involved in the construction of such an ever-growing corpus.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Finger</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , Paixão de Souza,
          <string-name>
            <given-names>M. C.</given-names>
            ,
            <surname>Namiuti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Monte</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. M.</given-names>
            ,
            <surname>Costa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. S.</given-names>
            ,
            <surname>Serras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. R.</given-names>
            ,
            <surname>Sturzeneker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. L.</given-names>
            ,
            <surname>Guets</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. P.</given-names>
            ,
            <surname>Mesquita</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. M.</given-names>
            ,
            <surname>Crespo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. C. R. M.</given-names>
            ,
            <surname>Rocha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. L. S. J.</given-names>
            ,
            <surname>Palma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. F.</given-names>
            ,
            <surname>Silva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. M.</given-names>
            ,
            <surname>Brasil</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.</surname>
          </string-name>
          <article-title>Carolina: a General Corpus of Contemporary Brazilian Portuguese with Provenance and Typology Information</article-title>
          .
          <article-title>Language resources and evaluation, submitted paper (</article-title>
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Baroni</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bernardini</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferraresi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zanchetta</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          :
          <article-title>The WaCky wide web: a collection of very large linguistically processed web-crawled corpora</article-title>
          .
          <source>Language resources and evaluation</source>
          ,
          <volume>43</volume>
          (
          <issue>3</issue>
          ),
          <fpage>209</fpage>
          -
          <lpage>226</lpage>
          (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bernardini</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baroni</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Evert</surname>
          </string-name>
          , E.:
          <article-title>A WaCky introduction</article-title>
          . In: Baroni,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Bernardini</surname>
          </string-name>
          , S. (eds.)
          <article-title>WaCky! working papers on the web as corpus</article-title>
          , pp.
          <fpage>9</fpage>
          -
          <lpage>40</lpage>
          . GEDIT, Bologna (
          <year>2006</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Ferraresi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bernardini</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Picci</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baroni</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Web corpora for bilingual lexicography: A pilot study of English/French collocation extraction and translation</article-title>
          .
          <source>In: Using Corpora in Contrastive and Translation Studies</source>
          , pp.
          <fpage>337</fpage>
          -
          <lpage>362</lpage>
          . Cambridge Scholars Publishing,
          <string-name>
            <surname>Newcastle</surname>
          </string-name>
          (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Boos</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prestes</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Villavicencio</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Padró</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>brWaC: a wacky corpus for Brazilian Portuguese</article-title>
          . In: Baptista,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Mamede</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Candeias</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Paraboni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            ,
            <surname>Pardo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.A.S.</given-names>
            ,
            <surname>Volpe</surname>
          </string-name>
          <string-name>
            <surname>Nunes</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.G. (eds.) PROPOR</surname>
          </string-name>
          <year>2014</year>
          .
          <article-title>LNCS</article-title>
          , vol.
          <volume>8775</volume>
          , pp.
          <fpage>201</fpage>
          -
          <lpage>206</lpage>
          . Springer, Heidelberg (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Sardinha</surname>
            ,
            <given-names>T. B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Filho</surname>
            ,
            <given-names>J. L. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alambert</surname>
          </string-name>
          , E.: Manual Córpus Brasileiro, https://www.linguateca.pt/Repositorio/manual_cb.pdf,
          <source>last accessed</source>
          <year>2021</year>
          /12/13.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Suárez</surname>
            ,
            <given-names>P. J. O. S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sagot</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Romary</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Asynchronous Pipeline for Processing Huge Corpora on Medium to Low Resource Infrastructures</article-title>
          .
          <source>In: Proceedings of the Workshop on Challenges in the Management of Large Corpora (CMLC-7)</source>
          <year>2019</year>
          , pp.
          <fpage>9</fpage>
          -
          <lpage>16</lpage>
          .
          <article-title>Leibniz-Institut für Deutsche Sprach</article-title>
          ,
          <string-name>
            <surname>Mannheim</surname>
          </string-name>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Davies</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferreira</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          : Corpus do Português: Web/Dialetics, https://www.corpusdoportugues.org/web-dial/,
          <source>last accessed</source>
          <year>2021</year>
          /12/13.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Corpus</surname>
            <given-names>ReLi</given-names>
          </string-name>
          , https://www.linguateca.pt/Repositorio/ReLi/, last accessed
          <year>2022</year>
          /03/10.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10. CETENFolha, https://www.linguateca.pt/CETENFolha/, last accessed
          <year>2022</year>
          /03/10.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>