<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Diachronic Italian Corpus based on “L'Unita` ”</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Pierpaolo Basile</string-name>
          <email>pierpaolo.basile@uniba.it</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Annalina Caputo</string-name>
          <email>annalina.caputo@dcu.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tommaso Caselli</string-name>
          <email>t.caselli@rug.nl</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pierluigi Cassotti</string-name>
          <email>pierluigi.cassotti@uniba.it</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rossella Varvara</string-name>
          <email>rossella.varvara@unifi.it</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>ADAPT Centre, School of Computing, Dublin City University</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>CLCG, University of Groningen</institution>
          ,
          <country country="NL">Netherlands</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>DILEF, University of Florence</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Dept. of Computer Science, University of Bari</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>English. In this paper, we describe the creation of a diachronic corpus for Italian by exploiting the digital archive of the newspaper “L'Unita`”. We automatically clean and annotate the corpus with PoS tags, lemmas, named entities and syntactic dependencies. Moreover, we compute frequency-based time series for tokens, lemmas and entities. We show some interesting corpus statistics taking into account the temporal dimension and describe some examples of usage of time series.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Diachronic linguistics is one of the two major
temporal dimensions of language study proposed by
de Saussure in his Cours de languistique ge´ne´rale
and has a long tradition in Linguistics. Recently,
the increasing availability of diachronic corpora as
well as the development of new NLP techniques
for representing word meanings has boosted the
application of computational models to investigate
historical language data
        <xref ref-type="bibr" rid="ref2 ref6 ref7">(Hamilton et al., 2016;
Tahmasebi et al., 2018; Tang, 2018)</xref>
        . This
culminated in SemEval-2020 Unsupervised Lexical
Semantic Change Detection
        <xref ref-type="bibr" rid="ref5">(Schlechtweg et al.,
2020)</xref>
        , the first attempt to systematically evaluate
automatic methods for language change detection.
      </p>
      <p>Italian is a Romance language which has
undergone lots of changes in its history. Its official
adoption as a national language occurred only
after the Unification of Italy (1861), having
previously been a literary language. Diachronic corpora
of Italian are currently available and accessible to
the public (e.g., DiaCORIS and MIDIA).
Unfortunately, restricted access/distribution of these
resources limits their utilisation. This actually
prevents the investigation of more recent NLP
methods to the diachronic dimensions.</p>
      <p>To obviate this limit, we collect and make freely
available1 a new corpus based on the
newspaper “L’Unita`”. Founded by Antonio Gramsci on
February, 12th 1924, “L’Unita`” was the official
newspaper of the Italian Communist Party (PCI 2,
henceforth). The newspaper had a troubled
history: with the dissolution of PCI in 1991, the
newspaper continued to live as the official
newspaper of the new Democratic Party of the Left
(PDS/DS) until July, 31th 2014. After that date,
it ceased its publication until June, 30th 2015, and
it was definitely closed on June, 3rd 2017.</p>
      <p>Since 2017, the historical archive of “L’Unita`”
has been made again visible and available on the
Web.3 One of the main issues of this resource is
the lack of information about who owns the rights
of the original archive. To our knowledge, the
online version of the archive was legally obtained by
downloading the original archive before the
closure of the newspaper. The current archive,
available online, does not contain the local editions of
the newspaper and the photographic archive.</p>
      <p>The main contribution of this work lies in the
1https://github.com/swapUniba/unita/
2It is the acronym of Partito Comunista Italiano.
3https://archivio.unita.news/
resource itself and its accessibility to the research
community at large. The corpus is distributed in
two formats: raw text and pre-processed. The
validity of the corpus for the automatic study of
language change is currently tested as part of the
DIACR-Ita task 4 at EVALITA 2020. However,
we illustrate some further potential applications of
the use of the corpus.</p>
    </sec>
    <sec id="sec-2">
      <title>2 Italian diachronic corpora</title>
      <p>
        Various Italian diachronic corpora are currently
available and accessible to the public.
DiaCORIS 5
        <xref ref-type="bibr" rid="ref4">(Onelli et al., 2006)</xref>
        comprises
written Italian texts produced between 1861 and
1945, for a total of 100 million words, while
MIDIA 6
        <xref ref-type="bibr" rid="ref1">(Gaeta et al., 2013)</xref>
        covers written
documents in Italian between the beginning of the XIII
century and the first half of the XX century, for a
total of 7,5 million words over 800 texts belonging
to different genres. The Corpus OVI dell’Italiano
antico7 consists of 1948 texts from the XII to the
XIV centuries, for a total of 536.000 words. The
LIZ8 database comprehends 1,000 literary texts
from the XIII to the XX century. Lastly, the
Corpus of Alcide de Gasperi’s public documents
        <xref ref-type="bibr" rid="ref8">(Tonelli et al., 2019)</xref>
        includes 1,762 documents
(newspaper articles, propaganda documents,
official letters, parliamentary speeches, for a total of
3.000.000 tokens) written from the Italian
politician Alcide De Gasperi and published between
1901 and 1954.
      </p>
      <p>These existing resources differ from each other
and from the present corpus in different ways.
First, the span of time the texts come from.
The OVI Corpus considers texts from the early
stages of the Italian language, with a time span of
three centuries. The MIDIA corpus and the LIZ
database cover 7 centuries, from the XIII to the
first half of the XX century. DiaCORIS, the De
Gasperi’s corpus and L’Unita` corpus contain texts
from a shorter and more recent period of time.
However, the time span considered in L’Unita`
corpus is interesting for the study of the Italian
language because of the deep changes that occurred
4https://diacr-ita.github.io/
DIACR-Ita/</p>
      <p>5http://corpora.dslo.unibo.it/
DiaCORIS/
6www.corpusmidia.unito.it
7http://gattoweb.ovi.cnr.it
8https://www.zanichelli.
it/ricerca/prodotti/
liz-4-0-letteratura-italiana-zanichelli
in that period. Indeed, the second half of the XX
century has seen a wider spread and use of Italian
among all the social classes.</p>
      <p>Second, these corpora differ for the genres
represented. The DiaCORIS and MIDIA corpora
have been designed as representative and balanced
samples of written Italian (considering, among
other genres, academic prose, fiction, press, legal
texts, etc). The OVI corpus and the LIZ database
comprehend only literary texts. The De Gasperi’s
corpus is representative of political text from a
single author. L’Unita` corpus is representative only of
press language, but this restriction may be an
advantage in the study of diachronic lexical change.
Indeed, observed semantic changes cannot be
attributed to attestation from different genres in
different periods, but can be interpreted as true
semantic shifts.</p>
      <p>Lastly, even if most of the corpora can be
queried online (with the exception of the LIZ
database), only the De Gasperi’s corpus can be
freely downloaded. This restriction affects the
usability of these resources for the NLP community.
With L’Unita` corpus we aim at releasing a new
diachronic resource that is freely available and that
can be used in the theoretical and computational
study of language change.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Corpus Creation</title>
      <p>The corpus creation consists of several steps:
Downloading All PDF files are downloaded
from the source site and stored into a folder
structure that mimics the publication year of each
article.</p>
      <p>
        Text extraction The text is extracted from the
PDF files by using the Apache Tika library.9 First,
the library tries to extract the embedded text if
present in the PDF; otherwise the internal OCR
is exploited. It is important to notice that during
this step several OCR errors occur. In particular,
during the processing of the early years, the
newspaper has an unconventional format where a few
large pages contain many articles split into several
columns. Due to this format, the OCR is not able
to correctly identify the column boundaries.
Cleaning In this step, we try to fix some text
extraction issues. The previous step leaves an empty
9https://tika.apache.org/
line when the end of a paragraph is reached.
However, a paragraph can be composed of multiple
lines which sometimes contain a word break at
the end of the line. We manage word breaks in
order to obtain a paragraph on a single text line;
we still retain the empty line for delimiting
paragraphs. Moreover, we remove noisy text by
adopting two heuristics: (1) paragraphs must contain at
least five tokens composed by only letter
characters; (2) 60% of the paragraph must contain words
that belong to a dictionary. The dictionary is built
by extracting words that occur into the Paisa`
corpus
        <xref ref-type="bibr" rid="ref3">(Lyding et al., 2014)</xref>
        taking into account only
words composed by letters. The output of this
process is a plain text file for each year where each
paragraph is separated by an empty line.
Processing All plain text files produced by the
cleaning step are processed by a Python script that
splits each paragraph into sentences and analyses
each sentence by performing several natural
language processing tasks. We rely on the spaCy10
Python library for performing: tokenization,
PoStagging, lemmatization, named entity recognition
and dependency parsing. The spaCy library
provides performance comparable to the
state-of-theart approaches with a good processing speed when
compared to other NLP tools.11 We also
provide the plain text in order to allow the
processing with other tools. Each plain text file is
analysed and transformed in vertical format adding
two tags: &lt;p&gt;...&lt;/p&gt; for the begin and the
end of a paragraph, and &lt;s&gt;...&lt;/s&gt; for
delimiting sentences. The vertical format is compliant
to the CONLL representation standard and the
tagset for the Italian12 is automatically mapped to the
10https://spacy.io/
11https://spacy.io/usage/facts-figures
12https://spacy.io/api/annotation
Universal Dependencies scheme13.
      </p>
      <p>Feature Description
Position The token position in the sentence
starting from 1
Token The token
Lemma The lemma
PoS-tag The PoS tag
Tag Additional tags, such as morphological tags
Dependency Dependency type
Head position Head position of the dependency
IOB2 NE IOB2 tag of the named entity
Punctuation Boolean indicating if punctuation
Space Boolean indicating if space character
Stop word Boolean indicating if stop word
Shape pTuhnecwtuoartdiosnh,adpiegi–tscapitalisation,</p>
      <p>Table 2: Description of token features.</p>
      <p>The corpus spans 67 years from 1948 to 2014.
For each year, we provide two files: (1) the plain
text file containing the cleaned text extracted from
PDF where each paragraph is delimited by an
empty line; (2) a vertical file. In the vertical file
format, exemplified in Table 1, each paragraph is
split in sentences and tokens occurring in each
sentence are annotated with 12 features, whose
symbols and descriptions are reported in Table 2.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Corpus Statistics</title>
      <p>In this section, we report some corpus statistics.
Table 3 illustrates the total number of occurrences
and the dictionary size for each feature (token,
lemma, and named entity, respectively).</p>
      <p>dict. size occurrences
token 4,177,128 425,833,098
lemma 4,053,561 425,833,098
named entity 5,429,470 26,330,273
Table 3: Dictionary size and total number of occurrences.
13http://universaldependencies.org/u/
pos/</p>
      <p>The corpus contains more than 400 million
occurrences and more than 25 million named
entities occurrences. The most frequent entities are
Italia, Roma and PCI. This result is expected since
“L’Unita`” was the newspaper of the Italian
Communist Party.</p>
      <p>Figure 1 shows the PoS-tags14 frequency over
time for open-class tags: NOUN, VERB,
ADJective, ADVerb and PROPer Noun. The most
frequent tag is NOUN followed by VERB, PROPN,
ADJ and ADV. We observe that the frequency of
PoS-tags is almost constant over time (excluding
PROPN) underlying a stable language style that
is typical for the news domain. We observe a
variable usage of proper nouns, that may be
related to the different types of events narrated over
time that do not depend on a particular language
style. Moreover, after the 1976, we observe a
complementary trend between the adjectives and
adverbs frequencies: the former slightly increase
over time, while the latter decrease. This may
denote a change in the language style that has varied
to prefer the usage of adjectives over adverbs in
more contemporary writing.</p>
      <p>An interesting analysis concerns the tokens
occurrences per year, whose result is plotted in
Figure 2. We observe a low number of occurrences in
the period (1948-1970), probably due to two
factors: (1) the first period contains many OCR errors
and noise removed during the cleaning step; (2)
the number of pages of the newspaper increases
over time. The latter may also explain the lower
number of tokens for some of the years, such as
1981, 1995, 2000, 2007-2008, 2014. In
particular, the latest years are characterised by
management issues (e.g. the newspaper liquidation in July
2000) that were reflected in the newspaper format.</p>
      <p>We also compute the time series of normalised
occurrences (frequency) for each token, lemma,
and named entity. All the aforementioned
statistics are distributed in separate files together with
the corpus.</p>
      <p>As an illustrative example of the potential use
of the corpus, in Figure 3 we plot the time series
for two keywords. The first, comunismo
[comunism], is assumed to be pivotal to this corpus due
to the specific role played by the newspaper in
relation to the PCI. The second keyword,
antipolitica [anti-politics], is particularly interesting as it is
14The used tag-set is described here https:
//universaldependencies.org/u/pos/
a term used to describe the current state of the
political life in Italy, characterised by a high level of
distrusts in parties and, more generally, in politics.
The lifespan of comunismo [comunism] appears to
be extremely influenced and characterised by
history. We observe two big spikes in the time series.
The first is around 1962, one of the harshest year
of the Cold War, witnessing the Cuban missile
crisis. The second spike is between 1989 and 1991,
corresponding to the beginning of the worldwide
crisis of the communist movement and the
dissolution of PCI. After 1991, the frequency of the term
constantly decreases. Interestingly, the frequency
for comunismo [comunism] is low between 1968
and 1988, a period of time that witnessed a cultural
hegemony of leftist movements and strong
criticism against the U.S.S.R. On the other hand, we
observe that antipolitica [anti-politics] is a recent
term whose first appearance dates back to 1977.
The word frequency starts to increase slowly from
1999 and it reaches its peak in 2012 with the
unexpected electoral success of the populist 5 Star
Movement at the local elections in May.</p>
      <p>Using the same approach, we plot the time
series for two named entities: PCI and Berlusconi.
We notice that the frequency of PCI start
dropping in 1986, few years before its dissolution in
1991, while the name Berlusconi has a substantial
increase in 1994 when he became the Italian Prime
Minister.</p>
      <p>Finally, we investigate how the vocabulary
changes between two periods: T1 = [1948 1958]
and T2 = [2004 2014]. For each period we
build the vocabulary Vi taking into account only
words that occur at least 10 times. Then, we
compute the differences between the two dictionaries,
V1 n V2 and V2 n V1, and sort the words in
descending order by occurrences. We observe that
the words agrari, imperialisti, mezzadri,
monarchici15 appear frequently in T1 and never appear in
T2, conversely the words euro, centrosinistra,
centrodestra, immigrati16 appear only in T2. A
similar analysis was executed on named entities17 and
shows that Scelba, D.C., PSI, U.R.S.S. are specific
to T1, while Berlusconi, PD, Bush, Obama to T2,
revealing differences in topics and people covered
15In English: agrarians, imperialists, sharecroppers,
monarchists.</p>
      <p>16In English: euro, centre-left politics, centre-right
politics, immigrants.</p>
      <p>17In this case we consider only entities that appear at least
5 times.
In this paper, we describe an Italian diachronic
corpus based on the newspaper “L’Unita`”. The
corpus spans 67 years (1948-2014) and is provided
both in plain text and in an annotated format that
includes PoS-tags, lemmas, named entities, and
syntactic dependencies. We compute some
statistics and time series for each token, lemma and
named entity. We think that the corpus and the
precomputed data are a valuable source of
information both for linguists and researchers interested
in diachronic analysis of the Italian language, and
for historians, political scientists, and journalists
as a digital resource enriched with automatic text
analysis technologies.</p>
      <p>However, the corpus has some issues that we
plan to fix in the future, such as OCR errors and
logical document structure recognition. We also
plan to process the corpus by exploiting other
Italian NLP pipelines in order to understand the
differences between the output of different tools.
Finally, we are working on generating and
making available temporal word embeddings for each
year.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Livio</given-names>
            <surname>Gaeta</surname>
          </string-name>
          , Iacobini Claudio, Ricca Davide, Angster Marco, De Rosa Aurelio, and
          <string-name>
            <given-names>Schirato</given-names>
            <surname>Giovanna</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Midia: a balanced diachronic corpus of italian</article-title>
          .
          <source>In 21st International Conference on Historical Linguistics.</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>William L. Hamilton</surname>
            , Jure Leskovec, and
            <given-names>Dan</given-names>
          </string-name>
          <string-name>
            <surname>Jurafsky</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Diachronic word embeddings reveal statistical laws of semantic change. In 54th Annual Meeting of the Association for Computational Linguistics</article-title>
          ,
          <article-title>ACL 2016 - Long Papers</article-title>
          , volume
          <volume>3</volume>
          , pages
          <fpage>1489</fpage>
          -
          <lpage>1501</lpage>
          , may.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Verena</given-names>
            <surname>Lyding</surname>
          </string-name>
          , Egon Stemle, Claudia Borghetti, Marco Brunello, Sara Castagnoli, Felice Dell'Orletta, Henrik Dittmann, Alessandro Lenci, and
          <string-name>
            <given-names>Vito</given-names>
            <surname>Pirrelli</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>The paisa'corpus of italian web texts</article-title>
          .
          <source>In 9th Web as Corpus Workshop (WaC-9)@ EACL</source>
          <year>2014</year>
          , pages
          <fpage>36</fpage>
          -
          <lpage>43</lpage>
          . EACL (
          <article-title>European chapter of the Association for Computational Linguistics)</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Corinna</given-names>
            <surname>Onelli</surname>
          </string-name>
          , Domenico Proietti, Corrado Seidenari, and
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Tamburini</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>The DiaCORIS project: a diachronic corpus of written Italian</article-title>
          .
          <source>In Proceedings of the Fifth International Conference on Language Resources and Evaluation (LREC'06)</source>
          , Genoa, Italy, May.
          <source>European Language Resources Association (ELRA).</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Dominik</given-names>
            <surname>Schlechtweg</surname>
          </string-name>
          ,
          <string-name>
            <surname>Barbara</surname>
            <given-names>McGillivray</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Simon</given-names>
            <surname>Hengchen</surname>
          </string-name>
          , Haim Dubossarsky, and
          <string-name>
            <given-names>Nina</given-names>
            <surname>Tahmasebi</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Semeval-2020 task 1: Unsupervised lexical semantic change detection</article-title>
          . arXiv preprint arXiv:
          <year>2007</year>
          .11464.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Nina</given-names>
            <surname>Tahmasebi</surname>
          </string-name>
          , Lars Borin, and
          <string-name>
            <given-names>Adam</given-names>
            <surname>Jatowt</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Survey of Computational Approaches to Lexical Semantic Change</article-title>
          . 1st International Workshop on Computational Approaches to Historical Language Change
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Xuri</given-names>
            <surname>Tang</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>A state-of-the-art of semantic change computation</article-title>
          .
          <source>Natural Language Engineering</source>
          ,
          <volume>24</volume>
          (
          <issue>5</issue>
          ):
          <fpage>649</fpage>
          -
          <lpage>676</lpage>
          , sep.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Sara</given-names>
            <surname>Tonelli</surname>
          </string-name>
          , Rachele Sprugnoli, Giovanni Moretti, and Fondazione Bruno Kessler.
          <year>2019</year>
          .
          <article-title>Prendo la parola in questo consesso mondiale: A multi-genre 20th century corpus in the political domain</article-title>
          . In CLiC-it.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>