<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Developing a Technology Allowing (Semi-) Automatic Interpretative Transcription</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Daniela Gîfu</string-name>
          <email>daniela.gifu@iit.academiaromana-is.ro</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mihaela Onofrei</string-name>
          <email>mihaela.onofrei@iit.academiaromana-is.ro</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Faculty of Computer Science, University ―Alexandru Ioan Cuza‖</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Institute of Computer Science</institution>
          ,
          <addr-line>Romanian Academy - Iasi Branch</addr-line>
          ,
          <country country="RO">Romania</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper responds to the great interest to humanities researchers who are concerned with the study of the Romanian language in its diachronic evolution: developing a set of tools allowing (semi-)automatic interpretative transcription of scanned Romanian documents written in Cyrillic, in print as well as manuscript forms. The corpus contains old data, belonging to the 19th20th centuries, in order to develop an automatic recognition and interpretative transcription of Romanian historical newspapers from Cyrillic (Cy) into Latin (La), in both manuscript and printed forms. We think that the present study will have an important impact the humanities research, including that of paleography, history, archaeology and that field of linguistics interested in the study of the language in diachrony, but it will also help the researchers in the field of computational linguistics that develops models for old language, in order to elaborate a diachronic POS tagger so necessary to recover old lemmata.</p>
      </abstract>
      <kwd-group>
        <kwd>diachronic corpus</kwd>
        <kwd>transliteration</kwd>
        <kwd>interpretative transcription</kwd>
        <kwd>technology for old language analysis</kwd>
        <kwd>statistics</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>It is well known that the operation of interpretative transcription of texts written in Cyrillic is
extremely laborious, but it will solve a problem of great interest to humanities researchers who
are concerned with the study of the Romanian language (including also Bessarabia texts) in its
diachronic evolution.</p>
      <p>From the perspective of Digital Humanities (history, paleo-linguistics, to name only a few
disciplines), this study is innovative because it will open a huge field of research, making
feasible automatic indexing and online content-based search in collections of old Romanian
documents. The transcripts produced by the machine will be complemented by linguistic
annotations, such that the researcher will be able to seize simultaneously: the original
CyrillicRomanian script, its Latin alphabet transcription, annotated elements as in modern language
dictionaries (tokens and lemmas), and even elements of grammar, such as syntactic structures,
etc.</p>
      <p>
        The novelty of this study includes two major components: developing a diachronic corpus,
called RODICA (ROmanian DIachonic Corpus with Annotations)1, still in its infancy (here,
approximatively 4.5 million words) and defining a method to implement a set of tools allowing
(semi-)automatic interpretative transcription of scanned Romanian documents written in
Cyrillic [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], using the part of our corpus that also contains texts written in Cyrillic and transliterated
in Latin (see Table 2), using the transcription rules described in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Note that the team at
Institute of Mathematics and Computer Science from Chişinău succeeded tolizefotrramnsacription
rules over the standards approved by national authority in Republic of Moldova and Romania.
[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>Research will focus on the automatic transliteration of Cyrillic Romanian texts belonging to
the 19th-20th centuries, from journalistic genre (written in the Cyrillic alphabet and
transliterated in the Latin alphabet). In order to define a methodology to investigate Romanian old
language, we consider that the corpus RODICA responds well. It is sufficient to allow an
analytical demarche that aims to identify the deviations from the norm that occur in a language, in
epochs that are themselves automatically identified statistically.</p>
      <p>This study is focused on semi-automatic interpretative transcription according to the way
Romanian language written in Cyrillic could be conserved. In this paper, the main objective is
the creation of an electronic corpus of old Romanian texts written with the Cyrillic alphabet,
from the 19th to the 20th centuries, belonging to journalistic genre in both manuscript and
printed types in order to develop an automatic recognition and interpretative transcription of
Romanian language tool from Cyrillic into Latin. As training data in the recognition process,
we will use 60% of our corpus in Cyrillic alphabet.</p>
      <p>The rest of the paper will be organized as follows; in section 2, we mention a few works
related to resources and tools related to the analysis of the old language. In section 3, we will
state the problem, present our methodology for the developing an automatic recognition and
interpretative transcription of Romanian historical heritage writings from Cyrillic into Latin,
using the corpus called RODICA. Finally, the paper contains some conclusive statements and
suggestions for future work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Previous Work</title>
      <p>The future of a language depends on early exposure and on a large number of people who has
access to it.</p>
      <p>
        A technology as the one proposed in this project has never been created before for
Romanian old texts. Optical character recognition (OCR) has made remarkable progress in the last
decade, current systems almost reaching the performance of human readers who are ignorant on
the target language. In particular, the character set of the Latin alphabet is recognized at high
rates, which decrease in the case of other alphabets (among them – the Cyrillic one [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]), when
umlauts appear or when the alphabets are not standard. For Latin scripts, Holley [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] reports
accuracy for recognition of printed 19th and early 20th-century characters in the range 81% to
99%.
      </p>
      <p>Recognition becomes much more problematic in the case of handwritten characters, with
their quasi-infinite diversity of forms, being the next phase in our research. Especially
problem</p>
      <sec id="sec-2-1">
        <title>1 http://profs.info.uaic.ro/~daniela.gifu/LR/</title>
        <p>
          atic is the Cyrillic handwriting recognition. Promising results have been obtained recently in
recognizing isolated characters and cursives [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. The major recognition difficulty in the case of
continuous writing is the fact that traditional methods require pre-segmentation of data prior to
the classification process. For this type of recognition, the best results were obtained using
Multidimensional (MD) Long Short-Term Memory (LSTM) type networks [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. The MD-LSTM
networks go through the data set from multiple directions and decide whether, in a meeting
point, a symbol should be issued or not. They learn dependencies in a variable length contextual
window, which gives them greater flexibility when changing the training data set. The model
outlined in [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] implements a multi-layer network that combines recurrent layers with
feedforward layers. A Connectionist Temporal Classification (CTC) type layer makes the decision
on the emission of symbols. The major advantage comes from training the network on images
and direct transcripts, rendering manual segmentation at letter level superfluous.
        </p>
        <p>
          Very few results are known about OCR of Romanian printed with Cyrillic. In [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]
encouraging results obtained by using an Adobe solution followed by the application of a rule-based
transliteration method are reported. As for old Romanian Cyrillic manuscripts, they pose
problems even for human readers and fully automatic recognition is an unattained goal so far. This
is why we envisage an interactive OCR-ing solution, where the expert is in the loop, playing
also a decision-making role. As more fragments of manuscripts will be interpretatively
transcribed, they will be used as training data for innovative DL algorithms with the expected result
that automatic transcription suggestions become more precise.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology</title>
      <p>
        Our method opens a new perspective for the study of our historical heritage, as conveyed by
Cyrillic Romanian, in both manuscript and printed form, by using full text search technology. It
will enable adaptation of modules that perform linguistic processing: segmentation of the text at
the word level (tokenization), morphosyntactic tagging, syntactic parsing, recognition and
classification of proper names, disambiguation of word senses, and others. The modern deep
learning (DL) technologies, based on neural networks, which allowed us [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] to develop basic
language processing tools for more than 50 languages and language variants, will be further
refined and enhanced to deal with old Cyrillic Romanian written texts.
      </p>
      <p>
        An adequate metaphor for this research is a bridge that covers the long way from pixels to
content. Indeed, unstructured grouping of pixels in images representing pages of old
RomanianCyrillic documents will be interpreted and their inner messages deciphered. In order to
acquire a collection of digitised Romanian-Cyrillic resources with their corresponding metadata
and interpretative transcriptions; organise training sets, we will establish the set of norms
(representation formats of intermediary steps in the process of transforming an image of a page
(viewed as a sequence of pixels) into a structured display of textual content. Among these
representation formats, we mention: the metadata describing a source document and XML
representations of the original Cyrillic characters and the final Latin transcribed content (as much as
possible, conforming to TEI [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]). These pieces of content refer to titles, running titles and
inter-titles, text placed in columns and lines, extra-linear writing (characters inserted above
supra - or under-infra-lines, with the indication of their position with respect to the main
elements of the text, marginal additions, literal and Arabic numbers (as for instance, those
indicating verses), etc.
      </p>
      <p>After we acquire an important collection of digitized resources (printed, semi-uncial and
handwritten) containing Romanian language in Cyrillic writings, covering all historical periods,
of various conditions of quality (noise level, uneven characters, etc.), with and without
supralinear writing, we will add metadata to this corpus, which will provide details about: main
language (which must be Romanian), second language(s) (if the document includes words or
passages of text in other languages), year of publication, document script (printed, semi-uncial
writing, or cursive manuscript), document source (typography), author, level and types of noise
(degraded pages, ink stains, creases, dirt, etc.), inclusion or not of supra-linear writing, if there
is any critical edition of the text (with indication of source), etc.</p>
      <p>Also, in order to build the parallel corpus of original page images in Cyrillic and their
interpretative transcriptions in Latin Romanian (UTF-8), we will annotate this corpus with respect to
Elements of Content (EoC) in the layout: characters, words, glued words (scriptio continue),
lines of writing, paragraphs, supra- and infra-linear writing (words/characters placed above and
under lines), border notations, etc. Extract out of this parallel corpus sample images of
characters in context, together with equivalent codes, to be used both in training and in evaluation.</p>
      <p>To improve the quality of original images, to segment images down to elementary EoCs, to
recognize the language different words or sequences of words are written in, to index and
search documents based on their EoCs, and to evaluate the recognition processes, we will
develop or adapt state-of-the-art visual segmentation software to distinguish EoCs in context.
Experiments will be carried out with commercial and open source packages (ABBYY
FineReader, for instance). The segmentation software should allow alignment of the EoCs identified
in the original digital format of the document and their deciphered textual equivalents. These
pointers have a triple role: to link the textual index back into the source document, to support
annotations of the expert users in the original document, and to support their corrections related
to the interpretative transcription.</p>
      <p>We will also train a language recognition system to distinguish among (sequences of) words
those belonging to different languages (often used in old Romanian texts: Romanian, Slavonic,
Hungarian, Greek, Latin, etc.). This will enable to distinguish foreign words in human
translations.</p>
      <sec id="sec-3-1">
        <title>3.1 Romanian Historical Corpus</title>
        <p>
          RODICA is a lexical resource developed based on an important newspapers collection [
          <xref ref-type="bibr" rid="ref11 ref12">11, 12</xref>
          ],
playing a significant aspect in the process of the literary Romanian language modernization,
especially in the 19th century, exemplified and analysed in many studies [
          <xref ref-type="bibr" rid="ref13 ref14 ref15">13, 14, 15</xref>
          ]. This
corpus structured in four historical regions (Bessarabia, Moldavia, Wallachia, and
Transylvania) is statistically described in Table 1.
        </p>
        <p>Importantly, the corpus RODICA represents a first iteration towards building a Romanian
Gold corpus, centred on diachronic meta-annotation, and contains over 4.5 million lexical
tokens in Latin. The punctuation, the words with less than two characters and the number from
the ―Total words‖ have been rNemotoevtehda.t part of this corpus has been transliterated from
Cyrillic to Latin, see Table 2.</p>
        <sec id="sec-3-1-1">
          <title>Province</title>
        </sec>
        <sec id="sec-3-1-2">
          <title>Bessarabia Moldavia Wallachia Transylvania</title>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>Total</title>
        <sec id="sec-3-2-1">
          <title>Period</title>
          <p>1817-2015
1829-2015
1829-2015
1837-2015
Total
words
643084
959010
1372610
1609230
4583934
Total
unique old
words
53029
56790
67050
210180</p>
        </sec>
        <sec id="sec-3-2-2">
          <title>Province</title>
        </sec>
        <sec id="sec-3-2-3">
          <title>Bessarabia Moldavia Wallachia Transylvania</title>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>Total</title>
        <sec id="sec-3-3-1">
          <title>Period</title>
          <p>1817-2015
1829-2015
1829-2015
1837-2015</p>
        </sec>
        <sec id="sec-3-3-2">
          <title>Total</title>
          <p>Chirilic
words
51084
18010
32610
89230
190934
% (Total</p>
          <p>Chirilic
words /Total
words)
7,94
1,88
2,38
5,54
For illustration, the Figure 1 contains some examples of unconventional writing. For instance,
in Figure 1, it can be observed that period does not always mark the end of a sentence, also, it
can be noticed that Arabic numbers are used instead of Romans, that a Cyrillic character must
be transcribed in two Latin letters, depending on the letters preceding that character. In addition,
the capital letters do not always mark the beginning of a phrase or their own name but are often
used without a grammatical explanation. There are also missing letters due to mistakes made by
the scribe.
Ш ш
Ш ш
Ъ ъ
Ь ь
Ѣ ѣ
Ю ю
Ѩ ѩ
Ѥ ѥ
Ѧ ѧ
Ѫ ѫ
Ѯ ѯ
Ѱ ѱ
Ѳ ѳ</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusions and Perspectives</title>
      <p>We consider that this research responds well both for applicative goals (for enabling effective
language chronology analysis using different lexical resources) and for scientific objectives (for
exploring the evolution of journalistic language).</p>
      <p>Automatic transliteration to the current Latin script and added annotation referring to
modern Romanian language are two highly challenging objectives, from both the technological and
linguistic points of views, and will open unprecedented research avenues for Romanian
scientists and not only.</p>
      <p>The success or failure of the study will be estimated according to a combination of the
temporal criteria, genre (journalistic) and printed script criteria, as follows: for each historical
period of 50 years, a random-per-script sample of 30 pages will be considered.</p>
      <p>In the future, we will expand this study for texts belonging to the 16th - 19th centuries, in
order to testing the automatic recognition and interpretative transcription of Romanian historical
heritage writings from Cyrillic into Latin, in printed as well as manuscript forms.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <p>This survey was published with the support of the PN-II-PT-PCCA-2013-4-1878 Partnership
PCCA 2013 grant, having as partners „Alexandurzua‖ IoUanniveCrsity of Iași, SIVECO
Romania, and „Ștefan Cel Mare‖ University of Suceava and of the granta- of the
tional Authority for Scientific Research and Innovation, CNCS/CCCDI – UEFISCDI, project
number PN-III-P2-2.1-BG-2016-0390, within PNCDI III.</p>
      <p>Romanian</p>
      <p>N
Procee</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Onofrei</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gifu</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bolea</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <year>2017</year>
          .
          <article-title>Old Geographical Corpora: a methodology for interpretative transcription at the 9th SpeD 2017, July 6-9</article-title>
          , Bucharest, Romania.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Petic</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Gifu</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <article-title>Transliteration and Alignment of Parallel Texts from Cyrillic to Latin</article-title>
          .
          <source>In: Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14)</source>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Calzolari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Choukri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Declerck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Loftsson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Maegaard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mariani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Moreno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Odijk</surname>
          </string-name>
          , S. Piperidis (eds.),
          <source>European Language Resources Association (ELRA)</source>
          ,
          <fpage>26</fpage>
          -
          <lpage>31</lpage>
          May
          <year>2014</year>
          , Reykjavik (Iceland), pp.
          <fpage>1819</fpage>
          -
          <lpage>1823</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Boian</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cojocaru</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ciubotaru</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Colesnicov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Malahov</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Petic</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2013</year>
          ).
          <article-title>Language Technology and Resources for cultural and historic heritage digitization</article-title>
          .
          <source>In: Proceedings of the 2nd International Conference on Intelligent Information Systems 2013, August 20-23</source>
          ,
          <year>2013</year>
          , Chișinău, Republic oofva,
          <source>Mppo.ld64-73.</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Smith</surname>
            <given-names>R.W.</given-names>
          </string-name>
          (
          <year>2013</year>
          )
          <article-title>History of the Tesseract OCR engine: what worked and what didn't. In Document Recognition</article-title>
          and
          <string-name>
            <surname>Retrieval</surname>
            <given-names>XX</given-names>
          </string-name>
          ,
          <string-name>
            <surname>edited by R. Zanibbi</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <article-title>dC-oüasnon, ings of SPIE-IS&amp;T Electronic Imaging</article-title>
          , SPIE Vol.
          <volume>8658</volume>
          ., doi:10.1117/12.2010051.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Holley</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          (
          <year>2009</year>
          ).
          <article-title>How Good Can It Get? Analysing and Improving OCR Accuracy in Large Scale Historic Newspaper Digitisation Programs</article-title>
          .
          <string-name>
            <surname>D-Lib Magazine</surname>
          </string-name>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Cireșan</surname>
            ,
            <given-names>D.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Meier</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          , GambardMel.l,aandL.Schmidhuber,
          <string-name>
            <surname>J.</surname>
          </string-name>
          (
          <year>2011</year>
          ).
          <article-title>Convolutional Neural Network Committees for Handwritten Character Classification, 11th</article-title>
          <source>Conference ICDAR</source>
          <year>2011</year>
          , Beijing, China.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Graves</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Schmidhuber</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2009</year>
          ).
          <article-title>Offline handwriting recognition with multidimensional recurrent neural networks</article-title>
          .
          <source>In Adv. in Neural Inform. Process. Systems</source>
          (pp.
          <fpage>545</fpage>
          -
          <lpage>552</lpage>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Ciubotaru</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cojocaru</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Colesnicov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Demidov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Malahova</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>Regeneration of Cultural Heritage: Problems Related to Moldavian Cyrillic Alphabet</article-title>
          ,
          <source>in Proceedings of the 11th International Conference ―Linguistic Resources and Romanian Language‖</source>
          ,
          <string-name>
            <surname>Iași</surname>
          </string-name>
          ,-272N6ov., p.
          <fpage>177</fpage>
          -
          <lpage>184</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Boroș</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Dumitrescu</surname>
          </string-name>
          , Ș.D. (
          <year>2017</year>
          ).
          <article-title>A Convolutional Approach to s-Multiword sion Detection Based on Unsupervised Distributed Word Representations and Taskdriven Embedding of Lexical Features</article-title>
          .
          <source>In The 18th International Conference on Engineering Applications of Neural Networks (EANN</source>
          <year>2017</year>
          ). Athens, Greece,
          <year>August</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Ide</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <article-title>Corpus Encoding Standard: Document CES 1, version 1</article-title>
          .4,
          <string-name>
            <surname>October</surname>
          </string-name>
          . http://www.cs. vassar.edu/CES/,
          <year>1996</year>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Gîfu</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <year>2017</year>
          .
          <article-title>Recovering Old Romanian Lemm1a3tath</article-title>
          ,
          <source>Inatternathtieonal Scientific Conference eLearning and Software for Education, ELSE</source>
          , Bucharest, April 27-
          <issue>28</issue>
          ,
          <year>2017</year>
          .
          <source>In: Proceedings of eLSE</source>
          <year>2017</year>
          , Ion Roceanu (ed.),
          <string-name>
            <surname>Carol I NDU Publishing</surname>
          </string-name>
          <article-title>House</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Gîfu</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <year>2016</year>
          . Lexical Semantics in Text Processing. ContraisctiSvetudiDesiaocnhron Romanian Language,
          <source>PhD thesis</source>
          , ―Alexandru Ioan Cuza‖ University of Iași,
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Diaconescu</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          (
          <year>1974</year>
          ).
          <article-title>Elemente de istorie a limbii romoâdneerne.litPearartreea I.m Probleme de normare a limbii române literare mo-d1e8r8n0e), (1B8u3c0ureşti</article-title>
          , pp-
          <fpage>6</fpage>
          . . 5
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Andriescu</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <year>1979</year>
          .
          <article-title>Limba presei Româneşti în secolul-leal, XEdIX</article-title>
          . Junimea, Iaşi.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Drăgan</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          (
          <year>1996</year>
          ).
          <article-title>Paradigme ale comunicării în masă</article-title>
          , Ed. Șansa, București.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>