<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Global Variants in the Czech Language</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>JaroslavaHlaváčová</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lukáš Kyjánek</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>MagdaŠevčíková</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics Malostranské náměstí 25</institution>
          ,
          <addr-line>118 00 Prague</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>identification. There are words written in several diferent ways in Czechl,aem.gp.i,on ∼ lampión (lampion). This variability may occur in either some inflectional wordforms (inflectional variantsh)r,acdfu. ∼ hradě in the locative case of the nhoruand (castle), or across the inflectional wordforms and derivatives (global variafnatnsta),zcijfn.í ∼ fantasijní in the adjective derived from the nounfantazie ∼ fantasie (fantasy). It is reasonable to distinguish the global variants as diferent words but to have formal means that interconnect them in the Natural Language Processing systems and resources. In this paper, we describe the identification of global variants in the Czech vocabulary and summarise new changes in the MorfFlex CZ dictionary and DeriNet lexicon concerning this type of variants. We reviewed several typical patterns within global variants captured in the available resources and combined a set of regular expressions with manual annotations to achieve the highest precision of the global and inflectional variant, morphology, word derivation, Czech ITAT'22: Information technologies - Applications and Theory, SeptemWorkshop Proce dings CEUR</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>seum ∼ muzeum (museum), peepshow ∼ peepšou ∼ pípšou
(peepshow).
eral slightly diferent ways in so-calolretdhographic
The written form is one of the possible representationnsouonfobchod ∼ vobchod (shop), and to the derived
veorbblanguages. Czech speakers must learn and use a relevcahnotdovat ∼ vobchodovat (to trade), and the derived
adjecscript, rules, and regularities of the respective wrtiitvienogbchodní ∼ vobchodní (commercial). All those words
system because of its substantial standardisation andmacondif-est the same diference in every single wordform of
ification. However, some words can be still written in stevh-eir inflectional paradigms (for instance, genitive cases
(spelling) variants, e.g., citron ∼ citrón (lemon), mu- ního ∼ vobchodního)). On the other hand, somewhere
of the nounob(chodu ∼ vobchodu) and
adjectiveob(chodbetween the two defined types, there are also several
cases in which the variability is limited to a few forms,
the variability in all the inflectional forms, often also
in derived words, e.g., the prothevt-2icattached to the
sources and all sorts of NLP applications.
tation of Czech, language development, and languidaegnetical.</p>
      <p>The emergence of orthographic variants in Czechcfi.sthe infinitive and past participles of themveyrslbet
influenced by various aspects like the spoken represen∼-myslit (to think), while the remaining wordforms are
contact. Some cases of the variability are only tempToh-e inflectional variants are captured in the
morphorary until the use of one of the orthographic varialnotgsiciasl dictionary MorfFlex CZ (hereafter MorfFlex) by
established and codified as the preferred one (which cmaenans of the 15th position in the morphological tag
detake years or decades). However, codified or not, mansycribing morphological categories of a given wordform
of the orthographic variants appear in the texts prod[2u]c. edUntil the 2020 edition of MorfFlex, there was no
by speakers, which complicates work with language dries-tinction between the description of global and
inflectional variants. All the variants were marked at the 15th</p>
      <p>Adhering to the current decisions on annotatingpotshiitsion of the Prague positional ta1g]s.eItn[the last
verphenomena in the corpus PDT-C (cf. its manual1i0n])[,</p>
      <p>sion, the global variants are annotated by means of links
we distinguish two types of orthographic varIinan-tsb.etween them3. In MorfFlex, one word from t h-teuple
flectional variants refer to relatively regular variaonftvsariants is selected as the basic one. All other
variwithin a set of wordforms of a given word, e.g., the laonctas- are linked with the basic one by means of additional
tive case of some masculine inanimate nounsolbikcheod
(shop): obchodu ∼ obchodě. Global variants1 address contains not only the basic variant, but also the (rough)
pieces of information in their lemma. This information
CEUR
htp:/ceur-ws.org
ISN1613-073</p>
      <sec id="sec-1-1">
        <title>1Global variants are also caflulleld-paradigm variants in the [4], [5], [6]).</title>
        <p>complete description of the morphological annotation of the corpus
PDT-C [10, pp. 36–42]. We will stick to the term global, as it is</p>
      </sec>
      <sec id="sec-1-2">
        <title>2This variation originates from the common Czech and is not codified.</title>
      </sec>
      <sec id="sec-1-3">
        <title>3There are more ways how to interconnect the global variants (see</title>
        <p>style of the variant. There is no strict rule for such
selection, as no of the possible criteria is easy to formulatesoinrgular locative case of the lemobmcahod is
obcheck. Usually, the more common (frequent) or standardchodu ∼ obchodě.
(over non-standard) variant was taken as the basic one.</p>
        <p>The final selection of the basic variant depended on th—e Global variants
lexicographer’s opinion. are a pair (or genera ll-ytuple) of lemmas whose</p>
        <p>Until recently, no special attention was given to tdhifeerence in spellings propagate to all wordforms
completeness of global variants within the whole of tohfetheir inflectional paradigms and to most of their
dictionary. That is why we reviewed available resourcedserivationally related words;oeb.cgh.,od ∼ vobchod
addressing the issue of orthographic variability (Se→c-obchodní ∼ vobchodní (apart from the first letter,
tion2), and extracted typical patterns for global vainrflei-ctional paradigms of these words are identical).
ants from them. We applied these patterns to the set of
lemmas from MorfFlex to find as many global variant— Basic variant</p>
        <p>
          is the representative variant fo-rtaunple of global
candidates as possible. We also exploited the DeriNet
lexicon 1[
          <xref ref-type="bibr" rid="ref1">4</xref>
          ], which models word-formation relations invariants.
        </p>
        <p>Czech, to search for global variants within derivationally
related words (Secti3o)n.After manual filtration of the
obtained global variants, our work resulted in inte2r.coAn-vailable Resources Containing
necting global variants in MorfFlex. They will appear inVariants
its next version. They were also partly uploaded to the
DeriNet 2.1 lexicon. Since any researcher working on a lexical resource must</p>
        <p>
          The resulting data and categorisation (Se4ctainodn5s) inevitably process also the orthographic variants, there
have potential in two research directions. First, theyarleeasdome pieces of annotations of them in the existing
lanto the improvement of the Natural Language Procesgsiunagge resources for Czech. However, capturing variants
applications for which the morphological dictionaisrnieost the primary goal in any of the resources, so
systemserve as the background data, c1f1.], [[
          <xref ref-type="bibr" rid="ref5">13</xref>
          ], and [
          <xref ref-type="bibr" rid="ref4">12</xref>
          ]. atic care is more than needed in this kind of annotation.
Second, they contribute to a wider linguistic discusTshioenavailable digitised resources provide a good point of
on (not only orthographic) variability in Czech, especdiaelplayrture for the phenomenon of orthographic variants,
in the context of border cases between inflectional and
        </p>
        <p>but we see the following two
shortcomings/inconsistenglobal variants mentioned above and exemplified befocriees in them. The number of the captured orthographic
the conclusions of this paper. variants is often relatively low in the resources. The
orthographic variants are not treated across derivation, e.g.,
úřad ∼ ouřad (ofice ) but notúřadovat ∼ ouřadovat (to
oficiate ).
1.1. Terminology There are various language resources and grammar
books that include this kind of annotation; however, in
To sum up definitions of the basic terms used in the following paragraphs, we describe only those that
this paper, we provide this section. We frame itare available and digitised, and thus machine-readable.
to facilitate reading in case readers would like toMorfFlex 2.0 [2], the lexicon of the inflectional
moreasily remind some definitions while reading more phology of Czech, in its currently available version,
aladvanced parts of the text. ready includes annotation of global variants. It classifies
them into three types of variansttasn:dard (labelDD,
— Inflectional paradigm e.g., lavor ∼ lavór (pail)), common Czech/non-standard
is a set of all wordforms derived by means of inflec-(labelGC, e.g., oprášit ∼ voprášit (to dust down))
anddistion from a citation wordform (so-callelmemda); tortion/typo (labelDS, e.g., Dominigue ∼ Dominique).
e.g., the inflectional paradigm of the lemombachod However, as the results of the work herein show, we were
consists of wordformobschod, obchodu, obchodě, … able to find many more -tuples of variants not included
in the current version.</p>
        <p>VALLEX 3.0 [9] is the valency lexicon of Czech verbs
— Inflectional variants
are a pair (or genera ll-ytuple) of wordforms
belonging to the same inflectional paradigm of thethat also interconnects and labels several orthographic
variants of verbs. Their representations are systematic,
same lemma and having the same values of all mor-but the definition of being a variant seems diferent from
phological categories, but diferent spellings; e.g.,
ours. For instance, VALLEX marks alcshoytit ∼ chytnout
(to catch) as variants although its status is arguable.</p>
        <p>Slovník spisovného jazyka českého (SSJČ) [3], the the extractions from MorfFlex and VALLEX were easy,
explanatory dictionary of Czech digitised into a saesmti-hese resources are designed to be processed
autostructured form4acto,vers wide vocabulary and includemsatically, we had to use regular expressions to extract
annotations of orthographic variants without anycaandddiid-ates for variants from SSJČ because, in its digitised
tional classification. However, it uses various optiovnerssion, it stores data in a semi-structured file format
in which these variants are underlined in glosses of(steeheSection2). Consequently, we checked whether the
dictionary, e.g., by using words extracted from this resource are attested in the
vocabulary of MorfFlex to mitigate incorrectly extracted,
• v. = viz (see) in “obepsati v. opsati” (copy), non-existent, and archaic words. We also wrote down
ex• comma in “mysliti, mysleti” (to think), or amples and patterns from the relevant linguistic studies
• řidč. = řidčeji (rarely) in “zpěvánka, řidč. zpě- like those cited in the previous section.</p>
        <p>vanka, zpívánka” ([usually a short] song). When comparing the-tuples of variants extracted
from the resources, we have observed that MorfFlex
inExcept for these language resources, there is also a long</p>
        <p>cludes more than two thirds of the variants captured
tradition of linguistic studies on the topic of orthographic</p>
        <p>
          in SSJČ. The remaining third of the variants from SSJČ
variants, cf. prefixess-/se- andz-/ze- and orthography
of loanwords in general 1in5, [pp. 167–170], morpholog- seems questionable, e.gz.,výhodněný ∼? zvýhodnělý
(privical variability and its reflection to orthograp1h6y, inile[ged) in which the variability does not afect the
indipp. 268–276], orthography of loanwords wsi/tzhchar- vidual characters but the afixes (and thus can diverge
acters, such aasnalýza ∼ analýsa (analysis) in [
          <xref ref-type="bibr" rid="ref9">17</xref>
          ], and the word meaning). The extracted variants from VALLEX
terminological issues of the phenomenon7]i.nT[hey cover only verbs; most of the variants are not included
provide extensive lists of variants or patterns thainttahreeother resources.
shared across variants. We extracted some patterns from
the studies, which helped us to identify typical gen3e.r2a.l Formalising Regular Patterns
properties of global variants.
        </p>
        <sec id="sec-1-3-1">
          <title>Having relatively large amoun t-toufples with global</title>
          <p>variants, we considered whether to use pattern-matching
3. Searching for Global Variants algorithms or to formalise frequent patterns manually.</p>
          <p>We chose the latter way as it allowed us to have a better
We developed a semi-automatic procedure that consoisvtesrview of the processed data and to create more
comof four straightforward subsequent steps to searcphlicfaotred patterns that would take into account not only
Czech global variants. We exploited the availablcehraer-acter changes but also morpho-syntactic categories
sources mentioned in the previous section, and we oefx-words. For instance, this decision allowed us to avoid
tracted frequent patterns that appear in global vainritaenrtcso.nnection of the masculine animate variant pair
We applied these patterns to the set of lemmas from Mčoesrafč- ∼ česáč (a man harvesting apples) to the masculine
Flex 2.0 to obtain-tuples of global variants, sucehx-as inanimate variant pačiersač ∼ česáč (an instrument for
tremismus ∼ extrémismus ∼ extremizmus ∼ extrémizmus harvesting fruits of tall trees).
(extremism). We exploited various types of intersections of the
ex</p>
          <p>To get also derivationally related global variantsratchtaetd lists of var ia-ntutples and their sorting to be
were not identified on the basis of the extracted pattearbnles,to infer frequent patterns that occur in global
variwe included DeriNet into the process. Thanks to thatanwtes. We first looked at the global variants extracted
covered also-tuples likextremistický ∼ extrémistický from all three resources, then at those that occurred in
(extremist) for the already identified variants of the batseleast two resources, and only then at those that were
wordextremism mentioned above. The result in-gtuples in individual resources but not in the others. More than
were manually annotated to eliminate randomly simoinlaerhundred observed patterns were formalised into the
words. In the last step, the data were uploaded as afnoerwm of regular expressions, e∧.go..* ↔ ∧vo.* in
obtype of annotation into DeriNet, and, in parallel,chnoedw∼ vobchod (shop). The relevant morpho-syntactic
links were also added to MorfFlex. categories were also stored with the particular regular
expression.
3.1. Extracting Variants from the Existing</p>
          <p>Resources</p>
          <p>3.3. Applying Patterns to MorfFlex
We started with assembling the existing resources liWsteeadpplied the formalised patterns to all the lemmas from
in Section2 and extracting variants from them. WhMiloerfFlex (we did not search for inflectional variants, so
we did not have to take wordforms into account). To
citronový.ADJ</p>
          <p>lemony
citronově.ADV
lemonily
citron.NOUN
lemon
citroník.NOUN
little lemon
citronóvý.ADJ</p>
          <p>lemony
citrónově.ADV
lemonily
citrón.NOUN</p>
          <p>lemon
citróník.NOUN
little lemon
citron.NOUN</p>
          <p>lemon
citronový.ADJ</p>
          <p>lemony
citronóvý.ADJ</p>
          <p>lemony
citronově.ADV</p>
          <p>lemonily
citrónóvě.ADV
lemonily
citrón.NOUN
lemon
citroník.NOUN
little lemon
citróník.NOUN
little lemon
 -tuples of global variant candidates; the task wastthoirdgcoolumn). The first column lists sizes of  -tuples —  = 2
through the lists of variants and exclud e -tthuopslees is for pairs.
which only accidentally met a derivational pattern, but
that were not variants, e.g., thefiaplaair(wallflower ) ≁
same pattern works well in real variantnselainkdertalec
∼ neandrtálec (Neanderthal). During the manual worklemmas.
ifála (pinnacle) had to be excluded manually, although tBhoeth resources difer in the data structures they use for
storing their data, but they both share the same set of
tional variants with variant lemma — see the examplientoof so-calledderivational families. Each family of
on the global variant candidates we also identified inflecD-eriNet interconnects derivationally related words
the pair of verbmsyslit, myslet.
pairs ( = 2 ), triples (= 3 ) and4-tuples. The bigger ← učitel (teacher) ← učit (to teach).
notated in the MorfFlex 2.0, and how many were addaesdnodes while derivational relations as edges. In other
thanks to the new foun-dtuples. It is visible, that thweords, each derived word in DeriNet has at maximum
main increment is recorded for small,erspecially for one base word (antecedent), eu.gč.i,telka (female teacher)
words is represented in a form of rooted tree (in graph
values of remain the same.
analysis of the prototypical cases of global variantts h(Seeca-djectivceitrónový (related to lemon) could be
contion4) and, on the other hand, cases that we do not tnreecatted to the noucintron, although the noucintrón would
In the following sections, we present a more detavilaerdiants caused structural inconsistencies. For instance,</p>
        </sec>
        <sec id="sec-1-3-2">
          <title>In the rooted tree data structure, unidentified global as variants (Sectio5n).</title>
          <p>3.4. Global Variants into DeriNet
be a better antecedent (both global varialenmtosno).fTo
tackle this issue, identifying global variants is crucial.</p>
        </sec>
        <sec id="sec-1-3-3">
          <title>We considered two possible ways of representing</title>
          <p>global variants in the current rooted tree data structure
We uploaded the resulting global variants into the noefwDesetriNet. In the first approach (see F1ig,p.art A), the
version of DeriNet 2.1, and we intend to do so alsogfloorbal variants would create parallel branches in the tree,
the next version of the inflectional dictionary MorfeF.lge.,xc.itron → citronový → citronově parallel tcoitrón →
citrónový → citrónově. The major disadvantage of thi4s. Prototypical Cases of Global
approach is that the branches may be disconnected orVariants
contain gaps if any variant is missing in the vocabulary.</p>
          <p>In the second approach (see Fi1g, .part B), the globalIn this section, we will present the most common types
variants would be connected toboanseic variant to of global variants together with typical ex6amOpnles.
which the derivatives are connected, while the ootfhtehre important properties of global variants is that their
variants would not have any derivatives connectedd,eer.gi.v, atives can also become global variants.
[ citron ∼ citrón ] → [ citronový ∼ citrónový ] → [ citronově Example: The pair of verblsítat ∼ létat (to fly ) derives
∼ citrónově ]. The lack of global variants in this approaitcehrative verblístávat ∼ létávat, adjectiveslítající ∼
lédoes not disconnect word(s) from the tree. Thereforet,awjíceí (flying ), and/or verbal nou nlístání ∼ létání (the
chose this approach for DeriNet. lfying ). Derivatives in each of the pairs are also global</p>
          <p>The selection of the basic variant followed the sivmailraiarnts.
criteria that were applied in MorfFlex. We tried to do
so consistently across th-teuples that share the same
pattern. The final decision depended on a lexicogra5phe4r..1. Long and Short Vowels</p>
          <p>As a result, the data from our experiments with glIonbatlhis type of variants, words vary in the length of a
variants has been already uploaded into DeriNet 2.v1o. wIfel, either in the afix, or in the root.
words are variants in this lexicon, one of the worEdxsaimsple: Sufix variation in svíčkař ∼ svíčkář (someone
selected as the basic one and the other ones are wchoonm-akes candles), and root variationkivnikat ∼ kvíkat
nected directly to it by special relation that is labe(ltloeodinaks/squeak).</p>
          <p>Type=Variant. Fig. 2 illustrates words derived from the
variant paiúrřad ∼ ouřad (ofice ) from DeriNet 2.1; the
missing variants of the individual derivatives, such as
úřadek ∼ ouřadek, will be connected in the new release.</p>
          <p>DeriNet projects but we plan to make a unification.
5Unfortunately, this task was not coordinated between MorfFl6eTxhainsdoverview is by no means complete.
4.7. Variants of Foreign Names</p>
        </sec>
        <sec id="sec-1-3-4">
          <title>Most frequent foreign geographic names have usually a</title>
          <p>The consonants alternate in the root; the instancCezseacrhetranslation.
of diferent origins. Example: The Czech variant oPafris is Paříž, Moscow is
Example: vlaštovka ∼ vlašťovka (a swallow), student ∼ Moskva, Berlin is Berlín.
študent (student), mrazený ∼ mražený (frozen). Though both words can appear in Czech texts, they are
not considered global variants. Moreover, the original of
4.3. Soft and Hard Adjectives the foreign name is usually not inflected.</p>
          <p>Person names are typically not translated, but their
There are two types of adjectives — soft and hard, busptelling is often unusual. In addition, errors or typos
some of them can vary between the two types. This wfarsequently occur in their spelling. In such cases, they can
quite common in the past, as is visible from the additiobneaclonsidered variants. Sometimes, one of the variants
information attached usually to one of the variantsi—saitsipselling adapted to the pronunciation, as the long
often archaic or outdated. The basic variant can be soft avsariant in the following example.
well as hard, depending on the lexicographer’s decisEioxna.mple: Abdulah ∼ Abdullah ∼ Abduláh.
At the beginning of our work, this type of variants wTahsis is not applied to Slavic names with the en-dijing
not recorded. or-i which are sometimes translated with the en-ýd.ing
Example: Adjectival variation in the pnaáimrsezdný ∼ As the variation appears only in the nominative singular
námezdní (hired), přívodný ∼ přívodní (feed, inflow ... e.g. (lemma) and vocative singular, we consider this type of
pipe). variants as inflectional.</p>
          <p>Example: All the three variants of the nČaamjkeovský
4.4. Prothetic v- (Tchaikovsky), namely Čajkovský ∼ Čajkovskij ∼
Čajkovski, are inflectional variants of the singular
nominaMany Czech words starting with the vowexeilst also in tive and vocative cases. Other cases do not manifest this
the variation with the prothv-etaitctheir very beginningt. ype of variation. They are not global variants.
Though the latter variant is considered non-standard Sainmdilarly, names of ancient Greeks with the lemma
is used mainly in spoken Czech, it is very common andending-es or-és are not global variants, as this variation
penetrating into the written Czech, too. appears only in nominative singular. They are
inflecExample: okno ∼ vokno (window). This type of variantstional variants.
can appear not only at the beginning of words, but Eaxlsaomple: Empedokles, Empedoklés.
after a prefix that precedes theo; e.g., zotvírat ∼ zvotvírat
(to open step by step).</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>5. Non-variants</title>
      <p>4.5. Vocalized and Non-vocalised Prefixes The soft–hard type can seemingly be applied to soft and
The prefixes v-, s-, vz-, roz-, od-, pod-, nad-, ob-, před- can hard declension of nouns with feminine or masculine
be expanded bye (ve-, se-, vze-, roze-, ode-, pode-, nade-, gender. In reality, in such cases, we should rather speak
obe-, přede-). Nevertheless, some words can have botahbout a combined paradigm and merge the two variants
into one inflectional paradigm. This has been already
spellings, which makes them variants.</p>
      <p>Example: střást ∼ setřást (shake of ), rozsmutnit ∼ rozes- done for masculine declension of soft–hard pairs, both
mutnit (make sad), objet ∼ obejet (go around). animate and inanimate.</p>
      <p>Example: The lemma kužel (cone) can be inflected
either as a hard noun (following the traditional masculine
4.6. Stylistic Variants (ú ∼ ou, ý ∼ ej, th ∼ t, inanimate declension clahsrsad) as well as a soft noun
s ∼ z ) (following the traditional masculine inanimate
declension classtroj). It is reasonable to join wordforms of the
This type of variants usually puts into opposition
stan</p>
      <p>two inflected sets and to represent the whole set of
worddard and non-standard Czech, let it be archaic, colloquial</p>
      <p>forms by a single lemma. As there is only one lemma,
or other sort of style. The most frequent is the variation
betweens andz, especially within the sufixes-ismus and these words cannot be global variants either.
-izmus. The feminine gender is diferent, as there the lemmas
Example: mechanismus ∼ mechanizmus (mechanism), difer. However, the diference is always within the
endvytékat ∼ vytejkat (flow/leak out ), úzký ∼ ouzký (narrow), ing, so according to the definition of global variants, they
ortopedie ∼ orthopedie (orthopedics). are not global variants. Though they are often viewed as
global variants, there is rather one inflectional paradigm
with inflectional variants afecting all the wordforms.</p>
      <p>The new pattern for this type of variation should Tbheis project included lots of manual work, as the topic
added and all the wordforms merged into a single inoflefc-variants is very variable and there are no rules for
tional paradigm with inflectional variants even fortthheereally strict distinction of what are and what are not
lemma. variants. Thus, the manual work discovered some border
Example: The lemmas kapuce ∼ kapuca (hood) have cases where it had to be decided from scratch. In general,
diferent inflectional paradigms, but the individual tawges adopted very strict rules, e.g. we do not consider
vari(combinations of number and grammatical case) difearnts those words that contain formally diferent afixes
only in endings. (see the examplezvýhodněný ≁ zvýhodnělý (privileged)).</p>
      <p>Similar cases are variants with diferent genders. All those cases are to be researched in greater detail in
Example: brambora (fem.), brambor (masc. inan.) (both the future.
potato); ribstole (fem.), ribstol (masc. inan.) (bothwall
bars).</p>
      <p>Again, the variation manifests itself only in endingAs, csoknowledgement
they cannot be considered global variants. The solution</p>
      <p>This work was supported by the Grant No. GA19-14534S
proposed for the nouns with the same gender
(merging the inflectional paradigms) cannot be applied heoref, the Czech Science Foundation, and the Grant No.</p>
      <p>START/HUM/010 of Grant schemes at Charles University
because of the so-call“ePdrinciple of morphological difer- (reg. No. CZ.02.2.69/0.0/0.0/19_073/0016935), and
LINentiation” introduced in10[]. One of its requirements is
that the gender of a noun should stay the same wiDtAhTin/CLARIAH-CZ project of the Ministry of Education
(LM2015071, LM2018101).
the whole inflectional paradigm. These examples reveal
that the Principle is questionable; it would probably be
advisable to reconsider it. References</p>
      <p>During the work on global variants in Czech resources,
we came across several peculiarities. [1] Hajič, J. 2004. Disambiguation of Rich Inflection
Example: The pairpécéčko ∼? písíčko (a sort of abbre- (Computational Morphology of Czech).
Nakladatelviation opfersonal computer / PC). Is it a pair of global ství Karolinum, Charles University, Czechia.
variants, or not? For the time being, the two lemmas[2a]reHajič, J.; Hlaváčová, J.; Mikulová, M.; Straka,
not interlinked. M.; Štěpánková, B. 2020. MorfFlex CZ 2.0.</p>
      <p>
        Sometimes, we found sets of seeming variants, that LINDAT/CLARIAH-CZ digital library at the Institute
had a typical variant pattern, but they were not varianotfs Formal and Applied Linguistics (ÚFAL), Faculty
because of diferent meaning. of Mathematics and Physics, Charles University,
Example: valečka (biol. sort of grass) ≁ válečka (someone Czechia. URL:http://hdl.handle.net/11234/1-3.185
[fem.] who rolls something), and/orstudenský (adjective [3] Havránek, B. (ed.) 1960–1971. Slovník spisovného
to the town of Studená) ≁ studénský (adjective to the town jazyka českého. Academia, Prague, Czechia.
of Studénka). [
        <xref ref-type="bibr" rid="ref1">4</xref>
        ] Hlaváčová, J. 2009. Formalizace systému české
      </p>
      <p>Neither we interconnected the onomatopoeic or ex-morfologie s ohledem na automatické zpracování
pressive words. českých textů. Ph.D. thesis, FF UK, 146 pp.
Example: ďoubnout ≁ ďubnout (expr. to push). [5] Hlaváčová, J. 2011 Problém variantních tvarů slov při
automatickém zpracování jazyka. In: Information
Technologies – Applications and Theory, pp. 75-78.
6. Conclusion [6] Hlaváčová J. 2019. Aggregates and Variants in Two
Czech Morphological Approaches. In: Proceedings
The paper presented the specialised project of looking of the 19th Conference ITAT 2019: Slovenskočeský
for global variants in available resources of Czech lexicNalLP workshop (SloNLP 2019), pp. 120-124.
data. The main aim was to make an “inventory” of Cze[c7h] Hrbáček, J. 1974. Lexikální ekvivalenty, dublety a
global variants and to annotate them. Special attentivoanrianty Naše řeč 57(1), pp. 28–33.
was paid to the distinction between the global and in[8fle]c-Křen, Michal et al. 2016. Corpus SYN, version 4.
tional ones. This distinction has been already capturedPrague, Institute of the Czech National Corpus,
Facin the new edition of MorfFlex 2.0, in which many pairs ulty of Arts, Charles Univershityt;p://www.korpus.
of global variants still remained unlinked. In particulacrz.,
it was necessary to reflect the existence of global var[i9a]ntLopatková, M.; Kettnerová, V.; Bejček, E.;
 -tuples in DeriNet. The newly identified global variants Vernerová, A.; Žaborktský, Z. 2016. VALLEX 3.0.
now are captured in the recent edition of DeriNet 2.1L. INDAT/CLARIAH-CZ digital library at the Institute
Comprising of the new annotation of global variants intoof Formal and Applied Linguistics (ÚFAL), Faculty
MorfFlex is planned for a future edition.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          4.2. Alveolar vs.
          <source>Postalveolar/Palatal Consonants of Mathematics and Physics</source>
          , Charles University, Czechia. URL:http://hdl.handle.net/11234/1-2.
          <fpage>307</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Mikulová</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Hajič</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Hana</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ; Hanová,
          <string-name>
            <given-names>H.</given-names>
            ;
            <surname>Hlaváčová</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ;
            <surname>Jeřábek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ;
            <surname>Štěpánková</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.; Vidová</given-names>
            <surname>Hladká</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ;
            <surname>Zeman</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          <year>2020</year>
          .
          <article-title>Manual for Morphological Annotation</article-title>
          . Revision for Prague Dependency Treebank -
          <article-title>Consolidated 2020 release</article-title>
          .
          <source>Technical Report TR-2020-64. Institute of Formal and Applied Linguistics (ÚFAL)</source>
          ,
          <source>Faculty of Mathematics and Physics</source>
          , Charles University, Czechia. ISSN:
          <fpage>1214</fpage>
          -
          <lpage>5521</lpage>
          . URL: https://ufal.mff.cuni.cz/techrep/tr6.4.pdf
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Richter</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Straňák</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Rosen</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>2012</year>
          .
          <article-title>Korektor - A System for Contextual Spell-checking and Diacritics Completion</article-title>
          .
          <source>In: Proceedings of the 24th International Conference on Computational Linguistics (Coling</source>
          <year>2012</year>
          ), pp.
          <fpage>1</fpage>
          -
          <lpage>12</lpage>
          .
          <article-title>Coling 2012 Organizing Committee</article-title>
          , Mumbai, India.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Straka</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Straková</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <year>2017</year>
          . Tokenizing, POS Tagging,
          <article-title>Lemmatizing and Parsing UD 2.0 with UDPipe</article-title>
          .
          <source>In: Proceedings of the CoNLL 2017 Shared Task: Multilingual Parsing from Raw</source>
          Text to Universal Dependencies, pp.
          <fpage>88</fpage>
          -
          <lpage>99</lpage>
          . Association for Computational Linguistics, Vancouver, Canada.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Straková</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ; Straka,
          <string-name>
            <given-names>M.</given-names>
            ;
            <surname>Hajič</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          <year>2014</year>
          .
          <article-title>Open-Source Tools for Morphology, Lemmatization, POS Tagging and Named Entity Recognition</article-title>
          .
          <source>In: Proceedings of 52nd Annual Meeting of the Association for Computational Linguistics: System Demonstrations</source>
          , pp.
          <fpage>13</fpage>
          -
          <lpage>18</lpage>
          . Association for Computational Linguistics, Baltimore, Maryland.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Vidra</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Žabokrtský</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Kyjánek</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Ševčíková</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ; Dohnalová, Š.;
          <string-name>
            <surname>Svoboda</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Bodnár</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <year>2021</year>
          .
          <article-title>DeriNet 2.1. LINDAT/CLARIAH-CZ digital library at the Institute of Formal and Applied Linguistics (ÚFAL)</article-title>
          ,
          <source>Faculty of Mathematics and Physics</source>
          , Charles University, Czechia. URL:http://hdl.handle.net/11234/1-3.
          <fpage>765</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [15]
          <article-title>Mluvnice češtiny 1: Fonetika, Fonologie, Morfonologie a morfemika, Tvoření slov</article-title>
          .
          <year>1986</year>
          .
          <article-title>Academia, nakladatelství Československé Akademie věd</article-title>
          , Prague, Czechia.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <source>[16] Mluvnice češtiny 2: Tvarosloví</source>
          .
          <year>1986</year>
          .
          <article-title>Academia, nakladatelství Československé Akademie věd</article-title>
          , Prague, Czechia.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [17]
          <article-title>Pravopis a výslovnost přejatých slov se s - z. Internetová jazyková příručka [online] (</article-title>
          <year>2008</year>
          -
          <fpage>2022</fpage>
          ).
          <article-title>Ústav pro jazyk český AV ČR</article-title>
          , Prague, Czechia. Cit.
          <volume>28</volume>
          . 5. 
          <year>2022</year>
          . URL: https://prirucka.ujc.cas.cz
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>