<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Design of the User's Interface of Virtual Lexicographic Laboratory for Explanatory Dictionary of the Spanish Language</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>n Kuprii</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mykyt</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Ukrainian Lingua-Information Fund, NAS of Ukraine</institution>
          ,
          <addr-line>03039 Kyiv</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University “Kharkiv Polytechnic In</institution>
          ,
          <addr-line>st6i1tu0t0e2” Kharkiv</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>One of the most effective tools to work with dictionaries in digital environment is virtual lexicographic laboratories (VLL). Unlike electronic dictionaries, they are intended mostly for professional linguists. The paper shares the authors' experience in elaborating the ofinthteerfvaicretual lexicographic laboratory for Explanatory dictionary of the Spanish language (DLE 23). Using the theory of lexicographic systems a formal model of DLE 23 was elaborated. On the basis of the model the database structure and VLL interface elements were defined. The current version of VLL DLE 23 interface has the following advantages: 1) making an inventory of language units in the dictionary or in a sample; 2) conducting DLE 23-based linguistic researches to reveal lexicalsemantic, etymological, grammatical and usage properties of the Spanish language; and 3) building of secondary lexicographic objects or sub-dictionaries on the basis of DLE 23, for example: sub-dictionary of morphemes, homonyms, collocations, etc.</p>
      </abstract>
      <kwd-group>
        <kwd>User's Interfac</kwd>
        <kwd>eVirtual Lexicographic Laboratories</kwd>
        <kwd>Explanatory Dictionary</kwd>
        <kwd>Digital Environment</kwd>
        <kwd>Interface Design</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 National Technical</title>
      <p>
        For the last two decades, lexicography has undergone significant changes. They relate
to the implementation of a wide range of approaches to comprehensive analysis of
language vocabulary, creation of integrated lexicographic systems combining
different linguistic facts by nature, construction of universal digital lexicographic
environments, etc. In this regard, there is a growing interest in the needs and skills of digital
dictionary users, which encourages the developers to focus on the user properties of
lexicography objects. This interest has led to the fact that most experts now
understand the necessity of creating dictionaries tailored to the users’ needs and skills [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
Now the most of printed dictionaries have electronic versions (CDs, online, etc.).
However, there are dictionaries which do not have a printed version, but exist only in
digital format [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
      <p>
        However, some well-known, classic dictionaries that have been traditionally made
in paper format are gradually changing to digital format. Among them is Oxford
English Dictionary (OED) [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], one of the most reputable academic dictionaries published
by Oxford University Press. It traces historical development of the English language
(containing all the words that existed or existed in English literary and spoken
language since 1150), providing a comprehensive resource for common users and
researchers, and describing the usage of its variants around the world.
      </p>
      <p>
        Another example is “Diccionario de la lengua es”p(aDñoLlEa, Dictionary of the
Spanish language) which has been published by Academia Re(aSl panEisphañola
Royal Academy) since 1780 [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The dictionary went through twenty three editions,
changing over time into main reference source of the Spanish language. The most
recent, the 23rd saw the light in October 2014. The first electronic version of DLE
appeared in CD-ROM at the end of 1995, containing 21st edition and the recent one in
2015, representing 23rd edition (DLE 23) and being available at www.dle.rae.es. Since
March 2017, the Academy has been working on 24th edition, which will be digital
only [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Holding the discussion on further publication of the dictionary in printed
form, Academician Alvarez de Miranda gives the following figures: the number of
requests to online version of DLE 23 has reached 750 million at the end of 2017,
which is an average of 65 million monthly [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        In our paper, explanatory dictionary is considered as a comprehensive source of
information to be used for language researches. Research potential of dictionary is fully
developed in digital environment. When working with printed dictionaries, especially
multi-volume publications, the researcher has to spend a lot of time for searching,
analyzing and summarizing the information he needs. Due to the great amount,
elaborated structure and completeness of a lexicographic description such dictionaries are
carriers of a huge number of implicitly-defined linguistic, cognitive, logical and other
relationships which are difficult or near impossible to be investigated with traditional
methods. Also, in paper versions, search is usually limited to headword list only [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <p>It is important for digital dictionary to provide access to any structural element of
dictionary entry and give an opportunity of selecting particular entries corresponding
to the user’s presets. In other words, thpeoressibmiliutyst of bfeinding relevant
information without knowing the headword. The OED interface offers a powerful tool
for researching the vast amount of information in the dictionary. Although OED has
extensive capabilities (search by vocabulary classes, categories, thesaurus, etc.,
advanced search), but it doesn’t provitdhe necessary tools for the researcher. As for
DLE 23 online, the interface is limited to headword list only and some filters (e.g.
“word”, “lemma”, “contains”, “starts with”, “endsc.).wTitoh”conedtuct extensive
researches the entire text, including its meta-language elements, of a dictionary will
be required.</p>
      <p>
        With reference to the above the Ukrainian Lingua-Information Fund has developed
and implemented in his dictionary projects a new approach which is called virtual
lexicographic laboratories or shortly VLL. In contrast to electronic dictionaries, a
VLL is intended for working with the whole dictionary text with clearly defined
structure. It means that the user has a full access to all structural elements of the entry
including meta-language information related to a headword (e.g. definition type,
availability of word usage examples, headword structure etc.) [
        <xref ref-type="bibr" rid="ref11 ref12">11, 12</xref>
        ]. The paper deals with
the elaboration of the interface for the virtual lexicographic laboratory for the
Explanatory dictionary of the Spanish language (VLL DLE 23) and focuses on its research
potential.
2
2.1
      </p>
      <sec id="sec-1-1">
        <title>DLE 23 as an Object of the Research</title>
        <sec id="sec-1-1-1">
          <title>Choosing the Object of the Research</title>
          <p>
            Our interest in DLE 23 is arisen by the following reasons: 1) international status of
the Spanish language; 2) credibility and academic status of the dictionary; 3) other
school of lexicography which differs from other schools, such as Ukrainian and
English. In addition, the dictionary in question is of interest to translation lexicography
while creating translation systems: Spanish-Ukrainian and Ukrainian-Spanish. In this
context, the Russian linguist L. Ščerba states that each language pair needs four
dictionaries, one for explanation and one for translation in each direction – one for
comprehension and one for production purposes for each speech community [
            <xref ref-type="bibr" rid="ref9">9</xref>
            ]. Another
and also important reason of choosing DLE 23 is the availability of its digital version
that supports HTML5 format. The letter guarantees the authenticity of the dictionary
text and allows us to focus our attention on the structure of dictionary entry.
          </p>
          <p>
            DLE 23 is a fundamental lexicon containing standard vocabulary to be widely used
both in Spain and in Latin America. The purpose of the dictionary isn’t limited
giving a reference on lexical meaning of language units. Each entry also includes
detailed description of their grammatical, syntactic and pragmatic features. It should
be noted that the dictionary ideology is based on the conception of the Spanish
lexicographer J. Casares. According to this conception: the dictionary isn’ta- a
betically arranged entries but it’s also a tool providing a user with
for searching appropriate words and phrases he may need during communication
process [
            <xref ref-type="bibr" rid="ref2">2</xref>
            ].
          </p>
          <p>The headword list covers language units which: 1) belong to Spanish vocabulary
and 2) have come from other modern languages (English, German, French etc.). The
former ones are in normal font and the letter ones are in italics: “gato”, “kilobyte”,
“ojo”, s“oftware” e.tcThe list also includes frequently used abbreviations (“DNA”,
“ONU”) and acronyms (“radar”, “laser”). Furthermore, the dictionary contains
prefixes (“a-”, “pre-”, “contra-”, “pro-” etc.), suffixes (“-aico”, -“ino”, -“ivo” etc.),
derivational elements of Greek and Latin origin (“archi-”, h“idro-”, -“ónimo”) and
Latin expressions ad(“hoc”, a“priori”).
2.2</p>
        </sec>
        <sec id="sec-1-1-2">
          <title>Lexicographic Description of Spanish Language in DLE 23</title>
          <p>The main parts of the dictionary entry, as shown in fig.1, are: a headword with its
feminine flection (1); headword information (2); a set of definitions (3); a set of
subto
set of alph
necessary resources
entries of headwordc’osllocations and set expressions (4); cross-references to other
entries (5). A brief description of each element is given below.</p>
          <p>The first element, lemma, (1) can be of masculine, feminine or both forms. In letter
case the masculine form goes first, and then the feminine form, represented by
respective ending, follows after. For example: “alcalde, desa”, “duque, quesa”, “gato, ta”
should be read as alcalde, alcaldesa; duque, duquesa; gato, gata.</p>
          <p>Headword information (2) includes lemma variants, etymology, word flection and
orthography. Some examples of headword information are shown in Table 2. It
should be noted that some or all of these information elements aren’t always provided
in the dictionary. But for building the lexicographic model of DLE 23 they are
considered to be present in the entry structure.
The set of definitions (3) contains the explanation of lexical meanings together with
different labels indicating particular limitations on using the word. There are six types
of the labels in DLE 23 to denote grammatical class (“adj.”, “adv.”, “pron.”), usage</p>
        </sec>
        <sec id="sec-1-1-3">
          <title>Definition type</title>
          <p>Standard</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Contextual</title>
    </sec>
    <sec id="sec-3">
      <title>By synonym</title>
    </sec>
    <sec id="sec-4">
      <title>Explanatory</title>
    </sec>
    <sec id="sec-5">
      <title>Others</title>
      <p>(“iron.”, “despect.”, “peyo),r.s”tyle (“coloq.”, “estud.”, “fe)s,t.d”omain (“Comp.”,
“Dep.”, D“er.” etc.), region (“And.”, “Cat.”, “Cád.”, “Vall.”, “Arg.”, “Am. Cent.”
Etc.) and time period (“ant.”, “desus.”, “germ.”, “p.). Tuhse.”re are five types of
definitions used in DLE 23 to state the meaning of lemmas: standard, by synonym,
explanatory, and others not particularly stated by the dictionary makers. A brief
characteristic of each definition type is given in Table 3.
The definitions may be supplied together with examples to illustrate the usage of the
word or represent a pattern to make up a sentence with the headword. Another
peculiarity of DLE 23 is indicating additional grammatical or usage features of a headword
in a form of the comments, like “U. (Uts.adoc. tasm.bimén.” como sustantivo
masculino – Also used as a noun masculine) or “U. t. en sen(tU.safdigo.” también en
sentido figural – Also used in figurative meaning).</p>
      <p>The headword’s collocations are described in separate subentries (4) and treated in
the same way as the headword itself. In DLE 23 they fall into two types: 1) noun +
adjective: agua bendita, agua blanca, agua corriente; and 2) others: a) verbal:
agarrar alguien un agua, bañarse en agua rosad,abeber agua un buque etc.;
b) adverbial: como agua de mayo, como agua para chocolate etc.; c) adjectival: de
agua y lana, de arte y ensayo etc.</p>
      <p>The last element of the entry is the list of the cross-references (5) to other
collocations described in the other entries.
3</p>
      <sec id="sec-5-1">
        <title>Implementation of VLL DLE 23 Project</title>
        <p>The project of VLL DLE 23 is planned to be implemented in two stages: 1) creating a
shortened version of VLL with minimum interface elements to test some
technological solutions and 2) developing fully functional application with expanded interface.
Currently VLL DLE 23 is at end of the first stage and it demonstrates more
capabilities for working with the dictionary in digital environment than the original online
version of DLE 23.</p>
        <p>The first task was building up a conceptual model which would serve as a basis for
elaborating database and interface elements. The conceptual model of the dictionary
has been built on the basis of HTML text taken from online version of DLE 23. This
text shows much deeper and more transparent mark-up of entry structure than printed
version does. With conceptual model it is easy to determine a set of information
elements of the dictionary to be accessible in database.</p>
        <p>The second task consisted in the development of database and choosing database
type for VLL DLE 23. According to our experience, relational databases proved to be
inappropriate for developing efficient digital lexicographic systems. In this case the
data is stored implicitly as a set of several tables and relationships between them.
Operating separate tables as a single object requires building a powerful software
infrastructure. Furthermore, the evolutionary potential of such a digital object is
limited by database opacity.</p>
        <p>Since the dictionary entries are the elements of lexicographic system with a strictly
defined structure, it is logical to represent them as classes in object-oriented
programming languages, with subsequent processing, editing, and storage in explicit
form. This possibility is provided by the so-called NoSQL databases (document type
databases). For databases of this type the main element to be stored and processed is a
document (object) with a strictly described structure.</p>
        <p>The main advantage of NoSQL databases for our project is their capability to store
explicitly lexicographic objects without altering their internal structure, which opens
up direct access to each element of the lexicographic object and greatly simplifies the
possibility of its editing and modification (expansion).</p>
        <p>When choosing a specific NoSQL database, we were guided by the following
criteria: 1) ease of use; 2) support for transaction mechanisms; 3) support for parallelism;
and 4) free of charge for scientific purposes. Taking into consideration these criteria,
LiteDB (http://www.litedb.org/) was chosen. This is a relatively simple, free copy of
the shareware MongoDB database. An additional advantage of this database is the
ease of installation and connection, since LiteDB is implemented as a single library
file (dll) and one settings file (xml), rather than a whole software package.</p>
        <p>For efficient operation each language unit is put into correspondence with the set
of parameters: 1) headword variants; 2) headword structure; 3) headword type;
4) homonymy; 5) number of collocations “Noun + Adjectivaen”d 6) number of
collocations of other types. In our opinion, this set of parameters will be enough for
selection of the entries and analysis of their structure and content.</p>
        <p>The final task was to make a Web application to work with the database of
VLL DLE 23. The application based on .Net Core 2.1 technology was created. For
easy creation and further editing of interface elements, a set of HTML, CSS templates
and JavaScript Bootstrap scripts were used.</p>
      </sec>
      <sec id="sec-5-2">
        <title>Method and Technology</title>
        <p>
          The modern digital lexicography has turned to a multidisciplinary field and refers to:
a) applied research area originated at the intersection of linguistics and computer
science that studies the application of methods and techniques of information science
and technologies to creating a wide range of lexicographic systems; and b) a branch
of computer industry which is developing rapidly mainly due to the fact that the
lexicographic description is one of the efficient ways to obtain and disseminate
knowledge [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ].
        </p>
        <p>In this context, we consider dictionary text primarily not as a reference system, but
as a way of transferring linguistic knowledge. First of all, this statement concerns
fundamental lexicons, i.e. big explanatory dictionaries. This sets the problem of
providing dictionaries with appropriate tools. Obviously, it can be achieved only in
digital environment.</p>
        <p>
          Effective solution of this problem requires general theoretical basis to describe the
widest possible range of lexicographic objects. As such, we use the theory of
lexicographic systems [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ].
        </p>
        <p>
          The theory is based on a rather universal phenomenological principle,
characteristic of any system where information processes take place. These processes are called
lexicographic effect in information systems [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ].
        </p>
        <p>We consider the lexicographic system (L-system) as a special informational
(semiotic and semantic) system, in which a lexicographic effect (or a certain combination
of lexicographic effects) is induced. The formal representation in simplified form is as
follows:
,
( )
(
( ))
The elements of the model get the following interpretation according to the theory of
lexicographic systems:
 D is a modeled object (area);
 Q denotes lexicographic effect which induces a class of relatively stable
information entities;
 I0Q(D) = {xi} refers to a class of elementary information units in relation to
lexicographic effect Q;
 V(I0(D)) designates a set of descriptions (interpretations) of elementary information
units;
  is a subset of structural elements composing V(I0(D));
 [] denotes a separate substructure generated by a -operator within ;
 Red[V(IQ(D)] indicates a process of a recursive reduction that decomposes the
structure of lexicographic system into its fine elements.</p>
        <p>We consider the dictionary as a lexicographic system of special type where I0Q(D) is a
set of headwords. By V(I0(D)) = {V(xi)} we mean a set of entry texts, V(xi) is the entry
describing the headword (lemma) xi. The subset of structural elements () are entry
text fragments containing a piece of lexicographic description. If we apply [] to
V(xi), we shall get all -elements, i.e. lexicographic descriptions, related to the
headword xi only.</p>
        <p>One of the main aspects in the definition of an L-system as an information system
of a special type is the concept of its architecture. We use ANSI/X3/SPARK
architecture consisting of three levels of data representation: conceptual, internal and external.
Conceptual model (conceptual level of representation) of the subject area is a
semiotic, semantic model in which the concepts of the subject area are integrated in an
unambiguous, final and consistent way. Internal model (the internal level of
presentation) specifies the types, structures and formats of presentation, storage and
manipulation of data, algorithmic base and software environment in which the conceptual
model is implemented. External model (external level of presentation) reflects the end
users’ views (and, therefore, applied programmtheers)subjoenct area. The model
implements a set of tools that enable a user to manipulate the data represented at the
internal level. One conceptual model may correspond to several internal and external
models.</p>
        <p>
          As it was mentioned above, the Ukrainian Lingua-Information Fund has
developed the software systems to support creating, maintaining and functioning of
the dictionaries in digital environment named virtual lexicographic laboratories
(VLL). Moreover, VLLs permit corporate work on large lexicographic projects
by dictionary makers living in different cities and even different counties but
having equal access to the dictionary and working tools. The first VLL was
created in the Ukrainian Lingua-Information Fund in 2001 and is being used to create
and maintain the multivolume explanatory Dictionary of the Ukrainian language.
At present, the Ukrainian Lingua-Information Fund has designed over forty virtual
lexicographic laboratories for lexicographic systems of various types
(http://lcorp.ulif.org.ua) [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ].
        </p>
        <p>The advantages of VLL include the following: almost unlimited potential for the
integration of various linguistic facts in one object, ability to reflect language
dynamics, efficiency of navigation through the structural elements, possibility for
computational experiments. The possibility of multiple use of once formed lexicographic
structures and arrays by many professionals (linguists, linguistic technologists and
publishers) provided by the digital environment is also of great importance.</p>
        <p>
          All entry elements represented in conceptual model are displayed in database
structure. Any VLL provides direct access to each structural element of the
lexicographic system and offers the possibility to construct different index schemes. As a
rule lexicographic database of a dictionary displays only basic structures. Further
expansion is determined by two factors: the allocation of more subtle structural
elements of the basic structures and the introduction new parameters of a
dictionary entry. The increasing of parameters of lexicographic system sets two tasks for a
computer tool system: to create a form for the effective representation of
parameters and to develop an interface circuits to work with them [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ].
        </p>
        <p>The virtual laboratories have been initially developed to support dictionary-making
process in digital environment. They don’t fit for comprehensive analysisn- of dictio
ary text due to the lack of respective tools. Therefore the problem is to build up a
VLL provided with tools to perform dictionary text analysis and linguistic researches
on the basis of the dictionary.
5</p>
      </sec>
      <sec id="sec-5-3">
        <title>Conceptual Model of DLE 23</title>
        <p>The lexicographic data model is used as a conceptual model, in which the
structural elements of the dictionary entry and the relations between them are fixed.
All selected elements of the dictionary entry are displayed on the structure of a
computer database, and that provides both direct access to each of them and abi
lity to build a variety of index schemes. The electronic text and the interface of
online version gave us necessary material for constructing the model.</p>
        <p>In dictionary structure we can distinguish a set of headwords W = {х} which serve
as identifiers of corresponding entries V(x). As we have already mentioned, the list of
the headwords includes morphemes, words, abbreviations and collocations. For
convenience, all language units that compose the list will be considered as headwords.</p>
        <p>In its turn, the structure of each entry V(x) is decomposed into left part L(x) which
contains headword characteristics, and right part P(x) where the semantics of a
headword x is given. As the object of lexicographic description covers the headwords of
two types (words and collocations) it would be logical to represent the formal model
of DLE 23 in the following way:
( )
( ) ⋃ *⋃ ( ) ⋃ ( )
( )+,
Here VLex(x) is a lexicographic description of the headword x; Vi jFras(x) is a
description of the j-th collocation of i-th type; m(i) is the number of phrases of i-th type, and
n(x) is the number of collocation types in the dictionary entry V(x). As it was
mentioned in section 2, there are two types of collocations: noun + adjective and others,
i.e. with other parts of speech. Both lexicographic descriptions and is put
in correspondence with a basic structure composed of left part and right part,
respectively:
(
),
In case of V = V Lex(x), L0 refers to entry headword with its lexicographic description.
As for V = Vi Fras(x), L0 denotes a collocation with its description. The right part P0 for
headword and collocation is the same by its structure.</p>
        <p>
          Based on the DLE 23 text analysis [
          <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
          ], we distinguish the following parameters
for L0 component: RR (headword with gender forms or collocation with its variants),
DUPL (regional variant), ETYM (etymology), MORPHO (word flection), ORTHO
(orthography) and UNCRT (uncertain). We introduce another parameter PARHWi (i-th
parameter of a headword) for any parameter regardless of its type. UNCRT refers to a
type of lexicographic description which can’t be identified by formal
can be regarded as an information element of headword. Each parameter is
represented in our model by text line (authentic original text).
(2)
(3)
markers
but it
        </p>
        <p>The mandatory parameter is RR which describes a headword with its gender forms
(or collocation with its variants), and the others are optional. Our model doesnm’t- i
pose any restrictions on the number of parameters. Moreover, we assume that the
entry may include several parameters of the same type. Each parameter is a text line
of particular structure that carries an element of lexicographic description. The
examples of parameter contents are given below:</p>
        <p>Example 1.</p>
        <p>RR
bikini.</p>
        <p>DUPL
Tb. biquini.</p>
        <p>Example 2.</p>
        <p>RR
poeta, tisa.</p>
        <p>ETYM</p>
        <p>Del lat. poēta, y este del gr. ποιητής poiētḗs; para la forma f., cf. fr. mediev.
poétisse.</p>
        <p>MORPHO
Para el f., u. t. la forma poeta.</p>
        <p>Example 3.</p>
        <p>RR
inmaculado, da.</p>
        <p>ETYM
Del lat. immaculātus.</p>
        <p>ORTHO
Escr. con may. inicial en acep. 2.</p>
        <p>Example 4.</p>
        <p>RR
ONG.</p>
        <p>UNCRT
Sigla de organización no gubernamental.</p>
        <p>The right part (P0) is described by MNGN (meaning number), REM (block of
labels), DEF (definition), ED (encyclopedic data), COMM (comment) and IL
(illustration). All these parameters denote the information elements which compose the text of
the right part.</p>
        <p>Let us consider the structure of the label block (REM). The text of the block is
subdivided sequentially into smaller fragments, each one representing a label of a certain
type: REM-GR (grammar); REM-PR (pragmatics); REM-ST (stylistics); REM-SF
(domain); REM-REG (geographic region); REM-WHU (where-used). Grammar
remark REM-GR is mandatory remark, and it goes first and determines the part of
speech the headword belongs to.</p>
        <p>As a rule, lexical meaning in entry text is described by the component DEF. The
comments COMM correlate with definitions. Each definition and each comment can
be accompanied by its illustrations IL. Each lexical meaning can be described by
several DEF, COMM and IL. Let us demonstrate the decomposition of text into
fragments for headword cómico:</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Example 5.</title>
      <p>MNGN
1.</p>
      <p>REM
adj.</p>
      <p>DEF
Que divierte y hace reír.</p>
      <p>IL
Situación cómica.</p>
      <p>MNGN
2.</p>
      <p>REM
adj.</p>
      <p>DEF
Perteneciente o relativo a la comedia.</p>
      <p>MNGN
3.</p>
      <p>REM
adj.</p>
      <p>DEF
Dicho de un actor: Que representa papeles cómicos.</p>
      <p>COM
U. t. c. s.</p>
      <p>The implementation of conceptual model in a form of computer database will
allow the selection of dictionary entries according to their profile. For example, the user
will have options for selecting entries with a certain number of meaning; without
illustrations; with or without comments; entries with two (or more) definitions;
meaning with two (or more) comments etc.
6
6.1</p>
      <sec id="sec-6-1">
        <title>Results and Discussions</title>
        <sec id="sec-6-1-1">
          <title>Description of VLL DLE 23 Interface</title>
          <p>VLL DLE 23 is available at https://services.ulif.org.ua:44359 and accessible using the
user’s lo-gin and password. Since the application is at the development stage, the
interface language is Ukrainian but in final version English and Spanish will be also
added. The main window (Fig. 2) consists of the following interface elements
necessary to perform research works:
 Menu bar;
 Search panel;
 Word list panel;
 Entry text box.
The headword list panel is composed of the list of alphabetically arranged lemmas
and navigation bar to flick through the list. For convenience the headword list has
been split into the pages each one consisting of 150 elements. The user can user- “Fo
ward” and “Back” buttons or enter a page numbetro gient toteaxptprofpireiladte
part of the list.</p>
          <p>Entry text box displays dictionary entries in HTML format. The way of entry
representation is the same as in the online version. The main window has also a text box
(not shown in fig. 2) to view HTML text of the entry selected or copy text fragments
for full-text search.</p>
          <p>The interface of VLL DLE 23 allows the following modes to work with the
dictionary:
 Headword list;
 Entry profile;
 Full-text search.</p>
          <p>Headword List. According to the standards of Ukrainian Lingua-Information Fund,
the list of the headwords should represent all the words even those which haven’t
been initially included by dictionary makers. In contrast to original DLE 23, the
headword list of VLL DLE 23 is completed by feminine forms and regional variants
of the headwords. Regardless of the working mode, the information on the number of
headwords is always displayed. The current version runs to 106323 units.</p>
          <p>The headword can be selected either by clicking on it on the list, or entering a
sequence of characters that match exactly match exactly a word in search. The search
filters enhance quick search in cases when the user doesn’t know
of the headword. To enter diacritic symbols the virtual keyboard (fig. 3) can be also
used. The symbol –” “ has been diunctreod to search for the morphemes bearing
graphical accent. For example, the morphemes -cóla, -éo or -fágo must be typed as
сola, eo, fag.o
the
correct spelling</p>
          <p>Entry Profile. This mode is good for making a sample of entries which satisfy the
parameters of entry elements represented in DLE 23 conceptual model. In current
version the choice is limited to headword and some definition &amp; collocation
parameters. In future all the entry elements will be available to work with through the
interface.</p>
          <p>The mode is activated by clicking the menu “Sample” after which the
appears. The dialog box has two tabs “Headword” and I“nEnfitrsyt”t.he user can
choose headword parameters by which the entries are to be selected:
 Headword variants: lemma, masculine, feminine, regional variant, not defined.
 Headword structure: word, collocation, morpheme, not defined.
 Headword type: foreign word, abbreviation, acronym, not defined.
 Homonymy: yes (≥1) / n.o (0)
The second tab is intended for selecting the entries which correspond to definition /
collocation parameters:
 Number of definitions: numerical value and additional options (&gt;, ,≥ = &lt;≤).,
 Number of collocations “Noun + :Anduj.m”erical value and additional options
(&gt;, ,≥ = &lt;≤).,
 Number of collocations of other types: numerical value and additional options
(&gt;, ,≥ = &lt;≤).,
 Number of cross-reference: numerical value and additional options (&gt;, ≥, = &lt;≤).,
dialog
box
The entries are possible to be selected by headword and definition / collocation
parameters at the same time. The figure 4 shows the sample of the entries containing
homonymic polysemous morphemes. The sample corresponds to the following
parameters:</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>1. Headword structure: morpheme; 2. Homonymy: ≥1; 3. Number of definitions: &gt;1.</title>
      <p>Each of these parameters can be discarded by checking respective checkbox inm- “Sa
ple” dialog box. As for homonymy, polysemy and collocations quantitative
parameters are provided: amount (to be entered by user) and conditions to limit the amount.
As it was mentioned above, statistical calculations are made for each sample. In this
case the sample is made up by 3 elements.</p>
      <p>Full-Text Search. This mode is required when it is necessary to select entries by
specific meta-language elements of DLE 23: labels of different kind, characters that
make up additional comments (“U. m. en se,nmt.etaf-ilga.n”g)uage markers (“Tb.”,
“Voz”) e.tcFurthermore, the text string for search may include both the text of a
dictionary entry and the elements of HTML code (“&lt;abbr title="Usado solo en infinitivo
y en imperativo"&gt;”) taken from the text box in bottom of the main window. The
figure 5 shows the way of selecting the verbs marked as “U. solo en infinit.” (Used only
in infinitive).</p>
      <p>Fig. 5. List of the entries
which
contain the
verbs
marked
as “Usado. solo
In next subsections of the paper we consider the application of the interface for
different tasks: getting statistics data on language material, conducting linguistic researches
and creating sub-dictionaries on the basis of DLE 23.
6.2</p>
      <sec id="sec-7-1">
        <title>Statistics</title>
        <p>In each mode, statistical information is generated for each sample by clicking menu
“Statistics”. A general view of statistics window is shown in fig.6.</p>
        <p>Fig. 6. General view
of “Statistics”
window.</p>
        <p>For example, we can get quantitative characteristics of all language units composing
DLE 23:
 Total number of entries, including referential entries: 87436;
 Referential entries: 151;
 Homonyms: 5027;
 Morphemes: 467;
 Entries with collocations: 9127;
 Entries without collocations: 78309;
 Collocations “Noun + Adjective”: 12035;
 Collocations of other types: 13372.
6.3</p>
      </sec>
      <sec id="sec-7-2">
        <title>Linguistic Researches</title>
        <p>The interface of VLL DLE 23 allows conducting linguistic researches on the entire
text of the dictionary. On their basis the user can make certain conclusions regarding
lexical-semantic, etymological, grammatical and usage peculiarities of Spanish
language units. The linguistic information to be drawn from DLE 23 text using the
interface is shown in Table 4.
Let us show a concrete example of linguistic research by means of our interface. The
task is to get a set of lemmas which are hyponyms to the word embarcación (ship,
vessel). To achieve this goal it will be necessary to build up a sample of DLE 23
entries in which lemmas are explained by indicating the broader term embarcación in
their definitions. In this case both tools “Sampl-et”ext ansdearc“hF”ull will e-be
quired. So, the sequence of steps is as follows:
r
1. In the dialog box “Sample”, tick the checkbox “Word” (tab
ters”) and click the bMutatokne sa“mple”;
2. Select “Fu-ltlext search” in the combo box “Search parameters”;
3. Enter the word “Embarcaciótnh”e seinarch box and then click the search button.
As a result, we get a sample of 109 entries, the lemmas of which are hyponyms to the
word embarcación: aljibe (cistern), almadía (plank boat), barca (cockle boat), barcón
(warship), barquía (rowboat), etc. In definitions we can outline components of lexical
meaning by which the lemmas differ: 1) purpose; 2) design; 3) characteristics such as
shape and size; 4) geographic region; and 5) time period. It should be noted that some
definitions may represent several components. The results obtained, in our opinion,
are possible to be used in compiling glossaries to lessons, vocabulary quizzes and
exercises to memorize new words.
“eH-eadword
param
The samples created by means of VLL DLE 23 can be considered as sub-dictionaries
to be built on the basis of DLE 23. The table 5 shows the sub-dictionaries with their
volume and language units they describe. The volume has been caal-culated
tistics” tool.</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>Dictionary of morphemes</title>
      <p>Dictionary of homonyms
Dictionary of Latin expressions
Dictionary of acronyms and abbreviations
Dictionary of foreign words
Dictionary of foreign collocations
Dictionary of collocations “Noun
Dictionary of set expressions
Dictionary of words of common gender
Dictionary of monosemantic words
Dictionary of polysemantic words
+</p>
      <p>Volume
By activating the parameters of headword and entry at the same time in the dialog box
“Sample”, the user can get a co m-dbiicntieodnarsyu.bUsing “Statistics” tool we
count all the units contained in the sub-dictionary.
can</p>
      <p>For example, we want to make a sub-dictionary of monosemantic words which
have the collocationsNo“un + Adjective”. In the dialog box “Sample” the following
options must be selected:
1. In the tab “Headword parameters,” tick the checkbox “ W(“Horeda”dword
strcuture” pane;l)
2. In the tab “Entry parameters” select the number of definitions “= nu1m”,ber of
collocations Noun + Adjective “≥ 1” and numbceorlloocfations of other types “=
0”.</p>
      <p>The sub-dictionary has in total 509 monosemantic headwords, out of which 476 are
non-homonymic and 43 are homonymic. The total amount of the collocations “noun
adjective” is 724.
+
7</p>
      <sec id="sec-8-1">
        <title>Conclusions and Future Works</title>
        <p>DLE 23, like other fundamental lexicons, is an exhaustive source of information on
the lexical-grammatical and lexical-semantic properties of Spanish language units.
Therefore, it can be useful not only to ordinary users as a reference book, but also to
linguists as a means for studying the language. For convenient use of the dictionary
for research works, a virtual lexicographic laboratory VLL DLE 23 providing access
not only to the world list, but also to the text of dictionary entries was developed.
Unlike the online version of DLE 23, the virtual lexicographic laboratory has the
following advantages:
 Practically unlimited potential for integrating various linguistic facts in one object;
 Ability to reflect language dynamics;
 Possibility of selecting linguistic information from dictionary text by using sample
parameters (in current version the number of parameters is limited to lemma
parameters and some definition / collocation parameters);
 Availability of three working modes: headword list, dictionary profile and full-text
search.</p>
        <p>The virtual lexicographic laboratory provides the users with the tools necessary for
studying grammatical, semantic, pragmatic, and other features of the Spanish linguage
units. The current version of VLL DLE 23 allows conducting the following researches
on the basis of DLE 23:
 Statistical calculations both on the entire dictionary and on a separate sample of
dictionary entries;
 Studying Spanish vocabulary and dictionary entry texts to extract various linguistic
facts recorded in DLE 23 (semantics, etymology, grammar, word usage);
 Building secondary lexicographic objects (or sub-dictionaries) that describe the
specific lexical composition of the Spanish language.</p>
        <p>Using VLE DLE 23 tools, it is possible to extract linguistic information, which may
serve as a basis for compiling learner’s dictionaries, preparing didactic materials and
tests for those who study Spanish.</p>
        <p>In final version of VLL DLE 23 the structural profile of dictionary entry will be
determined by all the structural elements of the conceptual model. The user will be
able to select the entries by indicating obligatory presence or absence of a structural
element. Additionally the user will have the possibility of specifying specific content
of the structural elements. The final stage of our elaboration will also include testing
of the fully developed user’s interface with representatives and real users.
, panhispánico</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1. Cadenaser.com.:
          <article-title>La RAE regala ejemplares del diccionario en papel porque nadie los compra</article-title>
          , https://cadenaser.com/ser/2018/07/04/cultura/1530712888_864828.html,
          <source>last accessed 11.04</source>
          .20.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Casares</surname>
            <given-names>J.:</given-names>
          </string-name>
          <article-title>Nuevo concepto del diccionario</article-title>
          .
          <source>Editorial CSIC</source>
          , Madrid (
          <year>1992</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3. Diccionario de la lengua esphatñtposl:a/,/dle.rae.es/,
          <source>last accessed</source>
          <year>2020</year>
          /02/23.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Kupriianov</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Akopiants</surname>
          </string-name>
          , N.:
          <article-title>Developing Linguistic Research Tools for Virtual Lexicographic Laboratory of the Spanish Language Explanatory Dictionary</article-title>
          . In: Lytvyn,
          <string-name>
            <given-names>V.</given-names>
            ,
            <surname>Sharonova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Hamon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Cherednichenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            ,
            <surname>Grabar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Kowalska</surname>
          </string-name>
          <string-name>
            <surname>Styczen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Vysotska</surname>
          </string-name>
          , V. (Eds.)
          <source>Computational Linguistics and Intelligent Systems. Proc. 3rd Int. Conf. COLINS 2019</source>
          , Vol. I: Main Conference. Kharkiv, Ukraine,
          <source>April 18-19</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>43</fpage>
          -
          <lpage>52</lpage>
          , CEUR-WS.org, http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2362</volume>
          /, last accessed
          <year>2020</year>
          /02/23.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Kupriianov</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <article-title>Lexicographic system of the Spanish language: Phenomenology of integral description</article-title>
          . Ukrainian
          <string-name>
            <surname>Lingua-Information</surname>
            <given-names>Fund</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kyiv</surname>
          </string-name>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Lew</surname>
            ,
            <given-names>R:</given-names>
          </string-name>
          <article-title>Online dictionary skills</article-title>
          . In: Kosem,
          <string-name>
            <given-names>I.</given-names>
            ,
            <surname>Kallas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Gantar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Krek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Langemets</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Tuulik</surname>
          </string-name>
          , M. (eds.):
          <article-title>Electronic lexicography in the 21st century: thinking outside the paper</article-title>
          .
          <source>Proceedings of the eLex 2013 conference, 17-19 October</source>
          <year>2013</year>
          , pp.
          <fpage>16</fpage>
          -
          <lpage>31</lpage>
          , Tallinn, Estonia (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7. Oxford English Dictionary, https://www.oed.com/,
          <source>last accessed</source>
          <year>2020</year>
          /02/23.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8. Rae.es.:
          <article-title>El nuevo diccionario académico será digital y más https://www.rae.es/noticias/el-nuevo-diccionario-academico-sera-digital-y-mas-panhispanico</article-title>
          ,
          <source>last accessed</source>
          <year>2020</year>
          /02/23.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Ščerba</surname>
            ,
            <given-names>L</given-names>
          </string-name>
          :.
          <article-title>Towards a General Theory of Lexicography</article-title>
          . In: R. R. K. Hartmann (ed.),
          <source>Lexicography. Critical Concepts</source>
          ,
          <source>Vol. 3</source>
          . pp.
          <fpage>11</fpage>
          -
          <lpage>50</lpage>
          . Routledge, New York ([
          <year>1940</year>
          ]
          <year>2003</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Shyrokov</surname>
          </string-name>
          , V.: Computer lexicography,
          <source>Kyiv</source>
          (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Shyrokov</surname>
          </string-name>
          , V (Ed.):
          <article-title>Computer linguistic studies: Proceedings of the Ukrainian LinguaInformation Fund NAS of Ukraine</article-title>
          . Vol.
          <volume>1</volume>
          :
          <article-title>Research paradigm and basic language information structures</article-title>
          . Ukrainian
          <string-name>
            <surname>Lingua-Information</surname>
            <given-names>Fund</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kyiv</surname>
          </string-name>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Shyrokov</surname>
          </string-name>
          , V (Ed.):
          <article-title>Computer linguistic studies: Proceedings of the Ukrainian LinguaInformation Fund NAS of Ukraine</article-title>
          . Vol.
          <volume>5</volume>
          :
          <article-title>Virtualization of linguistic technologies</article-title>
          . Ukrainian
          <string-name>
            <surname>Lingua-Information</surname>
            <given-names>Fund</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kyiv</surname>
          </string-name>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Vincze</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alonso</surname>
            <given-names>Ramos</given-names>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          :
          <article-title>Testing an electronic collocation dictionary interface: Diccionario de Colocaciones del E</article-title>
          . sIpna:ñoKlosem,
          <string-name>
            <given-names>I.</given-names>
            ,
            <surname>Kallas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Gantar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Krek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Langemets</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Tuulik</surname>
          </string-name>
          , M. (eds.):
          <article-title>Electronic lexicography in the 21st century: thinking outside the paper</article-title>
          .
          <source>Proceedings of the eLex 2013 conference, 17-19 October</source>
          <year>2013</year>
          , pp.
          <fpage>328</fpage>
          -
          <lpage>337</lpage>
          , Tallinn, Estonia (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>