<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>M. Ramos);</journal-title>
      </journal-title-group>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>TEI Lex-0 to Ontolex-Lemon</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Bruno Almeida</string-name>
          <email>brunoalmeida@fcsh.unl.pt</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rute Costa</string-name>
          <email>rute.costa@fcsh.unl.pt</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ana Salgado</string-name>
          <email>anasalgado@fcsh.unl.pt</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Margarida Ramos</string-name>
          <email>mvramos@fcsh.unl.pt</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Laurent Romary</string-name>
          <email>laurent.romary@inria.fr</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fahad Khan</string-name>
          <email>fahad.khan@ilc.cnr.it</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sara Carvalho</string-name>
          <email>carvalho@ua.pt</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mohamed Khemakhem</string-name>
          <email>mohamed.khemakhem@inria.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Raquel Silva</string-name>
          <email>raq.silva@fcsh.unl.pt</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Toma Tasovac</string-name>
          <email>ttasovac@humanistika.org</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>ArcaScience</institution>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>BCDH - Belgrade Center for Digital Humanities</institution>
          ,
          <country country="RS">Serbia</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>NOVA CLUNL - Centro de Linguística da Universidade Nova de Lisboa</institution>
          ,
          <country country="PT">Portugal</country>
        </aff>
      </contrib-group>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>This paper describes ongoing work in the modelling of usage information in the context of the MORDigital project. The latter is based on the encoding and publication as linked data of Diccionario da Lingua Portugueza, a Portuguese legacy dictionary authored by António de Morais Silva, whose first edition was published in 1789. In this paper, we focus on the TEI Lex-0 encoding and Ontolex-Lemon modelling of lexicographic articles from the Morais Silva dictionary that feature usage information. The approach described in this paper should be reusable for other projects involving the encoding and linked data publication of legacy dictionaries.</p>
      </abstract>
      <kwd-group>
        <kwd>legacy dictionaries</kwd>
        <kwd>usage information</kwd>
        <kwd>lexicography</kwd>
        <kwd>digital humanities</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org
(T. Tasovac)
CEUR
Workshop
Proceedings</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        The publication of Diccionario da Lingua Portugueza in 1789, authored by António de Morais
Silva, marks the beginning of contemporary Portuguese lexicography, following the model set by
several modern language dictionaries published in Europe in the 17th and 18th centuries [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ].
As the first Portuguese monolingual dictionary, it had a fundamental role in the standardisation
of this language, and constitutes a reference for studying the evolution of the Portuguese lexicon
[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The first edition of the dictionary had two volumes (Vol. 1, 752 p. and Vol. 2, 541 p.). Morais
directly oversaw the 2nd and 3rd editions (published, respectively, in 1813 and 1823). This work
was greatly revised and updated over the years, culminating in the 10th edition, which was
published in 12 volumes from 1949 to 1959.
      </p>
      <p>
        The MorDigital project1 aims at digitising and publishing in open access the first three
editions of the dictionary by Morais [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Our methodology involves the reuse of digitised
versions of the dictionary, available in the public domain as PDF files. The latter are currently
undergoing a re-OCRisation process to ensure the quality of the final output of the project. The
digitised versions of the dictionary will be structured by means of several open standards for
encoding and modelling lexical and dictionary data, which will facilitate interoperability with
existing systems and datasets.
      </p>
      <p>
        The encoding of the dictionary’s editions will be carried out in TEI Lex-0 [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], a baseline
XML encoding for machine-readable dictionaries based on the guidelines of the Text Encoding
Initiative (TEI). The TEI Lex-0 encoding of the Morais Silva dictionary will be the basis for a LMF
(Lexical Markup Framework) version, which should be facilitated by the ongoing convergence
between TEI and the LMF standard [
        <xref ref-type="bibr" rid="ref7">6</xref>
        ]. The TEI Lex-0 encoding of the Morais Silva dictionary
will be further transformed to RDF based on Ontolex-Lemon [
        <xref ref-type="bibr" rid="ref8">7</xref>
        ], a model originally developed
for enriching ontologies with lexical information, which has become a de facto standard for
publishing lexical resources as linked data [
        <xref ref-type="bibr" rid="ref9">8</xref>
        ]. The recently developed lexicography module,
lexicog [
        <xref ref-type="bibr" rid="ref10">9</xref>
        ], facilitates the application of Ontolex to dictionary data. An XSLT-based tool,
tei2ontolex2, will be used for the conversion process. Examples such as those presented in
this paper will be the basis for a wider coverage of features of this tool.
      </p>
      <p>While the work shown in this paper is focussed on digital lexicography, specifically on the
retrodigitisation of legacy dictionaries, technologies such as TEI and Ontolex are relevant in
many other domains of the digital humanities involving text encoding, analysis, and publishing,
including discourse analysis and other fields of linguistics, digital literary studies, cultural
heritage, and digital archives. TEI, for instance, includes several communities of practice whose
activity is centred on encoding text in a standardised, flexible, and interoperable framework. This
paper could, therefore, be useful for scholars in those communities, especially for those whose
work also involves applying linked data models for publishing digital humanities resources.</p>
      <sec id="sec-2-1">
        <title>1.1. Related work</title>
        <p>
          Relevant work for the approach described in this paper has been carried out recently. Much
of this work is focussed on relating TEI Lex-0 with Ontolex-Lemon, augmented with other
ontologies, for the linked data modelling of lexicographic resources, emphasising applications
in retrodigitisation projects3. The importance of the above-mentioned formats and models is
made clear in the context of the ELEXIS project4, a European infrastructure for interoperable
lexicographic resources, in which TEI Lex-0 and Ontolex-Lemon are two of the main formats
for publishing and interlinking dictionary data [
          <xref ref-type="bibr" rid="ref13">12</xref>
          ].
        </p>
        <p>
          Khan and Salgado [
          <xref ref-type="bibr" rid="ref14">13</xref>
          ] describe a novel approach to the modelling and publication of
lexicographic resources as linked data. This approach consists of using Ontolex-Lemon and lexicog in
conjunction with the CIDOC-CRM aligned FRBRoo ontology for representing diferent levels of
description of lexicographic resources (i.e., work, expression, manifestation) corresponding to
the diferent views of dictionaries explained in the TEI Guidelines [
          <xref ref-type="bibr" rid="ref15">14</xref>
          ], namely typographical
(the layout of the pages), editorial (the text of the dictionary) and lexical (the conceptual and
linguistic content of the dictionary). The work carried out in this paper pertains to the lexical
view of the Morais dictionary, in which elements from Ontolex and lexicog are used to model
usage information following an interoperable approach to that of Khan and Salgado [
          <xref ref-type="bibr" rid="ref14">13</xref>
          ] (see
Section 6 of this paper).
        </p>
        <p>
          In addition to the work described above, current research has focussed on lexicography and
digital humanities, including the application of ontologies and knowledge organisation. Costa et
al. [
          <xref ref-type="bibr" rid="ref16 ref6">15</xref>
          ] show how domain labels can be modelled through an OWL ontology in the medical and
health sciences (OntoDom-Lab-Med5), whose classes can be applied to the semantic annotation
of usage information in TEI Lex-0 encoded dictionaries. This can be done solely with TEI
elements and attributes, including the ontology class URI within the usage information element,
or with the XML Linking Language (XLink), which also allows to describe the role played by
ontology class URI, providing more complex domain information.
        </p>
        <p>
          In turn, Costa et al. [
          <xref ref-type="bibr" rid="ref17">16</xref>
          ] show how SKOS (Simple Knowledge Organization System), a W3C
recommendation for modelling knowledge organisation schemes, can be employed in digital
humanities projects for modelling linguistic/lexicographic categories represented in dictionaries’
lists of abbreviations (e.g., part of speech, grammatical gender, register). SKOS modelling, for
knowledge organisation purposes, is shown to be complementary to the TEI Lex-0 encoding of
dictionary articles.
        </p>
        <p>
          Salgado et al. [
          <xref ref-type="bibr" rid="ref18">17</xref>
          ] further highlight the importance of domain label modelling through
terminological methods, namely by structuring domain labels. The resulting taxonomies or
classifications of domains can be included directly within the &lt;teiHeader&gt; element of
TEIencoded dictionaries, whose categories can be linked to individual dictionary articles through the
TEI usage information element, while still retaining the text values that occur in the dictionary
articles for human readability purposes.
3See for example [
          <xref ref-type="bibr" rid="ref11 ref12">10, 11</xref>
          ]
        </p>
        <p>
          Section 5 of this paper illustrates the simpler option, as described in Costa et al. [
          <xref ref-type="bibr" rid="ref16 ref6">15</xref>
          ], for
linking the TEI encoding of usage information to OntoDomLab-Med and to the MorDigital
Domain Classification. The latter, described in Section 4, follows notions laid out in Costa et
al. [
          <xref ref-type="bibr" rid="ref17">16</xref>
          ] with regard to the complementarity between TEI encoding and SKOS modelling, and
Salgado et al. [
          <xref ref-type="bibr" rid="ref18">17</xref>
          ] regarding the application of terminological methods in lexicography for
structuring domain labels.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>2. Background</title>
      <sec id="sec-3-1">
        <title>2.1. Usage information in lexicography</title>
        <p>
          In lexicographic theory, usage or diasystematic6 information is understood as a set of constraints
or restrictions on the use of words, or their senses, to certain contexts or to a subset of language
users (e.g., [
          <xref ref-type="bibr" rid="ref20 ref21">19, 20</xref>
          ]). Dictionaries traditionally include usage information in the lexicographical
articles as labels (often abbreviated), or in more verbose forms, such as notes or as part of the
lexicographic definitions themselves. As Svensén notes, diasystematic marking in lexicographic
articles implies that “a certain lexical item deviates in a certain respect from the main bulk
of items described in a dictionary” [20, p. 315]. This notion of deviation from the lexicon of
standard varieties of languages is, therefore, the cornerstone of usage marking in dictionaries.
        </p>
        <p>
          As noted by Salgado et al., usage labels are devices whose simplicity is only apparent, since
they “often conceal the complexity of dynamic sociolinguistic, cultural, and ideological
processes that they are meant to illustrate” [21, p. 134]. In a digital humanities context, such
as the retrodigitisation of dictionaries, labels are also a challenge for the interoperability of
lexicographic datasets from diferent sources, ranging over diferent cultures, languages, and
periods. Indeed, since at least the 17th century, lexicographers have included labelling/marking
in dictionaries, describing a wide variety of usage information, whose treatment in theoretical
and practical lexicography has not always been consistent [
          <xref ref-type="bibr" rid="ref23">22</xref>
          ].
        </p>
        <p>
          There have been several surveys of usage information in theoretical and practical lexicography,
such as Ptaszyński [
          <xref ref-type="bibr" rid="ref23">22</xref>
          ] and Vrbinc and Vrbinc [
          <xref ref-type="bibr" rid="ref24">23</xref>
          ]. A more recent and comprehensive survey has
been carried out by Salgado et al. [
          <xref ref-type="bibr" rid="ref22">21</xref>
          ]. These studies note the dificulties caused by the variety,
and partial overlapping, of usage information types put forward by lexicographers, and the
diferent terminology employed by them. Hausmann [
          <xref ref-type="bibr" rid="ref25">24</xref>
          ] put forward the most comprehensive
classification, with 11 types of usage information. This classification was later adopted by
Bergenholtz and Tarp [
          <xref ref-type="bibr" rid="ref26">25</xref>
          ] and Svensén [
          <xref ref-type="bibr" rid="ref21">20</xref>
          ]. Landau [
          <xref ref-type="bibr" rid="ref20">19</xref>
          ], whose manual was first published
in the same year as Hausmann’s proposal, distinguished 9 types of usage information. In a
later study, Milroy and Milroy [
          <xref ref-type="bibr" rid="ref27">26</xref>
          ] included 5 types of usage information, distinguishing ‘group
labels’, which pertain to a subset of language users, from ‘register labels’, pertaining to specific
social and communicative contexts. Jackson [
          <xref ref-type="bibr" rid="ref28">27</xref>
          ] put forward a classification including 7 types
of usage information. Atkins and Rundell [
          <xref ref-type="bibr" rid="ref29">28</xref>
          ] considered 9 types of usage information, which
they call ‘linguistic labels’.
6The term ‘diasystem’, originating from dialectology [
          <xref ref-type="bibr" rid="ref19">18</xref>
          ], designates a general language system encompassing
several dialects. In the context of lexicography, the term ‘diasystematic’ applies to the marking of lexical items
whose usage deviates from that of the lexicon of the standard variety of a language.
t/rrceeyunm li/eeaggoon liitraaavon ,littfceoyunn itt/rrseegy ittrrsceeod lltsceaaxgoou lseaaggdnn ,littfceoyunn itt/rrseegy ,littfceoyunn itt/rrseegy lircceaohn liiitrseeadnm ,lltt/sseyu litreaavyn trscaouu lev
c r c — s e r s a s e s e t c — in ito tre ts le —
        </p>
        <p>n
C y n e e
ico txT ia equ tdu
o e m r i
s t o f tt
d a
y
c
y
c
n
e
2.2. TEI P5 and TEI Lex-0
The Text Encoding Initiative, or TEI, is an international organisation with a long history in
the development of guidelines, and associated schemas, for encoding machine-readable text in
social sciences and humanities. The current release of the guidelines, TEI P5, include a module
(i.e., a set of XML elements and attributes) for encoding dictionaries and other lexical resources,
e.g., glossaries and word lists included in other documents [14, 9: Dictionaries]. The common
characteristic of these resources is that they consist of entries/articles describing lexical items in
a language (or languages). Since the characteristics of these resources may vary widely, TEI P5
includes the &lt;entry&gt; element for encoding conventional dictionary articles, the &lt;entryFree&gt;
element for encoding unstructured entries in generic lexical resources, and the &lt;superEntry&gt;
element for grouping several lexical entries.</p>
        <p>In a straightforward TEI encoding of print dictionaries, the main body of text (encoded
through the &lt;body&gt; element) should include a set of &lt;entry&gt; elements, corresponding to the
dictionary’s articles. Among other possibilities, each &lt;entry&gt; element may contain:
• Information about written and/or spoken forms of the headword (through the &lt;form&gt;
element).
• Grammatical information (within the &lt;gramGrp&gt; element).
• Information about the headword’s senses (through the &lt;sense&gt; element).
• Cited quotations (through the &lt;cit&gt; element, part of the core TEI elements).
• Usage information (through the &lt;usg&gt; element).</p>
        <p>As defined in TEI P5, the &lt;usg&gt; element may appear in several components and positions of the
lexicographic article structure (&lt;entry&gt; element). The &lt;usg&gt; element has an optional @type
attribute for indicating types of usage information, and the guidelines include 16 sample values,
which have been adopted in many projects.</p>
        <p>
          The flexibility of TEI P5, allowing numerous possibilities and combinations of elements
for encoding dictionaries, is a hurdle for interoperability of dictionary data emanating from
diferent projects and for its use by NLP applications. TEI Lex-0 is a more recent initiative
at establishing a baseline encoding and target format for TEI dictionaries [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. It introduces a
number of constraints on the encoding of lexicographical articles, such as limiting the possible
occurrences of the &lt;def&gt; element (used for definition texts) to &lt;sense&gt;, while in TEI P5 &lt;def&gt;
could appear directly within &lt;entry&gt; and other elements.
        </p>
        <p>
          With regard to usage information, in TEI Lex-0 the &lt;usg&gt; element may still occur in several
points of the entry hierarchy, but the @type attribute is made mandatory. A more concise list of
10 values for possible usage types was also introduced, following Salgado et al. [
          <xref ref-type="bibr" rid="ref22">21</xref>
          ]7:
7The TEI Lex-0 reference document includes a table showing the correspondence between the suggested values
of usg/@type in TEI P5, their required values in TEI Lex-0 and some examples of real dictionary data [5, sec. 8.2.
Types of usage].
• "temporal". Marks the usage of a lexical item in a scale from old to new (this is known
as diachronic information in lexicographic literature).
• "geographic". Marks the place or region where a lexical item is mostly used (diatopic
information).
• "domain". Marks the subject field in which the lexical item is mostly used ( diatechnical
information).
• "frequency". Marks the relative occurrence of a lexical item (diafrequential
information).
• "textType". Marks the typical discource type or genre where a lexical item is mostly
used (diatextual information).
• "attitude". Marks the speaker’s subjective point of view regarding the referent of a
lexical item (diaevaluative information).
• "socioCultural". Marks the social groups (diastratic information) and/or
communicative situations (diaphasic information) where a lexical item mostly occurs.
• "meaningType". Marks a semantic extension of the sense of a lexical item8.
• "normativity". Marks the usage of a lexical item as non-standard or incorrect
(dianormative information).
        </p>
        <p>• "hint". Marks a non-specified usage of a lexical item (default value of &lt;usg&gt;).
TEI Lex-0 efectively restricts the scope of &lt;usg&gt; to put it more in line with lexicographic
theory, deprecating, for example, the encoding of lexical relations and etymological information
through the &lt;usg&gt; element, which are both allowed in TEI P5.</p>
      </sec>
      <sec id="sec-3-2">
        <title>2.3. Ontolex-Lemon, lexicog and LexInfo</title>
        <p>
          The Lexicon Model for Ontologies (Ontolex-Lemon) was put forward by the W3C
OntologyLexicon community group for the enrichment of ontologies with linguistic information in
the Semantic Web [
          <xref ref-type="bibr" rid="ref8">7</xref>
          ]. This model has since become a de facto standard for modelling lexical
resources as RDF and publishing them as linguistic linked data [
          <xref ref-type="bibr" rid="ref9">8</xref>
          ]. Ontolex-Lemon was heavily
inspired by the core model of the Lexical Markup Framework (LMF), an ISO standard for
machine-readable lexical resources [
          <xref ref-type="bibr" rid="ref30">29</xref>
          ], having transposed several LMF classes for modelling
linguistic information. These include the following:
• LexicalEntry. A lexical entry is a unit of a lexicon consisting of a set of grammatically
related forms associated with a collection of senses (e.g., cat in English, including both
singular and plural forms, which are associated with several senses in this language).
• Form. A form is a grammatical realisation of a lexical entry (e.g., the singular form ‘cat’,
which is the lemma or canonical form for representing the entry).
• LexicalSense. A sense associated with a lexical entry, which can be described, e.g., in a
dictionary definition.
8Labels for marking the semantic extension of senses (e.g., ‘figurative’, or ‘fig.’) are not considered in the typologies
of usage information of lexicographic literature. Nevertheless, it remains a possible usage type in TEI Lex-0,
maintaining interoperability with the style usage type of TEI P5.
• Lexicon. A lexicon is a collection of lexical entries for a particular language.
        </p>
        <p>
          While the core model has enough elements to describe information about the lexicon of
individual languages, it lacks expressive power to properly describe lexical resources, such as
dictionaries. Indeed, lexicographic articles often include information about forms shared by
diferent parts of speech, which necessarily correspond to diferent lexical entries in
OntolexLemon. Lexicog, the Ontolex-Lemon Lexicography Module, aims to address these issues, based
on several experiences in converting to linked data existing lexicographic resources [
          <xref ref-type="bibr" rid="ref10">9</xref>
          ]. Lexicog
introduces classes and properties that enable the distinction of the lexicon and its lexicographic
description. The following are the most relevant classes of lexicog:
• Entry. An entry is an element of a dictionary’s microstructure, corresponding to a
lexicographic article.
• LexicographicComponent. A lexicographic component is an element for describing
substructures of dictionary entries (e.g., senses, sense groups or subentries in a lexicographic
article).
• LexicographicResource. A lexicographic resource is a collection of lexicographic
articles.
        </p>
        <p>With these elements, it becomes possible to distinguish between lexical entries (which must
belong to the same part of speech, such as noun or adjective) and dictionary entries (in which
diferent parts of speech may be conflated). Sense groupings and subentries can be modelled as
lexicographic components, which in turn describe either lexical entries or lexical senses in the
language’s lexicon.</p>
        <p>
          The means for modelling usage information fall within the core Ontolex-Lemon model, which
includes the usage property. The latter allows to represent modulations in the meaning of lexical
entries determined by “usage conditions or pragmatic implications” [
          <xref ref-type="bibr" rid="ref8">7</xref>
          ], such as due to register,
connotations, etc. The domain of usage is restricted to instances of LexicalSense. This is
an important distinguishing feature of Ontolex-Lemon when compared to the TEI abstract
model: while in the latter usage information may appear directly in several parts of an entry, in
Ontolex-Lemon usage is necessarily associated with lexical senses.
        </p>
        <p>While Ontolex-Lemon includes the usage property, it does not specify how to represent usage
information. Furthermore, Ontolex-Lemon does not include elements for representing types of
usage information, or other linguistic categories (e.g., part of speech). Ontolex-Lemon relies on
external vocabularies for describing the properties of linguistic objects. The LexInfo ontology9,
in particular, was created to provide linguistic categories for modelling in Ontolex-Lemon.
The former declares 13 sub-properties of the Ontolex-Lemon usage property, which adopt the
above-mentioned usage types of TEI Lex-010.
10LexInfo and Ontolex-Lemon include three additional sub-properties of usage, (condition,
normativeAuthorization and register) but they pertain to domains of application other than lexicography,
namely argument structure and terminological databases.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>3. Usage information in the Morais Silva dictionary</title>
      <sec id="sec-4-1">
        <title>3.1. Typology of usage information</title>
        <p>In the Morais Silva dictionary, usage information is mostly marked by abbreviated labels, the
full form of which is given in the dictionary’s explanation of abbreviations11. The analysis
of the latter provides valuable insights into the usage information that was relevant to the
late 18th century lexicographer. Our analysis resulted in the following typology of labels (the
corresponding types of usage in TEI Lex-0 and LexInfo, are shown in parentheses):
• Diatechnical information ("domain"). E.g., Med., for medical terms).
• Diatextual information ("textType"). E.g., Poet., for poetic words.
• Diastratic information ("socioCultural"). E.g., Vulg., for words associated with the
common people.
• Diaphasic information ("socioCultural"). E.g., Fam., for words used in a familiar
context.
• Diatopic information ("geographic"). E.g., Asiat., for words used in the former
Portuguese colonies in India.
• Diachronic information ("temporal"). E.g., Ant., for antiquated words.
• Diaintegrative information ("hint"). E.g., Lat., for Latin words integrated in
Portuguese.
• Diafrequential information ("frequency"). E.g., P. us., for rarely used words.
• Semantic extension information ("meaningType"). E.g., f. and fig. , marking figurative
usages of lexical items).</p>
      </sec>
      <sec id="sec-4-2">
        <title>3.2. Examples of usage information in the Morais Silva dictionary</title>
        <p>11Although the listed abbreviations did not change significantly in the first three editions of the dictionary, it should
be noted that these abbreviations are not always used in the dictionary’s articles, in which the relevant information
is often marked using full forms or non-listed abbreviations. For example, ‘f.’ is explained as an abbreviation of
‘femenino’ (the feminine grammatical gender), although in some articles ‘f.’ is also used to mark figurative senses.
‘Fig.’ is also used, although it is not listed in the explanation of abbreviations.
associated to what can be construed as a citation example, surrar peles (‘to flesh hides’). The last
sense corresponds to the pronominal form of the verb, surrar-se, which means ‘to run away’,
‘to remove oneself’. This sense is marked with usage information (t. ch. or termo chulo in its
expanded form) indicating an informal communicative situation in which the participants joke
or mock one another or a third party, which can be thought of as diaphasic information12. As
we can see from these examples, the labels may appear at diferent points of the lexicographic
article.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. The MorDigital classification of domain labels</title>
      <p>Our approach pays special attention to the encoding and modelling of domain labels in
lexicographic resources, following previous work (cf. Section 1.1). This is justified in our approach
by the fact that domain labels constitute an interface between lexicography and terminology,
which is an important part of the interdisciplinary framework of the MorDigital project. In this
context, the following complementary approaches are possible:
• Declare a taxonomy/classification of the domains referred to by the dictionary labels in
the &lt;teiHeader&gt; element of the TEI Lex-0 encoding of the Morais Silva dictionary.
• Model the domain classification independently by means of W3C-maintained technologies,
such as RDF, OWL and SKOS.</p>
      <p>While the former approach facilitates browsing and querying the TEI encoding of the dictionary
based on the structure of the domain classification, the latter is relevant for the linked data
publication of the Morais Silva dictionary data, in which the URI of each domain in the classification
can be used for identifying and querying RDF data. Furthermore, the standalone modelling
and linked data publication of the domain classification facilitates its reuse, enabling, e.g., the
alignment with other KOS and domain classifications from other e-lexicography projects.
12The adjective ‘chulo’, to which the ch. abbreviation corresponds, is defined in the Morais Silva dictionary as “being
used in familiar conversation, joking, mocking or talking fresh, as they say” [30, vol. 1, p. 170].</p>
      <p>
        The list of domain labels appears in the front matter of the dictionary. To overcome the
deficiency of flat representation of labels in general-language dictionaries, TEI-Lex 0 now
recommends a hierarchical representation, an approach described in Salgado et al. [
        <xref ref-type="bibr" rid="ref18">17</xref>
        ], in
which the classification is included in the &lt;teiHeader&gt; element, as previously mentioned.
      </p>
      <p>The standalone modelling of the domain classification should be carried out in SKOS, the
W3C model for classifications and other KOS. This model has the advantage of allowing for
the straightforward modelling of hierarchies and associated networks of concepts, which can
be documented with notes, designated by lexical labels (either abbreviated or in full form)
and aligned with external KOS. The SKOS model can also easily be expanded with further
classes and properties for expressing more information13. The development of this SKOS
classification will be carried out in conjunction with the modelling of domain ontologies for
knowledge representation in the domains included in the dictionary’s list of abbreviations, of
which OntoDomLab-Med is the reference for medical and health sciences.</p>
      <p>A model is being developed for the consistent representation of information about each
domain in the SKOS classification. Figure 2 shows how the model is applied to the domain of
artillery (artilharia). In this example, we can see information about the designations of the
domain in the dictionary, most importantly the label(s) that occur in the lists of abbreviations
through the :usageLabel property. Both the abbreviated form and the full form, as explained
in the dictionary’s list of abbreviations, are included through standard RDF properties. This
example also shows the definition of artilharia taken from the dictionary’s first edition (definitions
from the other editions would also be included).
13The use of SKOS as an underlying model for the classification has the advantage of leveraging the relationship
with NOVA FCSH (School of Social Sciences and Humanities of NOVA University of Lisbon) through a controlled
vocabulary repository it manages in the context of the ROSSIO Infrastructure for social sciences, arts, and
humanities (http://vocabs.rossio.fcsh.unl.pt/).</p>
    </sec>
    <sec id="sec-6">
      <title>5. TEI Lex-0 encoding of usage information in lexicographic articles</title>
      <p>ical and health sciences (http://www.semanticweb.org/OntoDomLab-Med#Medicine
and http://www.semanticweb.org/OntoDomLab-Med#MedicalAndHealthSciences)
and concepts of the MorDigital Domain Classification in SKOS (e.g.,
http://vocabs.rossio.unl.pt/morais_domains/0025 for medicine).</p>
      <p>This approach allows for a more straightforward TEI encoding of dictionary articles, since the
alignment to external KOS and ontologies is centralised within the header element, requiring
less information in the articles themselves. The role of external KOS and ontologies remains
essential as the end results of terminological work. This enables both the hierarchical structuring
of the domain taxonomy and a richer conversion from TEI to linked data, as we will see in the
following section.</p>
    </sec>
    <sec id="sec-7">
      <title>6. Ontolex-Lemon modelling of usage information in lexicographic articles</title>
      <p>The core Ontolex-Lemon model already provides most of the necessary elements for modelling
information associated with lexical entries. The lexicographic module, lexicog, includes
additional elements for information pertaining to lexicographic articles. As we saw in Section 2.3, the
LexInfo ontology provides data categories for Ontolex, including several usage sub-properties
aligned with TEI Lex-0 (e.g., "domain", "socioCultural"). Finally, the domain classification
in SKOS allows to organise the subject fields corresponding to the domain labels and align them
with external knowledge bases and KOS.</p>
      <p>Figure 5 shows the metástase article modelled in Ontolex, along with elements from the
above-mentioned ontologies14. The senses are the key structural components of the article. As
such, they are modelled as lexicographic components, instances of lexicog’s respective class.
This allows to distinguish between the dictionary article and the metástase noun that it describes.
The noun’s senses are constrained to diferent domains (medicine and rhetoric), through the
lexinfo:domain object property. These domains are modelled as classes/concepts in external
KOS and ontologies, such as the OntoDomLab-Med ontology for medical and health sciences
and the MorDigital Domain Classification, whose URI are present in the taxonomy of domains
encoded in the TEI header, as seen in the previous section.</p>
    </sec>
    <sec id="sec-8">
      <title>7. Future work</title>
      <p>This paper provided an overview of usage information in lexicography and its respective
encoding, through TEI Lex-0, and linked data modelling by means of Ontolex-Lemon and
related ontologies. The work described in this paper is relevant in the context of the
TEIencoding and linked open data publishing of legacy dictionaries in the context of a digital
humanities project.</p>
      <p>
        Examples of lexicographic articles of the Morais Silva dictionary with usage information
were provided, along with ongoing work in the encoding and modelling of this information.
Further examples are being worked on to facilitate the conversion from the TEI Lex-0 encoding
to Ontolex-Lemon by means of a XSLT-based tool, such as tei2ontolex. More recent work on
the conversion between TEI Lex-0 and Ontolex in the context of the MorDigital project was
presented by Khan et al. [
        <xref ref-type="bibr" rid="ref32">31</xref>
        ]. This work is dependent on having quality TEI-encoded data,
obtained upstream, which is a fundamental task for the project.
14The representation of grammatical and semantic information was omitted to simplify the diagram.
      </p>
      <p>An important component of MorDigital pertains to the modelling of domain ontologies
covering the subject fields referred to by the domain labels of the Morais Silva dictionary. While
OntoDomLab-Med already covers domains in the medical and health sciences (e.g., medicine,
surgery, pharmacy), further work is being carried out in modelling the remaining domains15.
At the same time, work is being carried out in the SKOS modelling of the MorDigital domain
classification, whose first version was already published as linked open data 16. The latter started
from the non-hierarchical list of the 35 domains referred to in the Morais Silva dictionary’s
list of abbreviations, which will be gradually structured in articulation with the work carried
out for the domain ontologies. Further terminological and ontological work will be required
to put forward a list of superdomains and hierarchical structure for the domain classification.
These are major challenges in the project due to the inherent diversity of specialised fields
referred to in the lexicographic articles, including domains absent from the dictionary’s list of
abbreviations.</p>
    </sec>
    <sec id="sec-9">
      <title>Acknowledgments</title>
      <p>This paper is supported by (1) the MORDigital – Digitalização do Diccionario da Lingua
Portugueza de António de Morais Silva [PTDC/LLT-LIN/6841/2020] project financed by the
Portuguese National Funding through the FCT – Fundação para a Ciência e Tecnologia (2)
Portuguese National Funding through the FCT – Fundação para a Ciência e Tecnologia as part of
the project Centro de Linguística da Universidade NOVA de Lisboa – UID/LIN/03213/2020.
16Available from: https://vocabs.rossio.fcsh.unl.pt/morais_domains/</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J. P.</given-names>
            <surname>Silvestre</surname>
          </string-name>
          ,
          <article-title>Bluteau e as origens da lexicografia moderna</article-title>
          , INCM, Lisboa,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>T.</given-names>
            <surname>Verdelho</surname>
          </string-name>
          ,
          <string-name>
            <surname>O dicionário de Morais</surname>
          </string-name>
          <article-title>Silva e o início da lexicografia moderna, in: História da língua e história da gramática: actas do encontro</article-title>
          ,
          <source>Universidade do Minho, Braga</source>
          ,
          <year>2003</year>
          , pp.
          <fpage>473</fpage>
          -
          <lpage>490</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Correia</surname>
          </string-name>
          , Os dicionários portugueses, Caminho, Lisboa,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>R.</given-names>
            <surname>Costa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Salgado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Khan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Carvalho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Romary</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Almeida</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ramos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Khemakhem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Tasovac</surname>
          </string-name>
          , R. Silva,
          <article-title>MORDigital: the advent of a new lexicographical Portuguese project</article-title>
          , in: I. Kosem,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cukr</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jakubíček</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kallas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Krek</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          Tiberius (Eds.),
          <article-title>Electronic lexicography in the 21st century: post-editing lexicography</article-title>
          .
          <source>Proceedings of eLex</source>
          <year>2021</year>
          ,
          <string-name>
            <given-names>Lexical</given-names>
            <surname>Computing</surname>
          </string-name>
          , Brno,
          <year>2021</year>
          , pp.
          <fpage>312</fpage>
          -
          <lpage>324</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>T.</given-names>
            <surname>Tasovac</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Romary</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Banski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bowers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Does</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Depuydt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Erjavec</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Geyken</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Herold</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Hildenbrandt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Khemakhem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Lehečka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Petrović</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Salgado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Witt</surname>
          </string-name>
          , TEI Lex-
          <article-title>0: a baseline encoding for lexicographic data</article-title>
          .
          <source>Version 0.9.0</source>
          ,
          <year>2018</year>
          . URL: https: //dariah-eric.github.io/lexicalresources/pages/TEILex0/TEILex0.html.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <article-title>15For example, OntoDomLab-Math was recently developed for modelling the superdomain of mathematical sciences: https://github</article-title>
          .com/GuidaRamos/OntoDomLab-Math.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>L.</given-names>
            <surname>Romary</surname>
          </string-name>
          ,
          <article-title>TEI and LMF crosswalks</article-title>
          ,
          <source>Journal for language technology and computational linguistics 30</source>
          (
          <year>2015</year>
          )
          <fpage>47</fpage>
          -
          <lpage>70</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>P.</given-names>
            <surname>Cimiano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Buitelaar</surname>
          </string-name>
          , Lexicon Model for Ontologies:
          <source>Community Report, Technical Report</source>
          , W3C Ontology-Lexicon Community Group,
          <year>2016</year>
          . URL: https://www.w3. org/
          <year>2016</year>
          /05/ontolex/.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>P.</given-names>
            <surname>Cimiano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Chiarcos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gracia</surname>
          </string-name>
          ,
          <source>Linguistic Linked Data: Representation, Generation and Applications</source>
          , Springer, Berlin,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bosque-Gil</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gracia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Cimiano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Stolk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Khan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Depuydt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Does</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Frontini</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Kernerman</surname>
          </string-name>
          ,
          <article-title>The OntoLex Lemon Lexicography Module: final community report</article-title>
          ,
          <source>Technical Report</source>
          , W3C Ontology-Lexicon Community Group,
          <year>2019</year>
          . URL: https: //www.w3.org/
          <year>2019</year>
          /09/lexicog/.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>F.</given-names>
            <surname>Khan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bellandi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Boschetti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Monachini</surname>
          </string-name>
          ,
          <article-title>The Challenges of Converting Legacy Lexical Resources to Linked Open Data using Ontolex-Lemon</article-title>
          , in: LDK Workshops,
          <year>2017</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          . URL: http://ceur-ws.
          <source>org/</source>
          Vol-1899/OntoLex_2017_paper_4.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>R.</given-names>
            <surname>Stanković</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Stijović</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Vitas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Krstev</surname>
          </string-name>
          ,
          <string-name>
            <surname>O. Sabo,</surname>
          </string-name>
          <article-title>The Dictionary of the Serbian Academy: from the Text to the Lexical Database</article-title>
          ,
          <source>in: Proceedings of the XVIII EURALEX International Congress: Lexicography in Global Contexts</source>
          , Ljubljana University Press, Ljubljana,
          <year>2018</year>
          , pp.
          <fpage>941</fpage>
          -
          <lpage>949</lpage>
          . URL: https://dais.sanu.ac.rs/123456789/4927.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Tiberius</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Khan</surname>
          </string-name>
          , I. Kernerman,
          <string-name>
            <given-names>T.</given-names>
            <surname>Declerck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Krek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Monachini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ahmadi</surname>
          </string-name>
          ,
          <article-title>The ELEXIS Interface for Interoperable Lexical Resources</article-title>
          , in: I. Kosem,
          <string-name>
            <given-names>T. Zingano</given-names>
            <surname>Kuhn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Correia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ferreira</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jansen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pereira</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kallas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jakubíček</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Krek</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          Tiberius (Eds.),
          <source>Electronic lexicography in the 21st century: proceedings of the eLex 2019 conference, Lexical Computing, Brno</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>642</fpage>
          -
          <lpage>659</lpage>
          . URL: http: //hdl.handle.net/10379/15512.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>F.</given-names>
            <surname>Khan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Salgado</surname>
          </string-name>
          ,
          <article-title>Modelling Lexicographic Resources using CIDOC-CRM, FRBRoo</article-title>
          and Ontolex-Lemon, in: A.
          <string-name>
            <surname>Bikakis</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Ferrario</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Jean</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Markhof</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Mosca</surname>
            ,
            <given-names>M. N.</given-names>
          </string-name>
          <string-name>
            <surname>Asmundo</surname>
          </string-name>
          (Eds.),
          <source>Proceedings of the International Joint Workshop on Semantic Web</source>
          and
          <article-title>Ontology Design for Cultural Heritage co-located with the Bolzano Summer of Knowledge 2021 (BOSK 2021), CEUR-</article-title>
          <string-name>
            <surname>WS</surname>
          </string-name>
          ,
          <year>2021</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>12</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>TEI</given-names>
            <surname>Consortium</surname>
          </string-name>
          ,
          <article-title>TEI P5: Guidelines for Electronic Text Encoding and Interchange</article-title>
          .
          <source>Version 4.4.0. Last updated on 19th April</source>
          <year>2022</year>
          , revision
          <year>f9cc28b0</year>
          ,
          <year>2022</year>
          . URL: https://tei-c.org/ release/doc/tei-p5-doc/en/html/index.html.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>R.</given-names>
            <surname>Costa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Carvalho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Salgado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Simões</surname>
          </string-name>
          , T. Tasovac, Ontologie des marques de domaines appliquée aux dictionnaires de langue générale,
          <source>Langue(s) &amp; Parole</source>
          <volume>5</volume>
          (
          <year>2020</year>
          )
          <fpage>201</fpage>
          -
          <lpage>230</lpage>
          . URL: https://revistes.uab.cat/languesparole/article/view/v5-costa-etal.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>R.</given-names>
            <surname>Costa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Salgado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Almeida</surname>
          </string-name>
          ,
          <article-title>SKOS as a key element for linking lexicography to digital humanities</article-title>
          , in: K. Golub,
          <string-name>
            <given-names>Y. H.</given-names>
            <surname>Liu</surname>
          </string-name>
          (Eds.),
          <source>Information and Knowledge Organisation in Digital Humanities : Global Perspectives</source>
          , Routledge, London,
          <year>2021</year>
          , pp.
          <fpage>178</fpage>
          -
          <lpage>204</lpage>
          . URL: https://doi.org/10.4324/9781003131816-9.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>A.</given-names>
            <surname>Salgado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Costa</surname>
          </string-name>
          , T. Tasovac,
          <article-title>Applying terminological methods to lexicographic work: terms and their domains</article-title>
          , in: A.
          <string-name>
            <surname>Klosa-Kückelhaus</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Engelberg</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Möhrs</surname>
          </string-name>
          , P. Storjohann (Eds.),
          <source>Dictionaries and Society. Proceedings of the XX EURALEX International Congress</source>
          , IDS-Verlag, Mannheim,
          <year>2022</year>
          , pp.
          <fpage>181</fpage>
          -
          <lpage>195</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>U.</given-names>
            <surname>Weinreich</surname>
          </string-name>
          ,
          <article-title>Is a Structural Dialectology Possible?</article-title>
          ,
          <source>WORD</source>
          <volume>10</volume>
          (
          <year>1954</year>
          )
          <fpage>388</fpage>
          -
          <lpage>400</lpage>
          . doi:https: //www10.1080/00437956.
          <year>1954</year>
          .
          <volume>11659535</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>S.</given-names>
            <surname>Landau</surname>
          </string-name>
          ,
          <article-title>Dictionaries: the art and craft of lexicography</article-title>
          , 2nd ed ed., Cambridge University Press, Cambridge,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>B.</given-names>
            <surname>Svensén</surname>
          </string-name>
          ,
          <article-title>A handbook of lexicography: the theory and practice of dictionary-making</article-title>
          , Cambridge University Press, Cambridge,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>A.</given-names>
            <surname>Salgado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Costa</surname>
          </string-name>
          , T. Tasovac,
          <article-title>Improving the Consistency of Usage Labelling in Dictionaries with TEI Lex-0, Lexicography 6 (</article-title>
          <year>2019</year>
          )
          <fpage>133</fpage>
          -
          <lpage>156</lpage>
          . doi:
          <volume>10</volume>
          .1007/s40607- 019- 00061- x.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>M. O.</given-names>
            <surname>Ptaszyński</surname>
          </string-name>
          ,
          <article-title>Theoretical Considerations for the Improvement of Usage Labelling in Dictionaries:</article-title>
          A Combined
          <string-name>
            <surname>Formal-Functional</surname>
            <given-names>Approach</given-names>
          </string-name>
          ,
          <source>International Journal of Lexicography</source>
          <volume>23</volume>
          (
          <year>2010</year>
          )
          <fpage>411</fpage>
          -
          <lpage>442</lpage>
          . doi:
          <volume>10</volume>
          .1093/ijl/ecq029.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>M.</given-names>
            <surname>Vrbinc</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Vrbinc</surname>
          </string-name>
          ,
          <article-title>Diasystematic Information in the ”Big Five”: A Comparison of Print Dictionaries, CD-ROMS/ DVD-ROMS and Online Dictionaries</article-title>
          ,
          <source>Lexikos</source>
          <volume>25</volume>
          (
          <year>2015</year>
          )
          <fpage>424</fpage>
          -
          <lpage>445</lpage>
          . doi:
          <volume>10</volume>
          .5788/25- 1- 1306.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>F. J.</given-names>
            <surname>Hausmann</surname>
          </string-name>
          ,
          <article-title>Die Markierung in einem allgemeinen einsprachigen Wörterbuch: eine Übersicht</article-title>
          , in: F.
          <string-name>
            <given-names>J.</given-names>
            <surname>Hausmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Reichmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. E.</given-names>
            <surname>Wiegand</surname>
          </string-name>
          , L. Zgusta (Eds.),
          <source>Wörterbücher. Ein internationales Handbuch zur Lexikographie. Erster Teilband</source>
          , Walter de Gruyter, Berlin,
          <year>1989</year>
          , pp.
          <fpage>649</fpage>
          -
          <lpage>657</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>H.</given-names>
            <surname>Bergenholtz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Tarp</surname>
          </string-name>
          ,
          <source>Manual of Specialised Lexicography: The Preparation of Specialised Dictionaries</source>
          , John Benjamins, Amsterdam,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>J.</given-names>
            <surname>Milroy</surname>
          </string-name>
          , L. Milroy, Authority in Language: Investigating Standard English, Routledge, London,
          <year>1990</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>H.</given-names>
            <surname>Jackson</surname>
          </string-name>
          ,
          <string-name>
            <surname>Lexicography</surname>
          </string-name>
          : An Introduction, Routledge, London,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>B.</given-names>
            <surname>Atkins</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rundell</surname>
          </string-name>
          , The Oxford Guide to Practical Lexicography, Oxford University Press, New York,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [29]
          <article-title>ISO 24613-1, Language resource management - Lexical markup framework (LMF) - Part 1: Core model</article-title>
          , ISO, Geneva,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [30]
          <string-name>
            <surname>A. M. Morais Silva</surname>
          </string-name>
          ,
          <article-title>Diccionario da lingua portugueza composto pelo padre D. Rafael Bluteau, reformado</article-title>
          , e accrescentado por Antonio de Moraes Silva, natural do Rio de Janeiro, Oficina de Simão Thaddeo Ferreira, Lisboa, 1789. URL: https://purl.pt/29264.
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>F.</given-names>
            <surname>Khan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Salgado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Costa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Carvalho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Romary</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Almeida</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Khemakhem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Silva</surname>
          </string-name>
          , T. Tasovac,
          <article-title>Interlinking lexicographic data in the MORDigital project</article-title>
          , Mykolas Romeris University, Vilnius,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>