<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>HARTAes-vas: Lexical combinations for an academic writing aid tool in Spanish and Basque</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Margarita Alonso-Ramos</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Igone Zabala</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Universidad del País Vasco/Euskal Herriko Unibertsitatea</institution>
          ,
          <addr-line>Barrio Sarriena s/n, Leioa, 48940</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Universidade da Coruña and CITIC</institution>
          ,
          <addr-line>Campus da Zapateira s/n, A Coruña, 15071</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <fpage>22</fpage>
      <lpage>25</lpage>
      <abstract>
        <p>Academic writing has become a priority object of study especially in English, for which there are already many resources to help novice writers. This is not the case for Spanish university students who do not have many writing aids at their disposal. Here we focus on routinized lexical combinations that characterise academic discourse in Spanish and Basque. The aim is to extract these combinations from two academic corpora in order to build a writing aid tool serving both languages.</p>
      </abstract>
      <kwd-group>
        <kwd>1 Academic writing</kwd>
        <kwd>collocations</kwd>
        <kwd>discourse functions</kwd>
        <kwd>writing aid</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>The HARTAes-vas project is funded by the
Ministry of Science and Innovation in the 2019
call for R&amp;D Knowledge Generation Projects. It
is a project coordinated between the Universidad
del País Vasco / Euskal Herriko Unibertsitatea
(UPV/EHU) and the Universidade da Coruña
(UDC) and, in some objectives, it is a continuation
of previous projects related to academic writing in
Spanish. In this new project, we are tackling a
contrastive approach with two different languages
from both a typological and a sociolinguistic point
of view. The research team is made up of
members of the LyS group at the UDC and the Ixa
group at the UPV/EHU together with researchers
from the Foundation Elhuyar.</p>
      <p>
        In recent years, academic writing has become
a priority object of study, especially in English
([
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] among others). In order for members of
the academic community to produce knowledge,
they must be able to write in the conventional
forms of academic texts. However, when students
enter university, they are confronted with new
written genres for which they are not provided
with tools to facilitate the production of texts.
Moreover, university students in Spain must be
able to show proficiency in several languages and,
paradoxically, Spanish students have more
resources to help them with academic English
than with the other languages of the state. One of
the keys to this competence in writing lies in the
mastery of certain routine expressions that give it
its specific character: academic lexical
combination (ALC), ranging from collocations
(extraer conclusiones, ondorioak atera ‘draw
conclusions’), to discourse markers (en
conclusión, ondorioz ‘in conclusion’) and also
formulas such as parece razonable concluir que
(‘it seems reasonable to conclude that’), ondorioz
esan daiteke (‘consequently we can say’); all
these expressions are ALC which we have in order
to express a conclusion in Spanish and Basque.
      </p>
      <p>
        Before developing the tool that would help the
students learn to write in this academic style, a
diagnosis of the current written productions of our
university students is needed. In previous research
we have compiled a corpus of written productions
of Spanish academic novices made up of
Bachelor’s and Master’s theses ([
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]; hereafter
the Spanish novice corpus) and during this project
we have compiled a comparable corpus of written
productions of academic novices for Basque
(hereafter the Basque novice corpus). The
different sociolinguistic status of Basque with
respect to Spanish forces different strategies: on
the one hand, there is no academic corpus of
expert academic writing in Basque available as a
reference; on the other hand, Basque has not had
enough time for the stabilisation of academic
registers [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], which suggests as a starting
hypothesis that ALCs will have a lower degree of
fixation and recurrence. Likewise, the
agglutinative nature of Basque poses a challenge
to the usual techniques for extracting
combinations.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Goals</title>
      <p>The overall goal is to create a bilingual tool (or
two coordinated monolingual tools), focused on
the use of ALCs, combining a dictionary and a
corpus. We aim to build a tool where the user can
choose the language and find help in choosing the
appropriate lexical strategies according to
different discourse needs.</p>
      <p>More specifically, the project aims to:
• develop a model of ALCs that includes
the characteristics of agglutinative languages
such as Basque where different lexicographic
and discursive classifications will be
established;
• analyse the learners' use of such
combinations in Spanish and Basque;
• investigate what kind of help related to
the phenomena of lexical combinations they
need when writing;
• develop corpus-based linguistic
technologies for the automatic identification of
ALCs.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology</title>
      <p>The project has multiple orientations:
lexicological (as far as the linguistic phenomena
studied are concerned); corpus linguistics and
computational linguistics (insofar as the corpora
are the fundamental source of data and the
techniques with which they are exploited come
from NLP) and didactics (following the approach
of so-called computer-assisted language learning
and, more particularly, the data-driven learning
methodology).</p>
      <p>The agglutinative nature of Basque inspired
the design of alternative ALC identification
techniques since the usual lexical bundle
extraction technique is not suitable in all cases for
Basque. The reason is that some formulas are
made up of a single word in Basque and it is
necessary to take into account the so-called
morphemic bundles to complement the results
obtained with the techniques used for inflectional
languages. For example: en resumen ‘in short’
laburbilduz ‘short+gather+INSTR’; por
consiguiente ‘therefore’- ondorioz ‘consequence
+ INSTR’.</p>
    </sec>
    <sec id="sec-4">
      <title>3.1. Extracting academic vocabulary lists with corpus linguistics and NLP techniques</title>
      <p>
        We analysed the Spanish novice corpus
morphologically and syntactically to extract
collocations with LinguaKit, Freeling and
UDPipe, following the same criteria we used in
the expert corpus [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. We extracted the following
syntactic patterns: Subject-Verb (objetivo se
centra ‘objective focuses’), Verb-Object
(alcanzar objetivo ‘reach an objective’),
NounModifier (objetivo fundamental ‘main objective’),
N of N (serie de objetivos ‘series of objectives’).
We also extracted lists of n-grams, applying
criteria of frequency and distribution by scientific
domains and assigned the discursive function
according to the typology established in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        A similar procedure was applied to the Basque
novice corpus which was morphologically
analysed using Eustagger. We started by
extracting an academic vocabulary based on the
criteria defined in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. We have used this word list
to identify collocations, without the need to
syntactically analyse the corpus [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. We have
extracted the following syntactic patterns:
Subject-Verb (datuek erakutsi 'data show'),
VerbObject (datuak bildu 'collect data', datuetan
oinarritu 'rely on data'), Noun-Modifier (datu
esanguratsu 'significant data'), N-N (datu sorta
'data set', datu-bilketa 'data collection'). To obtain
the formulas, we extracted lists of n-grams,
applying the same criteria of frequency and
dispersion and the same typology of discursive
functions described in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Once the formula
candidates have been validated, the variation was
analysed in order to identify prototypical formulas
and their variants.
      </p>
    </sec>
    <sec id="sec-5">
      <title>3.2. Testing semantics strategies distributional</title>
      <p>
        Once the two corpora of Spanish and Basque
novice academic writing are balanced, we can
exploit them as comparable corpora and apply
computational techniques of distributional
semantics in order to find correspondences
between the formulas of the two languages. With
the Spanish list, vector representations
(embeddings) of each formula can be generated
using non-compositional strategies, and we can
then use them to identify the Basque single word
equivalents of Spanish expressions in a previously
obtained cross-linguistic semantic space. In this
way, we may be able to relate por consiguiente
and ondorioz, or para terminar ‘to conclude’ and
bukatzeko, following the non-compositional
strategy used by [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
      <p>
        Monolingual distributional models, both
monolexical and polylexical, will be generated
with fastText, and mapped to a multilingual space
with vecmap. Since we find both compositional
and non-compositional expressions among the
formulas, we will use equivalent search strategies
adapted to each type of structure. For the
noncompositional ones, we will represent each
formula with a single vector, using the
noncompositional method presented in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. We
consider that the use of this multilingual strategy
can help in the identification of formulas, because
if a Basque expression has a high degree of both
internal cohesion and distributional similarity
with a Spanish formula, the probability that it is
indeed a formula in Basque is also very high.
Likewise, it seems interesting to explore whether
distributional models also identify a more
discursive meaning, such as that of the formulas.
      </p>
    </sec>
    <sec id="sec-6">
      <title>4. Results</title>
      <p>The quantitative data from the Spanish novice
corpus analysis are shown in Table 1. The data are
presented with normalised frequency per million
words due to the different size of the corpora.</p>
      <p>The results of a contrastive analysis with the
expert corpus show that novices use fewer
collocations than experts. Also, novices use more
collocations belonging to the general language.
With respect to formulas, we see that novices use
fewer types than experts, but almost as many
tokens</p>
      <p>
        As far as Basque is concerned, we have already
achieved the compilation of a corpus of novice
academic writing [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Although its analysis has
not yet been completed, we can already observe
some characteristics: the ALCs are less stable
compared to the Spanish novel corpus and a
higher number of ALCs are considered incorrect.
By validating the lists of ALCs in the Basque
corpus, we will be able to make a more thorough
comparison: contrasting formulas by functions
and verifying whether the same functions are
covered in the two languages and checking
whether the equivalent bases are linked to more or
fewer collocates in the different languages. This
comparison will be vital for the design of the
writing aid tool. Pending the aforementioned
further analysis, the quantitative data are shown in
Table 2.
      </p>
    </sec>
    <sec id="sec-7">
      <title>5. Conclusions and future work</title>
      <p>We have presented the main tasks we carried
out to obtain the data for an academic writing aid
tool. Next, we will explore the transfer strategies
for the automatic identification of ALCs in several
languages. We start from the hypothesis that a
cross-linguistic language model trained to identify
the formulas in Spanish could recognise
expressions with similar characteristics in
Basque. If the results obtained with this strategy
are adequate, we could, on the one hand,
automatically obtain new formulas in both
languages in other corpora and, on the other hand,
identify formulas in Basque that could be mapped
to those in Spanish. Pending the results of the
experiments with distributional semantics
techniques, we are making progress in the design
of the tool, which must meet two requirements: 1)
provide onomasiological access by discursive
function; 2) include a field of warnings where
examples will be provided as correction models.</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgements</title>
      <p>This work has been supported by the Xunta de
Galicia, through grant ED431C 2020/11, by the
CITIC of the UDC through grant ED431G 2019/0
and by the Spanish Ministry of Science and
Innovation through projects
PID2019109683GB-C21 and PID2019-109683GB-C22. I
would like to thank Olga Zamaraeva for her
valuable and constructive suggestions.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>K.</given-names>
            <surname>Hyland</surname>
          </string-name>
          , P. Shaw (Eds.)
          <article-title>The Routledge Handbook of English for Academic Purposes</article-title>
          , Routledge, London,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>K.</given-names>
            <surname>Tusting</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>McCulloch</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Bhatt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hamilton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Barton</surname>
          </string-name>
          , Academics Writing:
          <article-title>The Dynamics of Knowledge Creation, Routledge</article-title>
          , Abingdon, NY,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Alonso-Ramos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>García-Salido</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Garcia</surname>
          </string-name>
          ,
          <article-title>Exploiting a corpus to compile a lexical resource for academic writing: Spanish lexical combinations</article-title>
          , in: I.
          <string-name>
            <surname>Kosem</surname>
          </string-name>
          , et al. (Eds.),
          <source>Electronic Lexicography in the 21st Century. Proceedings of eLex 2017 Conference, Lexical Computing Brno</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>571</fpage>
          -
          <lpage>586</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>García-Salido</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Garcia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Villayandre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Alonso-Ramos</surname>
          </string-name>
          ,
          <article-title>A Lexical Tool for Academic Writing in Spanish based on Expert and Novice Corpora</article-title>
          , in: N.
          <string-name>
            <surname>Calzolari</surname>
          </string-name>
          et al. (Eds.),
          <source>Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC</source>
          <year>2018</year>
          ), Miyazaki,
          <year>2018</year>
          , pp.
          <fpage>260</fpage>
          -
          <lpage>265</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>I.</given-names>
            <surname>Zabala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.J.</given-names>
            <surname>Aranzabe</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Aldezabal</surname>
          </string-name>
          ,
          <article-title>Retos actuales del desarrollo y aprendizaje de los registros académicos orales y escritos del euskera</article-title>
          ,
          <source>Círculo de Lingüística Aplicada a la Comunicación</source>
          <volume>88</volume>
          (
          <year>2021</year>
          )
          <fpage>31</fpage>
          -
          <lpage>50</lpage>
          . doi:
          <volume>10</volume>
          .5209/clac.78295.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>García-Salido</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Garcia</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>AlonsoRamos, Identifying lexical bundles for an academic writing assistant in Spanish</article-title>
          , in: G. Corpas Pastor, R. Mitkov (Eds.),
          <source>Computational and Corpus-Based Phraseology. Europhras</source>
          <year>2019</year>
          , volume
          <volume>11755</volume>
          of Lecture Notes in Computer Sciences, Springer, Cham,
          <year>2019</year>
          , pp.
          <fpage>144</fpage>
          -
          <lpage>158</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -30135-4_
          <fpage>11</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>García-Salido</surname>
          </string-name>
          ,
          <source>Compiling an Academic Vocabulary List of Spanish. Available at: doi:10.13140/RG.2.2.27681.33123.</source>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Gurrutxaga</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Alegria</surname>
          </string-name>
          ,
          <article-title>Automatic extraction of NV expressions in Basque: Basic issues on cooccurrence techniques</article-title>
          ,
          <source>in: Proceedings of the Workshop</source>
          on Multiword Expressions:
          <article-title>from parsing and generation to the real world, Association for Computational Linguistics</article-title>
          , Portland,
          <year>2011</year>
          , pp.
          <fpage>2</fpage>
          -
          <lpage>7</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Garcia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>García-Salido</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>AlonsoRamos, Weighted compositional vectors for translating collocations using monolingual corpora</article-title>
          , in: G. Corpas Pastor, R. Mitkov (Eds.),
          <source>Computational and Corpus-Based Phraseology. Europhras</source>
          <year>2019</year>
          , volume
          <volume>11755</volume>
          of Lecture Notes in Computer Sciences, Springer, Cham,
          <year>2019</year>
          , pp.
          <fpage>113</fpage>
          -
          <lpage>128</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -30135-
          <issue>4</issue>
          _
          <fpage>9</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>M. J. Aranzabe</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Gurrutxaga</surname>
            ,
            <given-names>I. Zabala</given-names>
          </string-name>
          ,
          <article-title>Compilación del corpus académico de noveles en euskera HARTAvas y su explotación para el estudio de la fraseología académica</article-title>
          .
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>69</volume>
          (
          <year>2022</year>
          )
          <fpage>95</fpage>
          -
          <lpage>103</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>