<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>LINGUATEC-IA: Research and development of Artificial Intelligence for the low-resourced languages of the Pyrenees</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Itziar Aldabe</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Itziar Aduriz</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Xabier Arregi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aitzol Astigarraga</string-name>
          <email>a.astigarraga@orai.eus</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Myriam Bras</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pauline Charrier</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Itziar Cortes</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Urtzi Etxeberria</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Igor Leturia</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Matthieu Martel</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Miguel Angel Pelles</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aure Séguier</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jordi Suïls</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>LAMPS - Université Perpignan Via Domitia</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Perpignan</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>HiTZ Center - Ixa, University of the Basque Country UPV/EHU</institution>
          ,
          <addr-line>Donostia</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Low resource languages, Large Language Models</institution>
          ,
          <addr-line>Speech Synthesis, Machine Translation, Cross-border linguistic</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Orai NLP Technologies - Elhuyar</institution>
          ,
          <addr-line>Usurbil</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>SoGEL - University of Lleida (UdL)</institution>
          ,
          <addr-line>Lleida</addr-line>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Université Toulouse-Jean Jaurès</institution>
          ,
          <addr-line>Toulouse</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2025</year>
      </pub-date>
      <abstract>
        <p>This paper presents the LINGUATEC-IA project, an ongoing initiative (2024-2026) aimed at advancing artificial intelligence applications for low-resource languages in the POCTEFA region (Aragonese, Catalan, Basque, and Occitan). The project focuses on developing specialized language models that can operate efectively with limited linguistic resources. The key objectives are to enhance transcription, machine translation, and speech synthesis systems for these languages in conjunction with French and Spanish; creating an automatic subtitling and dubbing platform; establishing an online repository for Pyrenean language resources; and strengthening a cross-border network for language technology excellence.</p>
      </abstract>
      <kwd-group>
        <kwd>infrastructure</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>The rapid advancement of artificial intelligence (AI) has revolutionized various fields, particularly
natural language processing (NLP). However, most research and technological development focus on
widely spoken languages, which has left many regional and minority languages underrepresented. The
LINGUATEC-IA project1 aims to address this gap by developing neural language models tailored for
low-resource languages in the POCTEFA region, specifically Aragonese, Catalan, Basque, and Occitan.
The primary goal is to improve digital accessibility and establish a cross-border language technology
infrastructure to support multilingual communication and information access. The LINGUATEC-IA
project is part of the Interreg VI-A Spain-France-Andorra Programme (POCTEFA 2021-2027), which
aims to strengthen the economic and social integration of the Spain-France-Andorra border region.</p>
      <p>Accordingly, the LINGUATEC-IA project addresses three key challenges of POCTEFA: 1) Innovation
Challenge 1 by increasing territorial innovation through AI and language technologies; 2) Territorial
Challenge 1 by supporting cultural integration via digitalization of minority languages (Aragonese,
Catalan, Basque, and Occitan), preserving cultural heritage while enabling communication across six</p>
      <p>CEUR
Workshop</p>
      <p>ISSN1613-0073
languages; and 3) Social Challenge 3 by strengthening cross-border labor markets through improved
language learning tools that facilitate worker mobility. The POCTEFA Programme recognizes
languages as a key territorial strength, and this project aims to develop these resources through targeted
technological innovation.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Background: LINGUATEC</title>
      <p>
        The European project EFA 227/16/LINGUATEC ”Development of cross-border cooperation and
knowledge transfer in language technologies” (2018-2020)2 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] established a consortium composed of 6 entities
working together in the development and dissemination of new innovative language resources, tools
and applications to improve the digitalization of Aragonese, Basque and Occitan.
      </p>
      <p>As a result, a series of tools, applications, and linguistic resources were developed designed to facilitate
communication and break language barriers across various languages.3 For Basque, technologies for
speech recognition and improvements in Spanish-Basque automatic translation were created, in addition
to Euskara Eskuz Esku, a tool for consulting linguistic norms. In Aragonese, notable developments
include voice synthesis, TRADUZE to improve automatic translation, the online dictionary
ARAGONARIO, and various multilingual applications oriented toward tourism and the Camino de Santiago. In
Occitan, monolingual and bilingual lexicons were developed, along with morphosyntactic and syntactic
analysis tools, as well as advanced voice synthesis systems (VOTZ) and speech recognition (ReVOc),
together with improvements in French-Occitan automatic translation. Finally, in the multilingual
domain, automatic translation applications between the languages of the Pyrenees were implemented,
available on Google Play and AppStore, as well as browser extensions and translation tools for websites
and content management systems (CMS). These developments represented a significant advance in
cooperation between languages and access to information in a more inclusive digital environment in
2020.</p>
      <p>The level of development achieved in the project encouraged the institutions belonging to the
consortium to take a strategic step: creating a network of excellence in AI, to create a cross-border
linguistic infrastructure. The LINGUATEC-IA project is a result of this network. This new initiative
builds upon the success of the previous project, aiming to develop more sophisticated natural language
processing capabilities. Thus, cooperation has been extended to new entities, languages, and territories,
and new goals have been set in terms of innovation.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Consortium</title>
      <p>The LINGUATEC-IA project brings together a diverse consortium of academic institutions and
organizations focused on developing language technologies for the mentioned languages. The LINGUATEC-IA
consortium is composed of 8 partners, with ELHUYAR serving as the lead coordinator and LO
CONGRÈS PERMANENT DE LA LENGA OCCITANA as co-coordinator. The partnership includes five
academic institutions (University of the Basque Country’s HITZ center, Université Toulouse Jean Jaurès,
Université Perpignan Via Domitia, IKER-CNRS, and Universitat de Lleida) and one government entity
(Government of Aragon).</p>
      <p>The project requires these 3 profiles, because it is not only about generating knowledge, but as
stated in POCTEFA, this knowledge must be transferred and applied to address the challenges of the
cross-border region. One of these challenges is the need to combine the preservation of all the languages
of the POCTEFA space with the improvement of the intercommunication between all of them, taking
advantage of the opportunities ofered by digitalization.</p>
      <p>The roles and specializations of the consortium members are outlined below:</p>
      <p>ELHUYAR4: The primary beneficiary and project coordinator, handling administrative and
2https://linguatec-poctefa.eu
3https://linguatec-poctefa.eu/recursos/
4https://www.orai.eus/
ifnancial management. Its AI center, Orai, specializes in machine translation, linguistic resources,
text mining, and speech technologies, with a focus on Basque and customised projects for
companies and organisations.</p>
      <p>LO CONGRÈS PERMANENT DE LA LENGA OCCITANA5: Co-coordinator and the
interregional regulatory body for Occitan, focused on strengthening knowledge and codification of the
language through various tools.</p>
      <p>HITZ (University of the Basque Country)6: A reference center for language technologies
with 80+ specialists focusing on AI for language and speech, particularly for Basque and other
low-resource languages.</p>
      <p>Université Toulouse Jean Jaurès7: A major center for Occitan language research and teaching,
contributing expertise through its OCRE research group.</p>
      <p>Université Perpignan Via Domitia8: Represented by LAMPS, contributing research in
automatic language processing for minority languages, especially Catalan.</p>
      <p>IKER-UMR54789: The only research center specializing in Basque studies in France, bringing
expertise in Basque linguistics and language processing.</p>
      <p>Universitat de Lleida10: The reference university for western Catalan territory, contributing
research on geographical and social language dynamics through its SoGeL research group.
Government of Aragon11: Participating for the first time through its General Directorate of
Cultural Heritage, promoting the inclusion of Aragonese in digital language technologies.</p>
      <p>This consortium represents a cross-border collaboration aimed at advancing language technologies
for the preservation and development of Pyrenean languages.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Objectives and Scope</title>
      <p>The project focuses on multiple key objectives that will drive technological progress and linguistic
preservation that builds upon the work started in LINGUATEC:
1. Research on the development of neural language models to suit low-resource language settings.
2. Improve transcription systems, neural machine translation, and speech synthesis for Aragonese,
Catalan, Basque, and Occitan. Similarly, to develop multilingual models that facilitate
communication between these languages and major languages such as French and Spanish.
3. Develop a prototype of an automatic subtitling and dubbing platform between the project’s
languages.
4. Create an online repository of resources, technologies, and applications for the languages of the</p>
      <p>Pyrenees.</p>
      <p>5. Consolidate the ’Cross-Border Network of Excellence in Language Technologies.’</p>
    </sec>
    <sec id="sec-5">
      <title>5. Methodology</title>
      <p>To achieve these ambitious goals, the project employs a structured methodology integrating collaborative
research, iterative development, and coordinated implementation. The common methodology for all
activities is based on a structured and collaborative approach, beginning with the establishment of
a Working Team formed by technicians from partners and external experts. This team developed a
5https://locongres.org/
6https://www.hitz.eus/
7https://www.univ-tlse2.fr/
8https://lamps.univ-perp.fr/
9https://iker.cnrs.fr/iker/
10https://www.sogel.udl.cat
11https://www.aragon.es/organismos/departamento-de-educacion-cultura-y-deporte/direccion-general-de-patrimonio-cultural
detailed Work Plan that defined the phases, schedule, and responsibilities, holding regular meetings
for follow-up and preparing semi-annual Progress Reports that have to be validated by the Technical
Working Group. The tools and applications are developed iteratively, with continuous feedback cycles
to refine outputs, generating progressively improved versions of each tool.</p>
      <p>To implement this methodology efectively, five strategic actions have been defined that cover all
aspects of the project: from management, communication, and dissemination, to the development and
implementation of the systems. These actions are designed to ensure eficient coordination between the
diferent components of the project, ensure fluid communication among all participants, maximize the
dissemination of results, facilitate the technical development of the systems, and allow for successful
implementation that meets the established objectives.</p>
      <p>ELHUYAR acts as the leading partner and leads Actions 1 (Management) and 5 (Dissemination). LO
CONGRES leads Actions 2 (Communication) and 4 (Development) and HiTZ leads Action 3 (Research).
The other partner entities lead some of the Activities within the actions and participates in cooperation
in all actions and activities. The actions are detailed in greater detail below.</p>
      <sec id="sec-5-1">
        <title>5.1. Action 1: Project Management</title>
        <p>Project Management is structured in three core activities. The administrative management establishes
the necessary structure through committees and working groups, secures technical assistance, and
develops essential planning documentation. The financial component maintains fiscal discipline through
regular monitoring, standardized reporting procedures, and careful tracking of ERDF funds while
ensuring regulatory compliance. Finally, the project monitoring keeps everything on track through
systematic progress reports, a comprehensive dashboard of indicators, and regular quality checks on
deliverables. This integrated framework allows the consortium to efectively manage all administrative,
legal, and financial aspects while remaining in accordance with the POCTEFA requirements.</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Action 2: Communication</title>
        <p>The objective of the action is to attract the interest of key target groups by efectively promoting the
solutions developed within the project, as well as sharing its activities and results. The aim is to ensure
a broad dissemination and visibility in a realistic, achievable, and measurable way. The primary target
audience includes researchers in linguistic technologies, individuals involved in regional language
development, local media, and the general public.</p>
      </sec>
      <sec id="sec-5-3">
        <title>5.3. Action 3: Research</title>
        <p>This action focuses on researching and developing language models for low-resource languages such as
Basque, Catalan, Occitan, and Aragonese. A key objective is to explore new AI techniques that reduce
the need for large datasets and computational resources. This is especially important given that not
all these languages have the same amount of available data. Some languages, such as Aragonese and
Occitan, are significantly more under-resourced, requiring tailored strategies and adapted tools. The
work is structured into four main activities: 1) gathering and preparing as much quality text as possible
in each language, including cleaning and organizing the data for training and evaluation; 2) developing
generative language models tailored to each language’s specific context, with a strong experimental
component to account for data scarcity; 3) evaluating these models using language-specific benchmarks
to measure accuracy, usefulness, and linguistic quality—focusing on automated evaluation wherever
possible; and 4), applying the models in real-world scenarios by building optimized, user-friendly
demonstrators that show their practical value and make them accessible to end users.</p>
      </sec>
      <sec id="sec-5-4">
        <title>5.4. Action 4: Development</title>
        <p>
          This action aims to develop innovative resources and applications that support the use and preservation
of the POCTEFA region’s languages by leveraging cutting-edge neural models. The goal is to improve
translation, transcription, and speech synthesis, fostering multilingualism and digital inclusion. The
work is organized into four main activities. First, the digital roadmaps for each language have been
updated using the European Language Equality [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] framework, identifying current resources and
outlining priorities for the next three years. Second, the project will create multimodal linguistic
resources—text and speech data—especially for the most under-resourced languages, ensuring models
have the material needed for quality training. Third, new tools and technologies will be developed or
improved to enhance digitalization and interoperability, such as better machine translation systems,
speech technologies, and neural models tailored to each language. Finally, practical applications will
be built to boost the real-world use of these languages, including translation and speech tools, and
platforms for subtitling and dubbing—contributing to a dynamic, multilingual digital ecosystem across
the region.
        </p>
      </sec>
      <sec id="sec-5-5">
        <title>5.5. Action 5: Dissemination and Consolidation</title>
        <p>This Action aims to share the project’s key results and establish a “Center of Excellence in Language
Technologies of the Pyrenees” by turning linguistic research into practical multilingual tools for the
POCTEFA region. It targets researchers, institutions, media, and the general public interested in
technologies like machine translation, speech synthesis, and multilingual communication. Building
on the LINGUATEC project, it expands collaboration among regional universities and seeks broader
international engagement. The Action includes developing a resource platform, hosting dissemination
events, and strengthening the already established cross-border excellence network.</p>
        <p>Although each Action is led by one entity, at least two partner entities participate in all Activities.
Additionally, the actions and the activities follow a logical process in which all are interrelated and
necessary to achieve the expected results.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Preliminary Analysis</title>
      <p>At the beginning of the project, it was considered appropriate to carry out an analysis of the availability
of data, resources, and tools that would enable research and development in AI focused on these
languages. The digital roadmap described for each language is the result of this need. The preliminary
analysis revealed that the amount of text corpora available is limited for the development of robust
large language models (LLMs), especially for Occitan and Aragonese. In contrast, speech and parallel
corpora are comparatively more abundant, enabling the development of more advanced speech and
machine translation (MT) technologies. Based on this analysis, actions 3 and 4 have been undertaken,
whose status and objectives are described below.</p>
      <sec id="sec-6-1">
        <title>6.1. Action 3: Research</title>
        <p>
          Regarding the development of generative language models, the initial analysis confirmed the original
hypothesis: the situation of Catalan and Basque is not comparable to that of Occitan and Aragonese.
In the cases of Catalan and Basque, not only are data and tools available, but there are also research
and technological centers specialized in the field of language-centered AI. These centers have already
developed generative language models for Catalan and Basque and are actively collaborating on various
projects [
          <xref ref-type="bibr" rid="ref3 ref4 ref5">3, 4, 5</xref>
          ].
        </p>
        <p>In contrast, Occitan and Aragonese lack comparable support. With regard to data availability, it
became evident that the size of the existing text corpora is extremely limited. Table 1, Table 2 and
Table 3 show the volume of accessible textual corpora for Aragonese and Occitan. In the case of Occitan,
the wide dialectal variety is worth highlighting, as it constitutes an additional dificulty.</p>
        <p>These corpora contain significantly fewer tokens than typically required for training robust language
models, even in the context of low-resource languages. This situation raises complex and
thoughtprovoking challenges for the project’s advancement.</p>
        <p>Therefore, with respect to the development of generative language models for Occitan and Aragonese,
several issues arise. Continued pre-training based on a multilingual foundational model might be an
adequate way to address this challenge, but it is essential to increase the size of the textual corpus to
enable this approach. One technique to augment the available data would be the generation of synthetic
data through machine translation. With regard to the training process, it initially seems advantageous
to start from models that perform well in French and Catalan, given the linguistic proximity of these
languages to Occitan and Aragonese.</p>
        <p>The evaluation of the models developed in the project is another significant challenge. The scarcity
of data, specifically the lack of evaluation datasets, is once again a crucial factor. Moreover, given the
limitations, scenarios will need to be defined to evaluate the performance of the models on specific
tasks, such as translation or text summarization.</p>
      </sec>
      <sec id="sec-6-2">
        <title>6.2. Action 4: Development</title>
        <p>This action focuses on the development of speech and MT technologies, leveraging the comparatively
greater availability of such corpora across the languages studied. The digital roadmaps for each language
have already been updated.</p>
        <p>Work is currently underway to collect and expand speech corpora, which will be used for training and
improving speech technologies. In the case of Basque, there are already substantial resources available to
support high-quality ASR (Automatic Speech Recognition) and TTS (Text-to-Speech) systems. Current
eforts aim to further enhance these systems, particularly for informal and dialectal speech, and to
create emotional and dialectal speech corpora for TTS applications. For Occitan and Aragonese, the
focus is on collecting new audio recordings along with transcriptions to support ASR development.
This involves both manual transcription of media content and crowdsourced contributions. Special
attention is being given to the Aranese variant of Occitan, for which new transcribed audio data is
being gathered. Additional TTS recordings using new voices, particularly in Aranese, are also being
planned.</p>
        <p>
          Technology development eforts include improvements to the existing ASR system for Basque,
especially in handling informal and dialectal speech. This includes the integration of new corpora
and a migration from Kaldi[
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] to Whisper[
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. For TTS, work is focused on emotional and dialectal
speech synthesis, as well as voice cloning. For Occitan and Aragonese, existing ASR systems[8]
will be upgraded by incorporating additional speech corpora and migrating to Whisper technology.
A dedicated ASR system will also be developed for the Aranese variant. Given Whisper’s robust
multilingual pretraining, it ofers excellent performance for low-resource languages when fine-tuned
with relatively small datasets. Existing TTS systems for Occitan[9] and Aragonese will be migrated to
the FastPitch-HiFiGAN framework, and support for the Aranese variant will be added. The project also
aims to improve voice quality and diversity across all Occitan variants by building a single, multispeaker,
multilingual TTS model.
        </p>
        <p>Regarding MT system development, current models already demonstrate solid performance, but future
work will focus on further enhancing their quality. This will be achieved by expanding training data
through a strategic combination of authentic and synthetic resources, leveraging automatic translation
of high-resource corpora. Continued training using transformer architectures and systematic evaluation
on established benchmarks will guide ongoing improvements.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>7. Expected Impact</title>
      <p>The implementation of this project is expected to bring benefits to both linguistic communities and the
ifeld of technological research. One of the key impacts will be to help on the digital preservation of
regional languages, as the development of AI-driven language tools will help preserve the relevance and
usability of Aragonese, Catalan, Basque, and Occitan in the digital era. This contributes to the digital
preservation of cultural heritage of these languages, making them accessible to future generations.
Additionally, the enhanced multilingual communication enabled by neural translation, speech synthesis,
and other technologies will foster smoother interactions among diferent linguistic groups, promoting
greater inclusivity and understanding.</p>
      <p>From a technological point of view, the project will develop new data and tools for low-resource
languages, where there is often a lack of such resources. The algorithms and models developed through
this initiative will not only serve the POCTEFA region but will also contribute to the global AI research
community, ofering new insights into the challenges and solutions for under-resourced languages.</p>
      <p>On a broader level, the project is expected to have a positive economic and social impact. By improving
linguistic accessibility, it will support cross-border collaboration, foster cultural exchange, and create
new economic opportunities within the POCTEFA region. This could ultimately help bridge gaps
between communities and enhance regional cohesion, contributing to a more integrated and dynamic
multilingual environment.</p>
    </sec>
    <sec id="sec-8">
      <title>8. Conclusion</title>
      <p>This project represents an efort to bridge the digital divide for low-resource languages using
cuttingedge AI technologies. By developing language models, enhancing translation and speech synthesis
systems, and fostering a cross-border linguistic infrastructure, this initiative plays a crucial role in the
preservation and modernization of the languages spoken across the POCTEFA region. As AI continues
to evolve, projects like this try to ensure that all languages, regardless of their resource availability,
have a place in the digital world.</p>
    </sec>
    <sec id="sec-9">
      <title>Acknowledgments</title>
      <p>The LINGUATEC-IA project has been 65% co-financed by the European Union through the Interreg VI-A
Spain-France-Andorra Programme (POCTEFA 2021-2027). The objective of POCTEFA is to strengthen
the economic and social integration of the Spain-France-Andorra border region.</p>
    </sec>
    <sec id="sec-10">
      <title>Declaration on Generative AI</title>
      <p>The author(s) have not employed any Generative AI tools.
Proceedings of Machine Learning Research, PMLR, 2023, pp. 28492–28518. URL: https://proceedings.
mlr.press/v202/radford23a.html.
[8] I. Morcillo, I. Leturia, A. Corral, X. Sarasola, M. Barret, A. Séguier, B. Dazéas, Automatic speech
recognition for gascon and languedocian variants of occitan, in: Proceedings of the 2024 Joint
International Conference on Computational Linguistics, Language Resources and Evaluation
(LRECCOLING 2024), 2024, pp. 1969–1978.
[9] A. Corral, I. Leturia, A. Séguier, M. Barret, B. Dazéas, P. B. de Mareüil, N. Quint, Neural
text-tospeech synthesis for an under-resourced language in a diglossic environment: the case of gascon
occitan, in: Proceedings of the 1st Joint SLTU (Spoken Language Technologies for Under-resourced
languages) and CCURL (Collaboration and Computing for Under-Resourced Languages) Workshop
«Language Resources and Evaluation Conference–Marseille–11–16 May 2020», European Language
Resources Association (ELRA), 2020.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>I.</given-names>
            <surname>Aldabe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Aztiria</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Beltrán</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ceberio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. Cortes</given-names>
            ,
            <surname>J.-B. Coyos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Dazeas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Esher</surname>
          </string-name>
          , G. Labaka, I. Leturia,
          <string-name>
            <given-names>K.</given-names>
            <surname>Sarasola</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Séguier</surname>
          </string-name>
          , J. Sibille, LINGUATEC: Desarrollo de recursos lingüísticos para avanzar en la digitalización de las lenguas de los pirineos,
          <source>Procesamiento del lenguaje natural 63</source>
          (
          <year>2019</year>
          )
          <fpage>159</fpage>
          -
          <lpage>162</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>I.</given-names>
            <surname>Aldabe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dunne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Farwell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Gallagher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Gaspari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Giagkou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hajic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>Kückens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lynn</surname>
          </string-name>
          , G. Rehm, G. Rigau,
          <string-name>
            <given-names>K.</given-names>
            <surname>Marheinecke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Piperidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Resende</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Vojtěchová</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Way</surname>
          </string-name>
          ,
          <article-title>Overview of the ELE project</article-title>
          ,
          <source>in: Proceedings of the 23rd Annual Conference of the European Association for Machine Translation</source>
          , p.
          <fpage>353</fpage>
          -
          <lpage>354</lpage>
          ,
          <year>2022</year>
          . URL: https://aclanthology.org/
          <year>2022</year>
          .eamt-
          <volume>1</volume>
          .66/.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Gonzalez-Agirre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pàmies</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Llop</surname>
          </string-name>
          , I. Baucells,
          <string-name>
            <given-names>S. D.</given-names>
            <surname>Dalt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Tamayo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. J.</given-names>
            <surname>Saiz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Espuña</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Prats</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Aula-Blasco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rubio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Shvets</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sallés</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Lacunza</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Pikabea</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Palomar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Falcão</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Tormo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Vasquez-Reina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Marimon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Ruíz-Fernández</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Villegas</surname>
          </string-name>
          ,
          <source>Salamandra technical report</source>
          ,
          <year>2025</year>
          . URL: https://arxiv.org/abs/2502.08489. arXiv:
          <volume>2502</volume>
          .
          <fpage>08489</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Corral</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. S.</given-names>
            <surname>Antero</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Saralegi</surname>
          </string-name>
          ,
          <article-title>Pipeline analysis for developing instruct LLMs in low-resource languages: A case study on Basque</article-title>
          , in: L.
          <string-name>
            <surname>Chiruzzo</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Ritter</surname>
          </string-name>
          , L. Wang (Eds.),
          <source>Proceedings of the</source>
          <year>2025</year>
          <article-title>Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), Association for Computational Linguistics</article-title>
          , Albuquerque, New Mexico,
          <year>2025</year>
          , pp.
          <fpage>12636</fpage>
          -
          <lpage>12655</lpage>
          . URL: https://aclanthology.org/
          <year>2025</year>
          .
          <article-title>naacl-long</article-title>
          .
          <volume>629</volume>
          /.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Etxaniz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Sainz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Miguel</surname>
          </string-name>
          , I. Aldabe,
          <string-name>
            <given-names>G.</given-names>
            <surname>Rigau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Agirre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ormazabal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Artetxe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Soroa</surname>
          </string-name>
          ,
          <string-name>
            <surname>Latxa:</surname>
          </string-name>
          <article-title>An open language model and evaluation suite for Basque</article-title>
          , in: L.
          <string-name>
            <surname>-W. Ku</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Martins</surname>
          </string-name>
          , V. Srikumar (Eds.),
          <source>Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume</source>
          <volume>1</volume>
          :
          <string-name>
            <surname>Long</surname>
            <given-names>Papers)</given-names>
          </string-name>
          ,
          <source>Association for Computational Linguistics</source>
          , Bangkok, Thailand,
          <year>2024</year>
          , pp.
          <fpage>14952</fpage>
          -
          <lpage>14972</lpage>
          . URL: https://aclanthology.org/
          <year>2024</year>
          .
          <article-title>acl-long</article-title>
          .
          <volume>799</volume>
          /. doi:
          <volume>10</volume>
          .18653/v1/
          <year>2024</year>
          .acl- long.799.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>D.</given-names>
            <surname>Povey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ghoshal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Boulianne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Burget</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Glembek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hannemann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Motlicek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Qian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Schwarz</surname>
          </string-name>
          , et al.,
          <article-title>The Kaldi speech recognition toolkit</article-title>
          ,
          <source>in: IEEE 2011 Workshop on Automatic Speech Recognition and Understanding</source>
          ,
          <source>IEEE Signal Processing Society</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Radford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. W.</given-names>
            <surname>Kim</surname>
          </string-name>
          , T. Xu,
          <string-name>
            <given-names>G.</given-names>
            <surname>Brockman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Mcleavey</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Sutskever</surname>
          </string-name>
          ,
          <article-title>Robust speech recognition via large-scale weak supervision</article-title>
          , in: A.
          <string-name>
            <surname>Krause</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Brunskill</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Cho</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Engelhardt</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Sabato</surname>
          </string-name>
          , J. Scarlett (Eds.),
          <source>Proceedings of the 40th International Conference on Machine Learning</source>
          , volume
          <volume>202</volume>
          of
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>