<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>SSN-MLRG at Text to Picto 2024: A BERT-Based Approach for Mapping French Sentences to Pictogram Terms</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Bhavana Anand</string-name>
          <email>bhavana2110584@ssn.edu.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Themozhi J</string-name>
          <email>themozhi2110992@ssn.edu.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shreyas Sai R</string-name>
          <email>shreyassai2110425@ssn.edu.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Charumathi P</string-name>
          <email>charumathi2110213@ssn.edu.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mirnalinee TT</string-name>
          <email>mirnalineett@ssn.edu.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Sri Sivasubramaniya Nadar College of Engineering</institution>
          ,
          <addr-line>Chennai, Tamil Nadu</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <abstract>
        <p>Language impairment arising from diferent factors such as genetic diseases or incidents such as a car accident or stroke can impair language development skills and may lead to a partial or complete loss of the ability to communicate in written or spoken language. Augmentative and Alternative Communication refers to the means of communication used to substitute or replace spoken or written language. This paper focuses on the semantic mapping of French sentences to corresponding AAC pictograms. For this task, a transformer model utilizing CamemBERT embeddings, a French BERT model fused with a contrastive learning technique, was implemented. The model has obtained a PictoER score of 141.909, a BLEU score of 3.419, and a METEOR score of 14.351.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Augmentative and Alternative Communication</kwd>
        <kwd>Semantic Mapping</kwd>
        <kwd>Transformer Models</kwd>
        <kwd>CamemBERT</kwd>
        <kwd>Natural Language Processing</kwd>
        <kwd>Multimodal Communication</kwd>
        <kwd>Contrastive Learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Communication is the act of giving and receiving information about that person’s needs, desires,
perceptions, knowledge, or afective states of another person. Language is the structured conventional
system that is used to communicate with one another. AAC, by definition, is a therapeutic approach [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]
that employs manual signs, symbol-based communication boards, and speech-generating computerized
devices, integrating a person’s entire spectrum of communication abilities.
      </p>
      <p>Pictograms are visual communication tools that adeptly convey meanings, particularly efective in
disambiguating homophones and other linguistically confounding terms. They provide the advantage
of enabling communication from a foundational level—suitable for individuals with low cognitive
abilities or those in early developmental stages—to a rich and advanced level although not with the
same completeness and flexibility as written language [2].</p>
      <p>The motivation behind this research stems from the communication gap between AAC users and
others in society. There is a lack of awareness about such alternative communication methods. This
leads to the need for a tool that converts modalities like speech and text into a sequence of pictograms
to bridge this divide.</p>
      <p>Inspired by the successful application of transformer models in natural language processing, we
participated in the ImageCLEF 2024 ToPicto task [3] under the ImageCLEF 2024 evaluation campaign
[4]. This task involves the semantic mapping of French sentences to AAC pictograms. The data set for
this task was constructed from the Traitement de Corpus Oraux en Français (TCOF) corpus [5], which
features a wide range of conversations, including arguments, commonplace events, and medical advice
among various demographics in French. The translation process involves transforming raw French
source sentences into their corresponding target sentences, where each word is reduced to its root form
and semantically mapped to the most suitable pictogram.</p>
      <p>Section 1 provides an overview of the need for text-to-pictogram translation tools. Section 2 discusses
existing research work in this area. Section 3 discusses the provided dataset. Section 4 presents
the methodology of the proposed model. Section 5 showcases the results and includes a discussion
summarizing the key findings and highlighting potential future research directions.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Works</title>
      <p>In the domain of AAC (Augmentative and Alternative Communication), the translation of text into
pictograms is essential to assist individuals with communication impairments. Historically, various
technologies have been explored, ranging from rule-based systems to sophisticated neural network
approaches, each aimed at enhancing the eficacy and accessibility of AAC solutions [6].</p>
      <p>An innovative approach is Sevens et al.’s text-to-pictogram translation system [7] for emails, which
leverages lexical-semantic databases to map Dutch text to pictograms using both direct and semantic
translation routes to address linguistic challenges. The AraTraductor [8] system employs syntactic
analysis to improve the accuracy of pictogram generation from text, showcasing the impact of syntactic
parsing on the understandability of AAC content. Further, PictoBERT [9], a BERT derivative, predicts
pictogram sequences on AAC boards by adapting transformer architecture to utilize word-sense data,
moving beyond traditional n-gram models. The PrAACT [10] methodology adapts large transformer
models for AAC, focusing on customization and adaptability, thus improving user-specific
communication. Additionally, the BabelDr system [11]integrates the Unified Medical Language System (UMLS)
with neural machine translation techniques to convert medical dialogue into pictograms, enhancing
patient-doctor communication. This system streamlines complex medical interactions into pictographs
by leveraging a neural architecture that parses and translates speech into UMLS-based semantic glosses,
significantly improving understanding through intuitive visual representations.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Dataset</title>
      <p>The dataset for the Text to Pictogram task was built from the Traitement de Corpus Oraux en Français
(TCOF) Corpus [5]. Daily life conversations encompassing a wide range of categories between AAC
users and their caregivers are present in this dataset.</p>
      <p>In the train dataset, for each unique utterance, the dataset consists of a source sentence that is
transcribed from speech. Each of these utterances is characterised by a unique ID. The source sentences
are converted to target sentences that typically represent the base form of each word in the source
corresponding to a sequence of pictogram terms. These words are then mapped to a list of pictogram
identifiers linked to each pictogram term. The pictograms are taken from ARASAAC, a collection
featuring over 25,000 pictograms. The test dataset contains a series of source sentences that are oral
transcriptions of utterances and the unique ID corresponding to them. The proposed model is used to
generate a hypothesis of target sentences obtained for each of these utterances.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Methodology</title>
      <p>This paper introduces a unique application of contrastive learning and transformer-based language
models to improve the semantic mapping of French sentences to AAC pictograms. This method
enhances the capabilities of CamemBERT, improving its adaptability to various linguistic contexts and
its learning ability from minimal examples. Our approach addresses the limitations of previous systems
by efectively handling complex semantic relationships and adapting to diverse linguistic contexts
without the need for extensive training data.</p>
      <sec id="sec-4-1">
        <title>4.1. Embeddings</title>
        <p>The proposed is anchored in the transformative capabilities of BERT (Bidirectional Encoder
Representations from Transformers) [12], a groundbreaking model in natural language understanding. The dataset
is tokenized and passed to the CamemBERT model [13] to generate fixed-size embedding vectors (768 for
the proposed model). A compressed version of the CamemBERT model, specifically designed for French
language processing tasks. Distilled CamemBERT [14] inherits its capabilities from CamemBERT, which
is built upon the RoBERTa architecture [15] and trained on a large corpus of French text, achieving
state-of-the-art performance in various NLP tasks. The key innovation of BERT lies in its ability to
capture complex linguistic patterns and semantics in French text, utilizing bidirectional Transformers,
self-attention mechanisms, and dynamic masking techniques that are utilized by the proposed model.</p>
        <p>The purpose of distillation is to drastically reduce the complexity of the model while preserving
the performance. The training objective of the student model, Distilled CamemBERT, is to closely
approximate the behavior of the teacher, CamemBERT model. The training objective is composed of
3 parts, DistillLoss, Cosine Loss, and MLM Loss. The distillLoss measures the similarity between the
output probabilities of the student and teacher models using cross-entropy loss on the Masked Language
Modeling (MLM) task while the Cosine Loss Ensures alignment between the last hidden layers of the
student and teacher models by computing the cosine similarity. The MLM loss implements the MLM
task loss to train the student model according to the original task of the teacher model. The objective
function is thus a combination of the 3 losses.</p>
        <p>= 0.5 ×  + 0.3 ×  + 0.2 ×</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Contrastive Learning</title>
        <p>These embeddings are then passed on to a Neural Network, which reduces the embedding size to
384. The objective of this Neural Network is to learn a new representation such that it minimizes the
diference between similar representations, which in our case are the source and target representations.
Contrastive learning techniques are used for this task. Contrastive learning is a paradigm that focuses
on learning representations by contrasting positive pairs of similar instances against negative pairs of
dissimilar instances. It is based on the principle that semantically similar instances should be closer
together in the learned representation space compared to dissimilar ones. By optimizing a contrastive
loss function, the model learns to distinguish between positive and negative pairs, efectively pulling
similar instances closer while pushing dissimilar ones apart [16]. This approach has gained significant
traction due to its efectiveness in learning robust representations from unlabeled data, particularly for
vision and language tasks. Here, a contrastive loss function is used which penalizes larger distances
between similar vector embeddings and smaller distances between dissimilar vector embeddings. The
loss function is given as</p>
        <p>= (1 − ) × 2 +  × (0,  − )2
where y is the label for similar and dissimilar pairs, d is the Euclidean distance between the 2
embeddings and m is the margin, which is a constant defined for the minimum distance between dissimilar
embeddings. The overall Contrastive Loss is typically calculated as the mean of the individual losses
across all pairs in a batch or a dataset. During inference, the generated embeddings are compared with
similar ones in the same representation space using cosine similarity to generate the pictograms.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Results and Discussion</title>
      <p>To evaluate the efectiveness of our translation of text into pictograms, three metrics were utilized:
PictoER [17], BLEU [18], and METEOR [19]. Each of these metrics ofers a distinct perspective on the
quality of the translation, providing a comprehensive assessment. Our system achieved a PictoER score
of 141.909, a BLEU score of 3.419, and a METEOR score of 14.351. The breakdown of these individual
scores provides valuable insights into the strengths and areas for improvement in our translation
approach.
6. Conclusion
In this paper, we present a novel approach for mapping French sentences to corresponding AAC
pictograms using a BERT-based model, specifically utilizing CamemBERT embeddings combined with
contrastive learning techniques. Our methodology shows significant improvements in handling complex
semantic relationships and adapting to diverse linguistic contexts. The efective integration of these
techniques has demonstrated the potential to greatly enhance the communication capabilities of AAC
users, bridging the gap between text and pictographic representation.</p>
      <p>Future research on our Text-to-Pictogram conversion model could take several promising directions
to enhance its utility and performance. First, exploring model optimization techniques such as pruning
and quantization could improve the eficiency of our transformer-based model without impacting
performance, making it more suitable for resource-limited environments. Furthermore, expanding the
model to support multiple languages would increase its utility, particularly in diverse linguistic settings.
Further development of semantic analysis techniques could improve the model’s ability to handle
complex sentence structures and ambiguities, enhancing translation accuracy. These improvements
would not only extend the model’s applicability but also boost its efectiveness in real-world situations.
[2] Croix-Rouge, Communication alternative améliorée (caa) : la croix-rouge française dévoile sa
première étude d’impact social!, 2021.
[3] C. Macaire, E. Esperança-Rodier, B. Lecouteux, D. Schwab, Overview of 2024 imagecleftopicto tasks
– investigating the translation of natural language into pictograms, in: Experimental IR Meets
Multilinguality, Multimodality, and Interaction. CEUR Workshop Proceedings (CEUR-WS.org),
Grenoble, France, 2024.
[4] B. Ionescu, H. Müller, A.-M. Drăgulinescu, J. Rückert, A. Ben Abacha, A. García Seco de Herrera,
L. Bloch, R. Brüngel, A. Idrissi-Yaghir, H. Schäfer, C. S. Schmidt, T. M. Pakull, H. Damm, B. Bracke,
C. M. Friedrich, A.-G. Andrei, Y. Prokopchuk, D. Karpenka, A. Radzhabov, V. Kovalev, C. Macaire,
D. Schwab, B. Lecouteux, E. Esperança-Rodier, W.-w. Yim, Y. Fu, Z. Sun, M. Yetisgen, F. Xia, S. A.
Hicks, M. A. Riegler, V. Thambawita, A. Størås, P. Halvorsen, M. Heinrich, J. Kiesel, M. Potthast,
B. Stein, Overview of imageclef 2024: Multimedia retrieval in medical applications, in:
Experimental IR Meets Multilinguality, Multimodality, and Interaction, Proceedings of the 15th International
Conference of the CLEF Association (CLEF 2024), Springer Lecture Notes in Computer Science
LNCS, Grenoble, France, 2024.
[5] V. André, Canut, Mise à disposition de corpus oraux interactifs : le projet tcof (traitement de
corpus oraux en français), Pratiques. Linguistique, littérature, didactique (2010) 35–51.
[6] Cataix-Nègre, Communiquer autrement: Accompagner les personnes avec des troubles de la
parole ou du langage : les communications alternatives, De Boeck Supérieur, 2017.
[7] V. Vandeghinste, I. S. Sevens, F. Van Eynde, Translating text into pictographs, Natural Language
Engineering 23 (2017) 217–244. doi:10.1017/S135132491500039X, [Online]. Available: https:
//doi.org/10.1017/S135132491500039X.
[8] S. Bautista, R. Hervás, A. Hernández-Gil, C. Martínez-Díaz, S. Pascua, P. Gervás, Aratraductor:
text to pictogram translation using natural language processing techniques, in: Proceedings of
the XVIII International Conference on Human Computer Interaction (Interacción ’17), ACM, New
York, NY, USA, 2017. doi:10.1145/3123818.3123825, [Online]. Available: https://doi.org/10.
1145/3123818.3123825.
[9] J. A. Pereira, D. Macêdo, C. Zanchettin, A. L. I. de Oliveira, R. d. N. Fidalgo, Pictobert: Transformers
for next pictogram prediction, Expert Systems with Applications 202 (2022) 117231. doi:10.1016/
j.eswa.2022.117231, [Online]. Available: https://doi.org/10.1016/j.eswa.2022.117231.
[10] J. A. Pereira, C. Zanchettin, R. d. N. Fidalgo, Praact: Predictive augmentative and alternative
communication with transformers, Expert Systems with Applications 240 (2024) 122417. doi:10.
1016/j.eswa.2023.122417, [Online]. Available: https://doi.org/10.1016/j.eswa.2023.122417.
[11] J. Mutal, P. Bouillon, M. Norré, J. Gerlach, L. O. Grijalba, A neural machine translation approach
to translate text to pictographs in a medical speech translation system - the babeldr use case,
in: Proceedings of the 15th biennial conference of the Association for Machine Translation in
the Americas (AMTA 2022), volume 1, Orlando, USA, 2022, pp. 252–263. [Online]. Available:
Association for Machine Translation in the Americas.
[12] J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, Bert: Pre-training of deep bidirectional transformers
for language understanding, 2019. arXiv:1810.04805.
[13] L. Martin, B. Muller, P. J. Ortiz Suárez, Y. Dupont, L. Romary, Villemonte de la Clergerie, D. Seddah,
B. Sagot, Camembert: a tasty french language model, CoRR abs/1911.03894 (2019). URL: http:
//arxiv.org/abs/1911.03894. arXiv:1911.03894.
[14] C. Delestre, A. Amar, Distilcamembert: a distillation of the french model camembert, 2022.</p>
      <p>arXiv:2205.11111.
[15] Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, V. Stoyanov,</p>
      <p>Roberta: A robustly optimized bert pretraining approach, 2019. arXiv:1907.11692.
[16] T. Gao, X. Yao, D. Chen, Simcse: Simple contrastive learning of sentence embeddings, 2022.</p>
      <p>arXiv:2104.08821.
[17] J. P. Woodard, J. T. Nelson, An information theoretic measure of speech recognition performance,
in: Workshop on standardisation for speech I/O technology, Naval Air Development Center,
Warminster, PA, 1982.
[18] K. Papineni, S. Roukos, T. Ward, W. J. Zhu, Bleu: a method for automatic evaluation of machine
translation, in: Proceedings of the 40th annual meeting of the Association for Computational
Linguistics, 2002, pp. 311–318.
[19] S. Banerjee, A. Lavie, Meteor: An automatic metric for mt evaluation with improved correlation
with human judgments, in: Proceedings of the ACL workshop on intrinsic and extrinsic evaluation
measures for machine translation and/or summarization, 2005, pp. 65–72.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Romski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. A.</given-names>
            <surname>Sevcik</surname>
          </string-name>
          ,
          <article-title>Augmentative communication and early intervention: Myths and realities</article-title>
          ,
          <source>Infants &amp; Young Children</source>
          <volume>18</volume>
          (
          <year>2005</year>
          )
          <fpage>174</fpage>
          -
          <lpage>185</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>