<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Leveraging LDA Topic Modeling and BERT Embeddings for Thematic Unsupervised Classification of Tourism News in Rest-Mex Competition</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Erika Rivadeneira-Pérez</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Cipriano Callejas-Hernández</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Mathematics Research Center (CIMAT)</institution>
          ,
          <addr-line>Guanajuato</addr-line>
          ,
          <country country="MX">Mexico</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <abstract>
        <p>In this work we present a solution to the Thematic Unsupervised Classification , a track presented in this year Rest-Mex competition, in which four topic-based groups are to be found from a unlabeled text item news related to tourism in Mexico. Our approach includes LDA topic modeling on a term-document representation, as well as BERT embeddings. Our BERT-based approach achieved a second place among all participating teams in the competition, demonstrating the efectiveness of mixing pre-trained models with traditional machine learning techniques.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Sentiment Analysis</kwd>
        <kwd>Deep Learning</kwd>
        <kwd>Transformer Models</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <sec id="sec-1-1">
        <title>1.1. Unsupervised Text Classification</title>
        <p>
          A very common method of unsupervised learning is clustering, which aims to identify distinct
groups in data, that is, we seek to learn something about the structure and patterns inherited by
the data. On other hand, unsupervised text classification aims to perform categorization without
using annotated data during training and therefore ofer the potential to reduce annotation
costs. In this direction little research has been conducted on unsupervised text classification, see
for instance the survey [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] . Moreover, Braun et al [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] show that similarity-based approaches
are the most popular technique, however for our approaches we followed a segmentation-based
one, where either we used clustering methods (or topic modeling) considering that the number
of clusters (or topics) is known in advance.
        </p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Thematic Unsupervised Classification Track</title>
      <p>For this task, 50,000 news items were collected on 4 diferent topics related to tourism. The
challenge is to identify the groups in an automatic way. We call our approach a
segmentationbased one, as detailed below.</p>
      <sec id="sec-2-1">
        <title>2.1. Corpus Description</title>
        <p>All data was obtained from google news over the last two years regarding 4 unrevealed touristic
topics, downloaded and tagged. Figure 1 shows a news example of the corpus.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Data Preprocessing</title>
        <p>As illustrated in Figure 2 we decided to homogenize the text in the following way:
• Removing hyperlinks.
• Removing numerical and special characters due to the large amount of numeric data that
does not provide relevance to the topics, such as dates or monetary amounts, for example.
• Converting the entire text to lowercase.</p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Approach 1. Topic Modelling</title>
        <p>
          The first proposal is to use of Latent Dirichlet Allocation (LDA), a widely used statistical model
in NLP for unsupervised text clustering. LDA aims to discover the underlying topics in a set
of text documents and group them into coherent categories [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. This model is based on the
premise that each document is a combination of various topics, and each topic is characterized
by its distribution of words, this is known as a word frequency method. Our pipeline in this
approach is the following:
• Text preprocessing: Text is preprocessed according to the specifications in section 2.2.
• Text Vector Representation: Text documents are map into a vector where each element
represents the count of a specific word in the document.
• Topic inference: Topic inference involves calculating the probability of each document
belonging to each topic based on the words it contains. In other words, topics are assigned
to documents based on the probability of belonging to each of them. Table 1 shows the
Top-10 most frequent words per Topic among all news.
• Interpretation of results: Finally, we analyze the assignment of topics to documents and
understanding the discovered patterns.
        </p>
        <sec id="sec-2-3-1">
          <title>Topic</title>
          <p>Topic 1
Topic 2
Topic 3
Topic 4</p>
        </sec>
        <sec id="sec-2-3-2">
          <title>Palabras más frecuentes</title>
          <p>autoridades, fiscalia, noticias, leon, personas, elementos, seguridad, policia, dos, mas
parte, asi, ciudad, ser, si, notificaciones, tambien, hace, mexico, mas
lectura, anuncio, min, articulo, seguridad, foto, tiempo, mas, mexico, nacional
personas, millones, tambien, gobierno, mexico, guanajuato, yucatan, pesos, mil, mas</p>
          <p>Note that in Table 1 frequent words overlap among diferent topics identified by LDA, which
can make it challenging to diferentiate them clearly. This overlap of words can be attributed to
various factors, such as thematic similarity between topics or the presence of shared vocabulary.
Hence, accurately discerning the boundaries and distinctive characteristics of each topic can
be dificult. In our particular case, we have observed that the LDA model faces dificulties in
accurately and eficiently diferentiating more specific topics due to the intersections of words
among the identified topics, see Figure 3 where the overlap between clusters is shown.</p>
        </sec>
      </sec>
      <sec id="sec-2-4">
        <title>2.4. Approach 2. Machine Learning Techniques</title>
        <p>For the second approach, we considered classical clustering methods: k-means and hierarchical
clustering, using the same word frequency representation of Section 2.3. The resulting
topic groups of k-means and hierarchical clustering can be seen in Figure 4a and Figure 4b
respectively. The figures show the two-dimensional SVD representation of the text and the
corresponding topics to which each news belongs. We can observe that the observations of the
obtained topics, for the most part, do not overlap with each other. However this traditional
approach shows a more, at least visually speaking, clear segmentation of our text data.</p>
      </sec>
      <sec id="sec-2-5">
        <title>2.5. Approach 3 ML techniques on Bert’s corpus representation</title>
        <p>Contextualized word and sentence embeddings produced by pre-trained models such as BERT
have demonstrated the state-of-the-art results in NLP tasks, recently Selva et al. [15] showed
that using BERT versus BOW as document representation surpass in topic coherence, moreover,
the use of term frequency to select topic words fails to capture the semantics of clusters precisely
because of words with high frequency may be common across diferent clusters, as seen in
(a) Visualization of corpus representation and
clustering results obtained with K-means
algorithm.
(b) Visualization of corpus representation and
clustering results obtained with
Hierarchical Clustering</p>
      </sec>
      <sec id="sec-2-6">
        <title>3.1. Evaluation Metrics</title>
        <p>
          To evaluate each system in the unsupervised classification task, an alignment must first be done.
Given the Gold Standard, the output of each k system must be renumbered so that the themes
correspond. This is because the only restriction that the participating teams have is that they
must identify 4 groups with the news shared in the competition [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. This means that the labels
do not necessarily coincide for the same groups expected in the Gold Standard. For this reason,
a re-labeling will be done for each system using the Gold Standard label that shares the most
instances with each of the groups resulting from the k system. Once the alignment is done, it
will be evaluated with a macro F-measure as shown in Equation 1.
1 ∑|︁|
|| =1
Thematic() =
()
(1)
1st Javilonso-Team 0.2827
        </p>
        <sec id="sec-2-6-1">
          <title>2nd CIMAT-Team_run3 0.2400</title>
          <p>3rd JCMQ-Team_run_5 0.2182
HM MCE-Team_2ndIterKmeans 0.2031</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>4. Conclusions</title>
      <p>As showed in Figures 4 and 4a, a term-frequency representation plus traditional clustering
methods achieves a good segmentation, at least visually speaking. However as noted in Table
2, pre-trained contextualized embeddings surpass the simple term-frequency representation
nonetheless in Figure 5 this segmentation is not clear enough, at least visually. We hypothesize
that this might be related to the overlap between news items in terms of tourism-related
keywords, such as COVID, since this created a big change in the market. We believe that more
features could have been included in this approach, such is the case of better handling the
scope of bert-embeddings, straightforward we passed news items through the pre-trained
model all-MiniLM-L6-v2 sentence transformer [16] cutting of up to 128 items, resulting in a
384 dimensional dense vector space. Because of the little known results on this topic and being
this the first time of this track in the Rest-Mex 2023 competition we wanted to try something
simpler.</p>
    </sec>
    <sec id="sec-4">
      <title>5. Acknowledgments</title>
      <p>The authors thank Dr. Johan Johan VanHorebeek from Mathematics Research Center (CIMAT)
for his support and guidance in creating this project.
Delgado, R. Guerrero-Rodríguez, L. Bustio-Martínez, Overview of rest-mex at iberlef 2022:
Recommendation system, sentiment analysis and covid semaphore prediction for mexican
tourist texts, Procesamiento del Lenguaje Natural 69 (2022).
[10] E. Olmos-Martínez, M. Á. Álvarez-Carmona, R. Aranda, A. Díaz-Pacheco, What does the
media tell us about a destination? the cancun case, seen from the usa, canada, and mexico,
International Journal of Tourism Cities (2023)
[11] M. A. Alvarez-Carmona, R. Aranda, A. Rodriguez-Gonzalez, D. Fajardo-Delgado, M. G.</p>
      <p>A. Sanchez, H. Perez-Espinosa, J. Martinez-Miranda, R. Guerrero-Rodriguez, L.
BustioMartinez, A. D. Pacheco, Natural language processing applied to tourism research: A
systematic review and future research directions, Journal of King Saud University-Computer
and Information Sciences (2022).
[12] Author, A.-B.: Contribution title. In: 9th International Proceedings on Proceedings, pp.</p>
      <p>1–2. Publisher, Location (2010)
[13] Haj-Yahia, Z., Sieg, A., &amp; Deleris, L. A. (2019, July). Towards unsupervised text classification
leveraging experts and word embeddings. In Proceedings of the 57th annual meeting of the
Association for Computational Linguistics (pp. 371-379).
[14] LNCS Homepage, http://www.springer.com/lncs. Last accessed 4 Oct 2017
[15] Selva Birunda, S., &amp; Kanniga Devi, R. (2021). A review on word embedding techniques
for text classification. Innovative Data Communication Technologies and Application:
Proceedings of ICIDCA 2020, 267-281.
[16] Reimers, N., &amp; Gurevych, I. (2019). Sentence-bert: Sentence embeddings using siamese
bert-networks. arXiv preprint arXiv:1908.10084.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Á</surname>
          </string-name>
          .
          <article-title>Álvarez-Carmona, Á</article-title>
          . Díaz-Pacheco,
          <string-name>
            <given-names>R.</given-names>
            <surname>Aranda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. Y.</given-names>
            <surname>Rodríguez-González</surname>
          </string-name>
          , L. BustioMartínez, V.
          <string-name>
            <surname>Muñis-Sánchez</surname>
            ,
            <given-names>A. P.</given-names>
          </string-name>
          <string-name>
            <surname>Pastor-López</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Sánchez-Vega</surname>
          </string-name>
          ,
          <article-title>Overview of rest-mex at iberlef 2023: Research on sentiment analysis task for mexican tourist texts</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>71</volume>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Schopf</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Braun</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Matthes</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          (
          <year>2022</year>
          ).
          <article-title>Evaluating unsupervised text classification: zero-shot and similarity-based approaches</article-title>
          .
          <source>arXiv preprint arXiv:2211</source>
          .
          <fpage>16285</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Blei</surname>
            ,
            <given-names>D. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ng</surname>
            ,
            <given-names>A. Y.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Jordan</surname>
            ,
            <given-names>M. I.</given-names>
          </string-name>
          (
          <year>2003</year>
          ).
          <article-title>Latent dirichlet allocation</article-title>
          .
          <source>Journal of machine Learning research, 3(Jan)</source>
          ,
          <fpage>993</fpage>
          -
          <lpage>1022</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Author</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Author</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Author</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Book title</article-title>
          .
          <source>2nd edn. Publisher</source>
          ,
          <string-name>
            <surname>Location</surname>
          </string-name>
          (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Aschauer</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Egger</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          (
          <year>2023</year>
          ).
          <article-title>Transformations in tourism following COVID-19? A longitudinal study on the perceptions of tourists</article-title>
          .
          <source>Journal of Tourism Futures.</source>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Hudson</surname>
            , Simon, and
            <given-names>Brent</given-names>
          </string-name>
          <string-name>
            <surname>Ritchie</surname>
          </string-name>
          .
          <article-title>"Understanding the domestic market using cluster analysis: A case study of the marketing eforts of Travel Alberta."</article-title>
          <source>Journal of Vacation Marketing 8.3</source>
          (
          <year>2002</year>
          ):
          <fpage>263</fpage>
          -
          <lpage>276</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Egger</surname>
          </string-name>
          , Roman.
          <source>Applied Data Science in Tourism: Interdisciplinary Approaches</source>
          , Methodologies, and
          <string-name>
            <surname>Applications</surname>
          </string-name>
          .
          <source>Zeitschrift für Tourismuswissenschaft</source>
          ,
          <year>2021</year>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Diaz-Pacheco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Á</surname>
          </string-name>
          .
          <string-name>
            <surname>Álvarez-Carmona</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Guerrero-Rodríguez</surname>
            ,
            <given-names>L. A. C.</given-names>
          </string-name>
          <string-name>
            <surname>Chávez</surname>
            ,
            <given-names>A. Y.</given-names>
          </string-name>
          <string-name>
            <surname>Rodríguez-González</surname>
            ,
            <given-names>J. P.</given-names>
          </string-name>
          <string-name>
            <surname>Ramírez-Silva</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Aranda</surname>
          </string-name>
          ,
          <article-title>Artificial intelligence methods to support the research of destination image in tourism. a systematic review</article-title>
          ,
          <source>Journal of Experimental &amp; Theoretical Artificial Intelligence</source>
          (
          <year>2022</year>
          )
          <fpage>1</fpage>
          -
          <lpage>31</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Á</surname>
          </string-name>
          .
          <article-title>Álvarez-Carmona, Á</article-title>
          . Díaz-Pacheco,
          <string-name>
            <given-names>R.</given-names>
            <surname>Aranda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. Y.</given-names>
            <surname>Rodríguez-González</surname>
          </string-name>
          , D. Fajardo-
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>