<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>BERT4EVER at ADoBo 2021: Detection of Borrowings in the Spanish Language Using Pseudo-label Technology</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Shengyi Jiang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tong Cui</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yingwen Fu</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nankai Lin</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jieyi Xiang</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Guangzhou Key Laboratory of Multilingual Intelligent Processing, Guangdong University of Foreign Studies</institution>
          ,
          <addr-line>Guangzhou</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>School of Information Science and Technology, Guangdong University of Foreign Studies</institution>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper, we report the solution of the team BERT 4EVER for the automatic detection of borrowings in the Spanish Language task in IberLeF 2021, which aims to detect lexical borrowings that appear in the Spanish press. We adopt the CRF model to tackle the problem. In addition, we introduce pseudolabel technology and ensemble learning to improve the generalization capability. Experimental results demonstrate the effectiveness of CRF model and pseudolabel technology.</p>
      </abstract>
      <kwd-group>
        <kwd>Automatic Detection of Borrowings</kwd>
        <kwd>CRF</kwd>
        <kwd>Pseudo-label Technology</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Lexical borrowing is a word formation that is widely used in many languages. Previous
work on computational detection of lexical borrowings has relied mostly on dictionary
and corpora lookup [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ][
        <xref ref-type="bibr" rid="ref2">2</xref>
        ][
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], with the limitation coming from the original dictionary
or corpora. On the other hand, computational approaches to mixed-language data have
usually framed the task of identifying the language of a word as a sequence labeling
task, where every word in the sequence is attached to a language tag [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ][
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        IberLeF 2021 proposes the task “Automatic Detection of Borrowings in the Spanish
Language” [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Our team, BERT 4EVER, also participates in this task. In this report,
we will review our solution to this task, namely, the CRF model aided by pseudo-label
technology and ensemble learning.
      </p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        Linguistic borrowing is the process of copying elements and patterns from another
language into one [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. This classification system is based on two processes: import and
substitution. Input is the incorporation into the recipient's language of a foreign form
that may or may not contain a meaning. Substitution refers to the substitution of foreign
phonemes or morphemes by foreign phonemes or morphemes of the recipient language
so as to localize the foreign form. Both processes can occur in the same borrowings.
Thus, linguistic borrowing involves communication between two languages and has
been extensively studied in the field of contact linguistics [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Various typologies have
been proposed to classify language loanwords according to different criteria, such as
typological features, linguistic hierarchy involved, integration of loanwords elements
in the recipient's language, etc [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ][
        <xref ref-type="bibr" rid="ref10">10</xref>
        ][
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
      <p>Now that English has established itself as the global lingua franca, many languages
are currently undergoing the process of importing new loanwords from English. In the
past decade, English has produced a large number of lexical loanwords in many
European languages, especially in the press.</p>
      <p>
        Previous work on computational detection of lexical borrowings have relied mostly
on dictionary and corpora lookup. Studies on anglicization have begun to use a
multimillion-word corpus [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ][
        <xref ref-type="bibr" rid="ref13">13</xref>
        ][
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Alex [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] combined lexicon lookup and a search
engine module that used the web as a corpus to detect English inclusions in a German
text corpus and compared the proposed model with a maxent Markov model. Furiassi
and Hofland [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] explored corpora lookup and character n-grams to extract false
anglicisms from an Italian newspaper corpus. Andersen [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] used dictionary lookup, regular
expressions and lexicon-derived frequencies of character n-grams to detect anglicism
candidates in the Norwegian Newspaper Corpus (NNC). In computational approaches
to mix-language data, the task, aiming for the identification of the language of a word,
has usually been assumed as a tagging problem which needs every word in the sequence
to be tagged [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        The large amount of available data presents methodological challenges to data
processing for English language research. Corpus-based studies of English borrowings in
Spanish media have traditionally relied on manual evaluation of either previously
compiled general corpora such as CREA [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], or new tailor-made corpora designed to
analyze specific genres, varieties or phenomena. In Spanish, Serigos [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] extracted
anglicisms from an Argentinian newspaper corpus by combining dictionary lookup (aided
by TreeTagger and the NLTK lemmatizer) with automatic filtering of capitalized words
and manual inspection. In Serigos [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], a character n-gram module was added in the
dictionary lookup method to estimate the probabilities of a word being English or
Spanish. Moreno Fernandez and Moreno Sandoval[
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] used different pattern-matching
filters and lexicon lookup to extract anglicism candidates from a tweet corpus in US
Spanish.
      </p>
    </sec>
    <sec id="sec-3">
      <title>Method</title>
      <p>In the automatic detection of borrowings in the Spanish Language task, we train five
CRF models based on the five-fold data and then use the trained CRF models to predict
unlabeled samples. We gather the pseudo-labeled dataset together with the training set
to train the new CRF model.
3.1</p>
      <p>CRF</p>
      <p>There are two random variables,  is a random variable on the sequence of data to
be labeled, and  is a random variable on the corresponding sequence of data to be
labeled. The random variables  and  are under the common distribution, but we
construct a conditional model  ( | ) from paired observation and label sequences in a
discriminative framework, without explicitly modeling the marginal  ( ).</p>
      <p>Let</p>
      <p>= ( ,  ) be a graph such that  = (  ) ∈ ，so that  is indexed by the
vertices of  .Then ( ,  ) is a conditional random field in case, when conditioned on  , the
random variables   obey the Markov property with respect to the graph:
 (  | ,   ,</p>
      <p>≠  ) =  (  | ,   ,  ~ )
where  ~ means that 
and  are neighbors in  , 
≠  means all vertices
except  .   and   are random variables corresponding to  and  .</p>
      <p>The collocation of the current word and the next (or last) word</p>
      <p>We conduct five-fold cross-validation for the training data and then train five models
based on the five-fold data. Each model predicts the test data separately. For each token
 , the predicted output of the model is</p>
      <p>=   ( )
in which  is the token  ’s feature representation,   is the  -th CRF model and
  is the output of  -th CRF model. Therefore, the output of the five models is
 = [ 1,  2,  3,  4,  5]</p>
      <p>We consider the label that appears most in  as the label of  .
3.3 Pseudo-label technology</p>
      <p>
        We use a pseudo-label strategy [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ][
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] to generate labeled data that does not require
manual labeling, as shown in Figure 2. We first use the competition open training set
to train CRF models, and then use the trained CRF models to predict unlabeled samples,
the predicted results as the sample label. And then we screen all the predicted samples
to filter out the sentences without lexical borrowings, only the sentences with lexical
borrowings exist. The unlabeled samples we use from GlobalVoices (Spanish portion
of GlobalVoices)2 and News-Commentary11 (Spanish portion of NCv11)3. We gather
the filtered sentence set together with the training set to train the new CRF model, and
although the sample quality obtained through data enhancement is not high, the new
model has higher generalization capability to some extent because the new model
trained with more data.
2 http://opus.nlpl.eu/GlobalVoices.php
3 http://opus.nlpl.eu/News-Commentary-v11.php
We first explore the performance of different prefixes/suffixes, and the results are
shown in Table 2. The feature of first (and last) 2 letters in the current word has the
greatest impact on the task, with an increase of 37.39 in the F value. When all the
prefix/suffix features are used together, the effect is the best, and the result based on
fivefold cross-validation has reached 53.42%.
      </p>
      <p>As shown in Table 3 and Table 4, the recall of CRF based on pseudo-label
technology is significantly improved, which proves that the pseudo-label technology can
improve the generalization performance of the model. On the final test set, the F value of
the CRF model reached 37.99%, and the F value of the CRF model based on
pseudolabel technology is 40.25%, which shows that the pseudo-label technology has a
significant impact on detection of borrowings in the Spanish language task.
5</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusion</title>
      <p>In the automatic detection of borrowings in the Spanish Language task in IberLeF 2021,
we adopt the CRF model aided by pseudo-label technology and ensemble learning. In
addition, we also explore the impact of different features on the task. In the future, we
will try to combine pseudo-label technology with deep learning models in order to
achieve better results on the detection of borrowings tasks.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgements</title>
      <p>This work was supported by the Key Field Project for Universities of Guangdong
Province (No. 2019KZDZX1016), the National Natural Science Foundation of China (No.
61572145) and the National Social Science Foundation of China (No. 17CTQ045). The
authors would like to thank the anonymous reviewers for their valuable comments and
suggestions.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Alex</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Automatic detection of English inclusions in mixed-lingual data with an application to parsing</article-title>
          . University of Edinburgh (
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Andersen</surname>
          </string-name>
          , G.:
          <article-title>Semi-automatic approaches to Anglicism detection in Norwegian corpus data</article-title>
          .
          <source>The anglicization of European lexis</source>
          ,
          <fpage>111</fpage>
          -
          <lpage>130</lpage>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Serigos</surname>
            ,
            <given-names>J. R. L.</given-names>
          </string-name>
          :
          <article-title>Applying corpus and computational methods to loanword research: new approaches to Anglicisms in Spanish</article-title>
          . University of Texas at Austin (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Molina</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>AlGhamdi</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ghoneim</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , et al.:
          <article-title>Overview for the second shared task on language identification in code-switched data</article-title>
          .
          <source>In: Proceedings of the Second Workshop on Computational Approaches</source>
          to Code Switching,
          <fpage>40</fpage>
          -
          <lpage>49</lpage>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Solorio</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blair</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maharjan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , et al.:
          <article-title>Overview for the first shared task on language identification in code-switched data</article-title>
          .
          <source>In: Proceedings of the First Workshop on Computational Approaches</source>
          to Code Switching, pp.
          <fpage>62</fpage>
          -
          <lpage>72</lpage>
          . (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Alvarez</given-names>
            <surname>Mellado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Espinosa</surname>
          </string-name>
          <string-name>
            <surname>Anke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Gonzalo</surname>
          </string-name>
          <string-name>
            <surname>Arroyo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Lignos</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Porta</given-names>
            <surname>Zamorano</surname>
          </string-name>
          , J.:
          <article-title>Overview of ADoBo 2021 shared task: Automatic Detection of Unassimilated Borrowings in the Spanish Press</article-title>
          .
          <source>Procesamiento del Lenguaje Natural</source>
          ,
          <volume>67</volume>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Haugen</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          :
          <article-title>The analysis of linguistic borrowing</article-title>
          .
          <source>Language</source>
          <volume>26</volume>
          (
          <issue>2</issue>
          ),
          <fpage>210</fpage>
          -
          <lpage>231</lpage>
          (
          <year>1950</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Weinreich</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          :
          <article-title>Languages in contact</article-title>
          .
          <source>Findings and Problems</source>
          (
          <year>1953</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Haspelmath</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Tadmor</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          :
          <article-title>Loanwords in the world's languages: a comparative handbook</article-title>
          . Walter de Gruyter (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Matras</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Sakel</surname>
          </string-name>
          , J.:
          <article-title>Grammatical borrowing in cross-linguistic perspective</article-title>
          .
          <source>Walter de Gruyter</source>
          <volume>38</volume>
          (
          <year>2007</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Thomason</surname>
            ,
            <given-names>S. G.</given-names>
          </string-name>
          and Kaufman, T.:
          <article-title>Language contact, creolization, and genetic linguistics</article-title>
          . Univ of California Press (
          <year>1992</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Andersen</surname>
          </string-name>
          , G.:
          <article-title>Pragmatic borrowing</article-title>
          .
          <source>Journal of Pragmatics</source>
          <volume>67</volume>
          ,
          <fpage>17</fpage>
          -
          <lpage>33</lpage>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Balteiro</surname>
          </string-name>
          , I.:
          <article-title>A reassessment of traditional lexicographical tools in the light of new corpora: sports Anglicisms in Spanish</article-title>
          .
          <source>International Journal of English Studies</source>
          <volume>11</volume>
          (
          <issue>2</issue>
          ),
          <fpage>23</fpage>
          -
          <lpage>52</lpage>
          (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Zenner</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Speelman</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Geeraerts</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Cognitive Sociolinguistics meets loanword research: Measuring variation in the success of anglicisms in Dutch</article-title>
          .
          <source>Cognitive Linguistics</source>
          <volume>23</volume>
          (
          <issue>4</issue>
          ),
          <fpage>749</fpage>
          -
          <lpage>792</lpage>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Núñez</surname>
            ,
            <given-names>N. E. E.</given-names>
          </string-name>
          :
          <article-title>Anglicisms in CREA: a quantitative analysis in Spanish newspapers</article-title>
          .
          <source>In: Language design: journal of theoretical and experimental linguistics 18</source>
          ,
          <fpage>215</fpage>
          -
          <lpage>242</lpage>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Alex</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Comparing Corpus-based to Web-based Lookup Techniques for Automatic English Inclusion Detection</article-title>
          .
          <source>In: Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC'08)</source>
          ,
          <source>European Language Resources Association (ELRA)</source>
          (
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Furiassi</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Hofland</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>The retrieval of false anglicisms in newspaper texts</article-title>
          .
          <source>Corpus Linguistics 25 Years On</source>
          ,
          <fpage>347</fpage>
          -
          <lpage>363</lpage>
          (
          <year>2007</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Serigos</surname>
          </string-name>
          , J.:
          <article-title>Using distributional semantics in loan word research: A concept-based approach to quantifying semantic specificity of anglicisms in Spanish</article-title>
          .
          <source>International Journal of Bilingualism</source>
          <volume>21</volume>
          (
          <issue>5</issue>
          ),
          <fpage>521</fpage>
          -
          <lpage>540</lpage>
          (
          <year>2007</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Moreno</surname>
            <given-names>F. F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moreno</surname>
            <given-names>S. A.</given-names>
          </string-name>
          :
          <article-title>Configuración lingüística de anglicismos procedentes de Twitter en el español estadounidense</article-title>
          .
          <source>Revista signos</source>
          <volume>51</volume>
          (
          <issue>98</issue>
          ),
          <fpage>382</fpage>
          -
          <lpage>409</lpage>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <string-name>
            <surname>Pseudo-Label</surname>
          </string-name>
          :
          <article-title>The Simple and Efficient Semi-Supervised Learning Method for Deep Neural Networks</article-title>
          .
          <source>In: ICML 2013 Workshop: Challenges in Representation Learning</source>
          . pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          . (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Shi</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gong</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ding</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          , et al.:
          <article-title>Transductive semi-supervised deep learning sing minmax features</article-title>
          .
          <source>In Proceedings of the European Conference on Computer Vision</source>
          . pp.
          <fpage>299</fpage>
          -
          <lpage>315</lpage>
          . (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>