<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Returning the L in NLP: Why Language (Variety) Matters and How to Embrace it in Our Models</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Barbara Plank</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Computer Science Department IT University of Copenhagen</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>NLP's success today is driven by advances in modeling together with huge amounts of unlabeled data to train language models. However, for many application scenarios like low-resource languages, non-standard data and dialects we do not have access to labeled resources and even unlabeled data might be scarce. Moreover, evaluation today largely focuses on standard splits, yet language varies along many dimensions [3]. What is more is that for almost every NLP task, the existence of a single perceived gold answer is at best an idealization. In this talk, I will emphasize the importance of language variation in inputs and outputs and its impact on NLP. I will outline ways on how to go about it. This includes recent work on how to transfer models to low-resource languages and language variants [5, 6], the use of incidental (or fortuitous) learning signals such as genre for dependency parsing [2] and learning beyond a single ground truth [1, 3, 4]. Biography. Barbara Plank is Professor in the Computer Science Department at ITU (IT University of Copenhagen). She is also the Head of the Master in Data Science Program. She received her PhD in Computational Linguistics from the University of Groningen. Her research interests focus on Natural Language Processing, in particular transfer learning and adaptations, learning from beyond the text, and in general learning under limited supervision and fortuitous data sources. She (co)-organised several workshops and international conferences, amongst which the PEOPLES workshop (since 2016) and the first European NLP Summit (EurNLP 2019). Barbara was general chair of the 22nd Northern Computational Linguistics conference (NoDaLiDa 2019) and workshop chair for ACL in 2019. Barbara is member of the advisory board of the European Association for Computational Linguistics (EACL) and vice-president of the Northern European Association for Language Technology (NEALT).</p>
      </abstract>
    </article-meta>
  </front>
  <body />
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Tommaso</given-names>
            <surname>Fornaciari</surname>
          </string-name>
          , Alexandra Uma, Silviu Paun, Barbara Plank, Dirk Hovy, and
          <string-name>
            <given-names>Massimo</given-names>
            <surname>Poesio</surname>
          </string-name>
          . Beyond Black &amp;
          <article-title>White: Leveraging Annotator Disagreement via Soft-Label MultiTask Learning</article-title>
          .
          <source>In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , pages
          <fpage>2591</fpage>
          -
          <lpage>2597</lpage>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Max</given-names>
            <surname>Mu</surname>
          </string-name>
          <article-title>¨ller-</article-title>
          <string-name>
            <surname>Eberstein</surname>
            , Rob van der Goot, and
            <given-names>Barbara</given-names>
          </string-name>
          <string-name>
            <surname>Plank</surname>
          </string-name>
          .
          <article-title>Genre as Weak Supervision for Cross-lingual Dependency Parsing</article-title>
          .
          <source>In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing</source>
          , pages
          <fpage>4786</fpage>
          -
          <lpage>4802</lpage>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Barbara</given-names>
            <surname>Plank</surname>
          </string-name>
          .
          <article-title>What to do about non-standard (or non-canonical) language in NLP</article-title>
          .
          <source>In Proceedings of KONVENS</source>
          <year>2016</year>
          , Ruhr-University Bochum. Bochumer Linguistische Arbeitsberichte,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Barbara</given-names>
            <surname>Plank</surname>
          </string-name>
          , Dirk Hovy, and
          <string-name>
            <given-names>Anders</given-names>
            <surname>Søgaard</surname>
          </string-name>
          .
          <article-title>Learning part-of-speech taggers with interannotator agreement loss</article-title>
          .
          <source>In Proceedings of the 14th Conference of the European Chapter of the Association for Computational Linguistics</source>
          , pages
          <fpage>742</fpage>
          -
          <lpage>751</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Barbara</given-names>
            <surname>Plank</surname>
          </string-name>
          , Kristian Nørgaard Jensen, and Rob van der Goot. DaN+:
          <article-title>Danish nested named entities and lexical normalization</article-title>
          .
          <source>In Proceedings of the 28th International Conference on Computational Linguistics</source>
          , pages
          <fpage>6649</fpage>
          -
          <lpage>6662</lpage>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Rob</surname>
            <given-names>van der Goot</given-names>
          </string-name>
          , Ibrahim Sharaf, Aizhan Imankulova, Ahmet U¨ stu¨n, Marija Stepanovic´,
          <string-name>
            <surname>Alan</surname>
            <given-names>Ramponi</given-names>
          </string-name>
          , Siti Oryza Khairunnisa, Mamoru Komachi, and
          <string-name>
            <given-names>Barbara</given-names>
            <surname>Plank</surname>
          </string-name>
          .
          <article-title>From Masked Language Modeling to Translation: Non-English Auxiliary Tasks Improve Zeroshot Spoken Language Understanding</article-title>
          .
          <source>In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , pages
          <fpage>2479</fpage>
          -
          <lpage>2497</lpage>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>