<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Exploring Combining Training Datasets for the CLIN 2019 Shared Task on Cross-genre Gender Detection in Dutch</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Bibliography</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Swedish / Sprakbanken University of Gothenburg</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>We present our entries to the Shared Task on Cross-genre Gender Detection in Dutch at CLIN 2019. We start from a simple logistic regression model with commonly used features, and consider two ways of combining training data from di erent sources. Our in-genre models do reasonably well, but the cross-genre models are a lot worse. Post-task experiments show no clear systematic advantage of one way of combining training data sources over the other, but do suggest accuracy can be gained from a better way of setting model hyperparameters.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Detection of binary author gender can be done from text alone with impressive
e ectiveness. For instance, Van der Goot et al. (2018) report an accuracy of 80%
on Dutch tweets for their system, and the top systems for four other languages
reported in Rangel et al. (2017) also perform in the low 80% accuracy range. These
results concern systems that were trained on and applied to Twitter data, with
multiple documents (i.e., tweets) per author. It is to be expected that performance
su ers when such systems are applied to another genre than they were initially
trained on. The CLIN 2019 Shared Task on Cross-genre Gender Detection in
Dutch therefore invites authors to investigate \gender prediction within and
across di erent genres in Dutch."1 This paper reports on our participation this
shared task.</p>
      <p>The shared task consists of an in-genre setting and a cross-genre setting.
In the former, models are trained on and applied to data from the same
genre/data source. In the latter, models are trained on data from one or more
sources, and applied to data from a genre that is assumed to be completely
unknown during training. The genres supplied in the shared task were News,
Twitter and YouTube comments.</p>
      <p>Having access to training data from multiple sources raises the question of
whether we can use that fact to construct models that generalize better and
therefore perform better in a cross-genre setting. Our contribution to the shared
task is a small investigation of the e ect of how the multiple sources are combined.
Building upon a basic logistic regression model with features taken from existing
research on author pro ling in general and gender identi cation in particular, we
compare two ways of combining training data sources.</p>
      <p>Section 2 gives a formal description of the used model and introduces an
alternative objective that combines training datasets in a principled way.
Section 3 describes the used features and gives some implementation details. The
results for our systems in the shared task are presented and brie y discussed in
Section 4. These results prompt a set of post-task experiments, whose outcome
and implications are re ected upon in Section 5
2</p>
    </sec>
    <sec id="sec-2">
      <title>Description of the models</title>
      <p>
        The core of our approach is a logistic regression model. To t a logistic regression
model with L2 regularization, we need to nd an intercept b0 and feature weights
B that minimize the sum of a) the normalized negative log-likelihood of the
model given the data, and b) the squared magnitude of those weights. To be
precise, we minimize
1
jDj
log L(b0; BjD) +
2
jBj
X bi2;
i=1
(1)
where D is the training data, and a hyperparameter that lets us set the strength
of the regularization factor.
        <xref ref-type="bibr" rid="ref7 ref8">(See, e.g., Hastie et al., 2009; Malouf, 2010 for proper
introductions.)</xref>
        For the in-genre setting, we simply apply this formulation of the
objective.
      </p>
      <p>
        In the cross-genre setting, we are able to train on a combination of two or
more datasets from di erent genres. Ideally, we would be able to leverage this fact
to nd models that generalize better and therefore fare better when applied to a
new genre. Handling data from di erent sources is well-studied in the domain
adaptation literature. However, there one often has access to training data from
both source and target domains
        <xref ref-type="bibr" rid="ref3 ref4 ref7">(like in Daume's frustratingly easy method,
Daume III, 2007, or equivalently, multilevel regression, Finkel and Manning,
2009)</xref>
        , or to source training data and distributional information about the target
domain (e.g., work on targeting di erent kinds of distributional shifts). In this
shared task, however, we have two or more sources, but are supposed to assume
no knowledge of the target genre. Therefore, such domain adaptation methods
do not apply directly.
      </p>
      <p>d
o
o
h
li
e
k
i
-l
g
o
l
e
v
it
a
g
e
n
d
e
z
il
a
m
r
o
N</p>
      <sec id="sec-2-1">
        <title>Source D₁</title>
        <p>Source D₂ (10× larger)
Pooled data model
Average objective
LSE objective (k=1,4,16,64)
(minima)
0.0
0.2
0.4
0.6</p>
      </sec>
      <sec id="sec-2-2">
        <title>Parameter setting 0.8 1.0</title>
        <p>A direct way to combine training data from multiple sources is simply to pool
the data. We can then proceed to train according to Equation 1, as we would
with a single dataset. When we are dealing with source datasets of unequal size,
we can also consider normalizing the respective negative log-likelihoods to the
sizes of the datasets, so that the larger dataset does not dominate the model. Just
summing/averaging normalized negative log-likelihoods for each source could
still lead to one source dominating, if that source is much easier to model. We
therefore combine the negative log-likelihoods by taking their maximum. The
log-sum-exp function gives us the kind of smooth maximum we need in order
to be able to use standard optimization algorithms. In the formulation we use,
log-sum-exp also takes a scaling parameter:
lse(x; y; k) =
log(ekx + eky):
With positive k, lse(x; y; k) is always greater than max(x; y). A higher k makes
lse less smooth but closer to max.</p>
        <p>The objective for two datasets would thus be
lse</p>
        <p>1
jD1j
1
k
1
jD2j
2
jBj
X b2;</p>
        <p>i
i=1
(2)
(3)
log L(b0; BjD1);
log L(b0; BjD2); k
+
where and k are hyperparameters. This is trivially extended to more datasets.</p>
        <p>Figure 1 illustrates the three ways of combining two source datasets graphically.
The hope is that by combining sources in a balanced way, we nd models that
generalize better to new genres. This is on the premise that we do not know
anything about the target genre. If we had information that the new genre is
more like one of the source genres, we might be better o building an `unfair'
model.
3</p>
        <p>
          Features, data preparation, and implementation
A linear model using surface form-based n-gram features has been shown to be
very e ective in (in-domain) gender identi cation
          <xref ref-type="bibr" rid="ref2">(Basile et al., 2018)</xref>
          , and we
will follow this method here, too, albeit in simpli ed form. The results presented
in the cited paper suggest the lion's share of accuracy is contributed by simple
unigram features, and Bamman et al. (2014) present an investigation of which
(classes of) lexical unigrams di erentiate male from female authors on Twitter.
We therefore only use unigram token occurrences in our model. Van der Goot
et al. (2018) show the e ectiveness of `bleached' lexical features when doing
crosslingual gender detection. Inspired by their approach, we include word lengths as
features. Finally, character n-grams are a common ingredient in author pro ling.
Zechner (2017), on authorship attribution, shows that even character unigram
frequencies carry identifying information. These therefore constitute our nal
feature subset. Keeping the feature set small and simple allows us to focus on
the e ects of model combination.
        </p>
        <sec id="sec-2-2-1">
          <title>X-val accuracy Eval accuracy .639 .6316</title>
          <p>
            We follow common praxis in authorship attribution
            <xref ref-type="bibr" rid="ref10">(see e.g., Smith and
Aldridge, 2011)</xref>
            , in restricting the set of features to just the most frequently
occurring types. The cut-o points where chosen on the basis of non-systematic
trial-and-error investigation of in-genre classi cation. Feature frequencies are
estimated from the training data, by (macro-)averaging frequency distributions
from di erent training data sources. Also following results from authorship
attribution, we use z-scores of frequencies as feature values. All in all, the feature
vector for a document is made up of z-scores for the 2500 most frequent words,
z-scores for the 50 most frequent characters and z-scores for the 10 most frequent
word lengths.
          </p>
          <p>
            Texts were tokenized using Cutter
            <xref ref-type="bibr" rid="ref6">(Graen et al., 2018)</xref>
            , with some provisions
to treat ascii emoticons `:P', repeated punctuation marks `???' and sequences
of unicode emoji ` ' as single words. Other punctuation was included as
any other other `word'. All text was lower-cased before constructing the feature
vectors.
          </p>
          <p>Fitting the logistic regression models was done with L-BFGS using the facilities
supplied by SciPy2 and Autograd.3
4</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Entries to the shared task</title>
      <p>For the in-genre models, we entered two groups: full models according to the
speci cations above (`in-genre 1'), and models that only uses the 2500 lexical
features (`in-genre 2'). We set the regularization hyperparamater by 5-fold
cross-validation.</p>
      <p>The results are in Table 1. Cross-validation gives fair estimates of the task
evaluation results (except for one overestimate). In the task evaluation, the full
models do better than the reduced. We speculate that the character and word
length features, being less sparse, make the models more robust. Compared to
the other entries to the shared task, both kinds of model perform reasonably well,
2 scipy.optimize.minimize, see scipy.org
3 See github.com/HIPS/autograd</p>
      <sec id="sec-3-1">
        <title>Cross-genre 1 News</title>
      </sec>
      <sec id="sec-3-2">
        <title>Twitter</title>
      </sec>
      <sec id="sec-3-3">
        <title>YouTube | Average</title>
      </sec>
      <sec id="sec-3-4">
        <title>Cross-genre 2 News</title>
      </sec>
      <sec id="sec-3-5">
        <title>Twitter</title>
      </sec>
      <sec id="sec-3-6">
        <title>YouTube</title>
        <p>| Average</p>
        <p>X-val acc
log10</p>
        <p>News Twitter YouTube Avg Eval acc Eval rank
with accuracies consistently in the top half. It should be noted that all entries
to the shared task perform well below the 80% mentioned in the introduction.
This is probably related to the training data set sizes and the fact that the task
requires prediction on the basis of just one document.</p>
        <p>For the cross-genre models, we also entered two groups of models: one group
combining the datasets using the lse-based formulation of the objective
(`crossgenre 1') and one pooling the data (`cross-genre 2'). The scaling hyperparameter
k for lse was kept constant at 20, no attempts where made to optimize it, and
was set by 5-fold cross-validation. The hyperparameter setting for the model
with the highest macro-averaged accuracy between data sources was chosen.</p>
        <p>The results are in Table 2. The cross-validated accuracies are all lower than
in the in-genre case: Apparently the models su er more from having to deal
with two di erent sources than they bene t from having larger training datasets.
Comparing accuracies of cross-genre 1 (lse) to cross-genre 2 (pooling), we can
observe that the pooled data model tends to cater better for the larger source
dataset (viz., Twitter or YouTube), although this is only very pronounced in the
model combing News and Twitter training data. The average cross-validation
accuracies of the two combination methods are very similar, except in this last
case, where lse has a slight advantage. The task evaluation accuracies are also
very similar between the two model types. As is to be expected from the shift in
genre, the cross-validation accuracy here is a very poor indication of evaluation
accuracy: the latter is on average almost 8 percentage points lower. Compared to
the other entries in the shared task, we now do a lot worse, performing in the
bottom segment. The poor results on the News genre { the least similar genre {
are to blame for this, as we are in the middle bracket for the other two genres.
5</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Post-task experiments</title>
      <p>As mentioned, the cross-genre models do not fare as well as the in-genre ones in
the shared task. In addition, the di erences between lse and pooling is also small.
Pooled data
LSE model (k=1,2,4,..,128)
(x-val setings)</p>
      <p>Pooled data
LSE model (k=1,2,4,..,128)
(x-val setings)
To see to what extent these results depend on our choice of hyperparamaters, we
trained models on two genres at di erent levels of and k, and evaluated on a
third genre. We only used the shared task's training data for these experiments.
The results are in Figure 2. Note that the results of the shared task evaluation are
not in here, since we did not use the task evaluation data. There does not seem to
be a clear, systematic di erence between accuracies for the two ways of combining
data. However, we can see that the hyperparameter settings from cross-validation
are suboptimal for both methods in all three datasets. In addition, choosing the
right hyperparameter setting for k can make a real di erence in performance,
although overall it seems that a higher k is preferable.
6</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>We have presented our e orts in the Cross-genre Gender Detection shared task,
where we aimed to compare two ways of combining data sources: simply pooling
the data vs optimizing an objective that combines the respective negative
loglikelihoods with log-sum-exp. The two methods perform similarly, and we have
not seen evidence of a real advantage of using the more involved method. However,
a set of post-task experiments does show that there is performance to be gained
from a better way of picking the hyperparamaters in both methods.</p>
      <p>In this work, we have not focussed on the feature set de nition nor studied the
e ectiveness of di erent kinds of features in any depth. In theory, these issues are
orthogonal to what we presented in our report. We thus reserve the investigation
of model combination methods in the context of known state of the art feature
sets for future work.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>David</given-names>
            <surname>Bamman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Jacob</given-names>
            <surname>Eisenstein</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Tyler</given-names>
            <surname>Schnoebelen</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Gender identity and lexical variation in social media</article-title>
          .
          <source>Journal of Sociolinguistics</source>
          ,
          <volume>18</volume>
          (
          <issue>2</issue>
          ):
          <volume>135</volume>
          {
          <fpage>160</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Angelo</given-names>
            <surname>Basile</surname>
          </string-name>
          , Gareth Dwyer, Maria Medvedeva, Josine Rawee, Hessel Haagsma, and
          <string-name>
            <given-names>Malvina</given-names>
            <surname>Nissim</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Simply the best: Minimalist system trumps complex models in author pro ling</article-title>
          .
          <source>In Experimental IR Meets Multilinguality, Multimodality, and Interaction</source>
          , pages
          <volume>143</volume>
          {
          <fpage>156</fpage>
          ,
          <string-name>
            <surname>Cham</surname>
          </string-name>
          . Springer International Publishing.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Hal Daume</surname>
            <given-names>III.</given-names>
          </string-name>
          <year>2007</year>
          .
          <article-title>Frustratingly easy domain adaptation</article-title>
          .
          <source>In Proceedings of the 45th Annual Meeting of the Association of Computational Linguistics</source>
          , pages
          <volume>256</volume>
          {
          <fpage>263</fpage>
          , Prague, Czech Republic.
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Jenny</given-names>
            <surname>Rose</surname>
          </string-name>
          Finkel and
          <string-name>
            <given-names>Christopher D.</given-names>
            <surname>Manning</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Hierarchical bayesian domain adaptation</article-title>
          .
          <source>In Proceedings of Human Language Technologies: The 2009 Annual Conference of the North American Chapter of the Association for Computational Linguistics</source>
          , pages
          <volume>602</volume>
          {
          <fpage>610</fpage>
          ,
          <string-name>
            <surname>Boulder</surname>
          </string-name>
          , Colorado. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Rob van der Goot</surname>
            , Nikola Ljubesic, Ian Matroos, Malvina Nissim, and
            <given-names>Barbara</given-names>
          </string-name>
          <string-name>
            <surname>Plank</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Bleaching text: Abstract features for cross-lingual gender prediction</article-title>
          .
          <source>In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)</source>
          , pages
          <fpage>383</fpage>
          {
          <fpage>389</fpage>
          .
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Johannes</given-names>
            <surname>Gra</surname>
          </string-name>
          en, Mara Bertamini, and
          <string-name>
            <given-names>Martin</given-names>
            <surname>Volk</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Cutter { a universal multilingual tokenizer</article-title>
          .
          <source>In Proceedings of the 3rd Swiss Text Analytics Conference - SwissText</source>
          <year>2018</year>
          , pages
          <fpage>75</fpage>
          {
          <fpage>81</fpage>
          ,
          <string-name>
            <surname>Winterthur</surname>
          </string-name>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Trevor</given-names>
            <surname>Hastie</surname>
          </string-name>
          , Robert Tibshirani, and
          <string-name>
            <given-names>Jerome</given-names>
            <surname>Friedman</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>The Elements of Statistical Learning: Data Mining, Inference, and</article-title>
          <string-name>
            <surname>Prediction</surname>
          </string-name>
          ,
          <source>Second Edition</source>
          . Springer-Verlag, New York.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Rob</given-names>
            <surname>Malouf</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Maximum entropy models</article-title>
          .
          <source>In Alex Clark</source>
          , Chris Fox, and Shalom Lappin, editors,
          <source>Handbook of Computational Linguistics and Natural Language Processing</source>
          , pages
          <volume>133</volume>
          {
          <fpage>155</fpage>
          . Wiley Blackwell.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Francisco</given-names>
            <surname>Rangel</surname>
          </string-name>
          , Paolo Rosso,
          <string-name>
            <given-names>Martin</given-names>
            <surname>Potthast</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Benno</given-names>
            <surname>Stein</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Overview of the 5th Author Pro ling Task at PAN 2017: Gender and Language Variety Identi cation in Twitter</article-title>
          .
          <source>In Working Notes Papers of the CLEF 2017 Evaluation Labs.</source>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Peter</given-names>
            <surname>Smith</surname>
          </string-name>
          and
          <string-name>
            <given-names>W.</given-names>
            <surname>Aldridge</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Improving authorship attribution: Optimizing burrows' delta method</article-title>
          .
          <source>Journal of Quantitative Linguistics</source>
          ,
          <volume>18</volume>
          (
          <issue>1</issue>
          ):
          <volume>63</volume>
          {
          <fpage>88</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Niklas</given-names>
            <surname>Zechner</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>A Novel Approach to Text Classi cation</article-title>
          .
          <source>Ph.D. thesis</source>
          , Umea University.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>