<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>AmritaNLP@PAN-RusProfiling : Author Profiling using Machine Learning Techniques</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Vivek Vinayan</string-name>
          <email>vivekvinayan82@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Naveen J R</string-name>
          <email>naveenaksharam@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Harikrishnan NB</string-name>
          <email>harikrishnannb07@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anand Kumar M</string-name>
          <email>m_anandkumar@cb.amrita.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Soman KP</string-name>
          <email>kp_soman@amrita.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>@BorisVasilevski3 главных вопроса для постановки целей</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>ACM Reference Format: Vivek Vinayan</institution>
          ,
          <addr-line>Naveen J R, Harikrishnan NB, Anand Kumar M and Soman KP. 2017</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Author Pro ling</institution>
          ,
          <addr-line>Russian Language, Text Classi cation, Semi-supervised Classi ers</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Center for Computational Engineering and Networking Amrita University</institution>
          ,
          <addr-line>Coimbatore</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper illustrates work done on "Gender Identi cation in Russian texts (RusPro ling)" shared task, hosted by PAN in conjunction with FIRE 2017. The task is to predict the author's gender, based on the Twitter data corpus which is in Russian. We will give a brief introduction to the task at hand, elaborate on the data-set provided by the competition organizers, discuss various feature selection methods, provided experimental analysis that we followed for feature representation and show comparative outcomes of di erent classi ers that we used for validation. We submitted a total of 3 models and their respective prediction for each test data-set with slightly di erent pre-processing technique based upon the test corpus content. As each of the test corpus were sourced from various platforms, this made it challenging to stick to one representation alone. As per the global ranking published for the shared task[6] our team secured 2nd position overall (Concatenating all Data-set) and our 3rd submission model performed the best among the 3 submission models from the overall test data corpus. Further under extended work we discuss in brief how hyper parameter tuning of certain attributes extend our validation accuracy by 6% from baseline.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>The Internet, it is a vast platform where anyone can have access
to myriads of information, from online news media articles to
various social media platforms, from personal blogs to personalized
websites, all this literally at the end of our ngertips, and in this
present age, life is becoming unimaginable without it. With the
availability to all this resources, people are writing and share
information more avidly over the internet than ever before, and it also
provides a certain degree of anonymity while doing so. Access to
such multitudinous information brings in certain set of problems
like theft of identity/content and plagiarism to name a few and this
we are trying to address with tasks such as "Author Pro ling".</p>
      <p>
        In "Author pro ling" [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] sharedtask, we examine the style of an
individual author and thus distinguish between classes of authors by
studying their sociolect aspects. In an even broader manner it helps
in predicting an author’s demographic, personality, education and
socio-networks through classi cation of texts into classes, based
on the stylistic choices of the author.
      </p>
      <p>
        With this paper on RusPro ling sharedtask, we focus on cross-genre
gender identi cation in Russian texts[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], which is becoming a part
of, one of the most upcoming trending task in NLP domain, under
"Author Pro ling"[
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ].
      </p>
      <p>
        In this task we have Twitter as training data corpus and as test
data corpus we have dataset from multiple social media domain
platforms like Twitter, Facebook, online reviews (where texts are
describing images, or letters to a friend), product and services. The
focus with this task is on gender pro ling in social media and the
main interest is in the everyday language and on how the basic
social and personality skills re ects on their writing [
        <xref ref-type="bibr" rid="ref3 ref5">3, 5</xref>
        ].
      </p>
      <p>The main challenge in this task is the language itself, as it is not
a native, thus we used certain pre-processing methods and built
our baseline representation on which we implemented classical
machine learning algorithms for this text classi cation task.
2</p>
    </sec>
    <sec id="sec-2">
      <title>CORPUS</title>
      <p>The Corpus for the training data was mainly sourced from social
media website Twitter and the labels were annotated for each of the
document with the author gender "male" or "female". The training
corpus is a collection of 600 data le in XML format which consists
of exactly half female and half male genre documents, the le name
are annotated by their associated gender label in a separate le
called "truth" which is in text format.</p>
      <p>Training Dataset</p>
      <p>Total number of documents
Total number of male documents
Total number of female documents</p>
      <p>A cursory analysis of the training corpus reviled that each
training data le has a combination of di erent tags and hyperlinks,
further the documents varied in count of content words i.e one
document went from no text to others over 3000 plus words in a
single document. Few of the les had mixed data of Russian and
English, where as few other where completely in English language.</p>
      <p>The test corpus is presented in 5 folders varied by the category
of di erent sources. Each set contains di erent amount of les, the
count of documents for each category varies from 96 to 776 les
each. On further inspection the text format provided in each folders
apart from the 3rd folder is di erent when comparing with the
training corpus, namely o ine texts, Facebook, Twitter, product
and online reviews and gender imitation corpus in order of their
folder number respectively as shown in Table 1-2.</p>
      <p>We have also taken the statistical data of the complete
vocabulary size that we gained from grid search of attributes namely
n-gram_range and min_df count. In each combination their
respective corpus size is found, and the statistics are tabulated in Table
3.
3</p>
    </sec>
    <sec id="sec-3">
      <title>METHODOLOGY</title>
      <p>
        The Figure 1 gives a rudimentary picture of the architecture that
we have implemented for our 3 submissions, in all of these models
we mainly focused on data pre-processing methods to incorporate
various features and build upon each one of them to improve the
feature representation. We started from a simple count based model,
the same methods are discussed next.
The feature selection was a process in which we started by building
a baseline model [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and improved on the accuracy of the model
with step by step empirical procedure of combining and modifying
the existing feature representation [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <sec id="sec-3-1">
        <title>Count Based Matrix :</title>
        <p>The 1st approach from the dataset was to form a simple
count based Term Document(TD) and Term Frequency
Inverse Document Frequency (TFIDF) matrix which became
the baseline for our accuracy and further went with adding
general features to previous representation.</p>
      </sec>
      <sec id="sec-3-2">
        <title>Feature Extraction :</title>
        <p>With the knowledge of the social media network "Twitter",
we essentially narrowed down our focus on search for
features to tags, like ’@’ which is mainly used to address
people/gathering and hash tag ’#’ which is based on the context
or the image of the adjoining content. Moving on, we found
that URLs were being used widely across most of the dataset
which linked to various internet sources, so we then
incorporated these as a feature to the earlier feature representation
which proved to show slight improvement on all of the
classi cation algorithms, captured below in Figure 3-4.</p>
      </sec>
      <sec id="sec-3-3">
        <title>Data Normalization :</title>
        <p>On further analysis we found that, individual URL’s in itself
as a feature seemed fruitless, thus only considering the
hyperlink itself, we focused on normalizing these across the
dataset and went with the count of the URL and those of the
tag’s as feature to represent a document. It proved to increase
the accuracy little more, This further led to normalizing of
various emoticons represented by a keyword and various
other punctuation like the exclamation mark ’!’, period ’.’ and
hyphen ’-’ which occurred multiple time or in continuous
repetition were converted to a single instance of each.</p>
      </sec>
      <sec id="sec-3-4">
        <title>Word Average: :</title>
        <p>As we were not familiar with the language, we considered
the average word length as the total number of character
per document to the total number of feature instance in that
document and appended that list as an average word length
'</p>
        <p>Training Data:
was not at this game and did not see the game, now
Processing Data:
&amp;
%
per document making it an independent feature. This is to
accommodate for the fact that the average vocabulary word
length that gender used can also be taken as a discriminative
feature between the 2 classes.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4 EXPERIMENT AND DISCUSSIONS</title>
      <p>As a part of experimental analysis we manually ran over few
random training documents based on the individual sizes of the le to
gather a glimpse of the overall change in data, then ran snippets
on these training set data to see the scale of, improvement of,
accuracy with the above considered parameters. Thus to distinguish for
better feature representation for the classi cation.</p>
      <p>After going through various transitions, the selected certain
features were extracted and then used as a part of pre-processing of
the entire training corpus these features were individually added
one set at a time to show the increase or decrease in their
accuracy corresponding to various classi ers by cross validation with
di erent classical ML classi ers, namely Logistic regression (LR),
Support Vector Machine (SVM) using linear kernel, Decision tree
$
(1) A simple count based matrix is taken to achieve baseline
accuracy, for this we have considered both TD and TFIDF
matrix representation from which we set a base-line accuracy
of 80.5% ( We randomly initialized few attributes like
ngram_range = 2 , min_df = 3 and used a linear SVM classi er
to get the baseline).
(2) Count of ’http’ and ’https’ are taken and converted to a single
key word ’https’ as this will help in adding feature to see
the usage of URLs between the 2 class distinguishing which
gender base might have used more number of hyperlinks
within their tweets.
(3) Count of ’#’ tags was further attributed to the previous
representation.
(4) Replaced emoticons with keyword.
(5) Took the average word length in a document i.e count of
character to number of feature instances as the language,
this we chose as a preferred method.
5</p>
    </sec>
    <sec id="sec-5">
      <title>FEATURE REPRESENTATIONS MODELS</title>
      <p>We submitted a total of 3 models/run’s and for each individual run
the following pre-processing method have been followed:
Submission 1 : We have considered feature representation 2,
3, 4, 5 and also the normalization of ’@’ followed by content
tags to simple key word( Splitting tags from their context
otherwise to preserve the word content in particular did
not show much di erence in validation accuracy), and used
SVM classi er for classi cation. Based on learned model
from training corpus the prediction for the test corpus’s
were taken.</p>
      <p>Submission 2 : The same feature representation as the 1st
run was considered, but we used a di erent classi er, we
took Adaboost based training model and the prediction for
the test corpus’s were taken
Submission 3 : In this run, we considered mostly with
regard to the other test datasets 1,2,4 and 5 where the content
are in longer and in paragraph form rather than the shorter
version and there was less to no use of tags and or hyperlinks.
Thus to normalize this we disregarded the above used tags
and reduced any extended repeat of punctuation’s to a single
count(e.g:’...’ is shortened to a single ’.’)</p>
      <p>A sample of this is shown in Figure 2.
6</p>
    </sec>
    <sec id="sec-6">
      <title>RESULTS</title>
      <p>
        As per the global ranking published for the shared task by the
organizers[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] our team secured 2nd position overall (Concatenating
all dataset). From the rankings our 3rd submission performed the
best compared to our team’s previous 2 submissions by a margin
of 1%, 2% respectively whereas from the leading team we trailed
by margin of 6%, this was w.r.t to the facts that we mentioned in
our submission 3 and also we got better validation accuracy for the
submission Model 3 for datasets 1,2,4 and 5.
      </p>
      <p>Individually, submission 3 gaining 2nd best accuracy in "o line
texts (picture descriptions, letter to a friend etc.)from RusPro ling
corpus" whereas the submission 1 gained our team 2nd place for
"gender imitation corpus" and 3rd in "product and service online
reviews corpus".
7</p>
    </sec>
    <sec id="sec-7">
      <title>EXTENDED WORK</title>
      <p>In our earlier experiments we randomly initialized our attributes
like max feature length, n_gram and min_df with 10000, 2 and 3
respectively. As a motive to increase validation accuracy we
performed a grid search for the hyper-parameter namely word count,
n_gram and min_df based values with the TFIDF model, where we
considered the following range of data values for each:
Word count : 10000 - 50000
ngram-range : 2 - 6
Min-df : 1 - 4</p>
      <p>After applying grid search we pushed the baseline accuracy to
82.5% when initializing max_feature with 10,000, n_gram with 2
and min_df with 1 and applying a linear SVM classi er. We further
pushed our validation to 86.49% by applying Adaboost classi er.
Over all we found that the trend of accuracy of TD feature
representation model decreased with increase in all the attributes, and the
accuracy of TFIDF feature representation model increased but
saturated after n_gram value exceeds 6 and the min_df value exceeds 4,
the same is show in Figure 5 where the best of the attributes, feature
combination were taken for each TD and TFIDF representation.
The challenge in this shared task we faced was the fact that we
were working on a language corpus that is non native to us, thus we
mainly focused on pre-processing and normalizing the data corpus
to get improved feature representation. We built from a basic count
representation and incorporated simple modi cation on iterating
feature representation and observed the various accuracy changes
involved with those features. Based on the experimental analysis
and further discussion on optimizing of the various attributes in
the extended work, we could make an inference that the baseline
can further be increased which could better improve the prediction,
fetching us better gender identi cation model.</p>
      <p>
        As a future study we considered making various embedded
representation for the Russian corpus and use deep learning techniques
for categorizing author gender [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. As these methods require more
number of training instances we are considering including certain
additional corpus provided by PAN [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] for this task and also
consider certain portions of labelled test dataset based on the variety
of the source that they are taken from.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>H.B.</given-names>
            <surname>Barathi Ganesh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. Anand</given-names>
            <surname>Kumar</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.P.</given-names>
            <surname>Soman</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Statistical semantics in context space: Amrita CEN@author pro ling</article-title>
          .
          <source>CEUR Workshop Proceedings</source>
          <volume>1609</volume>
          (
          <year>2016</year>
          ),
          <fpage>881</fpage>
          -
          <lpage>889</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>H.B.</given-names>
            <surname>Barathi Ganesh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.</given-names>
            <surname>Reshma</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M. Anand</given-names>
            <surname>Kumar</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Author identi cation based on word distribution in word space</article-title>
          .
          <source>2015 International Conference on Advances in Computing, Communications and Informatics</source>
          ,
          <string-name>
            <surname>ICACCI</surname>
          </string-name>
          <year>2015</year>
          (
          <year>2015</year>
          ),
          <fpage>1519</fpage>
          -
          <lpage>1523</lpage>
          . https://doi.org/10.1109/ICACCI.
          <year>2015</year>
          .7275828
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Celli</surname>
          </string-name>
          , Bruno Lepri,
          <string-name>
            <surname>Joan-Isaac</surname>
            <given-names>Biel</given-names>
          </string-name>
          , Daniel Gatica-Perez,
          <string-name>
            <given-names>Giuseppe</given-names>
            <surname>Riccardi</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Pianesi</surname>
          </string-name>
          .
          <year>2014</year>
          . The Workshop on Computational Personality Recognition. (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Tatiana</given-names>
            <surname>Litvinova</surname>
          </string-name>
          , Olga Litvinlova, Olga Zagorovskaya, Pavel Seredin, Aleksandr Sboev, and
          <string-name>
            <given-names>Olga</given-names>
            <surname>Romanchenko</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>" Ruspersonality": A Russian corpus for authorship pro ling and deception detection</article-title>
          .
          <source>In Intelligence, Social Media and Web (ISMW FRUCT)</source>
          ,
          <source>2016 International FRUCT Conference on. IEEE</source>
          , 1-
          <fpage>7</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Tatiana</given-names>
            <surname>Litvinova</surname>
          </string-name>
          and
          <string-name>
            <given-names>Olga</given-names>
            <surname>Litvinova</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Authorship Pro ling in RussianLanguage Texts</article-title>
          .
          <source>In Proceedings of 13th International Conference on Statistical Analysis of Textual Data (JADT</source>
          <year>2016</year>
          ), University Nice Sophia Antipolis, Nice.
          <fpage>793</fpage>
          -
          <lpage>798</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Tatiana</given-names>
            <surname>Litvinova</surname>
          </string-name>
          , Francisco Rangel, Paolo Rosso, Pavel Seredin, and
          <string-name>
            <given-names>Olga</given-names>
            <surname>Litvinova</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Overview of the RUSPro ling PAN at FIRE Track on Cross-genre Gender Identi cation in Russian</article-title>
          .
          <source>In Notebook Papers of FIRE</source>
          <year>2017</year>
          , FIRE-2017, Bangalore, India, December 8-10, CEUR Workshop Proceedings. CEUR-WS.org
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Francisco</given-names>
            <surname>Rangel</surname>
          </string-name>
          , Paolo Rosso, Moshe Koppel, Efstathios Stamatatos, and
          <string-name>
            <given-names>Giacomo</given-names>
            <surname>Inches</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Overview of the author pro ling task at PAN 2013</article-title>
          .
          <article-title>Notebook Papers of CLEF (</article-title>
          <year>2013</year>
          ),
          <fpage>23</fpage>
          -
          <lpage>26</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Francisco</given-names>
            <surname>Rangel</surname>
          </string-name>
          , Paolo Rosso, Ben Verhoeven, Walter Daelemans,
          <string-name>
            <given-names>Martin</given-names>
            <surname>Potthast</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Benno</given-names>
            <surname>Stein</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Overview of the 4th author pro ling task at PAN 2016: cross-genre evaluations</article-title>
          .
          <source>Working Notes Papers of the CLEF</source>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Francisco</given-names>
            <surname>Manuel Rangel Pardo</surname>
          </string-name>
          , Fabio Celli, Paolo Rosso, Martin Potthast, Benno Stein, and
          <string-name>
            <given-names>Walter</given-names>
            <surname>Daelemans</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Overview of the 3rd Author Pro ling Task at PAN 2015</article-title>
          .
          <article-title>In CLEF 2015 Evaluation Labs</article-title>
          and Workshop Working Notes Papers.
          <volume>1</volume>
          -
          <fpage>8</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Aleksandr</surname>
            <given-names>Sboev</given-names>
          </string-name>
          , Tatiana Litvinova, Dmitry Gudovskikh, Roman Rybka, and
          <string-name>
            <given-names>Ivan</given-names>
            <surname>Moloshnikov</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Machine Learning Models of Text Categorization by Author Gender Using Topic-independent Features</article-title>
          .
          <source>Procedia Computer Science</source>
          <volume>101</volume>
          (
          <year>2016</year>
          ),
          <fpage>135</fpage>
          -
          <lpage>142</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Aleksandr</surname>
            <given-names>Sboev</given-names>
          </string-name>
          , Tatiana Litvinova, Irina Voronina, Dmitry Gudovskikh, and
          <string-name>
            <given-names>Roman</given-names>
            <surname>Rybka</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Deep Learning Network Models to Categorize Texts According to Author's Gender and to Identify Text Sentiment</article-title>
          .
          <source>In Computational Science and Computational Intelligence (CSCI)</source>
          ,
          <source>2016 International Conference on. IEEE</source>
          ,
          <fpage>1101</fpage>
          -
          <lpage>1106</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>