<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Author Profile Prediction Using Trend and Word Frequency Based Analysis in Text</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jamal Ahmad Khan</string-name>
          <email>J_Ahmadkhan@Yahoo.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Argentina</institution>
          ,
          <addr-line>Chile, Colombia, Mexico, Peru, 600 each Spain</addr-line>
          ,
          <country country="VE">Venezuela</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Computer Science and Software Engineering, International Islamic University</institution>
          ,
          <addr-line>Islamabad</addr-line>
          ,
          <country country="PK">Pakistan</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <abstract>
        <p>PAN 2017 Author Profiling task include two target predictions, one is to predict the gender of text authors and second is to predict the language variety. The presented approach analyzed trends and topics followed in training dataset e.g. Authors discussing Politics, Tech, Religion, Nature etc. in their respective tweets. Along with that single words and word pair frequencies were also taken into account. A cross-lingual, general, simple and flexible approach was created that could be applied over all languages without any changes according to each task language.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>Social media nowadays has become an important indicator for trend based analysis
and become a reflection of not only our personality but also a powerful way of
expression. We express our feelings about different events or persons through text in
the form of short expressive sentences that carry information about a person’s likes,
dislikes, beliefs and interests.</p>
      <p>
        The benefits of Author profiling tasks include the right idea of author’s gender,
age, regions and traits, and this possibility of finding people’s traits is of growing
importance [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] especially for the business organizations that are interested to invest in
areas following trend based analysis over different types of social media tools like
Twitter, Facebook, WhatsApp, SnapChat etc. where one or different types of
socializing tools are more used in different regions of world and marketing
organizations may also be interested to find out the reviews submitted by people of
different genders and origins in order to make strategic decisions [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        Also the task of Author profiling deals with the aspects of forensics where the
gender and specific regions of countries are important, and this task of author
profiling task of PAN@CLEF 2017 [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] deals with this aspect where one can create a
prediction model for two aspects of a tweeted text; one is the gender prediction
having two classes male or female and second is to predict language variety which in
turn can predict the specific region or regional affiliation of the person tweeting the
text.
      </p>
      <p>
        By identifying general topics discussed in social networks can provide us better
understanding of collective interests [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and trends and locations (location in the sense
of language variety) are interdependent or correlated [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The Submitted system uses
the approach of trend/topic based analysis combining with it are the word frequencies
and word pair frequencies to classify language varieties and gender.
      </p>
      <p>
        Five different trends including Politics, Technology, Nature, Travel and Religion
were analyzed. Frequent words appearing in texts of different languages were also
analyzed, because people use social media, like twitter to follow different topics and
issues that are trending in their respective geographical areas and of their specific
interests. They also use specific expressive words that are common or more frequent
in their regional areas, so term frequencies and word pair frequencies were also taken
into account because both are an important indicator for text based analysis [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
Hence both these aspects can be handy in predicting the language variety and gender
of different persons.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2 Dataset</title>
      <p>
        Like previous year, the training dataset of PAN CLEF 2017 for author profiling
task contained xml files for each language representation with an exception that
demographic traits in current dataset are processed side by side [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. A total of four
languages were presented including Arabic, English, Portuguese and Spanish. All
languages have a range of varieties shown in Table.1. Each language variety
represents a region in the world where it is used and each variety had 600 xml
documents the training dataset. Also an equal gender representation for each language
was provided in training dataset as shown in Table. 2.
Total
Examples
600 each
      </p>
      <sec id="sec-2-1">
        <title>Arabic</title>
      </sec>
      <sec id="sec-2-2">
        <title>English</title>
      </sec>
      <sec id="sec-2-3">
        <title>Portuguese</title>
      </sec>
      <sec id="sec-2-4">
        <title>Spanish Table 2. XML Documents Representation for Task for Gender</title>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3 System Methodology</title>
      <p>In order to fulfill the task of Author Profiling at PAN CLEF 17, a general system
was built so that it may be able to perform equally irrespective of language shifts from
one to another or having to jump from one basic approach to another. By performance
it’s meant that the system would predict both gender and language variety for each
author in Twitter corpus without using complex linguistic analysis methods like Part
of Speech (POS) analysis or stylistic approaches from variety to variety and for
gender prediction as well.</p>
      <p>A simple scoring approach is used where each document is assigned scores in
order to classify it as one of the classes that are created as preprocessing of training
dataset. Each class represents exactly one variety in each language representation with
two subclasses under each variety representing gender i.e. male and female.</p>
      <p>The experimental setup is divided into following steps



</p>
      <p>Trend word-lists Preparation
Dataset Preprocessing
Language Variety Classification</p>
      <p>Gender Classification</p>
    </sec>
    <sec id="sec-4">
      <title>3.1 Trend word-lists Preparation</title>
      <p>A total of five general topics were chosen including Politics, Technology, Religion,
Nature and Romance to create lists of 500 words in each language i.e. Arabic,
English, Portuguese and Spanish and for each chosen topic. These word lists are
referred as trend list in presented system, where L is the main language variety and
i is variable representing specific trend. Different sources were used in this regard.</p>
    </sec>
    <sec id="sec-5">
      <title>3.2 Dataset Preprocessing</title>
      <p>The system performs following steps in order to preprocess data for creation of
language variety classes and gender subclasses.</p>
      <p>All xml based twitter text was extracted from each language variety training
dataset discarding Hash-tags, html/xml tags, html/web links and extra white
spaces in order to get plain text from each file and all text was combined as a
single document</p>
      <p>representing a language variety. Where v is the
represented language variety in document.
2.</p>
      <p>All extracted text was combined in order to find top 100 most frequent words
and word pairs</p>
      <p>each, where t is the id for specific term
(1)
calculated.
3. Each common term occurring in both trend list
and document
is scored
according to term occurrence and Trend score
for each document is
4. Steps 2 and 3 are repeated in order to find gender based term frequency, word
pair frequency and Trend scores calculation under each language variety v.
5. Steps 1 to 4 are repeated for each Language L.
6.</p>
      <p>Hence we have created different language variety classes
and gender
subclasses</p>
      <p>, where g represents the gender i.e. male or female as shown in
7. The System loops through each language folder of dataset and every document
is processed in the same manner as a separate class with unknown
variety label u and Author Identity id having its own , and</p>
    </sec>
    <sec id="sec-6">
      <title>3.3 Language Variety Classification</title>
      <p>Following Steps are taken by the system in order to assign the class
language variety label while looping through each variety class .
with a
1. For each single word and word pair common in each
is increased for class as shown in equation 4.
and
score
∑
∑
(4)
(5)
2. For each trend score in and in
between both is calculated as shown in equation 5.
absolute difference</p>
      <p>The least trend difference score is added to class variety score and in
this way a class with unknown language variety is assigned a language variety v
having highest score .</p>
    </sec>
    <sec id="sec-7">
      <title>3.4 Gender Classification</title>
      <p>Once the Language variety of Document class has been decided, the system
repeats similar steps as shown in section 3.3 to find out the gender for class but
this time the system has to decide among only two gender subclasses of variety class
v. The gender subclass having highest score is assigned to class .</p>
    </sec>
    <sec id="sec-8">
      <title>4 Results</title>
      <p>
        A 10-fold cross validation criteria was used during the initial validation of the
system and no different approach was used for prediction of language varieties among
main languages or gender prediction as shown in above section. Once all parameters
were adjusted during system validation over training datasets, the system was ready to
run over full training and test datasets provided at TIRA [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] in order to predict both
gender and language variety.
      </p>
      <p>Following are the evaluator results shown in table 3.</p>
      <p>The results consistency distributed over different languages in above table for both
training and test datasets clearly indicate the generalization of used system; it’s
because of the designed approach that denied over-fitting while training stages. The
performance of model is good where it has to predict among smaller group of
variables either it be gender or language variety, a typical example of which is the
results of Portuguese language where the system’s prediction accuracy is highest for
both gender and variety because each has only 2 values to predict from. As the
number of values in language variety increases, the performance of system decreases
and hence the overall results are affected.</p>
    </sec>
    <sec id="sec-9">
      <title>4 Conclusion</title>
      <p>In this paper a new and generalized approach is presented which may be able to
perform with same attributes over all given languages in dataset without changing its
modes or functionality with language change. The results however indicate a major
flaw in the system over which more work is required. This flaw is the decrease in
prediction when there is an increase in number of language varieties. This can
however be improved either by increasing the quantity of single words and word pairs
while the creation of variety classes or by increasing the trends.</p>
      <p>The future work will certainly emphasize to improve the prediction eminence of
the system by following the suggestions discussed above.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1. Francisco Rangel, Paolo Rosso, Ben Verhoeven, Walter Daelemans, Martin Pottast,
          <string-name>
            <given-names>Benno</given-names>
            <surname>Stein</surname>
          </string-name>
          .
          <source>Overview of the 4th Author Profiling Task at PAN</source>
          <year>2016</year>
          :
          <article-title>Cross-Genre Evaluations</article-title>
          . In: (Eds.)
          <article-title>CLEF Labs and Workshops, Notebook Papers</article-title>
          .
          <source>CEUR Workshop Proceedings. CEUR-WS.org</source>
          , vol.
          <volume>1609</volume>
          , pp.
          <fpage>750</fpage>
          -
          <lpage>784</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2. Francisco Rangel, Paolo Rosso, Moshe Koppel, Efstatios Stamatatos,
          <string-name>
            <given-names>Giacomo</given-names>
            <surname>Inches</surname>
          </string-name>
          .
          <article-title>Overview of the Author Profiling Task at PAN 2013</article-title>
          . In: (Eds.)
          <article-title>Notebook Papers of CLEF LABs and Workshops</article-title>
          .
          <source>CEUR-WS.org</source>
          , vol.
          <volume>1179</volume>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3. Francisco Rangel, Paolo Rosso,
          <source>Martin Potthast and Benno Stein. Overview of the 5th Author Profiling Task at PAN</source>
          <year>2017</year>
          :
          <article-title>Gender and Language Variety Identification in Twitter</article-title>
          . In: (Eds.)
          <article-title>CLEF Labs and Workshops, Notebook Papers</article-title>
          .
          <source>CEUR Workshop Proceedings. CEUR-WS.org</source>
          , vol.
          <volume>10456</volume>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Ceren</given-names>
            <surname>Budak</surname>
          </string-name>
          , Divyakant Agrawal, Amr El Abbadi.
          <article-title>Structural Trend Analysis for Online Social Networks</article-title>
          .
          <source>Proceedings of VLDB Endowment (PVLDB)</source>
          ,
          <volume>4</volume>
          (
          <issue>10</issue>
          ):
          <volume>646656</volume>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Ceren</given-names>
            <surname>Budak</surname>
          </string-name>
          , Theodore Georgiou, Divyakant Agrawal, Amr El Abbadi.
          <source>GeoScope: Online Detection of Geo-Correlated Information Trends in Social Networks. PVLDB 7</source>
          (
          <issue>4</issue>
          ):
          <fpage>229</fpage>
          -
          <lpage>240</lpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Yuen-Hsien</surname>
            <given-names>Tseng</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chi-Jen</surname>
            <given-names>Lin</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu-I Lin</surname>
          </string-name>
          .
          <article-title>Text mining techniques for patent analysis</article-title>
          .
          <source>Information Processing &amp; Management</source>
          . Volume
          <volume>43</volume>
          ,
          <string-name>
            <surname>Issue</surname>
            <given-names>5</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pages</surname>
          </string-name>
          1216-
          <fpage>1247</fpage>
          (
          <issue>Sep</issue>
          ,
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7. [Online] http://www.tira.io/tasks/pan/ (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>