<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>N. N. A. Balaji);bharathib@ssn.edu.in(B. Bharathi)
orcid:</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>SSNCSE_NLP@Authorship Identification of SOurce COde (AI-SOCO) 2020</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nitin Nikamanth AppiahBalaj</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>i B. Bharathi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of CSE, Sri Siva Subramaniya Nadar College of Engineering</institution>
          ,
          <addr-line>Tamil Nadu</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>As the amount of data and software applications increases, it becomes important to identify the true authors for ownership and liability of the work. Issues such as plagiarism in academic activities, open-source contributions, and identification of the creators of malware applications can be done using automatic authorship identification models. In this work, the performance of Character Count vectorization and TFIDF models are studied on the AI-SOCO data-set. We achieved a significant improvement from the baseline with 85% accuracy on the test-set and 92% accuracy on the dev-set.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Natural Language Processing</kwd>
        <kwd>Plagiarism detection</kwd>
        <kwd>Machine Learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>The authorship identification can solve two major problems: identification of plagiarism by
interpreting a program writing based on the previous works, and identification of the creators
of malicious applications. Both are equally intimidating and the only solution is to device a
foolproof model for identifying the author of a particular code snippet by understanding the
users’ style in writing programs. This model aims to safeguards the rights of one’s work in
academics, open-source contributions, projects, online coding contests, and all around the open
internet.</p>
      <p>Authorship detection on academic activity is important, for the identification of students
cheating on their assignments. With online classes and online submission systems being
promoted, it adds up to the need for the detection of plagiarism. In addition to this, recruitment
processes for companies opting for online coding rounds escalate the need for identification
of cheating. As other alternatives such as invigilation during academic exams and proctoring
during recruitment processes are resource-intensive and require a large labor force, authorship
identification becomes an eficient alternative.</p>
      <p>Malware software is annoying enough to ruin a person’s or even an entire organization’s time.
Even though laws against malware applications exists, it is hard to find the person involved
in the program, to stop further malicious programs, and punish the person. Devising author
identification models could instill fear in the first place, hence reducing the attempt for even
indulging in coding malicious programs.</p>
      <p>Data Set Number of users Number of programs per user Total samples
Train-set 1000 50 50,000
Dev-set 1000 25 25,000
Test-set 1000 25 25,000</p>
      <p>Total 1000 100 100,000</p>
      <p>In this work feature extraction techniques and machine learning modeling techniques are
analyzed. Char count vectorization with the Random Forest model is proposed for the authorship
identification task. The AI-SOCO data-set containing C++ programs and corresponding user-id
is used to analyze the performance of the models.</p>
      <p>The remaining of the paper is organized in the following fashion: The data-set description
and baseline analysis in Sectio2n,followed by the model architecture in Secti3o, nresults and
discussion in Section4 and conclusions in the Sectio5n.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Data-set Description and Baseline Analysis</title>
      <p>The data-set is a collection of C++ programs by 1000 users fromctohdeeforces programming
platform. The programs are coded in diferent versions of C++. It contains 100 (50 for training,
25 for development, and 25 for testing) diferent code snippets of each user. It contains a
mapping of user-ids (uid) to the program-ids (pid) and the set of programs. These program
ifles are verified and ready to compile, working programs that were submitted to the coding
platform. As the data-set contains an equal number of samples for each testing individual, the
data-set is fairly distributed, without any imbalance. The distribution of the data-set among the
train, test, and dev-set is explained in Tab1l.e</p>
      <p>The baseline model considered by the AI-SOCO challenge is a char Count Logistic regression
model which gives an accuracy of 29.252%. In addition to this model, a TFIDF feature extractor
with KNN based baseline with 10k features and k=25 is considered. This model shows an
accuracy of 62.128%1[].</p>
    </sec>
    <sec id="sec-3">
      <title>3. Architecture and Evaluation Scheme</title>
      <p>The architecture included three parts - feature extraction, classification, performance evaluation.
The C++ program file contains raw code, hence Count and TFIDF vectorizers are considered
for feature extraction. These vectors and their corresponding uids now become a numerical
classification problem, which is trained with Naive Bayes (NB) and Random Forest (RF) classifiers.
As the data-set is well balanced the accuracy score is considered for comparing and evaluating
the performance of the model. The scikit-lear2n] [text feature extractor is used for converting
the text into numerical vectors and scikit-learn2’]sR[F and NB implementations are used for
classification.</p>
      <p>C++ Program Code</p>
      <p>from data-set
Char (Count/TFIDF)</p>
      <p>Vectorization
n-gram = 2-5</p>
      <p>RF/NB</p>
      <p>Classifier</p>
      <p>User id labels</p>
      <p>Accuracy calculation</p>
      <sec id="sec-3-1">
        <title>3.1. Count Vectorization</title>
        <p>The count vectorization provides a primitive but powerful and simple way to transform the
program documents to numerical vectors, as could be seen from previous w3o]r.kT[he count
vectorization generates vectors of length equal to the length of the vocabulary the vectorizer is
trained on and each value presents the number of instances the particular character or word
appears in a document. This is helpful in the case of author prediction as some authors would
mostly use a particular set of variable names for frequent mundane tasks. For instance for
looping some show preference tfoor loops, whereas some prefewrhile loops, so is the bias
with if-else andswitch-case statements. Hence the set of characters or sequences of character
diference is an important factor for studying a person’s individual style of programming.</p>
        <p>The count vectorizer is trained on the AI-SOCO corpus and this transformation function is
used to fit the classification model on the train-set. Character Count vectorization with diferent
n-gram ranges was analysed and the range of 2-5 was the best performing. The random forest
proved to be the best model for fitting the large sparse matrix generated by the count vector
transformation.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. TF-IDF vectorization</title>
        <p>
          Term Frequency Frequency Inverse Document Frequency (TF-IDF) method is an extension of
the Count vectorization technique. This method has shown significant results with feature
extraction from program code4s,[
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. The particular diference between the two techniques is
that the TF-IDF method has an additional inverse frequency term to give importance to the rare
words/characters in the documents. It gives a weight for each character based on its frequency
of occurrence in all the provided documents. This reduces the significant impact of too banal
words on the vector output. This could be of great importance as the common syntax of a C++
program includes the tokens and keywords limkeain, struct, class, int, float, etc., These are used
by all the users and could be of lesser importance and hence is removed for concentrating only
on the important lexicons.
        </p>
        <p>Similar to the count vectorizer, the TF-IDF vectorizer also shows good performance with an
n-gram range of 2-5 with a random forest classifier.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results and Observations</title>
      <p>In this section the vectorization techniques are compared based on the performance on the
dev-set and test-set accuracy scores. Both the techniques generated a significant improvement
when compared to the baseline results. The random forest classifier is observed to perform
better than the Naive Bayes classifier. With respect to the Char Logistic regression model there
is a huge increase in accuracy from 29% to 92% on the dev-set. Similarly with respect to the
TFIDF-KNN baseline model there is a increase in accuracy from 62% to 92% with around 48%
improvement on the dev-set. The detailed scores on the dev-set and test-set are shown in Table
2 and Table3 respectively.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>Authorship identification from program code is becoming important in recent times due to
the rapid increase in the number of free-to-use applications, online academic and research
contributions. This could help to regulate plagiarism and detect the creators of malware
softwares. In this work, Char count based feature extraction technique along with a Random
Forest classifier is proposed for the AI-SOCO data-set. Our system showed significant increase
in performance with respect to the baseline model with an accuracy of 85% on the test-data and
92% on the dev-set.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Fadel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Musleh</surname>
          </string-name>
          , I. Tufaha,
          <string-name>
            <given-names>M.</given-names>
            <surname>Al-Ayyoub</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Jararweh</surname>
          </string-name>
          , E. Benkhelifa,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          ,
          <article-title>Overview of the PAN@FIRE 2020 task on Authorship Identification of SOurce COde (AI-SOCO), in: Proceedings of The 12th meeting of the Forum for Information Retrieval Evaluation (FIRE</article-title>
          <year>2020</year>
          ), CEUR Workshop Proceedings, CEUR-WS.org,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>F.</given-names>
            <surname>Pedregosa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Varoquaux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gramfort</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Michel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Thirion</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Grisel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Blondel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Prettenhofer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Weiss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Dubourg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Vanderplas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Passos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Cournapeau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Brucher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Perrot</surname>
          </string-name>
          , E. Duchesnay,
          <article-title>Scikit-learn: Machine learning in Python</article-title>
          ,
          <source>Journal of Machine Learning Research</source>
          <volume>12</volume>
          (
          <year>2011</year>
          )
          <fpage>2825</fpage>
          -
          <lpage>2830</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ramırez-de-la Cruz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Ramırez-de-la Rosa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Sánchez-Sánchez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Luna-Ramırez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Jiménez-Salazar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Rodrıguez-Lucatero</surname>
          </string-name>
          , Uam@ soco
          <year>2014</year>
          :
          <article-title>Detection of source code reuse by means of combining diferent types of representations</article-title>
          ,
          <source>FIRE [4]</source>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Phani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lahiri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Biswas</surname>
          </string-name>
          ,
          <article-title>Personality recognition in source code working note: Team besumich</article-title>
          .,
          <source>in: FIRE (Working Notes)</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>16</fpage>
          -
          <lpage>20</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Giménez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Paredes</surname>
          </string-name>
          ,
          <article-title>Prhlt at pr-soco: A regression model for predicting personality traits from source code</article-title>
          .,
          <source>in: FIRE (Working Notes)</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>38</fpage>
          -
          <lpage>42</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>