<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Novel Approach for Fake Comments and Reviews Detection on the Online Social Networks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Akshat Gaurav</string-name>
          <email>akshat.gaurav@ronininstitute.org</email>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
          <xref ref-type="aff" rid="aff7">7</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>B.B. Gupta</string-name>
          <email>gupta.brij@gmail.com</email>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kwok Tai Chui</string-name>
          <email>jktchui@ouhk.edu.hk</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff7">7</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dragan Peraković</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Priyanka Chaurasia</string-name>
          <email>p.chaurasia@ulster.ac.uk</email>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ching-Hsien Hsu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff7">7</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science and Information Engineering, Asia University, Taiwan &amp; Department of Computer Science and Information Engineering, National Chung Cheng University, Taiwan &amp; Department of Medical Research, China Medical University Hospital, China Medical University</institution>
          ,
          <country country="TW">Taiwan</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Hong Kong Metropolitan University (HKMU)</institution>
          ,
          <addr-line>Hong Kong</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>International Conference on Smart Systems and Advanced Computing</institution>
          ,
          <addr-line>Syscom-2021</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>National Institute of Technology Kurukshetra</institution>
          ,
          <addr-line>Kurukshetra-136119, Haryana</addr-line>
          ,
          <country country="IN">India &amp;</country>
          <institution>Asia University</institution>
          ,
          <addr-line>Taichung 413</addr-line>
          ,
          <institution>Taiwan &amp; Stafordshire University</institution>
          ,
          <addr-line>Stoke-on-Trent ST4 2DE</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Ronin Institute</institution>
          ,
          <addr-line>Montclair, New Jersey 07043, U.S</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>Ulster University</institution>
          ,
          <addr-line>Magee campus, Londonderry</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff6">
          <label>6</label>
          <institution>University of Zagreb</institution>
          ,
          <country country="HR">Croatia</country>
        </aff>
        <aff id="aff7">
          <label>7</label>
          <institution>2021 Copyright for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International</institution>
          ,
          <addr-line>CC BY 4.0</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>As the primary source of information dissemination, social media networks have surpassed traditional news organisations for the first time. Nonetheless, as the number of people who use social media websites grows, they become more susceptible to the spread of misinformation, making it increasingly dificult to distinguish between real news and false news in real time. In this paper, we proposed a machine learning technique for the detection of fake comments in social networks. According to the results of the experiment, it is clear that the machine learning technique eficiently detects the fake comments.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Deep learning</kwd>
        <kwd>Fake comments</kwd>
        <kwd>Machine learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        For many reasons, social media has overtaken email as the most important distribution and
consumption medium for news and information. For starters, getting news via social media
is usually faster and less expensive than getting news from conventional sources[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Further,
engaging and interacting with other readers by commenting, debating, and fighting with
them is a great method to get one’s points through while also encouraging participation and
participation. Despite these developments, spreading real-time information with the use of
social media has contributed to the spread of disinformation, often referred to as fake comments
and reviews.
      </p>
      <p>
        Recommendations and feedback are becoming more important as internet communication
technology advances. In today’s world, people are increasingly relying on internet reviews
to assist them make a purchasing choice. For company owners, internet comments are a way
to develop and better their enterprises. Product enhancements may be made based on user
input via online comments. However, internet remarks are not always honest, and false online
comments are common [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. There are company owners that will pay for nice or poor evaluations
regarding their rivals’ items to be written. Consumers are misled by these phoney reviews and
end up purchasing inferior goods because of it. Both consumers and sellers depend on honest
evaluations to guide their purchasing choices, and false testimonials may have a significant
impact on both. Innocent clients may sufer financial losses as a result of this. As a result, many
people are interested in learning how to spot fake comments. Most shopping websites, on the
other hand, have solely addressed the issue of negative ratings and comments. As a result, the
detection of false reviews is critical in both corporate and academic settings[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>Several academics have ofered a number of methods for identifying various cyber threats
[4, 5, 6, 7, 8] as well as phoney reviews and comments[9, 10]. In this part, we discuss some of
the most extensively used fraudulent review and comment detection techniques.</p>
      <p>There are a number of popular social networking sites, like Facebook and Twitter, that are
used by internet users to obtain information on the World Wide Web[11]. Spammers and
hackers who transmit harmful depicted objects in the form of spam through social networking
platforms are examples of content protection[12, 13].</p>
      <p>Online social networking has developed as one of the most popular methods to exchange
information and communicate with others in the course of everyday life. Online social
networking is a fun and convenient way to meet new people and keep in touch with old friends. There
are, however, worries about user privacy and account security. During events, people use social
media sites like Twitter and Facebook to disseminate false information. Author in [14] presented
a method to identify phoney accounts by studying the features of dangerous information that
spreads in real-time. Fake profiles are created by stealing the personal information of a real user
and utilising that information to establish a new profile. As time goes on, the profile gets hacked
such that it may send friend requests to a friend of the original account holder. The suggested
method outlines our chrome extension-based architecture for detecting bogus Twitter accounts
by examining several attributes.</p>
      <p>Web applications are automatically scanned for XSS attack vectors using XSS-explorer, a
universal and automated server-side flexible framework proposed by the author in [ 8]. Extensive
XSS attack investigations are generated for each web application’s injection locations, which may
be explored using the built-in XSS-explorer tool. This strategy relies on approaches that allow
for the accurate filling of injection locations in forms with relevant information. Identification
of these points allows us to search for every possible web page of the application, allowing us
to look for more attack vectors and speeding up the process of finding them.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Setup</title>
      <p>The suggested method uses six core machine learning techniques to identify bogus news.
Tokenization and stop phrasing are the first steps in the recommended technique. Stastical
approaches are used to assess the performance of each ML methodology.</p>
      <p>It might take a long time and be dificult to determine if a remark is real or not. Because
of this, a previously gathered and recognised dataset of bogus news was deployed. We drew
inspiration for our project from the Kaggle database. The dataset includes a header, title and
text columns, as well as a fake/real comment item flag.</p>
      <p>It is necessary to remove stop words from text before adding it into machine learning models.
Our models will perform better if we can use these methods to find and optimise the most
relevant terms. There are a lot of useless words and strange characters in our datasets since
we utilise real-world news items. Our data collection was simplified as a consequence of the
removal of these extraneous characters. Stop words are removed as the last step in preprocessing.
They were thus omitted from all of our testing due to the potential for excessive noise they
would have generated.</p>
      <p>In the last section, we used machine learning models to pre-process the data. In order to get
over this limitation, we employ a counter vectorized to translate the text data into vector form.
Finally, we evaluate the performance of multiple machine learning models using statistical
techniques.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Results and Disscusion</title>
      <p>It is possible to improve the classification accuracy of various techniques by varying the quantity
of labelled data used in training . When just a tenth of the labelled data is utilised, the suggested
detection methods demonstrate an improvement in performance of up to three percent, while
simultaneously improving the accuracy and lowering the false postive. After extracting and
transforming the undesired content using tokenization algorithms. Following that, we trained
distinct models using machine learning techniques. The performance of these machine learning
models is evaluated using following eqations, and result is represented in figure 1.
 =</p>
      <p>+  
  +   +  ′ +  ′
 =</p>
      <p>+  ′
  =</p>
      <p>+  ′</p>
      <p>× 
 1 −  = 2 ×  + 
(1)
(2)
(3)
(4)
(a) Accuracy</p>
      <p>(b) Precision
(c) Recall
(d) F-1 Score</p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>We examined social media reviews and comments in this research and created a novel approach
for detecting fraudulent reviews and comments. We used machine learning technologies to
identify fake reviews and comments. To show our system’s usefulness and eficiency, we
compared it to a variety of other temporal outlier detection approaches. Detecting fraudulent
reviews using review data involves various challenges. Our study could not conclusively
establish when a product is most likely to be the subject of fake reviews and comments, which
is an exciting topic for more research.
[4] A. Gaurav, A. K. Singh, Light weight approach for secure backbone construction for
manets, Journal of King Saud University-Computer and Information Sciences (2018).
[5] D. P. Ivan Cvitić, G. Praneeth, Digital forensics techniques for social media networking,</p>
      <p>Insights2Techinfo (2021) 1.
[6] Z. Zhou, A. Gaurav, B. Gupta, H. Hamdi, N. Nedjah, A statistical approach to secure
health care services from ddos attacks during covid-19 pandemic, Neural Computing and
Applications (2021) 1–14.
[7] Z. Zhou, A. Gaurav, B. B. Gupta, M. D. Lytras, I. Razzak, A fine-grained access control and
security approach for intelligent vehicular transport in 6g communication system, IEEE
Transactions on Intelligent Transportation Systems (2021).
[8] S. Gupta, B. B. Gupta, Robust injection point-based framework for modern applications
against xss vulnerabilities in online social networks, International Journal of Information
and Computer Security 10 (2018) 170–200.
[9] J. Zhao, H. Wang, Detecting fake reviews via dynamic multimode network, International</p>
      <p>Journal of High Performance Computing and Networking 13 (2019) 408–416.
[10] A. Gaurav, B. Gupta, A. Castiglione, K. Psannis, C. Choi, A novel approach for fake
news detection in vehicular ad-hoc network (VANET), in: International conference on
computational data and social networks, 2020, pp. 386–397. Tex.organization: Springer.
[11] S. R. Sahoo, B. B. Gupta, Classification of various attacks and their defence mechanism in
online social networks: a survey, Enterprise Information Systems 13 (2019) 832–864.
[12] S. R. Sahoo, B. B. Gupta, Classification of spammer and nonspammer content in online
social network using genetic algorithm-based feature selection, Enterprise Information
Systems 14 (2020) 710–736.
[13] S. R. Sahoo, B. B. Gupta, Hybrid approach for detection of malicious profiles in twitter,</p>
      <p>Computers &amp; Electrical Engineering 76 (2019) 65–81.
[14] S. R. Sahoo, B. Gupta, Real-time detection of fake account in twitter using machine-learning
approach, in: Advances in computational intelligence and communication technology,
Springer, 2021, pp. 149–159.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>K.</given-names>
            <surname>Kumari</surname>
          </string-name>
          ,
          <article-title>Online social media threat and it's solution</article-title>
          ,
          <source>Insights2Techinfo</source>
          (
          <year>2021</year>
          )
          <article-title>1</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>H.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Luo</surname>
          </string-name>
          , Y. Liu,
          <string-name>
            <given-names>G.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <article-title>Semisupervised learning based fake review detection</article-title>
          ,
          <source>in: 2017 IEEE International Symposium on Parallel and Distributed Processing with Applications and 2017 IEEE International Conference on Ubiquitous Computing and Communications (ISPA/IUCC)</source>
          , IEEE,
          <year>2017</year>
          , pp.
          <fpage>1278</fpage>
          -
          <lpage>1280</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>W.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>He</surname>
          </string-name>
          , S. Han,
          <string-name>
            <given-names>F.</given-names>
            <surname>Cai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <article-title>A method for the detection of fake reviews based on temporal features of reviews and comments</article-title>
          ,
          <source>IEEE Engineering Management Review</source>
          <volume>47</volume>
          (
          <year>2019</year>
          )
          <fpage>67</fpage>
          -
          <lpage>79</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>