<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Using information retrieval to evaluate trustworthiness assessment of eshops</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Claudio Carpineto</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giovanni Romano</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Davide Lo Re</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Fondazione Ugo Bordoni</institution>
          ,
          <addr-line>Rome</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>To protect consumers from online counterfeiting, several systems have been recently made available that check whether an ecommerce website is trustworthy or not. In this paper we study how to evaluate and compare trust checkers using an information retrieval methodology to gather suitable data and build a ground truth test collection. The main ndings of our experimental evaluation are that the inter-checker agreement was moderate and that the single trust checkers presented relative advantages and disadvantages. Essentially, review-based systems were very precise but largely incomplete, whereas feature-based systems provided assessments for any ecommerce website being submitted but were more prone to errors.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Online counterfeiting continues to be a thorn in the side of both consumers and
enterprises, and there are signs that the problem is worsening despite growing
e orts to combat it [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. One particular but important aspect of this
phenomenon is represented by ecommerce websites selling counterfeit goods. These
fake websites can attract purchasers and enhance their attempts at deception by
using a number of channels of promotion and advertising such as email, social
networking, web ads, and search engine optimization techniques. On the other
side, better technical countermeasures have begun to appear.
      </p>
      <p>
        As an anti-counterfeiting aid, some desktop web browsers now remove the
full URL from view in the URL bar and display only the domain name, which
is usually short and clear for legitimate websites (as opposed to long and messy
strings for fake ones). To better protect consumers from being led into a
swindle, several research and commercial systems have been recently developed that
explicitly assess whether a given website is trustworthy or not. Although these
systems employ di erent paradigms, algorithms, and information sources, they
can be roughly grouped in two main categories, namely those based on user
reviews and white/black lists (e.g., WOT mywot.com, Trustpilot trustpilot.com,
Webutation webutation.net, Scamvoid scamvoid.com), and those making use of
website features; e.g., Scamadviser scamadviser.com, [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>The relative availability of trust checkers raises the question of their
evaluation and comparison. To the best of our knowledge, this issue has not been
addressed so far. In this paper we focus on two main research questions: do
current trust checkers agree on each other?, which is the best trust checker?</p>
      <p>To answer these questions, we present an information retrieval methodology
consisting of three main steps. We rst collect data retrieved by major search
engines in response to search queries with brand names; then we identify
ecommerce websites (whether legitimate or fake) in search results with a suitable
classi er and use them for measuring inter-checker agreement; and nally build
a ground truth dataset containing manually-labeled legitimate and fake
ecommerce websites used for measuring the accuracy of trust checkers.
2</p>
      <p>
        Why assessing trustworthiness of ecommerce websites
is a di cult task
Assessing whether an ecommerce website is trustworthy or not may be di cult
even for humans. In [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], it is shown that non-experts were inconsistent in
discriminating between legitimate and fake ecommerce websites containing discounted
o ers. To illustrate the di culty of this task, in Figure 1 we show two websites
selling products of two well known luxury brands, namely `Hugo Boss' (a) and
`Iceberg' (b). At rst glance, website (a) has a nice look and feel, o ers
products with reasonably discounted prices, and shows a well designed navigation
system for nding products. It also presents other features that might reassure
shoppers that the store is trustworthy. Security seals are in place, a live chat is
provided, a Facebook label is displayed. Further, on a more technical side, the
brand name (i.e., Hugo Boss) is in the URL's path and not in the domain name,
as customary in legitimate ecommerce retailers (except for the o cial website of
the brand). However, on closer inspection, we realized that the security label was
appropriated without signing up with its vendor, that the store did not have its
own Facebook page, and the live chat did not work. Also, more important, the
website did not provide any contact information except for a contact form to be
lled in by shoppers. This is a clear sign that the website may be fake. Turning to
website (b), we see that it has several nice characteristics to recommend itself.
In particular, very detailed contact information are displayed in the footer to
increase trust of shoppers, including telephone number and address of physical
stores. However, we found that the provided contact information were contained
in an image and they turned out to be dummy contacts, on more inquiry; the
website did not even o er an email address. This was probably another scam
website.
      </p>
      <p>
        Trust checkers may use people's evaluations and reviews in their assessments,
or detect speci c mechanisms employed by fake websites such as cloaking [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]
and search redirection [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. A more comprehensive approach consists of training a
classi er on a large set of learning features, possibly including the earlier
mechanisms. Two recent trust classi ers are [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], the latter of which makes
use of 33 learning features spanning product o er, merchant information,
payment methods, website registration data, ecommerce-speci c SEO, and relative
behavior of website. Unlike users, automatic checkers can take into account a
large number of a website's features simultaneously, and they can easily access
useful external information that may be not readily available to consumers; e.g.,
white/black lists, consumer reviews, WHOIS data, Alexa metrics, other checker's
assessments, etc. On the other hand, fake websites continue to take steps to
increase their similarity to genuine ones, so that sometimes it may be required a
deeper understanding of what is behind a feature.
3
      </p>
    </sec>
    <sec id="sec-2">
      <title>Gathering legitimate and fake ecommerce websites</title>
      <p>
        First of all we needed to collect data containing (legitimate and) fake ecommerce
websites. One of the major promotion channels used by counterfeiters is
manipulation of search engine results. As brand names are extensively searched on
the web, usually with a shopping intent [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], online sellers of counterfeits can use
search engine optimization techniques to achieve high listings in search results.
This way, they may increase tra c on their websites by fraudulently attracting
purchasers seeking to buy genuine products or `complicit' shoppers who are
willing to buy replicas. Here we talk about entirely bogus websites, not legitimate
sellers that may occasionally sell counterfeit products. Their massive presence
in search-engine results has been observed in some recent studies [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>We collected search results for 39 major shoe brands (e.g., `Armani',
`Christian Louboutin', `Gucci', `Prada', `Valentino') through the following procedure.
We used three types of queries. One neutral query consisting of the brand name
followed by `shoes', to help with ambiguity in brand names and focus on the same
product for multi-product brands; one biased query were we added `cheap' to
the brand name and `shoes', to emphasize that shoppers were seeking discounted
o ers; one complicit query formed by adding `replica' to the brand name and
`shoes', to clearly indicate that users were happy with counterfeit products. The
queries were automatically submitted to three search engines (i.e., Bing, Google,
and Yahoo!) set to the English language, and the rst 100 results were saved. In
all we collected about 35000 search results. The collected URLs were then
postprocessed. We performed URL normalization and grouped together the URLs
with a same domain name because they usually refer to distinct o ers from the
same eshop. We kept as a group representative one randomly selected group
member. After this operation, the number of URLs reduced to about 24000.</p>
      <p>
        Search results for brand queries usually contain many ecommerce webistes
but also many non-ecommerce websites such as shopbots, product catalogues,
shop locators, ecommerce blogs, etc. The next step was to identify proper
ecommerce websites in search results. This is a research task in its own. We used
a speci c classi er described in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] that makes use of 24 classi cation features
about various aspects of the website including product navigation and search,
product display, purchase management, and customer service information. In an
experimental evaluation, this classi er was able to discriminate between
ecommerce and non-ecommerce websites with a classi cation accuracy of about 90%.
By running this classi er on the search results we identi ed about 10000
ecommerce websites (whether legitimate or fake), with some approximation error.
weband
March,
      </p>
      <p>The next step was to build a ground-truth dataset containing eshops
labeled as legitimate or fake. As this is a time-consuming process that requires
trained personnel, the size of the ground truth dataset was small. Starting from
the large dataset containing eshops described above, we manually labeled
randomly selected items until a balanced dataset containing 255 legitimate and fake
ecommerce websites was found (by downsampling the class containing legitimate
websites).
4</p>
    </sec>
    <sec id="sec-3">
      <title>Trust checkers</title>
      <p>
        We used four trust checkers: WOT, Trustpilot, Scamadviser, and RI.SI.CO. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].1
WOT and Trustpilot are mainly based on lists of websites and on user reviews.
Scamadviser essentially uses WHOIS data, but the details of both the features
and the algorithm are not disclosed. RI.SI.CO. is an SVM classi er with a large
set of features. The interfaces to the four systems are shown in Figure 2. The
three commercial system take an URL as an input and return a weighted
assessment of thrustworthiness, whereas the input to RI.SI.CO. is a brand name
and its output consists of a list of fake eshops. In Figure 2, we show (top left)
the output of RI.SI.CO. for the query `Christian Louboutin' (i.e., a luxury shoe
brand). We also show the outputs of Scamadviser (right top) and WOT (left
bottom) for one (highlighted) result of RI.SI.CO. Scamadviser rated it as
moderately unsafe, with a score of 52% in a range from 0% to 100% safeness), while
for WOT it was clearly untrustworthy; i.e., 9 out of 100. Trustpilot was not able
to assess the website retrieved by RI.SI.CO.; we show the output of Trustpilot
(right bottom) for Zalando, a legitimate website receiving four stars out of ve.
      </p>
      <p>
        Strictly speaking, the assessment weights returned by WOT, Trustpilot, and
Scamadviser are scores, not probabilities. In order to compare the four
systems, we had to convert assessment weights into binary classi cation values.
For WOT and Scamadviser we used a simple splitting criterion based on equal
width intervals; i.e, threshold = 0.5. While more powerful, supervised methods
are conceivable [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], we will see in Section 1 that it yielded results close to the
theoretical upper bound of performance. For Trustpilot, we converted any score
into the class "legitimate" (the rationale is explained below).
5
      </p>
    </sec>
    <sec id="sec-4">
      <title>Inter-checker agreement</title>
      <p>The rst experiment was aimed to measure the consistency of rating across
different checkers. One practical constraint concerned the time necessary to gather
the rates. Because APIs were available only for WOT, it was not possible to
run all the checkers through the whole dataset containing about 10000
ecommerce websites. To keep the times manageable, we used a subset containing 632
1 RI.SI.CO. is available at http://uibm-ici.fub.it/risico with a password-protected
access. It was developed for and in cooperation with the Directorate-General for the
Fight against Counterfeiting - Italian Patent and Trademark O ce.</p>
      <p>
        (a)
(c)
(b)
(d)
randomly extracted URLs. We submitted these URLs to the checkers and
collected the binarized assessments, used to measure the inter-checker agreement.
First of all we noted that Trustpilot and WOT were able to assess only a subset
of the URLs, respectively 81 and 351 URLs. We removed Trustpilot from this
evaluation and considered the 351 URLs assessed by all three remaining
systems. The percent agreement (calculated by averaging the number of agreement
scores divided by the total number of scores for each URL) was 88%, which can
be seen as a moderate agreement for this simple measure of interrater reliability
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. More speci cally, we found that only for 232 URLs out of 351 the three
checkers returned the same class label, which means that in more than one third
of cases they were not able to make a unanimous decision. An analysis of
pairwise consistency showed that RI.SI.CO. and WOT agreed 323 times, while the
agreement between RI.SI.CO. and Scamadviser was equal to that between WOT
and Scamadviser and was much lower; i.e., 246. From these ndings, it is clear
that Scamadviser made more unique decisions. On the whole, this experiment
suggests that the predictions made by trust checkers are di erent, but it does
not tell us if they are correct. This question is answered in the next section.
6
      </p>
    </sec>
    <sec id="sec-5">
      <title>Classi cation accuracy on ground-truth dataset</title>
      <p>Using the ground truth dataset, we evaluated the classi cation accuracy of the
four checkers. The results are shown in Table 1. For the systems returning
trustworthiness scores, we show the performance values using both a prede ned
threshold value = 0.5 and the value that maximizes the (a posteriori) global
classi cation accuracy.2 The main ndings of this experiment are that the
reviewbased systems (Trustpilot and WOT) made their assessments on only a subset
of the URLs, in proportions similar to those reported above in the inter-checker
agreement analysis, but they were more precise than the feature-based systems
(Scamadviser and RI.SI.CO.). In particular, Trustpilot (under the extensive
interpretation) and WOT achieved an overall accuracy of, respectively, 100% and
93%, versus 77% of Scamadviser and 87% of RI.SI.CO.. Table 1 also suggests
that evaluation of trustworthiness was, in general, more di cult for fake than for
legitimate eshops. One possible interpretation is that legitimate eshops usually
have only legitimacy features in place, whereas fake websites may or may not
have illegitimacy features.</p>
      <p>Choosing a system-speci c threshold value may have a great impact not
only on Trustpilot, as already mentioned, but also on Scamadviser, whose global
accuracy improved from 77% to 81% with a threshold value equal to 0.75 rather
than 0.5. By contrast, it did not a ect much the performance of WOT. Its score
distribution was such that the classi cation accuracy with a threshold = 0.5 was
nearly the same as that with the optimal threshold value (i.e., 0.52), with as
2 We noticed that the optimal threshold of Trustpilot is set assuming that whenever
the system returns an assessment the website is trustworthy, regardless of the number
of stars. In other words, it seems that Trustpilot assesses the convenience of purchase
or the quality of service for otherwise legitimate ecommerce websites.
many as 49 fake websites receiving a WOT score = 1 (in a range from 1 to 100).
Note also that optimizing the threshold value for global accuracy will lower the
recall on legitimate or fake eshops.</p>
      <p>WOT Trustpilot Scamadviser RI.SI.CO.</p>
      <p>Number of assessments 55% 14% 100% 100%</p>
      <p>ACCURACY OF ACTUAL ASSESSMENTS
Legitimate and fake eshops 93% (94%) 66% (100%) 77% (81%) 87%
Only legitimate eshops 96% (95%) 100% (100%) 92% (86%) 89%
Only fake eshops 92% (93%) 0% (0%) 64% (77%) 85%
7</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusions</title>
      <p>Using an evaluation methodology inspired by information retrieval, we found
that trustworthiness assessments of eshops made by existing checkers are
characterized by moderate inter-consistency and varying classi cation accuracy. This
research suggests that there is much room for performance improvement and that
combination of existing methods holds potential for developing better solutions.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Claudio</given-names>
            <surname>Carpineto</surname>
          </string-name>
          and
          <string-name>
            <given-names>Giovanni</given-names>
            <surname>Romano</surname>
          </string-name>
          .
          <article-title>Learning to detect and measure fake ecommerce websites in search-engine results</article-title>
          . Submitted.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Je rey P. Dotson</surname>
          </string-name>
          , Ruixue Rachel Fan,
          <source>McDonnel Feit Elea</source>
          , Je rey
          <string-name>
            <given-names>D.</given-names>
            <surname>Oldham</surname>
          </string-name>
          , and
          <string-name>
            <surname>Yi-Hsin Yeh</surname>
          </string-name>
          .
          <article-title>Brand attitudes and search engine queries</article-title>
          .
          <source>Journal of Interactive Marketing</source>
          ,
          <volume>37</volume>
          :
          <fpage>105</fpage>
          {
          <fpage>116</fpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Elizabeth</surname>
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Freeman</surname>
            and
            <given-names>Gretchen G.</given-names>
          </string-name>
          <string-name>
            <surname>Moison</surname>
          </string-name>
          .
          <article-title>A comparison of the performance of threshold criteria for binary classi cation in terms of predicted prevalence and kappa</article-title>
          .
          <source>Ecological Modelling</source>
          ,
          <volume>217</volume>
          :
          <fpage>48</fpage>
          {
          <fpage>58</fpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>N.</given-names>
            <surname>Leontiadis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Moore</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N.</given-names>
            <surname>Christin</surname>
          </string-name>
          .
          <article-title>Measuring and analyzing searchredirection attacks in the illicit online prescription drug trade</article-title>
          .
          <source>In Proceedings of USENIX Security</source>
          <year>2011</year>
          , San Francisco, CA, USA,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Mary L. McHugh</surname>
          </string-name>
          .
          <article-title>Interrater reliability</article-title>
          .
          <source>Biochemia Medica</source>
          ,
          <volume>22</volume>
          (
          <issue>3</issue>
          ):
          <volume>276</volume>
          {
          <fpage>282</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6. NetNames.
          <article-title>The risks of the online counterfeit economy</article-title>
          .
          <source>Technical report</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7. OECD/EUIPO.
          <article-title>Trade in Counterfeit and Pirated Goods: Mapping the Economic Impact</article-title>
          .
          <source>OECD Publishing</source>
          , Paris,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>John</given-names>
            <surname>Wadleigh</surname>
          </string-name>
          and Jake Drew andTyler Moore.
          <article-title>The e-commerce market for "lemons": Identi cation and analysis of websites selling counterfeit goods</article-title>
          .
          <source>In In WWW '15</source>
          , pages
          <fpage>1188</fpage>
          {
          <fpage>1197</fpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>D. Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Der</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Karami</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Saul</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>McCoy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Savage</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G. M.</given-names>
            <surname>Voelker</surname>
          </string-name>
          .
          <article-title>Search + seizure: The e ectiveness of interventions on seo campaigns</article-title>
          .
          <source>In In IMC'14</source>
          ,
          <string-name>
            <surname>Ney</surname>
            <given-names>York</given-names>
          </string-name>
          , NY, USA, pages
          <volume>359</volume>
          {
          <fpage>372</fpage>
          . ACM Press,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>