<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>A. N. Shaikh, A. M. Shabut, and M. A.
Hossain. A literature review on phish-
ing crime, prevention review and inves-
tigation of gaps. In</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Use of Interpersonal Deception Theory in Counter Social Engineering</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Grace Hui Yang</string-name>
          <email>huiyang@cs.georgetown.edu</email>
          <email>huiyang@cs.georgetown.edu Georgetown University</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yue Yu</string-name>
          <email>yy476@georgetown.edu Georgetown University</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>The InfoSense Group, Department of Computer Science</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2016</year>
      </pub-date>
      <volume>10</volume>
      <abstract>
        <p>Social engineering attacks exploit human vulnerabilities rather than computer vulnerabilities. Ranging from straightforward spam emails to sophisticated context-aware social engineering, social engineering has demonstrated rich varieties. Surprisingly, even the simplest type of attacks are able to fool numerous innocent people. The more sophisticated ones are even more \successful" in achieving their malicious purposes. In order to mitigate and combat these attacks, we need better automated counter social engineering algorithms and tools. In this position paper, we propose a reinforcement learning framework that incorporates interpersonal deception theory to ght against social engineering attacks on social media sites.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Social engineering attacks have become increasing
threats in online life. These attacks are made
possible by the rapid increase in usage of social networks,
mobile devices, and web platforms. Social
engineering attacks are maliciously planned by online
criminals. Social engineering attacks, sometimes also called
phishing attacks, aim to collect sensitive and
personal information about users, including identi cation
numbers, passwords, and bank account information
[AZ17, SSH16] via emails, chatrooms, dating sites and
other forms of social media.</p>
      <p>Copyright © CIKM 2018 for the individual papers by the papers'
authors. Copyright © CIKM 2018 for the volume as a collection
by its editors. This volume and its papers are published under
the Creative Commons License Attribution 4.0 International (CC
BY 4.0).</p>
      <p>Early social engineering attacks include only
cosmetic deceptions where the appearance of a graphical
user interface (GUI) are imitated. For instance, part
of a social media platform user interface may contain
a component that could be used for social
engineering. The victims trust that these components on the
GUI are used as they intended. However, the
malicious attacks exploit this trust by showing the known
components but embedding malicious codes in them
to obtain users' sensitive information.</p>
      <p>Recent attacks have become more sophisticated.
They focus more on mimicking a legitimate system's
behavior or a legitimate user's behavior rather than
imitating GUI components. Online social networks,
such as Facebook, Twitter, LinkedIn, as well as dating
sites, are used to chat with the victims rst, gain their
trusts, and then obtain their sensitive information via
these interpersonal communications. There are also
hybrid attacks that combine both aspects of \looking
like" and \behaving like" known values to make the
deception even more convincing.</p>
      <p>Social engineering is within the scope of
interpersonal deception [BB96]. It is a type of interpersonal
communication where the messages knowingly
transmitted by a sender to foster a false belief or conclusion
by the receiver and to obtain sensitive or personal
information of the receiver. It belongs to the type of
strategic behaviors with a clear goal-oriented nature.
The deception happens when communicators control
the information contained in their messages to
convey a fake meaning. In this research, we propose to
employ general theories and understandings developed
in Interpersonal Deception Theory (IDT) [BB96]
for modeling and combating social engineering from a
broader scope and a deeper level.</p>
      <p>Detecting social engineering attacks can be
complex. Many counter social engineering methods share
a few common stages. The rst stage is planning or
orchestration. The following questions are asked { \how
is the victim or the target chosen?" \How does the
attack reach the target?" and \Is the attack
automated?". The second stage is about \Is it behavior
that deceives the target?" and \Does the deception
occur in the system or external?". The third stage
cares about \Does the deception take one step or
multiple steps?" \Is it persistent?" [HL15]. Counter social
engineering is a type of task that requires good
understanding of the cognitive and behavioral patterns of
the criminals and that of the victims.</p>
      <p>Interpersonal communications are governed by a set
of cognitive and behavioral theories. As a subtype of
interpersonal deception, social engineering makes no
exception. In this position paper, we propose to make
use of known theories on interpersonal deception and
model them into a multi-agent reinforcement learning
framework. We aim to bridge the understandings in
psychology with modern machine learning algorithms
and tools. We make use of the in uence during
interactions, pre-interactions and post-interactions among
the victims, the social engineering attackers and our
counter social engineering agents.</p>
      <p>The proposed reinforcement learning framework is
exible. We can design the states and actions to re ect
the factors that have studied and proved to be useful
in IDT. It would be important to model completeness,
directness, knowledge, clarity of the messages as well
as to model the personality, vulnerability, arousal,
negative a ect, cognitive e ort, suspicion, and attempted
control of the message senders. Inspired by IDT, we
also explore the impact of context (e.g. personal and
contextual data about individuals) and relationships
(e.g. friend network) to social engineering. We not
only model them but also implement tools to collect
context and make use of social network information
to both detect social engineering attacks and generate
counter social engineering messages and activities to
investigate the attackers.</p>
      <p>In this position paper, we discuss the possibilities
for modeling theories about interpersonal deception to
create new counter social engineering strategies for
underlying arti cial intelligence (AI) and machine
learning (ML) algorithms. The research is proposed to be
built on top of a multi-agent reinforcement learning
framework with the capabilities to model and use
interpersonal deception theories for combating the
attacks.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>Previous research has investigated the psychological
side of the problem { why phishing works [DTH06].
The largest factors have to do with people's
misconceptions about computer security and attack templates.
There is a general lack of knowledge about computer
systems, computer security, and security indicators (or
the absence of security indicators). Many social
media users have only basic or even incorrect assumptions
and heuristics when deciding how to respond to emails
or chat messages asking for sensitive information. For
instance, some assume that once a business already has
their personal information, it is safe to give it again.</p>
      <p>As a result, educating social media users about
making the right decisions when receiving social
engineering attacks is a very important component in
preventing such attacks. Unfortunately, most
existing approaches only focus on awareness training and
hope users make better decisions the next time they
are faced with a questionable email or chat message.
The e ectiveness of this strategy is limited. Instead,
our research focuses on performing automated active
detection and intervention for potential attacks,
relieving users of the pressure of self-protection.</p>
      <p>The automated approaches include sandboxing,
authorisation-authentication-accounting (AAA),
monitoring via Honeypots, integrity checking, machine
learning [HL15]. Those countermeasures aim to
prevent and detect attacks before and after the victim
data are collected and used [AZ17].</p>
      <p>Existing machine learning techniques used for
counter social engieering have focused on classi cation
algorithms. Both linear and non-linear supervised
machine learning models are used to make the decision
on whether an email or a website is social engineering.
Algorithms such as Support Vector Machines (SVM),
Nave Bayes, and k-Nearest Neighbor are widely used.
The challenging here is to identify the good features
that are able to distinguish normal emails and text
messages from the engineering ones.</p>
      <p>We summarize a list of popular features and cues
used in the machine learning appraoches:</p>
      <p>URL features such as IR address characteristics,
geographic properties, domain names. [AZ17]
Content-based features which examine how
suspicious the content is, e.g. asking for money, asking
for a bank account, asking for a password. [AZ17]
Document structure features, including a web
page's main page, component les, DOM
structures etc.</p>
      <p>Linguistic cues for deceptions, such as the length
of unique words, the length of sentences, word
diversity, type-token ratio, six-letter words, the
number of verbs being used, tentative words,
modal verbs. [HBGMS15]
Complexity of language use, such as
exclusive words, causation relations, certainty,
negations, negative emotions including anger, sadness
words and pleasant and unpleasantness words.
[HBGMS15]</p>
      <p>Unsupervised approaches such as clustering,
mining and statistical language models are also used in
social engineering detection. For instance, simple
techniques such as term frequency and inverse document
frequency (TF-IDF), regular expressions representing
social engineering text patterns, as well as latent
semantic analysis (LSA) [AZ17], are still used in
social engineering attack detection. Zhou and Shi have
shown in [ZSZ08] that with two n-gram statistical
language models (SLM), one using deception data and
the other using legitimate data, together with the
Kneser-Ney smoothing technique, the language
modeling approach can outperform SVM. In this proposed
research, we fully explore the features and cues,
including n-grams, and take advantage of already proposed
features in the literature.</p>
      <p>There are also patents invented for detecting and
ghting against social engineering attacks [KS16,
SKB+16]. Some patents propose developing decoy
systems to trap the attackers. The decoy systems
contain hardware components as well as decoy
documents and other digital information. They have a
more realist understanding of how a deception system
works. For instance, a deception system could
generate receipts, tax documents, and other form-based
documents with credentials, names, emails, addresses
or login information collected online or within an
organization [SKB+16]. These patents are quite
complex in terms of their designs. Even though no e
ectiveness metrics are reported, these approaches sound
quite practical. However, within each component of
these patented system, the detailed features do not
seem as e ective as what has been studied in the
research community. For example, the linguistic features
mentioned in those patents are quite naive. Keywords
such as "top secret" and "privileged" are hard-coded
into the decoy system and no advanced machine
learning techniques nor more exible methods are used to
make the patented system scalable to large scales.</p>
      <p>Spear phishing is the form of social engineering that
deceives the victims by creating emails, text messages
or chats with context relevant to the victim. The
relevant content is collected from the Internet. An
adversary can digitally \stalk" a victim (a Web user)
and discover as much information as possible about
the victim, either through direct observation of posted
information or by inferring knowledge using simple
inference logic. Such knowledge includes a person's race,
relationship status, estimated income level, and
religion [SYS+15a, SYS+15b]. Current technologies for
counter spear phishing are still in its infancy. Most
existing techniques overlap largely with the privacy
community and data mining community in terms of
understanding how the contextual information is crawled
and collected for an individual from publicly available
online data.</p>
      <p>In the Arti cial Intelligence community, research in
dialogue-based systems and adversarial search are
relevant to counter social engineering in terms of their
common goal of being interactive and adaptive.
However, there is no prior study on counter social
enineering in this context. The work by Banerjee and Peng
[BP03] is perhaps the most similar to what we
propose here. They proposed a multi-agent reinforcement
learning framework for countering deception.
However, their work is in the domain of gaming and
adversarial search. Moreover, they do not show how
to incorporate existing strategies into a reinforcement
learning algorithm, which is our focus.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Multi-Agent ing</title>
    </sec>
    <sec id="sec-4">
      <title>Reinforcement</title>
    </sec>
    <sec id="sec-5">
      <title>Learn</title>
      <p>This research develops a general reinforcement
learning framework for modeling two teams of agents, the
social engineering attackers, and counter social
engineering agents, into one interactive, dynamic
environment. We propose elements, framework,
placeholders for interpersonal deception theories, and ways to
model human-programmable policies for modeling and
combating social engineering attacks.</p>
      <p>Multi-agent learning (MAL) lies at the intersection
of distributed arti cial intelligence and reinforcement
learning. A multi-agent system (MAS) contains
multiple agents as its name suggests. Agents in a MAS
typically operate in large, complex, dynamic and
unpredictable environments which is a key di erence
between MAL and typical supervised machine learning.
MAL is also an area where game theory meets with
reinforcement learning. Game theory has been
extensively studied for adversarial search in arti cial
intelligence as well as for concurrent reinforcement learning.
Most algorithms for multi-agent reinforcement
learning have been proposed mostly in the space of
stationary environment. That is, one agent is explicitly
formulated based on stationary policies for self-play
and the other agents are following stationary policies
and assuming explicit knowledge of the agent and the
domain. Nonetheless, there are quite few MAL
algorithms available.</p>
      <p>In the problem of counter social engineering, there
are multiple agents. There are one or more
attackers. There are also one or more victims. They are
the agents in the game. Here we also assume a clear
division of attackers and victims as two sides of the
agents. Di erent from the dynamic search problem
that we have just mentioned, counter-deception is
uncooperative, which makes it closer to the traditional
AI problem of adversarial search where both teams of
players would like to win the other team. When we
model the counter social engineering problem, its
uncooperative nature needs to be taken into account and
to be properly represented in the model.</p>
      <p>In the multi-agent stochastic game (SG), we
propose it to be a tuple &lt; S; Aph; Ac; f; Rph; Rc &gt;, where
S is the discrete set of states, Aph is the set of actions
that the phishers take, Ac is the set of actions that
the counter social engineering agents take. Both
actions Aph and Ac yield a joint action set A = Aph Ac.
f is the state transition probabilistic function and is
de ned over S A S ! [0; 1]. The phisher reward
function Rph is de ned over S Aph S ! R and
counter social engineering agent reward function Rc is
de ned over S Ac S ! R.</p>
      <p>States S is a discrete set of states.</p>
      <p>Actions A is a discrete set of actions that an agent
can take. For instance, the criminal's actions include
searching for a name and collecting context for an
individual.</p>
      <p>Observations is a discrete set of observations
that an agent makes about the states. O is the
observation function which represents a probabilistic
distribution for making observation o given action a and
landing in the next state s0.</p>
      <p>Transitions T is the state transition function
T (si; a; sj ) = P r(si; a; sj ) ranging from 0 to 1. It is the
probability of starting in state si, taking action a, and
ending in state sj . The sum over all actions give the
total state transition probability T (si; sj ) = P r(si; sj ).</p>
      <p>Reward r = R(s; a) is the immediate reward, also
known as reinforcement. It gives the expected
immediate reward of taking action a at state s. An agent in
an MDP usually maximizes its own long-term reward.</p>
      <p>Long term reward is the sum of all past and
future rewards in the entire process: P1
t=1 r. It can be
optionally discounted for the future states: Pt1=1 trt,
where is the discount factor.</p>
      <p>A policy describes the behaviors of an agent. A
non-stationary policy is a sequence of mapping from
states to actions. It is also the solution that we usually
seek in a Markov Decision Process. A policy makes a
decision that which action should be taken for a state.
We optimize to decide how to move around in the
state space in order to optimize the long-term reward
Pt1=1 r in the entire process. The policy studies :
S ! A, such that optimizes the long-term reward
that is represented in a value function V .</p>
      <p>The goal for counter social engineering can
be de ned as the following: select or suggest the
most suitable strategy c for counter social engineering
agent c to best ful ll the long-term expected rewards.
Given that di erent counter social engineering
strategies demonstrate signi cantly di erent algorithms and
result representations, this research concentrates on
how to de ning the elements of the multiple-agent
reinforcement learning framework: its states, actions,
rewards, etc.
4</p>
    </sec>
    <sec id="sec-6">
      <title>Mathematical Modeling</title>
      <p>The research proposed is a new attempt for
studying interactions in the interpersonal deception
process. Here we assume a game between the two groups
of agents, the social engineering attackers, and the
counter social engineering agents. When the game
turns into the case that the agents have di erent goals,
it is very interesting for us to see how to detect the
social engineering attacks and perform e ective counter
social engineering activities. In this research, we focus
on how to represent theories in interpersonal
deception into policies that a machine can understand and
execute.</p>
      <p>A multi-agent reinforcement learning algorithm is
modeled as a stochastic game with a set of joint actions
A = A1 A2 A3 :::An, where each set of actions Ai
is the possible actions of agent i. The goal of the ith
agent is assumed to nd a strategy or a policy i which
maximizes the agent's expected sum of discounted long
term rewards for state s, i.e. the value function,
1
v i (s) = X iE(rtij i;
i=0
i; s0 = s)
where rti is the reward for the ith agent at time t, s0
is the initial joint state, i is the strategy of the ith
agent's opponent, and is the discount factor.</p>
      <p>Here we propose a general framework for policy
presentation, where both the social engineering
strategy and the counter social engineering strategy can be
present in the same framework. For instance, we could
model the two sides of strategies as a bimatrix game,
in which a part of matrices, M1 and M2. The entries in
the matrices Mk(a1; a2) are used to represent the
payo of the kth agent for the joint actions (a1; a2). Here
the two matrices are of size jA1j jA2j each. Here 1
means social engineering agents, and 2 means counter
social engineering agents. If our game is a zero-sum
game, then the matrices can be written as</p>
      <p>M1 =</p>
      <p>If we consider a simple two agent game here at a
single stage, we further assume that both learners are
naive Q-learners that maintain a Q-table for the values
of their possible actions with the updating function
where is the learning rate ranging from 0 to 1 and
rt is the reward at time t. The policy that the agent
takes would output the action at as the maximizing
action at = arg maxb Qt(b), where b is also an action.</p>
      <p>Note that the above is only for two agents. For
agents more than two, there would be other
possibilities and other solutions. We here study and investigate
new presentations for the even complex settings,
especially how agents form two teams to combat with each
other.</p>
      <p>The transitions T and the rewards R are the main
interests of a designed policy. For instance, one
strategy for detecting social engineering attacks is to see if
an attacker sends similar social engineering messages
to multiple victims, probably in the same
organization. This strategy begins with putting any sender
of emails into the picture. Then it is expected for the
email sender to send multiple email messages (with the
\multiple" action enabled), possibly with form letter
writing skills; and those messages are received by
multiple receivers. The received messages are validated
by their content similarity. If the similarity passes a
threshold, the messages are further validated to see if
majority of those message are \asking" for sensitive
information such as money. If this action is further
con rmed, then we decide it is probably a social
engineering attack.
5</p>
    </sec>
    <sec id="sec-7">
      <title>Conclusion</title>
      <p>Social engineering is a common type of malicious
attack on social media. It is a complex process. Its
complexity comes from the involvement of many factors
ranging from cognitive and behavioral theories to the
latest web technologies, and to the new digital lifestyle
of everyone. Social engineering attacks have escalated
in complexity. Recent types of attacks, such as the
context-aware attacks, are more di cult to detect than
social engineering attacks from a few years ago.</p>
      <p>In this position paper, we discuss a new counter
social engineering method that oversees many factors
during social engineering attacks, coordinates
strategy optimization to pro-act appropriately at various
stages in detecting and combating the attacks. In
particular, we propose a new framework that incorporate
interpersonal deception theories into multi-agent
reinforcement learning to combat social engineering
attacks. Our approach helps bridge the gap between the
human understanding of interpersonal deceptions and
machine discovered knowledge, improving reaction and
response to novel attack types and enabling the use of
known interpersonal deception strategies and wisdom.</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgements</title>
      <p>This research is supported by National Science
Foundation grant IIS-145374. Any opinions, ndings,
conclusions, or recommendations expressed in this paper
are of the authors, and do not necessarily re ect those
of the sponsor.
[AZ17]
[BB96]
[BP03]
[DTH06]</p>
      <sec id="sec-8-1">
        <title>Ahmed Aleroud and Lina Zhou. Phishing environments, techniques, and countermeasures: A survey. 68, 04 2017.</title>
      </sec>
      <sec id="sec-8-2">
        <title>David B Buller and Judee K Burgoon. Interpersonal deception theory. Communication theory, 6(3):203{242, 1996.</title>
      </sec>
      <sec id="sec-8-3">
        <title>Bikramjit Banerjee and Jing Peng. Coun</title>
        <p>tering deception in multiagent
reinforcement learning. In Proceedings of the
Workshop on Trust, Privacy, Deception
and Fraud in Agent Societies at
AAMAS03, Melbourne, Australia, pages 1{5,
2003.</p>
      </sec>
      <sec id="sec-8-4">
        <title>Rachna Dhamija, J Doug Tygar, and Marti Hearst. Why phishing works. In</title>
        <p>Proceedings of the SIGCHI conference on
Human Factors in computing systems,
pages 581{590. ACM, 2006.
[HBGMS15] Valerie Hauch, Iris Blandon-Gitlin,
Jaume Masip, and Siegfried L Sporer.</p>
        <p>Are computers e ective lie detectors? a
meta-analysis of linguistic cues to
deception. Personality and Social Psychology</p>
        <p>Review, 19(4):307{342, 2015.
[HL15]
[KS16]
[SKB+16]</p>
      </sec>
      <sec id="sec-8-5">
        <title>Ryan Heart eld and George Loukas. A</title>
        <p>taxonomy of attacks and a survey of
defence mechanisms for semantic social
engineering attacks. ACM Comput. Surv.,
48(3):37:1{37:39, December 2015.</p>
      </sec>
      <sec id="sec-8-6">
        <title>A.D. Keromytis and S.J. Stolfo. Sys</title>
        <p>tems, methods, and media for
generating bait information for trap-based
defenses, September 22 2016. US Patent
App. 15/155,790.</p>
      </sec>
      <sec id="sec-8-7">
        <title>S.J. Stolfo, A.D. Keromytis, B.M.</title>
        <p>Bowen, S. Hershkop, V.P. Kemerlis, P.V.
Prabhu, and M.B. Salem. Methods,
systems, and media for baiting inside
attackers, November 22 2016. US Patent
9,501,639.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [SYS+15a]
          <string-name>
            <given-names>Lisa</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <surname>Grace Hui Yang</surname>
          </string-name>
          , Micah Sherr, Andrew Hian-Cheong, Kevin Tian, Janet Zhu, and Sicong Zhang.
          <article-title>Public information exposure detection: Helping users understand their web footprints</article-title>
          .
          <source>2015 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM)</source>
          ,
          <volume>00</volume>
          :
          <fpage>153</fpage>
          {
          <fpage>161</fpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [SYS+15b]
          <string-name>
            <given-names>Lisa</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <surname>Hui Yang</surname>
          </string-name>
          , Micah Sherr, Yifang Wei, Andrew Hian-Cheong, Kevin Tian, Janet Zhu, Sicong Zhang, Tavish Vaidya, and
          <string-name>
            <given-names>Elchin</given-names>
            <surname>Asgarli</surname>
          </string-name>
          .
          <article-title>Helping users understand their web footprints</article-title>
          .
          <source>In Proceedings of the 24th International Conference on World Wide Web, WWW '15 Companion</source>
          , pages
          <volume>117</volume>
          {
          <fpage>118</fpage>
          , New York, NY, USA,
          <year>2015</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [ZSZ08]
          <string-name>
            <given-names>Lina</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Yongmei</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Dongsong</given-names>
            <surname>Zhang</surname>
          </string-name>
          .
          <article-title>A statistical language modeling approach to online deception detection</article-title>
          .
          <source>IEEE Transactions on Knowledge and Data Engineering</source>
          ,
          <volume>20</volume>
          (
          <issue>8</issue>
          ):
          <volume>1077</volume>
          {
          <fpage>1081</fpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>