<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Are you (Google) Home? Detecting Users' Presence through Tra c Analysis of Smart Speakers</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>D. Caputo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>L. Verderame</string-name>
          <email>luca.verderame@unige.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>A. Merlo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>A. Ranieri</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>L. Caviglione</string-name>
          <email>luca.caviglioneg@ge.imati.cnr.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Computer Security Lab Department of Informatics</institution>
          ,
          <addr-line>Bioengineering</addr-line>
          ,
          <institution>Robotics and Systems Engineering University of Genoa</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Institute for Applied Mathematics and Information Technologies National Research Council of Italy</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Smart speakers and voice-based virtual assistants are core building blocks of modern smart homes. For instance, they are used to retrieve information, interact with other devices, and command a variety of Internet of Things (IoT) nodes. To this aim, smart speakers and voice-based assistants typically take advantage of cloud architectures: vocal commands of the user are sampled, sent through the Internet to be processed and transmitted back for local execution, e.g., to perform an automation task or activate an IoT device. Even if privacy and security is enforced by means of encryption, features of the tra c, such as the throughput, the size of protocol data units or the IP addresses, can leak important information about the habits of the users as well as the number and the type of IoT nodes deployed. In this perspective, the paper showcases risks of machine learning techniques to develop black-box models to automatically classify tra c and implement privacy leaking attacks. We prove that such tra c analysis allows to detect the presence of a person in a house equipped with a Google Home device, even if the same person does not interact with the smart device. Experimental results collected in a realistic scenario are presented and possible countermeasures are discussed.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Smart speakers and voice-based virtual assistants are important building blocks of modern smart
homes. For instance, they can ben used to retrieve information, interact with other devices, and
command a wide range of Internet of Things (IoT) nodes. Moreover, they can be used as hubs
for managing IoT deployments or implementing device automation services, e.g., to perform
routines in smart lighting or provide remote connectivity for domestic appliances. According
to [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], there are over 200 million of smart speakers installed in private properties (with the
wide acceptation inside private houses and small o ce settings), and the trend is expected to
culminate in 2030 when the number will exceed 500 million of units. In general, smart speakers
and voice-based virtual assistants take advantage of cloud-based architectures: vocal commands
of the user are sampled and sent through the Internet to be processed. As a result, the smart
speaker or the appliance running the virtual assistant receives a textual representation as well
as optional, companion multimedia data. Then, it executes the command or route it to a proper
hub, e.g., to communicate via ZigBee or Bluetooth links with IoT nodes. To enforce privacy and
security, the prime mechanism is the encryption of tra c (see, e.g., reference [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ] and references
therein). However, features of the ows such as, the throughput, the size of protocol data units
or (address, port) tuples, can leak important information about the habits of the users [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] or
the number and the type of IoT nodes [
        <xref ref-type="bibr" rid="ref22 ref4">4, 22</xref>
        ]. As a consequence, an attacker can collect tra c
from the local IEEE 802.11 wireless loop or between the home gateway and the Internet and
then try to guess the type of IoT nodes and the state of sensors and actuators. With such a
knowledge, the malicious entity can launch a wide array of o ensive campaigns, such as pro le
users, plan attacks to the physical space or perform social engineering campaigns [
        <xref ref-type="bibr" rid="ref22 ref4">4, 22</xref>
        ].
      </p>
      <p>
        Despite the underlying technology or the complexity of the deployment, there is an increasing
interest in investigating risks arising from the statistical analysis of the tra c exchanged by
a smart speaker and the cloud. For instance, in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] authors showcase how passive network
analysis can be used to identify devices and correlate some user activities, e.g., tra c ows
produced by switches and health monitors can leak the sleep cycle of a user. In [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], the
tra c produced by state transitions of home devices (i.e., a thermostat and a carbon dioxide
detector) can be used to infer if a user is present in the home. Such idea is further re ned in
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], where passive measurements are used to develop models of the daily routine of individuals
(e.g., leaving/arriving home). Concerning works aiming at identifying devices, possibly by
adopting machine learning or statistical tools, in [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] several machine learning techniques are
used to identify IoT devices by exploiting \poor" information like the length of packets produced
during normal operations. Additionally, in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] the risks of HTTP-based communications are
discussed, both from the perspective of inferring data about the devices (e.g., the state or the
intensity of a light source) and performing session-highjacking attacks.
      </p>
      <p>
        In addition, sensitive data contained in IoT nodes and smart speakers can be relevant for
forensics investigations [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], and tra c patterns can be manipulated by malware to ex ltrate
data, for instance through information hiding schemes [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] or covert channels [
        <xref ref-type="bibr" rid="ref10 ref19">10, 19</xref>
        ]. In this
vein, the paper discusses risks of machine learning techniques to develop black-box models for
automatically classifying tra c and to implement privacy leaking attacks. Di erently from
previous works [
        <xref ref-type="bibr" rid="ref1 ref12 ref22 ref3">1, 3, 12, 22</xref>
        ], we focus on understanding whether it is possible to recognize the
presence of a user when no queries are performed. In fact, when a request is sent towards the
Internet, the produced tra c volumes or the appearance of speci c network addresses trivially
leak the presence of a human operator in the house. To this aim, we empirically prove how
it is possible to detect the presence of a person in a house by analysing the tra c produced
by a Google Home device under the assumption that the person is not interacting with it.
Nevertheless, attention will be devoted in proposing ideas to mitigate such kind of threats by
acting at the tra c level. In fact, the design of suitable mitigation techniques is often neglected
(see, e.g., [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] for a notable exception) or addressed at an API-permission level [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], which is
de nitely out of the scope of the paper.
      </p>
      <p>Summing up, the contributions of this work are: i ) to review the architectural blueprint
used by smart speakers and voice-based virtual assistants and elaborate an e ective model to
conduct privacy leaking attacks, and ii ) to evaluate the e ectiveness of using machine learning
techniques for black-box modelling of tra c. We also sketch some design rules to mitigate
identi cation attacks in Section 5.</p>
      <p>The remainder of the paper is structured as follows. Section 2 discusses the general
architecture used by smart speakers to control IoT devices, introduces the threat model and the
machine learning mechanisms that can be exploited by the attacker. Section 3 deals with the
testbed used to collect data, while Section 4 presents numerical results and Section 5 concludes
the paper and showcases some possible future directions.</p>
    </sec>
    <sec id="sec-2">
      <title>Smart Speakers: Architecture and Threat Model</title>
      <p>As hinted, smart speakers and voice-based virtual assistants are a core foundation for smart
homes. In essence, they provide a user interface to issue requests or commands in a natural
manner, i.e., by simply talking. Such devices can be also used as hubs for other IoT nodes and
network appliances or to perform tasks like playing music and video, purchasing items, and to
make recommendations. Besides, smart speakers and virtual assistants can provide a variety of
information including directions and weather forecasts.</p>
      <p>As today, the most popular smart speakers implementing the aforementioned features are
Google Home1, HomePod2 and Amazon Echo3, whereases virtual assistants are Amazon Alexa4,
Apple Siri5 and Google Assistant6. Literature still lacks of a uni ed terminology for this class of
devices and services. In fact, smart speakers and virtual assistants are identi ed as intelligent
personal assistants, virtual personal assistants, home digital voice assistants, voice-enabled
speakers, smart speakers, and voice-based virtual assistants, just to mention the most popular
names. Therefore, in the following, we only use the terms smart speakers or Intelligent Virtual
Assistant (IVA) interchangeably, except when doubt may arise.</p>
      <p>Even if each smart speaker is characterized by speci c design choices and some setups are
implemented via a complex interplay of technologies and services, the core architectural blueprint
is quite standard and depicted in Figure 2. The overall set of components is often de ned as the
ecosystem as to emphasize the end-to-end pipeline at the basis of such services, i.e., hardware or
software entities allowing the interaction of end users, computing and communication services,
and software running in IoT nodes. Even if each vendor usually implements its own blueprint,
the typical one is composed of four major components:</p>
      <p>
        Smart Speaker or IVA: it is in charge of collecting vocal commands, sample them and
transmit the data trough the Internet to a backend. Upon receiving a response, the smart
speaker or the software IVA agent can provide a feedback to the user or directly interact
with other devices. For instance, the smart speaker could start the playback of a music
stream received from a Content Delivery Network (CDN) or send through a ZigBee link
a command to a smart lightbulb. In some cases, it can also act as a sort of \router",
thus delivering commands to the suitable hub. To avoid security and privacy threats,
communications are encrypted via the Secure Socket Layer (SSL) [
        <xref ref-type="bibr" rid="ref15 ref8">8, 15</xref>
        ].
      </p>
      <p>Client and IoT Devices: they are the targets of commands of the ecosystem. Typical
nodes deployed in a smart home are sensors, actuators, Bluetooth/ZigBee bridges, wireless
speakers or IoT-capable appliances. As previously said, some entities belonging to this
class can be colocated within the smart speaker.</p>
      <p>
        IVA Cloud: it is the backend in charge of processing data and delivering back
text/binary representations of commands to be executed, including additional contents like
multimedia streams, geographical information or JSON les containing a composite
variety of information. With the advent of open ecosystems promoting the interaction among
services provided by multiple vendors, the borders of the IVA cloud are blurring [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ],
[
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. For instance, vocal stimuli could be processed in a datacenter and sent back to
1https://store.google.com/product/google_home
2https://www.apple.com/homepod/
3https://www.amazon.com/echodot
4https://developer.amazon.com/alexa
5https://developer.apple.com/siri/
6https://assistant.google.com/
the IVA while contents can be delivered by a third-part CDN and some IoT nodes could
establish a direct point-to-point connection with the computing infrastructure of their
manufacturer.
      </p>
      <p>
        Network: it connects the smart speaker or the IVA with IoT nodes as well as the Internet.
Typical deployments use a single local (wireless) network connected via a router/gateway
to the Internet. However, in most complex scenarios, di erent networks could be present,
e.g., a local access cabled network for some IoT nodes and hubs and multiple wireless
loops to connect smart devices and grant access to the user via a smartphone. Concerning
protocols used to exchange data between the IVA and the cloud, the TCP is the main
choice, with the multipath variant to optimize performances and reduce delays [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. A
notable exception is the Google ecosystem. In fact, it exploits QUIC [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], a protocol
originally engineered to improve performance issues of HTTP/2 and based on transport
streams multiplexed over UDP. We point out that the presence of QUIC can represent a
signature to ease the identi cation of the ecosystem (e.g., Apple HomeKit vs. Google).
However, this requires to understand its behaviors, which can be highly in uenced by the
underlying network conditions (see, e.g., [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] for a sensitivity/performance analysis of the
SPDY counterpart in di erent wireless settings).
      </p>
      <p>
        Concerning the typical usage scenario, smart speakers rely upon a microphone to sense
commands, which are processed by a vocal interpreter running locally. In fact, only
wakeup commands are executed within the device, while others are transmitted remotely to the
cloud. Each IVA is activated via its own phrase or keyword and the most popular are \Ok
Google", \Alexa", and \Hey Siri", for the case of Google Assistant, Amazon Echo/Alexa and
Apple/HomeKit ecosystem, respectively. As it will be detailed later on, a relevant fragility is
due to the continuous data exchange from the IVA and the cloud. Even if several frameworks
could be considered \secure" both from the architectural and technological viewpoints, still
they are prone to a variety of privacy-breaking attacks targeting a composite set of features
observable within the encrypted tra c ows [
        <xref ref-type="bibr" rid="ref2 ref27 ref4 ref8">2, 4, 8, 27</xref>
        ].
2.1
      </p>
      <sec id="sec-2-1">
        <title>Threat Model</title>
        <p>
          We aim at investigating the class of attacks targeting the encrypted tra c in a black-box
manner, i.e., without trying to decipher the payload of protocol data units. Literature abounds
of works dealing with techniques against SSL ows, for instance, [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] provides an extensive
survey on Man-in-the-Middle (MitM) attacks for SSL/TLS conversations as well as techniques
to highjack or spoof di erent protocol entities and nodes (e.g., BGP routes, ARP/RARP caches,
and access points). Moreover, [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] reports an MitM attack expressly crafted for the Alexa IVA.
Speci cally, it targets \skills", which are extensions introduced to integrate third-part devices
and services in the Amazon ecosystem. An attacker can redirect the voice input of the victim
to a malicious node, thus highjacking the conversation. However, such attacks are de nitely
outside of the scope of this paper. Rather, we consider an adversary wanting to pro le the
user, for instance, for reconnaissance purposes or to plan a physical attack. To this aim, the
adversary can exploit the tra c to infer \behavioral" information, e.g., when the victim is not
at home. Figure 3 depicts the reference threat model.
        </p>
        <p>
          In more detail, we assume an adversary (denoted as malicious user in the gure) that can
only perform a passive attack, i.e., he/she can observe and acquire the tra c produced by the
victim but cannot alter or manipulate it. To this aim, the adversary should access the home
router. However, this is not a tight constraint as he/she can abuse the IEEE 802.11 wireless
loop to gather information to be sent to the IVA (see, e.g., [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] for an analysis of threats that
can be done by moving throughout the attack surface). We also assume that the adversary is
not able to use the contents of the packets to launch the attack: in other words, he/she is not
able to attack the TLS/SSL or VPN schemes usually deployed. Therefore, by inspecting the
tra c produced by the smart speaker, the adversary can only rely on statistics and metadata
of conversations. As an example, the attacker inspects (or computes by performing suitable
operations) values like the throughput, the size of protocol data units, the IP address, the
number of di erent endpoints, ags within the headers of the packets, or the behavior of the
congestion control of the TCP. Finally, as it's usually done in similar works, we assume that the
attacker is able to isolate and recognize tra c that comes from di erent IoT devices [
          <xref ref-type="bibr" rid="ref20 ref24 ref5">5, 20, 24</xref>
          ].
        </p>
        <p>
          Even if the deployment of encryption schemes is not su cient to prevent the leakage of
important information about the habits of users [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] and the number or the type of IoT nodes
deployed [
          <xref ref-type="bibr" rid="ref1 ref3 ref4">1, 3, 4</xref>
          ], this was a suitable countermeasure to mitigate a wide variety of threats.
Alas, the advent of computational-e cient statistical tools brings into a feasibility zone a new
wave of attacks. As a prototypal example, the work in [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ] demonstrated how to leak the
language of the talker by inspecting the bit rate of VoIP conversations. In essence, authors
used a sort of \signatures" produced by the variable bit rate codec to feed various classi ers,
such as the k -Nearest Neighbors, Hidden Markov Models, and Gaussian Mixture Models and
a computational-e cient variant of the 2 classi ers, to identify the language with di erent
performances (e.g., they can discriminate between English and Hungarian and from Brazilian
Portuguese and English with a 66:5% and 86% of accuracy, respectively). We then review
the most suitable tools that the adversary can use to extract information obtained from the
gathered tra c and then unhinge the privacy of smart speakers and part of the IoT subsystem.
2.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Machine Learning Techniques for Attacking the IoT Ecosystem</title>
        <p>
          Nowadays, gathering and analyzing tra c is a core technique used during the reconnaissance
phase of an attack [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ], e.g., to enumerate devices or to ngerprint hosts for searching known
vulnerabilities. In this work we consider the attacker wanting to infer high-level information, for
instance to launch social engineering campaigns or plan physical attacks. Literature showcases
di erent machine learning approaches and their adoption to solve networking duties is becoming
a de-facto standard (see, [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ] for a recent survey on the use of deep learning for di erent tra c
classi cation problems). However, in the perspective of endowing an attacker with the suitable
tools to gather information on the state of the smart speaker or the IVA, we shortlisted the
following most promising algorithms.
        </p>
        <p>
          Decision Tree (DT) is a family of non-parametric supervised learning methods suitable
for classi cation and regression problems [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. The DT builds classi cation or regression
models in the form of a tree structure. To this aim, it breaks down the data into smaller
subsets while developing an associated decision tree. The process is iterated by further
splitting the dataset and the nal result is a tree with decision nodes and leaf nodes.
Adaptive Boosting - AdaBoost (AB) exploits the the idea of creating a highly accurate
prediction rule by combining many relatively weak and inaccurate rules [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. AdaBoost
can be used in conjunction with many other types of learning algorithms to improve
performance. In this case, the output of the other learning algorithms (de ned as weak
learners) is combined into a weighted sum that represents the nal output of the boosted
classi er.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Experimental Testbed</title>
      <p>To prove the e ectiveness of privacy threats of smart speakers leveraging machine learning
techniques, we developed an experimental testbed. Due to the lack of public datasets
containing network tra c of smart speakers, we have also developed an automated framework for
generating and collecting the relevant network tra c.</p>
      <p>Concerning the device under investigation, we used a Google Home Mini7 since it is one of
the most popular smart appliances. Our version is equipped with an IEEE 802.11 L2 interface,
an internal microphone to sense commands and the surrounding environment, and a loudspeaker
for audio playback and LEDs for visual feedbacks. The con guration of the device must be
done via a companion application8. To this aim, we provided the SSID and the password of
our test network, which allowed the smart speaker to communicate remotely with the cloud
running Google services and to exchange data with other devices connect to the same network
(e.g., smart tv, smart light bulbs, etc.). We did not performed other tweaks as to reproduce an
average installation usually accounting for the device deployed by the user in an out-of-the-box
avor.</p>
      <p>Since we are focusing on privacy leakages related to the behavior of the microphone when
disabled or when sensing various situations, i.e, the presence of humans or a quiet condition,
we performed three di erent measurement campaigns, each one lasting 3 days. In particular,
for the rst round of tests, the microphone of the smart speaker was manually set o as to
investigate the tra c exchanged between the device and the remote cloud datacenter. Then,
for the second round, the microphone was manually set on and the device put in a quiet
condition, i.e., the microphone did not receive any stimuli from the surrounding environment,
which was completely without noise or voices. For the last round of tests, we set the microphone
on and we simulated the presence of humans speaking each others or background noise. We
underline that human talkers will not issue the \Ok Google" phrase or will not inadvertently
activate the smart speaker. In the following, we denote the di erent tests as mic off for the
case when the microphone is disabled, mic on and mic on noise for tests with the microphone
active and the smart speaker placed in a silent or noisy environment, respectively. To the aim
of having proper audio patterns, we selected videos from YouTube in order to stimulate the
smart speaker with a wide variety of talkers and settings (e.g., female and male speakers of
di erent ages).</p>
      <p>To capture data, we prepared a standard computer to act as the IEEE 802.11 access point
and we deployed ad-hoc scripts for running tshark9, i.e., the command line interface provided
by the Wireshark tool. To process the dataset and perform computations, we used a computer
with an Intel Core i7-3770 processor, with 16 GB of RAM running the Ubuntu 16.04 LTS
operating system.</p>
      <p>To implement the machine learning algorithms presented in Section 2.2, we used the
scikitlearn10 library. In essence, it is an open-source library developed in Python that contains the
implementation of the most popular machine learning algorithms.
3.1</p>
      <sec id="sec-3-1">
        <title>Data Handling</title>
        <p>As said, we only collected tra c without performing any operation aimed at breaking the
encryption scheme. In other words, we consider a worst-case scenario where the attacker is
7https://store.google.com/it/product/google_home_mini
8https://play.google.com/store/apps/details?id=com.google.android.apps.chromecast.app
9https://www.wireshark.org/docs/man-pages/tshark.html
10http://scikit-learn.org/
not able to perform deep packet inspection or more sophisticated actions (e.g., pinning of SSL
certi cates). Instead, the threat model we investigate deals with a malicious entity wanting
to infer the smart speaker state by only using statistical information observable within the
encrypted network tra c. To this aim, the attacker can extract/compute indicators by using
two di erent \grouping" schemes, as depicted in Figure 1. In more detail, we computed the
desired metrics by considering a suitable amount of packets obtained according to the windowing
mechanisms considered as follows:
t
time spans of length</p>
        <p>t (see Figure 1a);
bursts of a xed length of N (see Figure 1b).</p>
        <p>Δt =  t2-t1
t1
t2 t
1
2</p>
        <p>N - 1</p>
        <p>N
(a) Packets grouped in a window of t seconds
(b) Packets grouped in a window of N data units</p>
        <p>We point out that the size of the windows a ect the amount of information to be processed by
the machine learning algorithm. In fact, even if the dataset sill remains unchanged, the number
of windows is directly proportional to the volume of information o ered to the statistical tool
(i.e., for each window a statistical indicator is computed). Concerning the statistical indicators
that an attacker can obtain from the tra c exchanged between the IVA and the cloud, we
consider:</p>
        <p>Number of TCP, UDP and ICMP packets: allow to quantify the composition of the
tra c in terms of observed protocols. For instance, UDP datagrams indicate the presence
of signaling carried by the QUIC protocol, whereas TCP segments can represent the
exchange of additional data such as multimedia material.</p>
        <p>Number of di erent IP addresses and TCP/UDP ports: the presence of di erent endpoints
could be used to spot interaction between the smart speaker and the IVA cloud, including
actions requiring to contact third-part entities or providers, IoT nodes, private datacenters
or CDN facilities.
per -window Inter packet time (IPT) or packet count: allow to consider how tra c
distributes within the two windows used to group packets described in Figure 1.
Aggressiveness of the source could be used to reveal user activity or stimuli triggered by a vocal
input.</p>
        <p>Average value and standard deviation of the TCP window: describe the behavior of
the ow in terms of burstiness and bandwidth usage. Such information could lead to
indications about how the IVA and its cloud exchange data.</p>
        <p>Average value and standard deviation of the IPT: similarly to the previous case, they
can be used to complete information inferred from the packet rate. For instance, the</p>
        <p>IPT could be used to recognize whether a ow is generated by an application with some
real-time constraints.</p>
        <p>Average value and standard deviation of the packet length: hint at the type of the
application layer, for instance, small packets can suggest the presence of voice-based activities
requiring a low (bounded) packetization delay.</p>
        <p>Average value and standard deviation of the TTL: can be used to mark ow belonging to
di erent portions of the network and possibly indicating that the smart speaker has been
activated for a task also requiring the interaction with additional providers or actuators
(e.g., IoT nodes).</p>
        <p>
          We point out that many indicators are intrinsically \privacy leaking" as they allow a
malicious observer to infer some information about the smart home hosting the device [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. For
instance, counting di erent conversations and the number of protocol data units in a timeframe
could reveal the presence of speci c IoT nodes or the type of the requested operation, e.g.,
retrieving a summary of the news. At the same time, considering such values could impact on
the performance of the classi cation framework owing to the exploitation of interactions among
the di erent architectural components, which are di cult to forecast.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Preliminar Results</title>
      <p>In this section, we showcase numerical results obtained in our trials. First, we provide an
overview of the collected dataset, then we present the performances of machine learning
algorithms used to leak privacy of users with particular attention on the time needed for the
training phase.
4.1</p>
      <sec id="sec-4-1">
        <title>Dataset Overview</title>
        <p>As presented in Section 3, the dataset has been generated in a 9 day long measurement
campaign composed of three trials of 3 days with di erent conditions of the microphone of the
smart speaker. Speci cally, for the mic off case, we collected 203; 596 packets for a total size
of 69 Mbytes. Instead, when the microphone is active, we collected 216; 456 packets in the
mic on scenario and 282; 656 packets mic on noise one, for a total size of 74 and 173 Mbytes,
respectively. The overall dataset has been processed with the StandardScaler, thus leading to
a statistical population with average equal to 0 and standard deviation equal to 1.</p>
        <p>Figure 4 depicts the average values characterizing the dataset in each scenario. It is worth
noting that the average packet length and the average size of the TCP window for the mic off
and mic on cases are very similar. Instead, for the mic on noise case, the average packet length
doubles, whereas the average TCP window size halves.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Classifying the State of the Smart Speaker</title>
        <p>We now show the results obtained when trying to classify the state of the smart speaker to
conduct a privacy leaking attack.</p>
        <p>The rst experiment aimed at investigating whether it is possible to identify if the
microphone of the smart speaker or the device hosting the IVA is turned on or o . We point out that
this can be also viewed as a sort of side-channel, where the attacker can identify if users are
in the proximity of the device. In this perspective, Figure 5 shows the accuracy of the
classiers adopted to infer from the tra c whether the microphone is ON or OFF, i.e., discriminate
among mic on or mic off cases. To better understand the performances, we also investigated
when the di erent \grouping" strategies presented in Section 2.2 are used to feed the machine
learning algorithms.</p>
        <p>As shown, best results are achieved by using the AdaBoost algorithm (denoted as AB in the
gure). However, it is important to note that, for identifying the state of the microphone with
an acceptable level of accuracy, the attacker has to collect about 500 s of tra c or 500 packets.
Therefore, a real-time classi cation could be not possible in the sense that the attacker has
to wait a non-negligible amount of time before he/she has the knowledge to launch the attack
(e.g., force the physical perimeter where the smart speaker is deployed).</p>
        <p>
          The second experiment aimed at discriminating between the two di erent behaviors of the
surrounding environment, i.e., the mic on and mic on noise states. We recall that such states
can be used by the attacker to infer if the smart speaker operates in a silent environment or in
the presence of noise, e.g., people are talking to each other or the television is turned on. In
both cases, there is not a direct interaction, that is, in the case of Google Home, any user did
not issue the \Ok Google" phrase. Then, the malicious user cannot exploit \macro" features of
the tra c, such as the number of TCP connections, the IP range or the tra c volume [
          <xref ref-type="bibr" rid="ref22 ref4">4, 22</xref>
          ].
        </p>
        <p>Figure 6 depicts the obtained results. Compared to the previous experiment, to reach a
good level of accuracy, it is su cient to use a reduced amount of packets. As an example, for
the case of the Decision Tree, good degrees of accuracy to decide whether the smart speaker is
in the mic on or mic on noise states are achieved by using time-windows with t = 15 seconds
or a burst of N = 20 packets. From the perspective of understanding the security and privacy
of voice-based appliances, this result reveals a potential exploitable hazard. In fact, when the
user does not directly interact with the smart speaker (e.g., the \Ok Google" phrase is not
issued), the tra c generated towards the remote cloud should be the same for both the mic on
and mic on noise conditions. In other words, it is expected that the network tra c does not
exhibit any signature. Even if we did not have access to the internals of the Google Home
Mini used in our testbed, the di erent tra c behaviors could be due to the fact that the smart
speaker is always in an \awake" mode and selected stimuli are sent to the cloud as to identify
activation phrases like \Ok Google" or \Hey Siri". However, this could partially contradict the
believing that such phrases are completely handled locally by the smart speaker or the IVA.</p>
        <p>To assess the performances of the di erent classi ers in a comprehensive manner, Figure
7 shows the confusion matrices of the AdaBoost and Decision Tree classi ers when used to
discriminate between the mic on - mic on noise cases. It is possible to notice how the confusion
matrices show the goodness of the chosen algorithms having the highest values distributed on
the diagonal. Similar considerations can be done for the other techniques but they have been
omitted here for the sake of brevity.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions and Future Works</title>
      <p>
        In this paper, we investigated the feasibility of adopting machine learning techniques to breach
the privacy of users interacting with smart speakers or voice assistants. Di erent from other
works discovering the presence of the user via intrinsically privacy-leaking activities (e.g., the
activation of a IoT node and the related tra c ow), we concentrated on discriminating how
the internal microphone is used. Results indicate the e ectiveness of our approach, thus making
the management of silence and noise epoque as major privacy concerns. To increase the user's
privacy a possible countermeasure could be the insertion of suitable padding inside the packets
to normalize the average length, as well as the standard deviation moreover, using a unique
protocol for the transport, could add another layer of privacy. Another possible countermeasure
could be an of appropriate noise, for instance by exploiting some form of tra c camou age or
morphing [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Therefore, suitable tra c morphing or protocol manipulation techniques should
be put in place within the device or, at least, in-home routers as to reduce the attack surface
that can be exploited by malicious entities.
      </p>
      <p>Future work will aim at re ning our framework by considering smart speakers from other
vendors. Besides, we are working towards the implementation of a sort of \warden" able to
normalize tra c generated towards the IVA cloud.
A.1</p>
    </sec>
    <sec id="sec-6">
      <title>Appendix</title>
      <sec id="sec-6-1">
        <title>General smart home scenario and threat model</title>
        <p>Intelligent Virtual Assistant (IVA)
Cloud
NETWORK
NETWORK
 User voice-command  </p>
        <p>Response
Smart Speaker Devices
User
Client and IoT Devices</p>
        <p>A.2</p>
      </sec>
      <sec id="sec-6-2">
        <title>Dataset Overview and Classi ers Results</title>
        <p>(a) Mean Packet Length
(b) Mean TCP window
(c) Mean IPT
(d) Mean TTL</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Abbas</given-names>
            <surname>Acar</surname>
          </string-name>
          , Hossein Fereidooni, Tigist Abera, Amit Kumar Sikder, Markus Miettinen, Hidayet Aksu, Mauro Conti,
          <string-name>
            <surname>Ahmad-Reza Sadeghi</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A. Selcuk</given-names>
            <surname>Uluagac</surname>
          </string-name>
          .
          <article-title>Peek-a-</article-title>
          <string-name>
            <surname>Boo: I See Your Smart Home Activities</surname>
          </string-name>
          , Even Encrypted!
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Efthimios</given-names>
            <surname>Alepis</surname>
          </string-name>
          and
          <string-name>
            <given-names>Constantinos</given-names>
            <surname>Patsakis</surname>
          </string-name>
          . Monkey Says,
          <source>Monkey Does: Security and Privacy on Voice Assistants. IEEE Access</source>
          ,
          <volume>5</volume>
          :
          <fpage>17841</fpage>
          {
          <fpage>17851</fpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Yousef</given-names>
            <surname>Amar</surname>
          </string-name>
          , Hamed Haddadi, Richard Mortier, Anthony Brown, James Colley, and
          <string-name>
            <given-names>Andy</given-names>
            <surname>Crabtree</surname>
          </string-name>
          .
          <article-title>An Analysis of Home IoT Network Tra c</article-title>
          and Behaviour.
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Noah</given-names>
            <surname>Apthorpe</surname>
          </string-name>
          , Dillon Reisman, and
          <string-name>
            <given-names>Nick</given-names>
            <surname>Feamster</surname>
          </string-name>
          .
          <article-title>A Smart Home is No Castle: Privacy Vulnerabilities of Encrypted IoT Tra c</article-title>
          .
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Lei</given-names>
            <surname>Bai</surname>
          </string-name>
          , Lina Yao, Salil S Kanhere,
          <string-name>
            <given-names>Xianzhi</given-names>
            <surname>Wang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Zheng</given-names>
            <surname>Yang</surname>
          </string-name>
          .
          <article-title>Automatic Device Classi cation From Network Tra c Streams of Internet of Things</article-title>
          .
          <source>In 43rd Conference on Local Computer Networks</source>
          , pages
          <fpage>1</fpage>
          <article-title>{9</article-title>
          . IEEE,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Prasenjeet</given-names>
            <surname>Biswal</surname>
          </string-name>
          and
          <string-name>
            <given-names>Omprakash</given-names>
            <surname>Gnawali</surname>
          </string-name>
          .
          <source>Does QUIC Make the Web Faster? In 2016 IEEE Global Communications Conference</source>
          , pages
          <fpage>1</fpage>
          <article-title>{6</article-title>
          . IEEE,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Andrea</given-names>
            <surname>Cardaci</surname>
          </string-name>
          , Luca Caviglione, Alberto Gotta, and
          <string-name>
            <given-names>Nicola</given-names>
            <surname>Tonellotto</surname>
          </string-name>
          .
          <article-title>Performance Evaluation of SPDY Over High Latency Satellite Channels</article-title>
          .
          <source>In International Conference on Personal Satellite Services</source>
          , pages
          <volume>123</volume>
          {
          <fpage>134</fpage>
          . Springer,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Luca</given-names>
            <surname>Caviglione</surname>
          </string-name>
          .
          <article-title>A First Look at Tra c Patterns of Siri</article-title>
          .
          <source>Transactions on Emerging Telecommunications Technologies</source>
          ,
          <volume>26</volume>
          (April):
          <volume>664</volume>
          {
          <fpage>669</fpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Luca</given-names>
            <surname>Caviglione</surname>
          </string-name>
          , Mauro Coccoli, and
          <string-name>
            <given-names>Alessio</given-names>
            <surname>Merlo</surname>
          </string-name>
          .
          <article-title>A Taxonomy-based Model of Security and Privacy in Online Social Networks</article-title>
          .
          <source>International Journal of Computer Sciences and Engineering</source>
          ,
          <volume>9</volume>
          (
          <issue>4</issue>
          ):
          <volume>325</volume>
          {
          <fpage>338</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Luca</surname>
            <given-names>Caviglione</given-names>
          </string-name>
          , Maciej Podolski, Wojciech Mazurczyk, and
          <string-name>
            <given-names>Massimo</given-names>
            <surname>Ianigro</surname>
          </string-name>
          .
          <source>Covert Channels in Personal Cloud Storage Services: The case of Dropbox. IEEE Transactions on Industrial Informatics</source>
          ,
          <volume>13</volume>
          (
          <issue>4</issue>
          ):
          <year>1921</year>
          {
          <year>1931</year>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Mauro</surname>
            <given-names>Conti</given-names>
          </string-name>
          , Nicola Dragoni, and
          <string-name>
            <given-names>Viktor</given-names>
            <surname>Lesyk</surname>
          </string-name>
          .
          <article-title>A Survey of Man in the Middle Attacks</article-title>
          .
          <source>IEEE Communications Surveys &amp; Tutorials</source>
          ,
          <volume>18</volume>
          (
          <issue>3</issue>
          ):
          <year>2027</year>
          {
          <year>2051</year>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Bogdan</surname>
            <given-names>Copos</given-names>
          </string-name>
          , Karl Levitt, Matt Bishop, and
          <string-name>
            <given-names>Je</given-names>
            <surname>Rowe</surname>
          </string-name>
          . Is Anybody Home?
          <article-title>Inferring Activity from Smart Home Network Tra c</article-title>
          .
          <source>In IEEE Security and Privacy Workshops</source>
          , pages
          <volume>245</volume>
          {
          <fpage>251</fpage>
          . IEEE,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Wenrui</surname>
            <given-names>Diao</given-names>
          </string-name>
          , Xiangyu Liu,
          <string-name>
            <surname>Zhe Zhou</surname>
          </string-name>
          , and Kehuan Zhang. Your Voice Assistant is Mine:
          <article-title>How to Abuse Speakers to Steal Information and Control Your Phone</article-title>
          .
          <source>In Proceedings of the 4th ACM Workshop on Security and Privacy in Smartphones &amp; Mobile Devices</source>
          , pages
          <volume>63</volume>
          {
          <fpage>74</fpage>
          . ACM,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Kevin</surname>
            <given-names>P Dyer</given-names>
          </string-name>
          , Scott E Coull, Thomas Ristenpart, and Thomas Shrimpton.
          <article-title>Peek-a-Bboo, I Still See You: Why E cient Tra c Analysis Countermeasures Fail</article-title>
          .
          <source>In IEEE Symposium on Security and Privacy</source>
          , pages
          <volume>332</volume>
          {
          <fpage>346</fpage>
          . IEEE,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>Marcia</given-names>
            <surname>Ford</surname>
          </string-name>
          and
          <string-name>
            <given-names>William</given-names>
            <surname>Palmer</surname>
          </string-name>
          . Alexa, Are You Listening to Me?
          <article-title>An Analysis of Alexa Voice Service Network Tra c</article-title>
          .
          <source>Personal and Ubiquitous Computing</source>
          ,
          <volume>23</volume>
          (
          <issue>1</issue>
          ):
          <volume>67</volume>
          {
          <fpage>79</fpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>T.</given-names>
            <surname>Hastie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Tibshirani</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Friedman</surname>
          </string-name>
          .
          <source>The Elements of Statistical Learning: Data Mining, Inference, and Prediction</source>
          . Springer New York,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>B.</given-names>
            <surname>Kinsella</surname>
          </string-name>
          . Smart Speaker Prevision. https://voicebot.ai/
          <year>2019</year>
          /04/15/smart-speaker
          <string-name>
            <surname>-</surname>
          </string-name>
          installedbase-to-surpass-200
          <string-name>
            <surname>-</surname>
          </string-name>
          million-in-2019
          <string-name>
            <surname>-</surname>
          </string-name>
          grow-to-500
          <string-name>
            <surname>-</surname>
          </string-name>
          million-in-2023
          <source>-canalys. Last Accessed: Sept</source>
          .
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>Shancang</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Shancang</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <surname>Kim-Kwang Raymond</surname>
            <given-names>Choo</given-names>
          </string-name>
          , Qindong Sun,
          <string-name>
            <given-names>William J</given-names>
            .
            <surname>Buchanan</surname>
          </string-name>
          , and Jiuxin Cao. IoT Forensics:
          <article-title>Amazon Echo as a Use Case</article-title>
          .
          <source>IEEE Internet of Things Journal</source>
          ,
          <volume>14</volume>
          (
          <issue>8</issue>
          ):1{
          <issue>1</issue>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>W.</given-names>
            <surname>Mazurczyk</surname>
          </string-name>
          and
          <string-name>
            <given-names>L.</given-names>
            <surname>Caviglione</surname>
          </string-name>
          .
          <article-title>Information Hiding as a Challenge for Malware Detection</article-title>
          .
          <source>IEEE Security Privacy</source>
          ,
          <volume>13</volume>
          (
          <issue>2</issue>
          ):
          <volume>89</volume>
          {
          <fpage>93</fpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Yair</surname>
            <given-names>Meidan</given-names>
          </string-name>
          , Michael Bohadana, Asaf Shabtai, Juan David Guarnizo,
          <article-title>Mart n Ochoa, Nils Ole Tippenhauer, and Yuval Elovici. Pro lIoT: A Machine Learning Approach for IoT Device Identi - cation Based on Network Tra c Analysis</article-title>
          .
          <source>In Proceedings of the Symposium on Applied Computing</source>
          , pages
          <volume>506</volume>
          {
          <fpage>509</fpage>
          . ACM,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Richard</surname>
            <given-names>Mitev</given-names>
          </string-name>
          , Markus Miettinen, and
          <string-name>
            <surname>Ahmad-Reza Sadeghi</surname>
          </string-name>
          . Alexa Lied to Me:
          <article-title>Skill-based Manin-the-Middle Attacks on Virtual Assistants</article-title>
          .
          <source>In Proceedings of the 2019 ACM Asia Conference on Computer and Communications Security</source>
          , pages
          <volume>465</volume>
          {
          <fpage>478</fpage>
          . ACM,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Anto^nio J Pinheiro</surname>
          </string-name>
          ,
          <string-name>
            <surname>Jeandro de M Bezerra</surname>
          </string-name>
          , Caio AP Burgardt, and
          <string-name>
            <surname>Divanilson R Campelo. Identifying IoT</surname>
          </string-name>
          <article-title>Devices and Events Based on Packet Length from Encrypted Tra c</article-title>
          .
          <source>Computer Communications</source>
          ,
          <volume>144</volume>
          :8{
          <fpage>17</fpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>S.</given-names>
            <surname>Rezaei</surname>
          </string-name>
          and
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          .
          <article-title>Deep Learning for Encrypted Tra c Classi cation: An Overview</article-title>
          .
          <source>IEEE Communications Magazine</source>
          ,
          <volume>57</volume>
          (
          <issue>5</issue>
          ):
          <volume>76</volume>
          {
          <fpage>81</fpage>
          , May
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Musta zur R Shahid</surname>
          </string-name>
          , Gregory Blanc, Zonghua Zhang, and Herve Debar.
          <article-title>IoT Devices Recognition Through Network Tra c Analysis</article-title>
          .
          <source>In EEE International Conference on Big Data</source>
          , pages
          <volume>5187</volume>
          {
          <fpage>5192</fpage>
          . IEEE,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Siraj</surname>
            <given-names>A Shaikh</given-names>
          </string-name>
          , Howard Chivers, Philip Nobles, John A Clark, and
          <string-name>
            <given-names>Hao</given-names>
            <surname>Chen</surname>
          </string-name>
          .
          <source>Network Reconnaissance. Network Security</source>
          ,
          <year>2008</year>
          (
          <volume>11</volume>
          ):
          <volume>12</volume>
          {
          <fpage>16</fpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Charles</surname>
            <given-names>V Wright</given-names>
          </string-name>
          , Lucas Ballard, Fabian Monrose, and
          <string-name>
            <surname>Gerald M Masson. Language</surname>
          </string-name>
          <article-title>Identi cation of Encrypted VoIP Tra c: Alejandra y Roberto or Alice and Bob? In USENIX Security Symposium</article-title>
          , volume
          <volume>3</volume>
          , pages
          <fpage>43</fpage>
          {
          <fpage>54</fpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <surname>Yuchen</surname>
            <given-names>Yang</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Longfei Wu</surname>
            , Guisheng Yin,
            <given-names>Lijie</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>and Hongbin</given-names>
          </string-name>
          <string-name>
            <surname>Zhao</surname>
          </string-name>
          .
          <article-title>A Survey on Security and Privacy Issues in Internet-of-Things</article-title>
          .
          <source>IEEE Internet of Things Journal</source>
          ,
          <volume>4</volume>
          (
          <issue>5</issue>
          ):
          <volume>1250</volume>
          {
          <fpage>1258</fpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>