<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Exploring Storage Bottlenecks in Linux-based Embedded Systems</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Russell Joyce</string-name>
          <email>russell.joyce@york.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Neil Audsley</string-name>
          <email>neil.audsley@york.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Real-Time Systems Research Group, Department of Computer Science, University of York</institution>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2015</year>
      </pub-date>
      <abstract>
        <p>With recent advances in non-volatile memory technologies and embedded hardware, large, high-speed persistent-storage devices can now realistically be used in embedded systems. Traditional models of storage systems, including the implementation in the Linux kernel, assume the performance of storage devices to be far slower than CPU and system memory speeds, encouraging extensive caching and bu ering over direct access to storage hardware. In an embedded system, however, processing and memory resources are limited while storage hardware can still operate at full speed, causing this balance to shift, and leading to the observation of performance bottlenecks caused by the operating system rather than the speed of storage devices themselves. In this paper, we present performance and pro ling results from high-speed storage devices attached to a Linux-based embedded system, showing that the kernel's standard le I/O operations are inadequate for such a set-up, and that `direct I/O' may be preferable for certain situations. Examination of the results identi es areas where potential improvements may be made in order to reduce CPU load and increase maximum storage throughput.</p>
      </abstract>
      <kwd-group>
        <kwd>Linux</kwd>
        <kwd>storage</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Categories and Subject Descriptors</title>
      <p>D.4.7 [Operating Systems]: Organization and Design|
Real-time systems and embedded systems</p>
    </sec>
    <sec id="sec-2">
      <title>1. INTRODUCTION</title>
      <p>Traditionally, access to persistent storage has been orders
of magnitude slower than volatile system memory, especially
when performing random data accesses, due to the high
latency and low bandwidth associated with the mechanical
operation of hard disk drives, as well as the constant
increase in CPU and memory speeds over time. Despite the
deceleration of single-core CPU scaling in recent years, the
main bottleneck associated with accessing non-volatile
storage in a general-purpose system is still typically the storage
device itself.</p>
      <p>Linux (along with many other operating systems) uses a
number of methods to reduce the impact that slow storage
devices cause on overall system performance. Firstly, main
memory is heavily used to cache data between block device
accesses, avoiding unnecessary repeated reads of the same
data from disk. This also helps in the e cient operation
of le systems, as structures describing the position of les
on a disk can be cached for fast retrieval. Secondly, bu ers
are provided for data owing to and from persistent storage,
which allow applications to spend less time waiting on disk
operations, as these can be performed asynchronously by the
operating system without the application necessarily waiting
for their completion. Finally, sophisticated scheduling and
data layout algorithms can be used to optimise the data
that is written to a device, taking advantage of idle CPU
time caused by the system waiting for I/O operations to
complete.</p>
      <p>For a general-purpose Linux system, these techniques can
have a large positive e ect on the e cient use of storage {
memory and the CPU often far outperform the speed of a
hard disk drive, so any use of them to reduce disk accesses
is desirable. However, this relationship between CPU,
memory and storage speeds does not hold in all situations, and
therefore these techniques may not always provide a bene t
to the performance of a system.</p>
      <p>The limited resources of a typical embedded system can
skew the balance between storage and CPU speed, which can
cause issues for a number of embedded applications that
require fast and reliable access to storage. Examples of these
include applications that receive streaming data over a
highspeed interface that must be stored in real-time, such as data
being sent from sensors or video feeds, perhaps with
intermediate processing being performed using hardware
accelerators.</p>
      <p>This paper considers e ects that the limited CPU and
memory speeds of an embedded system can have on a fast
storage device { due to the change in balance between
relative speeds, the system cannot be expected to perform in the
same way as a typical computer, with certain performance
bottlenecks shifting away from storage hardware limitations
Device
Buffers
Device
Buffers
Stored
Data
Stored
Data
Application Space</p>
      <p>Kernel Space
Data</p>
      <p>Standard I/O</p>
      <p>CDaDcahaetata
Data</p>
      <p>Direct I/O
and into software operations.</p>
      <p>Results are presented in section 3 from basic testing of
storage devices in an embedded system, showing that
sequential storage operations experience bottlenecks caused by
CPU limitations rather than the speed of the storage
hardware if standard Linux le operations are used. Removing
reliance on the page cache (through direct I/O) is shown to
improve performance for large block sizes, especially on a
fast SSD, due to the reduction in the number of times data
is copied in main memory.</p>
      <p>Potential solutions brie y presented in section 5 suggest
that restructuring the storage stack to favour device accesses
over memory and CPU usage in this type of system, as well
as more radical changes such as the introduction of
hardware accelerators, may reduce the negative e ects of CPU
limitations on storage speeds.</p>
    </sec>
    <sec id="sec-3">
      <title>PROBLEM SUMMARY</title>
      <p>
        Recent advances in ash memory technology have caused
the widespread adoption of Solid-State Drives (SSDs), which
o er far faster storage access compared to mechanical hard
drives, along with other bene ts such as lower energy
consumption and more-uniform access times. It is anticipated
that over the next several years, further advances in
nonvolatile memory technologies will accelerate the increasing
trend in storage device speeds, potentially allowing for large,
non-volatile memory devices that operate with similar
performance to volatile RAM. At a certain point, fast storage
speeds, relative to CPU and system memory speeds, will
cause a critical change in the balance of a system,
requiring a signi cant reconsideration of an operating system's
approach to storage access [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <p>At present, this shift in the balance of system performance
is beginning to a ect the embedded world, where processing
and memory speeds are typically low due to constraints such
as energy usage, size and cost, but where fast solid-state
storage still has the potential to run at the same speed as in
a more powerful system. For example, an embedded system
consisting of a slow, low-core-count CPU and slowly-clocked
memory connected to a high-end, desktop-grade SSD has
a far di erent balance between storage, memory and CPU
than is expected by the operating system design. While
such a system may run Linux perfectly adequately for many
tasks, it will not be able to take advantage of the full speed
of the SSD using traditional methods of storage access, due
to bottlenecks elsewhere in the system.</p>
      <p>Before fast solid-state storage was common, non-volatile
storage in an embedded system would often consist of slow
ash memory, due to the high energy consumption and low
durability of faster mechanical media, meaning the
potential increase in secondary storage speeds provided by SSDs
is even greater in embedded systems than many
generalpurpose systems. An increase in the general storage
requirements and expectations for systems, driven by elds such as
multimedia and `big data' processing, have also accelerated
the adoption of fast solid-state storage in embedded systems.
2.1</p>
    </sec>
    <sec id="sec-4">
      <title>Buffered vs Direct I/O</title>
      <p>
        The Linux storage model relies heavily on the bu ering
and caching of data in system memory, typically requiring
data to be copied multiple times before it reaches its
ultimate destination. The kernel provides the `direct I/O' le
access method to reduce the amount of memory activity
involved in reading and writing data from a block device,
allowing data to be copied directly to and from an
application's memory space without being transferred via the page
cache. While this allows applications more-direct access to
storage devices, it can also create restrictions and have a
severe negative impact on storage speeds if used incorrectly.
In the past, there has been some resistance to the direct
I/O functionality of Linux [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], partly due to the bene ts of
utilising the page cache that are removed with direct I/O,
and the large disparity between CPU/memory and storage
speeds meaning there were rarely any situations where the
overhead of additional memory copies was signi cant enough
to cause a slowdown. However, when storage is fast and the
speed of copying data around memory is slow, using direct
I/O can have a signi cant performance improvement if
certain criteria are met.
      </p>
      <p>Figure 1 shows the basic principles of standard and direct
I/O, with direct I/O bypassing the page cache and removing
the need to copy data from one area of memory to another
between storage devices and applications.</p>
      <p>One of the main issues with direct I/O is the large
overhead caused when dealing with data in small block sizes.
Even when using a fast storage device, reading and writing
small amounts of data is far slower per byte than larger sizes,
due to constant overheads in communication and processing
that do not scale with block size. Without kernel bu ers
in place to help optimise disk accesses, applications that use
small block sizes will su er greatly in storage speed when
using direct I/O, compared to when utilising the kernel's data
caching mechanisms, which will queue requests to more e
ciently access hardware. The performance of accessing large
block sizes on a storage device does not su er from this
issue, however, so applications that either inherently use large
block sizes, or use their own caching mechanisms to emulate
large block accesses, can use direct I/O e ectively where
required.</p>
      <p>
        A further issue with the implementation of direct I/O in
the Linux kernel is that it is not standardised, and is not
part of the POSIX speci cation, so its behaviour and safety
cannot necessarily be guaranteed for all situations. The
formal de nition of the O_DIRECT ag for the open() system
call is simply to \try to minimize cache e ects of the I/O to
and from this le" [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], which may be interpreted di erently
(or not at all) by various le systems and kernel versions.
2.2
      </p>
    </sec>
    <sec id="sec-5">
      <title>Real-time Storage Implications</title>
      <p>Many high-performance storage applications require
consideration of real-time constraints, due to external producers
or consumers of data running independently of the system
storing it { if a storage system cannot save or provide data at
the required speed then critical information may be lost or
the system may malfunction. While solid-state storage
devices have far more consistent access times than mechanical
storage, making them more suitable for time-critical
applications, if the CPU of a system is proving to be a bottleneck
in storage access times, the ability to maintain a consistent
speed of data access relies heavily on CPU utilisation.</p>
      <p>If the CPU can be removed as far as possible from the
operation of copying data to storage, the impact of other
processes on this will be reduced, increasing the
predictability of storage operations and making real-time guarantees
more possible. This could be achieved through methods
such as hardware acceleration, as well as simpli cation of
the software storage stack.
2.3</p>
    </sec>
    <sec id="sec-6">
      <title>Motivation</title>
      <p>A number of examples exist where fast and reliable access
to storage is required by an embedded system, which may be
limited by CPU or memory resources when standard Linux
le system operations are used.</p>
      <p>
        Embedded accelerators are increasingly being investigated
for use in high-performance computing environments, due to
their energy e ciency when compared to traditional server
hardware [
        <xref ref-type="bibr" rid="ref7 ref8">8, 7</xref>
        ]. CPU usage when performing storage
operations has also been identi ed as an issue in server situations,
using large amounts of energy compared to storage devices
themselves, and motivating research into how storage
systems can be made to be more e cient [
        <xref ref-type="bibr" rid="ref4 ref9">4, 9</xref>
        ].
      </p>
      <p>
        Standalone embedded systems that use storage devices
also create motivation for e cient access to storage, for
applications such as logging sensor data and recording
highbandwidth video streams [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Often, external data sources
will have constraints on the speeds required for their
storage, for example, with the number of frames of video that
must be stored each second, so any methods that can help
to meet these requirements while keeping energy usage at a
minimum are desirable.
      </p>
      <p>Consider a basic Linux application that reads a stream
of data from a network interface and writes it to a
continuous le on secondary storage using standard le operations.
Disregarding any other system activity and additional
operations performed by the le system, data will be copied a
minimum of six times on its path from network to disk:
1. From the network device to a bu er in the device driver
2. From the driver's bu er to a general network-layer
kernel bu er
3. From the kernel bu er to the application's memory
space
4. From the application's memory space to a kernel le
bu er
5. From the le bu er to the storage device driver
6. From the driver's bu er to the storage device itself
This process has little impact on overall throughput if
either storage or network speed is slow relative to main
memory and CPU, however as soon as this balance changes, any
additional memory copying can have a severe impact.
Techniques such as DMA can help to reduce the CPU load related
to copying data from one memory location to another,
however this relies on hardware and driver support, and does not
fully tackle the ine ciencies of unnecessary memory copies.</p>
      <p>One advantage of the kernel using its page cache to store a
copy of data is the ability to access that data at a later time
without having to load it from secondary storage, however
this will have no bene t if data is solely being written to or
read from a disk as part of a streaming application, because
by the time the data is needed a second time it is likely that
it has already been purged from the cache.
3.</p>
    </sec>
    <sec id="sec-7">
      <title>EXPERIMENTAL WORK</title>
      <p>In order to examine the e ects that a slow system can
have on the performance of storage devices, and to
identify the potential bottlenecks present in the Linux storage
stack, we performed a number of experiments with storage
operations while collecting pro ling and system performance
information.
3.1</p>
    </sec>
    <sec id="sec-8">
      <title>Experimental Set-up</title>
      <p>The experimental set-up consisted of a an Avnet
ZedBoard Mini-ITX development board connected to storage
devices using its PCI Express Gen2 x4 connector. The
ZedBoard Mini-ITX provides a Xilinx Zynq-7000
system-onchip, which combines a dual-core ARM Cortex-A9 processor
(clocked at 666MHz) with a large amount of FPGA fabric,
alongside 1GiB of DDR3 RAM and many other on-board
peripherals.</p>
      <p>The system uses Linux 3.18 (based on the Xilinx 2015.2
branch) running on the ARM cores, while an AXI-to-PCIe
bridge design is programmed on the FPGA, to provide an
interface between the processor and PCI Express devices.</p>
      <p>To provide a range of results, two storage devices were
tested with the system: a Western Digital Blue 500GB SATA
III hard disk drive, connected though a Startech SATA III
RAID card; and an Intel SSD 750 400GB. While both
devices use the same PCI Express interface for their physical
connection to the board, the RAID card uses AHCI for its
logical storage interface, whereas the SSD uses the more
efcient NVMe interface.</p>
      <p>Due to limitations of the high-speed serial transceiver
hardware on the Zynq SoC, the speed of the SSD interface is
limited to PCI Express Gen2 x4 (from its native Gen3 x4),
reducing the maximum four-lane bandwidth from 3940MB/s
to 2000MB/s. While this is still far faster than the 600MB/s
maximum of the SATA-III interface used by the HDD, it
means the SSD will never achieve its advertised maximum
capable speed of 2200MB/s in this hardware set-up.
3.2</p>
    </sec>
    <sec id="sec-9">
      <title>Data Copy Tests</title>
      <p>To determine an indication of the operating speeds of the
storage devices at various block sizes with minimal external
overhead, we performed basic testing using the Linux dd
utility. For write tests, /dev/zero was used as a source le,
and for read tests, /dev/null was used as a destination.
Both storage devices were freshly formatted with an ext4
le system before each test.</p>
      <p>For each block size and storage device, four tests were
performed: reading from a le on the device, writing to a le
on the device, and reading and writing with direct I/O
enabled (using the iflag=direct and oflag=direct operands
of dd respectively). Additionally, read and write tests were
performed with a 512MiB tmpfs RAM disk (/dev/shm) in
order to determine possible maximum speeds when no
external storage devices or low-level drivers were involved. Each
test was performed with and without the capture of system
resource usage and collection of pro ling data, so results
could be gathered without any additional overheads caused
by these measurements.</p>
      <p>With the secondary storage devices, data was recorded for
20GiB sequential transfers, and with the RAM disk, 256MiB
transfers were used due to the lower available space.
Sequential transfers are used as they represent the type of
problem that is likely to be encountered when requiring
highspeed storage in an embedded system { reading and
writing streams of contiguous data { as well as being simple to
implement and test. Storage devices, especially mechanical
hard disks, generally perform faster with sequential transfers
than with random accesses, and operating and le system
overheads are also likely to be greater for non-sequential
access patterns, so further experimentation will be necessary
to determine whether the same e ects are present when
using di erent I/O patterns.
3.3</p>
    </sec>
    <sec id="sec-10">
      <title>System Resource Usage and Profiling</title>
      <p>
        To collect information about system resource usage during
each test, we used the dstat utility [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] to capture memory
usage, CPU usage and storage device transfer speeds each
second.
      </p>
      <p>
        Additionally, to determine the amount of execution time
that is spent in each relevant function within the user
application, kernel and associated libraries during the tests,
we ran dd within the full-system pro ler, operf, part of
the OPro le suite of tools [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The impact on performance
caused by pro ling is kept to a minimum through support
from the CPU hardware and the kernel performance events
subsystem, however slight overheads are likely while the
proler is running, potentially causing slower speeds and slight
di erences in observed data.
3.4
      </p>
    </sec>
    <sec id="sec-11">
      <title>Results</title>
      <p>The following results were gathered using the system and
methods described above, in order to investigate various
aspects of storage operations.
3.4.1</p>
      <sec id="sec-11-1">
        <title>Read and Write Speeds</title>
        <p>Figure 2 shows the average read and write speeds for a
number of block sizes when transferring data to or from the
storage devices and RAM disk.</p>
        <p>The standard read and write speeds for both storage
devices are very similar, with the SSD only performing slightly
faster than the HDD for all block sizes tested. This suggests
that bottlenecks exist outside of the storage devices in the
test system, either caused by the CPU or system memory
bandwidth, as it is expected that the SSD should perform
signi cantly faster than the HDD in both read and write
speed.</p>
        <p>For the 512B block size, speeds are slower on both
devices, however there is little di erence in speed once block
sizes increase above this. This slow speed could be due to
the signi cant number of extra context switches at a low
block size being a bottleneck, rather than the factors
limiting storage operations at 4KiB and above.</p>
        <p>The consistently slightly higher speeds seen with the SSD
are likely to be caused by it using an NVMe logical interface
to communicate with the operating system, compared to
the less-e cient AHCI interface used by the SATA HDD.
If the storage operations are indeed experiencing a CPU
bottleneck, then the more-e cient low-level drivers of NVMe
would allow for this higher speed. This could be con rmed
by repeating the tests with a SATA SSD connected to the
same RAID card as the HDD, instead of using a separate
NVMe device.</p>
        <p>When both reading and writing using standard I/O, RAM
disk performance is far higher than both non-volatile
storage devices. This was expected even when bottlenecks exist
outside of the storage devices themselves, as the kernel
optimises accesses to tmpfs le systems by avoiding the page
cache, thus requiring fewer memory copy operations.</p>
      </sec>
      <sec id="sec-11-2">
        <title>3.4.2 Impact of Direct I/O</title>
        <p>When performing the same tests with direct I/O enabled,
speeds to the storage devices are generally higher when the
block size is su ciently large to overcome the overheads
involved, such as increased communication with hardware.
512B and 4KiB block sizes are slower than the standard
write tests, as the kernel cannot cache data and write it to
the device in larger blocks, but larger block sizes are faster.</p>
        <p>It appears that the HDD is limited by other factors when
block sizes of 512KiB and above are used, which may be due
to the ine ciencies of AHCI, or simply the speed limitations
of the disk itself. This is reinforced by the HDD direct I/O
read speeds being approximately equal to, or lower than
standard I/O speeds to the device, rather than seeing the
performance increases of the SSD.</p>
        <p>For the SSD, maximum direct write speeds are over
double those of standard I/O, and maximum direct read speeds
also show a signi cant improvement, however these are both
still far lower than the rated speeds of the device. A further
bottleneck appears to be encountered between 1MiB and
16MiB direct I/O block sizes, suggesting that at this point
55
)%50
(
sn 45
o
itc 40
n
uF 35
oyp 30
yC25
r
om20
eM15
ien 10
im 5
T
0
HDD Read
HDD Write
HDD Read Direct I/O</p>
        <p>HDD Write Direct I/O
SSD Read
SSD Write</p>
        <p>SSD Read Direct I/O</p>
        <p>SSD Write Direct I/O
the block size is large enough to overcome any
communication and driver overheads and the earlier limitations
experienced with non-direct I/O are once again a ecting speeds.
This speed limit (at around 230MiB/s) also matches the
write speed limit of the RAM disk when using block sizes
between 512KiB and 16 MiB, suggesting that both the SSD
and RAM disk are experiencing the same bottleneck here.
3.4.3</p>
      </sec>
      <sec id="sec-11-3">
        <title>CPU Usage</title>
        <p>Figure 3 shows the mean single-core kernel CPU usages
across block sizes for each test, where single-core gures are
calculated as the maximum of the two cores for each sample
recorded. In general, it can be seen that a large amount of
CPU time is spent in the kernel across the tests, with all but
HDD direct I/O using an entire CPU core of processing for
large block sizes, strongly suggesting that the bottlenecks
implied by the speed results are caused by inadequate
processing power.</p>
        <p>The low system CPU usage of the HDD direct I/O tests
suggests that the bottleneck may indeed be the disk itself,
unlike the SSD tests, which show more clear, consistent
limits in their transfer speeds.</p>
        <p>For the SSD direct I/O write test, the 16MiB block size
where the speed bottleneck begins corresponds to where
CPU usage reaches 100%, further suggesting that the
bottleneck is caused by processing on the CPU.</p>
        <p>Further experimentation to test the direct impact that
CPU speed has on the storage speeds could be carried out
by repeating the tests while altering the clock speed of the
ARM core, or limiting the number of CPU cores available
to the operating system.
3.4.4</p>
      </sec>
      <sec id="sec-11-4">
        <title>Profiling Results</title>
        <p>Results from pro ling show that for both read and write
tests, a large amount of CPU time is spent copying data
between user and kernel areas of memory. Figure 4 shows the
percentage of total execution time spent in the kernel
functions __copy_to_user (for read) and __copy_from_user (for
write), used for copying data to and from user space
respectively. The direct I/O write tests spend no time in these
functions, but instead a large amount of time is spent
ushing the CPU data cache in the v7_flush_kern_dcache_area
function.</p>
        <p>Both read and direct I/O operations, which rely on more
immediate access to storage devices, additionally spend a
large amount of CPU time waiting for device locks to be
released in the _raw_spin_unlock_irq function.
4.</p>
      </sec>
    </sec>
    <sec id="sec-12">
      <title>RELATED WORK</title>
      <p>There are several areas of related work that suggest the
current position of storage in a system architecture needs
rethinking, due to the introduction of fast storage
technologies, and due to ine ciencies in the software storage stack
and le systems. Their focus is not entirely on embedded
systems, but also on the increasing demand for e cient and
fast storage in high-performance computing environments.</p>
      <sec id="sec-12-1">
        <title>Refactor, reduce, recycle.</title>
        <p>
          The discussion in [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] advocates the necessary simpli
cation of software storage stacks through refactoring and
reduction, in order to make them able to fully utilise
emerging high-speed non-volatile memory technologies, and lower
the relative processor impact caused by fast I/O. It
demonstrates that due to the large increase in storage speeds
available with these devices, the traditional balance between slow
storage and fast CPU speed is broken, and that
improvements can be made through changes to the way storage is
handled by the operating system. The work does not
explicitly reference embedded systems, however the same theories
apply in greater measure, due to even greater restrictions on
processing resources.
        </p>
      </sec>
      <sec id="sec-12-2">
        <title>Hardware file systems.</title>
        <p>
          E orts such as [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] and [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] attempt to improve the
storage performance in an embedded system by o oading
certain le system operations to hardware accelerator cores on
an FPGA. While the hardware le system implementation
in [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] is quite limited and specialised in its operation, it is
motivated by a similar need to optimise storage access
beyond what was capable by the CPU in the target system.
These hardware le system accelerators are aimed more
towards usage in high-performance computing environments
than stand-alone embedded systems.
        </p>
      </sec>
    </sec>
    <sec id="sec-13">
      <title>DISCUSSION AND FURTHER WORK</title>
      <p>The results presented in section 3 show that when CPU
resources are su ciently constrained, there are clear
bottlenecks in storage operations, besides the access times of
storage devices themselves. In order to utilise the full
potential of high-speed storage devices in an embedded Linux
environment, and to avoid their use degrading the operation
of other tasks running in the system, changes must be made
to the storage stack to optimise how they are accessed.
5.1</p>
    </sec>
    <sec id="sec-14">
      <title>Potential Solutions</title>
      <p>There are several potential solutions to the problems
covered, ranging from optimisations in existing software
implementations to more radical system architecture changes.
5.1.1</p>
      <sec id="sec-14-1">
        <title>VFS Optimisations</title>
        <p>It may be possible to reduce storage overheads through
restructuring the storage stack in Linux to better optimise
it for high-speed storage with lower CPU usage. Results
from pro ling may be used to identify the areas of the
storage stack that are performing particularly ine ciently, or
that are simply unnecessary for the required tasks. Such
optimisations would potentially require large changes to the
structure of the Linux kernel.
5.1.2</p>
      </sec>
      <sec id="sec-14-2">
        <title>Improved Direct I/O</title>
        <p>The performance results show that using direct I/O can
give a major boost to performance, especially with the fast
SSD, if block sizes are above a reasonable threshold for disk
access operations, but can also severely reduce performance
if used for small block sizes.</p>
        <p>Given its potential bene ts, a reimplemented pseudo-direct
I/O could operate with the bene ts of direct I/O for large
block sizes, but attempt to e ciently bu er storage
device accesses when block sizes are below a practical limit.
Standardising direct I/O so its operation can be guaranteed
across le systems and kernel versions would also allow its
usage to be more widely accepted.
5.1.3</p>
      </sec>
      <sec id="sec-14-3">
        <title>Hardware Acceleration</title>
        <p>One possible method of relieving CPU load during storage
operations would be to introduce hardware acceleration into
a system, in order to perform some of the tasks associated
with the software storage stack in hardware instead. These
accelerators could range from simple direct memory access
(DMA) units, used to perform the expensive memory copy
operations without taking up CPU time, or more-complex
le-system-aware accelerators that access the storage device
directly, e ectively shifting the hardware/software divide
further up the storage stack.</p>
        <p>Introducing hardware that can access storage
independently of the CPU may give an advantage for applications
that use large streams of data, as more than just the
storage device can be attached to the hardware. For example,
a hardware accelerator could directly receive data from a
hardware video encoder and write it straight to a le on
persistent storage, with little CPU intervention and no bu ering
required in main system memory.
5.2</p>
      </sec>
    </sec>
    <sec id="sec-15">
      <title>Further Work</title>
      <p>There is potential for much deeper investigation into the
operation of fast storage devices in an embedded Linux
environment, in order to fully understand the bottlenecks
involved and propose more comprehensive solutions.</p>
      <p>While the results presented in section 3 highlight some
examples of circumstances where storage speeds are heavily
limited by areas other than storage devices themselves, they
only focus on basic tests working with sequential data on
a single le system type and clean devices. Further
experiments will be carried out with other I/O patterns, such as
random reads/writes, and benchmarks based on real-world
usage patterns, in order to better gauge the scope of the
issue and the focus for improvements.</p>
      <p>Further work will also involve modifying areas of the test
platform, such as the CPU clock speed and the number of
available cores, in order to give insight on the direct a ect
this has on results. Alternative platforms, such as
morepowerful server hardware, can be used to test exactly how
much of a limiting e ect the embedded hardware has on
storage capabilities.</p>
      <p>As well as experimental work on existing implementations,
practical work to test the feasibility of solutions suggested
above will be necessary in order to improve on the
current situation. Modelling storage system operation based
on experimental results may assist in implementation work,
through the identi cation of areas that can be improved and
giving a base on which to test solutions in a more abstract
way.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <article-title>[1] open(2) { Linux Programmer's Manual</article-title>
          .
          <source>Release</source>
          <volume>4</volume>
          .
          <fpage>02</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Dstat</surname>
          </string-name>
          : Versatile resource statistics tool, Mar.
          <year>2012</year>
          . Online: http://dag.wiee.rs/home-made/dstat/.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <article-title>[3] OPro le { A System Pro ler for Linux, Aug</article-title>
          .
          <year>2015</year>
          . Online: http://oprofile.sourceforge.net/.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Caul</surname>
          </string-name>
          eld et al.
          <article-title>Understanding the impact of emerging non-volatile memories on high-performance, IO-intensive computing</article-title>
          .
          <source>In Proc. 2010 ACM/IEEE Int. Conf. High Performance Computing</source>
          , Networking, Storage, and
          <string-name>
            <surname>Analysis</surname>
          </string-name>
          , New Orleans, LA, Nov.
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Mendon</surname>
          </string-name>
          .
          <article-title>The case for a Hardware Filesystem</article-title>
          .
          <source>PhD thesis</source>
          , University of North Carolina at Charlotte, NC,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>National</given-names>
            <surname>Instruments</surname>
          </string-name>
          .
          <article-title>Data acquisition: I/O for embedded systems</article-title>
          . White Paper, Oct.
          <year>2012</year>
          . Available: http://www.ni.com/white-paper/7021/en/.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Putnam</surname>
          </string-name>
          et al.
          <article-title>A recon gurable fabric for accelerating large-scale datacenter services</article-title>
          .
          <source>In Proc. 41st Int. Symp. Computer Architectures</source>
          ,
          <year>June 2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>R.</given-names>
            <surname>Sass</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Kritikos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Schmidt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Beeravolu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Beeraka</surname>
          </string-name>
          .
          <article-title>Recon gurable computing cluster (RCC) project: Investigating the feasibility of FPGA-based petascale computing</article-title>
          .
          <source>In Proc. 15th Annu. IEEE Symp. Field-Programmable Custom Computing Machines, Apr</source>
          .
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>P.</given-names>
            <surname>Sehgal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Tarasov</surname>
          </string-name>
          , and
          <string-name>
            <surname>E. Zadok.</surname>
          </string-name>
          <article-title>Evaluating performance and energy in le system server workloads</article-title>
          .
          <source>In Proc. 8th USENIX Conf. File and Storage Technologies</source>
          , Feb.
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S.</given-names>
            <surname>Swanson</surname>
          </string-name>
          and
          <string-name>
            <surname>A. M.</surname>
          </string-name>
          <article-title>Caul eld. Refactor, reduce, recycle: Restructuring the I/O stack for the future of storage</article-title>
          .
          <source>Computer</source>
          ,
          <volume>46</volume>
          (
          <issue>8</issue>
          ):
          <volume>52</volume>
          {
          <fpage>59</fpage>
          ,
          <string-name>
            <surname>Aug</surname>
          </string-name>
          .
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>L.</given-names>
            <surname>Torvalds</surname>
          </string-name>
          . Re:
          <article-title>O DIRECT question</article-title>
          .
          <source>Linux Kernel Mailing List, Jan</source>
          .
          <year>2007</year>
          . Available: https://lkml.org/lkml/2007/1/11/121.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>V.</given-names>
            <surname>Varadarajan</surname>
          </string-name>
          ,
          <string-name>
            <surname>S. K. R</surname>
          </string-name>
          , A. Nedunchezhian, and
          <string-name>
            <given-names>R.</given-names>
            <surname>Parthasarathi</surname>
          </string-name>
          .
          <article-title>A recon gurable hardware to accelerate directory search</article-title>
          .
          <source>In Proc. IEEE Int. Conf. High Performance Computing, Dec</source>
          .
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>