<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Software Design Approaches for Mastering Variability in Database Systems</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>David Broneske</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sebastian Dorok</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Veit Köppen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andreas Meister</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>author names are in lexicographical order</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>General Terms</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Database, Software Engineering</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Otto-von-Guericke-University Magdeburg Institute for Technical and Business Information Systems Magdeburg</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2014</year>
      </pub-date>
      <abstract>
        <p>For decades, database vendors have developed traditional database systems for di erent application domains with highly di ering requirements. These systems are extended with additional functionalities to make them applicable for yet another data-driven domain. The database community observed that these \one size ts all" systems provide poor performance for special domains; systems that are tailored for a single domain usually perform better, have smaller memory footprint, and less energy consumption. These advantages do not only originate from di erent requirements, but also from di erences within individual domains, such as using a certain storage device. However, implementing specialized systems means to reimplement large parts of a database system again and again, which is neither feasible for many customers nor e cient in terms of costs and time. To overcome these limitations, we envision applying techniques known from software product lines to database systems in order to provide tailor-made and robust database systems for nearly every application scenario with reasonable e ort in cost and time.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        In recent years, data management has become increasingly
important in a variety of application domains, such as
automotive engineering, life sciences, and web analytics. Every
application domain has its unique, di erent functional and
non-functional requirements leading to a great diversity of
database systems (DBSs). For example, automotive data
management requires DBSs with small storage and memory
consumption to deploy them on embedded devices. In
contrast, big-data applications, such as in life sciences, require
large-scale DBSs, which exploit newest hardware trends,
e.g., vectorization and SSD storage, to e ciently process
and manage petabytes of data [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Exploiting variability to
design a tailor-made DBS for applications while making the
variability manageable, that is keeping maintenance e ort,
time, and cost reasonable, is what we call mastering
variability in DBSs.
      </p>
      <p>
        Currently, DBSs are designed either as one-size- ts-all
DBSs, meaning that all possible use cases or functionalities
are integrated at implementation time into a single DBS,
or as specialized solutions. The rst approach does not
scale down, for instance, to embedded devices. The second
approach leads to situations, where for each new
application scenario data management is reinvented to overcome
resource restrictions, new requirements, and rapidly
changing hardware. This usually leads to an increased time to
market, high development cost, as well as high maintenance
cost. Moreover, both approaches provide limited capabilities
for managing variability in DBSs. For that reason, software
product line (SPL) techniques could be applied to the data
management domain. In SPLs, variants are concrete
programs that satisfy the requirements of a speci c application
domain [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. With this, we are able to provide tailor-made
and robust DBSs for various use cases. Initial results in the
context of embedded systems, expose bene ts of applying
SPLs to DBSs [
        <xref ref-type="bibr" rid="ref17 ref22">17, 22</xref>
        ].
      </p>
      <p>The remainder of this paper is structured as follows: In
Section 2, we describe variability in a database system
regarding hardware and software. We review three approaches
to design DBSs in Section 3, namely, the one-size- ts-all, the
specialization, and the SPL approach. Moreover, we
compare these approaches w.r.t. robustness and maturity of
provided DBSs, the e ort of managing variability, and the level
of tailoring for speci c application domains. Because of the
superiority of the SPL approach, we argue to apply this
approach to the implementation process of a DBS. Hence, we
provide research questions in Section 4 that have to be
answered to realize the vision of mastering variability in DBSs
using SPL techniques.
2.</p>
      <p>VARIABILITY IN DATABASE SYSTEMS
Variability in a DBS can be found in software as well as
hardware. Hardware variability is given due to the use of
di erent devices with speci c properties for data processing
and storage. Variability in software is re ected by di
erent functionalities that have to be provided by the DBS
for a speci c application. Additionally, the combination of
front-side bus</p>
      <p>APU
CPU</p>
      <p>MainMemory
memory
bus</p>
      <p>I/O
controller
HDD</p>
      <p>SSD
hardware and software functionality for concrete application
domains increases variability.
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>Hardware</title>
      <p>
        In the past decade, the research community exploited
arising hardware features by tailor-made algorithms to achieve
optimized performance. These algorithms e ectively utilize,
e.g., caches [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] or vector registers of Central Processing
Units (CPUs) using AVX- [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ] and SSE-instructions [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ].
Furthermore, the usage of co-processors for accelerating data
processing opens up another dimension [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. In the
following, we consider processing and storage devices and sketch
the variability arising from their di erent properties.
2.1.1
      </p>
      <sec id="sec-2-1">
        <title>Processing Devices</title>
        <p>
          To sketch the heterogeneity of current systems, possible
(co-)processors are summarized in Figure 1. Current
systems do not only include a CPU or an Accelerated
Processing Unit (APU), but also co-processors, such as Many
Integrated Cores (MICs), Graphical Processing Units (GPUs),
and Field Programmable Gate Arrays (FPGAs). In the
following, we give a short description of varying processor
properties. A more extensive overview is presented in our recent
work [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ].
Central Processing Unit: Nowadays, CPUs consist of
several independent cores, enabling parallel execution of
different calculations. CPUs use pipelining, Single Instruction
Multiple Data (SIMD) capabilities, and branch prediction
to e ciently process conditional statements (e.g., if
statements). Hence, CPUs are well suited for control intensive
algorithms.
        </p>
        <p>Graphical Processing Unit: Providing larger SIMD
registers and a higher number of cores than CPUs, GPUs o er
a higher degree of parallelism compared to CPUs. In
order to perform calculations, data has to be transferred from
main memory to GPU memory. GPUs o er an own memory
hierarchy with di erent memory types.</p>
        <p>Accelerated Processing Unit: APUs are introduced to
combine the advantages of CPUs and GPUs by including
both on one chip. Since the APU can directly access main
memory, the transfer bottleneck of dedicated GPUs is
eliminated. However, due to space limitations, fairly less GPU
cores t on the APU die compared to a dedicated GPU,
leading to reduced computational power compared to dedicated
GPUs.</p>
        <p>Many Integrated Core: MICs use several integrated and
interconnected CPU cores. With this, MICs o er a high
parallelism while still featuring CPU properties. However,
similar to the GPU, MICs su er from the transfer
bottleneck.</p>
        <p>GPU</p>
        <p>MIC
PCIe bus</p>
        <p>FPGA</p>
        <p>Field Programmable Gate Array: FPGAs are
programmable stream processors, providing only a limited
storage capacity. They consist of several independent logic cells
consisting of a storage unit and a lookup table. The
interconnect between logic cells and the lookup tables can be
reprogrammed during run time to perform any possible
function (e.g., sorting, selection).
2.1.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Storage Devices</title>
        <p>Similar to the processing device, current systems o er a
variety of di erent storage devices used for data processing.
In this section, we discuss di erent properties of current
storage devices.</p>
        <p>Hard Disk Drive: The Hard Disk Drive (HDD), as a
non-volatile storage device, consists of several disks. The
disks of an HDD rotate, while a movable head reads or writes
information. Hence, sequential access patterns are well
supported in contrast to random accesses.</p>
        <p>
          Solid State Drive: Since no mechanical units are used,
Solid State Drives (SSDs) support random access without
high delay. For this, SSDs use ash-memory to persistently
store information [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ]. Each write wears out the ash cells.
Consequently, the write patterns of database systems must
be changed compared to HDD-based systems.
        </p>
        <p>
          Main-Memory: While using main memory as main
storage, the access gap between primary and secondary storage
is removed, introducing main-memory access as the new
bottleneck [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]. However, main-memory systems cannot omit
secondary storage types completely, because main memory is
volatile. Thus, e cient persistence mechanisms are needed
for main-memory systems.
        </p>
        <p>To conclude, current architectures o er several di erent
processor and storage types. Each type has a unique
architecture and speci c characteristics. Hence, to ensure high
performance, the processing characteristics of processors as
well as the access characteristics of the underlying storage
devices have to be considered. For example, if several
processing devices are available within a DBS, the DBS must
provide suitable algorithms and functionality to fully utilize
all available devices to provide peak performance.
2.2</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Software Functionality</title>
      <p>
        Besides hardware, DBS functionality is another source
of variability in a DBS. In Figure 2, we show an excerpt
of DBS functionalities and their dependencies. For
example, for di erent application domains di erent query types
might be interesting. However, to improve performance
or development cost, only required query types should be
used within a system. This example can be extended to
other functional requirements. Furthermore, a DBS
provides database operators, such as aggregation functions or
joins. Thereby, database operators perform di erently
depending on the used storage and processing model [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. For
example, row stores are very e cient when complete tuples
should be retrieved, while column stores in combination with
operator-at-a-time processing enable fast processing of single
columns [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. Another technique to enable e cient access
to data is to use index structures. Thereby, the choice of an
appropriate index structure for the speci c data and query
types is essential to guarantee best performance [
        <xref ref-type="bibr" rid="ref15 ref24">15, 24</xref>
        ].
      </p>
      <p>Note, we omit comprehensive relationships between
functionalities properties in Figure 2 due to complexity. Some
functionalities are mandatory in a DBS and others are
optional, such as support for transactions. Furthermore, it is
possible that some alternatives can be implemented together
and others only exclusively.
2.3</p>
    </sec>
    <sec id="sec-4">
      <title>Putting it all together</title>
      <p>So far, we considered variability in hardware and software
functionality separately. When using a DBS for a speci c
application domain, we also have to consider special
requirements of this domain as well as the interaction between
hardware and software.</p>
      <p>
        Special requirements comprise functional as well as
nonfunctional ones. Examples for functional requirements are
user-de ned aggregation functions (e.g., to perform genome
analysis tasks directly in a DBS [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]). Other applications
require support for spatial queries, such as geo-information
systems. Thus, special data types as well as index structures
are required to support these queries e ciently.
      </p>
      <p>
        Besides performance, memory footprint and energy e
ciency are other examples for non-functional requirements.
For example, a DBS for embedded devices must have a small
memory footprint due to resource restrictions. For that
reason, unnecessary functionality is removed and data
processing is implemented as memory e cient as possible. In this
scenario, tuple-at-a-time processing is preferred, because
intermediate results during data processing are smaller than
in operator-at-a-time processing, which leads to less memory
consumption [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ].
      </p>
      <p>
        In contrast, in large-scale data processing, operators should
perform as fast as possible by exploiting underlying
hardware and available indexes. Thereby, exploiting underlying
hardware is another source of variability as di erent
processing devices have di erent characteristics regarding
processing model and data access [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. To illustrate this fact,
we depict di erent storage models for DBS in Figure 2. For
example, column-storage is preferred on GPUs, because
rowstorage leads to an ine cient memory access pattern that
deteriorates the possible performance bene ts of GPUs [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
      <p>APPROACHES TO DESIGN
TAILORMADE DATABASE SYSTEMS</p>
      <p>
        The variability in hardware and software of DBSs can
be exploited to tailor database systems for nearly every
database-application scenario. For example, a DBS for
highperformance analysis can exploit newest hardware features,
such as SIMD, to speed up analysis workloads. Moreover,
we can meet limited space requirements in embedded
systems by removing unnecessary functionality [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ], such as the
support for range queries. However, exploiting variability is
one part of mastering variability in DBSs. The second part
is to manage variability e ciently to reduce development
and maintenance e ort.
      </p>
      <p>In this section, rst, we describe three di erent approaches
to design and implement DBSs. Then, we compare these
approaches regarding their applicability to arbitrary database
scenarios. Moreover, we assess the e ort to manage
variability in DBSs. Besides managing and exploiting the
variability in database systems, we also consider the robustness
and correctness of tailor-made DBSs created by using the
discussed approaches.
3.1</p>
    </sec>
    <sec id="sec-5">
      <title>One-Size-Fits-All Design Approach</title>
      <p>One way to design database systems is to integrate all
conceivable data management functionality into one single DBS.
We call this approach the one-size- ts-all design approach
and DBSs designed according to this approach one-size-
tsall DBSs. Thereby, support for hardware features as well
as DBMS functionality are integrated into one code base.
Thus, one-size- ts-all DBSs provide a rich set of
functionality. Examples of database systems that follow the
one-sizets-all approach are PostgreSQL, Oracle, and IBM DB2. As
one-size- ts-all DBSs are monolithic software systems,
implemented functionality is highly interconnected on the code
level. Thus, removing functionality is mostly not possible.</p>
      <p>DBSs that follow the one-size- ts-all design approach aim
at providing a comprehensive set of DBS functionality to
deal with most database application scenarios. The claim for
generality often introduces functional overhead that leads to
performance losses. Moreover, customers pay for
functionality they do not really need.
3.2</p>
    </sec>
    <sec id="sec-6">
      <title>Specialization Design Approach</title>
      <p>
        In contrast to one-size- ts-all DBSs, DBSs can also be
designed and developed to t very speci c use cases. We call
this design approach the specialization design approach and
DBSs designed accordingly, specialized DBSs. Such DBSs
are designed to provide only that functionality that is needed
for the respective use case, such as text processing, data
warehousing, or scienti c database applications [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ].
Specialized DBSs are often completely redesigned from scratch
to meet application requirements and do not follow common
design considerations for database systems, such as locking
and latching to guarantee multi-user access [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]. Specialized
DBSs remove the overhead of unneeded functionality. Thus,
developers can highly focus on exploiting hardware and
functional variability to provide tailor-made DBSs that meet
high-performance criteria or limited storage space
requirements. Therefore, huge parts of the DBS (if not all) must
be newly developed, implemented, and tested which leads
to duplicate implementation e orts, and thus, increased
development costs.
3.3
      </p>
      <p>Software Product Line Design Approach
In the specialization design approach, a new DBS must
be developed and implemented from scratch for every
conceivable database application. To avoid this overhead, the
SPL design approach reuses already implemented and tested
parts of a DBS to create a tailor-made DBS.</p>
      <p>
        To make use of SPL techniques, a special work ow has to
be followed which is sketched in Figure 3 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. At rst, the
domain is modeled, e.g., by using a feature model { a
treelike structure representing features and their dependencies.
With this, the variability is captured and implementation
artifacts can be derived for each feature. The second step,
the domain implementation, is to implement each feature
using a compositional or annotative approach. The third step
of the work ow is to customize the product { in our case,
the database system { which will be generated afterwards.
      </p>
      <p>By using the SPL design approach, we are able to
implement a database system from a set of features which are
mostly already provided. In best case, only non-existing
features must be implemented. Thus, the feature pool
constantly grows and features can be reused in other database
systems. Applying this design approach to DBSs enables
to create DBSs tailored for speci c use cases while
reducing functional overhead as well as development time. Thus,
the SPL design approach aims at the middle ground of the
one-size- ts-all and the specialization design approach.
3.4</p>
      <p>Characterization of Design Approaches
In this section, we characterize the three design approaches
discussed above regarding:
a) general applicability to arbitrary database applications,
b) e ort for managing variability, and
c) maturity of the deployed database system.</p>
      <p>
        Although the one-size- ts-all design approach aims at
providing a comprehensive set of DBS functionality to deal
with most database application scenarios, a one-size- ts-all
database is not applicable to use cases in automotive,
embedded, and ubiquitous computing. As soon as tailor-made
software is required to meet especially storage limitations,
one-size- ts-all database systems cannot be used. Moreover,
specialized database systems for one speci c use case
outperform one-size- ts-all database systems by orders of
magnitude [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]. Thus, although one-size- ts-all database systems
can be applied, they are often not the best choice regarding
performance. For that reason, we consider the applicability
of one-size- ts-all database systems to arbitrary use cases
as limited. In contrast, specialized database systems have
a very good applicability as they are designed for that
purpose.
      </p>
      <p>
        The applicability of the SPL design approach is good as
it also creates database systems tailor-made for speci c use
cases. Moreover, the SPL design approach explicitly
considers variability during software design and implementation
and provides methods and techniques to manage it [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. For
that reason, we assess the e ort of managing variability with
the SPL design approach as lower than managing variability
using a one-size- ts-all or specialized design approach.
      </p>
      <p>We assess the maturity of one-size- ts-all database
systems as very good, as these systems are developed and tested
over decades. Specialized database systems are mostly
implemented from scratch, so, the possibility of errors in the
code is rather high, leading to a moderate maturity and
robustness of the software. The SPL design approach also
enables the creation of tailor-made database systems, but
from approved features that are already implemented and
tested. Thus, we assess the maturity of database systems
created via the SPL design approach as good.</p>
      <p>In Table 1, we summarize our assessment of the three
software design approaches regarding the above criteria.</p>
      <sec id="sec-6-1">
        <title>Criteria</title>
      </sec>
      <sec id="sec-6-2">
        <title>a) Applicability b) Management e ort c) Maturity</title>
      </sec>
      <sec id="sec-6-3">
        <title>One-SizeFits-All ++</title>
      </sec>
      <sec id="sec-6-4">
        <title>Approach</title>
      </sec>
      <sec id="sec-6-5">
        <title>Specialization SPL ++ +</title>
        <p>+
+</p>
        <p>The one-size- ts-all and the specialization design approach
are each very good in one of the three categories
respectively. The one-size- ts-all design approach provides robust
and mature DBSs. The specialization design approach
provides greatest applicability and can be used for nearly every
use case. Whereas the SPL design approach provides a
balanced assessment regarding all criteria. Thus, against the
backdrop of increasing variability due to increasing variety
of use cases and hardware while guaranteeing mature and
robust DBSs, SPL design approach should be applied to
develop future DBSs. Otherwise, development costs for yet
another DBS which has to meet special requirements of the
next data-driven domain will limit the use of DBSs in such
elds.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>ARISING RESEARCH QUESTIONS</title>
      <p>
        Our assessment in the previous section shows that the
SPL design approach is the best choice for mastering
variability in DBSs. To the best of our knowledge, the SPL
design approach is applied to DBSs only in academic
settings (e.g., in [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]).Hereby, the previous research were based
on BerkeleyDB. Although BerkeleyDB o ers the essential
functionality of DBSs (e.g., a processing engine), several
functionality of relational DBSs were missing (e.g.,
optimizer, SQL-interface). Although these missing
functionality were partially researched (e.g., storage manager [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] and
the SQL parser [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]), no holistic evaluation of a DBS SPL
is available. Especially, the optimizer in a DBS (e.g., query
optimizer) with a huge number of crosscutting concerns is
currently not considered in research. So, there is still the
need for research to fully apply SPL techniques to all parts
of a DBS. Speci cally, we need methods for modeling
variability in DBSs and e cient implementation techniques and
methods for implementing variability-aware database
operations.
4.1
      </p>
    </sec>
    <sec id="sec-8">
      <title>Modeling</title>
      <p>
        For modeling variability in feature-oriented SPLs, feature
models are state of the art [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. A feature model is a set
of features whose dependencies are hierarchically modeled.
Since variability in DBSs comprises hardware, software, and
their interaction, the following research questions arise:
RQ-M1: What is a good granularity for modeling a
variable DBS?
In order to de ne an SPL for DBSs, we have to model
features of a DBS. Such features can be modeled with di erent
levels of granularity [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Thus, we have to nd an
applicable level of granularity for modeling our SPL for DBSs.
Moreover, we also have to consider the dependencies
between hardware and software. Furthermore, we have to nd
a way to model the hardware and these dependencies. In
this context, another research questions emerges:
RQ-M2: What is the best way to model hardware and
its properties in an SPL?
Hardware has become very complex and researchers demand
to develop a better understanding of the impact of
hardware on the algorithm performance, especially when
parallelized [
        <xref ref-type="bibr" rid="ref3 ref5">3, 5</xref>
        ]. Thus, the question arises what properties of
the hardware are worth to be captured in a feature model.
      </p>
      <p>
        Furthermore, when thinking about numerical properties,
such as CPU frequency or amount of memory, we have to
nd a suitable technique to represent them in feature
models. One possibility are attributes of extended
feature-models [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], which have to be explored for applicability.
4.2
      </p>
    </sec>
    <sec id="sec-9">
      <title>Implementing</title>
      <p>
        In the literature, there are several methods for
implementing an SPL. However, most of them are not applicable to
our use case. Databases rely on highly tuned operations
to achieve peak performance. Thus, variability-enabled
implementation techniques must not harm the performance,
which leads to the research question:
on inheritance or additional function calls, which causes
performance penalties. A technique that allows for variability
without performance penalties are preprocessor directives.
However, maintaining preprocessor-based SPLs is horrible,
which accounts this approach the name #ifdef Hell [
        <xref ref-type="bibr" rid="ref10 ref11">11, 10</xref>
        ].
So, there is a trade-o between performance and
maintainability [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ], but also granularity [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. It could be bene cial
for some parts of DBS to prioritize maintainability and for
others performance or maintainability.
      </p>
      <p>
        RQ-I2: How to combine di erent implementation
techniques for SPLs?
If the answer of RQ-I1 is to use di erent implementation
techniques within the same SPL, we have to nd an
approach to combine these. For example, database operators
and their di erent hardware optimization must be
implemented using annotative approaches for performance
reasons, but the query optimizer can be implemented using
compositional approaches supporting maintainability; the
SPL product generator has to be aware of these di erent
implementation techniques and their interactions.
RQ-I3: How to deal with functionality extensions?
Thinking about changing requirements during the usage of
the DBS, we should be able to extend the functionality in
the case user requirements change. Therefore, we have to
nd a solution to deploy updates from an extended SPL
in order to integrate the new requested functionality into a
running DBS. Some ideas are presented in [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], however,
due to the increased complexity of hardware and software
requirements an adaption or extension is necessary.
4.3
      </p>
    </sec>
    <sec id="sec-10">
      <title>Customization</title>
      <p>In the nal customization, features of the product line are
selected that apply to the current use case. State of the art
approaches just list available features and show which
features are still available for further con guration. However, in
our scenario, it could be helpful to get further information of
the con guration possibilities. Thus, another research
question is:
RQ-C1: How to support the user to obtain the best
selection?
In fact, it is possible to help the user in identifying suitable
con gurations for his use case. If he starts to select
functionality that has to be provided by the generated system,
we can give him advice which hardware yields the best
performance for his algorithms. However, to achieve this we
have to investigate another research question:
RQ-C2: How to nd the optimal algorithms for a
given hardware?
To answer this research question, we have to investigate the
relation between algorithmic design and the impact of the
hardware on the execution. Hence, suitable properties of
algorithms have to be identi ed that in uence performance
on the given hardware, e.g., access pattern, size of used data
structures, or result sizes.</p>
    </sec>
    <sec id="sec-11">
      <title>CONCLUSIONS</title>
      <p>RQ-I1: What is a good variability-aware
implementation technique for an SPL of DBSs?
Many state of the art implementation techniques are based
DBSs are used for more and more use cases. However,
with an increasing diversity of use cases and increasing
heterogeneity of available hardware, it is getting more
challenging to design an optimal DBS while guaranteeing low
implementation and maintenance e ort at the same time. To solve
this issue, we review three design approaches, namely the
one-size- ts-all, the specialization, and the software
product line design approach. By comparing these three design
approaches, we conclude that the SPL design approach is a
promising way to master variability in DBSs and to provide
mature data management solutions with reduced
implementation and maintenance e ort. However, there is currently
no comprehensive software product line in the eld of DBSs
available. Thus, we present several research questions that
have to be answered to fully apply the SPL design approach
on DBSs.</p>
    </sec>
    <sec id="sec-12">
      <title>ACKNOWLEDGMENTS</title>
      <p>This work has been partly funded by the German BMBF
under Contract No. 13N10818 and Bayer Pharma AG.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>D. J.</given-names>
            <surname>Abadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. R.</given-names>
            <surname>Madden</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N.</given-names>
            <surname>Hachem</surname>
          </string-name>
          .
          <article-title>Column-stores vs. Row-stores: How Di erent Are They Really? In SIGMOD</article-title>
          , pages
          <volume>967</volume>
          {
          <fpage>980</fpage>
          . ACM,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Apel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Batory</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>Kastner, and</article-title>
          <string-name>
            <given-names>G.</given-names>
            <surname>Saake.</surname>
          </string-name>
          Feature-Oriented
          <source>Software Product Lines</source>
          . Springer,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>C.</given-names>
            <surname>Balkesen</surname>
          </string-name>
          , G. Alonso,
          <string-name>
            <given-names>J.</given-names>
            <surname>Teubner</surname>
          </string-name>
          , and
          <string-name>
            <surname>M. T.</surname>
          </string-name>
          <article-title>Ozsu. Multi-Core, Main-Memory Joins: Sort vs</article-title>
          .
          <source>Hash Revisited. PVLDB</source>
          ,
          <volume>7</volume>
          (
          <issue>1</issue>
          ):
          <volume>85</volume>
          {
          <fpage>96</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D.</given-names>
            <surname>Benavides</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Segura</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Ruiz-Cortes</surname>
          </string-name>
          .
          <source>Automated Analysis of Feature Models 20 Years</source>
          Later:
          <string-name>
            <given-names>A Literature</given-names>
            <surname>Review. Inf</surname>
          </string-name>
          . Sys.,
          <volume>35</volume>
          (
          <issue>6</issue>
          ):
          <volume>615</volume>
          {
          <fpage>636</fpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>D.</given-names>
            <surname>Broneske</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Heimel</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Saake</surname>
          </string-name>
          .
          <article-title>Toward Hardware-Sensitive Database Operations</article-title>
          .
          <source>In EDBT</source>
          , pages
          <volume>229</volume>
          {
          <fpage>234</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>D.</given-names>
            <surname>Broneske</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bre</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Saake. Database Scan</surname>
          </string-name>
          <article-title>Variants on Modern CPUs: A Performance Study</article-title>
          .
          <source>In IMDM@VLDB</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>K.</given-names>
            <surname>Czarnecki</surname>
          </string-name>
          and
          <string-name>
            <given-names>U. W.</given-names>
            <surname>Eisenecker</surname>
          </string-name>
          .
          <article-title>Generative Programming: Methods, Tools, and Applications</article-title>
          . ACM Press/Addison-Wesley Publishing Co.,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S.</given-names>
            <surname>Dorok</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bre</surname>
          </string-name>
          , H. Lapple, and
          <string-name>
            <given-names>G.</given-names>
            <surname>Saake</surname>
          </string-name>
          .
          <article-title>Toward E cient and Reliable Genome Analysis Using Main-Memory Database Systems</article-title>
          . In SSDBM, pages
          <volume>34</volume>
          :
          <article-title>1{34:4</article-title>
          . ACM,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Dorok</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bre</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Saake. Toward E cient Variant</surname>
          </string-name>
          <article-title>Calling Inside Main-Memory Database Systems</article-title>
          .
          <source>In BIOKDD-DEXA. IEEE</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J.</given-names>
            <surname>Feigenspan</surname>
          </string-name>
          , C. Kastner, S. Apel,
          <string-name>
            <given-names>J.</given-names>
            <surname>Liebig</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Schulze</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Dachselt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Papendieck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Leich</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Saake. Do Background</surname>
          </string-name>
          <article-title>Colors Improve Program Comprehension in the #ifdef Hell? Empir</article-title>
          . Softw. Eng.,
          <volume>18</volume>
          (
          <issue>4</issue>
          ):
          <volume>699</volume>
          {
          <fpage>745</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J.</given-names>
            <surname>Feigenspan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Schulze</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Papendieck</surname>
          </string-name>
          , C. Kastner,
          <string-name>
            <given-names>R.</given-names>
            <surname>Dachselt</surname>
          </string-name>
          , V. Koppen, M. Frisch, and
          <string-name>
            <given-names>G.</given-names>
            <surname>Saake</surname>
          </string-name>
          .
          <source>Supporting Program Comprehension in Large Preprocessor-Based Software Product Lines. IET Softw</source>
          .,
          <volume>6</volume>
          (
          <issue>6</issue>
          ):
          <volume>488</volume>
          {
          <fpage>501</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>B.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. K.</given-names>
            <surname>Govindaraju</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Luo</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P. V.</given-names>
            <surname>Sander</surname>
          </string-name>
          .
          <source>Relational Query Coprocessing on Graphics Processors. TODS</source>
          ,
          <volume>34</volume>
          (
          <issue>4</issue>
          ):
          <volume>21</volume>
          :1{
          <fpage>21</fpage>
          :
          <fpage>39</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>B.</given-names>
            <surname>He</surname>
          </string-name>
          and
          <string-name>
            <given-names>J. X.</given-names>
            <surname>Yu</surname>
          </string-name>
          .
          <source>High-throughput Transaction Executions on Graphics Processors. PVLDB</source>
          ,
          <volume>4</volume>
          (
          <issue>5</issue>
          ):
          <volume>314</volume>
          {
          <fpage>325</fpage>
          ,
          <string-name>
            <surname>Feb</surname>
          </string-name>
          .
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>C.</given-names>
            <surname>Ka</surname>
          </string-name>
          stner, S. Apel, and
          <string-name>
            <given-names>M.</given-names>
            <surname>Kuhlemann</surname>
          </string-name>
          .
          <article-title>Granularity in Software Product Lines</article-title>
          .
          <source>In ICSE</source>
          , pages
          <volume>311</volume>
          {
          <fpage>320</fpage>
          . ACM,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>V.</given-names>
            <surname>Ko</surname>
          </string-name>
          <article-title>ppen, M. Schaler, and</article-title>
          <string-name>
            <given-names>R.</given-names>
            <surname>Schro</surname>
          </string-name>
          <article-title>ter. Toward Variability Management to Tailor High Dimensional Index Implementations</article-title>
          .
          <source>In RCIS</source>
          , pages
          <volume>452</volume>
          {
          <fpage>457</fpage>
          . IEEE,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>T.</given-names>
            <surname>Leich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Apel</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Saake</surname>
          </string-name>
          .
          <article-title>Using Step-wise Re nement to Build a Flexible Lightweight Storage Manager</article-title>
          .
          <source>In ADBIS</source>
          , pages
          <volume>324</volume>
          {
          <fpage>337</fpage>
          . Springer-Verlag,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>J.</given-names>
            <surname>Liebig</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Apel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Lengauer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Leich</surname>
          </string-name>
          .
          <source>RobbyDBMS: A Case Study on Hardware/Software Product Line Engineering</source>
          . In FOSD, pages
          <volume>63</volume>
          {
          <fpage>68</fpage>
          . ACM,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>A.</given-names>
            <surname>Lu</surname>
          </string-name>
          <article-title>bcke, V. Koppen, and</article-title>
          <string-name>
            <given-names>G.</given-names>
            <surname>Saake</surname>
          </string-name>
          .
          <article-title>Heuristics-based Workload Analysis for Relational DBMSs</article-title>
          . In UNISCON, pages
          <volume>25</volume>
          {
          <fpage>36</fpage>
          . Springer,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>S.</given-names>
            <surname>Manegold</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. A.</given-names>
            <surname>Boncz</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M. L.</given-names>
            <surname>Kersten</surname>
          </string-name>
          .
          <article-title>Optimizing Database Architecture for the New Bottleneck: Memory Access</article-title>
          . VLDB J.,
          <volume>9</volume>
          (
          <issue>3</issue>
          ):
          <volume>231</volume>
          {
          <fpage>246</fpage>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>R.</given-names>
            <surname>Micheloni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Marelli</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Eshghi</surname>
          </string-name>
          .
          <article-title>Inside Solid State Drives (SSDs</article-title>
          ). Springer,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>M.</given-names>
            <surname>Rosenmu</surname>
          </string-name>
          <article-title>ller</article-title>
          .
          <source>Towards Flexible Feature Composition: Static and Dynamic Binding in Software Product Lines. Dissertation</source>
          , University of Magdeburg, Germany,
          <year>June 2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>M.</given-names>
            <surname>Rosenmu</surname>
          </string-name>
          <article-title>ller</article-title>
          , N. Siegmund,
          <string-name>
            <given-names>H.</given-names>
            <surname>Schirmeier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sincero</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Apel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Leich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Spinczyk</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Saake.</surname>
          </string-name>
          FAME-DBMS:
          <article-title>Tailor-made Data Management Solutions for Embedded Systems</article-title>
          . In SETMDM, pages
          <article-title>1{6</article-title>
          . ACM,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>M.</given-names>
            <surname>Saecker</surname>
          </string-name>
          and
          <string-name>
            <given-names>V.</given-names>
            <surname>Markl</surname>
          </string-name>
          .
          <article-title>Big Data Analytics on Modern Hardware Architectures: A Technology Survey</article-title>
          . In eBISS, pages
          <volume>125</volume>
          {
          <fpage>149</fpage>
          . Springer,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>M.</given-names>
            <surname>Scha</surname>
          </string-name>
          <article-title>ler, A</article-title>
          . Grebhahn, R. Schroter,
          <string-name>
            <given-names>S.</given-names>
            <surname>Schulze</surname>
          </string-name>
          ,
          <string-name>
            <surname>V.</surname>
          </string-name>
          <article-title>Koppen, and</article-title>
          <string-name>
            <surname>G. Saake.</surname>
          </string-name>
          <article-title>QuEval: Beyond High-Dimensional Indexing a la Carte</article-title>
          .
          <source>PVLDB</source>
          ,
          <volume>6</volume>
          (
          <issue>14</issue>
          ):
          <volume>1654</volume>
          {
          <fpage>1665</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>M.</given-names>
            <surname>Stonebraker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Madden</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. J.</given-names>
            <surname>Abadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Harizopoulos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Hachem</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Helland</surname>
          </string-name>
          .
          <article-title>The End of an Architectural Era (It's Time for a Complete Rewrite)</article-title>
          .
          <source>In VLDB</source>
          , pages
          <volume>1150</volume>
          {
          <fpage>1160</fpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>S.</given-names>
            <surname>Sunkle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kuhlemann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Siegmund</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Rosenmuller, and</article-title>
          <string-name>
            <given-names>G.</given-names>
            <surname>Saake. Generating Highly Customizable SQL Parsers. In</surname>
          </string-name>
          <string-name>
            <surname>SETMDM</surname>
          </string-name>
          , pages
          <volume>29</volume>
          {
          <fpage>33</fpage>
          . ACM,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>T.</given-names>
            <surname>Willhalm</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Oukid</surname>
          </string-name>
          ,
          <string-name>
            <surname>I.</surname>
          </string-name>
          <article-title>Muller, and</article-title>
          <string-name>
            <given-names>F.</given-names>
            <surname>Faerber</surname>
          </string-name>
          .
          <article-title>Vectorizing Database Column Scans with Complex Predicates</article-title>
          .
          <source>In ADMS@VLDB</source>
          , pages
          <volume>1</volume>
          {
          <fpage>12</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhou</surname>
          </string-name>
          and
          <string-name>
            <given-names>K. A.</given-names>
            <surname>Ross</surname>
          </string-name>
          .
          <article-title>Implementing Database Operations Using SIMD Instructions</article-title>
          .
          <source>In SIGMOD</source>
          , pages
          <volume>145</volume>
          {
          <fpage>156</fpage>
          . ACM,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>M.</given-names>
            <surname>Zukowski</surname>
          </string-name>
          .
          <article-title>Balancing Vectorized Query Execution with Bandwidth-Optimized Storage</article-title>
          .
          <source>PhD thesis</source>
          , CWI Amsterdam,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>