<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Ph.D. Workshop, August</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Self-Driving: From General Purpose to Specialized DBMSs</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jan Kossmann Supervised by:  Dr. Hasso Plattner</string-name>
          <email>Jan.Kossmann@hpi.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Hasso Plattner Institute August-Bebel-Str.</institution>
          <addr-line>88 Potsdam</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <volume>27</volume>
      <issue>2018</issue>
      <fpage>2</fpage>
      <lpage>5</lpage>
      <abstract>
        <p>Large data sets, variable workloads of high complexity, and exible cloud infrastructure make the administration of database systems at the same time more challenging and more important. In the future, self-driving database systems will utilize workload-driven optimization and machine learning techniques to generate forecasts of future workloads, decide upon which actions to take to process workloads most efciently and to incorporate knowledge from past decisions into future ones without human intervention. But database systems are typically not designed having such capabilities in mind. We propose the architecture of a generalized framework that enables seamless integration of self-driving capabilities into database systems. Thereby, general-purpose database systems can transform themselves into systems tailored to a speci c use case. Furthermore, we present our scalable approach to nding solutions for large problem instances of a physical design challenge, i.e., the index selection problem. Both concepts have been implemented and partly evaluated with real-world data on the research database system Hyrise.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION &amp; MOTIVATION</title>
      <p>
        Today's relational database systems are general purpose
systems. They are designed to handle most real-world use
cases and workloads su ciently. To achieve this, such
systems rely on generic con gurations and algorithms.
However, to achieve cost e ciency, fully utilize existing hardware
resources, and achieve optimal performance such systems
need to be optimized. Database administrators (DBAs) are
responsible for manually tuning and con guring systems in
order to meet the speci ed Service Level Agreements (SLAs)
and deliver optimal performance. Tuning and con guration
can include decisions upon the available physical resources
(e.g., available network bandwidth, CPUs, main memory),
knob con guration (e.g., bu er pool size, number of
concurrent threads, di erent concurrency mechanisms), and
physical database design (e.g., partitioning or indexes). The
recent work of Kraska et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] on learned index data
structures goes one step further. The authors do not only present
very interesting results by outperforming traditional data
structures that were tuned for decades. Also, the
demonstrated solution of replacing core components of database
systems with machine learning models opens up
opportunities to rethink the tuning and con guration of database
systems in general. This might lead to a point where database
systems could autonomously optimize themselves and adjust
their con guration for the currently processed workload. In
Section 4 of this work we will detail how we equipped the
relational database system Hyrise with self-driving capabilities
and preliminarily evaluated our approach with real-world
data.
      </p>
      <p>There are three main trends which strengthen and
motivate the need for such self-driving database systems: (i)
more variable and complex workloads, (ii) the shift from
onpremise to cloud deployments, and (iii) an increased
number of available options and (hardware) con guration
dimensions. We will detail these trends in the following
paragraphs.</p>
      <p>
        First, today's relational database systems have to handle
a broad variety of complex, volatile, and combined
workloads. HTAP workloads contain transactional and
analytical queries. The recent increase of machine learning
applications on all sorts of data sets increases the workload
complexity even further. Data is not necessarily extracted
to process data- and computation-intensive machine
learning tasks, but the computations are executed directly on
a database system [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Thereby, workloads are more
variable, harder to assess and, in the end, manual tuning gets
increasingly challenging. Simultaneously, the need for high
performance to ful ll these complex workloads makes proper
con guration and tuning even more important.
      </p>
      <p>Second, the shift from on-premise to cloud database
solutions is another motivation. Nowadays, many cloud providers
o er the opportunity to deploy (traditional) relational
database systems on cloud infrastructure. There is a variety of
reasons to choose cloud over on-premise database
deployments. Flexibility caused by scalability and elasticity,
coste ectiveness because of better resource utilization, and the
avoidance of the need to own and maintain physical
infrastructure to name a few. Both parties cloud providers, and
cloud customers have an interest in self-driving database
systems. Cloud providers maintain a large number of
systems. Cost-e ciency and optimal resource utilization are a
necessity to be competitive. Hence, proper tuning and
conguration of their systems are of high importance. Manual
tuning of these systems by DBAs would be cumbersome,
time-consuming, and, hence very expensive. On the other
hand, cloud customers do also have an interest in optimized
systems. They usually pay for units of computation time
per machine and they also pay more for larger machines.
Therefore, they have an interest in renting fewer machines,
smaller machines and renting these for the shortest possible
periods. Thus, utilizing the rented resources in the most
performant way without investing manual administration e ort
is necessary to avoid unnecessary costs.</p>
      <p>
        Furthermore, database systems as well as the underlying
hardware o er increasingly more con guration and tuning
options. The number of knobs of database systems is
growing with every release of a new version. In 2016, MySQL
o ered more than 500 tunable knobs. In addition,
changing the state of one knob might a ect the impact of other
knobs [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Hence, they cannot be considered independently.
Some new hardware mechanisms can be dynamically
congured which adds further tuning dimensions, e.g., Intel's
Cache Allocation Technology [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. To exploit such
mechanisms to their full extent they need to be used taking the
currently processed and soon to be processed workload into
account.
      </p>
      <p>
        All the above-mentioned aspects demand specialized
congurations for each database deployment, but the e ort to
build those would be too high. If we leverage
workloaddriven optimization and machine learning techniques
(ranging from simple techniques, for example, linear regression
and decision trees to complex neural networks), we can
transform general purpose database systems into database
systems tailored to a speci c use case. This might cause
computationally intensive calculations. However, the computation
power of new CPUs, GPUs, and TPUs [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] enables e cient
processing of large problem instances.
2.
      </p>
    </sec>
    <sec id="sec-2">
      <title>RESEARCH ISSUES</title>
      <p>In this section, we want to highlight research problems
originating from self-driving databases. Developing
self-driving database systems is a complex problem involving
multiple (database) components. Therefore, there is no single
research problem, but a whole variety of problems arise. We
see three main areas that must be mastered to enable
selfdriving database systems.</p>
      <p>
        1. Tuning: The core problem of self-driving database
systems. The main task of such systems is to take certain
actions to increase performance with regards to, e.g.,
latency, memory consumption, or energy consumption.
Manifold reasons make tuning di cult. First, the sheer
complexity and problem size for large database
deployments. Furthermore, tuning decisions might
depend on other tuning decisions and considering them
separately leads often to suboptimal results [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Also,
heavyweight operations to put the tuning choices into
place, for example, repartitioning of large tables, can
easily negatively a ect system performance.
2. Feedback loop: Self-driving database systems have to
evaluate themselves and learn based on their past
tuning decisions in order to be independent of human
oversight. This is particularly hard with volatile workloads
that change frequently. The system's actions need to
be assessed even though the new workload might look
completely di erent as previously anticipated.
3. Forecasting: Tuning choices for the current workload
can be a ected if the system's workload is going to
shift in the next few minutes. Workloads can change,
e.g., based on seasonal e ects or load peaks caused by
high usage. Therefore, detailed forecasts are necessary
to anticipate future workloads.
      </p>
      <p>In addition, today's database systems were not designed
with self-driving capabilities in mind. Thus, we want to
investigate how a database design could support the above
presented main requirements without sacri cing performance
or well-established functionalities.</p>
      <p>
        Furthermore, we work on nding e cient and robust1
solutions, even for large problem instances, for tuning
problems. With more volatile and diverse workloads [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] the
requirements for high-performance operation of the database
system are constantly changing. Finding the most
promising set of actions to achieve the best performance for the
currently processed workload is of high importance. The
amount of time it takes to nd such a set of actions is not
less important. If the solution of a problem instance takes
large amounts of time, the workload could have changed,
thus rendering the solution out-dated.
3.
      </p>
    </sec>
    <sec id="sec-3">
      <title>RELATED WORK</title>
      <p>
        This section should demonstrate why the currently
existing solutions are not su cient. The beginnings of
selfdriving databases can be seen as early as in the 1970s where
self-adaptive databases could automatically tune parts of
their physical design [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Later, commercial database
systems [
        <xref ref-type="bibr" rid="ref12 ref13 ref3">3, 13, 12</xref>
        ] introduced tools and advisors that supported
DBAs in tuning tasks. In most cases, these tools do not
apply foresight to anticipate future workloads but mostly take
reactive measures. In addition, advisors still involve manual
human e ort to take the nal tuning decision.
      </p>
      <p>
        The modern term of self-driving databases [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] targets a
much wider and more holistic goal. These systems should be
completely self-su cient and maintain all aspects that are
necessary to make database systems run permanently and
e ciently without human intervention. The recent work
of Aken et al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] presents a system that combines
multiple machine learning techniques to automatically manage a
database system by tuning its knob con guration. Tuning
decisions are based on a large repository of previously tuned
systems that processed a similar workload. This is a great
step forward, but we see multiple open challenges here. For
example, the approach cannot yet work holistically. It only
handles knob con gurations. Physical database design is not
targeted at all. Instead, it expects a reasonable physical
design to be in place. This is a fair assumption at this point of
time, but in our opinion separating knob con guration and
physical design might, in the end, lead to suboptimal results.
Therefore, we favor a holistic approach. Furthermore, in
order to ful ll this tuning process, the system's workload is
characterized. Their system characterizes workloads based
on the database system's internal runtime metrics, e.g.,
acquired locks or written/read pages and not on the logical
query level. We argue that only because workloads show
the same runtime metrics they cannot necessarily be
classied as equal or similar. Especially high variable, complex
workloads that combine, for example, HTAP and machine
learning queries will be hard to characterize with such
approaches.
      </p>
      <p>
        There are state-of-the-art solutions for many tunable
aspects, e.g., indexes [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], views [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] or knob con guration [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
1Robust solutions o er not necessarily the best performance,
but aim to provide acceptable performance in most cases.
Evaluator
      </p>
      <p>Selector</p>
      <p>Executor
R</p>
      <p>Tuner</p>
      <p>Evaluations</p>
      <p>Choices
Runtime KPIs
Constraints
However, for large problem instances, these solutions show
unacceptable runtimes or limit candidate sets a priori
resulting in far from optimal results. If database systems should
become fully autonomous, minutes of runtime for tuning
a single aspect is not acceptable if dozens of tuning
decisions have to be taken. We identify this as an opportunity
for further research especially when problem instances grow
through ever-increasing database sizes.</p>
    </sec>
    <sec id="sec-4">
      <title>RESEARCH PLAN</title>
      <p>In this section, we describe the approach we pursue, which
work is currently conducted and planned for the future, and
how we plan to evaluate the proposed solution.
4.1</p>
    </sec>
    <sec id="sec-5">
      <title>Approach</title>
      <p>We have developed the architecture of a generalized
selfdriving framework and implemented this in the new version
of the Hyrise2 database system. Hyrise is a relational
mainmemory database system that stores tables in column-major
format. The recent rewrite of the Hyrise codebase gave us
the opportunity to take all design decisions with self-driving
capabilities in mind and plan the system's architecture
accordingly.</p>
      <p>Figure 1 depicts how we integrate the necessary
components to enable a self-driving database system into Hyrise.
For simplicity reasons not all interfaces and components can
be visualized in this gure. The Query Optimizer generates
e cient query plans from SQL strings. These plans are fed
into the Query Plan Cache. In case a query plan has
already been cached, the planning phase can be omitted. The
query plan cache plays an important role because it holds a
representation of current and past workloads. This is
necessary to generate forecasts for future workloads. The
central component responsible for self-driving is the driver. It
is responsible for ensuring that the database system most
e ciently processes the incoming workload while
considering speci ed constraints and at the same time not wasting
resources. The Workload Analyzer generates workload
forecasts of future workloads based on the data from the Query
Plan Cache. These forecasts have two main purposes. First,
they serve as input for the Organizer. The Organizer
controls and supervises the tuning process: it determines when
to start or abort tuning based on workload forecasts. It
also enforces the time constraints to ensure that the tuning
2https://github.com/hyrise/hyrise
process itself does not consume too many resources.
Second, they are used by the tuner. It nds its decisions based
on workload forecasts. The Tuner itself consists of three
components: The Evaluator component evaluates all
possible actions and assigns a negative or positive desirability
expressing the expected bene t for the system, a con dence
that is associated with the desirability, and a cost. We
expect one evaluator to exist for each tunable entity, for
example, indexes, partitions, knobs. Evaluators might consult
the query optimizer for (hypothetical) cost estimations. Its
output is consumed by a selector which selects the
mostbene cial actions while considering a cost budget. There
could be multiple exchangeable selectors each with a
different strategy, e.g., greedy, heuristic-based or solver-based
strategies, which show their strengths in di erent
scenarios. In the end, the Executor takes care of executing the
proposed actions by changing the con guration. This
architecture o ers the necessary exibility while avoiding
redundant components at the same time. The selector can
be exchanged by the organizer based on past tuning
experiences or current needs. While we need one evaluator
implementation per tunable aspect, multiple instances of a
single selector implementation can be used for multiple
tunable aspects. The driver does also draw conclusions from
past tuning runs. Therefore, it continuously monitors the
database system's runtime KPIs and observes whether the
desired e ects occurred. It might also change the query plan
cache to trigger the eviction of plans that became out-dated
after a tuning run. Constraints contain hardware resource
restrictions and service level agreements.</p>
      <p>Our self-driving framework can bene t from Hyrise's
exible architecture. Tables are physically partitioned into small
(consisting of around 100 000 rows) chunks. Small chunks
offer high exibility regarding data placement/movement,
indexing, compression and re-encoding. All of the above could
be triggered by tuning decisions of a self-driving database
system. Hyrise's architecture shows further bene ts.
Normally, for systems storing the data in large monolithic blocks
containing dozens of millions of tuples the time necessary for
automatic tuning can be hard to estimate. Also, applying
changes, e.g., re-encoding a large table, is a heavyweight,
time-consuming operation. Chunks divide the problem in
multiple smaller problems, allowing more lightweight
operations and better estimations of the necessary e ort. In
addition, chunks are append-only containers. Once they reach
their capacity they become immutable. Therefore, the
selfdriving database systems can make assumptions about the
0
underlying data with a higher certainty which leads in the
end to better tuning results.
4.2</p>
    </sec>
    <sec id="sec-6">
      <title>Current and Future Work</title>
      <p>The concept of the generalized self-driving framework is
currently implemented in Hyrise and under evaluation. We
have conducted rst experiments on real-world enterprise
data and the results look promising.</p>
      <p>Furthermore, we have started to investigate how large
problem instances of physical design challenges can be solved
e ciently without restricting the candidate set or relying
on external solvers. In the beginning, we studied the
wellresearched index selection problem. For a given workload
and memory budget the best index con guration is searched.
The preliminary results, also depicted in Figure 2, indicate
that we outperform current state-of-the-art approaches in
both scalability and solution quality for real-world as well
as synthetic workloads. Our approach relies on a
generalized recursive solution principle utilizing the query
optimizer's cost model. It is completely integrated with the
above presented generalized self-driving framework.
However, the number of optimizer calls is greatly reduced by the
approach's recursive structure.</p>
      <p>The next step is to create reliable workload forecasts on
the basis of the query plan cache data. Precise workload
forecasts do not only enable the tuner to nd the best
solutions for a certain workload, but also to nd robust
solutions that work acceptably well in cases where workloads
that were assumed to be less likely need to be processed.
Afterwards, we want to investigate how a feedback loop could
be established. The driver needs to monitor and assess the
e ect of past tuning decisions and in uence the tuner's
behavior for future decisions.
4.3</p>
    </sec>
    <sec id="sec-7">
      <title>Planned Evaluation</title>
      <p>Current evaluations of our approach were conducted with
data and workloads from a productive real-world Enterprise
Resource Planning application of a Fortune Global 500
company as well as with synthetic benchmarks. At the moment,
we are in the process of getting access to several productive
cloud database systems. These systems would o er both,
the possibility to learn based on data from real-world
systems and the opportunity to evaluate our approach in
practice for a large number of systems. We envision to
additionally take more complex synthetic benchmarks, e.g., the
TPC-DS benchmark into account to provide reproducible
examples.</p>
    </sec>
    <sec id="sec-8">
      <title>CONCLUSION</title>
      <p>We have presented our architecture of a generalized
selfdriving framework that integrates with typical database
components. The framework is capable of handling the three
main challenges for self-driving database systems:
forecasting of future workloads, adjusting the con guration
accordingly, and learning based on these adjustments. In addition,
we have demonstrated our concept and evaluations for the
e cient solution of large problem instances of physical
design challenges. Upcoming tasks include the implementation
of workload forecasting, assessing the impact of past con
guration decisions, and an extensive evaluation with more
real-world systems.
6.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <article-title>[1] Introduction to Cache Allocation Technology</article-title>
          . https://software.intel.com/en-us/articles/ introduction-to
          <article-title>-cache-allocation-technology</article-title>
          .
          <source>Accessed: 2018-04-10.</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D. V.</given-names>
            <surname>Aken</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Pavlo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. J.</given-names>
            <surname>Gordon</surname>
          </string-name>
          , and
          <string-name>
            <surname>B. Zhang.</surname>
          </string-name>
          <article-title>Automatic database management system tuning through large-scale machine learning</article-title>
          .
          <source>In Proceedings of the SIGMOD Conference</source>
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Chaudhuri</surname>
          </string-name>
          and
          <string-name>
            <given-names>V. R.</given-names>
            <surname>Narasayya</surname>
          </string-name>
          .
          <article-title>Self-tuning database systems: A decade of progress</article-title>
          .
          <source>In Proceedings of the 33rd VLDB Conference</source>
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D.</given-names>
            <surname>Dash</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Polyzotis</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Ailamaki</surname>
          </string-name>
          .
          <article-title>Cophy: A scalable, portable, and interactive index advisor for large workloads</article-title>
          .
          <source>PVLDB</source>
          ,
          <volume>4</volume>
          (
          <issue>6</issue>
          ):
          <volume>362</volume>
          {
          <fpage>372</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hammer</surname>
          </string-name>
          .
          <article-title>Self-adaptive automatic data base design</article-title>
          .
          <source>In American Federation of Information Processing Societies</source>
          : 1977 National Computer Conference.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>N. P.</given-names>
            <surname>Jouppi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Young</surname>
          </string-name>
          , and
          <string-name>
            <surname>N. P.</surname>
          </string-name>
          et al.
          <article-title>In-datacenter performance analysis of a tensor processing unit</article-title>
          .
          <source>In Proceedings of the 44th ISCA</source>
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>T.</given-names>
            <surname>Kraska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Beutel</surname>
          </string-name>
          , and
          <string-name>
            <surname>E. H. C.</surname>
          </string-name>
          et al.
          <article-title>The case for learned index structures</article-title>
          .
          <source>CoRR, abs/1712.01208</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>Kru</surname>
          </string-name>
          ger, C. Tinnefeld,
          <string-name>
            <given-names>M.</given-names>
            <surname>Grund</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zeier</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.</given-names>
            <surname>Plattner</surname>
          </string-name>
          .
          <article-title>A case for online mixed workload processing</article-title>
          .
          <source>In Proceedings of the Third DBTest</source>
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Luo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z. J.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. N.</given-names>
            <surname>Gubanov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. L.</given-names>
            <surname>Perez</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C. M.</given-names>
            <surname>Jermaine</surname>
          </string-name>
          .
          <article-title>Scalable linear algebra on a relational database system</article-title>
          .
          <source>In 33rd IEEE ICDE</source>
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>I.</given-names>
            <surname>Mami</surname>
          </string-name>
          and
          <string-name>
            <given-names>Z.</given-names>
            <surname>Bellahsene</surname>
          </string-name>
          .
          <article-title>A survey of view selection methods</article-title>
          .
          <source>SIGMOD Record</source>
          ,
          <volume>41</volume>
          (
          <issue>1</issue>
          ):
          <volume>20</volume>
          {
          <fpage>29</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>A.</given-names>
            <surname>Pavlo</surname>
          </string-name>
          , G. Angulo, and
          <string-name>
            <surname>J. A.</surname>
          </string-name>
          et al.
          <article-title>Self-driving database management systems</article-title>
          .
          <source>In CIDR</source>
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>K.</given-names>
            <surname>Yagoub</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Belknap</surname>
          </string-name>
          , and
          <string-name>
            <surname>B. D.</surname>
          </string-name>
          et al.
          <article-title>Oracle's SQL performance analyzer</article-title>
          .
          <source>IEEE Data Eng. Bull.</source>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>D. C.</given-names>
            <surname>Zilio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Rao</surname>
          </string-name>
          , and
          <string-name>
            <surname>S. L.</surname>
          </string-name>
          et al.
          <article-title>DB2 design advisor: Integrated automatic physical database design</article-title>
          .
          <source>In Proceedings of the Thirtieth VLDB</source>
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>