<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Scaling User Preference Learning in Near Real-Time to Large Datasets</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ian Beaver</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Joe Dumoulin NextIT Corporation</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>W. Riverside Ave</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Spokane WA</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>ibeaver</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>jdumoulin}@nextit.com http://www.nextit.com</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>In previous research we have shown the architecture and application of a case-based reasoning (CBR) system used to discover user preferences in an existing mixed-initiative dialogue system. In this paper we apply this CBR system to increasingly large datasets to test its ability to maintain nearreal time performance in generating new user preferences. We also propose possible future applications of the system.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Next IT is a company in Spokane WA, USA that builds
natural language applications for the worldwide web and
for mobile devices. As a way to increase user satisfaction
and reduce the number of turns required to complete tasks
by returning users, we developed a scalable system based
on MapReduce for learning user preferences from past
experience in near real-time. In a prior publication
        <xref ref-type="bibr" rid="ref2">(Beaver
and Dumoulin 2013)</xref>
        we describe the architecture and
operation of this system as well as provide some preliminary
performance testing results. We refer to this CBR system as
the Learned Preferences Generation Services (LPGS) in the
prior paper as well as this one.
      </p>
      <p>
        Since the publication of that paper, MongoDB, the data
store we chose to implement the system on, has switched
to using the V8 JavaScript engine internally
        <xref ref-type="bibr" rid="ref8">(MDB 2013)</xref>
        .
Thus the developer preview version we originally tested,
2.4, is now stable and in production. This allows us to test
the scaling performance of the LPGS on a multi-threaded
MapReduce engine, to see how it handles larger dataset sizes
that we would expect to see in a real world deployment.
      </p>
      <p>In the next section we briefly review the LPGS design and
application, all of which is covered in more detail in the
original publication. We then cover the testing and performance
of the system, followed by some future applications we hope
to soon support.</p>
    </sec>
    <sec id="sec-2">
      <title>System Overview and Application</title>
      <p>
        The LPGS was implemented using a CBR approach due to
the fact that we are attempting to partially automate a
conversation on behalf of a returning user leveraging specific
knowledge of previous conversations with the same user.
This is a key differentiator of CBR from other major AI
approaches that focus more on drawing generalizations and
associations from data and then applying them to specific
cases
        <xref ref-type="bibr" rid="ref1">(Aamodt and Plaza 1994)</xref>
        . By looking at specific
instances in a user’s history and reusing that information to
minimize the number of steps required for the user to repeat
the same tasks we can make the Natural Language System
(NLS) more efficient and increase the user satisfaction over
time.
      </p>
      <p>
        The CBR system architecture we designed is made up of a
Data Store, the LPGS, and a Search Service. The Data Store
is implemented in MongoDB, chosen for its schema
flexibility
        <xref ref-type="bibr" rid="ref3">(Berube 2012)</xref>
        and its ability to easily scale as the
number of cases increases
        <xref ref-type="bibr" rid="ref4">(Bonnet et al. 2011)</xref>
        while also
providing a built-in MapReduce framework eliminating the need to
deploy a separate MapReduce system. The Data Store
contains:
      </p>
      <p>User inputs for analysis (case memory). These are
individual task-related user interactions with the NLS and
include the user input text and meta data such as input
means, timestamp, and the NLS conversation state
variables.</p>
      <p>Learned preferences (case-base). Rules created from
successful cases that have been reviewed and retained for use
in future cases.</p>
      <p>User defined settings. Settings such as if the use of
preferences are enabled for a user, and per user thresholds of
repetitive behaviour before creating a preference solution.</p>
      <p>The Search Service is implemented as a lightweight
HyperText Transfer Protocol service that translates requests for
prior learned behaviours from the NLS into efficient queries
against MongoDB and returns any matching cases.</p>
      <p>The LGPS is implemented as a pipeline of MapReduce
jobs and filter functions. The inputs are the users
conversational history over a specific prompt in a specific task
from case memory, and outputs are any learned preferences,
which we refer to as rules, that can be assumed for that
prompt. A prompt by the NLS is an attempt to fill in a slot
in a form of information needed for the NLS to complete a
specific task. By pre-populating slots with a users past
answers, a task can be completed faster and with less back and
forth prompting and responding with the user.</p>
      <p>This MapReduce pipeline consists of two jobs. The first
MapReduce job, as seen in Figure 1, compresses continuous
user inputs that are trying to complete the same slot within
the same task. It may take the user several interactions with
the NLS to resolve a specific slot since the user may give
incorrect or incomplete data, or respond to the system prompt
with a clarifying question of their own.</p>
      <p>The compression is done by keeping track of when the
user was first prompted for the slot, and when the slot was
either filled in or abandoned. If the slot is eventually filled
in, the prompting case’s slot value and starting context are
combined with the final case’s slot value and ending context.
If the slot is never satisfied the conversation is thrown out as
there is no final answer to be learned from it.</p>
      <p>When this first job completes, the compressed cases are
stored along with the cases where the slot was resolved in a
single interaction. As shown in Figure 2, the second
MapReduce job is then started on the first job’s results. This job
attempts to count all of the slot outcomes for this user that are
equivalent. First by grouping all of the cases by end context
and then merging them into a single case, containing a list
of all starting states that created the end context.</p>
      <p>These final cases are passed through a set of functions
that determine if the answer was given often enough to
create a rule based on the users current settings. Cases that meet
these conditions are saved in the case-base as a learned
preference.</p>
      <p>This pipeline is applied to a user’s history any time the
user has a new interaction with the NLS containing tasks
that have ’learnable’ slots or whenever a user changes their
thresholds in the application settings. Domain experts
defining the set of tasks in the NLS also define slots within those
tasks that may be learned. Certain slots should never be
saved as cases for learning such as arrival or departure dates
or the body of a text message.</p>
    </sec>
    <sec id="sec-3">
      <title>Testing and Performance</title>
      <sec id="sec-3-1">
        <title>System Evaluation</title>
        <p>The primary measure of success for the learning system is
reducing the number of steps required for the user to
complete a task in the future. To evaluate this measure we needed
to ensure that when a user repeats a task as many times as
needed based on their settings, a rule is created and that rule
is found on the next attempt to complete the task. The
evaluation was done following these steps:
1. Create a new user account
2. Choose custom threshold settings or use system defaults
3. Walk through a task in the system conversationally
4. Repeat the conversation enough times to meet the set
thresholds
5. Assert that on the next attempt to complete the task a
prompt to validate a learned preference appears
6. Assert that on the next attempt to complete the task no
prompt appears but the task is completed using the learned
preference
Once the system was shown to be working correctly for a
single user, we released access to the UI in the form of a
mobile application to a limited group of 35 testers. The testers
had the ability to enable and disable the use of the learned
preferences during their conversation to compare the change
in experience. In our limited release testing user feedback
was very positive. One user commented that “Using the
application without learning enabled is annoying”, compared
to the experience with it enabled. This was due to the
decrease in prompting by the NLS on repeat uses with the
learning system enabled. An example of a conversation
collected from this evaluation is shown in Table. 1.</p>
      </sec>
      <sec id="sec-3-2">
        <title>System Performance</title>
        <p>One of the goals of the system is that the creation of new
preferences for a specific user must happen in near real-time
from when a user input is received. The definition of near
real-time in this context is driven purely by user experience.</p>
      </sec>
      <sec id="sec-3-3">
        <title>Initial Conversation(s)</title>
        <sec id="sec-3-3-1">
          <title>User: text Fred</title>
          <p>NLS: I found more than one phone number for
’Fred’. Which do you mean?</p>
          <p>Home - ###
Mobile - ###</p>
          <p>Work - ###</p>
        </sec>
        <sec id="sec-3-3-2">
          <title>User: Home number NLS: What do you want the text to say?</title>
        </sec>
      </sec>
      <sec id="sec-3-4">
        <title>Verification of New Rule</title>
        <sec id="sec-3-4-1">
          <title>User: text Fred</title>
          <p>NLS: I notice that in the past when you’ve asked me
to send a text message to Fred you chose to use
the phone number ###. Would you like me to
assume you mean ### now and in the future?
User: Yes.</p>
          <p>NLS: Great! I’ll make a note of it. What do you want
the text to say?</p>
        </sec>
      </sec>
      <sec id="sec-3-5">
        <title>Rule in Use</title>
        <sec id="sec-3-5-1">
          <title>User: text Fred NLS: What do you want the text to say?</title>
          <p>Since the system allows users to tailor the number of
repetitions required before learning an answer to a prompt, there is
an expectation by the user that, for example when the
number of times required is three, after the third time booking
a flight it will not ask them for their departure airport that
they have given the last three times in a row. If the system
were then to ask them for that information, their
expectations would not be met. In this example the definition of
near real-time must be less than a realistic window of time
before the user would repeat this task.</p>
          <p>In this domain of booking flights, several hours may be
an acceptable time frame since it is rare that users would
book multiple flights in a several hour period leaving from
the same airport. There may be domains where the same
tasks are completed many times a day, as in a personal
assistant domain where the user wants the system to learn that
a nickname is associated to a specific contact they write text
messages to often. In this domain the acceptable time frame
may be only a matter of minutes. Therefore we recognize
that since this acceptable time frame varies by domain and
expectations of the user base, we can only show how the
system scales within the limits of the testing hardware
available to us and know there will be larger computing capacity
needed to cover domains with fast preference availability
expectations or large concurrent user bases.</p>
          <p>
            Scaling Concerns A consideration in the initial system
design was to make it easy to scale the system to meet the
demands of an ever growing Data Store of user histories,
and an ever increasing user base. To meet this need we
selected the MapReduce programming model as it was
designed to run on large clusters of commodity hardware and
automatically partition the data across the machines
            <xref ref-type="bibr" rid="ref5">(Dean
and Ghemawat 2008)</xref>
            . A motivation of MapReduce is to
push the data closer to the processing. As the processing is
distributed across commodity servers the data is distributed
along with it, allowing the data size to continue to grow
without greatly impacting performance. There are many
different MapReduce engines available, Hadoop being perhaps
the best known. It has been well proven in industry with
Hadoop clusters over 5,000 nodes in size existing in
production (Morgan 2013). However, as MongoDB includes a
MapReduce engine, we use it instead of adding additional
complexity by requiring an external engine such as Hadoop.
          </p>
          <p>
            In this architecture, as the size of the Data Store grows, the
load on an individual server can remain constant by simply
adding more servers to the cluster and letting the data
rebalance across them. MongoDB handles this data partitioning
through a mechanism called Sharding, where a single
collection of data is distributed across multiple servers or shards
            <xref ref-type="bibr" rid="ref9">(MDB 2014)</xref>
            . Each shard is an independent database that can
execute MapReduce functions on its partition of the data.
For example, if the Data Store contains 1 terabyte of data,
and there are 4 shards in the cluster, then each shard only
has to operate on 256GB of data. If there are 40 shards in
the cluster, then each shard only needs to operate on around
25GB of data.
          </p>
          <p>To test that the LPGS was capable of scaling to large
numbers of cases, we needed to create a test data set in
incremental sizes and show how performance degrades. We measure
the average time it took to execute a single MapReduce job
across the conversation history data, and the average time to
analyse a single user for the complete set of tasks defined in
the NLS for each case memory size. Since the LPGS only
works on users that added new cases since the last time it
ran, running against all users would be a test of the worse
case scenario in the system.</p>
        </sec>
      </sec>
      <sec id="sec-3-6">
        <title>Test Dataset Creation</title>
        <p>In order to create large data sets for the purpose of load
testing, we used actual conversations from users of an existing
NLS in the personal assistant domain. These conversations
were inserted into the case memory directly, truncated in a
way that the number of inputs or cases per user would form
a Normal Distribution where = 365; = 168 with
negatives remapped as
f (n) =
2 1</p>
        <p>Z
10</p>
        <p>1
&lt; 1</p>
        <p>Where 1 Z 10 is generated at random. This larger
distribution between 1. . . 10 is to simulate users that try the
NLS out of curiosity with no intention of accomplishing any
task and then abandon it. We chose and values based on
projected usage expectation in the personal assistant domain
after reviewing historical NLS usage in current production
environments. We then simply marked all of the users as
having new cases available, this way the LPGS would have
to look at all users at the same time, creating the maximum
load on the system given the size of the Data Store.
the chunks and shards and a more predictable performance
curve.</p>
      </sec>
      <sec id="sec-3-7">
        <title>Testing Environment</title>
        <p>The MongoDB cluster was constructed with 8 homogeneous
servers with 2xE5450 CPUs, 16GB RAM and 2x73GB 15k
rpm drives with RAID0. The system OS is Ubuntu Server
12.04LTS and the database and MapReduce system is
MongoDB v2.4.5.</p>
        <p>MongoDB was configured as 4 shards of 2-node replica
sets. In the case memory and case-base collections, the ID
was used as the shard key. In the user settings collection the
UserID was used as the shard key. The LPGS was running on
a workstation with an Intel i7-3930K CPU and 64GB RAM
and was configured to use 32 worker threads, meaning 32
users’ histories would be analysed in parallel. This number
is configurable based on the computing power of the
machine the LPGS is running on. Multiple instances can be
started on multiple machines as well in order to reduce the
total analysis time as the number of concurrent users grows.</p>
      </sec>
      <sec id="sec-3-8">
        <title>Performance Results</title>
        <p>The case memory was initially populated with 1,000 unique
user histories. All users were flagged as having new data so
that the LPGS would process the entire case memory as a
test of a worse case scenario. The average MapReduce
wallclock time and average user analysis wall-clock time were
then recorded. After each run more users were imported in
case memory and the database cluster and LPGS were fully
restarted in order to clear out any cached data that may skew
the benchmarks. The next size tested was with 2,000 unique
users, followed by 5,000, at which point the case memory
was increased by 5,000 users each run. We stopped when we
had reached 40,000 unique users, which comprised around
14.7 million cases, as we had reached the limits of the disk
space available on our MongoDB cluster.</p>
        <p>Anomalies The results of the testing can be seen in Figure
4 and Figure 5. In both figures, there can be seen an anomaly
at case memory sizes of 5,000 users and 25,000 users. This
is due to the fact that MongoDB’s balancer uses by default
a range function to partition the data across the shards. As
the shard key used was the MongoDB ObjectID the most
significant bits represent a time stamp. This means that they
increment in a regular and predictable pattern.</p>
        <p>
          This monotonically increasing number when inserted
through the range function causes the inserted cases to be
written into the same chunk of data, until the chunk size
limit is reached and the chunk is split into two and moved
          <xref ref-type="bibr" rid="ref9">(MDB 2014)</xref>
          . Also as each user history was inserted
sequentially when populating the case memory, the majority
of a single user’s history will reside in a single chunk on
a single server. These issues create imbalances in the shard
distribution where at certain data set sizes more chunks may
reside on one shard than the others creating a
disproportionately higher load on that shard and increasing the
MapReduce execution times. For our use case it would be better
in the future to use a hash function to partition the data as
it would lead to a more random distribution of cases across
Memory Saturation The other finding shown in these
figures is that around the 25,000 user mark, the average
MapReduce time remains near constant for the rest of the
case memory sizes. This is due to the fact that around the
25,000 user mark the memory is saturated on the shards
and they must swap out to disk. At this threshold the
performance of the system in the worst case scenario does not
continue to degrade as the Data Store grows. This
threshold could be raised using the same server configuration by
adding more memory to the shards.
        </p>
        <p>Conclusions Therefore we conclude that given the
hardware configuration we tested, the worst it can perform is
around 500 milliseconds per MapReduce job given that we
are analysing 32 concurrent user histories. We can also
conclude that adding more shards to the system would reduce
the memory usage per shard and delay the point at which
this threshold is reached. An added benefit of adding more
shards as opposed to increasing existing shard memory is
that the processing load would be distributed across more
servers. What still needs to be evaluated is how the number
of concurrent users analysed affect the performance
degradation and worst case average MapReduce times.</p>
        <p>MongoDB handles the load of 32 parallel MapReduce
jobs on completely separate (meaning uncached) data very
well. The total time it takes to process 40,000 users would
be acceptable in most domains without needing to use
multiple instances of the LPGS. The Average User Analysis time
meets our definition of near real-time given the size of the
data we tested even as a worse case scenario. In a real world
case where 40,000 unique users need to be reviewed
concurrently would in all likelihood mean there was a great deal
more total users in the system.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Future Applications</title>
      <p>Given that the system performance is acceptable and
scalable, we have some ideas on using this same processing
pipeline to further enhance user experiences.</p>
      <p>
        Reasoning About Specific Users As CBR is a
methodology for both reasoning and learning
        <xref ref-type="bibr" rid="ref7">(Kolodner 1992)</xref>
        and we
are primarily using it for learning, we could use the same
data for reasoning about specific users as well. An example
would be if a user has the use of learning enabled, but has
rejected every potential new rule that has been found for a
specific task, we may assume that this user does not want us
trying to automate that task. This would allow the user to get
the benefit of learned preferences in other tasks without the
annoyance of occasionally invalidating new potential rules
for a task they do not wish to automate.
      </p>
      <p>Let Users Define Learnable Slots In the current system,
the slots that are watched for learning are defined by domain
experts when creating the set of tasks the system can
preform. As a way to make the learning system more
personalized, a user could add a slot to be learnable for them and
select context to a rule from the available system context.
This could be exposed through a UI element that is shown
next to prompts that are part of a task. When the user clicks
on the element, a dialogue could appear that lets them select
which context elements they think are relevant to the answer.</p>
      <p>External Information Sources A possibility we have
considered in travel domains is to use data from external sources
to take into account weather, delays, flight changes, and
other travel info and look at how that affects use of the NLS
by users. If we took into account this meta-data from sources
outside of users, we could begin to predict usage spikes in
the system when alerts like weather changes or flight delays
are present.</p>
      <p>Time and Location Awareness By adding time and
locality information as features in the conversation state, we can
look at what conversations are had during what times of day
and in what locations. For example, a personal assistant
application may notice that this specific user always listens to
their Workout playlist in the gym at 9AM Monday through
Friday. Therefore the system could learn that when the user
wants to listen to music around 9AM on a weekday, and
their current location is at the gym, it should just start
playing their Workout playlist.</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>
        We have shown that our CBR system used to discover user
preferences is able to scale to real world workloads and still
maintain acceptable performance. In a worst case scenario
the system was able to handle the growing case memory
size up to the limits of its disk space. Given that our test
cluster used 4 shards when MongoDB supports up to 1,000
        <xref ref-type="bibr" rid="ref6">(Horowitz 2011)</xref>
        , we are confident that the system would
continue to scale several orders of magnitude more than our
test data size. We also proposed some ideas taking advantage
of this scalable processing pipeline to leverage the same
system to do more than automate conversational tasks for users.
      </p>
      <p>Morgan, T. 2013. Big hadoop shops are
on a hockey stick growth curve.
EnterpriseTech Blog. Available online at http:
//www.enterprisetech.com/2013/10/30/
big-hadoop-shops-hockey-stick-growth-curve.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Aamodt</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Plaza</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <year>1994</year>
          .
          <article-title>Case-based reasoning: Foundational issues, methodological variations, and system approaches</article-title>
          .
          <source>AI</source>
          communications
          <volume>7</volume>
          (
          <issue>1</issue>
          ):
          <fpage>39</fpage>
          -
          <lpage>59</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Beaver</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Dumoulin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <year>2013</year>
          .
          <article-title>Applying mapreduce to learning user preferences in near real-time</article-title>
          .
          <source>In Case-Based Reasoning Research and Development</source>
          . Springer.
          <fpage>15</fpage>
          -
          <lpage>28</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Berube</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <year>2012</year>
          .
          <article-title>Encode video with mongodb work queues</article-title>
          .
          <source>IBM</source>
          .com. Available online at https://www.ibm.com/developerworks/ library/os-mongodb
          <article-title>-work-queues/ os-mongodb-work-queues-pdf</article-title>
          .pdf.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Bonnet</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Laurent</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Sala</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Laurent</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Sicard</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <year>2011</year>
          .
          <article-title>Reduce, you say: What nosql can do for data aggregation and bi in large repositories</article-title>
          .
          <source>In Database and Expert Systems Applications (DEXA)</source>
          ,
          <year>2011</year>
          22nd International Workshop on,
          <fpage>483</fpage>
          -
          <lpage>488</lpage>
          . IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Ghemawat</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <year>2008</year>
          .
          <article-title>Mapreduce: simplified data processing on large clusters</article-title>
          .
          <source>Communications of the ACM</source>
          <volume>51</volume>
          (
          <issue>1</issue>
          ):
          <fpage>107</fpage>
          -
          <lpage>113</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Horowitz</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <year>2011</year>
          .
          <article-title>The secret sauce of sharding</article-title>
          .
          <source>MongoSF</source>
          . Available online at http://www.mongodb.com/ presentations/secret-sauce-sharding.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Kolodner</surname>
            ,
            <given-names>J. L.</given-names>
          </string-name>
          <year>1992</year>
          .
          <article-title>An introduction to case-based reasoning</article-title>
          .
          <source>Artificial Intelligence Review</source>
          <volume>6</volume>
          (
          <issue>1</issue>
          ):
          <fpage>3</fpage>
          -
          <lpage>34</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>MDB.</surname>
          </string-name>
          <year>2013</year>
          .
          <article-title>Performance improvements</article-title>
          .
          <source>MongoDB 2</source>
          .4 Release Notes. Available online at http://docs.mongodb.org/manual/ release-notes/2.4/#v8-javascript-engine.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>MDB.</surname>
          </string-name>
          <year>2014</year>
          .
          <article-title>Sharding and mongodb</article-title>
          .
          <source>MongoDB Documentation Project</source>
          . Available online at http://docs.mongodb.org/master/ MongoDB-sharding-guide.pdf.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>