<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>ProCAKE: A Process-Oriented Case-Based Reasoning Framework</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ralph Bergmann</string-name>
          <email>bergmann@uni-trier.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lisa Grumbach</string-name>
          <email>grumbach@uni-trier.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lukas Malburg</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christian Zeyen</string-name>
          <email>zeyen@uni-trier.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Business Information Systems II, University of Trier</institution>
          ,
          <addr-line>54286 Trier</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents ProCAKE - the process-oriented casebased knowledge engine of the CAKE framework, which has evolved from several research projects at the University of Trier over the years. ProCAKE constitutes a domain-independent framework that can be used to implement diverse structural or process-oriented case-based reasoning applications for integrated process and knowledge management. This paper gives an overview of the main components and demonstrates their application by examples.</p>
      </abstract>
      <kwd-group>
        <kwd>Knowledge Management</kwd>
        <kwd>Process Management</kwd>
        <kwd>CaseBased Reasoning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Nowadays, workflow technology is widely used in business, scientific, and even
private domains. However, implementing the technology involves a significant
amount of knowledge, which is why integrated process and knowledge
management is of great importance. Process-Oriented CBR (POCBR) particularly
addresses this integration by applying and extending CBR methods for process and
workflow management [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. In general, many research prototypes were presented
in the field of CBR but only few frameworks are publicly available for
developing CBR applications. Recent frameworks such as myCBR [
        <xref ref-type="bibr" rid="ref1 ref8">1,8</xref>
        ] or jColibri [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]
mainly focus on structural and textual CBR. We are not aware of any
framework that is particularly tailored to the development of POCBR applications. To
this end, we present ProCAKE, a generic framework for building structural and
process-oriented CBR applications. The software is developed at the Department
of Business Information Systems II at the University of Trier. The source code
is freely available on our website1. ProCAKE constitutes the core of the CAKE
framework [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and builds the foundation for various past and ongoing research
activities of our research group. The developed prototypes include a broad range
of algorithms for retrieval and adaptation [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>Architecture Overview</title>
      <p>The ProCAKE framework is written in Java while it uses XML for
configuration and persistence. It has a pattern driven architecture and relies heavily on
interfaces and factories. ProCAKE provides its own data type structure (also
referred to as system classes) for defining the domain model and cases. The most
commonly used types are:
Base Classes The base classes include Atomic classes such as Boolean,
Numeric, String, Chronologic, and Void. Atomic classes can be combined with
composite classes such as Aggregate, Interval, or Collection. By this means,
cases for structural CBR can be represented.</p>
      <p>
        NEST Classes The NESTGraph class is a specific composite class for
representing workflows or processes as semantic graphs (cf. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]). It encloses further
composite classes for representing the graph elements. To name just a few,
graph edges are represented by NESTPartOfEdge, NESTControlflowEdge ,
and NESTDataflowEdge classes and graph nodes are represented by
NESTWorkflowNode , NESTTaskNode, and NESTDataNode, respectively. Graph
element classes are linked to user classes that describe their semantics using
custom composite classes.
      </p>
      <p>
        ProCAKE implements many syntactic and semantic similarity measures for
the various data types. Most of the measures are formally described in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. For
example, measures for Numeric classes use linear, exponential, and threshold
functions whereas measures for String classes apply Levenshtein or regular
expressions. Several taxonomic measures exist for semantic similarity assessment. For
the NESTGraph class, a measure is implemented that performs graph-matching
with an A* search algorithm [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Several algorithms for retrieval are implemented:
Besides a k-NN retrieval, several MAC/FAC approaches and an A* parallel
retriever are implemented for accelerating retrieval with large case bases. An
adaptation framework enables the integration and application of domain-dependent
adaptation methods.
      </p>
      <p>Both the user classes and similarity measures can be specified via XML. The
core components for instantiating ProCAKE are:
Configuration A benefit of the pattern driven architecture of ProCAKE is
the extensibility. For instance, the various implementations for persistence,
retrieval, and adaptation are organized in factories. Implementations can be
registered and configured in an XML file (usually named composition.xml)
that is read once at start-up.</p>
      <p>Data Model In addition to the system classes, custom-structured classes
(referred to as user classes) can be defined as sub-types of system classes. At
start-up, the system reads the model configuration files. Even though several
custom models can be defined, the usual practice is to use a single model
definition as default (usually named model.xml).</p>
      <p>Similarity Model To allow for comparing system and user classes, similarity
measures have to be defined for each class. Analogous to the data model,
several similarity models can be defined while usually a single model (named
sim.xml) is used as default. A similarity measure can be selected as default
for a class, so that the measure is applied for all sub-classes of that class.
It must be ensured that a similarity measure is defined for each system and
user class.</p>
      <p>Objects Analogous to the Java perspective, objects in ProCAKE represent the
concrete data. However, object classes are used to explicitly model the
relationship to the respective system and user classes. Every data used in a
case or query for retrieval has to be represented by such objects.
Consequently, the presented data type structure represents all possible data types
ProCAKE can operate on.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Example Application</title>
      <p>In the following, we demonstrate an exemplary application of ProCAKE. For
this purpose, we consider cooking recipes as a simple form of workflows. The
case base comprises 40 sandwich recipes2. A cooking workflow is represented as
a semantic graph in which each node is associated with a semantic description
(cf. Fig. 3). Preparation steps are represented as task nodes and ingredients
are represented as data nodes. A workflow node represents general information
about the recipe.
2 Recipes were extracted manually from https://allrecipes.com</p>
      <p>For the semantic descriptions of the graph nodes, we define specific attributes
with help of user classes in the data model (see excerpt in Fig. 1). The semantics
of the workflow node is expressed by a class named WorkflowSemantic. It includes
attributes such as name of type string or preparation time and calories, both of
type integer with lower and upper bound (see Fig. 1, lines 1–9). Data nodes
are described by a class named DataSemantic that includes attributes name
and amount (see lines 11–14). The latter is further divided into value and unit
(see lines 16–19). Unit can be one of predefined strings and value is an integer.
Possible values for the name attribute are also predefined and taxonomically
ordered (see lines 21f.). The description of task nodes consists of a string object
called name whose values are also taxonomically ordered.</p>
      <p>On the basis of this model, the similarity model (see excerpt in Fig. 2)
deifnes similarity measures for comparing the objects. Similarity measures are
determined for local attributes as well as for aggregate objects and whole workflow
graphs. For attributes like name, preparation time, or calories, we apply simple
measures such as Levenshtein distance for strings or a linear function for
numeric values (see Fig. 2, line 8). Ingredient types are compared on the basis of
the given taxonomy order and manually annotated similarity values (see lines
20f.). The similarity of aggregate objects is mostly computed by a weighted
average function (see lines 1–6 and 10–13). For the attribute amount, we apply
the aggregate minimum (cf. line 15), because if the unit types are not equal (the
similarity is 0), the value is not directly comparable. In this event, the minimum
function ensures that the similarity of the value attribute is ignored. To compute
the similarity of the workflow graphs, we apply the A* similarity measure.</p>
      <p>An exemplary query and the corresponding local similarities to an example
case are depicted in Fig. 3. Workflow nodes are represented as rhombuses, task
nodes as rectangles, and data nodes as ovals. Semantic descriptions of the nodes
are written in grey rectangles. Solid edges with description po indicate part-of
edges whereas cf and df denote control-flow and data-flow edges, respectively.</p>
      <p>Query</p>
      <p>WQ
prep. time (min): 35
calories: 500
po
po</p>
      <p>t1
df name: shred
( , ) = 0.75
( ,1
... cf mmax(t1) cf ...
name: slice po</p>
      <p>df
mmax(d1)</p>
      <p>po
name: mozzarella
amount:
unit: piece
value: 1/4</p>
      <p>Case
WC
prep. time (min): 20
calories: 1000
[...]</p>
      <p>The query consists of three constraints: The desired preparation time is set
to 35 minutes, whereas the desired amount of calories is set to 500. This
information is annotated at the workflow node WQ. Furthermore, shredded cheese
is desired, which is represented by the partial workflow containing a data node
cheese (d1) as input to the task shred (t1). The query graph can be used as
input for retrieving suitable workflow graphs from the case base. For
determining the similarity between two workflow graphs, an A* algorithm searches for
the best mappings between the nodes and edges. Each mapping is rated with a
similarity. Figure 3 depicts the similarities (dashed arrows) of the best possible
mappings (mmax) between nodes. Please note that edges are also mapped and
rated with similarities between the graphs. The global similarity of the
worklfows is obtained by aggregating the local similarities between the mappings of
all elements of the query graph to the elements of the case graph.</p>
      <p>The similarities of mapped elements are computed locally by comparing the
semantic descriptions. In the given example, the data nodes are considered to
be equal (sim(d1; mmax(d1)) = 1), as mozzarella is a child node of cheese in the
taxonomy and further semantic attributes are not given in the query. The local
similarity of the task nodes, here shred and slice, originates from the given
values in the taxonomy. In the example, we assume that the common parent node
is annotated with a similarity value of 0.5, i.e., sim(t1; mmax(t1)) = 0:5. The
similarity of the workflow nodes is composed of an aggregated average of
single similarity values for preparation time and calories attributes. The preparation
time given in the query should not be exceeded. Thus, we apply a threshold
measure, which sets the similarity value to 0, if the limit is reached. Since preparation
time of the case is lower than that of the query, the obtained similarity is 1. To
compute the similarity of the attribute calories, we use a linear numeric measure,
leading to a similarity of 0.5. In conclusion, the overall similarity of the
worklfow nodes results from the sum of single weighted similarities: sim(WQ; WC ) =
0:5 simtime(WQ; WC ) + 0:5 simcal(WQ; WC ) = 0:5 0:5 + 0:5 1 = 0:75.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Future Work</title>
      <p>We are continuously improving and extending features and the documentation.
Our goal is to integrate domain-independent algorithms into the generic
ProCAKE framework. For instance, current work focuses on transferring
generalization and specialization methods for domain-specific workflow representations
to arbitrary user classes. The latest version of ProCAKE already includes many
features of our research prototypes. It particularly provides various similarity
measures, several retrieval methods, and a generic adaptation manager. The
example application of ProCAKE is publicly available to foster the
implementation of new applications. To further facilitate the development, we are currently
working on a graphical tool for visualizing and creating workflow graphs. We
appreciate any suggestions or feedback and are open to extensions.
Acknowledgements. This work is funded by the German Research Foundation
(DFG) under grant no. BE 1373/3-3.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bach</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Althof</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Developing case-based reasoning applications using mycbr 3</article-title>
          .
          <source>In: Case-Based Reasoning Research and Development - 20th Int. Conf., ICCBR 2012. Proceedings. LNCS</source>
          , vol.
          <volume>7466</volume>
          , pp.
          <fpage>17</fpage>
          -
          <lpage>31</lpage>
          . Springer (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bergmann</surname>
          </string-name>
          , R.:
          <source>Experience Management: Foundations</source>
          ,
          <string-name>
            <given-names>Development</given-names>
            <surname>Methodology</surname>
          </string-name>
          , and
          <string-name>
            <surname>Internet-Based</surname>
            <given-names>Applications</given-names>
          </string-name>
          , LNCS, vol.
          <volume>2432</volume>
          . Springer (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bergmann</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gessinger</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Görg</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Müller</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>The Collaborative Agile Knowledge Engine CAKE</article-title>
          .
          <source>In: Proc. of the 18th Int. Conf. on Supporting Group Work</source>
          ,
          <year>2014</year>
          . pp.
          <fpage>281</fpage>
          -
          <lpage>284</lpage>
          . ACM (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Bergmann</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gil</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Similarity assessment and eficient retrieval of semantic workflows</article-title>
          .
          <source>Information Systems</source>
          <volume>40</volume>
          ,
          <fpage>115</fpage>
          -
          <lpage>127</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Bergmann</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Minor</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Müller</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schumacher</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <string-name>
            <surname>Project</surname>
            <given-names>EVER</given-names>
          </string-name>
          :
          <article-title>Extraction and Processing of Procedural Experience Knowledge in Workflows</article-title>
          .
          <source>In: Proc. of ICCBR 2017 Workshops. CEUR Proc.</source>
          , vol.
          <year>2028</year>
          , pp.
          <fpage>137</fpage>
          -
          <lpage>146</lpage>
          . CEUR-WS.org (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Minor</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Montani</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Recio-García</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          :
          <source>Process-Oriented Case-Based Reasoning. Information Systems</source>
          <volume>40</volume>
          ,
          <fpage>103</fpage>
          -
          <lpage>105</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Recio-García</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>González-Calero</surname>
            ,
            <given-names>P.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Díaz-Agudo</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>jcolibri2: A framework for building Case-based reasoning systems</article-title>
          .
          <source>Sci. Comput. Prog</source>
          .
          <volume>79</volume>
          ,
          <fpage>126</fpage>
          -
          <lpage>145</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Stahl</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roth-Berghofer</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Rapid Prototyping of CBR Applications with the Open Source Tool myCBR</article-title>
          .
          <source>In: Advances in Case-Based Reasoning, 9th European Conf, ECCBR</source>
          <year>2008</year>
          ,
          <article-title>Proceedings</article-title>
          . LNCS, vol.
          <volume>5239</volume>
          , pp.
          <fpage>615</fpage>
          -
          <lpage>629</lpage>
          . Springer (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>