<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Social Mining as a Knowledge Management Solution</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Hasselt University</institution>
          ,
          <addr-line>Martelarenlaan 42, 3500 Hasselt</addr-line>
          ,
          <country country="BE">Belgium</country>
        </aff>
      </contrib-group>
      <fpage>57</fpage>
      <lpage>64</lpage>
      <abstract>
        <p>We introduce knowledge management and identify common problems and challenges. We then introduce social mining, a technique from the eld of process mining. We show that the techniques used in social mining could have added value for the knowledge management challenges. We also identify areas for improvement in social mining. The rst of these concerns the joint activity metric. We argue that for this metric only correlation measures should be used and the distance measures should be re-evaluated. Next we propose the more detailed use of time information to add value to the existing metrics. Lastly we suggest a way of validating the connection in a social network. As a conclusion we nd that social mining can be enhanced in these areas and then used as a tool to tackle knowledge management problems.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Knowledge management is the process of capturing, distributing, and e
ectively using knowledge(Davenport,1994)[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. A more formal de nition is given
by Duhon[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]:
"Knowledge management is a discipline that promotes an integrated
approach to identifying, capturing, evaluating, retrieving, and sharing all
of an enterprise's information assets. These assets may include databases,
documents, policies, procedures, and previously un-captured expertise
and experience in individual workers."
      </p>
      <p>We focus on the second part of this de nition, meaning we consider the
expertise and experience in individual workers. The purpose of this paper is to
introduce some of the classical challenges in knowledge management and identify
how process mining might be used to provide solutions.</p>
      <p>Process mining is the bridge between data mining and process modelling.
Starting from an event log we use data mining on this log to get process
information. The main focus of process mining is process discovery, leading to an
increasing amount of discovery algorithms. We can however also use it to nd
social networks, this is called social mining. Social mining is what we will use
to establish the link with knowledge management. For this paper we use the
Copyright c by the paper's authors. Copying permitted only for private and academic
purposes.
term social mining. If there is a chance for confusion with the practice of mining
information o of social media, we suggest the term social process mining.</p>
      <p>However, limitations of the current techniques are preventing a direct
application of the social mining algorithms in the context of knowledge management.
We discuss the potential areas of improvement to leverage the use of social
mining for knowledge management challenges.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Knowledge management challenges</title>
      <p>
        In knowledge management there are a few classic challenges[
        <xref ref-type="bibr" rid="ref4 ref9">9,4</xref>
        ]. We identify
the ones that are interesting and later show techniques from social mining that
might be applied to them.
      </p>
      <p>A common activity in knowledge management is the generation of a
knowledge map that shows where knowledge is located in the company. This might
contain information about knowledge repositories but also about who knows
what. This can be constructed using surveys but it's hard to keep such a map
up to date. People might change what they work on or their contact information
might change.</p>
      <p>Another topic is team selection and the selection of the leader for such a
team. For the latter activity, we require an up to date list of the skills of our
employees. It might also be interesting to know who has worked together before
or which employees already communicate often. To select a leader we need to
nd a person that is central in both the trust and the advice network. If both
these networks are up to date it will be easier to nd a leader.</p>
      <p>Part of knowledge management is the creation of new ideas. An often used
technique for this is fusion. This is a concept where people with a di erent
background (but with some common ground) are put together in a room in a
bid to generate new ideas. It's based on the idea that innovation occurs at the
boundaries between mind-sets, not within. Again it appears to be important to
know what skills people have and what they have recently been working on.</p>
      <p>A big problem in knowledge management is brain drain. This is a
phenomenon that appears when valuable knowledge leaves the company. The main
solution for this problem as identi ed by knowledge management literature is
making sure there is enough transfer of knowledge. This both includes
appointing an experienced mentor to a novice entering the company and transcribing
or codifying the knowledge of people leaving the company. Sharing experience
is another aspect of this. If everyone shares their experiences this means that
knowledge will propagate throughout the company.
3</p>
    </sec>
    <sec id="sec-3">
      <title>The details of social mining</title>
      <p>In general process mining focusses on the discovery of a process model for a
given event log. Another branch of process mining uses the event log to create
social networks. This is called social mining. Since this concerns the creation of
social networks and the analysis of people and the activities they perform, we
argue it can be applied to knowledge management.</p>
      <p>In process mining an event log is a log listing the events that have occurred.
Here, each event is an atomic registration of an action in a process and contains
at least a timestamp and a link to a case.</p>
      <p>For the purpose of extracting a social network, a resource eld for each
activity is also mandatory. This resource eld indicates the person (or machine)
that undertook the activity. Using the resource eld we can construct a social
network based on the information in the log. In this paper we assume the
resource eld always indicates a unique employee. The terms resource, person and
employee are used interchangeably.</p>
      <p>
        A social network is a graph in which each vertex represents a person. The
edges between these vertices then show the connections between these people.
An example of such a social graph can be seen in Figure 1. There are multiple
ways to create these social graphs based on bibliographical information, surveys
or even mailing data. [
        <xref ref-type="bibr" rid="ref7 ref8">7,8</xref>
        ]
Case Activity Resource
1 Register Bob
1 Negotiate Sally
1 Sign Dan
1 Sendo Mary
2 Register Bob
2 Negotiate Sally
2 Sign Dan
2 Negotiate Sally
2 Sendo Mary
      </p>
      <p>The social network in Figure 1 is created by applying the handover of work
technique to the log found in Table 1. This technique is a part of the possible
causality metric of social mining.</p>
      <p>
        Social mining was introduced by Song and Van Der Aalst[
        <xref ref-type="bibr" rid="ref1 ref10 ref11">1,10,11</xref>
        ]. They use
the resource eld and other information from the log to determine if there should
be a connection between a pair of resources or not. There are di erent ways to
calculate these connection, they are grouped in the following four metrics.
1. Possible causality (containing handover of work)
2. Joint cases
3. Joint activities
4. Special event type
      </p>
      <p>The rst metric assumes that when people handover cases to one another,
they have a connection. This means that following events lead to a connection
between these resources in the social graph. Part of this metric is the handover
of work metric which counts immediate handovers between people. Another
example is subcontracting, which counts cases that get handed from person A to
person B and then get handed back to person A. In a knowledge management
context, this metric can be used to identify the ow of information between
resources.</p>
      <p>The second metric establishes a connection between two resources if they
work on the same case.It is harder to establish relevance for this metric in the
described knowledge management problems. However if we assume that a case
corresponds to a project, this metric can be used to nd out who worked on
projects together. That information could then be used when we are selecting a
team or a team leader.</p>
      <p>The third metric yields a connection if two resources perform the same
activities. Typically if a company is divided in functional departments this is the
case.</p>
      <p>This metric is promising for a lot of the knowledge management problems.
This metric can be used to create pro les that relate people with the activities
they perform, identifying the skills of a person. Knowing these skills can help,
not only in team selection, but also when we are generating a knowledge map.
It can also identify the risk of a brain drain if it is know that an activity is
only performed by one employee. These pro les can also be used when trying to
generate new ideas. Employees can be selected with both di erent skills and a
few common skills to simulate fusion.</p>
      <p>The last metric relies on special event types. Imagine an event that represents
the activity of assigning a task to someone else. In this case it is known that the
resource gives a task to someone else, so there is a connection between the two.
Since it's so speci c this metric does not have a clear knowledge management
application.</p>
      <p>
        The eld of social mining dates back to 2004, having been originally
introduced by Van Der Aalst[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].However no recent work builds on it. The last paper
we could nd that uses process mining as a tool of creating a social network
stems from 2008[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. We think this proves a chance to pick up an area of process
mining that is now less used.
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>Social mining's areas of improvement</title>
      <p>In the previous section we have given the details of social mining linked to
the knowledge management challenges. There are opportunities for added value,
however the current state of social mining still shows some room for
improvement. There are a few areas which we rst wish to improve before applying social
mining to knowledge management.</p>
      <p>We identify three main areas for improvement. The rst area is of a smaller
scale, concerning the distance measures used in the joint activity metric. As a
next improvement we suggest that added value can be created by taking
timestamps into account. A last improvement opportunity is the validation of the
connections in the social network. We suggest using email tra c as a way of
verifying the connection between two people. We will now look at these possible
improvements in more detail.
4.1</p>
      <sec id="sec-4-1">
        <title>Joint activity metric</title>
        <p>
          The joint activity metric assumes that there is a relation between people who
perform the same activities. To calculate this metric, a performer-activity matrix
is created. In this matrix every row corresponds to a resource and every column
to an activity. The cells of the matrix represent the number of times the resource
performs the activity.An example of such a matrix, related to the log of Table
1, can be found in Figure 2. To determine if there is a connection between two
resources, the distances between the rows in the matrix are calculated. To do
this calculation a few distance measures and a correlation measure are proposed
[
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. The result of these measures is then used to determine if there should be a
connection between the resources corresponding to these rows.
        </p>
        <p>Bob
Sally
Dan
Mary</p>
        <p>
          The measures proposed (based on the implementation of the metrics in ProM
6 [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] a process mining toolkit) are euclidean distance, Pearson correlation coe
cient, Hamming distance and a similarity coe cient. However the original paper
also mentions Minkowksi distance. The main improvement that can be made
here is to add clarity for the end user. Using the current implementation, mining
the same log with a di erent measure can lead to opposite results. An example
of this can be found in Figure 3. Here we show the social networks mined from
the same log, once using euclidean distance and once using the Pearson
correlation coe cient. Using the euclidean distance connects every one in the network
except for Dan and Sally whereas using the Pearson correlation coe cient only
connects Dan and Sally.
        </p>
        <p>
          The problem with using a mix of correlation and distance measures is that
while a correlation measure yields a result in [ 1; 1] a distance measure yields a
result in the range [0; 1[. A threshold needs to be determined such that values
above( or below) this threshold signify a connection. When using the Pearson
correlation coe cient, a threshold value of 0.0 is suggested, but no threshold
values are given for the distance measures [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ].
        </p>
        <p>Establishing a threshold value for the distance measures proves to be more
di cult. Consider for example the Hamming and the euclidean distance. The
(a) the social network using euclidean dis-(b) The social network using the Pearson
tance
correlation coe cient
Hamming distance has a maximum that corresponds with the number of
activities in the matrix. The euclidean distance however has a maximum depending
on the amount of times activities are preformed. When applying these measure
to the example from Figure 2, we can use these metrics to nd the distance
between Bob and Sally. The Hamming distance between them is 2 and the
euclidean distance between them is p13 3:6. Assume now that instead of
3, Sally performs the negotiate activity 10000 times. Using the Hamming
distance as currently implemented in ProM, every non-zero value of the matrix is
rescaled to one. This means the Hamming distance between both of them stays
unchanged. However with the new values the euclidean distance between them
becomes p(100000004) = 26p(147929) 10000. This makes it very clear that
both distance measures operate on a di erent scale.</p>
        <p>Our suggestion is to drop the use of distance measures and instead focus
on correlation measures. We should look at the di erent correlation measures to
determine which ones are most appropriate.It is not guaranteed that the Pearson
correlation coe cient is the most appropriate one. Using these measures there is
already a possible threshold value at 0:0 and there is no danger of them operating
on a di erent scale since they stay between 1 and 1. It will also make it easier
for the end user to identify what measure to use since they will have to choose
between the same sort of measures.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Timestamps</title>
        <p>The original approach only uses timestamp information to order the activities.
We believe the addition of more timestamp related information can add value to
the results obtained. The addition of timestamps can be considered an
improvement of a di erent scale since it a ects more than one metric. There are a few
reasons why the addition of this information to the metrics might add value.</p>
        <p>We can assume that people who are performing the same activity at the same
time have a better chance of knowing each other than people performing the
same activity at another time of day. Similarly we can assume that when there
is an instant follow up, that this indicates a stronger connection. This means
the handover of an activity from person A to person B with only 5 minutes
in between corresponds to a stronger connection than the same handover with
hours or even days in between.</p>
        <p>Time related information is also interesting when it comes to joint activities
or cases. It could be used to nd out when a resource last performed an activity
or when they were last in a case with another resource. This is important when
trying to determine the skills of a person. A person might have done activity
A a thousand times, but never in the last ve years. Whereas another person
might have done the same activity only ve hundred times, but more recently.
The latter person will probably be able to answer our questions about activity
A more rapidly than the former.</p>
        <p>We thus propose to examine all existing metrics and develop expanded
metrics making use of the time related information. This information can then be
used to add weight to the connections in the social network or to determine what
happened more recently and thus is more important.
4.3</p>
      </sec>
      <sec id="sec-4-3">
        <title>Validation of connections</title>
        <p>The third area where we see room for improvement relates to the validation of
connections in our social network. The di erent networks (created using the
different metrics) are successful at showing the ow of information and the grouping
of resources according to activities. It would be interesting to be able to
conclude that there is a certain ow of communication based on the information
in these networks. Right now this is not always the case. Consider for example
two people who perform sequential activities, meaning a connection would be
established when using the possible causality metric, but that are grouped in
di erent departments and don't know one another.</p>
        <p>
          Our proposal here is to enrich the social networks with information from the
email system. Mining the emails sent in a company we can see actual
communication between people [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ][
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. This allows us to validate connections in our social
network by checking if there is email activity between the connected people.
5
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>We started o by introducing knowledge management and some challenges from
the domain. The main challenges are the brain drain, the creation of knowledge
maps and team selection. We then proposed social mining as a possible way to
add value. We did this by liking the social mining metrics with the knowledge
management challenges.</p>
      <p>We suggest some areas of improvement for social mining to be addressed
before applying social mining to knowledge management. There are three areas
of improvement. The rst addresses the distance and correlation measures used in
the joint activity metric. Here we suggest to drop the distance measures and focus
on correlation metrics. A second suggestion concerns timestamp information.
This information is only used to order the events. We see chances to add value by
using more details of this information. A last improvement lies in the validation
of connections in the social network. Here we propose the use of existing email
mining algorithms to validate the relation between connections in the social
network and communication.</p>
      <p>Brining these results together we have identi ed possible value improvement
for knowledge management by using social mining techniques. We have also
identi ed some steps to be taken rst to improve the existing social mining
techniques.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Van der Aalst</surname>
            ,
            <given-names>W.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Song</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Mining Social Networks: Uncovering interaction patterns in business processes</article-title>
          .
          <source>In: Business Process Management</source>
          , pp.
          <volume>244</volume>
          {
          <fpage>260</fpage>
          . Springer (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Culotta</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bekkerman</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McCallum</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Extracting social networks and contact information from email and the web (</article-title>
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Davenport</surname>
          </string-name>
          , T.H.:
          <article-title>Saving IT's Soul: Human-Centered Information Management</article-title>
          .
          <source>Harvard business review 72(2)</source>
          ,
          <volume>119</volume>
          {
          <fpage>31</fpage>
          (
          <year>1994</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Davenport</surname>
            ,
            <given-names>T.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prusak</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Working knowledge: How organizations manage what they know</article-title>
          . Harvard Business Press (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Duhon</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>It's all in our heads</article-title>
          .
          <source>Inform</source>
          <volume>12</volume>
          (
          <issue>8</issue>
          ),
          <volume>8</volume>
          {
          <fpage>13</fpage>
          (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Farnham</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Portnoy</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Turski</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Using email mailing lists to approximate and explore corporate social networks</article-title>
          .
          <source>In: Proceedings of the CSCW</source>
          . vol.
          <volume>4</volume>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Fisher</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dourish</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Social and temporal structures in everyday collaboration</article-title>
          .
          <source>In: Proceedings of the SIGCHI conference on Human factors in computing systems</source>
          . pp.
          <volume>551</volume>
          {
          <fpage>558</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Kazienko</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brdka</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Musial</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gaworecki</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>Multi-layered social network creation based on bibliographic data</article-title>
          .
          <source>In: Social Computing (SocialCom)</source>
          ,
          <year>2010</year>
          IEEE Second International Conference on. pp.
          <volume>407</volume>
          {
          <fpage>412</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Koenig</surname>
            ,
            <given-names>M.E.</given-names>
          </string-name>
          :
          <article-title>What is KM? Knowledge Management Explained</article-title>
          (May
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Song</surname>
          </string-name>
          , M.,
          <string-name>
            <surname>van der Aalst</surname>
            ,
            <given-names>W.M.P.</given-names>
          </string-name>
          :
          <article-title>Towards comprehensive support for organizational mining</article-title>
          .
          <source>Decision Support Systems</source>
          <volume>46</volume>
          (
          <issue>1</issue>
          ),
          <volume>300</volume>
          {317 (Dec
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Van Der Aalst</surname>
            ,
            <given-names>W.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reijers</surname>
            ,
            <given-names>H.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Song</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Discovering social networks from event logs</article-title>
          .
          <source>Computer Supported Cooperative Work (CSCW) 14(6)</source>
          ,
          <volume>549</volume>
          {
          <fpage>593</fpage>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Verbeek</surname>
            ,
            <given-names>H.M.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buijs</surname>
          </string-name>
          , J.,
          <string-name>
            <surname>Van Dongen</surname>
            ,
            <given-names>B.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>van der Aalst</surname>
          </string-name>
          , W.M.
          <article-title>: Prom 6: The process mining toolkit</article-title>
          .
          <source>Proc. of BPM Demonstration Track</source>
          <volume>615</volume>
          ,
          <issue>34</issue>
          {
          <fpage>39</fpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>