<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Process Mining on Video Data⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Arvid Lepsien</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jan Bosselmann</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andreas Melfsen</string-name>
          <email>amelfsen@ilv.uni-kiel.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Agnes Koschmider</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Group Process Analytics, Computer Science Department Kiel University</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Institute of Agricultural Engineering Kiel University</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <fpage>56</fpage>
      <lpage>62</lpage>
      <abstract>
        <p>Disciplines like life and natural sciences could gain high benefits from process mining in terms of identifying anomalies in the process or supporting predictive analytics in what is being measured. These disciplines, however, mostly work with data at a much lower level of abstraction and the data does not directly relate to high-level business process concepts as required for process mining. This paper discusses an approach for process mining on video data. As a use case, we applied our approach on video surveillance data of pigpens. Although, our process analytics pipeline from raw video data to a discovered process model has not yet been fully implemented, we are convinced that our approach is an essential contribution towards a (semi)automatic technique aiming to replace manual work.</p>
      </abstract>
      <kwd-group>
        <kwd>process mining</kwd>
        <kwd>activity recognition</kwd>
        <kwd>video labeling</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Process mining is an established technique to give insights into data in terms of a
structured order of activities (i.e., a process model). In this way, process mining
allows identifying bottlenecks or compliance issues in business events. Mainly,
process mining relies on business event data that is used as input to process
mining algorithms and thus the data is expected to be on a high (business)
abstraction level. Despite the success of process mining in the business context,
process mining can provide an additional benefit to disciplines dealing with high
volume and veracity of data. These disciplines like life or natural science have a
high demand for a structured approach to answer process related questions like
(1) what unknown processes are acting (i.e., did we find all processes that exist)
and (2) whether the found processes actually work as thought.</p>
      <p>
        Previously, we suggested approaches to discover process models from sensor
event data [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and ”raw” time series data [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] with the purpose to give new insights
into the data in terms of the identification of anomalies in the process flow aiming
to prevent unintended consequences. This paper presents our approach to discover
process models from video data. As a use case, we applied our approach on video
surveillance data of pigpens. So far, the behavior of pigs has been studied manually.
Therefore, our approach aims to provide a (semi)automatic approach for pig
behavior analysis in terms of health monitoring and understanding animal welfare.
In this way, our approach makes a contribution to both questions (1) and (2)
mentioned above.
      </p>
      <p>The next section motivates why the use case is an appropriate starting point
to develop techniques for process mining on video data.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Challenges</title>
      <p>Compared to low-level raw data like sensor event data and time series used as
input for activity recognition, video surveillance data of pigpens on the one hand
eases the extraction of process activities, but on the other hand several challenges
as mentioned below have to be overcome. Reasons facilitating the analysis are:
(1) the behavior of pigs is limited to a few activities, which significantly simplifies
activity detection compared to recognition of human activities in smart homes or
smart factories. (2) A distinction between individual pigs is not necessary. This
significantly simplifies the entity-centricity, which is challenging in smart homes
where usually multiple objects are moving that need to be distinguished from
each other.</p>
      <p>
        To apply process mining on video data, however, requires bridging the following
challenges: (1) no appropriate reference data set and labeled data exist. The
freely accessible video-based data sets are mainly for object detection of other use
cases like autonomous driving. Large computer vision libraries like Facebook AI
Research’s Detectron2 [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] grant access to trained neural networks, however, the
detection of pigs is not covered by the commonly used COCO (Common Objects
in Context) [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and ImageNet [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] datasets. We found two pig-specific data sets
for detecting positions and orientation [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] and tracking [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], but these data sets
do not suit process discovery purposes. Almost no process-specific data exist in
the data set. Therefore a high manual efort is required since neither labeled
data nor an appropriate data analysis pipeline exist for our use case. (2) Image
quality significantly correlates with the analysis results. Image quality is afected
by image resolution, camera angle and camera quality. We initially received a
data set of very low quality. In addition, the data set was not representative (i.e.,
too short image sequences). Therefore, recording of a new data set was necessary
with a camera installation from a diferent angle. (3) Image noise (e.g., due to
randomly switching from day to night mode, camera pollution and distortion
due to neighboring pigpens). Finally, we recorded a new representative data set
of higher image quality and less image noise.
      </p>
      <p>The next section presents our approach aiming to address the challenges
mentioned above.</p>
    </sec>
    <sec id="sec-3">
      <title>Approach</title>
      <p>– extract related video data from original data set: we observed four pigpens,
each with ten to twelve pigs, over a period of a few weeks. We recorded video
material with a resolution of 1920x1080 pixels and 12.5 fps every day from
6:00 a.m. to 6:00 p.m. Mostly, the pig behavior does not change. Instead the
pigs are in a kind of dormant phase. Many interesting actions only take place
over a very short period of time, sometimes lasting just a few seconds. To
detect related actions in our data set, we developed an algorithm measuring
the movement intensity of a video sequence, which makes it easy to recognize
the active phases of the pigs (see Figure 2). A spike in the chart indicates a
new action.
– mine domain-specific knowledge : in this step we aim to identify context-related
information that enhances action and object recognition. For instance, the
location of the movement areas varies from groups of pigs. A group of young
pigs would divide the pigpen diferently than a group of older pigs. Thus,
context information in terms of pig specificity is necessary in order to not
distort the analysis results. Although, multiple data mining techniques have
been used to mine domain-specific knowledge, again no specific technique
exists for our use case. Therefore, the techniques have been tailored to our use
case. First, we aimed to identify areas of high (visual) actions. The algorithm
presented before has been enhanced to identify active movement areas. In
general, a pigpen is divided into these three areas: sleeping/resting area,
defecation area and feeding area. To automatically detect these areas, we
used a slightly modified version of our motion intensity detection algorithm.
We divided the images of a video into an area of 20x20 tiles and calculated
the intensity of each tile over the entire video. Next, we converted the results
into a 20x20 heatmap and easily identified the active areas. Figure 3 shows</p>
      <p>an example. Then, knowledge of the positions of all pigs over time is used
to create a heatmap of common pig positions. To do this, we calculate the
midpoint of each bounding box detected on the video. The position heatmap
is then constructed from the relative frequency of midpoints per heatmap
bin (see Figure 4 for a log-normalized example output of this analysis).
Tracking traces have been clustered to find common movement patterns (and
paths between common areas). Figure 5 shows an example of 150 movement
trajectories extracted from one video. Diferent movement patterns can be
observed, e.g. the pigs are mostly stationary in their resting area.
(a) Original movement intensity heatmap of (b) Smoothed movement intensity heatmap
a video divided into 20x20 rectangles. of the original version.</p>
      <p>Fig. 3: Example of our algorithm to explore the movement intensity of areas in
the video.</p>
      <p>– object recognition: in this step, the video data is prepared for further analysis.</p>
      <p>
        First, an object detection is applied on the video. We chose YOLOv5 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] due
to its ease of use and ability to produce appropriate results with a relatively
small amount of hand-labeled training data. Based on the object detection,
(multi) object tracking has been used. This allows to analyze the same pig
over multiple frames. We chose the DeepSORT (Deep Simple and Online
Realtime Tracking) [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] algorithm to implement the tracking. DeepSORT is a
well-established algorithm. The algorithm has been shown to work in a similar
context to ours [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and performs reasonably well on our data set without
additional training. In the future, improved solutions for object detection and
tracking could be applied to improve the quality of tracking results. However,
many other solutions for the multiple object tracking problem require labeled
tracking data for training. Since we aim to reduce manual labeling efort,
the implementation of other tracking algorithms should be in proportion to
manual efort. While a tracking dataset is available for pigs [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], it does not
match our camera setup exactly. Also, there is no any labeled tracking data
available when applying the analysis process in a diferent domain. If it was
on purpose, the tracking results could be even used to localize individual pigs
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
– recognize activities in video: The prepared video sequences and the associated
position data from the tracking can be used as input for activity detection.
In this step, also a model to learn visual features could be used. The learning
process would have to be designed in a way where the features correspond
to low-level events of the underlying process of the video. These low-level
events can then be used to create event logs. While several techniques for
pig activity recognition exist [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], they are either very specific to the unique
properties of pigs or very specific to one type of activity (i.e. lying, standing,
aggression). We choose not to use pig-specific techniques in this step to keep
the approach generic.
– discover process model: the activities from the last step need to be
aggregated/abstracted and enhanced with domain-specific knowledge (see step 2).
Then, a case ID has to be created, e.g., according to the movement areas. A
process model can then be mined from the event log.
– refinement : use the quality of the resulting process model to optimize the
activity recognition and process model discovery.
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>Summary and Outlook</title>
      <p>
        In studies of agricultural science alterations in behavior processes of pigs can be
a helpful tool for analyzing and evaluating animal behavior, animal health and
environmental impact. However, most approaches on identifying pig behavior
based on video data only focus on single activities like e.g., lying, eating without
analyzing the process. This paper suggested a process mining-pipeline to extract
a process model from video data. As a use case, we applied our approach on video
surveillance data of pigpens. Beside animal health, welfare and thermal comfort
state, our approach can be used as a helpful indicator to evaluate and adjust
climate conditions in mechanically ventilated barns. Likewise observations of
activity and feed intake, which will vary depending on diferent climate conditions,
supports the control of the above [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. We see further use cases for our approach in
medicine and material science that also handle large volume and veracity of data.
Our approach of process mining on video data might be in medicine and material
science for predictive analytics and outlier detection, which we believe to be
more challenging than the current use case. Both assumptions that facilitate the
analysis (i.e., low number of activities and entity-centricity) need to be bridged
for an eficient solution.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Cowton</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kyriazakis</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bacardit</surname>
          </string-name>
          , J.:
          <article-title>Automated individual pig localisation, tracking and behaviour metric extraction using deep learning</article-title>
          .
          <source>IEEE Access 7</source>
          ,
          <fpage>108049</fpage>
          -
          <lpage>108060</lpage>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Deng</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dong</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Socher</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>L.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fei-Fei</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Imagenet: A large-scale hierarchical image database</article-title>
          .
          <source>In: 2009 IEEE conference on computer vision and pattern recognition</source>
          . pp.
          <fpage>248</fpage>
          -
          <lpage>255</lpage>
          . Ieee (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Janssen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mannhardt</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koschmider</surname>
            , A., van Zelst,
            <given-names>S.J.</given-names>
          </string-name>
          :
          <article-title>Process model discovery from sensor event data</article-title>
          .
          <source>In: Process Mining Workshops - ICPM 2020 International Workshops, Padua, Italy, October 5-8</source>
          ,
          <year>2020</year>
          . LNBIP, vol.
          <volume>406</volume>
          , pp.
          <fpage>69</fpage>
          -
          <lpage>81</lpage>
          . Springer (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Jocher</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stoken</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chaurasia</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Borovec</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <source>NanoCode012</source>
          , TaoXie, et al.:
          <source>ultralytics/yolov5: v6</source>
          .
          <fpage>0</fpage>
          -
          <lpage>YOLOv5n</lpage>
          '
          <article-title>Nano' models, Roboflow integration, TensorFlow export</article-title>
          ,
          <source>OpenCV DNN support (Oct</source>
          <year>2021</year>
          ). https://doi.org/10.5281/zenodo.5563715, https://doi.org/10.5281/zenodo.5563715
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maire</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Belongie</surname>
            ,
            <given-names>S.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bourdev</surname>
            ,
            <given-names>L.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Girshick</surname>
            ,
            <given-names>R.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hays</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , et al.:
          <article-title>Microsoft COCO: common objects in context</article-title>
          .
          <source>CoRR abs/1405</source>
          .0312 (
          <year>2014</year>
          ), http://arxiv.org/abs/1405.0312
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Nasirahmadi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hensel</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Edwards</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sturm</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>A new approach for categorizing pig lying behaviour based on a delaunay triangulation method</article-title>
          .
          <source>Animal</source>
          <volume>11</volume>
          (
          <issue>1</issue>
          ),
          <fpage>131</fpage>
          -
          <lpage>139</lpage>
          (
          <year>2017</year>
          ). https://doi.org/https://doi.org/10.1017/S1751731116001208
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Psota</surname>
            ,
            <given-names>E.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mittek</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>P</surname>
          </string-name>
          ´erez,
          <string-name>
            <given-names>L.C.</given-names>
            ,
            <surname>Schmidt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Mote</surname>
          </string-name>
          ,
          <string-name>
            <surname>B.:</surname>
          </string-name>
          <article-title>Multi-pig part detection and association with a fully-convolutional network</article-title>
          .
          <source>Sensors</source>
          <volume>19</volume>
          (
          <issue>4</issue>
          ),
          <volume>852</volume>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Psota</surname>
            ,
            <given-names>E.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schmidt</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mote</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , P´erez, L.C.
          <article-title>: Long-term tracking of grouphoused livestock using keypoint detection and map estimation for individual animal identification</article-title>
          .
          <source>Sensors</source>
          <volume>20</volume>
          (
          <issue>13</issue>
          ),
          <volume>3670</volume>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Wojke</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bewley</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paulus</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Simple online and realtime tracking with a deep association metric</article-title>
          .
          <source>In: 2017 IEEE international conference on image processing (ICIP)</source>
          . pp.
          <fpage>3645</fpage>
          -
          <lpage>3649</lpage>
          . IEEE (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kirillov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Massa</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lo</surname>
          </string-name>
          , W.Y.,
          <string-name>
            <surname>Girshick</surname>
          </string-name>
          , R.:
          <source>Detectron2</source>
          . https:// github.com/facebookresearch/detectron2 (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xiao</surname>
            ,
            <given-names>D.:</given-names>
          </string-name>
          <article-title>A review of video-based pig behavior recognition</article-title>
          .
          <source>Applied Animal Behaviour Science</source>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Ziolkowski</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koschmider</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schubert</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Renz</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Process mining for time series data</article-title>
          .
          <source>Technical report</source>
          (
          <year>2022</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>