<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>PREDICTIVE ANALYTICS AS AN ESSENTIAL MECHANISM FOR SITUATIONAL AWARENESS AT THE ATLAS PRODUCTION SYSTEM</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>M.A. Titov</string-name>
          <email>mikhail.titov@cern.ch</email>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>M.Y. Gubin</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>A.A. Klimentov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>F.H. Barreiro Megino</string-name>
          <xref ref-type="aff" rid="aff5">5</xref>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>M.S. Borodin</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>D.V. Golubkov</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>on behalf of the ATLAS Collaboration</string-name>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Brookhaven National Laboratory</institution>
          ,
          <addr-line>P.O. Box 5000, Upton, NY, 11973</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Institute for High Energy Physics</institution>
          ,
          <addr-line>1 Ulitsa Pobedy, Protvino, Moscow region, 142280</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>National Research Centre “Kurchatov Institute”</institution>
          ,
          <addr-line>1 Akademika Kurchatova pl., Moscow, 123182</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>National Research Tomsk Polytechnic University</institution>
          ,
          <addr-line>30 Lenin Avenue, Tomsk, 634050</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>University of Iowa</institution>
          ,
          <addr-line>108 Calvin Hall, Iowa City, IA, 52242</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>University of Texas at Arlington</institution>
          ,
          <addr-line>701 South Nedderman Drive, Arlington, TX, 76019</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff6">
          <label>6</label>
          <institution>2017 Mikhail A. Titov, Maksim Y. Gubin, Alexei A. Klimentov, Fernando H. Barreiro Megino</institution>
          ,
          <addr-line>Mikhail S. Borodin, Dmitry V. Golubkov</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <fpage>61</fpage>
      <lpage>67</lpage>
      <abstract>
        <p>The workflow management process should be under control of a specific service that is able to forecast the processing time dynamically according to the status of the processing environment and workflow itself, and to react immediately on any abnormal behavior of the execution process. Such situational awareness analytic service would provide the possibility to monitor the execution process, to detect the source of any malfunction, and to optimize the management process. The stated service for the second generation of the ATLAS Production System (ProdSys2, an automated scheduling system) is based on predictive analytics approach. Its primary goal is to estimate the duration of the data processings (in terms of ProdSys2, it is task and chain of tasks) with possibility for later usage in decision making processes. Machine learning ensemble methods are chosen to estimate completion time (i.e., “Time-To-Complete”, TTC) for every (production) task and chain of tasks, and “abnormal” task processing times would warn about possible failure state of the system. This is the primary phase of the service development that also includes the strategy for its precision enhancement. The first implementation of such analytic service is designed around Task TTC Estimator tool and it provides a comprehensive set of options to adjust the analysis process and possibility to extend its functionality.</p>
      </abstract>
      <kwd-group>
        <kwd>situation awareness</kwd>
        <kwd>predictive analytics</kwd>
        <kwd>production system</kwd>
        <kwd>Apache Spark</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The key point of any control system is its ability to process a large amount of data and to
extract the most valuable set of parameters and corresponding connections between them that would
describe the behavior of the controlled system and forecast its future state with a certain level of
confidence. The complexity of such control systems encapsulates a comprehensive model it is based
on. A theoretical model of situation awareness contains workflow description and corresponding
procedures regarding all above requirements. This model is viewed as “a state of knowledge” [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ].
      </p>
      <p>
        The formal definition of situation awareness is presented as “the perception of the elements
in the environment within a volume of time and space, the comprehension of their meaning and the
projection of their status in the near future” [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. It describes three essential levels / phases in
achieving the situation awareness. The first phase (Level 1) is about perceiving the crucial factors of
the environment (selection process of the most significant parameters to describe the controlled
system’s behavior). The second phase (Level 2) is aimed at understanding the meaning of the
collected data and their possible impact on the whole controlled system. And the third phase (Level
3) gives an estimation of the (near) future state of the controlled system. Each of the phases is based
on the outcome of the previous one, and it is a cyclic process.
      </p>
      <p>
        The original scope of this model of situation awareness was the dynamic human decision
making (in a variety of domains), and its structure is presented at the Figure 1 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. But also, this
model represents a basic strategy to control the behavior of the system and might be used as a
strategy (primary actions) for anomaly detection.
      </p>
      <p>
        Specifically, this schematic model is considered as a workflow for the control service of the
ATLAS Production System (2nd generation, i.e., ProdSys or ProdSys2) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. ProdSys is a part of the
distributed production/analysis and data management workflow in the ATLAS Experiment [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] (see
Figure 2). ProdSys is a distributed task management and automated scheduling system with the
current processing rate of about 2M tasks per year. Its major responsibilities include the following
workflows: the central production of Monte-Carlo data, highly specialized production for physics
groups, as well as data pre-processing and analysis using such facilities as grid infrastructures, clouds
and supercomputers.
      </p>
      <p>
        The system contains such core components / layers as Ref. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]: i) web UI to manage tasks
(that consists of jobs that all run the same program) and production requests (i.e., group of tasks); ii)
DEfT (Database Engine for Tasks) to formulate tasks, chains and groups of tasks; iii) JEDI (Job
Meta-data handling
      </p>
      <p>AMI
pyAMI
Central, physics group
production requests</p>
      <p>Distributed
Data Management</p>
      <p>
        Rucio
Execution and Definition Interface, which is an intelligent component in the PanDA system [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]) to
manage task-level workload with such key feature as dynamic job definition and execution for
resources usage optimization (see Figure 3). The 4th layer is represented by PanDA system
(workload management system) that is responsible for job execution process.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Problem statement</title>
      <p>The possibility of controlling / tracking the state of a particular system requires being aware
of its crucial parameters and processes, and their regular behavior. Such parameters and processes
should be categorized following the potential failures of the controlled system. List of potential
failures: i) overload in the system (that can cause for stuck handling processes); ii) malfunctioning of
computing resources (that can cause failure of data processing); iii) components misconfiguration
(that can cause improper operational processes); iv) general malicious activities (that can cause
incorrect data results).</p>
      <p>The following parameters, which are actual ProdSys parameters or derived ones, reflect the
specific state of the system and indicates a strong correlation with its failures of a particular category
presented above. Controlled parameters are:
 Duration of task (chain of tasks) execution (forecast and control Time-To-Complete);
 Rate of task submission over processing during a limited period of time;
 Cumulative number of failures or/and errors occurred;
 Metrics of resources utilization (which are compared with optimal values).</p>
      <p>These parameters are considered as the initial set that has to be maintained, and should be extended
later to improve the operational process.</p>
      <p>The goal of the designed control system (service) is being able to track the defined set of
parameters of the controlled system (ProdSys), to distinguish their non-valid values according to the
running state of the system, to forecast their next values considering the state of the computing
environment and system itself, and base on corresponding metrics and controlled parameters to detect
the source of malfunctioning. The current project is focused on description of the task TTC
estimation as a part of the control service infrastructure (i.e., as a part of a proof-of-concept).</p>
    </sec>
    <sec id="sec-3">
      <title>3. Technology and methods</title>
      <sec id="sec-3-1">
        <title>Predictive analytics</title>
        <p>Predictive analytics is a form of conjunction of advanced analytics and decision optimization
that uses both new and historical data to forecast future activity, behaviour and trends. Advanced
analytics includes in itself: statistical analysis, machine learning, modeling, data visualization, and
reporting; while decision optimization is represented by: scoring engine, rules engine, and
recommendation engine. Thus, multiple variables are combined into a predictive model that is
capable to assess future probabilities with a certain level of confidence. These predictive models
(modeling) are used to look for correlations between different data elements (features). Predictive
models are presented by two types: i) classification (predict the class membership for the considered
object); ii) regression (predict a number for the specific feature of the object). Three of the most
widely used predictive modeling techniques are decision trees, regression and neural networks.</p>
        <p>These techniques (or a combination of them) is a foundation for the Level 3 of the model of
situation awareness. Predictive analytics should be applied to the parameters which are under the
control, thus, the forecasting of their next value (state) will trigger the control system to follow the
predefined actions.</p>
      </sec>
      <sec id="sec-3-2">
        <title>Processing framework and machine learning methods</title>
        <p>
          Apache Spark [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] parallel processing framework and Spark.MLlib [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] machine learning
library are used as a prediction tools for the control system / service. Apache Spark is a framework
providing speedy and parallel processing of distributed data in real time (sophisticated analytics,
realtime streaming and machine learning on large datasets). It provides powerful cache and persistence
capabilities. Because of these features Apache Spark is applied to defined data analytics problems.
There are two regression ensemble methods (ensemble of decision trees) that were considered as the
initial implementation for the predictive analytics: i) Gradient-Boosted Trees (GBT) regression
method; ii) Random Forests (RF) regression method. Both methods were evaluated with the current
service implementation, and RF outreached GBT in key characteristics such as less prone to
overfitting, marginally better performance of parallel implementations, robust when the training set
contains outliers.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Conducted research</title>
      <p>The initial focus in controlling crucial parameters was set on estimation of the duration of
task (chain of tasks) execution. The ability of being aware of the task execution process assists in
forecasting the near-future state of the controlled system (as a basic representation of the system’s
state). Three core stages for estimation of the task execution duration were highlighted, where each
of the stage is applied at a certain time of the task lifecycle with available set of parameters (and
applied corresponding techniques) and with a certain level of estimation accuracy: i) thresholds
definition (based on statistical analysis); ii) initial / static predictions (based on descriptive data of
task); iii) dynamic predictions (based on static and dynamic data of the running task).</p>
      <p>The general workflow of the designed control system with its main components is presented
at the Figure 4. It shows that the management process is concentrated in the “Control node”, while
DEfT/JEDI database
ShRBD
M
it
w
gaen</p>
      <p>Sqpoo tcxaaehD
data processing is executed in the “Analytics cluster”. (Blocks “Collector” and “Model Handler” are
responsible for the task TTC estimation.)</p>
      <sec id="sec-4-1">
        <title>Task Time-To-Complete by threshold definition</title>
        <p>The set of thresholds of the task execution duration is calculated based on the statistical
analysis of the historical data for the last several months (6 months was used for the conducted
research, see Figure 5). This stage is considered as a rough estimation of the duration of task of a
particular type for the possible and valid values. The type of the task is a combination of its features
such as &lt;projectType&gt; (e.g., “mc” – Monte-Carlo simulated data, “data” – real data collected from
the physical experiment) and &lt;productionStep&gt; (e.g., evgen, simul, recon, merge, etc.).</p>
        <p>At Figure 5, there are plots with corresponding thresholds and columns per productionStep
that are represented by: i) red vertical lines, which correspond to the values of maximum task
execution duration for 95% of tasks that are finished during the defined period of time; ii) blue
vertical lines are the same as red ones, but for 80% of finished tasks; iii) green vertical lines represent
mean values of the task execution duration among all tasks of the defined type; iv) columns of a
particular color represent the number of tasks that were started in a particular month (i.e., number in
the legend corresponds to the number of the month) finished within a certain period of time (i.e., task
duration). Thereby, the threshold for execution duration of 95% of tasks of a particular type was
chosen as a rough estimation (other 5% of tasks were pruned as exceptions, which evaluation should
be done separately), and which defines the range of valid values (i.e., “no warning” values).</p>
      </sec>
      <sec id="sec-4-2">
        <title>Task Time-To-Complete by predictive modeling</title>
        <p>The process of predictive modeling assumes to solve the problem (forecasting the value of a
particular parameter) by creating a model from the sample data (training data based on historical
data), thus take data with known relationships, and using learned relationships to make predictions on
new data (to discover a certain outcome, e.g., numerical value of task execution duration).</p>
        <p>There are several categories of prediction approaches (where sub-categories are named
according to temperature states to represent the time point of the task lifecycle when it is applied):
 Initial / static predictions
o “Cold”-prediction is based on task definition parameters that categorize the
average execution process for the defined task type (with particular conditions).</p>
        <p>It gives the estimation of the task duration during its definition.
 Dynamic predictions
o “Warm”-prediction is based on description and state of scout jobs that are used
to check the processing environment. It gives the prediction for the task duration
immediately after task is launched.
o “Hot”-prediction is based on the current state of the task processing (states of
environment and corresponding jobs). It gives the prediction adjustment during
the task execution (this prediction is calculated multiple times during the task
execution process).</p>
        <p>The example of the static prediction based on Random Forest regression method is shown at
the Figure 6. The following data was used for the prediction process: i) predictive models are based
on “training” data (“mc-all” + “mc-recon” types) for the defined 3 months; ii) “test” data
(“mcrecon” type) for the next month (after training data is collected) was used to generate predicted task
durations and compare it with real values. This result shows that predictions for tasks (execution
duration of 95% of tasks of the defined type is under 27d) are actually following the real values of
task durations (the accuracy for the defined example is presented with the following absolute error
values: mean=1.57d, std=6.67d, RMSE=6.85d). Since the prediction approach of this category does
not consider dynamic parameters and the state of the environment, the accuracy of predictions is low
and can be considered as a rough estimation, same as for thresholds (further research will explore
generated predictions based on different set of parameters for the training data). Predictions of
“dynamic” category will increase the final accuracy (this is under the development).</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Acknowledgements</title>
      <p>We thank ProdSys/PanDA team for providing the data for the conducted analysis and for
their continued support. This work has been carried out using computing resources of the federal
collective usage center Complex for Simulation and Data Processing for Mega-science Facilities at
NRC “Kurchatov Institute”, http://ckp.nrcki.ru/. Also, this work was funded in part by the Russian
Ministry of Science and Education under contract No. 14.Z50.31.0024.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>Techniques and methods of predictive analytics would benefit the control and monitoring
processes for the controlled system. The situational awareness analytic service (based on predictive
analytics techniques) would provide the possibility to detect the source of any malfunction, and to
optimize the whole management process.</p>
      <p>The obtained results facilitate filtering the regular processes of tasks execution (with certain
level of confidence) and form the basic layer for further sophisticated processes of highlighting
abnormal operations and executions. The further improvement of the prediction process will increase
the accuracy of the selection process of potentially failure task (chain of tasks), and the extension of
this approach to other system’s computing blocks and data will keep the system aware of potential
malfunction.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Endsley</surname>
            <given-names>M.R.</given-names>
          </string-name>
          <article-title>Design and evaluation for situation awareness enhancement</article-title>
          .
          <source>// Proceedings of the Human Factors and Ergonomics Society Annual Meeting</source>
          ,
          <year>1988</year>
          , vol.
          <volume>32</volume>
          , no.
          <issue>2</issue>
          , pp.
          <fpage>97</fpage>
          -
          <lpage>101</lpage>
          . DOI:
          <volume>10</volume>
          .1177/154193128803200221
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Endsley</surname>
            <given-names>M.R.</given-names>
          </string-name>
          <string-name>
            <surname>Toward</surname>
          </string-name>
          <article-title>a Theory of Situation Awareness in Dynamic Systems</article-title>
          . // Human Factors Journal,
          <year>1995</year>
          , vol.
          <volume>37</volume>
          no.
          <issue>1</issue>
          , pp.
          <fpage>32</fpage>
          -
          <lpage>64</lpage>
          . DOI:
          <volume>10</volume>
          .1518/001872095779049543
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Borodin</surname>
            <given-names>M.</given-names>
          </string-name>
          et al.
          <article-title>Scaling up ATLAS production system for the LHC Run 2 and beyond: project ProdSys2</article-title>
          . // Journal of Physics: Conference Series,
          <year>2015</year>
          , vol.
          <volume>664</volume>
          , no.
          <issue>6</issue>
          ,
          <issue>062005</issue>
          . DOI:
          <volume>10</volume>
          .1088/
          <fpage>1742</fpage>
          -6596/664/6/062005
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>ATLAS</given-names>
            <surname>Collaboration</surname>
          </string-name>
          .
          <source>The ATLAS Experiment at the CERN Large Hadron Collider // Journal of Instrumentation</source>
          ,
          <year>2008</year>
          , vol.
          <volume>3</volume>
          , no.
          <issue>8</issue>
          ,
          <issue>S08003</issue>
          . DOI:
          <volume>10</volume>
          .1088/
          <fpage>1748</fpage>
          -0221/3/08/S08003
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>De</surname>
            <given-names>K.</given-names>
          </string-name>
          et al.
          <article-title>Task Management in the New ATLAS Production System</article-title>
          . // Journal of Physics: Conference Series,
          <year>2014</year>
          , vol.
          <volume>513</volume>
          , no.
          <issue>3</issue>
          ,
          <issue>032078</issue>
          . DOI:
          <volume>10</volume>
          .1088/
          <fpage>1742</fpage>
          -6596/513/3/032078
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Barreiro</given-names>
            <surname>Megino F.H</surname>
          </string-name>
          . et al.
          <article-title>PanDA: Exascale Federation of Resources for the ATLAS Experiment at the LHC</article-title>
          . // EPJ Web of Conferences,
          <year>2016</year>
          , vol.
          <volume>108</volume>
          , 01001. DOI:
          <volume>10</volume>
          .1051/epjconf/201610801001
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Apache</given-names>
            <surname>Spark</surname>
          </string-name>
          : https://spark.apache.org
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Meng</surname>
            <given-names>X.</given-names>
          </string-name>
          et al.
          <source>MLlib: Machine Learning in Apache Spark. // Journal of Machine Learning Research</source>
          ,
          <year>2016</year>
          , vol.
          <volume>17</volume>
          , no.
          <issue>1</issue>
          , pp.
          <fpage>1235</fpage>
          -
          <lpage>1241</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>