<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Inferring Dependencies Among Web Services with Predictive and Statistical Analysis of System Logs</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ethem Utku Aktas</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mehmet Cagri Calpur</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Umit Ulkem Yildirim</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Emrah Y ld r m</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Sabanci University Orta Mahalle</institution>
          ,
          <addr-line>34956 Tuzla, Istanbul</addr-line>
          ,
          <country country="TR">Turkey</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Softtech A.S. Research and Development Center Tuzla Piri Reis Cad.</institution>
          <addr-line>62, 34947 Istanbul</addr-line>
          ,
          <country country="TR">Turkey</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Software system behaviour analysis is a challenging research problem in software engineering. The main reason for this is the lack of real data from large industrial systems. Softtech Inc. is a subsidiary of a large private bank in Turkey and this study is aimed to analyse mapping the services architecture and the software system health of a particular department at Softtech using speci c software web service logs. The services that are the subject of this study consist of 196 web services related to credit and credit card application transactions from various channels. While these processes are related to similar applications, they call various web services that perform di erent operations in the background. Related services account for 2 million logs daily. We have conducted empirical and statistical analysis on the data, in order to infer the correlations and dependencies among the observed web services. Hypothetically, we think there are 3 types of dependencies between the web services. In our experiments, we used average response times and the number of times web services are called at speci c time intervals as input data. The results suggest that they can be used for inferring that there is a dependency between two web services. In this preliminary work for dependency inference from unstructured web services' log data, we have utilized simple statistical analysis tools to derive important insight about the collection of services under our observation. The results have encouraged us to carry on with a more detailed analysis approach to further advance our research e orts.</p>
      </abstract>
      <kwd-group>
        <kwd>Statistical Analysis</kwd>
        <kwd>Predictive Analytics</kwd>
        <kwd>Dependency In- ference</kwd>
        <kwd>Log Analysis</kwd>
        <kwd>Software Service Management</kwd>
        <kwd>Software Relia- bility</kwd>
        <kwd>Software Architecture</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Ethem Utku Aktas1;2, Mehmet Cagr Calpur1;2, Umit Ulkem Y ld r m1, and
Emrah Y ld r m1</p>
      <p>1 Softtech A.S.</p>
      <p>Research and Development Center
Tuzla Piri Reis Cad. 62, 34947 Istanbul, Turkey</p>
      <p>2 Sabanci University</p>
      <p>Orta Mahalle, 34956 Tuzla, Istanbul, Turkey
futku.aktas, cagri.calpur, umit.yildirim, emrah.yildirimg@softtech.com.tr
O zet. Yaz l m davran s analizi onemli yaz l m muhendisligi konular
ndand r ve endustriyel sistemlerden elde edilen gercek veri eksikligi nedeniyle
incelenmesi zor bir problemdir. Bu cal smada Turkiye'nin en buyuk ozel
bankalar ndan I_s Bankas 'n n yaz l m gelistirme ve bilgi islem alanlar nda
hizmet veren istiraki Softtech A.S.'nin bir departman n n gelistirdigi ve
yonettigi web servislerinin loglar ile web servisleri aras bag ml l k analizi
ve servis mimarisi c kar m yapmay hede iyoruz. Cal smam za konu olan
servisler cesitli kanallardan gelen kredi ve kredi kart basvuru islemleri
ile ilgili 196 web servisinden olusmaktad r. Bu islemler benzer
uygulamalarla ilgili olmakla birlikte, arka planda farkl islemler yapan cesitli
web servisleri cag rmaktad rlar. I_lgili servisler gunluk 2 milyon kutuk
kayd olusturmaktad r. Elde ettigimiz veriyi ureten servisler aras ndaki
korelasyon iliskisi c kar mlar nda bulunabilmek icin ampirik ve
istatisliksel analiz yontemleri kullan lm st r. Varsay msal olarak, web
servisleri aras nda 3 tur bag ml l k oldugunu dusunuyoruz. Deneylerimizde,
belli zaman aral klar ndaki ortalama yan t sureleri ve web servislerin
kac kere cagr ld klar girdi olarak kullan lm st r. Sonuclar bu iki bilginin,
web servisleri aras nda bir bag ml l k oldugu sonucuna varmak icin
kullan labilecegi ongorusunu ortaya koymaktad r. Bu oncu cal smam zda
istatistiksel analiz yontemleri kullanarak toplad g m z web servisi
verisinden degerli bilgiler c karmay basard k, bu sonuclar bizi arast rmam z
daha ileri goturmek konusunda da tesvik edicidir.</p>
      <p>Anahtar Kelimeler: I_statistiksel Analiz Tahminleme Analitigi Bag ml l k
C kar m Kutuk Analizi Yaz l m Servis Yonetimi Yaz l m Guvenilirligi</p>
      <p>Yaz l m Mimarisi</p>
    </sec>
    <sec id="sec-2">
      <title>Introduction</title>
      <p>
        Software systems evolve and as they evolve, their complexity increases. As a
result it gets harder to monitor them. However, reliability of such systems is
a key factor for customer satisfaction. One of the fundamental assumptions in
monitoring such complex software systems is that there exist patterns in the
behavior of executions. Many previous studies support this assumption [
        <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4 ref5 ref6 ref7 ref8">1-8</xref>
        ].
      </p>
      <p>Recent trends in software architecture such as microservices, magni es the
problem of maintaining large number of services. The maintenance problem is
really complex for large companies that rely on hundreds or even thousands
of web services. Documentation and records for ownership of the architecture
may be neglected. In this study, we propose web service dependency inference
techniques to map large scale software systems.</p>
      <p>
        One of the problems in studying this research area is the lack of real data from
such complex systems. Execution data such as system logs need to be collected
and analysed in these systems and we conjecture that it is possible to nd out
dependency relations between services by observing the log data [
        <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4 ref5 ref6 ref7 ref8">1-8</xref>
        ].
      </p>
      <p>In this preliminary study, our aim is to conduct a statistical analysis using
system logs of speci c web services and nd dependency relations among these
services. First, the architecture of the system under observation is presented.
Then, structure of the log data and the procedure to collect and preprocess this
data is explained. The methodology for the analysis of the data and the results
are given next. Finally, comments and conclusions on the study are presented.
2</p>
    </sec>
    <sec id="sec-3">
      <title>Related</title>
    </sec>
    <sec id="sec-4">
      <title>Work</title>
      <p>
        Lin and Siewiorek collected error log data from 13 le servers for a 22 month
period [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. They conducted a trend analysis and proposed a heuristic method to
predict failures.
      </p>
      <p>
        Hellerstein, Zhang and Shahabuddin proposed a predictive approach to
detect the probability of a threshold value to be violated for a production web
server and the occurrence time for the violation [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. In addition to modelling the
non-stationary behaviour of metrics, they model stationary, time-serial
dependencies.
      </p>
      <p>
        Weiss modelled a genetic based machine learning system to predict rare
events by identifying predictive temporal and sequential patterns within data
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        Vilalta et al. reported three case studies for failure prediction [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]: Long term
prediction of performance variables such as disk utilization, short term prediction
of abnormal behaviour such as threshold violations and short term prediction
of system events such as router failure. As a result, they show that predictive
algorithms perform successful results in the estimation of performance variables
and prediction of critical events.
      </p>
      <p>
        Vilalta and Ma described an approach to detect patterns in event sequences
[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. They aimed to predict the occurrence of target events such as computer
attacks on host networks.
      </p>
      <p>
        Levy and Chillarege conducted a case study to develop an approach for early
warnings of failure [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. One of their ndings in this study showed that the overall
counts of alarms rise which foretell an impending failure.
      </p>
      <p>
        Liang et al. investigated the characteristics of failure events, the correlation
between these events and non-fatal events [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        Fu and Zu developed a model to cluster failure events based on their
correlations and predict their future occurrences [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
3
      </p>
    </sec>
    <sec id="sec-5">
      <title>Breakdown of System and Service Architecture</title>
      <p>The system architecture consists of four layers: Top layer is the user interface
layer. Second layer is for load balancing the transaction requests from the
interface layer (IIS layer). Third layer contains the web services processing requests.
Final layer is the database layer (Figure 1).</p>
      <p>Tellers at branches send requests using Desktop/web applications or
customers send requests via alternative distribution channels (such as internet,
ATM, etc.).
The data is gathered from the round trip times of second layer and the third
layer, in other words, load balancing layer and web-service layer. It includes the
time stamp, physical server name (IIS server), name of the web-service called,
response time and the contents of the request and the response.</p>
    </sec>
    <sec id="sec-6">
      <title>Feature Engineering</title>
      <p>Data preparation is the crucial part of statistics and machine learning processes.
In the following section, the structure of the log data, how they are selected and
pre-processed are explained.
4.1</p>
      <sec id="sec-6-1">
        <title>Structure of Logs</title>
        <p>The research data is a subset of the original log output that is gathered from
the system. The information which had been used for our research was extracted
from the original logs, which are web-service name, time stamp, number of calls,
response time, system and business faults.
4.2</p>
      </sec>
      <sec id="sec-6-2">
        <title>Feature Selection</title>
        <p>The features for machine learning methods are prepared from the number of
calls data of the logs for speci c time intervals. The number of calls of a web
service is an important indicator for showing the healthiness of the service.
Increasing or decreasing number of calls may be indicating an upcoming problem
in the system. Each row of the data set contains the number of calls for the
same time slot of previous weeks, starting from 8 weeks prior to the prediction
date. Since banking operations are mostly conducted on weekdays and the web
services under observation are mostly used by the workforce in the main
operations headquarters and branches, the dataset is formed using only weekday
information.</p>
        <p>The features for statistical analysis are also prepared by aggregating the log
data. Both the number of calls and the response times for speci c time intervals
are used for statistical analysis. The number of times the web service is called at
speci c time intervals may show the dependency between the web services. For
example, if the number of times a service is called increases and at the same time
that of another web service also increases, this may be the result of a dependency
between the services. The details of the statistical analysis are given section 5.2.
4.3</p>
      </sec>
      <sec id="sec-6-3">
        <title>Preprocessing of Log Data</title>
        <p>Log Data Aggregation. The amount of data generated for all the banking
transactions would overwhelm any data processing system. The solution to this
problem is by applying data aggregation techniques to the log data. Five minute
periods of log data is aggregated into count of calls for that web-service, mean
and standard deviation of the response time, and count of system faults and
business faults.</p>
        <p>If no service call occurs in that time period, no data aggregation can be done
and so there would be missing data for those time periods. For such periods,
data having zero counts, averages, etc. are inserted afterwards, so as to ensure
the continuity of the time series.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Experiments and Results</title>
      <p>The research consists of application of machine learning techniques on the
aggregate log data to predict future response times and statistical analysis of the
data for inferring the web service dependency mapping.
5.1</p>
      <sec id="sec-7-1">
        <title>Predictive Analytics</title>
        <p>Linear Regression, Support Vector Machines (SVM) and Random Forest
techniques have been employed for the prediction of number of calls at speci c time
intervals from the aggregated data previously explained in Feature Selection
section. Average error, which is the average percentage of the di erence between
the true value and the predicted number of calls of the web service for the related
time period to the true value, is used as the performance metric. Random Forest
application is the best performing machine learning technique, SVM come in
second place with the average error value almost doubled and Linear Regression
is the worst performing algorithm. The average error results are in the Table 1.
The data consists of response times and number of calls for 196 web services,
our aim is to uncover the relationship between these web services. In order to
gain insight on the subject we have employed statistical methods and analysed
correlations between the services' behaviour.</p>
        <p>Empirical Analysis The purpose of empirical analysis of the data was to
conduct fast analysis of the data to lter out obvious unrelated services. We have
created pairs from each service and calculated Mean Percentage Similarity(MPS)
for each pair. MPS values are calculated using the number of counts of the web
services during speci c time periods and these values are expected to show us
the similarities between the web services. The error is based solely on web service
activity, which means if there is a call for the web service there should be log
data generated for that time window in our log les. Since the data is from
one of the biggest banks and approximately 2 million log records are generated,
there is a very high chance many of the services get calls in the same aggregate
window. Therefore we have ltered the MPS values to search for 99% and above
similarity for the empirical correlation observation. There are 196 pairs with
99% and higher MPS value, Table 2 shows some examples for the MPS based
correlation. The evaluation shows that we can observe the dependency relations
for these services as they represent the 3 dependency types that we have derived
from the actual system.</p>
        <p>Correlation Analysis The second correlation type is based on the Pearson
Correlation. The variables are the average response times for each aggregation
window for the services under observation. In time series analysis, a correlated
event may occur some time after the original event, such a delay is called lag
and application of lag to the correlation analysis may yield interesting outcomes,
which are normally hidden.</p>
        <p>Pearson analysis was conducted on the data with 0 lag value. Because, in
this work, the nature of our data collection strategy (aggregation of log data
in 5 minute intervals) introduces a pseudo-lag property to our data set. The
correlation is about the response times and the resulting R values suggest that
both positive and negative correlations can be observed for the response times
of the web services. We have ltered the Pearson Correlation results to observe
only the strongest correlations with 95% or higher for positive correlation and
negative correlation results.</p>
        <p>Table 3 includes sample values out of 1020 correlation rules surpassing the
95% lter. Information in Table 2 and 3 summarizes the results and shows that
empirical analysis produced rules are compatible to the Pearson Correlation.
Dependency Types Finding dependencies between web services is considered
to be useful to nd the possible sources of problems and the e ected services
when these problems occur. The architect of the systems-under-evaluation
analysed the correlation rules and we have conjectured 3 types of dependency
relations. The rst type of relation is Dependency on the Data Source, where two
unrelated services competing for the same database resource (Figure 2).</p>
        <p>The second type of relation is the Hierarchical Dependency, where a web
service is called in another executing web service. Apart from the rest of the
execution time for calling service, the response time is increased by the response
time of the inner web service (Figure 3).</p>
        <p>The third type of relation is the Serial Execution Dependency, where a web
service call follows another web service call in a procedure. We conjecture that
negative Pearson Correlations of service calls occur based on this kind of
dependency. If the rst called web service becomes unresponsive and the response
time increases, the second web service would not execute (Figure 4).</p>
        <p>The dependencies given above are the hypothesis that we think there are
between the web services. And the selected correlation results of R values most
possibly verify our hypothesis with our current knowledge of the related web
services. However, the results should be veri ed by increasing the granularity of
the data used.
6</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>Conclusion and Future Work</title>
      <p>The results of the research demonstrates the inadequacy of the data aggregation
method, because it is very likely that the real behaviour of the web services
are lost. For further research e ort, we will try to improve the granularity of
the data to be able to re-apply machine learning methods to the response time
predictions. Response time prediction and anomaly detection methods are very
important for large organizations utilizing numerous system for their operations.
Such research is important for companies with strict service level agreements and
availability goals.</p>
      <p>The statistical analysis methods generated some rules, which provides insight
about the system-under-observation. By observing the system architecture, we
think there are 3 types of dependencies between the web services. The results
suggest that "average response times" and "the number of times web services
are called at speci c time intervals" can be used for inferring that there is a
dependency between two web services. As a future research direction, we will
use the correlation results as a foundation and we will conduct software system
dependency mapping research.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Fu</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>C.Z.</given-names>
          </string-name>
          :
          <article-title>Quantifying temporal and spatial correlation of failure events for proactive management</article-title>
          .
          <source>In: Reliable Distributed Systems</source>
          ,
          <year>2007</year>
          .
          <source>SRDS</source>
          <year>2007</year>
          .
          <article-title>26th IEEE International Symposium on</article-title>
          . pp.
          <volume>175</volume>
          {
          <fpage>184</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Hellerstein</surname>
            ,
            <given-names>J.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shahabuddin</surname>
            ,
            <given-names>P.:</given-names>
          </string-name>
          <article-title>An approach to predictive detection for service management</article-title>
          .
          <source>In: Integrated Network Management</source>
          ,
          <year>1999</year>
          .
          <article-title>Distributed Management for the Networked Millennium</article-title>
          .
          <source>Proceedings of the Sixth IFIP/IEEE International Symposium on</source>
          . pp.
          <volume>309</volume>
          {
          <fpage>322</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Levy</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chillarege</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          :
          <article-title>Early warning of failures through alarm analysis a case study in telecom voice mail systems</article-title>
          .
          <source>In: Software Reliability Engineering</source>
          ,
          <year>2003</year>
          .
          <source>ISSRE</source>
          <year>2003</year>
          . 14th International Symposium on. pp.
          <volume>271</volume>
          {
          <fpage>280</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Liang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sivasubramaniam</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jette</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sahoo</surname>
          </string-name>
          , R.:
          <article-title>Bluegene/l failure analysis and prediction models</article-title>
          .
          <source>In: Dependable Systems and Networks</source>
          ,
          <year>2006</year>
          . DSN 2006. International Conference on. pp.
          <volume>425</volume>
          {
          <fpage>434</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>T.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Siewiorek</surname>
            ,
            <given-names>D.P.</given-names>
          </string-name>
          :
          <article-title>Error log analysis: statistical modeling and heuristic trend analysis</article-title>
          .
          <source>IEEE Transactions on Reliability</source>
          <volume>39</volume>
          (
          <issue>4</issue>
          ),
          <volume>419</volume>
          {
          <fpage>432</fpage>
          (
          <year>1990</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Vilalta</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Apte</surname>
            ,
            <given-names>C.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hellerstein</surname>
            ,
            <given-names>J.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ma</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weiss</surname>
            ,
            <given-names>S.M.</given-names>
          </string-name>
          :
          <article-title>Predictive algorithms in the management of computer systems</article-title>
          .
          <source>IBM Systems Journal</source>
          <volume>41</volume>
          (
          <issue>3</issue>
          ),
          <volume>461</volume>
          {
          <fpage>474</fpage>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Vilalta</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , Ma, S.:
          <article-title>Predicting rare events in temporal domains</article-title>
          . In: null. p.
          <fpage>474</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Weiss</surname>
            ,
            <given-names>G.M.:</given-names>
          </string-name>
          <article-title>Timeweaver: A genetic algorithm for identifying predictive patterns in sequences of events</article-title>
          .
          <source>In: Proceedings of the 1st Annual Conference on Genetic and Evolutionary Computation-Volume</source>
          <volume>1</volume>
          . pp.
          <volume>718</volume>
          {
          <fpage>725</fpage>
          . Morgan Kaufmann Publishers Inc. (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>