<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Research on Unsupervised Anomaly Detection of Gas Heating Energy Consumption Based on Ensemble Learning 1</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Lizhuo Gao</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Huihua Yang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yanzhu Hu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Beijing University of Posts and Telecommunications</institution>
          ,
          <addr-line>Beijing</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <fpage>188</fpage>
      <lpage>193</lpage>
      <abstract>
        <p>To improve the detection of abnormal data of heating energy of gas boilers, based on the gas heating energy consumption data of a place in Beijing in recent two years, this paper deeply studies and analyzes the detection effects, advantages and disadvantages of unsupervised anomaly detection algorithm, iForest, LOF and One-Class SVM algorithm models. Finally, based on the idea of ensemble learning, the weighted fusion of the above three algorithms is carried out, and the accuracy of anomaly detection in gas heating is successfully improved. The F1 value of the fusion model on the data set is about 95.8%.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Anomaly detection</kwd>
        <kwd>Gas</kwd>
        <kwd>Ensemble learning</kwd>
        <kwd>Unsupervised</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>There are abnormal conditions in the energy consumption of gas heating, such as natural gas pipeline
leakage, furnace pipe scaling or corrosion, water shortage and pipe explosion. With the increase of the
amount of data, there are more and more abnormal data, and it will become more and more difficult to
extract effective information from the data. Therefore, it is necessary to improve the accuracy of
abnormal detection with the help of machine learning method. At present, the research on abnormal
detection of gas heating energy using machine learning algorithm is almost a piece of white paper.
Machine learning methods are mainly divided into supervised learning and unsupervised learning.
Supervised learning requires high-quality labels. These labels need to be manually labeled for the
current data set, and then targeted training. Although the supervised detection effect is better, it has no
generalization, so the supervised method is not desirable.</p>
      <p>The algorithm mentioned in this paper are popular and unsupervised anomaly detection algorithms.
Meanwhile, in other application fields of anomaly detection, many researchers have proposed integrated
detection algorithms and achieved certain results. For example, Li Guocheng and others have proposed
an isolated forest power theft detection algorithm based on bagging quadratic weighted integration,
which improves the power theft detection effect[1]. Gao Xin et al. proposed an integrated algorithm for
anomaly detection of power dispatching data based on log interval isolated forest[2]. The proposed
method is progressiveness in the comprehensive performance of anomaly detection AUC value.</p>
      <p>In summary, this paper adopts the unsupervised learning method to make an exploratory research
on the abnormal detection of energy consumption of gas heating. Based on the idea of integration,
iForest, LOF and One-Class SVM are weighted integrated as the base classifiers, and the fusion model
has better effect on the gas data set.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Anomaly detection model</title>
    </sec>
    <sec id="sec-3">
      <title>2.1. Detection based on iForest algorithm</title>
      <p>Forest (Isolated forest) algorithm detects outliers by isolating sample points, and isolates samples by
using binary search tree structure of isolated tree[3]. In our given gas sampling point database, the vast
majority of data points are surrounded together in space, reflecting similar characteristics, while outliers
will be isolated from other data and isolated earlier.
2.1.1.</p>
    </sec>
    <sec id="sec-4">
      <title>Create iForest</title>
      <p>below:</p>
      <p>The establishment of iForest mainly depends on the generation of iTree. The core steps are described
Step1: The root node was established and 100 samples were randomly selected as the sample set.</p>
      <p>Step2: Specify the function variable and randomly select one important feature from the existing 12
parameters as the splitting basis of iTee.
specified variable.
cut-off point.</p>
      <p>Step3: Generates a random cutoff point in the selected maximum and minimum values of the
Step4: The sample data segment is divided into two subspaces by extending a hyperplane from the
Step5: Determine whether the samples contained in each subspace are the same samples, or the iTree
splitting times have reached log2n. If the conditions are met, splitting is terminated and an iTree is
constructed. Otherwise proceed to Step3, Step4, and Step5.</p>
      <p>According to the above steps, set the number of iTree N=50, and set the extraction method as
repeatable sampling. The whole iForest can be established by repeating the above steps.
2.1.2.</p>
      <p>
        Calculate outlier score
100 randomly selected data samples can be evaluated by the iForest tree. When the data samples are
spread all over iTree, the node depth of the samples in each tree is extracted, and the average depth of
the samples in iForest is calculated. The abnormal score of the samples can be obtained by further
transformation of the depth of the iTree. The formula for calculating abnormal scores is shown in (
        <xref ref-type="bibr" rid="ref1 ref4">3</xref>
        ):
 (,  ) = 2
 ( ) = 2 ( − 1 ) −
() = ( ) + 0.577
( )
( )
2( − 1 )
      </p>
      <p>Where, h(x) is the depth of the node of sample x in iTree, E represents the average value, and c(n)
represents the average length of the binary tree constructed by N sample points. In this application
model, n=100. The closer s(x) score is to 1, the more likely the sample is to be an outlier, and the closer
s(x) score is to 0, the more likely it is to be a normal sample.</p>
    </sec>
    <sec id="sec-5">
      <title>2.2. Detection based on LOF algorithm</title>
      <p>LOF( Local Outliers Factor) algorithm realizes anomaly detection by calculating the local density
deviation of a given data point relative to its neighborhood [4]. We select the data at any time from a
set of data as point p and calculate the local density of point p. The lower the density is, the more likely
it is to be an outlier.</p>
      <p>The k
th distance of the point p to be measured is defined as follows:</p>
      <p>( ) =  (,  )</p>
      <p>
        And : (
        <xref ref-type="bibr" rid="ref2">1</xref>
        ) There are at least k points o, that do not include p in the set. (
        <xref ref-type="bibr" rid="ref3">2</xref>
        ) There are at most k-1
Nk ( p) is the kth distance of p and all points including the k
th distance, so the number of k
th domain
points of p is greater than or equal to k. Reach-distance k( p, o) is defined as the reachable distance
from o to p, which is at least the k
      </p>
      <p>th distance of o, or the true distance of o, p.</p>
      <p>The local outlier of point p is the average of the ratio of the local accessible density of the domain
point Nk ( p) of point p to the local accessible density of point p. The significance of this formula is
that if the ratio is closer to 1, the density difference between point p and field point is smaller, and point
p and field point belong to the same class cluster. If this ratio is greater than 1, it indicates that the
density of point p is less than that of its domain points, and p may be an outlier. Which can be expressed
as:


( ) = ∑∈
( ) =
∑∈
( )
( )(
| ( )|</p>
      <p>
        () /
| ( )|
(, )
())
(
        <xref ref-type="bibr" rid="ref1 ref4">3</xref>
        )
(
        <xref ref-type="bibr" rid="ref5">4</xref>
        )
      </p>
    </sec>
    <sec id="sec-6">
      <title>2.3. Detection based on One-Class SVM algorithm</title>
      <p>One-Class SVM is an unsupervised learning method, that is, it does not need us to mark the output
label of the training set. There are many solutions for One-Class SVM to find the divided hyperplane
and support vector machine[5]. Here is only a special idea SVDD. For SVDD, we expect that all
samples that are not abnormal are positive categories. At the same time, it uses a hypersphere rather
than a hyperplane to divide. The algorithm obtains the spherical boundary around the data in the feature
space, hoping to minimize the volume of the hypersphere, so as to minimize the influence of abnormal
point data.</p>
      <p>
        Assuming that the parameters of the generated hypersphere are the center o and the corresponding
hypersphere radius r &gt; 0, the volume V (r) of the hypersphere is minimized, and the center o is a
linear combination of support vectors; Similar to the traditional SVM method, the distance from all
training data points xi to the center can be strictly less than r . But at the same time, a relaxation
variable with penalty coefficient C is constructed ζ i . The optimization problem is shown in the
following equation (
        <xref ref-type="bibr" rid="ref5">4</xref>
        ):
      </p>
      <p>≤  +  ,  = 1,2,3 …</p>
      <p>≥ 0,  = 1,2,3 …</p>
      <p>After solving with Lagrange duality, we can judge whether the new data point Z is included. If the
distance from Z to the center is less than or equal to radius r , it is not an abnormal point. If it is outside
the hypersphere, it is an abnormal point.</p>
    </sec>
    <sec id="sec-7">
      <title>2.4. Algorithm optimization</title>
      <p>Anomaly detection can be regarded as a binary classification problem of anomaly and normal, so
the three basic models are called base classifiers in this paper. Based on the idea of ensemble learning[6],
this paper takes iForest, LOF and One-Class SVM as three base classifiers for ensemble learning.</p>
      <p>The specific process is as follows:</p>
      <p>If a sample point in the first base classifier has been accurately classified, its weight will be reduced
when constructing the next classifier; On the contrary, if a sample point is not accurately classified, the
weight is increased. Because there are many ways of weight distribution of training data, which will not
be repeated here.</p>
      <p>Then, the weight updated samples are used to train the next classifier, and the whole training process
goes on like this. Finally, increase the weight of the base classifier with small classification error rate
to make it play a greater decisive role in the final classification function, while reduce the weight of the
base classifier with large classification error rate to make it play a smaller decisive role in the final</p>
      <p>Here is the calculation method of the weight coefficient of the base classifier, that is, first calculate
the detection error rate of the above three anomaly detection models, then calculate the weight
coefficient of the three base classifiers, and finally combine the three base classifiers.</p>
      <p>
        Step1: Calculation error rate 
, as shown in the following equation (
        <xref ref-type="bibr" rid="ref6">5</xref>
        ).

=
∑
 ℎ ( )
      </p>
      <p>( )</p>
      <p>Where I[] is the discriminant function, Xi represents the ith sample in the training set, a total of n
samples, t takes 1, 2 and 3 to represent the three basic classifiers of iForest, LOF and One-Class SVM
respectively, Dt (xi ) is the weight distribution of the training set of the t-th basic classifier, and ht (xi )
is the classification result of the i-th sample by the t-th basic classifier.</p>
      <p>Step2：Calculate the weight</p>
      <p>
        of base classifier ℎ , as shown in the following equation (
        <xref ref-type="bibr" rid="ref7">6</xref>
        ).
      </p>
      <p>Step3：Combine three base classifiers to get f (x) , as shown in the following equation (7). Thus,
the final classifier is obtained, as shown in the following equation (8). Among them, three base
classifiers are used, so T is 3.</p>
      <p>=log
1 −</p>
      <p>( ) =</p>
      <p>ℎ ( )
 ( ) = 
( ) =  
ℎ ( )
classification function. In other words, the classifier with low error rate accounts for a large proportion
in the final classifier, and vice versa (see Figure 1).</p>
    </sec>
    <sec id="sec-8">
      <title>3. Comparative experiment</title>
    </sec>
    <sec id="sec-9">
      <title>3.1. Data Description</title>
      <p>
        This paper collects the real heating data of boiler room a in Beijing from 2020 to 2022. The sample
size after logarithmic fusion is 500000 and the dimension is more than 200. After expert selection and
feature selection, the dimension is reduced to 53 dimensions. These data sets contain real time series,
and abnormal data has also been marked. Some of the marks are artificially designed and added fault
and alarm scenarios based on the experience of gas operation and maintenance experts. The abnormal
proportion of the finally processed data sets is about 11%. The amount of data and labels can support
the construction and evaluation of the model. Some of the real labels are shown in the Table 1.
(
        <xref ref-type="bibr" rid="ref6">5</xref>
        )
(
        <xref ref-type="bibr" rid="ref7">6</xref>
        )
(7)
(8)
      </p>
      <p>In the data preprocessing stage, the data with single point blank or only single point mutation of
continuous variables are supplemented by the average values of the previous second and the next second.
Discrete variables are directly filled with 0 value according to their meaning. At the same time, this
paper also constructs the instantaneous heat released by gas and the heat absorbed by boiler water. On
this basis, the change trend of heat proportion is estimated through calculation, which is used as the
auxiliary index of the anomaly detection model in this paper. In feature selection, through SelectKBest
in sklean, this paper uses the maximum information coefficient as the scoring function to select the
features most related to the label.</p>
    </sec>
    <sec id="sec-10">
      <title>3.2. Experimental Result</title>
      <p>The parameter settings of iForest model are as follows: Max_samples= 120，contamination=0.11.
The settings in the LOF model are as follows: n_neighbors=20, leaf_size=30, contamination=0.11.
Euclidean distance measurement is adopted. One-Class SVM adopts Gaussian kernel function, and the
parameter gamma of RBF kernel type is set to 0.1. The contamination represents the proportion of
outliers in the total data volume.</p>
      <p>The performance of anomaly detection using iForest, LOF or One-Class SVM is very general. In
fact, this is because each algorithm is good at dealing with different scenes and characteristics, as
follows: LOF believes that the outliers are the ones with large deviation from the local density of a
given data point relative to its neighborhood. In other words, it pays more attention to the local; iForest
believes that in the data space, the sparsely distributed areas mean that the probability of data occurring
in this area is very low, so it can be considered that the data falling in these areas is abnormal. In other
words, it pays more attention to the whole; One-Class SVM is an algorithm for anomaly detection based
on the characteristics of normal data. It considers all data with similar characteristics to normal data as
normal data, otherwise it is considered as abnormal data. However, in the complex scenario of energy
consumption of gas heating, the manifestations of anomalies are complex. For example, the abnormal
data of water pipe scaling is very similar to the normal data. In other words, the three models have their
places that are not well considered. Therefore, based on the idea of integration, this paper takes iForest,
LOF and One-Class SVM as base classifiers, combines them with reference to the weighted method of
AdaBoost, complements their advantages, and forms an anomaly monitoring model with better
performance in gas heating energy consumption detection(see Figure 2).</p>
      <p>The anomaly detection algorithm finally divides the data into normal data and abnormal data, which
is a binary classification problem. Therefore, the anomaly detection algorithm uses the evaluation
standard F-score value of the classification model to evaluate(see Table 2).</p>
      <p>It can be seen that the accuracy and recall rate of the model are very high. Among them, the recall
rate of One-class SVM model is very high, but the accuracy is low. Overall, the F1 value of the improved
model is about 0.96, and achieved good results.</p>
    </sec>
    <sec id="sec-11">
      <title>4. Conclusion</title>
      <p>Based on the idea of integration, this paper establishes an unsupervised anomaly detection model of
gas heating energy consumption through weak classifier integration, and compares it with other popular
anomaly detection models. This paper mainly makes two contributions: first, unsupervised anomaly
detection model is applied to the heating energy anomaly detection of gas boiler room for the first time,
and relevant exploratory research is carried out in this field, which has a strong reference value for the
application of gas heating anomaly detection combined with machine learning. Second, the detection
effect is improved by integrating the fusion model, and the F1 value reaches 0.5 on the existing data set
0.9579, which is 2 to 1 percentage points higher than the single anomaly detection model. Of course,
when this model is applied to different gas fired boilers in practical application, there are problems such
as insufficient universality, long training time and high complexity, which will be further improved in
the subsequent research.</p>
    </sec>
    <sec id="sec-12">
      <title>5. References</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>3.3. Evaluation</mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Li</given-names>
            <surname>Guocheng</surname>
          </string-name>
          . :
          <article-title>An isolated forest power theft detection algorithm based on Bagging quadratic weighted integration</article-title>
          [J].
          <source>Power System Automation</source>
          ,
          <year>2022</year>
          ,
          <volume>46</volume>
          (
          <issue>2</issue>
          ):
          <fpage>92</fpage>
          -
          <lpage>100</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Gao</given-names>
            <surname>Xin</surname>
          </string-name>
          .
          <article-title>: A method of power dispatching data anomaly detection based on log interval isolation [J] Power grid technology</article-title>
          ,
          <year>2021</year>
          ,
          <volume>45</volume>
          (
          <issue>12</issue>
          ):
          <fpage>10</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>He</given-names>
            <surname>Kun</surname>
          </string-name>
          .
          <article-title>: Study on identification of electrical anomalies in isolated forest based on PCA[J]</article-title>
          .
          <source>Computing Technology and Automation</source>
          ,
          <year>2021</year>
          ,
          <volume>40</volume>
          (
          <issue>2</issue>
          ) :
          <fpage>76</fpage>
          -
          <lpage>80</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Zhang</given-names>
            <surname>Shuo</surname>
          </string-name>
          . :
          <article-title>Outlier detection algorithm based on grid LOF and adaptive K-means</article-title>
          [J].
          <source>Command Information Systems and Technology</source>
          ,
          <year>2019</year>
          ,
          <volume>10</volume>
          (
          <issue>1</issue>
          ):
          <fpage>90</fpage>
          -
          <lpage>94</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Li</given-names>
            <surname>Chengliang</surname>
          </string-name>
          .
          <article-title>: Study on anomaly detection method of Airborne Tacang Ranging Information based on One-Class SVM</article-title>
          [J].
          <source>Modern navigation</source>
          ,
          <year>2015</year>
          (3):
          <fpage>282</fpage>
          -
          <lpage>285</lpage>
          ,
          <fpage>309</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Li</given-names>
            <surname>Yuan</surname>
          </string-name>
          .
          <article-title>: Research on fault detection based on K-means clustering and local outlier factor algorithm [J]</article-title>
          .
          <source>Chemical automation and instrumentation</source>
          ,
          <year>2019</year>
          ,
          <volume>46</volume>
          (
          <issue>10</issue>
          ):
          <fpage>816</fpage>
          -
          <lpage>821</lpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>