<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Variable Stars Classi cation with the Help of Machine Learning</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Kirill Naydenkin</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Konstantin Malanchev</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maria Pruzhinskaya</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Lomonosov Moscow State University, Sternberg Astronomical Institute</institution>
          ,
          <addr-line>Universitetsky pr. 13, Moscow, 119234</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>National Research University Higher School of Economics</institution>
          ,
          <addr-line>21/4 Staraya Basmannaya Ulitsa, Moscow, 105066</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Physics Faculty, Lomonosov Moscow State University</institution>
          ,
          <addr-line>Leninskii Gori 1, 119234</addr-line>
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <fpage>296</fpage>
      <lpage>303</lpage>
      <abstract>
        <p>With the appearance of modern technologies such as CCDmatrices, large telescopes and computer networks the precision of our observations increased immensely. On the other hand, such accurate and complex data formed TBs large data bases which are very fragile and unattainable for the treatment by classical methods. The scales of this problem can be seen especially in variable star sky surveys. For many terabytes of data one has to classify all the stars in catalog to nd stars of particular type of variability. This problem is known as very important since almost every part of modern astrophysics is interested in new objects to study. In some elds like cosmology, this question is very vital due to high demand for new data of model-anchors like Cepheids or supernova stars. To facilitate this task many machine learning based algorithms were proposed (Richards et al., 2011 [2]). In this study we perform a way to classify the Zwicky Transient Facility Public Data Release 1 catalog onto variable stars of di erent types. As the priority classes we set Cepheids, RR Lyrae and Scuti. \One vs all" classi cation technique revealed highly accurate results on validation data, concretely 0.90{0.95 with ROC-AUC metrics.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The Zwicky Transient Facility (ZTF) is a 48-inch Schmidt telescope
with a 47 sq. deg. eld of view at the Palomar Observatory in
California. This large eld of view ensures that the ZTF survey can scan
the entire northern sky every night. The ZTF survey started on 2018
March 17. During the planned three years survey, ZTF is expected to
acquire 450 observational epochs for 1.8 billion objects. Its main
scienti c goals are the physics of transient objects, stellar
variability, and solar system science (Graham et al. 2019 [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]; Mahabal et al.
2019 [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]).
      </p>
      <p>Copyright © 2020 for this paper by its authors. Use permitted under Creative</p>
      <p>In this paper we made an attempt to classify the rst data release
of ZTF survey (DR1) which contains data acquired between March
and December 2018, thus covering a timespan of around 290 days.
The rst data release includes more than 800 thousand light curves
observed in both zr and zg passbands.</p>
      <p>
        The further steps of classi cation ZTF DR1 imply highly
accurate data treatment in the beginning. To have more clear awareness
about the stage of data treatment it is important to divide it onto
small parts. First of all, in demand of supervised machine learning
algorithms it is important to prepare precise labeled data which in our
case consist of star-catalog with coordinates and types of 55
thousands variable stars of the General Catalog of Variable Stars (GCVS,
Samus et al. 2017 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]). Even though this number is relatively small
compared to the size of ZTF data release, there are some ways to
generate the data on the basis of accurate initial frames.
1.1
      </p>
      <p>Data Preparation
Since ZTF DR1 marks the same objects observed in di erent
passbands and/or di erent sky elds with di erent identi ers (IDs),
cross-match can yield more than one ZTF DR1 ID for given GCVS
object. Taking into account that ZTF ranged in 12m 21m, it is
reasonable to remove those stars which do not belong to that range
(with slight o set due to star magnitude uctuations with time).
After this ltering, 43 thousands of total 55 thousands GCVS stars
remained.</p>
      <p>Fig. 1 represents the distribution of number of cross-matched ZTF
DR1 objects. To nd the real range of di erence in coordinates
between catalogs we measured distances for the case when the only ID
was found for a given GCVS object and for the case of two IDs
separately in both lters. As the result, we received that only relatively
small amount of objects has more than 0.1" di erence. This
coordinates distinction can be explained by catalogs' positional inaccuracy,
which for GCVS is 0.1".</p>
      <p>To this point we had matched labeled GCVS with objects from
ZTF. As we mentioned before, ZTF DR1 contains photometry in
two passbands (zg, zr) which we can be used either separately or in
combination. We have to note that regardless shrinking the search
radius to 1.5" we still can receive multiple results | from a light
curve in only one lter up to a few series in both (Fig. 2). Taking
into account such asymmetry we made up three homogeneous sets:
19k objects in g band, 14k in r band, and 13k in combination (setting
up threshold for minimal amount of observations in a pair).
2</p>
    </sec>
    <sec id="sec-2">
      <title>Objects of Interest</title>
      <p>In our study we took some particular classes of interest among the
variable stars: Cepheids, RR Lyrae and Scuti. It is important to
note that we do not use multi-class classi cation for chosen types,
every type goes through one-vs-all approach.</p>
      <p>First of all, let's describe the physical nature of chosen types.
Cepheids are pulsating stars, which radius and brightness (as well as
the temperature) change with time. Cepheids stars are well known
for the dependency between luminosity and pulsating period, this
property makes them important indicators of cosmic distances. RR Lyrae
stars can also be used as standard candles for distance measurements,
though they do not follow a strict period-luminosity relationship at
visual wavelengths. Scuti as well as Cepheids are important
standard candles and have been used to establish the distance to many
large clusters around the center of our galaxy.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Analysis</title>
      <p>
        The next step toward the data classi cation is to engineer features
out of light curves. To start we created three most valuable sources of
information (Richards et al., 2011 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]): magnitude amplitude range,
the main peak period and power of Lomb{Scarge periodogram (Lomb
1976 [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]; Scargle 1982 [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]). For this purpose we used astropy [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]
library-based LombScargle() function which allows us to de ne
both positional coordinates of Lomb{Scargle periodograms peak.
      </p>
      <p>
        One possible way to select a set of variable stars out of ZTF DR1
is to use the Lomb{Scarge periodogram. Such attempts were made
recently and yielded to the strong results (Chen et al 2020 [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]).
      </p>
      <p>
        On the one hand, one could possibly observe the importance of
di erent features for di erent types of variable stars from (Richards
et al., 2011 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]). On the other hand, ZTF light curves can be slightly
di erent from the data studied in the article and this could possible
change the picture.
      </p>
      <p>
        After choosing metric for result evaluation we worked with
validation set to nd an optimal list of features for the classi cation task
in ZTF DR1. Richards et al., 2011 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] presented a table of pairwise
random forest feature importance for all basic variable types. As
one can see coordinates of rst peaks of Lomb{Scarge periodogram
and their ration (ratio of periodogram's frequencies) have a signi
cant in uence on majority of classes. Basic characteristics of a signal
such as amplitude, std, skew, median absolute deviation also have
a strong correlation with correct classi cation of all classes. For our
one-vs-all approach we took 17 the most strongest features out of
Richards et al., 2011 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] | coordinates of rst four periodogram's
peaks, frequencies ratio, mentioned basic signal properties and few
additional such as trend angle. To implement these features we used
python libraries | astropy and statistical analysis from scipy. It
is obvious, that application of this approach to one-vs-all classi
cation leads to imbalance in labeled data because even the largest types
consist out of less then 15% of the train set. To deal with imbalanced
data few possible options are available. First of all, we can use one
of the ways to generate rare sample, for example random sampling
from a chosen distribution. On the other hand, we can divide the
more abundant class into N distinct clusters and train each of N
classi ers on one of the distinct clusters and on all of the data from
the rare type. After that ensemble the result from all models. As far
as we address the imbalanced problem with signi cantly low
number of minor class examples, it is more e cient to use over-sampling
techniques instead of under-sampling. For this purpose we took
Synthetic Minority Oversampling Technique, or SMOTE. SMOTE rst
selects a minority class instance A at random and nds its k-nearest
minority class neighbors. The synthetic instance is then created by
choosing one of the k-nearest neighbors B at random and connecting
A and B to form a line segment in the feature space. The synthetic
instances are generated as a convex combination of the two chosen
instances A and B. Using SMOTE from imblearn python library we
updated the minority class by oversampling to have 10 percent the
number of examples of the majority class, then used random
undersampling to reduce the number of examples in the majority class to
have 50 per cent more than the minority class.
      </p>
      <p>
        For the rst trial of binary classi cation we used Cepheids. The
performance is shown in Table 3.
Comparing our approach with recent study of Chen et al 2020 [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ],
which performed a large number of new RR Lyrae, Cepheid and
      </p>
      <p>
        Scuti we used machine learning classi cation instead of directly
measured distances in parameter space. Machine learning way of
data classi cation can perform better results because it relies on both
successes and failures to estimate the reliability of certain candidate.
Moreover, applying hierarchical classi cation we have an ability to
track mistakes of our model on di erent levels | end up with highest
possible accuracy on large types which generally have some physical
di erence like binaries, pulsating, eruptive or rotating stars. On the
next step large types break into sub-types.
4.2 Future Plans
As the future structure of the research we can mark the following
steps. First of all, to obtain some initial results we have to
complete the process of validation of our model on GCVS data for our
classes of interest. After that we can start to classify ZTF data
either after ltrating with Lomb-Scarge periodograms or making
features directly for each light curve. As the next step we will use more
public available labeled catalogs such as ASAS-SN [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], ATLAS [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ],
Catalina catalog of periodic variable stars [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], the Gaia catalog of
RR Lyrae and Cepheids [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. At this point, we will also introduce
data generation which can signi cantly improve the accuracy of the
results.
      </p>
      <p>At the nal step we will take di erent classi cation tools of ML:
classical Random Forest and XGBoost models with the main hyper
parameters to choose and imbalanced learning to make hierarchical
classi cation.</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusions</title>
      <p>In this work we consider the machine learning technique as a
perspective approach for the classi cation task. We use ZTF data release 1
that contains zg and zr photometry of the variable objects. Our
preprocessing procedure includes several steps. First, we have to prepare
the datasets with labels to use them later for training the model and
testing the accuracy of the method. For this purpose the General
Catalogue of Variable Stars is used. After cross-matching the
catalog with ZTF DR1 we found 19k common objects in zg-band, 14k
in zr-band, and 13k in combination of passbands. Then, we have to
choose the appropriate features to describe the light curves. As a
starting point, we tried the magnitude amplitude range, the main
peak period and power of Lomb|Scarge periodogram.</p>
      <p>There are many types of variable stars di er by the underlying
physical processes or their observational appearance. As the objects
of interest we chose RR Lyrae, Cepheid and Scuti. We applied
binary classi cation technique to Cepheid stars and found quite good
results on the validation dataset. The work done is a preparatory step
towards the further thorough machine learning classi cation of the
variable stars in ZTF data.</p>
      <p>Acknowledgments. K. Malanchev and M. Pruzhinskaya are
supported by RBFR grant 20-02-00779. The authors acknowledge the
support by the Interdisciplinary Scienti c and Educational School of
Moscow University \Fundamental and Applied Space Research".</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Pruzhinskaya</surname>
            ,
            <given-names>M. V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Malanchev</surname>
            ,
            <given-names>K. L.</given-names>
          </string-name>
          , et. al.:
          <article-title>Anomaly detection in the Open Supernova Catalog</article-title>
          .
          <source>Monthly Notices of the Royal Astronomical Society</source>
          <volume>489</volume>
          (
          <issue>3</issue>
          ):
          <volume>3591</volume>
          {
          <fpage>3608</fpage>
          (
          <year>2019</year>
          ). https://doi.org/10.1093/mnras/stz2362.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Richards</surname>
          </string-name>
          et al.:
          <article-title>On Machine-Learned Classi cation of Variable Stars with Sparse and Noisy Time-Series Data</article-title>
          .
          <source>The Astrophysical Journal</source>
          .
          <volume>733</volume>
          (
          <issue>10</issue>
          ) (
          <year>2011</year>
          ).
          <volume>10</volume>
          .1088/
          <fpage>0004</fpage>
          - 637X/733/1/10
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Graham</surname>
            ,
            <given-names>M. J.</given-names>
          </string-name>
          , et. al.:
          <source>The Zwicky Transient Facility: Science Objectives. Publ. Astron. Soc. Pac</source>
          .
          <volume>131</volume>
          (
          <issue>1001</issue>
          ):
          <volume>078001</volume>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Samus</surname>
            ,
            <given-names>N.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kazarovets</surname>
            ,
            <given-names>E.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Durlevich</surname>
            ,
            <given-names>O.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kireeva</surname>
            ,
            <given-names>N.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pastukhova</surname>
            <given-names>E.N.</given-names>
          </string-name>
          :
          <source>General Catalogue of Variable Stars: Version GCVS 5.1. Astronomy Reports</source>
          <volume>61</volume>
          (
          <issue>1</issue>
          ):
          <fpage>80</fpage>
          -
          <lpage>88</lpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Mahabal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , et. al.:
          <article-title>Machine Learning for the Zwicky Transient Facility</article-title>
          .
          <source>The Astronomical Society of the Paci c</source>
          <volume>131</volume>
          (
          <issue>997</issue>
          ) (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Chen</surname>
          </string-name>
          , et. al.
          <article-title>The Zwicky Transient Facility Catalog of Periodic Variable Stars</article-title>
          . arXiv:
          <year>2005</year>
          .
          <article-title>08662 [astro-ph</article-title>
          .
          <source>SR]</source>
          (
          <year>2020</year>
          ).
          <volume>10</volume>
          .3847/
          <fpage>1538</fpage>
          -4365/ab9cae
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Lomb</surname>
            ,
            <given-names>N.R.</given-names>
          </string-name>
          :
          <article-title>Least-squares frequency analysis of unequally spaced data</article-title>
          .
          <source>Astrophys. Space Sci</source>
          .
          <volume>39</volume>
          :
          <issue>447</issue>
          {
          <fpage>462</fpage>
          (
          <year>1976</year>
          ). https://doi.org/10.1007/BF00648343
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Scargle</surname>
            ,
            <given-names>J. D.</given-names>
          </string-name>
          :
          <article-title>Studies in astronomical time series analysis. II. Statistical aspects of spectral analysis of unevenly spaced data</article-title>
          .
          <source>Astrophysical Journal</source>
          <volume>263</volume>
          :
          <fpage>835</fpage>
          -
          <lpage>853</lpage>
          (
          <year>1982</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>All-Sky Automated</surname>
          </string-name>
          Survey for Supernovae. https://asas-sn.osu.edu/
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <article-title>The Herschel ATLAS</article-title>
          . https://www.h-atlas.org/
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <article-title>The Catalina Surveys Data Release 2</article-title>
          . http://nesssi.cacr.caltech.edu/DataRelease/
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>12. Gaia Archive at ESA. https://gea.esac.esa.int/archive/</mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <given-names>Astropy</given-names>
            <surname>Project</surname>
          </string-name>
          . http://www.astropy.org
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>