<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>K. G. );</journal-title>
      </journal-title-group>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Learning based on AQI and Weather Features</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Arjun Shah</string-name>
          <email>arjun.a.shah244@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Varun Viswanath</string-name>
          <email>varunvis2903@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kashish Gandhi</string-name>
          <email>kashishgandhi6112003@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Workshop</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Solar Power Generation, Zero Inflated Model</institution>
          ,
          <addr-line>Power Transform, Time series, LSTM, Deep Learning</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Synapse, Computer Engineering, D.J. Sanghvi College of Engineering</institution>
          ,
          <addr-line>Mumbai</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>1952</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>This paper addresses the pressing need for an accurate solar energy prediction model, which is crucial for eficient grid integration. We explore the influence of the Air Quality Index and weather features on solar energy generation, employing advanced Machine Learning and Deep Learning techniques. Our methodology uses time series modeling and makes novel use of power transform normalization and zero-inflated modeling. Various Machine Learning algorithms and Conv2D Long Short-Term Memory model based Deep Learning models are applied to these transformations for precise predictions. Results underscore the efectiveness of our approach, demonstrating enhanced prediction accuracy with Air Quality Index and weather features. We achieved a 0.9691 R2 Score, 0.18 MAE, 0.10 RMSE with Conv2D Long Short-Term Memory model, showcasing the power transform technique's innovation in enhancing time series forecasting for solar energy generation. Such results help our research contribute valuable insights to the synergy between Air Quality Index, weather features, and Deep Learning techniques for solar energy prediction.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>4,5</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        In the modern world, it has become increasingly clear that eliminating fossil fuels is one of the huge
requirements to achieve a carbon-neutral future. The Working Group III Special Report on Renewable
Energy Sources and Climate Change Mitigation (SRREN) [19] suggests that consumption of fossil fuels
accounts for the majority of anthropogenic GHG emissions worldwide. It states that CO2 consumption
had risen to over 390 ppm, which was around 39% above pre-industrial levels by the end of 2010. In
the race to an eficient energy ecosystem, solar energy is a promising renewable resource, but its
intermittent nature poses challenges for integration into the grid. A recent study shows that India lost
29% of its utilizable global horizontal irradiance potential due to air pollution in the period between
2008 and 2018 [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Efective prediction of solar power generation is crucial for eficient planning and
management of solar resources. Renewable energy like solar power is said to benefit human beings
in a lot of diferent ways and the most important is in the health domain. Research by Galimova et al.
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] suggests that by 2050 if the world goes under a global transition and the energy sector emissions
drop by 92% then we can reduce premature deaths by air pollution by 97%. This study reinforced
the significance of considering environmental factors in our solar energy forecasting models. This
study examines the use of machine learning algorithms that incorporate Air Quality Index (AQI) and
meteorological features to improve forecast accuracy.
      </p>
      <p>In order to shed some light on the inconsistent patterns of solar generation data, a number of
regression models were initially utilised to predict the per-hour generation of solar power. We thus
benchmarked a number of regression models, of which the chief ones were Linear Regression, Lasso,
Ridge, ElasticNet, and ensemble models like RandomForest and XGBoost. These models utilize diferent
Event, Lucknow, India.
∗Corresponding author.</p>
      <p>CEUR</p>
      <p>
        ceur-ws.org
methodologies which were compared to determine the model with optimal performance to provide
valuable insights into the future of solar generation.[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] Additionally, this paper describes a few novel
methods such as implementing a Zero Inflated Model and scaling the data using Power Transform,
which can significantly improve solar power predictions irrespective of the irregularity in data. This
study also incorporates the Convolutional Long Short-Term Memory 2D (ConvLSTM2D) network in
the prediction process. ConvLSTM2D is a deep learning algorithm that combines the spatial processing
capabilities of Convolutional Neural Networks (CNNs) with the sequential processing capabilities of
Long Short-Term Memory (LSTMs)s. Using the spatiotemporal dependence of solar generation data,
the ConvLSTM2D network improves the solar energy generation forecast accuracy by building upon
historical data.
      </p>
      <p>The novel contributions of this work can be summarized as follows:</p>
      <p>This study hopes to investigate and optimise the performance of the aforementioned machine
learning algorithms and provide a clear and accurate picture of solar generation by considering weather
features and AQI data which are theorized to have an impact on the fluctuation of this data. It is
envisaged that the methodologies employed in this paper will contribute considerably to painting
a clearer picture of the sporadic nature of solar power and its influencing factors. The thorough
benchmarking and prediction pipeline provides significant benefits for the eficient and sustainable
utilization of solar resources by stakeholders, contributing to the adoption and use of practices that
ultimately power our society into a clean energy future.</p>
    </sec>
    <sec id="sec-3">
      <title>2. Survey of Literature</title>
      <p>
        Various studies have used machine learning algorithms to increase the understanding and
improvement of solar power forecasting models. Chuluunsaikhan et al. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] discusses the importance
of considering environmental factors such as climate and air pollution when predicting solar power
generation. It states that solar panels work best when there is sunlight and no partial shade. However,
factors such as weather conditions (e.g. clouds or rain) and air pollution (e.g. fine dust) can cause partial
shading and reduce the power output of solar panels. The authors propose a method to regulate the
power output of solar panels through machine learning. Machine learning models are developed with
three components: weather components, air pollution components, and combined meteorological and
air pollution components. The datasets used in the study were collected from 2017 to 2019 from the
Seoul province of South Korea. The paper describes the methodology used, including data acquisition,
feature extraction, model training, and power output prediction. The authors compare machine
learning models, such as linear regression, k-Nearest Neighbors (kNN), Support Vector Regression (SVR),
Multi-Level Perceptron (MLP), Random Forest Regressor (RFR), and Gradient-Boosting Regressor (GBR).
Models are evaluated using quantitative error methods such as MAE, Coeficient of Determination
(R2), and Root Mean Square Error (RMSE). Experimental results show that weather and air pollution
parameters can be efective predictors. This paper has been the main premise of our research.
      </p>
      <p>
        Zhou et al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] presents a diferent approach to forecasting short-term solar power output in smart
cities by using deep learning techniques. They used a combination of clustering, CNN, LSTM, and
attention mechanisms that obtained improved accuracy in predicting future energy generation. The
authors proposed that for future work one could develop training models with time series-based data for
further improvement and this is a proposal that we took under consideration. This study and research
by Zhou et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] motivated us to incorporate air quality index as a feature in our machine learning
models since they used community multiscale air quality in their research indicating how air pollutants
can contribute to soiling of PV panels afecting the solar power generation which is also mentioned by
Chiteka et al. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>
        Additionally, Jia et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] presented models to predict solar radiation; even though our research is
based on solar power generation, this paper gave us important insights regarding the use of machine
learning models in solar forecasting under various weather conditions. Along with this we also
considered how machine learning models are computationally better than physical modeling methods
since we can use historical data to train the model and predict new data which is dificult to do in
the physical modeling approach. In addition, the study by Jebli, Liu, Sweerts et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ][
        <xref ref-type="bibr" rid="ref7">7</xref>
        ][
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] helped
us understand the importance of addressing topics like Pearson correlation, air pollutant deposition
efects, and random forest optimization in our research. Incorporating this increased the accuracy of
the prediction models clearly indicating how diferent factors and approaches combined can enhance
solar power generation prediction.
      </p>
      <p>
        Along with machine learning models, there were a lot of studies that suggested the use of deep
learning methods for predicting solar power generation. Application of models like CNN’s and Recurrent
Neural Network (RNN)’s exhibits the efectiveness of these deep learning techniques in capturing
complex patterns and dependencies in solar generation signified by Lee et al. [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. The study by
Zazoum et al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] also evaluates the accuracy and reliability of deep learning methods in forecasting
solar PV power generation which is essential for efective grid integration and energy management.
To address the unique characteristics of the dataset, which exhibited an excess of zero values, the
researchers in the study proposed by Kim et al. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] explored alternative statistical models beyond the
traditional Poisson regression. In addition to the Poisson model, the zero-inflated model was employed,
acknowledging its ability to efectively handle datasets with an excessive proportion of observed zero
values. By employing the zero-inflated model, the researchers sought to capture the dual processes
contributing to the occurrence of zeros, distinguishing between structural zeros and excess zeros.
Thomas et al. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] also emphasizes the need for a modeling framework for univariate and multivariate
zero-inflated time series of counts. The basic modeling framework used is observation-driven Poisson
regression with a Generalized Linear Model (GLM) structure.
      </p>
      <p>
        The Zero-Inflated Poisson (ZIP) model is employed to capture the possibility of extra observed
zeros relative to the Poisson distribution, a common feature in count data. Using these insights we also
utilized zero-inflated models in our research. Yeom et al. [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] introduces a novel approach using deep
learning models, specifically ConvLSTM networks, to predict short-term solar radiation by incorporating
geostationary satellite images. The proposed model showed high accuracy in capturing cloud-induced
variations in ground-level solar radiation compared to the conventional Artificial Neural Network
(ANN) and RFR models. This paper led to us implementing ConvLSTM2D on our dataset too.
      </p>
      <p>To summarize, the reviewed papers have considerably contributed to solar power generation using
machine learning and deep learning techniques. Their research provided observations that helped us
build our research on and further enhance solar forecasting by utilizing AQI, time series-based data,
exploring novel approaches, and other diferent approaches to making solar power forecasting more
reliable and accurate.</p>
    </sec>
    <sec id="sec-4">
      <title>3. Methodology</title>
      <sec id="sec-4-1">
        <title>3.1. Dataset</title>
        <p>
          In order to accurately train our model on features that would help it efectively predict solar
generation, we needed a dataset that had high granularity, solar generation, irradiance, and
meteorological data. We thus utilized the UNISOLAR Solar Generation Dataset [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] which includes two years of
Photovoltaic solar energy generation data collected at an interval of 15 minutes at La Trobe University
Campus in Victoria, Australia. Weather data like apparent temperature, air temperature, dew point
temperature, wind speed, wind direction, and relative humidity were also provided by the dataset. We
curated the data by merging and selecting provided features such that a suitable dataset may be created
for our model. We did this by merging the data provided about the solar panels such as the number of
panels and type of inverters, with the aforementioned weather data features to obtain a comprehensive
and feature rich data of potential factors afecting solar data. In order to improve the input features of
our model, we utilized AQI data sourced from the aqicn.org website, which was captured by a station
located in Macleod, Victoria with the location showed in Fig. 1(a). This station was at a distance of 1.77
km from the University where the UNISOLAR dataset was collected (as shown in Fig. 1(b)).
        </p>
        <p>It was noticed that the solar data generated for 15-minute intervals had a number of irregularities
in the number of data collections per day. Switching to one-hour intervals eliminated this problem and
made the data intervals regular. Another reason for the use of 1 hour intervals was to reduce the noise
caused due to various external factors like equipment sensitivity, cloud cover and other atmospheric
factors.</p>
        <p>For our study, we implemented 70:30 dataset split for training and testing respectively. This
approach with 70% of the dataset dedicated to training, allows our model to efectively learn the complex
temporal dependencies and relationships inherent in solar data, thereby enhancing its predicting
accuracy. The remaining 30% served as an independent testing set to determine the eficacy of our
proposed model on unseen data.</p>
        <p>(a)
(b)</p>
      </sec>
      <sec id="sec-4-2">
        <title>3.2. Time-Series Approach</title>
        <p>Initially, a regression-based approach was utilized to predict the solar power generation based on
the factors present. However, this did not provide adequate information regarding the relationship
between these factors and solar power generation.</p>
        <p>This prompted us to try out a time series-based approach as we also had chronological data. This
is because a time series model is designed to capture patterns in sequential and discrete data points
over a set period in time. We made use of solar generation at a particular moment in time to predict
the generation at some time in the future as this type of approach would help us identify trends and
seasonality to make more accurate predictions. We decided to shape our prediction by creating a column
that would predict solar generation values 24, 48 and 72 hours out and compared the efectiveness of
these models.</p>
        <p>The Machine Learning models used for generating and comparing solar power generation for a
Time Series approach were:
• Linear Regression</p>
        <p>It is a statistical model used to establish a linear relationship between a dependent variable and
one or more independent variables. It allows for the prediction of coeficients that decide the
strength of the established linear relationship. It is used to predict the line of best fit which
minimizes the distance between actual values and predicted values.</p>
        <p>=  0 +  1 1 +  2 2 + … +     + 
(1)
• Gradient Boosting Regression</p>
        <p>It is an ensemble model which combines multiple weak prediction models like decision trees,
and takes the strongest combination of predictions to build a strong predictive model. Gradient
 ̂ = ∑   (  ) =  0(  ) +∑   ℎ (  )
• Random Forest Regression</p>
        <p>It is a supervised learning algorithm used for regression-based tasks. It combines concepts of
decision trees and ensemble learning to make precise predictions. It also reduces overfitting
due to an element of randomness which restricts individual trees from memorizing the training
set data provided to it. The decision of the algorithm is made by aggregating the decisions of
individual trees using majority voting.</p>
        <p>Boosting works by repeatedly training weak models to fit the negative gradient of the ongoing
prediction’s loss model, which helps improve its accuracy.</p>
        <p>=1

=1</p>
        <p>=1
 ̂ = 1
∑   (  )

=1

=1
(2)
(3)
(4)
• XGBoost Regression</p>
        <p>Extreme Gradient Boosting is an advanced gradient boosting algorithm that is widely used due
to the high accuracy of predictions. It achieves this by iteratively adding weak models to the
ensemble and ensuring they fit according to the current prediction. It also uses features like
column and row subsampling to further improve its performance.</p>
        <p>̂ = ∑   (  ) =  0(  ) +∑ ℎ (  )
• Random Forest Regression + XGBoost Regression</p>
        <p>Building upon the positive results displayed by Lokesh et al [20], we decided upon an ensemble
model consisting of a Random Forest Regressor and a XGBoost Regressor, considering that their
strengths are complementary in the sense that both are ensemble learners and use boosting and
bagging respectively. RFR can handle non-linear relationships between data points and XGBoost
can capture subtle patterns and temporal dependencies.
• ConvLSTM2D</p>
        <p>In this subsection, we delve into the specifics of the ConvLSTM2D model, a hybrid architecture
that merges CNNs and LSTM networks. This unique architecture is designed specifically for
analyzing spatio-temporal data, where understanding both spatial relationships and temporal
dynamics is crucial for accurate predictions.</p>
        <p>The architecture of the ConvLSTM2D model, illustrated in Figure 2, consists of layers that
facilitate the processing of spatio-temporal data. At its core are ConvLSTM units, which extend
the traditional LSTM cells by incorporating convolutional operations within the recurrent
structure. This innovative design enables the model to efectively capture both spatial and temporal
dependencies within the data. The utilization of ConvLSTM units makes the model adept at
handling sequences of spatio-temporal observations, a characteristic essential for tasks such as
video analysis, weather forecasting, and motion prediction.</p>
        <p>The training procedure of the ConvLSTM2D model follows a systematic approach tailored to
exploit its full potential. Initially, standard data preprocessing techniques, similar to those
employed for other machine learning models, are applied. However, due to the unique architecture
of ConvLSTM2D, additional reshaping of the data is necessary to conform to its input shape
requirements. The input data is reshaped into a 5-dimensional tensor, accommodating batches
of sequences, each comprising 2D matrices over time. This reshaping operation enables the
model to interpret the input data as a spatio-temporal sequence, facilitating efective learning
of complex patterns. Furthermore, to ensure robust training and prevent issues stemming from
skewed distributions, a power transformation is applied to scale input features appropriately.</p>
        <p>This normalization step fosters homogeneous feature scales, thereby preventing any individual
feature from disproportionately influencing the learning process.</p>
        <p>During the training phase, the model is optimized using the Adam optimization algorithm, with
a predefined learning rate. The optimization process involves iteratively adjusting the model’s
internal parameters to minimize the mean squared error loss between predicted and actual values.
This iterative optimization enables the model to discern intricate patterns within the data and
refine its predictive capabilities accordingly.</p>
        <p>The ConvLSTM2D model was chosen for its inherent capability to capture spatio-temporal
dependencies efectively, a critical requirement in our analysis. Through rigorous experimentation
and evaluation, it consistently outperformed alternative models, exhibiting superior performance
in terms of predictive accuracy and generalization capabilities. Its hybrid architecture, which
integrates both convolutional and recurrent operations, enables it to discern complex patterns
within spatio-temporal data, making it particularly well-suited for the tasks at hand.</p>
      </sec>
      <sec id="sec-4-3">
        <title>3.3. Zero-Inflated Models</title>
        <p>An analysis of the solar power data showed that the distribution was such that an unusually
high number of values amounted to zero, this was likely due to negligible solar generation due to the
intermittent nature of sunlight received by the solar panels. These zeros overshadowed the other values
in the distribution by a fair amount.</p>
        <p>The histogram shown in Figure 3 highlights this zero-inflated data of our initial dataset. It also
provides insights into the range of values of solar energy generation and their respective frequencies.</p>
        <p>
          Thus, on noticing this skewness of the solar energy generation data, we decided to switch to a
zero-inflated model which essentially diferentiates between structural zeros (due to nighttime and
genuine absences of solar generation) and zero inflation (additional factors that may influence the
probability of observing a zero, such as cloudy days or equipment failures). According to a study,
previous research indicates that if excessive zero is not accounted for, an unreasonable fit for both the
the zero-inflated nature of our dataset.
zeros and nonzero counts will occur (Perumean-Chaney et al. 2013). ZI (Lambert 1992) and hurdle
models (Mullahy 1986; Heilbron 1994) have been developed to model zero-inflation when the regular
count models such as Poisson or negative binomial are unrealistic [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. Thus it became a necessity
to apply such a zero-inflated model to our data that would best represent the state of the current
distribution and convert it into a distribution that might be more acceptable for our regression models
to make predictions with. After considering various models and comparing them with our data to
determine the best fit, we decided upon the Tweedie distribution.
        </p>
        <p>Tweedie Distribution
The Tweedie distribution(Figure 4) is characterized by the following components:
• Random Variable: Let  be the random variable representing the observed value.
• Mean Parameter:  is the mean parameter of the distribution, indicating the average value of  .
• Power Parameter:  is the power parameter of the distribution, controlling the variance structure.</p>
        <p>It determines the shape of the distribution and can take any positive value, excluding 1.
• Normalizing Constant: (, )</p>
        <p>is the normalizing constant that ensures the Probability Density
Function (PDF) integrates to 1 over the support of  . It accounts for the specific value of  and 
and is essential for properly defining the distribution.</p>
        <p>The probability density function of the Tweedie distribution is thus given by:
 ( ; , ) =
 −1 exp ((1−)
1−</p>
        <p>)
(, )
exp (−</p>
        <p>)</p>
        <p>(, )</p>
        <p>The Tweedie distribution encompasses various shapes, ranging from heavy-tailed with excess
zeros (for  &lt; 1 ) to symmetric and Gaussian-like (for 1 &lt;  &lt; 2 and  &gt; 2 ). It includes special cases
such as the Poisson distribution (when  = 1 ) and the gamma distribution (when  = 2 ). The choice
of  determines the specific characteristics of the Tweedie distribution, including its skewness, tail
behavior, and overall shape. This was suitable for our zero-inflated distribution, by using a modified
version of the above called a Zero Inflated Tweedie (ZIT) Model which accounts for the excess zeros
using a separate inflation component.
with probability 
with probability 1 − 


log (</p>
        <p>) =  0 +  1 1 +  2 2 + … +    
log() =  0 +  1 1 +  2 2 + … +</p>
        <p>) =  0 +  1 1 +  2 2 + … +    
where:
-  represents the response variable (count variable with excess zeros).
-  represents the positive count variable.
-  represents the probability of excess zeros.
-  represents the mean parameter.
-  represents the dispersion parameter.
-  1,  2, … ,   represent the predictor variables for the mean equation.
-  1,  2, … ,   represent the predictor variables for the dispersion equation.
-  1,  2, … ,   represent the predictor variables for the zero-inflation equation.
-  0,  1,  2, … ,   represent the coeficients for the mean equation.
-  0,  1,  2, … ,   represent the coeficients for the dispersion equation.
-  0,  1,  2, … ,   represent the coeficients for the zero-inflation equation.</p>
        <p>Now in order to accurately and eficiently apply a customized zero-inflated model to our data and
test out various regression models in Python, we utilized H2O, a Java-based software for data modeling
and general computing.</p>
        <p>H2O is an abstracted distributed processing engine that allows for simple horizontal scaling in
order to provide solutions faster and more eficiently. When it comes to model application, it consists
of numerous estimator functions like H2OGradientBoostingEstimator and H2ODeepLearningEstimator
served using REST API abstraction, each of which consists of a plethora of hyperparameters for
deep customization. For the scope of our project, we have chosen H2OGradientBoostingEstimator,
H2OXGBoostEstimator, H2ODeepLearningEstimator, and H2ORandomForestEstimator. The value of
the parameters distribution and tweedie_power were set to ‘tweedie’ and ‘1.5’ respectively for the models
and other parameters were optimized using GridSearchCV. For the Deep Learning Zero-inflated Model
(ZIM), 4 hidden layers were chosen of neuron count 100, 100, 50 and 50 respectively. Transforming our
data using a zero-inflated model resulted in a marked improvement in our solar power prediction with
a reduced MAE and RMSE.</p>
      </sec>
      <sec id="sec-4-4">
        <title>3.4. Power-Transform</title>
        <p>It is a data transformation and scaling technique which is another way to tackle the skewness of
the solar generation data; we found that using PowerTransformer was a significantly better fit for our
dataset than the zero-inflated model. This scaling method is applied feature-wise to make the data more
Gaussian or Gaussian-like which is inherently assumed by regression-based prediction models. It is
used when dealing with non-constant variance. There are 2 diferent methods of performing the power
transform, namely the Box-Cox transform and the Yeo-Johnson transform.</p>
        <p>The Yeo-Johnson power transform is given by the formula:
  =
⎧(( + 1)  − 1) /, if  ≠ 0,  ≥ 0
⎪ln( + 1), if  = 0,  ≥ 0
⎨− ((|| + 1) 2− − 1) /(2 − ), if  ≠ 2,  &lt; 0
⎪
⎩− ln(|| + 1), if  = 2,  &lt; 0</p>
        <p>Here,  represents the original variable, and  is a parameter that determines the type of power
transform applied. The transformed variable   is the result of applying the Yeo-Johnson power
transform. The aforementioned Yeo-Johnson power transform was thus applied to the data, and the
resulting distribution was notably found to resemble a Tweedie distribution.</p>
        <p>The solar energy generation data points are now normalized to make the distribution more Gaussian
as shown in Figure 5.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. Results</title>
      <p>For the prediction of solar energy generation using multiple methodologies, we have found that
the Power Transformed data led to the most accurate prediction in comparison to Regular Time Series
and Zero-Inflated models.</p>
      <p>Power Transformation of data is a particular method that stands out in comparison to the rest. This
is due to solar energy generation being dependent on various factors like temperature, seasonality, time
of day, and air quality of the region. These data points have non-linear relationships with each other.
The improvement compared to the prior models can be attributed to the power transform method’s
ability to make the data more Gaussian like. This addresses the inherent skewness present in solar data,
leading to an increase in performance across the board.</p>
      <p>ConvLSTM2D models outperform normal regression models as it combines convolutional
operations and LSTM memory cells, allowing for the modeling of both spatial and temporal dependencies
in data. It also helps capture long term dependencies eficiently in the time series data This makes
ConvLSTM2D well-suited for tasks where both spatial and sequential information are important, such
as solar power generation forecasting. Figure 7 shows the Loss vs Epoch plot for the ConvLSTM2D
model displaying the training progress.</p>
      <p>
        For evaluating the performance of the models on the gathered data, various statistical metrics were
considered, of which ultimately R2, MAE, and RMSE were used. Our key metric was R2, as the coeficient
of determination R2 is generally a better indicator of regression model performance when compared to
other metrics[
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. These metrics were tabulated and compared across the various models and for three
diferent time slots: 24 hours, 48 hours and 72 hours. There are three such tables corresponding to the
principal methodologies utilised: Regular time-series, Zero Inflated Model, and Power Transform.
      </p>
      <p>The range of target variables was between 0 and 320 for regular time series and Zero-Inflated models.
For the Zero Inflated Model, we performed five-fold cross-validation, and there was no significant
variance in the validation accuracy from the training accuracy.</p>
      <p>The range of target variables was between -0.9 and 1.8 after Power Transform was applied.</p>
      <p>After comparing all three methodologies, regression models, and variation in prediction time, the
best combination of factors respectively has been tabulated in Table V.</p>
      <sec id="sec-5-1">
        <title>Models</title>
      </sec>
      <sec id="sec-5-2">
        <title>Linear Regression</title>
      </sec>
      <sec id="sec-5-3">
        <title>GradientBoosting Regression</title>
      </sec>
      <sec id="sec-5-4">
        <title>XGBoost Regression</title>
      </sec>
      <sec id="sec-5-5">
        <title>RandomForest Regression</title>
      </sec>
      <sec id="sec-5-6">
        <title>RandomForest + XGBoost</title>
      </sec>
      <sec id="sec-5-7">
        <title>Models</title>
      </sec>
      <sec id="sec-5-8">
        <title>GradientBoosting Regression</title>
      </sec>
      <sec id="sec-5-9">
        <title>XGBoost Regression</title>
      </sec>
      <sec id="sec-5-10">
        <title>RandomForest Regression</title>
      </sec>
      <sec id="sec-5-11">
        <title>Deep Learning</title>
        <p>The solar power generation data when plotted monthly follows a specific pattern that can be
attributed to the seasonal cycle of the Australian landmass, where the dataset was sourced from. The
generation is noted to be maximum from November to February which coincides with the summer
months in Australia and reaches its minimum during the months of May to August which are the winter
months. This can be explained due to the sun being directly overhead during summer, leading to longer
days and more exposure to solar radiation. This phenomenon reverses during the winter months as is
shown by the histogram (Fig. 8).</p>
        <p>Our time-series based solar power prediction models also capture this phenomenon as seen in the
graphs plotted as shown in Fig. 9 and Fig. 10. There is a drop in solar production during the winter
months of 2020 and the peak production is reached in the summer of 2021.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>5. Conclusion</title>
      <p>This study investigates the use of various machine learning algorithms to efectively determine the
future solar generation in a region by utilizing a time series approach. The chief models employed were
Linear Regression, Lasso, Ridge, ElasticNet, ensemble models like RandomForest and XGBoost, and deep
learning models like ConvLSTM2D. Later, on conferring, we decided to switch to a time-series based
approach due to the seasonal fluctuation in the solar data. However, on noticing the skewness of the
solar energy generation data which had a high number of zeros, we decided to switch to a zero-inflated
model which helps ascertain the diference between the true zero data points and the inflated zeros.
This approach immediately yielded a higher accuracy of solar prediction with a lower mean standard
error. Another way to tackle the skewness of the solar generation data was to try diferent scaling
techniques; we found that using PowerTransformer was a significantly better fit for our dataset than
the zero-inflated model. This scaling method is applied feature-wise to make the data more Gaussian or
Gaussian-like which is inherently assumed by regression-based prediction models. In conclusion, this
study investigates the use of machine learning algorithms incorporating AQI and climate factors to
provide more accurate solar generation forecasts. Considering seasonal variations in the solar data, the
time series-based method was modified. In addition, a zero-inflated model and scaling techniques were
used to address the skewness of the solar generation data. The findings provide valuable insights for
solar stakeholders, contributing to the adoption and use of sustainable energy sources.</p>
    </sec>
    <sec id="sec-7">
      <title>6. Future Scope</title>
      <p>Solar energy generation forecasting is a dynamic field that will always develop and demand
exploration. There is a lot of scope in the future to improve the accuracy of solar forecasting. To make
this forecasting scalable and generalized for all geographical regions, the dataset can be collected from
diferent regions and time periods since our research primarily utilizes data from a specific region and
time zone. Additionally, data fusion and feature engineering can be done to enhance the power of
forecasting. Data fusion is the practice of collecting data from multiple sources like satellite imagery
and weather stations which can be combined together for better results. There was a limitation on the
dataset which we faced during implementation which was the fact that the AQI data we used wasn’t
from a data station at the exact geographical location of the data collection but rather another data
station away from the solar site. Thus in order to increase accuracy one can take AQI data from a
data station at maximum proximity to the data collection center in order to eliminate susceptibility to
geographical errors. Another possibility can be to collect our own data to create a localized dataset
which would likely be more accurate since we can customize it and remove any bottlenecks faced earlier
by having utilised an external dataset. Moreover, feature engineering can be employed to extract more
meaningful features from the data available, hence increasing the model performance. Features like
cloud cover, dust, and further manners of seasonality can also be taken into consideration for future
research.</p>
    </sec>
    <sec id="sec-8">
      <title>Declaration on Generative AI</title>
      <p>The authors have not employed any Generative AI tools.
2021 Jul 5;7:e623. doi: 10.7717/peerj-cs.623. PMID: 34307865; PMCID: PMC8279135.
[19] O. Edenhofer et al., Eds., Renewable Energy Sources and Climate Change Mitigation: Special Report
of the Intergovernmental Panel on Climate Change. Cambridge: Cambridge University Press, 2011.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Chuluunsaikhan</surname>
            ,
            <given-names>Tserenpurev.</given-names>
          </string-name>
          (
          <year>2021</year>
          ).
          <article-title>Predicting the Power Output of Solar Panels based on Weather and Air Pollution Features using Machine Learning</article-title>
          .
          <source>Journal of Korea Multimedia Society. 24. 222. 10</source>
          .9717/kmms.
          <year>2021</year>
          .
          <volume>24</volume>
          .2.222.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yan</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Du</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          (
          <year>2021</year>
          ).
          <article-title>Deep Learning Enhanced Solar Energy Forecasting with AI-Driven IoT</article-title>
          .
          <source>Wireless Communications and Mobile Computing</source>
          ,
          <year>2021</year>
          ,
          <string-name>
            <surname>Article</surname>
            <given-names>ID</given-names>
          </string-name>
          9249387. doi:
          <volume>10</volume>
          .1155/
          <year>2021</year>
          /9249387.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Ghosh</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dey</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ganguly</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roy</surname>
            ,
            <given-names>S. B.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Bali</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          (
          <year>2022</year>
          ).
          <article-title>Cleaner air would enhance India's annual solar energy production by 6-28 TWh</article-title>
          . Environmental Research Letters,
          <volume>17</volume>
          (
          <issue>5</issue>
          ), 054007. doi:
          <volume>10</volume>
          .1088/
          <fpage>1748</fpage>
          -9326/ac5d9a.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Galimova</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ram</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Breyer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          (
          <year>2022</year>
          ).
          <article-title>Mitigation of air pollution and corresponding impacts during a global energy transition towards 100% renewable energy system by 2050</article-title>
          .
          <source>Energy Reports</source>
          ,
          <volume>8</volume>
          ,
          <fpage>14124</fpage>
          -
          <lpage>14143</lpage>
          . ISSN 2352-
          <fpage>4847</fpage>
          . doi:
          <volume>10</volume>
          .1016/j.egyr.
          <year>2022</year>
          .
          <volume>10</volume>
          .343.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Jia</surname>
            ,
            <given-names>Dongyu</given-names>
          </string-name>
          &amp; Yang, Liwei &amp; Lv, Tao &amp; Liu,
          <string-name>
            <surname>Weiping</surname>
          </string-name>
          &amp; Gao, Xiaoqing &amp; Zhou,
          <string-name>
            <surname>Jiaxin.</surname>
          </string-name>
          (
          <year>2022</year>
          ).
          <article-title>Evaluation of machine learning models for predicting daily global and difuse solar radiation under diferent weather/pollution conditions</article-title>
          .
          <source>Renewable Energy</source>
          .
          <volume>187</volume>
          . 10.1016/j.renene.
          <year>2022</year>
          .
          <volume>02</volume>
          .002.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Jebli</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Belouadha</surname>
            ,
            <given-names>F.-Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kabbaj</surname>
            ,
            <given-names>M. I.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Tilioua</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2021</year>
          ).
          <article-title>Prediction of solar energy guided by Pearson correlation using machine learning</article-title>
          .
          <source>Energy</source>
          ,
          <volume>224</volume>
          , 120109. doi:
          <volume>10</volume>
          .1016/j.energy.
          <year>2021</year>
          .120109
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>Random forest solar power forecast based on classification optimization</article-title>
          .
          <source>Energy</source>
          ,
          <volume>115940</volume>
          . doi:
          <volume>10</volume>
          .1016/j.energy.
          <year>2019</year>
          .115940
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Sweerts</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pfenninger</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Folini</surname>
            , D., van der Zwaan,
            <given-names>B.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Wild</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>Estimation of losses in solar energy production from air pollution in China since 1960 using surface radiation data</article-title>
          .
          <source>Nature Energy. doi:10.1038/s41560-019-0412-4</source>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schwede</surname>
            ,
            <given-names>D. B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Appel</surname>
            ,
            <given-names>K. W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mangiante</surname>
            ,
            <given-names>M. J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wong</surname>
            ,
            <given-names>D. C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Napelenok</surname>
            ,
            <given-names>S. L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Whung</surname>
          </string-name>
          , P.-Y., &amp;
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>The impact of air pollutant deposition on solar energy system eficiency: an approach to estimate PV soiling efects with the Community Multiscale Air Quality (CMAQ) model</article-title>
          .
          <source>Science of the Total Environment, 651(Pt 1)</source>
          ,
          <fpage>456</fpage>
          -
          <lpage>465</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.scitotenv.
          <year>2018</year>
          .
          <volume>09</volume>
          .194
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Zazoum</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          (
          <year>2022</year>
          ).
          <article-title>Solar photovoltaic power prediction using diferent machine learning methods</article-title>
          .
          <source>Energy Reports</source>
          ,
          <volume>8</volume>
          (
          <issue>Supplement 1</issue>
          ),
          <fpage>19</fpage>
          -
          <lpage>25</lpage>
          . ISSN 2352-
          <fpage>4847</fpage>
          . doi:
          <volume>10</volume>
          .1016/j.egyr.
          <year>2021</year>
          .
          <volume>11</volume>
          .183.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>C.-H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>H.-C.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Ye</surname>
          </string-name>
          , G.-B. (
          <year>2021</year>
          ).
          <article-title>Predicting the Performance of Solar Power Generation Using Deep Learning Methods</article-title>
          .
          <source>Applied Sciences</source>
          ,
          <volume>11</volume>
          (
          <issue>15</issue>
          ),
          <volume>6887</volume>
          . doi:
          <volume>10</volume>
          .3390/app11156887
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Chiteka</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arora</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sridhara</surname>
            ,
            <given-names>S. N.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Enweremadu</surname>
            ,
            <given-names>C. C.</given-names>
          </string-name>
          (
          <year>2020</year>
          ).
          <article-title>A novel approach to Solar PV cleaning frequency optimization for soiling mitigation</article-title>
          .
          <source>Scientific African</source>
          ,
          <volume>8</volume>
          , e00459.
          <source>ISSN 2468-2276</source>
          . doi:
          <volume>10</volume>
          .1016/j.sciaf.
          <year>2020</year>
          .e00459.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>D.-W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Deo</surname>
            ,
            <given-names>R. C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Park</surname>
          </string-name>
          , S.-J.,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>J.-S.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>W.-S.</given-names>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>Weekly heat wave death prediction model using zero-inflated regression approach</article-title>
          .
          <source>Theoretical and Applied Climatology</source>
          ,
          <volume>137</volume>
          ,
          <fpage>823</fpage>
          -
          <lpage>838</lpage>
          . doi:
          <volume>10</volume>
          .1007/s00704-018-2632-4
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Thomas</surname>
            ,
            <given-names>S. J.</given-names>
          </string-name>
          (
          <year>2010</year>
          ).
          <article-title>Model-based clustering for multivariate time series of counts</article-title>
          . Rice University. ProQuest Dissertations Publishing. (Publication No.
          <volume>3421317</volume>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Yeom</surname>
            ,
            <given-names>J. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Deo</surname>
            ,
            <given-names>R. C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Adamowski</surname>
            ,
            <given-names>J. F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Park</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>C. S.</given-names>
          </string-name>
          (
          <year>2020</year>
          ).
          <article-title>Spatial mapping of shortterm solar radiation prediction incorporating geostationary satellite images coupled with deep convolutional LSTM networks for South Korea</article-title>
          .
          <source>Environmental Research Letters</source>
          ,
          <volume>15</volume>
          (
          <issue>9</issue>
          ), 094025. doi:
          <volume>10</volume>
          .1088/
          <fpage>1748</fpage>
          -9326/ab9467
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>S.</given-names>
            <surname>Wimalaratne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Haputhanthri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kahawala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Gamage</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Alahakoon</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Jennings</surname>
          </string-name>
          , ”
          <article-title>UNISOLAR: An Open Dataset of Photovoltaic Solar Energy Generation in a Large Multi</article-title>
          -Campus University Setting,”
          <source>2022 15th International Conference on Human System Interaction (HSI)</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>5</lpage>
          , doi: 10.1109/HSI55341.
          <year>2022</year>
          .986947
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Feng</surname>
            ,
            <given-names>C.X.</given-names>
          </string-name>
          <article-title>A comparison of zero-inflated and hurdle models for modeling zero-inflated count data</article-title>
          .
          <source>J Stat Distrib App</source>
          <volume>8</volume>
          ,
          <issue>8</issue>
          (
          <year>2021</year>
          ). https://doi.org/10.1186/s40488-021-00121-4
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Chicco</surname>
            <given-names>D</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Warrens</surname>
            <given-names>MJ</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jurman</surname>
            <given-names>G.</given-names>
          </string-name>
          <article-title>The coeficient of determination R-squared is more informative than SMAPE, MAE, MAPE, MSE and RMSE in regression analysis evaluation</article-title>
          .
          <source>PeerJ Comput Sci.</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>