<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Analyzing climate changes impacts using big data Hadoop</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Silvana Greca</string-name>
          <email>silvana.greca@fshn.edu.al</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ingrid Shehi</string-name>
          <email>ingrid.shehi@fshnstudent.info</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jonuz Nuhi</string-name>
          <email>jonuz.nuhi@fshnstudent.info</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Tirana, Faculty of Natural Sciences, Department of Informatics</institution>
          ,
          <addr-line>Bulevardi Zogu I, Tirana</addr-line>
          ,
          <country country="AL">Albania</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper aims to analyze climate change data using Hive and Hadoop, two big data processing frameworks. We collected data from various sources and used Hive to store and manage the data, and Hadoop to process it. By using these tools, we were able to perform complex queries and analysis on large datasets with ease. This paper also used Super Set, a data visualization tool, to create interactive dashboards that display the results of our analysis. The dashboards help users to explore the data and gain insights into climate change trends. Our findings show that the temperature in the city of Durrës has increased by 1.1°C since the pre-industrial era. This paper demonstrates the usefulness of big data processing tools for analyzing climate change data and provides valuable insights into the impact of global warming on our planet.</p>
      </abstract>
      <kwd-group>
        <kwd>1 Climate change</kwd>
        <kwd>big data processing</kwd>
        <kwd>Hadoop</kwd>
        <kwd>Hive</kwd>
        <kwd>Data analysis</kwd>
        <kwd>Data visualization</kwd>
        <kwd>Superset</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>In recent years, the rapid growth of data has
become a ubiquitous challenge in various
domains, including climate science. The scale
and complexity of climate data pose significant
challenges to traditional data processing and
analysis techniques. However, big data
processing tools, such as Hadoop and Hive, have
emerged as promising solutions to handle
largescale data processing and analysis.</p>
      <p>Hadoop is a distributed computing framework
that can store and process large datasets in
parallel, while Hive is a data warehousing tool
that provides a SQL-like interface for managing
and querying large datasets stored in Hadoop.</p>
      <p>These tools have already been used in various
domains, including climate science, to process
and analyze large datasets efficiently. In
particular, Hadoop's distributed processing
capabilities make it well-suited for analyzing
large climate datasets.</p>
      <p>In this paper, we present a study that leverages
Hadoop and Hive to analyze climate change data.
Our study collects data from various sources,
including temperature measurement data, and
uses Hadoop's distributed processing capabilities
to store and process the data in parallel. We then
use Hive to manage and analyze the data,
including running complex queries and creating
visualizations. Finally, we use Superset, a data
visualization tool, to create interactive
dashboards that provide insights into climate
change trends.</p>
      <p>The objective of this paper is to demonstrate
the usefulness of big data processing tools for
analyzing climate change data and to provide
valuable insights into the impact of global
warming on our planet. The paper is structured as
follows: in the next section, we review related
work on using big data processing tools for
climate science. We then describe the
methodology used in our study, including data
collection, processing, and analysis. Next, we
present our results and discuss the implications of
our findings. Finally, we conclude the paper with
a summary of our contributions and future
research directions.</p>
      <p>By using big data processing tools, we can
gain deeper insights into climate change patterns
and trends, which can help inform policy
decisions and mitigation strategies. We hope that
this study will inspire further research on the use
of big data processing tools for climate science
and contribute to the ongoing efforts to address
the urgent issue of climate change.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Literature overview</title>
      <p>The use of big data processing tools such as
Hadoop and Hive for analyzing climate change
impacts has been gaining increasing attention in
recent years. In this section, we will discuss some
of the notable studies that have leveraged these
tools to analyze climate data and investigate the
impacts of global warming on our planet.</p>
      <p>
        One example of such studies is the work of
Hossein Hassani, Xu Huang, Emmanuel Silva.
[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], who used Big Data to analyze climate data
collected from multiple sources. Their study
focused on identifying and characterizing
heatwaves and their impacts on human health. By
leveraging the processing power of Hadoop, they
were able to efficiently process large volumes of
data and identify significant trends and patterns.
Their results demonstrated the potential of big
data processing tools for analyzing the complex
and dynamic impacts of climate change.
      </p>
      <p>
        Another notable study is the work of
HuaDong Guo, Li Zhang, Lan-Wei Zhu. [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], who
used Big Data. Their study reviews the advances
of climate change studies based on Earth
observation big data and provides examples of
case studies that utilize Earth observation big data
in climate change research, such as synchronous
satellite–aerial–ground observation experiments,
which provide extremely large and abundant
datasets. With the era of global environment
change dawning, Earth observation big data will
underpin the Future.
      </p>
      <p>
        In a similar vein, the work of Thanos
Papadopoulos, M.E. Balta [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] leveraged big data
analytics to analyze climate data for various
challenges. Their research addresses the impact
of climate change on businesses, operations, and
supply chains by identifying and discussing how
these challenges and opportunities can be better
pursued. Their opinion paper proposes ideas
about future research on BDA and climate change
could follow to facilitate the transition to a
sustainable future.
      </p>
      <p>Overall, these studies highlight the potential
of big data processing tools for analyzing climate
change impacts and providing valuable insights
into the complex and dynamic relationships
between climate and the environment. By
leveraging the processing power of tools for Big
Data, researchers can efficiently process and
analyze large volumes of data, which can inform
policy decisions and mitigation strategies for
addressing the urgent issue of climate change.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Hadoop and it’s functionality</title>
      <p>
        Hadoop is an open-source framework that is
used to efficiently store large amount of data from
gigabytes to petabytes. Instead of using one large
computer to store and process the data, Hadoop
uses a set of clusters to analyze massive datasets
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The processing of datasets is done through
HDFS. HDFS enables the rapid transfer of data
between computer nodes. When the data is ready
to be processed, it is sent to Map Reduce and split
into smaller datasets. After the data is divided
each of those datasets get a mapping function.
Then the shuffling start which just send the data
from the mapper to the reducers. The Reducing
stage is about summarizing the effects of the
previous stages and reducing them to a small set
of values [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The use of Hadoop is very
costeffective because the commodity hardware is
very cheap. HDFS besides storing large amount
of data it can store them in different formats.
Since Hadoop process the data in parallel the
speed of our output is very fast and also Hadoop
creates a duplicate of the data in all the nodes so
if a node crash or burn, we have a copy of it so
data is not lost as it is shown in the figure 1.
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. The Hive</title>
      <p>
        Hive, an open-source data warehousing
solution tool built on top of Hadoop Distributed
File System (HDFS) that allows users to query
and analyze large datasets using a SQL-like
language called HiveQL [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>The connection between Hive and Hadoop is
established through the use of a driver program
that acts as a mediator between Hive and the
Hadoop cluster as is shown in the figure 2.</p>
      <p>
        The driver program is responsible for
translating HiveQL queries into MapReduce jobs,
submitting them to the Hadoop cluster for
processing, and returning the results back to
Hive. Results are aggregated and returned to the
user as the output of the Hive query [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
    </sec>
    <sec id="sec-5">
      <title>4.1 Hive components</title>
      <p>
        • Hive Client - Hive allows writing
applications in various languages. It supports
different types of clients such as: Thrift
Server, JDBC Driver, ODBC Driver.
• Metastore - The Metastore is a database
that stores metadata about the tables and
partitions in Hive. It keeps track of the
schema, location, and other properties of the
data stored in Hive. The Metastore supports
multiple databases, including MySQL,
PostgreSQL, and Derby.
• HiveQL - HiveQL is a SQL-like
language that is used to query data stored in
Hive. It provides a familiar syntax for SQL
users and supports many SQL operations,
including joins, aggregations, and subqueries.
• Driver - The Driver is responsible for
receiving queries from the user, generating an
execution plan, and coordinating the
execution of the plan with the other
components of Hive.
• Compiler - The Compiler takes the query
plan generated by the Driver and generates a
series of MapReduce or Tez jobs to be
executed on the Hadoop cluster. The Compiler
is responsible for optimizing the query plan to
improve performance.
• Execution Engine - The Execution
Engine is responsible for executing the
MapReduce or Tez jobs generated by the
Compiler. The Execution Engine runs on the
Hadoop cluster and processes the data stored
in HDFS.
• Hive Server - The Hive Server is a server
process that exposes a Thrift interface,
allowing clients to connect to Hive and
execute queries using various programming
languages, such as Java, Python, and R. Hive
Server 2 (HS2) is the preferred version of the
server [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
    </sec>
    <sec id="sec-6">
      <title>5. Using Superset to visualize the data</title>
      <p>Data is stored in tabular mode. To help
viewers to understand the data, a visualization
tool is needed. Apache Superset is a tool from
ASF (Apache Software Foundation) which
provides a way to data visualization and
exploration from simple line to highly detailed
geospatial charts. It allows integration with most
relational databases. Hive is used as a data source
from superset.</p>
    </sec>
    <sec id="sec-7">
      <title>5.1 Superset and Hive connection to visualize data</title>
      <p>The chart in the figure 4 shows the flow of our
work. First, we configured Hadoop in our
environment and after that we configured Hive, a
tool to work with our data. We extracted the data
in Hive and after that we configured Superset.
After Superset was successfully configured, we
connected it with Hive so it could visualize our
data and Superset also uses queries so we can
filter the data we want to visualize.</p>
      <p>
        The recommended way to connect Superset
with hive is by using the pyhive library. The
pyhive library is a python interface for interacting
with hive databases. For connection to a Hive
database, you will need to configure a new data
source in the Superset interface and specify the
connection details for the Hive database. Once
the connection is established, Superset can
execute HiveQL queries on the data stored in
Hive. Superset can’t store the data in itself it gets
the data from the existing database. When you
create a new chart or dashboard in Superset, you
can write a SQL query to pull the data from Hive,
or use the Superset Query Builder interface to
generate the query for you [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
    </sec>
    <sec id="sec-8">
      <title>5.2 Advantages of Superset</title>
      <p>
        As a visualization tool Superset stands out
because using it the user have total access over
the data. Using Superset you can add users to
your database, provide access to them and track
their behavior. To create a visualization of the
data we can create query in Superset and the
queries are created the same way as in any SQL
based database. Since Superset is a visualization
tool it is created for non-programmers so it can be
used by anyone with a basic understanding of
SQL. Superset is also accessible as web
application or app [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. The figure 5 shows the
connection between Superset and Hive.
      </p>
    </sec>
    <sec id="sec-9">
      <title>6. Big Data Climate Change and visualizing</title>
    </sec>
    <sec id="sec-10">
      <title>6.1 About climate change data</title>
      <p>
        To use Hadoop, we have created an ubuntu
environment and everything we have used is in a
live server. To secure our work environment we
have used a firewall. A live server is a term used
to describe a server that is currently in operation
and accessible on the internet or a network. It is a
server that is running and serving web pages,
applications, or other content to clients who are
requesting it. A firewall is a security system that
is designed to prevent unauthorized access to or
from a private network. It acts as a barrier
between a trusted network (such as a corporate or
home network) and an untrusted network (such as
the internet) [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        Some say climate change is the biggest threat
of our age. There are a lot of organizations that
collect and use climate trends data. We have
taken our data from NOAA’s MLOST, NASA’S
GISTEMP and UK’s HADCRUT [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
      <p>
        The newer data is collected by the Berkley
Earth, which is affiliated with Lawrence National
Library. The Berkeley Earth Surface
Temperature Study combines 1.6 billion
temperature reports from 16 pre-existing
archives. It is nicely packaged and allows for
slicing into interesting subsets (for example by
country). They publish the source data and the
code for the transformations they applied [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
    </sec>
    <sec id="sec-11">
      <title>6.2 Dataset Overview</title>
      <p>This Dataset focuses on data related to the city
of Durrës, Albania, covering a time period
between 1744 and 2013. There are six columns of
data: date, country, city, average temperature,
longitude and latitude. Each column has a
specific type of data associated with it. The date
column contains information about the time
period during which the climate data was
recorded. This column is typically represented as
a timestamp or date format, such as yyyy-mm-dd.
The country and city column contains
information about the geographic location where
the climate data was collected. This column is
typically represented as a string data type, with
the name of the country and city. These data
column is essential for analyzing climate data for
specific regions and identifying trends and
patterns in climate data for different countries.
The average temperature column contains
information about the temperature at the location
where the data was collected. This column is
typically represented as a numeric data type, with
the temperature value recorded in degrees
Celsius. The average temperature column is
essential for analyzing climate data for
temperature trends over time and identifying
changes in temperature patterns due to climate
change. The longitude column contains
information about the east-west position of a
point on the earth's surface. The latitude provides
information about the north-south position of a
particular location where the temperature data
was recorded. It is typically represented as a
numeric data type, with the latitude value
recorded in degrees. In summary, each column in
a climate data file contains specific types of data
that are critical for analyzing climate data and
identifying trends and patterns in temperature and
climate variables over time. The date, country,
average temperature, and longitude columns
provide essential information that enables
researchers to gain insights into the impact of
climate change on the environment and identify
strategies for mitigating its effects. The figure 6
shows the format of the data file.</p>
    </sec>
    <sec id="sec-12">
      <title>6.3 Visualizing data</title>
      <p>Superset is a popular open-source Business
Intelligence tool that allows users to easily create
visualizations and dashboards based on their data.
After data is stored successfully, we started using
superset so we could make data visualization.
Superset is deployed in docker. Docker is a
popular platform for deploying and managing
containerized applications. Running Superset in
Docker allows you to easily deploy and manage
Superset as a containerized application. While
working with this data we have made an
analyzation of the data for Albania like finding
the maximum, minimum and average of
temperatures in different times. Figures 7 through
10 display a selection of these examples.
The area chart graphic from Superset
visualization is a powerful tool for visualizing
data related to climate change as it is shown in
figure 11. By plotting the average temperature
from 1800, 1900, and 2000 for the city of Durres
in Albania, we can see the stark changes in
temperature over the past century. The graphic
shows a clear upward trend, with temperatures
steadily increasing from 1800 to 2013. This
visual representation of the data makes it easy to
see the significant impact that climate change has
had on the temperature in Durres. It highlights the
urgency of taking action to address this issue, as
well as the importance of using data visualization
tools like Superset to help us better understand
and communicate complex environmental issues.</p>
    </sec>
    <sec id="sec-13">
      <title>7. Conclusions</title>
      <p>In this work, we have demonstrated the
effectiveness of Hadoop, Hive, and Superset in
managing and analyzing data related to climate
change. The use of Hadoop allowed us to process
and store large amounts of data, while Hive
provided a powerful SQL-like interface for
querying and analyzing the data. Finally,
Superset enabled us to visualize the data in
interactive dashboards and charts, providing a
comprehensive view of our findings. Through our
analysis, we have identified significant patterns
and trends related to climate change. Our findings
indicate that global temperatures are rising, and
extreme weather events are becoming more
frequent. These changes have far-reaching
consequences for our planet, including the loss of
biodiversity, the displacement of communities,
and the exacerbation of social and economic
inequalities. Found out that the data have been
increased from 1700 to 2013 with an average of
about 1.1°C. In conclusion, our work
demonstrates the power of advanced technologies
like Hadoop, Hive, and Superset in analyzing
complex issues like climate change. By
leveraging these tools, we were able to gain
critical insights into the impact of climate change
and develop evidence-based strategies to address
it. Moving forward, we believe that the use of
advanced technologies and sound research
methods will be essential in addressing the
pressing issues facing our planet, including
climate change.</p>
    </sec>
    <sec id="sec-14">
      <title>8. References</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>White</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>Hadoop: The definitive guide (4th ed</article-title>
          .).
          <source>O'Reilly Media.</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2017</year>
          ).
          <article-title>Hadoop: Processing large datasets with HDFS and MapReduce</article-title>
          .
          <source>Journal of Big Data</source>
          ,
          <volume>4</volume>
          (
          <issue>1</issue>
          ),
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Hive</given-names>
            <surname>: A Data Warehousing</surname>
          </string-name>
          <article-title>Tool on Hadoop</article-title>
          .
          <source>Journal of Computer Science and Technology</source>
          ,
          <volume>34</volume>
          (
          <issue>1</issue>
          ),
          <fpage>1</fpage>
          -
          <lpage>21</lpage>
          . doi:
          <volume>10</volume>
          .1007/s11390-019-1909-
          <volume>3</volume>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <article-title>[4] Building a high-level dataflow system on top of Map-Reduce: The Pig experience</article-title>
          .
          <source>Proceedings of the VLDB Endowment</source>
          ,
          <volume>2</volume>
          (
          <issue>2</issue>
          ),
          <fpage>1414</fpage>
          -
          <lpage>1425</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Programming</given-names>
            <surname>Hive</surname>
          </string-name>
          :
          <article-title>Data Warehouse and Query Language for Hadoop" by Edward Capriolo, Dean Wampler</article-title>
          , and Jason Rutherglen
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Ganta</surname>
            ,
            <given-names>S. R.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Ranganatham</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2020</year>
          ).
          <article-title>Analyzing and Visualizing Big Data Using Hive and Superset</article-title>
          .
          <source>International Journal of Computer Sciences and Engineering</source>
          ,
          <volume>8</volume>
          (
          <issue>7</issue>
          ),
          <fpage>263</fpage>
          -
          <lpage>268</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          , Han,
          <string-name>
            <given-names>R.</given-names>
            , &amp;
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          (
          <year>2020</year>
          ).
          <source>Application of Superset in Big Data Analysis Platform. 2nd International Conference on Computer Science and Software Engineering</source>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Neupane</surname>
            ,
            <given-names>Kishan</given-names>
          </string-name>
          &amp; Haddad, Rami &amp; Chen,
          <string-name>
            <surname>Lei.</surname>
          </string-name>
          (
          <year>2018</year>
          ).
          <article-title>Next Generation Firewall for Network Security: A Survey. 1-6</article-title>
          .
          <fpage>10</fpage>
          .1109/SECON.
          <year>2018</year>
          .
          <volume>8478973</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Climate</given-names>
            <surname>Change: Earth Surface Temperature Data</surname>
          </string-name>
          https://www.kaggle.com/datasets/berkeleye arth/
          <article-title>climate-change-earth-surfacetemperature-data</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Hassani</surname>
            ,
            <given-names>Hossein</given-names>
          </string-name>
          &amp; Huang, Xu &amp; Silva,
          <string-name>
            <surname>Emmanuel.</surname>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>Big Data and Climate Change</article-title>
          .
          <source>Big Data and Cognitive Computing</source>
          .
          <volume>3</volume>
          (
          <issue>1</issue>
          ).
          <volume>10</volume>
          .3390/bdcc3010012.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Hua-Dong</surname>
            , Guo &amp; Zhang, Li &amp; Zhu,
            <given-names>Lanwei.</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>Earth observation big data for climate change research</article-title>
          .
          <source>Advances in Climate Change Research</source>
          .
          <volume>6</volume>
          . 10.1016/j.accre.
          <year>2015</year>
          .
          <volume>09</volume>
          .007.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Papadopoulos</surname>
            ,
            <given-names>Thanos</given-names>
          </string-name>
          &amp; Balta,
          <string-name>
            <surname>Maria.</surname>
          </string-name>
          (
          <year>2021</year>
          ).
          <article-title>Climate Change and big data analytics: Challenges and opportunities</article-title>
          .
          <source>International Journal of Information Management. 63. 102448. 10</source>
          .1016/j.ijinfomgt.
          <year>2021</year>
          .
          <volume>102448</volume>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>