<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Engineering and
Technology International Journal of
Computer and Information Engineering</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Algorithms for Better Decision Making</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Olta Llaha</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Azir Aliu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>prevention process. The Database Management</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2015</year>
      </pub-date>
      <volume>9</volume>
      <issue>10</issue>
      <fpage>0000</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>Nowadays, huge and accessible data is an ever increasing field of study. The change in technology is, in turn, increasing its degree of interactivity, configuring several scenarios of great complexity in which data is understood on the basis of our interaction with it at different levels. Data visualization involves presenting data in graphical or pictorial form which makes the information easy to take in. It helps to explain facts and determine courses of action. Criminology is an interesting application where data visualization plays an important role in terms of prediction and analysis. Crime analysis plays an important role in devising solutions to crime problems and formulating crime prevention strategies. The purpose of this paper is to evaluate the performance of machine learning algorithms, which can be used for analyzing data collected of the past crimes. We identified the most appropriate machine learning algorithm to analyze the collected data from sources specialized in crime prevention. This study helps the institutions against crime to better predict and classify it. Data Visualization, Machine learning, Decision making models that may be important in the crime System is designed for case management and overall crime counting and not for data analysis. The addition of machine learning algorithms to</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Crime is a complex social phenomenon that
has grown due to major changes in society. Law
enforcement agencies need to learn the factors
that lead to an increase in crime tendency. This
study focuses on crime prevention, which is an
important component of an overall strategy to
reduce crime and to strengthen public safety.</p>
      <sec id="sec-1-1">
        <title>Decision</title>
        <p>Making
in
crime
prevention
has
attracted a great concern and attention. Decision
making is very important in crime prevention in
order to
enforcement
decide
accurate
actions
and
law
strategies.</p>
        <p>Law
enforcement
agencies face a large volume of data that needs to
be processed and turned into useful information.
Data visualization approach has been exposed to
be
proactive
decision-making
concept in
preventing and predicting crime. By processing
criminal data, law enforcement agencies can use
Albania</p>
        <p>2023 Copyright for this paper by its authors. Use permitted under Creative
the
database
management system
enables a
system that functions as a criminology expert.
Crime analysis can produce a superior result by
integrating</p>
        <p>machine learning algorithms into
Decision Support System (DSS). The DSS is
required in crime analysis because it has the
capability to improve the quality of decision
making for crime prevention.</p>
        <p>Figure 1 shows the Conceptual framework of
DSS in crime prevention. In the system, the use of
machine learning algorithms is also noted, which
will make it possible to create models and as an
output will determine knowledge about crimes,
such as crime trends, crime location, etc.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Data visualization and machine learning for decision making</title>
      <p>
        The primary purpose for data visualization is
to assist people with processing large amounts of
information. Data volumes are large and human
cognitive capacities to remember and understand
data are limited [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Data visualization should be
made to simplify visualization as much as
possible to help people make more effective
decisions. Put simply, data visualization is a
method of producing an output so that all
problems and solutions can be clearly seen by the
domain experts [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Data visualization plays an
important role in terms of prediction and analysis.
models in order to identify problems, solve
problems and make decisions. Decision support
systems (DSS) are classically designed to serve
the management level of organizations [5]. They
help managers, in this case law enforcement
agencies make decisions that are semi structured,
unique or rapidly changing and are not easily
specified in advance. DSS use sophisticated
analysis and modeling tools. Data visualization
and machine learning extend the possibilities for
decision support by discovering patterns and
relationships hidden in data and therefore
enabling the inductive approach of data analysis
[6].
      </p>
      <p>The process of DSS in general is shown in fig.
3. DSS supports the ability to import various
training datasets and test ML models in .csv, .xls
and .json formats.</p>
      <p>To display the relationship between users and
the system, a diagram of cases was compiled (fig.
4). The main unary scenarios of user interaction
with the system are: selecting a data set, viewing
data statistics, updating the task, viewing models
rating and visualization.</p>
      <p>The paper introduces data visualization and
machine learning based on decision support
systems. It is designed to enable the using of
decision support systems by just having a basic
level of knowledge on data visualization and
machine learning algorithms.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology</title>
      <p>In this study we compare data visualization
techniques and machine learning algorithms to
discover the most suitable method or algorithm
for our data. The machine learning algorithms are
compared by applying them to the WEKA [8]
environment. The implemented algorithms are:
EM, COBWEB, DBCSCAN, Hierarchical
Cluster, Make Density Based Cluster, K-Mean,
Farthest First, Filtered Cluster. Algorithms have
been applied to these data to determine their
effectiveness in crime prediction and prevention.
The data analyzed is extracted from the database
of law enforcement agencies. The collected data
is stored into database for further process. The
number of instances or records is 90. There are 90
records because the data is very sensitive and we
could not get more data from the law enforcement
agencies, for this reason. This is also a limitation
of this article. Reduced number of data due to their
sensitivity.</p>
      <p>Table 1
Dataset details</p>
      <p>The data relates to areas where crimes occur
and to the information about the perpetrators.
Some of the features we have considered are: the
area where the crime occurred (urban or rural),
age (from 17 to 55 years old), employment status
(whether employed or not), gender, education
(middle school, high school, university), civil
status (whether married, single, or divorced) and
whether the person who committed the crime was
previously convicted or not. Crime dataset is in
csv format.
3.1.</p>
    </sec>
    <sec id="sec-4">
      <title>Clustering</title>
      <p>Clustering is a partitioning of data into groups
of similar objects. Presenting data from a few of
these groups certainly loses some details, but it
achieves simplicity. It models the data according
to his groupings. From a machine learning
perspective, groups correspond to hidden patterns,
group search is learned without supervision and
the final system presents a data concept.
Clustering is the main subject of active research
in various fields such as statistics, pattern
recognition and machine learning.</p>
      <sec id="sec-4-1">
        <title>1. Hierarchical clustering methods</title>
        <p>The method will create a hierarchical
decomposition of a given set of data objects.
Based on how the hierarchical decomposition is
formed, we can classify hierarchical methods.</p>
        <p>Agglomerative Approach is also known as
Button-up Approach [9]. Here we begin with
every object that constitutes a separate group. It
continues to fuse objects or groups close together.</p>
        <p>Divisive Approach is also known as the
TopDown Approach [9]. We begin with all the objects
in the same cluster. This method is rigid, i.e., it
can never be undone once a fusion or division is
completed.</p>
      </sec>
      <sec id="sec-4-2">
        <title>2. Partitioning based Methods</title>
        <p>Partition methods move the instances from one
group to another, starting from an initial partition.
The partition algorithm divides data into many
subsets. One of the most commonly used
algorithms is EM (Expectation – Maximization)
[10]. This algorithm tends to work with isolated
and compact groups. The basic idea is to find a
clustering structure that minimizes a certain error
criterion, which measures the "distance" of each
instance to its representative value.</p>
        <p>K-means algorithm is an iterative algorithm
that tries to partition the dataset into K pre-defined
distinct non-overlapping subgroups (clusters)
where each data point belongs to only one group.
It tries to make the inter-cluster data points as
similar as possible while also keeping the clusters
as different as possible. It assigns data points to a
cluster such that the sum of the squared distance
between the data points and the cluster’s centroid
is at the minimum [11]. The less variation we have
within clusters, the more homogeneous the data
points are within the same cluster.</p>
      </sec>
      <sec id="sec-4-3">
        <title>3. Density Based Methods</title>
        <p>The basic idea of "density-based" methods is
that for every instance of a group of zones near a
given radius must contain a minimum number of
instances. These methods identify the clusters and
the distribution of their parameters. The
algorithms produce clusters in a determined
location based on the high density of data set
participants. The DBSCAN (Density-Based
Spatial Clustering of Applications with Noise)
algorithm detects arbitrary groups and forms and
is efficient for large databases. This algorithm is
based on this intuitive notion of “clusters” and
“noise” [11]. The key idea is that for each point of
a cluster, the neighborhood of a given radius has
to contain at least a minimum number of points.</p>
      </sec>
      <sec id="sec-4-4">
        <title>4. Based Model Methods</title>
        <p>These methods use a hypothesized model
based on probability distribution. Model-based
clustering methods find characteristic
descriptions for each cluster, with each cluster
representing a concept or class. The COBWEB
algorithm yields a clustering dendrogram called
classification tree that characterizes each cluster
with a probabilistic description [12]. The
algorithm assumes that all attributes are
independent. It causes us to achieve a high
predictability of the values of the nominal
variables, given a set.</p>
        <p>This paper uses some data visualization
techniques such as charts and graphs.</p>
        <p>a. Charts</p>
        <p>What is the easiest method to display how one
or more data sets develop? It is a chart, of course.
Charts have a variety of forms, such as bar and
line charts, which may show relationships
between items over time. Pie charts can show how
the elements or portions relate together within a
whole.</p>
        <p>A line chart is created by connecting data
points within a data series using line segments.
Line charts are frequently employed for showing
trends in data that vary continuously over a period
of time or range [13].</p>
        <p>b. Graphs</p>
        <p>The use of graphs provides a general means to
transform the data and their relationships into an
abstract view for showing complex relationships
and improving data comprehension. Meanwhile,
graphs can also be adjusted flexibly to answer
specific questions based on the distinctive
characteristics of the data [14]. We demonstrate
the effectiveness of graph-based representations
by applying them to our data.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. Experimental results</title>
      <p>To conduct this study we used WEKA
software based on the approach and familiarity
with its use. The WEKA software package has
different programs for different techniques and
algorithms. WEKA is a collection of machine
learning algorithms and contains tools for data
pre-processing, classification, regression,
clustering, association rules, and visualization. It
is also well-suited for developing new machine
learning schemes. Table 2 presents a comparison
of the results of the algorithms applied to our data
in WEKA.</p>
      <p>In this paper we used some algorithms (Table
2) and among them is K-mean algorithm. This
algorithm provides clear results which are easy to
interpret. Model construction is done by
modifying the parameter values and this
algorithm groups the crime data in less time to
build the model. The K-mean algorithm was
applied to these data. The visualization of this
algorithm is shown in Figure 5.</p>
      <p>The number of clusters is two (0 and 1) and the
instances are grouped according to this scheme:
Cluster 0 has 75 instances or 83% cluster, while
cluster 1 has 15 instances or 17% cluster. The
number of iterations is 3.</p>
      <p>According to the results of the K-mean
algorithm, the persons who commit the most
crimes are jobless and with secondary education.
0.6
0.5
0.4
0.3
0.2
0.1
0
0.5</p>
      <p>The implementation of this algorithm has
clustered the data in the least amount of time to
build a model, exactly 0 seconds. This is shown in
Figure 6.</p>
      <p>Comparing this time with the time of other
algorithms for the same number of instances, this
algorithm has the shortest time, so it is faster.</p>
    </sec>
    <sec id="sec-6">
      <title>5. Conclusions</title>
      <p>This study presents a contribution in the field
of data visualization, with a focus on crime data.
This involves designing and creating a data set,
which will be used in data visualization and
machine learning using various visualization
techniques and machine learning algorithms. This
is about doing data visualization for those who
want to have knowledge of the data, interpret it
and take decisions from the data.</p>
      <p>The results of the experiments performed in
this study indicate that data visualization is
applicable in the field of criminology. The
Kmean algorithm aggregates the data in less time to
build a model, compared to other algorithms. This
algorithm shows promising results for the crime
prevention problem because the accuracy rate is
high in our experiments. The k-mean clustering
algorithm is easy to interpret and simple to
implement.</p>
      <p>Decision making is very important in crime
prevention in order to take accurate actions and
build law enforcement strategies. Through our
data analysis law enforcement agencies can create
strategies, operating in areas where most crimes
occur or for the perpetrators, their features (from
our study were those who were unemployed and
with secondary education). Data visualization
techniques and machine learning contribute to
predicting the likelihood of a crime occurring and
as a result to prevent it.</p>
    </sec>
    <sec id="sec-7">
      <title>6. References</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Maizura</surname>
            ,
            <given-names>Noor &amp; Ab</given-names>
          </string-name>
          <string-name>
            <surname>Hamid</surname>
            ,
            <given-names>Siti</given-names>
          </string-name>
          <string-name>
            <surname>Haslini</surname>
            &amp; Mohemad, Rosmayati &amp; Jalil, Masita &amp; Hitam,
            <given-names>Muhammad.</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>A Review on a Classification Framework for Supporting Decision Making in Crime Prevention</article-title>
          .
          <source>Journal of Artificial Intelligence. 8</source>
          .
          <fpage>17</fpage>
          -
          <lpage>34</lpage>
          .
          <fpage>10</fpage>
          .3923/jai.
          <year>2015</year>
          .
          <volume>17</volume>
          .34.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Mohd</surname>
            ,
            <given-names>Maseri</given-names>
          </string-name>
          &amp; Abdullah, Embong &amp; Mohamad
          <string-name>
            <surname>Zain</surname>
          </string-name>
          , Jasni. (
          <year>2010</year>
          ).
          <article-title>A Framework of Dashboard System for Higher Education Using Graph-Based Visualization Technique</article-title>
          .
          <volume>87</volume>
          .
          <fpage>55</fpage>
          -
          <lpage>69</lpage>
          .
          <fpage>10</fpage>
          .1007/978-3-
          <fpage>642</fpage>
          - 14292-
          <issue>5</issue>
          _
          <fpage>7</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Donohoe</surname>
            , David &amp; Costello,
            <given-names>Eamon.</given-names>
          </string-name>
          (
          <year>2020</year>
          ).
          <article-title>Data Visualisation Literacy in Higher Education: An Exploratory Study of Understanding of a Learning Dashboard Tool</article-title>
          .
          <source>International Journal of Emerging Technologies in Learning (iJET)</source>
          .
          <volume>15</volume>
          . 115. 10.3991/ijet.v15i17.
          <fpage>15041</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>(</given-names>
            <surname>Cho</surname>
          </string-name>
          , Wonhee &amp; Lim, Yoojin &amp; Lee,
          <string-name>
            <surname>Hwangro</surname>
          </string-name>
          &amp; Varma, Mohan &amp; Lee, Moonsoo &amp; Choi,
          <string-name>
            <surname>Eunmi.</surname>
          </string-name>
          (
          <year>2014</year>
          ).
          <article-title>Big Data Analysis</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>