<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Demand Forecasting for Inventory Management using Limited Data Sets: A Case Study from the Oil Industry</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jorge Ivan Romero-Gelvez</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Esteban Felipe Villamizar</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Olmer Garcia-Bedoya</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jorge Aurelio Herrera-Cuartas</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Universidad de Bogotá Jorge Tadeo Lozano</institution>
          ,
          <addr-line>Bogotá</addr-line>
          ,
          <country country="CO">Colombia</country>
        </aff>
      </contrib-group>
      <fpage>111</fpage>
      <lpage>119</lpage>
      <abstract>
        <p>This document's main focus is to present a way to solve forecasting issues using open source tools for time series analysis. First, we present an introduction to the hydrocarbon sector and time series analysis, later we focus on the solution methods based on supervised learning trained (support vector regression) with bio-inspired algorithms (Particle swarm optimization). We expose some benefits of use support vector machines and open source tools that focuses on variables like trend and seasonality. In this work, we chose the fb-prophet package and support vector regressor with scikit-learn as the primary tools because they have representative results dealing with limited data sets and Particle swarm optimization as training algorithm because of their speed and adaptability. Finally, we show the results and compare them with their RMSE obtained.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Hydrocarbon</kwd>
        <kwd>Forecasting</kwd>
        <kwd>small time-series</kwd>
        <kwd>support vector regressor</kwd>
        <kwd>particle swarm optimization</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>million, and Arauca USD 101 million.</p>
      <p>Also, the ACP projects that in 2020 there will be an investment in exploration and production
of oil and gas of USD 4,970 million, 23% higher than in 2019. Indeed, the largest company in
the country and the leading oil company in Colombia belong to the group of the 39 largest oil
companies in the world and is one of the top five in Latin America. They have hydrocarbon
extraction fields in the center, south, east and north of Colombia, two refineries, ports for
the export and import of fuel and crude oil on both coasts, and a transportation network of
8,500 kilometers of pipelines and pipelines to throughout the entire national geography, which
interconnect production systems with large consumption centers and maritime terminals.</p>
      <p>Given this panorama, the industries must plan strategies to manage and control their
inventories since their importance lies in obtaining profits. Inventory management plays a vital role
within the business chain, which is the bufer between two processes, supply, and demand. This
can be known or unknown, variable, or constant. The sourcing process contributes goods to
inventory while demand consumes the same inventory. This is necessary due to diferences in
rates and times between supply and demand, and this diference must be attributed to internal
or external factors. Endogenous factors are policy issues, but exogenous factors are
uncontrollable. Internal factors include economies of scale, smoothing of operation, and customer
service, the most important exogenous factor being uncertainty.</p>
      <p>Given that inventory control is a critical aspect for efective management and administration,
in order to guarantee the availability of equipment, spare parts, and materials to meet the
needs of expenses and projects with the expected quality, cost, and opportunity. Likewise,
the materials management process is to ensure that the materials that register stock in the
warehouse correspond to optimal levels of inventories, such quantity must fully meet the needs
of the company with the minimum investment.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Literature review</title>
      <p>
        There are applications in forecasts using SVR and it is more common to find them in recent
years, since there is a special interest in machine learning applications to make predictions on
time series. An example of this can be seen in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] with financial forecasts, [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ] over rainfall
predictions, electric load forecasting [
        <xref ref-type="bibr" rid="ref4 ref5 ref6">4, 5, 6</xref>
        ], forecasting carbon price [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] among many
others. Next, we present a brief introduction to the time series and later we will deal with the
comparison of the applied models.
      </p>
      <sec id="sec-2-1">
        <title>2.1. Time series analysis</title>
        <p>
          The analysis of time series and demand forecasts becomes the primary input for the MRP model.
For this reason, it is proposed to contrast diferent methods that allow considering seasonal
periods, and that may also include new demand observations to adjust the model in real-time.
According to [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] Time-series methods refers to a set of observations of real phenomena (like
mathematical, biological, social sciences, physical, economic among others) given as part of
a discrete set in time. The main thought is that past data can be utilized to generate future
estimations. In time series analysis is common to try to get patters of data as trend, seasonality,
cycles, and randomness as inputs to modeling the phenomena.
• Trend: The data set exhibits a stable pattern of growth or decrease.
• Seasonality: A seasonality pattern are those that are repeated at fixed intervals.
• Cycles: The variation of cycles is similar to seasonality, except that the duration and the
magnitude of the cycle varies.
• Randomness: A random series is where you do not have a recognized pattern of data.
        </p>
        <p>One can generate random series of data that have a specific structure. The data that
seems to have apparently a randomness , actually have a specific structure. Actually the
really random data fluctuate around a fixed average.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Support Vector machine and support vector regressor</title>
        <p>According with [9] In machine learning, support vector machine proposed by Vapnik is one
of the most popular approaches for supervised learning [10, 11]. This model resembles logistic
regression in that a linear function  ⊤</p>
        <p>+  drives both. Its main diference is that the logistic
regression produces probabilities, the support vector machine produces a class identity. The
SVM predicts that the positive class is present when  ⊤
 +  is positive. Likewise, it predicts
that the negative class is present when  ⊤</p>
        <p>+  negative. A notable feature of support vector
machines is the kernel, considering that many algorithms can be written in the form of a dot
product. As an example, the rewritten SVM linear function is shown as:
 ⊤ +  =  + ∑    ⊤ ( )

 =1
where  ( ) is a training example and  is a vector of coeficients. Rewriting the learning
algorithm this way allows us to replace  by the output of a given feature function  ( ) and
the dot product with a function  ( ,  ( )
represents an inner product analogous to  ( )⊤ (
) =  ( ) ⋅  (</p>
        <p>( )) called a kernel. The (·) operator
 ( )). For some feature spaces, we may not
use literally the vector inner product. In some infinite dimensional spaces, we need to use
other kinds of inner products, for example, inner products based on integration rather than
summation. After replacing dot products with kernel evaluations, we can make predictions
using the function
 ( ) =  + ∑    ( ,  ( )</p>
        <p>)

is linear. Also, the relationship between</p>
        <p>This function is nonlinear with respect to</p>
        <p>, but the relationship between  ( ) and  ( )
and  ( ) is linear. The kernel-based function is
exactly equivalent to preprocessing the data by applying  ( ) to all inputs, then learning a
linear model in the new transformed space. The kernel trick is powerful for two reasons. First,
it allows us to learn models that are nonlinear as a function of  using convex optimization
techniques that are guaranteed to converge eficiently. This is possible because we consider 
ifxed and optimize only  , i.e., the optimization algorithm can view the decision function as
being linear in a diferent space. Second, the kernel function  often admits an
implementation that is significantly more computational eficient than naively constructing two vectors
dot product.</p>
        <p>) = min (,</p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Support vector regressor</title>
        <p>and explicitly taking  ( ) their dot product. In some cases,  ( ) can even be infinite
dimensional, which would result in an infinite computational cost for the naive, explicit approach.
In many cases,  ( ,  ′) is a nonlinear, tractable function of x even when  ( ) is intractable.
As an example of an infinite-dimensional feature space with a tractable kernel, we construct
a feature mapping  ( ) over the non-negative integers  . Suppose that this mapping returns
a vector containing  ones followed by infinitely many zeros. We can write a kernel function
( )) that is exactly equivalent to the corresponding infinite-dimensional
According to [12] the basic idea of SVR is to map the data  into a high dimensional feature
space  by nolinear mapping  and to do linear regression in this space. Also, according to
[13] we can consider a set of training data  ( ) = (
⋅  ( )) +  , where each   ⊂  n denotes
the input space of the sample and has a corresponding target value   ⊂  for  = 1, … ,  , where
corresponds to the size of the training data. The idea of the regression problem is to determine
a function that can approximate future values accurately. The generic form of SVR can be seen
as follows
 ( ) = ( ⋅  ( )) + 
(3)
(4)
(5)
(6)
(7)
where,  ⊂ 
n
,  ⊂</p>
        <p>and  denotes a nonlinear transfor to high-dimensional space. Our goal
is to find the value of and such that values of can be determined by minimizing the regression
risk
points as
where  (⋅) is a cost function,  is a constant, and vector  can be written in terms of data
By substituting eq.5 into eq.3, the generic equation can be rewritten as
 reg( ) = ∑  ( (  ) −   ) +  ‖ ‖2
 = ∑ (  −   )  (  )</p>
        <p>∗

 =1

 =1
 (  ,  ) = exp − | −   |2</p>
        <p>}
{
 =1

 =1
 ( ) = ∑ (  −  ∗) ( (  ) ⋅  ( )) + 
∗
= ∑ (  −   )  (  ,  ) + 
In eq.6, the dot product is replaced with a kernel function  (  ,  ) = (
 (  ) ⋅ 
functions enable the dot product to be performed in high-dimensional feature space using
lowdimensional space data input without knowing the transformation  . All kernel functions must
satisfy Mercer’s condition that corresponds to the inner product of some feature space. The
(  )). Kernel
radial basis function RBF is commonly used as the kernel for regression</p>
      </sec>
      <sec id="sec-2-4">
        <title>2.4. Particle swarm optimization</title>
        <p>Particle swarm optimization is a computational technique that optimizes a problem by
iteratively attempting to promote a candidate solution concerning a given degree of quality. It
solves a problem by producing a population of candidate solutions called particles and moving
them around the search-space with particular position and velocity. Each particle’s movement
is influenced by its local best-known place but is also guided toward the best-known positions
in the search-space, which are renewed as another particle found better solutions. This is
expected to move the swarm toward the best solutions. According to [14] the main inputs for the
formulation of this algorithm can be seen as follows:</p>
        <p>intervals.
•  is the dimension of the search space.
•  is the search space, a hyperparallelepid defined as the euclidean product of  real

min, m ax]</p>
        <p>The standard form of the algorithm composes by a set of particles (called swarm), made of
a position in the search space, the fitness value at each position, a velocity for displacement,
a memory that contains the best position and last the fitness value of the previous best. The
search is performed in two phases, initialization of the swarm and a cycle of iterations. The
main steps can be seen as follows:
• Initialisation of the swarm: pick a random position in search space and pick a random
velocity.
• Stop .</p>
        <p>• Iteration: compute the new velocity, move and compute the new fitness.</p>
      </sec>
      <sec id="sec-2-5">
        <title>2.5. The prophet forecasting model</title>
        <p>Prophet is an open source tool for forecasting time series observations based on an additive
model in which nonlinear trends adjust to seasonality. Their results improve with data that
includes compromises with strong seasonal efects and a considerable amount of historical
data. We use Prophet open-source software in Python [15] based in a decomposable time series
model [16] components: trend, seasonality, and vacations. They are combined as follows:
 ( ) =  ( ) +  ( ) + ℎ( ) +</p>
        <p>Were  ( ) is the trend function which models non periodic changes in the value of the time
series,  ( ) represents periodic changes, and ℎ( ) represents the efects of vacations which occur
on potentially irregular schedules over one or more days. The error term  
idiosyncratic changes which are not accommodated by the model; We will make a parametric
represents any
assumption with normally distributed  .
(8)
(9)
2.5.1. Prediction model
The development of an algorithm is necessary to obtain a data history in which the breakdown
of the monthly inventory value is represented by the BSE material and transit corresponding
to two years. In this analysis of generated data, factors such as inventory value and material
dependence are taken into account. These have a significant influence on the model’s behavior
since they provide realism and specific variations of the trend line, which are of interest to
optimize its management. All these data have been obtained from the materials management
system. Subsequently, these data will be processed and analyzed in order to see the interaction
between them, carrying out a parameterization that allows characterizing the logic of
generation of the historical.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Data description</title>
      <p>The warehouse corresponds to the code assigned in the information system that represents an
organizational unit or warehouse, corresponds to the physical place where the materials are
stored, which allows diferentiation of material stocks. In this case, the logistics center 2000
is taken, as a Detail describes the types of warehouse: Imported: corresponds only to
material in transit of expenses or projects. Expenses: the types of warehouses are associated with
new material in good condition and repaired material in good condition. These materials are
characteristic of the operation and maintenance, such as spare parts, consumables, and
supplies for the operation, equipment, lubricants, consumption tools, parts from manufacturers,
spare parts. Projects: this material is part of the business investment with a physical location
in patios and covered warehouse, this material is acquired according to the project’s
requirement, given that the material is no longer required by the project, it is assigned as not required
material or in the process of sale which is ofered to other projects of the business group.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Solution method and results</title>
      <p>In order to solve the problem, we contrast two methods, the fb-prophet and PSO-SVM.
• Forecasting Model selection: In order to use the method that generates less error
  . First, we apply support vector machine with particle swarm optimizations as global
optimization algorithm. In addition, Fb-prophet (black box method) is also used. Next,
we select the method with the least error of them
• IDE: IPython/Jupyter notebooks and Google-Colab.</p>
      <sec id="sec-4-1">
        <title>4.1. Results</title>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions</title>
      <p>The main forecast topic with a variety of backgrounds should make more forecasts than they
can do manually—the first component of our forecast. The system is the new model that we
have developed in many prediction iterations of a variety of data in FbProphet. We use a
simple modular regression model that often works well with predetermined parameters, and
that allows you to select the components that are relevant to your forecast problem and quickly
make adjustments as needed. The success of model lies in its ability to adjust the positions of
all particles in an area of search space with satisfactory solutions.</p>
      <p>According to a determined objective function to minimize, in this case, the root mean square
error. It measures the dispersion of error in forecast, this value is the diference between real
demand and the forecast squaring, disabling those periods where the diference was higher
compared to others. From this calculation, decisions against forecast models and their results
are guided to the best choice. Likewise, it is established that this type of problems are solved by
evolutionary algorithms, the importance of using these algorithms as the swarm of particles
lies in the high eficiency in generating predictions with better performance, the formulation
is reduced to characterize the movement of the particles based on a speed operator that must
associate exploration and convergence by decomposing the speed into three components in
order to decipher the study behavior.
[9] I. Goodfellow, Y. Bengio, A. Courville, Deep Learning, MIT Press, 2016. http://www.</p>
      <p>deeplearningbook.org.
[10] B. E. Boser, I. M. Guyon, V. N. Vapnik, A training algorithm for optimal margin classifiers,
in: Proceedings of the fifth annual workshop on Computational learning theory, 1992, pp.
144–152.
[11] C. Cortes, V. Vapnik, Support-vector networks, Machine learning 20 (1995) 273–297.
[12] K.-R. Müller, A. J. Smola, G. Rätsch, B. Schölkopf, J. Kohlmorgen, V. Vapnik,
Predicting time series with support vector machines, in: International Conference on Artificial
Neural Networks, Springer, 1997, pp. 999–1004.
[13] C.-H. Wu, J.-M. Ho, D.-T. Lee, Travel-time prediction with support vector regression,</p>
      <p>IEEE transactions on intelligent transportation systems 5 (2004) 276–281.
[14] M. Clerc, Beyond standard particle swarm optimisation, in: Innovations and
Developments of Swarm Intelligence Applications, IGI Global, 2012, pp. 1–19.
[15] S. J. Taylor, B. Letham, Forecasting at scale, The American Statistician 72 (2018) 37–45.
[16] A. C. Harvey, S. Peters, Estimation procedures for structural time series models, Journal
of Forecasting 9 (1990) 89–108.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>C.-J. Lu</surname>
            ,
            <given-names>T.-S.</given-names>
          </string-name>
          <string-name>
            <surname>Lee</surname>
          </string-name>
          , C.-C. Chiu,
          <article-title>Financial time series forecasting using independent component analysis and support vector regression, Decision support systems 47 (</article-title>
          <year>2009</year>
          )
          <fpage>115</fpage>
          -
          <lpage>125</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A. D.</given-names>
            <surname>Mehr</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Nourani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. K.</given-names>
            <surname>Khosrowshahi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Ghorbani</surname>
          </string-name>
          ,
          <article-title>A hybrid support vector regression-firefly model for monthly rainfall forecasting</article-title>
          ,
          <source>International Journal of Environmental Science and Technology</source>
          <volume>16</volume>
          (
          <year>2019</year>
          )
          <fpage>335</fpage>
          -
          <lpage>346</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>C.</given-names>
            <surname>Balsa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. V.</given-names>
            <surname>Rodrigues</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Lopes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Rufino</surname>
          </string-name>
          ,
          <article-title>Using analog ensembles with alternative metrics for hindcasting with multistations</article-title>
          ,
          <source>ParadigmPlus</source>
          <volume>1</volume>
          (
          <year>2020</year>
          )
          <fpage>1</fpage>
          -
          <lpage>17</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , W.-C. Hong,
          <article-title>Electric load forecasting by complete ensemble empirical mode decomposition adaptive noise and support vector regression with quantum-based dragonfly algorithm</article-title>
          ,
          <source>Nonlinear Dynamics</source>
          <volume>98</volume>
          (
          <year>2019</year>
          )
          <fpage>1107</fpage>
          -
          <lpage>1136</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Che</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Sequential grid approach based support vector regression for short-term electric load forecasting</article-title>
          ,
          <source>Applied Energy</source>
          <volume>238</volume>
          (
          <year>2019</year>
          )
          <fpage>1010</fpage>
          -
          <lpage>1021</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , W.-C. Hong,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Electric load forecasting by hybrid self-recurrent support vector regression model with variational mode decomposition and improved cuckoo search algorithm</article-title>
          ,
          <source>IEEE Access 8</source>
          (
          <year>2020</year>
          )
          <fpage>14642</fpage>
          -
          <lpage>14658</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>B.</given-names>
            <surname>Zhu</surname>
          </string-name>
          , D. Han,
          <string-name>
            <given-names>P.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , Y.
          <string-name>
            <surname>-M. Wei</surname>
          </string-name>
          ,
          <article-title>Forecasting carbon price using empirical mode decomposition and evolutionary least squares support vector regression</article-title>
          ,
          <source>Applied energy 191</source>
          (
          <year>2017</year>
          )
          <fpage>521</fpage>
          -
          <lpage>530</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>F. R.</given-names>
            <surname>Jacobs</surname>
          </string-name>
          ,
          <article-title>Manufacturing planning and control for supply chain management</article-title>
          ,
          <source>McGrawHill</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>