<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Security Price Data Cleaning Technique: Reynold's Decomposition Approach</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Rachel V. Mok</string-name>
          <email>mokw@uah.edu</email>
          <email>rmok@mit.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Wai Yin Mok, Kit Yee Cheung</string-name>
          <email>mokw@uah.edu, kityeemok@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>College of Business Administration, University of Alabama in Huntsville</institution>
          ,
          <addr-line>Huntsville, Alabama, 35899</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Mechanical Engineering, Massachusetts Institute of Technology</institution>
          ,
          <addr-line>Cambridge, Massachusetts 02139</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <fpage>94</fpage>
      <lpage>98</lpage>
      <abstract>
        <p>We propose a security price data cleaning technique based on Reynold's decomposition that uses T0 (the time period of integration) to determine the de-noise level of the price data. The goal of this study is to find the optimal T0 that reveals an underlying price trend, possibly indicating the intrinsic value of the security. The DJIA (Dow Jones Industrial Average) Index and the thirty companies comprising the index are our fundamental interest. Preliminary results suggest that the graphs of ↵ (a key percentage measure) versus T0 of the thirty companies and the DJIA Index exhibit at least two properties: (1) ↵ drops exponentially as T0 increases when T0 / order of magnitude of 100 days, and (2) ↵ drops linearly as T0 increases when T0 ' order of magnitude of 100 days. For the DJIA Index itself, T0 is less than order of magnitude of 100 days. The result of applying our technique to each component stock of the DJIA parallels the result of the technique applied to the DJIA Index itself.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Understanding and analyzing financial data in
order to forecast and make cost-effective decisions
is challenging because of the complex and volatile
nature of security prices. The most recent
financial market meltdown in 2008-09 casted doubts
on financial data analysis and forecasting.
Inability to recognize or acknowledge financial
distress signaled by pertinent financial data was a
significant factor leading to these catastrophic
economic results
        <xref ref-type="bibr" rid="ref6">(Kaur, 2015)</xref>
        . Thus, veracity of
financial data takes priority in any data driven
decision making. Like any big data infrastructure,
veracity includes validation, noise level, deception,
detection, relevance and ranking of data collected
        <xref ref-type="bibr" rid="ref5">(Goes, 2014)</xref>
        . Depending on how collected
financial data are captured and processed in an analysis,
generated assessments can vary greatly from real
financial market performance. One has to look no
farther than the recent settlement of $77 million
between the SEC and Standard &amp; Poor credit
rating agency to see an example of how data analysis
can be misleading (http://www.sec.gov/
news/pressrelease/2015-10.html).
      </p>
      <p>
        Several financial computation models that
deal with cleaning financial data employ
similar methodologies, such as candlestick
strategies
        <xref ref-type="bibr" rid="ref4 ref5">(Detollenaere and Mazza, 2014)</xref>
        ,
multiplestage algorithm for detecting outliers in ultra
highfrequency financial market data
        <xref ref-type="bibr" rid="ref15">(Verousis and ap
Gwilym, 2010)</xref>
        , financial data filtering (http://
www.olsendata.com) and data-cleaning
algorithm
        <xref ref-type="bibr" rid="ref1 ref1 ref2 ref2">(Chung et al., 2004a; Chung et al., 2004b)</xref>
        .
Most data cleaning methodologies involve the
detection, distribution and/or the removal of outliers
        <xref ref-type="bibr" rid="ref11 ref13">(Shamsipour et al., 2014; Sun et al., 2013)</xref>
        .
However removing outliers in the dataset may have
a statistical distortion effect on the dataset itself
        <xref ref-type="bibr" rid="ref3">(Dasu and Loh, 2012)</xref>
        .
      </p>
      <p>
        To this end, we propose a data cleaning
technique based on Reynold’s decomposition in order
to decompose the price data into a mean part and
a fluctuating part. Fluctuations in stock prices are
perpetual and irrational in time because the weak
form of market efficiency and different types of
market participants create a complex dynamic of
behavioral finance
        <xref ref-type="bibr" rid="ref14">(Verheyden et al., 2015)</xref>
        .
Nevertheless, our approach could minimize part of the
effect of irrational price fluctuations by
incorporating and averaging fluctuation points (i.e.,
outliers) within a moving time period of integration,
T0. In essence, the length of T0 in the analysis
determines the level of veracity, with the larger
the T0, the lesser the influence of the fluctuation
points will be. We believe our data cleaning
technique is particularly applicable to security prices
due to the intense nature of security price changes
in relatively short periods, and it allows the user
to gauge different moving time periods of
integration to produce a unique set of statistical data for
targeted analysis.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Reynold’s Decomposition</title>
      <p>
        In the study of turbulence in fluid dynamics, each
component of the velocity is characterized by
fluctuations over time. One method to study the
dynamics in this regime is to perform a Reynold’s
decomposition such that the mean part of the
velocity is separated from the fluctuations. We
propose that this technique could also be used to study
financial data. In other words, we propose that the
price as a function of time, p(t), can be
decomposed into the following:
p(t) = p¯(t) + p0(t)
(1)
where p¯(t) is the mean portion and p0(t) is the
fluctuating portion of the price. We define p¯(t) to be
a moving time-average that can be found by
performing the following integral
p¯(t) =
1 Z t+T0/2
T0 t T0/2
p(t0)dt0
(2)
where T0 is the time period of integration. T0 must
be a time period that is greater than the time
period of the fluctuations, ⌧ , and less than the time
period of interest, T . T is dependent on each
particular analysis; for example, T could be weeks,
months, or years. Thus, ⌧ &lt; T 0 &lt; T .
Furthermore, the time-averaged value of the fluctuating
portion over the entire time period of interest is
zero
        <xref ref-type="bibr" rid="ref10 ref8">(M u¨ller, 2006; Mills, 1999)</xref>
        . As the time
period of integration increases, p¯(t) is farther away
from the actual p(t) and the magnitude of p0(t)
increases. Thus, the goal of this research is to find
the optimal time period of integration, T0, that
excludes the miscellaneous fluctuations and captures
the essential trend of the price data.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Methods</title>
      <p>
        In this study, we focus on the thirty
companies comprising the Dow Jones Industrial
Average (DJIA) as of May 13, 2015, and the DJIA
Index because, being the second oldest financial
index, the DJIA is the benchmark that tracks
financial market performance as a whole. Thus,
it represents a broad market, and its validity is
intensely scrutinized and followed by at least 10
Wall Street analysts
        <xref ref-type="bibr" rid="ref12 ref7 ref8 ref9">(Lee and Swaminathan, 1999;
Moroney, 2012; Stillman, 1986)</xref>
        . The ticker
symbols for the thirty companies that were studied in
this analysis are as follows: GS, IBM, MMM, BA,
AAPL, UTX, UNH, HD, DIS, CVX, NKE, TRV,
JNJ, MCD, CAT, XOM, PG, AXP, WMT, DD, V,
JPM, MRK, VZ, MSFT, KO, PFE, INTC, CSCO,
and GE. Because different companies can
comprise the DJIA Index at any point in time, we only
focus on the index as a whole when performing the
analysis for the DJIA Index itself.
      </p>
      <p>Daily adjusted close stock price data for the
thirty Dow Jones companies listed above from the
time of inception of the company to May 13, 2015,
are obtained from Yahoo! Finance. For the DJIA
Index, the daily adjusted close stock price data
from Jan. 29, 1985, to May 13, 2015, are also
obtained from Yahoo! Finance. The adjusted close
stock price is used because it accounts for stock
dividends and splits. Only days in which the stock
price is provided, i.e., business days, are
considered in this study. Thus, the time from Friday to
Monday is taken as only one (business) day.</p>
      <p>We estimate the time period of fluctuations to
be a day, ⌧ ⇠ 1 business day, and the time
period of interest to be the total number of business
days since the inception of the stock, T ⇠ 260 ⇥
n business days, where n represents the number of
years since the inception of the stock. Further, we
chose the following time periods of integration,
T0, for this study: 4 days, 10 days, 20 days, 30
days, 64 days, 130 days, 194 days, 260 days, 390
days, and 520 days, which roughly represent the
following time periods: one week, two weeks, one
month, one-and-a-half months, a quarter of a year,
half of a year, three-quarters of a year, one year,
one-and-a-half years, and two years, respectively.</p>
      <p>p¯(t) is calculated by only considering the
analysis time period from T0/2 after the day of
inception to T0/2 before May 13, 2015, such that for
each day p¯(t) is calculated, the full time period
of integration is used. To exemplify, consider the
case where T = 1000 days and T0 = 100 days.
Then the first 50 days (day 1 to day 50) are not
included in the analysis, and neither are the last 50
days (day 951 to day 1000). For each day in the
analysis time period, the integration stated in Eq.
(2) is performed numerically to find p¯(t) for that
day. p0(t) is found by subtracting p¯(t) from p(t),
the actual price, for that day.</p>
      <p>For each specific T0, the statistics of p0(t) are
analyzed. Specifically, a histogram with 25 bins
of p0(t) is created for each T0. As an example,
Fig 1 shows a histogram for GS (The Goldman
Sachs Group Inc). Note that like Fig 1, most of
the histograms are centered around 0, which
suggests that most of the fluctuations for the stocks
are nearly zero. Therefore, the actual stock price
is near or nearly equal to the local time-average for
most of the time period analyzed. For most stocks,
as T0 increases, the maximum height achieved by
the histogram decreases and the histogram tails
become heavier. Thus, as T0 increases, there are
more observations away from the center of the
distribution. This is observed because as the time
period of integration increases, more points are
considered in the average. Therefore, there is a
greater likelihood that p¯(t) is different from the
actual price.</p>
      <p>To measure the fidelity of p¯(t) to p(t), the
number of data points of p0(t) that are within 1
dollar from zero are counted and divided by the
total number of data points in the analysis period.
We will call this percentage measure ↵ , and this
measure should be as close as possible to 100% to
reflect that p¯(t) is a good approximation of p(t).
As stated previously, if p0(t) is near zero, that
means that p¯(t) is close to p(t) because p(t) =
p¯(t) + p0(t). As T0 increases, ↵ decreases
because the mean is farther away from the actual
price when the integration period is larger.
Using the MATLABr curve fitting tool, it is found
that for all of the thirty stocks the relationship
between ↵ and T0 is best represented by the
following equation
↵ (T0) = a1 eb1 T0 + c1 ed1 T0
(3)
where a1, b1, c1, and d1 are curve fitting
parameters. In fact, the lowest goodness-of-fit measure
R2 among all thirty stocks is 0.9909. As an
example, the curve fitting parameters for GS are a1 =
0.3613, b1 = 0.09018, c1 = 0.1034, and d1 =
0.003687. The first derivative of this equation is
d↵
dT0</p>
      <p>= a1 b1 eb1 T0 + c1 d1 ed1 T0
and the second derivative is
d2↵
dT02 = a1 (b1)2 eb1 T0 + c1 (d1)2 ed1 T0
For most of the stocks, it was discovered that
when T0 is fewer than 100 days, the measure ↵
drops exponentially as T0 increases. However, the
second derivative (Eq. (5)) becomes near zero in
a range from 96 days to 387 days for the thirty
stocks analyzed, with the most common being
approximately 125 days. Thus, when T0 is at least
(4)
(5)
an order of magnitude of 100 days, ↵ starts to
decrease linearly for nearly all of the stocks
analyzed. Fig 2 plots the curve fitted ↵ (T0) for all
thirty analyzed stocks. As we can see, the general
trend among the thirty stocks is that ↵ drops
exponentially when T0 is fewer than 100 days, but ↵
drops linearly when T0 is greater than 100 days.
An appealing fact is that the graph of ↵ against
T0 for the DJIA Index, Fig 3, also exhibits similar
trends in ↵ , as shown in Fig 2. Note the
different scales of the vertical axes of Fig 2 and Fig 3,
which means that Fig 3 is much flatter than Fig 2.</p>
      <p>Mathematically, we will define the point where
the slope is constant by the following</p>
      <p>d2↵
T0! tc dT02 = 0
lim
(6)
where tc is the time period of integration at which
the second derivative of ↵ approaches zero. Thus,
for the thirty stocks analyzed, tc is in the
following range 96 days &lt; tc &lt; 387 days. Therefore,
for time periods of integration larger than tc, the
change in ↵ will be relatively small.</p>
      <p>The standard deviation of the fluctuations
p0(t), defined as
v
u
= tu N
1
1</p>
      <p>N
X(p0(ti)
i=1
p¯0(t))2,
(7)
is also analyzed where N is the total number of
data points and p¯0(t) is the total time average of the
fluctuations. A large of the fluctuations reflects
that p¯(t) is not equal to p(t). To remove the
miscellaneous fluctuations of the price, should be as
large as possible. As indicated by the straight lines
in the log-log plot in Fig 4, and T0 are related by
a power law where the slope of the line indicates
the exponent.</p>
      <p>Using the MATLABr curve fitting tool, we
fit the following equation for each of the thirty
stocks:
= a2 T0b2 + c2
(8)
where a2, b2, and c2 are curve fitting
parameters. As an example, a2 = 2.194, b2 =
0.4317, and c2 = 1.425 for GS. The
lowest goodness-of-fit measure R2 among the thirty
stocks is 0.9849. b2, the exponent, varies from
0.35 to 0.69 for all thirty stocks. The average
exponent is 0.5.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Results and Conclusions</title>
      <p>This paper demonstrates preliminary results of
an ongoing security price data cleaning
research. We found that the graphs of ↵
versus T0 of the thirty companies and the DJIA
Index exhibit at least two properties: (1)
↵ drops exponentially as T0 increases when
T0 / order of magnitude of 100 days, and (2)
↵ drops linearly as T0 increases when T0 '
order of magnitude of 100 days. Thus, the
optimal T0 for the thirty companies studied is
approximately 100 days. For the DJIA Index itself, the
optimal T0 appears to be less than 100 days. One
of the possible explanations is that the DJIA Index
might show the counter measure effect of
fluctuation points among the thirty companies since the
DJIA is a composite of the thirty companies that
collectively provide a balance view of the market.
As a result, T0 might be even smaller for the
second derivative to approach zero. We also found
that and T0 are related by a power law. As for
future research, we plan to define mathematical
metrics in our study of security price valuations
and trading strategies.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Kee H. Chung</surname>
            , Chairat Chuwonganant, and
            <given-names>D. Timothy</given-names>
          </string-name>
          <string-name>
            <surname>McCormick</surname>
          </string-name>
          .
          <year>2004a</year>
          .
          <article-title>Order preferencing and market quality on NASDAQ before and after decimalization</article-title>
          .
          <source>Journal of Financial Economics</source>
          ,
          <volume>71</volume>
          (
          <issue>3</issue>
          ):
          <fpage>581</fpage>
          -
          <lpage>612</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Kee H. Chung</surname>
          </string-name>
          ,
          <string-name>
            <surname>Bonnie F. Van Ness</surname>
            ,
            <given-names>and Robert A. Van Ness.</given-names>
          </string-name>
          2004b.
          <article-title>Trading costs and quote clustering on the NYSE and NASDAQ after decimalization</article-title>
          .
          <source>Journal of Financial Research</source>
          ,
          <volume>27</volume>
          (
          <issue>3</issue>
          ):
          <fpage>309</fpage>
          -
          <lpage>328</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Tamraparni</given-names>
            <surname>Dasu</surname>
          </string-name>
          and Ji Meng Loh.
          <year>2012</year>
          .
          <article-title>Statistical distortion: Consequences of data cleaning</article-title>
          .
          <source>Proceedings of the VLDB Endowment</source>
          ,
          <volume>5</volume>
          (
          <issue>11</issue>
          ):
          <fpage>1674</fpage>
          -
          <lpage>1683</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Benoit</given-names>
            <surname>Detollenaere</surname>
          </string-name>
          and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Mazza</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Do Japanese candlesticks help solve the trader's dilemma?</article-title>
          <source>Journal of Banking and Finance</source>
          ,
          <volume>48</volume>
          :
          <fpage>386</fpage>
          -
          <lpage>395</lpage>
          , November.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Paulo</surname>
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Goes</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Big data and IS research</article-title>
          .
          <source>MIS Quarterly</source>
          ,
          <volume>38</volume>
          (
          <issue>3</issue>
          )
          <string-name>
            <surname>:</surname>
          </string-name>
          iii-viii,
          <year>September</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Inderjit</given-names>
            <surname>Kaur</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Early warning system of currency crisis : Insights from global financial crisis 2008</article-title>
          .
          <source>IUP Journal of Applied Economics</source>
          ,
          <volume>14</volume>
          (
          <issue>1</issue>
          ):
          <fpage>69</fpage>
          -
          <lpage>83</lpage>
          , January.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Charles M. C. Lee</surname>
            and
            <given-names>Bhaskaran</given-names>
          </string-name>
          <string-name>
            <surname>Swaminathan</surname>
          </string-name>
          .
          <year>1999</year>
          .
          <article-title>Valuing the Dow: A bottom-up approach</article-title>
          .
          <source>Financial Analysts Journal</source>
          ,
          <volume>55</volume>
          (
          <issue>5</issue>
          ):
          <fpage>4</fpage>
          -
          <lpage>23</lpage>
          , September.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Anthony F.</given-names>
            <surname>Mills</surname>
          </string-name>
          .
          <year>1999</year>
          .
          <article-title>Basic Heat and Mass Transfer</article-title>
          . Prentice Hall, second edition.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Richard</given-names>
            <surname>Moroney</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>What we're thinking add it up: Dow has further upside</article-title>
          .
          <source>Dow Theory Forecasts</source>
          ,
          <volume>68</volume>
          (
          <issue>9</issue>
          ):
          <fpage>2</fpage>
          -
          <lpage>3</lpage>
          , February.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Peter</given-names>
            <surname>Mu</surname>
          </string-name>
          ¨ller.
          <year>2006</year>
          .
          <article-title>The Equations of Oceanic Motions</article-title>
          . Cambridge University Press.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Mansour</given-names>
            <surname>Shamsipour</surname>
          </string-name>
          , Farshad Farzadfar, Kimiya Gohari, Mahboubeh Parsaeian, Hassan Amini, Katayoun Rabiei, Mohammad Sadegh Hassanvand, Iman Navidi, Akbar Fotouhi, Kazem Naddafi, Nizal Sarrafzadegan, Anita Mansouri, Alireza Mesdaghinia, Bagher Larijani, and
          <string-name>
            <given-names>Masud</given-names>
            <surname>Yunesian</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>A framework for exploration and cleaning of environmental data - Tehran air quality data experience</article-title>
          .
          <source>Archives of Iranian Medicine</source>
          ,
          <volume>17</volume>
          (
          <issue>12</issue>
          ):
          <fpage>821</fpage>
          -
          <lpage>829</lpage>
          , December.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Richard Joseph</given-names>
            <surname>Stillman</surname>
          </string-name>
          .
          <year>1986</year>
          .
          <article-title>Dow Jones Industrial Average : history and role in an investment strategy</article-title>
          .
          <source>Irwin Professional Pub.</source>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>W.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Whelan</surname>
          </string-name>
          , AB.
          <string-name>
            <surname>McBratney</surname>
            , and
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Minasny</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>An integrated framework for software to provide yield data cleaning and estimation of an opportunity index for site-specific crop management</article-title>
          .
          <source>Precision Agriculture</source>
          ,
          <volume>14</volume>
          (
          <issue>4</issue>
          ):
          <fpage>376</fpage>
          -
          <lpage>391</lpage>
          ,
          <year>August</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Tim</given-names>
            <surname>Verheyden</surname>
          </string-name>
          , Lieven De Moor, and Filip Van den Bossche.
          <year>2015</year>
          .
          <article-title>Towards a new framework on efficient markets</article-title>
          .
          <source>Research in International Business and Finance</source>
          ,
          <volume>34</volume>
          :
          <fpage>294</fpage>
          -
          <lpage>308</lpage>
          , May.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Thanos</given-names>
            <surname>Verousis</surname>
          </string-name>
          and Owain ap Gwilym.
          <year>2010</year>
          .
          <article-title>An improved algorithm for cleaning ultra high-frequency data</article-title>
          .
          <source>Journal of Derivatives &amp; Hedge Funds</source>
          ,
          <volume>15</volume>
          (
          <issue>4</issue>
          ):
          <fpage>323</fpage>
          -
          <lpage>340</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>