<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Estimating Speeds and Directions of Pedestrians in Real-Time Videos: A solution to Road- Safety Problem</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sultan Daud Khan</string-name>
          <email>sultan.khan@disco.unimib.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Complex Systems and Arti cial Intelligence Research Center Department of Informatics, Systems and Communication University of Milano-Bicocca</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Pedestrian injuries and fatalities are one of the most signi cant problems related to travel and road safety. Pedestrians are vulnerable users of roads and due to the very di erent velocities and mass when compared to vehicles like cars and trucks, and very often they undergo serious injuries in case of collisions. Older pedestrians are even more vulnerable to injuries and fatalities due to (i) their reduced mobility and re exes and (ii) their increased fragility when compared to young individuals. Crosswalks are the point where pedestrians face lower level of safety because they have to cross the street and must be aware of the incoming tra c. Such kind of awareness becomes di cult in case of old pedestrians because of their reduced physical and perceptive capabilities. Besides other factors, lower speed of an old pedestrian is an important factor that limits the mobility of old pedestrians and it also increases the risk of fatalities while crossing the road. In this paper, we developed vision based intelligent system that can detect low speeds and directions of pedestrians and can help him/her by (a) increasing the time associated to a green light for pedestrians, (b) using audible signals to help the pedestrians understanding that there are cars approaching the crossing.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Pedestrian injuries and fatalities are one of the most signi cant problem related
to travel and road safety. Pedestrians are the vulnerable users of road and due
to the very di erence in speed and mass when compared to vehicles like cars and
trucks, are very often they undergo serious accidents. Studies have shown that
more than one fth of the pedestrians killed while crossing the roads [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. About
273000 pedestrian were killed in road tra c crashes in 2010 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. 15% of the total
numbers of people killed were pedestrians in European Countries [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Walking is
a basic mode of transport in all societies around the world. Walking is good for
health and particular for the cardiovascular patients. Due to increase in number
of vehicles and population, people prefer to walk rather than taking a car
particularly in the cases when their destinations are near. Of course, these pedestrian
will use the roads but due to some risk factors like di erence speed, alcohol, lack
of road infrastructure for pedestrian and lack of other perceptive capabilities
will lead them to injuries and sometimes to fatalities. It is also stated that most
of these accidents took place in urban areas. Urban areas are most populated
places relative to rural areas. Urban areas are equipped with wide road and huge
tra c ow which makes it di cult for pedestrians to move. Pedestrian injuries
and fatalities also have psychological, socioeconomic and health costs. Although
there is no estimation for economic impact of pedestrian injuries but road tra c
crashes consumes 1% to 2% of gross national product [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Also the survivors of
tra c crashes, their families and friends often su er intense social, physical and
psychological e ects. Pedestrians form mixed group of people in terms of age,
gender and socioeconomic status. Studies have showed that pedestrian crashes
are related to risk factors and road geometrical factors [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. These factors include:
1) age and gender 2) pedestrian crossing time 3) pedestrian crossing speed 4)
crossing the street with red or green tra c light. Pedestrian crashes a ect the
people from di erent age group. Studies have shown that in United States in
2009, the fatality rate for pedestrians older than 75 years, higher than the
fatality rate of any other age group [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Literature suggests that pedestrians of age
older than 65 years have high accident risk than any other age group. There are
many reasons involved in the high fatality rate for pedestrian older than 65 years.
These include: 1) de cits in their physical abilities 2 sensory and perceptual
abilities 3) cognitive abilities. The aged population of most of developed countries
like japan and European Countries like Italy etc. are increasing. The rate of old
population is expected to increase by a 20% by year 2031 [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Older pedestrians
face major problems and accidents may occur as a result of age-related decline
skills used while crossing the road. These include 1) motion perception 2)
memory capacity 3) reaction time and physical mobility such as the ability to rotate
neck, walking and muscle control, balance and postural control [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] . There are
two components of motion; distance travelled and speed and it is believed that
age di erences in motion perceptions are the cause of certain accidents.
Crosswalks are the point where pedestrians face lower level of safety because they
have to cross the street and must be aware of the incoming tra c. Many studies
have examined the behavior of pedestrians crossing the road by analyzing several
factors. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] Analyzed the behavior of 1392 pedestrian in signalized crosswalks.
They made hypothesis that pedestrians are more optimistic of crossing the road
with red tra c light if another pedestrian crossed before him. Moreover, men are
more optimistic to cross the road with red tra c light than woman most unsafe
choices were taken by old pedestrians in [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].Old pedestrians face many
problems as other participants in the community, particularly in transport domain.
Older pedestrians often report inability to complete crossings in the time given
by pedestrian light. Keeping in view the above discussion, our work is motivated
by two factors. One is that the proposed problem has a great social meaning.
Secondly, our proposed solution, which makes it di erent from existing approaches,
is focused on detecting the old pedestrians crossing on the basis of the low speeds
and help him/her by (a) increasing the time associated to a green light for
pedestrians, (b) using audible signals to help the pedestrians understanding that there
are cars approaching the crossing. Previous research on tra c signal control was
mainly focused on vehicle monitoring, and very little literature can be found on
pedestrians side. An approach to detect and count pedestrians at an intersection
using xed camera is proposed in [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Background subtraction is employed
for motion segmentation and median ltering and erosion/dilation operations
are performed to reduce noise. Connected components are extracted and
information about the size and coordinates of each connected component is used to
compute the number of people in the scene. Computer vision based multi-agent
approach is presented in [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], where each agent makes decisions according to
local variables and information received from other agents. The problem of object
tracking in an uncontrolled urban environment is discussed in [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. A speci c
motion detection algorithm was used to detect objects such as pedestrians,
vehicles, etc. The motion detection algorithm proposed is based on construction
of a reference edge image of the background, composed of all stationary edges
in the scene. A single camera looking at an intersection point is used in [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
The authors focused on motion tracking. Motion segmentation is performed
using an adaptive background model that can gain robustness with respect to the
changes in illumination while tracking of objects is performed by computing the
overlap between bounding boxes. In the above works frame di erence method
are used which can not completely extract all the information regarding
foreground area, the central part of the target will be lost which ultimately result in
bad tracking. Alternative to other approaches a vision-based intelligent
pedestrian crossing system is developed in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] using stereo vision approaches which
detects the pedestrians and help those who need longer to cross the road. In
this paper, we propose a vision-based approach that can detect the pedestrians
e ciently and automatically. The proposed system is not only applicable to road
safety problem but can also be applicable to security systems inside building.
For robust foreground segmentation we use Lucas-Kanade optical ow [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] and
Gaussian Mixture Model [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. The perfect background can not be obtained by
optical ow and GMM methods individually. After foreground segmentation, we
apply Lucas- Kanade Tracker that track the points of pedestrian from frame to
frame and calculate the instantaneous velocity and average speed of pedestrian
is determined by calculating the scale factor. Also, in this paper, we estimate
the direction of the pedestrians. The rest of the paper is organized as follows. In
Section 2, we shall describe motion detection techniques. In Section 3, we discuss
tracking. In Section 4, we shall discuss our proposed framework and in Section
5, we shall discuss experimental results.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Motion Segmentation</title>
      <p>Motion segmentation is the most important pre-processing step for detecting the
moving objects from the video. Traditionally in video surveillance with a xed
camera, researchers tend to nd some sort of motion in the video. There are two
part of such of videos, background and foreground part. The object in motion
is the foreground part of the video and the rest static part is the background.
Motion detection is used to extract foreground part from the video. Such kind
of extraction is useful for detecting, tracking and understanding the behavior of
the object. A survey on motion detection techniques can be found in [18]. There
are two types of motion: 1) large-scale body movements like movement of head,
legs and arms [19], and 2) small scale body movements like hand gestures and
facial expressions [20], [21].</p>
      <p>A. Foreground Segmentation: A popular and traditional foreground object
segmentation method is a background subtraction. It calculates the di erence
between current image and background image and detects the foreground by
setting up the threshold value. Pixels in the current frame that deviate signi cantly
from the background are considered to be moving objects. These foreground
pixels are further processed for object localization and tracking. This technique is
prone to errors when there is a change in illumination in the video.</p>
      <p>B. Approximate Median: median lter is another technique for motion
segmentation. It bu ers N number of frames and median of these frames are
calculated which will be the background reference frame. This method is e ective
in some cases but many frames have to be stored for calculating median frame
which makes it not suitable in most of cases. Also this method will end up with
errors if there is change in illumination.</p>
      <p>C. Gaussian Mixture Model : The GMM is one of the most commonly used
methods for background subtraction in most of visual surveillance applications.
A mixture of Gaussians is maintained for each pixel in the image. with time
to time, new pixel values updates the mixture of Gaussians using an online
Kmeans approach. This updating of mixture of Gaussians is used to account for
illumination changes, sensor movements and noise.</p>
      <p>D. Temporal Di erencing : Another common approach for motion
segmentation is the temporal di erencing. In temporal di erencing, video frames are
separated by a constant time interval and compared to nd the regions that are
changed. A small time interval between the frames can increase the robustness
to illumination changes. Temporal di erencing approach is computationally
inexpensive but in some cases it fails to extract the shape of the object and cause
small holes.</p>
      <p>
        E. Optical Flow : optical ow estimates the motion by matching points on
objects over multiple frame using vectors. Optical ow technique gives more
accurate estimation of motion if the frame rate is high. Horn and Schunck [21],
Lucas and Kanade [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] , and Szeliski and Couglan [22] are popular techniques
for calculating optical ow. A comparison of these methods can be found in [23].
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Object Tracking</title>
      <p>Tracking is de ned as the problem of estimating the trajectories of objects in
image plane. Video surveillance has motivated many researchers to explore tracking
techniques. Tracking moving objects in a video sequence is a di cult job.
Occlusion makes tracking a di cult problem for the researchers. In video surveillance,
normally some features of the moving objects are extracted and tracking that
object using those features. Selecting good features that can be used for tracking
is very important, since the object appearance, color and orientation may change
from frame to frame. So we need to extract those features that can be tracked
for a long period of time.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Proposed Methodology</title>
      <p>Pedestrian injuries and fatalities are most signi cant problems related to road
safety. Crosswalks are the point where pedestrians face lower level of safety and
most of time they end up with serious accidents. Therefore, as a solution an
automatic monitoring system based on computer vision technology is needed
that can automatically detect the behaviors of pedestrians on crosswalks and
alarm the system in order to prevent pedestrian vehicle collisions. Figure 1 shows
the methodology, video is streamed from the camera to the system and system
processes the video frame by frame. From each frame, foreground is segmented
which represents the objects of interest (pedestrians). Blob analysis is performed
to nd out the independent blobs of a particular size. Corner points are extracted
from each bounding box. Later on, these points are tracked through number
of frames using Luckas Kanade point tracker. Instantaneous velocity of each
point related to bounding box is calculated and average speed of the object is
determined by equation discussed in the following section.
4.1</p>
      <sec id="sec-4-1">
        <title>Foreground Segmentation</title>
        <p>Identifying moving objects in video sequence is a fundamental and critical task
in video surveillance and gesture recognition in human-machine interface.
Foreground segmentation is an important pre-processing step for detecting moving
objects from the video. Traditionally, background subtraction method is used
for extracting moving objects from the video frame where pixels in the currents
frame that deviate signi cantly from the background are considered as part of
moving objects. Such kinds of methods are usually prone to errors due to
unpredicted and changing behavior of the pixels. In addition, this method can not
accurately detect fast moving or slow moving as well as multiple objects. Also
these methods are a ected by change in illumination in the video frame. Some
time change in illumination in static background will be detected as part of
moving object. Such errors and noise must be removed from the foreground objects
before applying blob analysis and tracking.In order to extract valid and
accurate foreground objects, we employed both Gaussian mixture model and Lucas
Kanade optical ow as in [24]. There are ve popular optical ow methods:
gradient based algorithms, block based algorithm, energy based, phase-based
algorithm and neuro-dynamic algorithms [25].As optical ow can not get rid of
change in illumination so we use Gaussian Mixture Model in combination with
optical ow to extract an accurate foreground objects. GMM is more robust to
light changes and slight movements for small image sequence. In order to obtain
a foreground object without noise, we make use of both LK optical ow and
GMM method as show in Figure 2.As shown in Figure 2, Lucas Kanade
optical ow and GMM are applied in parallel to the input image.LK optical ow
is applied to two adjacent images of video i.e. f (x, y, t1) and f (x, y, t) and
output of LK optical ow will be in the form of magnitudes. LK optical ow
tends to nd the corresponding points on f (x, y, t1) on next frame f (x, y,
t).we discuss LK optical ow in more detail in the following section. After
calculating LK optical ow between two frames, we use threshold Thof to segment
motion from static background. Thof value can be calculated experimentally.
The range of Thof is [0.005 0.002] which is calculated experimentally. The range
of Thof for di erent videos is di erent. For slow moving objects the value of
Thof will be low and for fast moving objects the range of Thof will be high. It is
matter of fact that fast moving objects generates optical ow vectors with high
magnitudes and slow moving objects generate ow vectors of lower magnitudes.
So for the videos, where objects moving with di erent speeds, mean value of all
ow vectors should be taken as Thof . In extracting moving part from the image,
the pixel with large magnitude than Thof will be classi ed as foreground while
the pixels whose magnitudes are less than Thof will be the part of background.
In the same way, we get foreground objects by applying GMM, but the GMM
ends up with errors as shown in Figure 2. So in order to extract accurate
foreground we apply logical product of foreground mask generated by LK optical
ow and GMM. Later on, we apply Morphological processes like morphological
opening and closing on the binary image generated by logical product of LK
optical ow and GMM. The morphological open operation is erosion followed by
dilation eliminates smooth contours and protrusions while morphological close
is dilation followed by erosion smooth the section of contours, eliminates small
holes and lls gaps in contours. These operations are dual to each other. Later
on, ood ll algorithm is applied to ll small holes. The output image f out (x,
y, t) from Morphological processing block contains accurate foreground objects
while will applied to Blob analysis block for detecting moving objects.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Blob Analysis</title>
        <p>Blobs are the connected regions in the binary image. The purpose of blob
analysis is to detect those points or regions in binary image that are di erent from
other part of image in terms of brightness or area etc. Following are the steps
for nding connected components in the binary image.
1. Search for unlabeled pixel
2. Label all the pixels in connected region containing p by ood ll algorithm
3. Repeat step 1 and step 2 until all pixels are labeled.</p>
        <p>In the next step, we measure the area of each connected component. Area of
connected component is the number of pixels in the region. There may be di erent
moving objects in the video frame with di erent area sizes. In Transportation
surveillance system, video frame contains multiple objects like pedestrians,
vehicles of di erent sizes. In this paper, we are interested only in pedestrians whom
we want to detect and track over multiple frames. The area of pedestrians is
normally contains less number of pixels than vehicles and trucks in other words
the size of pedestrian is less than vehicle and trucks. On the basis of this
assumption, we set the upper and lower bound of blobs area. The upper and lower
bound of blobs area can be found experimentally. The size of blob also depends
on the resolution of frame. In this paper, we use videos of resolution 576 x 768
pixels and lower and upper bound of blobs area is [1000 2500] pixels. The
connected component (object) will be classi ed as pedestrian if its area lies within
the upper and lower bound otherwise it will be discarded.
4.3</p>
      </sec>
      <sec id="sec-4-3">
        <title>Corner Points Detection and Lucas Kanade Point Tracker</title>
        <p>
          In order to track objects detected in blob analysis step, we use KLT feature
tracker [26] where motion is detected by using pyramidal Lucas- Kanade
optical ow method, using Shi and Tomasi feature detection algorithm [27]. They
pyramidal implementation of Lucas- Kanade gives more robustness against huge
movement with di erent speeds. There are two categories of optical ow
algorithms 1) sparse optical ow 2) dense optical ow. Dense optical ow methods
like Horn- Schunck [21] estimates the displacement of all pixels of image while
sparse optical ow algorithms, such as Lucas-Kanade approach estimates the
displacement for selected number of pixels. Sparse optical ow gives more
robustness towards noise. In sparse optical ow methods, the selection of pixels
should be done automatically and done wisely. There are various methods in
literature for selecting features to be traced correctly. Robust features can be
found and tracked over multiple frames using scale invariant feature transform
(SIFT) algorithm [28]. Speeded-Up Robust Features (SURF) [29] also presents
a method for extracting interest points and describe these points for fast
comparison. Canny edge detector [30] detects edges in image. Then image is divided
into small blocks and we search for the closest edge pixel from the center of the
block. If the pixel is found, it will be regarded as feature point for that block.
Harris Corner detector [31] detects the points where two edges are detected.
One of the disadvantage of SIFT and SURF algorithms is that they are
computationally expensive. The comparison of all these feature detection techniques
can be found in [32]. In this paper, we use Shi and Tomasi algorithm for
extracting corner points. Figure 3 (a) shows the pedestrian detected during blob
analysis step. In the next, we detect corner points of each bounding box using
Shi and Tomasi corner detector as shown in Figure 3 (b). But through our
experiments we realize that points detected in the rst frame may not be tracked
over multiple frames. This is due to dramatically change in the appearance of
the objects and change in intensity values of pixels. Such kind of change always
results in tracking failure. Here, it should be noted that our aim is to nd the
instantaneous velocity of valid pixels. A pixel will be valid pixel if its forward and
backward trajectory does not di er signi cantly. [33] Purposes a method that
can automatically detect the tracking failure by forward and backward tracking
of pixels. As shown in Figure 4 (a) , point p on frame t is to be tracked on frame
t + 1. Let p is a point detected on frame t + 1. This is a forward trajectory of
point p. in order to check the validity of the pixel, the point location p will track
back the point on frame t. this is backward trajectory. Let p is the point location
tracked during backward trajectory. The di erence between the two trajectories
is d. In ideal case, the value of d is zero. So here, we de ne a threshold value
for d. experimentally the value of d is in range of [
          <xref ref-type="bibr" rid="ref13">1 3</xref>
          ]. The higher the value of
d, the more the pixels will be in errors and hence speed of pedestrian will be
erroneous. It should be noted that before calculating the speed of pedestrian,
invalid pixels must be removed. The pixel will be invalid if its related d is greater
than a threshold value otherwise it will a valid pixel. Through our experiments,
we have observed that among group of valid pixels there are some pixels which
appear to be valid but in actual are invalid pixels. Such group of pixels are static
and do not move with the object. So for accurate result, we remove all invalid
and static pixels and consider only valid pixels.
        </p>
        <p>Figure 4 (b) shows the points of one of pedestrian in frame that were detected
by Shi Tomasi corner detector. 0 represents the points (pixels) whose d is
greater than threshold range when tracked over next frame and hence regarded
as invalid pixels while 1 represents the valid pixels. We consider only valid pixels
for nding the speed, direction and understanding the behavior of pedestrian.
4.4</p>
      </sec>
      <sec id="sec-4-4">
        <title>Estimation of Speed</title>
        <p>In order to nd the speed of pedestrian, we must use the valid points. To nd
the speed of pedestrian, instantaneous speed can be found by using successive
image frames of the video. This instantaneous speed for each pixel can be found
by Equation 1.</p>
        <p>v =
d= t
(1)
Where v is the instantaneous velocity of a valid point. d is the change in
displacement of valid point over successive frames. t is time interval between two
successive video frames and is equal to the frame capture rate of the camera. In
our experiments, t is 33.33 milliseconds. To nd the accurate speed of
pedestrian, only one valid point is not enough. We need to nd out valid points for the
same pedestrian, calculate the instantaneous velocities of those points. Then by
averaging instantaneous velocities of all valid points. Lets assume that n valid
points are selected from the pedestrian and let vi represents the instantaneous
velocity of point i. where i = 1::::n. By using those instantaneous velocity
vectors, we can nd the instantaneous velocity of pedestrian by Equation 2:
V iv(t) =
n
X vi(t)
i=1
(2)</p>
        <p>Where V iv is the instantaneous velocity of a pedestrian at time t, vi(t) is
the instantaneous velocity vector of ith point and n is the number of valid points
that tracked. vi(t) in Equation 2, is calculated in pixels per second because
we measure the displacement d between two frame in pixels which has no
correspondence with real displacement. As we know, that objects closer to the
camera will have large pixel displacement than the objects far from the camera
although both objects are moving with same speed in real world. Thus, it is very
di cult to nd the accurate speed of pedestrians using only optical ow. In order
to overcome this limitation, we calculate the displacement of each pedestrian in
the world coordinates. In our case, we calculate the distance between entry and
exit point of pedestrians. Later on, we calculate the displacement in pixels and
derive a linear scale factor to relate displacement in images to the motion in the
world.
4.5</p>
      </sec>
      <sec id="sec-4-5">
        <title>Estimation of Orientation</title>
        <p>At intersections and pedestrian crossings, pedestrians frequently change their
speed and directions. It is even more dangerous and prone to pedestrian vehicle
collisions if the pedestrians change their walking directions in the middle of the
road. Therefore to avoid such collisions, driver must know when the pedestrian is
going to change his/her walking direction to dangerous area. Therefore,
estimating the orientation of pedestrian on pedestrian crossing becomes very important.
In this paper, we estimate the walking direction of pedestrian by making use of
optical ow vectors of the valid points. Let vel (x, y, u, v) is the velocity vector
of a pixel. Where (x, y) is the coordinate of a pixel and u and v are the horizontal
and vertical movements of the pixel. As the optical ow eld of the foreground
image contains all the velocity vectors vel (x, y, u, v), therefore it is easy to
get the magnitude r (x, y) and angle. Let (x, y) be a direction of optical ow
vector at pixel (x, y) in frame t and is given by Equation 3.</p>
        <p>(x; y) = tan 1(u=v)
(3)
As we are interested in nding the orientation of the pedestrian, therefore we
consider angle information.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Experimental Results and Discussions</title>
      <p>We carried out our experiments on a PC of 2.6 GHz (Core i5) with 4.0 GB
memory and data set from UCF. As shown in Figure 5 (a), pedestrians cross the
road in the opposite directions while the car is moving towards the pedestrians.
Studying such kind of scenario becomes very important for understanding the
pedestrian/vehicle interactions. And in order to avoid collisions we develop a
system that automatically nds the speeds and directions of pedestrians and
vehicles. Figure 5 (a) shows a sample frame taken from a video sequence. Figure
5 (b) and (c) shows the objection detection and tracking results after applying
our algorithm mentioned in section 4. The most important result are shown
in Figure 6. Figure 6(a) shows the di erent speeds of pedestrians. The color
bar shows amount of speed a pedestrian is travelling with. The dark regions
shows that pedestrians are moving with high speed while cyan and yellow regions
represents relatively low speed. Figure 6(b) shows the directions of pedestrians.
As shown in Figure 6 (b) there are two dominant ows, one towards the East
and other towards west. Red color shows that pedestrians are moving towards
west while green color shows pedestrians moving towards east. The vehicle in
cyan color is moving towards south
6</p>
    </sec>
    <sec id="sec-6">
      <title>Conlcusions</title>
      <p>In this paper we proposed a framework to nd the speed and direction of people
moving in the video. Foreground objects (people) are extracted by applying
Gaussian mixture model and optical ow. Blob analysis is performed to detect
pedestrians. Lucas Kanade Tracker is used to track each point across multiple
frames. Later on, average speed (pixels/second) of each pedestrian is computed.
It is observed that proposed framework worked very well in low density scenarios.
We are developing methods and techniques that can automatically map image
coordinates to world coordinates and nd the actual speed(meter/seconds) of
pedestrians and is a part of our future works.
18. T. B. Moeslund and E. Granum, \A survey of computer vision-based human motion
capture," Computer Vision and Image Understanding, vol. 81, no. 3, pp. 231{268,
2001.
19. V. I. Pavlovic, R. Sharma, and T. S. Huang, \Visual interpretation of hand
gestures for human-computer interaction: A review," Pattern Analysis and Machine
Intelligence, IEEE Transactions on, vol. 19, no. 7, pp. 677{695, 1997.
20. B. Fasel and J. Luettin, \Automatic facial expression analysis: a survey," Pattern</p>
      <p>Recognition, vol. 36, no. 1, pp. 259{275, 2003.
21. B. K. Horn and B. G. Schunck, \Determining optical ow," Arti cial intelligence,
vol. 17, no. 1, pp. 185{203, 1981.
22. R. Szeliski and J. Coughlan, \Spline-based image registration," International
Journal of Computer Vision, vol. 22, no. 3, pp. 199{218, 1997.
23. J. L. Barron, D. J. Fleet, and S. S. Beauchemin, \Performance of optical ow
techniques," International journal of computer vision, vol. 12, no. 1, pp. 43{77,
1994.
24. W. Li, X. Wu, K. Matsumoto, and H.-A. Zhao, \Crowd foreground detection and
density estimation based on moment," in Wavelet Analysis and Pattern Recognition
(ICWAPR), 2010 International Conference on. IEEE, 2010, pp. 130{135.
25. M. Cristani, R. Raghavendra, A. Del Bue, and V. Murino, \Human behavior
analysis in video surveillance: a social signal processing perspective," Neurocomputing,
vol. 100, pp. 86{97, 2013.
26. C. Tomasi and T. Kanade, Detection and tracking of point features. School of</p>
      <p>Computer Science, Carnegie Mellon Univ., 1991.
27. J. Shi and C. Tomasi, \Good features to track," in Computer Vision and Pattern
Recognition, 1994. Proceedings CVPR'94., 1994 IEEE Computer Society
Conference on. IEEE, 1994, pp. 593{600.
28. D. G. Lowe, \Object recognition from local scale-invariant features," in Computer
vision, 1999. The proceedings of the seventh IEEE international conference on,
vol. 2. Ieee, 1999, pp. 1150{1157.
29. H. Bay, T. Tuytelaars, and L. Van Gool, \Surf: Speeded up robust features," in</p>
      <p>Computer Vision{ECCV 2006. Springer, 2006, pp. 404{417.
30. J. Canny, \A computational approach to edge detection," Pattern Analysis and</p>
      <p>Machine Intelligence, IEEE Transactions on, no. 6, pp. 679{698, 1986.
31. C. Harris and M. Stephens, \A combined corner and edge detector." in Alvey vision
conference, vol. 15. Manchester, UK, 1988, p. 50.
32. N. Nourani-Vatani, P. Borges, and J. Roberts, \A study of feature extraction
algorithms for optical ow tracking," in Australasian Conference on Robotics and
Automation, 2012.
33. Z. Kalal, K. Mikolajczyk, and J. Matas, \Forward-backward error: Automatic
detection of tracking failures," in Pattern Recognition (ICPR), 2010 20th
International Conference on. IEEE, 2010, pp. 2756{2759.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>W. H.</given-names>
            <surname>Organization</surname>
          </string-name>
          et al., \
          <article-title>Pedestrian safety: A road safety manual for decisionmakers and practitioners</article-title>
          .
          <source>" Geneva and Switzerland</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>J.</given-names>
            <surname>Whitelegg</surname>
          </string-name>
          ,
          <article-title>Quality of life and public management: Rede ning development in the local environment</article-title>
          .
          <source>Routledge</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>E.</given-names>
            <surname>Commission</surname>
          </string-name>
          , E. Commission et al.,
          <article-title>\Care community road accident database," Brussels</article-title>
          , EC: http://europa. eu. int/comm/transport/care,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>M.</given-names>
            <surname>Peden</surname>
          </string-name>
          , R. Scur eld,
          <string-name>
            <given-names>D.</given-names>
            <surname>Sleet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Mohan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Hyder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Jarawan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. D.</given-names>
            <surname>Mathers</surname>
          </string-name>
          et al.,
          <source>\World report on road tra c injury prevention,</source>
          "
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>A.</given-names>
            <surname>Dernellis</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Ashworth</surname>
          </string-name>
          , \
          <article-title>Pedestrian subways in urban areas: some observations concerning their use," Tra c engineering &amp; control</article-title>
          , vol.
          <volume>35</volume>
          , no.
          <issue>1</issue>
          , pp.
          <volume>14</volume>
          {
          <issue>18</issue>
          ,
          <year>1994</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>W. H.</given-names>
            <surname>Organization</surname>
          </string-name>
          et al.,
          <source>\Global status report on road safety</source>
          <year>2013</year>
          <article-title>: supporting a decade of action: summary,"</article-title>
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7. S. for Seniors Working Group (WA),
          <source>Safety for Seniors: Final Report on Pedestrian Safety</source>
          ,
          <year>1988</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8. T. Rosenbloom, \
          <article-title>Crossing at a red light: Behaviour of individuals and groups," Transportation research part F: tra c psychology and behaviour</article-title>
          , vol.
          <volume>12</volume>
          , no.
          <issue>5</issue>
          , pp.
          <volume>389</volume>
          {
          <issue>394</issue>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9. C. Holland and
          <string-name>
            <given-names>R.</given-names>
            <surname>Hill</surname>
          </string-name>
          , \
          <article-title>Gender di erences in factors predicting unsafe crossing decisions in adult pedestrians across the lifespan: A simulation study,"</article-title>
          <source>Accident Analysis &amp; Prevention</source>
          , vol.
          <volume>42</volume>
          , no.
          <issue>4</issue>
          , pp.
          <volume>1097</volume>
          {
          <issue>1106</issue>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>V.</given-names>
            <surname>Bhuvaneshwar and P. B. Mirchandani</surname>
          </string-name>
          , \
          <article-title>Real-time detection of crossing pedestrians for tra c-adaptive signal control,"</article-title>
          <source>in Intelligent Transportation Systems</source>
          ,
          <year>2004</year>
          .
          <source>Proceedings. The 7th International IEEE Conference on. IEEE</source>
          ,
          <year>2004</year>
          , pp.
          <volume>309</volume>
          {
          <fpage>313</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>C. Conde</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Serrano</surname>
          </string-name>
          , L. Rodr
          <string-name>
            <surname>guez-Aragon</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Perez</surname>
          </string-name>
          , and E. Cabello, \
          <article-title>An experimental approach to a real-time controlled tra c light multi-agent application,"</article-title>
          <source>in Proceedings of AAMAS-04 Workshop on Agents in Tra c and Transportation</source>
          ,
          <year>2004</year>
          , pp.
          <volume>8</volume>
          {
          <fpage>13</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <given-names>P.</given-names>
            <surname>Vannoorenberghe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Motamed</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.-M. Blosseville</surname>
          </string-name>
          , and J.-G. Postaire, \
          <article-title>Monitoring pedestrians in a uncontrolled urban environment by matching low-level features,"</article-title>
          <source>in Systems, Man, and Cybernetics</source>
          ,
          <year>1996</year>
          ., IEEE International Conference on, vol.
          <volume>3</volume>
          . IEEE,
          <year>1996</year>
          , pp.
          <volume>2259</volume>
          {
          <fpage>2264</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13. H.
          <string-name>
            <surname>Vceraraghavan</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          <string-name>
            <surname>Masoud</surname>
            , and
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Papanikolopoulos</surname>
          </string-name>
          , \
          <article-title>Vision-based monitoring of intersections,"</article-title>
          <source>in Intelligent Transportation Systems</source>
          ,
          <year>2002</year>
          . Proceedings.
          <source>The IEEE 5th International Conference on. IEEE</source>
          ,
          <year>2002</year>
          , pp.
          <volume>7</volume>
          {
          <fpage>12</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <given-names>A.</given-names>
            <surname>Fascioli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. I.</given-names>
            <surname>Fedriga</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Ghidoni</surname>
          </string-name>
          , \
          <article-title>Vision-based monitoring of pedestrian crossings,"</article-title>
          <source>in Image Analysis and Processing</source>
          ,
          <year>2007</year>
          .
          <source>ICIAP</source>
          <year>2007</year>
          . 14th International Conference on. IEEE,
          <year>2007</year>
          , pp.
          <volume>566</volume>
          {
          <fpage>574</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <given-names>B. D.</given-names>
            <surname>Lucas</surname>
          </string-name>
          , T. Kanade et al.,
          <article-title>\An iterative image registration technique with an application to stereo vision."</article-title>
          <source>in IJCAI</source>
          , vol.
          <volume>81</volume>
          ,
          <year>1981</year>
          , pp.
          <volume>674</volume>
          {
          <fpage>679</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>Stau er and</article-title>
          <string-name>
            <given-names>W. E. L.</given-names>
            <surname>Grimson</surname>
          </string-name>
          , \
          <article-title>Adaptive background mixture models for realtime tracking," in Computer Vision</article-title>
          and Pattern Recognition,
          <year>1999</year>
          . IEEE Computer Society Conference on.,
          <source>vol. 2</source>
          . IEEE,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <given-names>C.</given-names>
            <surname>Cedras</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Shah</surname>
          </string-name>
          , \
          <article-title>A survey of motion analysis from moving light displays," in Computer Vision</article-title>
          and Pattern Recognition,
          <year>1994</year>
          . Proceedings CVPR'
          <fpage>94</fpage>
          .,
          <source>1994 IEEE Computer Society Conference on. IEEE</source>
          ,
          <year>1994</year>
          , pp.
          <volume>214</volume>
          {
          <fpage>221</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>