<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Automatic activity estimation through recognizing and handling objects in video sequences for image annotation and retrieval</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Francisco E. Mart´ınez-Pe´rez</string-name>
          <email>fmartinezperez@acm.org</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>He´ctor G. Pe´rez-Gonza´lez</string-name>
          <email>hectorgerardo@uaslp.mx</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>J. Angel Gonzalez-Fraga</string-name>
          <email>angel_fraga@uabc.edu.mx</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mo´ nica Tentori</string-name>
          <email>mtentori@uabc.mx</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Facultad de Ciencias, Universidad Auto ́ noma de Baja California</institution>
          ,
          <addr-line>Ensenada, Me ́xico</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Facultad de Ciencias, Universidad Auto ́ noma de Baja California</institution>
          ,
          <addr-line>Ensenada, Me ́xico</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Facultad de Ingenier ́ıa, Universidad Auto ́ noma de Baja California</institution>
          ,
          <addr-line>Ensenada, Me ́xico</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Facultad de Ingenier ́ıa, Universidad Auto ́ noma de San Luis Potos ́ı, SLP</institution>
          ,
          <addr-line>Me ́xico</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2011</year>
      </pub-date>
      <volume>1</volume>
      <issue>1</issue>
      <fpage>11</fpage>
      <lpage>23</lpage>
      <abstract>
        <p>Automatic estimation of human activities is widely studied topic. However, the process becomes difficult when we want to estimate activities from a video stream, because human activities are dynamic and complex. Our contribution is focused on activity estimation based on object behavior through automatic analysis of video sequences. Another contribution is focused on providing a tool with the aim of monitoring activities in a health-care environment. Our activity estimation process was developed in four phases: The first phase includes the detection of the interactions in the setting by slit-scanning technique; the second phase includes object recognition by composite correlation filters; the third phase follows several criteria for activity estimation. When the behavior of the objects related to the activities is validated, the estimation of an activity is confirmed. Each activity is related to the handling of objects, date and time of the activity, and activity description. All this information is recorded in a database; and after this the last phase includes the activity representation, using indexes for image recovery related to each activity, allowing us to create an activity representation. The activities are estimated at a 92.72 percentage of accuracy including hygiene, feeding and taking of vital signs.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        The tracking of human movement (human tracking [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]) using video sequences and the recognition of
the type of human activities (human activity recognition [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]) are important tasks with multiple
applications for video surveillance: human computer interaction including teleconferencing and content-based
video-retrieval [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] from digital repositories and so forth. These areas have mainly focused on the
context retrieval based on data related to different kinds of environments [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. It has contributed to the
development of computational systems capable of interpreting automatically a video sequence and
extracting useful information. It is necessary to know which information is really relevant to the automation
process related to human activity recognition, in order to capture and include this kind of information in
the requirements analysis to develop specific algorithms.
      </p>
      <p>
        Currently, one of the main used techniques for activity recognition is computer vision. The use of this
technique is attributable to the increased computational power that allows huge amounts of video to be
processed and stored. However, automatic activity recognition is not a simple task due to the particular
way humans perform their activities and the different tools or objects used in those activities. Therefore,
the nature of human activities poses the following challenges: i) Recognizing concurrent activities; ii)
Recognizing interleaved activities; iii) Ambiguity of interpretation and iv) Recognizing multiple persons
as shown in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        Due to the availability of large and steadily growing amounts of visual and multimedia data,
Content Based Image Retrieval (CBIR) has created thematic access methods that offer more than simple
text-based queries based on matching exact database fields. For example, in the medical field, digital
images used for diagnostics and therapy, are produced in ever increasing quantities. In reference [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]
was reported a review that concentrates on image retrieval in the medical domain, that does a
systematic overview of techniques used, visual features employed, images indexed and medical departments
involved. Although the need for information in a quick and timely decision-making has been increasing
in hospital environment, a few research efforts have being reported for automatic activity retrieval. In
reference [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] was reported the need to maximize the attention span and decrease the spent time to record
health-care of patients where one option is the automation of recording of health-care activities. Many
systems propose to use text from the patient record [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] or studies [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] to search by content in distributed
data bases. Other researches classify the images to augment text-based search with visual information
analysis [
        <xref ref-type="bibr" rid="ref21 ref8">8, 21</xref>
        ] or a semi-automatic method for image annotation shown in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] . Basically all systems
that give details use color and grey level features, mostly in the form of a histogram [
        <xref ref-type="bibr" rid="ref17 ref21">17, 21</xref>
        ].
      </p>
      <p>The work cited above, it has assumed that images or video sequences are previously stored in several
data bases before the retrieval process; even images or video sequences are acquired or labeled in a
manual form. This work shows the whole process since acquisition to activity representation, and automatic
labeling of images in video sequences. We use the activity term as annotations in images that are linked
to one or several objects used in the health care activities.</p>
      <p>One contribution of this paper is the activity estimation based on object behavior through automatic
analysis of video sequences. This approach automatically recognizes the human interactions that happen
in a specific setting, through handling several objects in a base location, which leads us to infer the
activity in an automatic fashion. To check when an activity happens, we recognize the interactions taking
into account a whole view of the scene. The information obtained from the recognition and behavior of
objects is processed and used to obtain visual representations of activities.</p>
      <p>Another contribution is focused on providing a tool with the aim of monitoring activities in a particular
health-care environment. Specifically, we provide inferences and representations related to activities that
caregivers perform to the elderly. We identified four important activities related to health-care activities:
the taking of blood pressure; the measuring of blood glucose; the activity of providing patient hygiene
and the feeding activity.</p>
      <p>This paper is organized as follows. In section 2, we present the implemented methods. Section 3
discusses the interactions in the scene. In section 4, we present the object recognition using composite
correlation filters. In section 5 we show the activity classifications. In section 6 we present how
recovering information is related to the activities. In section 7 we present our results. Finally, Section 8 presents
the conclusions and directions for future work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Methods</title>
      <p>
        In our approach, activity recognition is focused on what is happening in the scene, where both object and
people are involved for their recognition, on tracking such objects from frame to frame, and on analyzing
object tracks to recognize their behavior in a scene. So, it is necessary to know and model the
context/environment that we are interested in monitoring as mentioned in [
        <xref ref-type="bibr" rid="ref20 ref22 ref5">20, 22, 5</xref>
        ]. Therefore, we developed
a case study that provided us functional requirements to create a monitoring system. The case study also
considered the user’s requirements related to recovery and representation of the information based on
performed activities.
      </p>
      <p>
        The study was conducted in a private nursing home in the city of Ensenada , Mexico. Figure 1 shows
the patient’s room being monitored and the areas where the interactions flow of humans and objects is
realized. We captured almost 400 hours of video sequences of the setting. The observational study was
recorded with two cameras as shown in Figures 1b and 1c.
This work is based on the hypothesis that human activities can be inferred by the interactions with objects
[
        <xref ref-type="bibr" rid="ref16">16</xref>
        ][
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].For this reason, the proposed process for activity estimation was developed in four phases: The
first phase includes the detection of interactions between humans and objects in the scene through the
implementation of the slit-scanning technique proposed by [
        <xref ref-type="bibr" rid="ref12 ref3">12, 3</xref>
        ]. If an interaction is detected, it means
that an activity has started and must be validated.
      </p>
      <p>Validation of an activity involves the recognition of objects that are handled in the setting, their
behavior and the duration of the activity. The second phase takes care of object recognition, which was
implemented with composite correlation filters. In the case study we identified the kind of objects that
were related to a particular activity and the duration of each activity; so, as long as we have these features
we will be able to create an activity description. All this information is recorded in a database.</p>
      <p>The third phase identifies objects behavior and it automatically links them to a specific activity. In
an activity, caregivers can use one, two or more objects, depending on the activity, the objects show a
different handling behavior, and so we need a recognition filter for each current object in the setting.</p>
      <p>
        Confirmation of activities allows us to create both indexing and linking of the video sequenced
images with each activity in the database, as well as recovering more information from the scene for
contextual representation of the activity. This kind of representation can be used as a tool for quick reference
video related to monitoring. The last phase explains how the indexes are used for information retrieval
related to the activities that happened in the setting and it allows us to create an activity representation
and description. These four stages were adapted to the general framework of visual surveillance
proposed by [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. We adapted the object detection model for an optical flow model. The second adaptation
is related to the elimination of object segmentation module of the original framework. This adaptation
was necessary because the composite correlation filter method it does not need image segmentation to
recognize and track objects. Figure 2 shows the general framework that includes these adaptations in
the second stage (mentioned above) and stages related to this work that will be discussed later .
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Detecting interactions</title>
      <p>
        In healthcare environments for elders it is very important to know what kinds of interactions are
performed by caregiver towards them. For example, it is important to check when healthcare activities are
performed; therefore, it is necessary to check that the activities are being performed in a correct form.
We will use the term interaction to refer to the action in which a caregiver is handling objects and the
actions when the caregiver is moving from base location to the patient’s bed or vice versa . This is with
the purpose to know interactions in the scene. According to figure 1, interactions can be developed in
zones A and B. Interactions in zone A is when caregivers are handling objects in a base location. An
interaction in zone B is when someone performs an activity from base location to the patient’s bed or
vice versa. Interactions detection in both zones A and B are obtained using slit-scanning technique. This
technique creates a composite image of video activities over time. Slit scanning, originally developed
in photography, exposes film to only a narrow slit from a scene; while panning the camera smoothly
captures a normal scene, interesting images are created by irregular panning (spatially distorted scenes),
or when objects moving in the scene are seen as motion over space. The same approach is realized in
video by video slicing. Video slicing first extracts a scan line (vector) from a video frame, and then adds
that line to a composite image over time as shown in figure 3. The contiguous video slices in each row
give observers a sense of the video history, where changes and perturbations are easily seen as presented
in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]
3.1
      </p>
      <sec id="sec-3-1">
        <title>Interactions as cues</title>
        <p>To get the interactions , we obtained two pixel vectors of each image in a video sequence as shown in
figure 4. Vector a belongs to the patient bed edge and the width size is a pixel (figure 4-(a)), and vector
b, belongs to the base location as shown in figure 4-(b). Position related to the two vectors is due to the
behavior that the caregivers showed in their activities. The main idea of handling of these two vectors
is to find drastic differences which are linked to interactions in the scene, being a sign that something
happens.</p>
        <p>The process to obtain the interactions is as follows: Each vector is normalized using its highest value,
because highest values show a significant change related to the person or object in the scene. From
vector 2 to 30, we obtain the sum of the absolute difference by the incoming vector with its predecessor.
If there is a difference greater than 2, then the first vector is stored in a temporary variable, and also it is
considered that there is an interaction in the scene. In this way, the system starts to record the behavior
and take the significant changes as an interaction cue that is performed. We take into account the previous
15 minutes before the interactions happen in the scene, as a threshold reference to compare the behavior
in the scene versus lighting variations that happen in a day. When highest values exceed the threshold,
the interaction is checked in the room. Figure 4 shows an example related to the execution of blood
pressure activity. The black line in figure 4b shows the result of the interactions performed in the zone
A. The red line in figure 4a shows the result of the interactions performed in the zone B. The process
starts without interactions in the room from frame 1 to 500 in figure 4. After that, a caregiver enters in
the room and moves to the base location where he puts an object on the table (from frame 501 to 550).
In this zone, the value of the line changes and remains high, and this is due to the object that was put
there as shown from frame 601 to 1750. In the range from frame 501 to 1801 it can be seen how the
lines are crossing, meaning the transition between zones from A to B. From frame 2101 to 2401 it can
be seen another example performed in the zone A. Finally, from frame 3150 to 3450 it can be seen two
interactions performed in the zone A but between them there is an interaction performed in the zone B.
The analysis showed that when there is not an activity performed in the room, the lines behavior is almost
constant and without exceeding a difference more than 25 points as shown in figure 4 from frame 1 to
frame 500. Several times peaks are visualized that belong to adjustments of the intensities of light that
affected the cameras as shown in figure 4 from frame 1950 to 1980 in black line (figure 4b). When a
caregiver enters and leaves at an instant in the room, the lines in the graph shows a significant variation
as can be seen with the highest peaks in figure 4 from 3101 to 3201 among other times. However, when
a caregiver remains immobile, his variation leads to stabilization within a range of at least 15 as shown
from frame 501 to 1801 of the line in figure 4a.</p>
        <p>Using slit-scanning technique it was possible to check the interactions performed in the setting. The
highest values are related to the activity performed when it starts or ends. Therefore, these highest values
represent a cue of interaction, namely something is happening in the scene. Based on these highest
values, the indexes related to these frames can be saved in a database as a cue of when the interaction
starts or ends. However, these data are just to begin the activity estimation, so in the following sections
are discussed the steps to consolidate the activity estimation.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Object recognition by correlation filters</title>
      <p>Based on our study, we decided to create tags for the recognition of six objects related to four impo rtant
health activities: the baumanometer object is related to the activity of taking blood pressure; the dextrose
kit with the sample is related to the activity of measuring blood glucose; the tray is related to feeding
and there are three objects involved with the activity of patient hygiene such as the toilet paper, saline
solution and a lotion for the body. These objects are manipulated in a base location so that we confirm
the hypothesis for the inference of activities from the handling of objects. We used the optical flow
advantage by composite correlation filters implementation as objects recognition and tracking method in
dynamic settings. To this end, we use MatLab as a programming language.
4.1</p>
      <sec id="sec-4-1">
        <title>Non-linear composite correlation filters</title>
        <p>
          Object recognition based on correlation filters computes a level of similarity between two images: i)
the reference images and ii) the test image or the captured setting frame as figure 5 shows. The image
of the workplace scene is used to test the filter in real time and it is matched with a reference image
previously recorded and used to train the filter. One of the advantages in using correlation filters is
that we can locate multiple objects without segmentation in a scene, reducing the processing time. We
implemented the kth-law synthetic-discriminant function because it has been shown that nonlinear filters
have tolerance to some object deformations and good performance in the presence of different types of
noise [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].
The kth-law composite filter in the frequency domain is:[
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]
        </p>
        <p>hˆk = Sˆk((Sˆk)+Sˆk) 1c
Where the S matrix of size PxN (number of pixels in each images and number of training images), has
the vector form of the Fourier transform of each training image as its i-th column, S+ is the complex
conjugate transpose of S and () 1 denotes the inverse matrix. Vector c contains the desired cross-correlation
peak value for each training image and factor k is a nonlinear operator affecting the module of each
Fourier transform in S, k is a real value between 0 and 1. For doing the correlation operation we
reordered the vector h into his 2D form, and finally we get the filter H(u; v). The correlation operation
showed in figure 5 can be defined in terms of Fourier transform by</p>
        <p>c(x; y) = T F 1(G(u; v)H (u; v));
where H (u; v) is the complex conjugate of the k-law filter, G(u; v) is the test scene preprocessed with
the nonlinear k-law factor, T F 1 is the inverse Fourier transform and c(x; y) is the correlation output.
(1)
(2)
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Activity classification</title>
      <p>
        The correlation values, obtained in the previous section, are converted to ones and zeros where1 stand
for an object that has been recognized and 0 when there is not any object in the scene. In this way we
converted the whole stream of video sequences in a train of pulses as can be seen in figure 6a and 6b.
Each activity can be composed by one, two or several objects. Each object has a different behavior
related and according to the activity it performs. An objects can appear in the setting one time, two times
or several times, producing a signature related to its behavior when entering and leaving the setting as
shown in figure 6; and is processed in a thread in which its behavior is analyzed. In order to relate each
object to an activity, it is necessary to validate three states; and every observed object must perform one
of them [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]:
i) Initialized, when an object is being set up and placed in its base location (e.g., the tools table).
ii) Activated, when object’s beat changes (e.g., from a motionless to a mobile state) or remains in the
same state (e.g., mobile) and;
iii) Suspended, when an artifact remains motionless and it is suspended.
      </p>
      <p>These three states are checked all the time. Namely, the initialization phase starts when the first object
appears in the base location. This occurrence is an action and is recorded in a temporal database. When
the object is recognized, the time and date in which the event started are stored; then, the activated
state starts and checks the objects’ behavior in the base location. This means that an object’s behavior
is happening again. The time and date in which the recognized objects enters and leaves are recorded
again. Finally, once the period of time after the object disappears or remains motionless longer than a
threshold, the activity is suspended. At this point, if the object’s behavior is similar to the programmed
behavior then the activity is inferred and data related to that activity (activity name, number of objects
recognized, date, activity time in which starts and ends, and the duration of the activity) are recorded in
the database. Criteria timeouts for mobility-immobility were proposed according to the average times
observed in the case study and is equal to the average total time of activity divided by the number of
appearances on the scene.</p>
      <p>In order to implement this behavior signature in the inference of the four activities, it was necessary to
use concurrent processes. For each process, we classified the activities independently which allowed us
to infer several activities simultaneously. Each process is responsible for retrieving information relevant
to the object, and the proper analysis of that behavior.</p>
      <p>Grouping objects and linking these groups to specific activities; and three states related to the object’s
behavior, allowed us to accomplish the early steps of the challenges related to activity recognition. Such
activities include concurrent activities, interleaved activities and ambiguity of interpretation, using vision
techniques. Each time an object enters or leaves the scene it will be recognized the index related to the
video sequence will be linked and will be recorded in the temporal database. This process will be
executed including the camera installed in the base location and the other cameras installed in the scene.
All this, allowed us to obtain a better representation related to the activity that is discussed in the next
section as an example of recovering a representation.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Recovering a representation</title>
      <p>Building a representation based on activity interactions is a hard work, because it is necessary to know
what information is relevant related to specific points. Such points in this approach are obtained through
indexes created in the inference of the activity. The activity includes where the starting and ending
indexes are and also the indexes related to the handling object. Therefore, indexes are used for
image recovery related to each activity, allowing us to create an activity representation. In the same way,
slit-scanning technique gave us an interaction representation related to the scene as mentioned above,
however there are a few elements of the event to highlight.</p>
      <p>We relied on our inference results where each activity is related to indexes and has a duration of
execution. We selected timeline technique to create a representation based on a span of time (duration of
activity) in an easy way. The timeline allows us to a very rapid and detailed exploration of the video
history where each slice can highlight details related to an activity. We extend the timeline
representation using slit-scanning technique by obtaining several vectors to complete a window size of 250, which
starts from 100 to 350 lines obtained of 250 frames of video sequence as shown in figure 7d. This kind
of view extends the exploration area, and in this case our representation allows us to show a better view
of the objects that are being used in a clear way as shown in figure 7b(4). Figure 7 shows an example in
which one hour was obtained (figure 7a) and four performed activities can be seen. Figure 7b shows a
specific activity performed in which can be seen the specific object used in the activity. After that, figure
7c shows a representation obtained in one minute. Finally, figure 7d shows the x position in which
starts getting the vectors and this position goes increasing and moves according to the frame in the video
sequence until the complete window size of 250 and restarts again at starting position of x and repeats
again the process.</p>
      <p>Another representation is based on windows that are created by each activity. Where images are
recovered that is based on indexes related to the activity of video sequence. Namely, for each window created,
we obtain two images that are located in the middle zone of each one of them; also two other images are
located in the start and end of each window. For example figure 7b shows seven windows created in a
whole activity, so we obtained 28 images to represent the activity. We know that with this representation
we can omit some details related to activity performance; however, our intention is to reduce the query
time in videos so that this does not become too tedious. Furthermore, this information can be configured
to send alarms related to risk events.
In this subsection, we present our approach gathering the phases as a system. First, we introduce the
interactions in the setting. This phase is strongly linked to the object recognition (phase 2). Namely, the
system checks any interactions that happen in the scene. Once detected, the system is waiting that the
object recognition module gets the first result with the aim to record it in the database. After the system
continues awaiting that next interaction will happen in the adjacent area. Once time the system identified
several movements in both areas, the activity is confirmed and something happens in the setting.
On the other hand, for each time object recognition is done, these results are sent to the phase 3. In which,
the results are converted to create the activity estimation based on object behavior and their manipulation.
Finally, as result of this phase, we obtained the indexes related to the video sequences and it is possible
to construct an activity representation and determines whether it is necessary to send an alarm or simply
save the representation as a reference tool.
7</p>
    </sec>
    <sec id="sec-7">
      <title>Results and discussion</title>
      <p>The system was evaluated in the usability lab at UABC campus Ensenada, Mexico, over a period of
5 days and where we replicate the handling of the objects used within the activities inside the nursing
home. In this evaluation, we performed 11 activities per day: 4 hygiene activities, 3 feedings, 2 blood
pressure measure and 2 samples of glucose (dextrose). In total, we carried out 55 activities that were
recorded and processed. Our system was able to recognize 51 activities, and the rate of effectiveness
for activities inference was 92.72 percent. The missing activities that were not identified, were because,
these did not exceed the established threshold. This was due to occlusions, shadows and light changes,
affecting the outputs of correlation filters and producing results below threshold. Another cause was the
speed how caregivers perform the activity because the time for object recognition is not enough. On the
other hand, the interaction was used as cue to detect whether an activity is performed, but the activity is
confirmed only if the objects have a known behavior as described in Activity Classification Section.
The execution of the activities was recorded by three cameras wvc53gca Linksys in the usability lab.
The video was captured using MPGe-4 format with a resolution of 320x240 pixels. One camera was
located on the top of the room; another one was located in proximity to the preparation table and the
third was located near to the left side of the patient. Figure 8 shows an activity performed related to the
blood pressure activity. Also figure 8a shows the video sequence captured in which it can be observed
that there are several objects however just one object is the target (dextrose kit object) and other ones
represent a kind of noise. The target object appears in the setting two times as shown in figure 8b. In
this process the three states of initialization, activation and suspension that are related to the object are
recognized. After that, the correlation results are converted in a train of pulses as shown in figure 8c.
Obtaining this conversion allowed us to create a meaning related to the activity and the train of pulses.
The information obtained of the conversion, belong to object recognition that in this case the object is the
baumanometer; number of occurrences; time and date in which the activity was performed; duration of
the activity; and a meaning that belongs to the behavior related to: 1) Object appears in the base location,
meaning that the user is interacting in the room; 2) the user is interacting with the patient because the
object is being used and 3) the object appears again and the user is recording the activity as shown in
figure 8 (1,2,3). Also, we obtained a representation that is based on the indexes of each frame, which was
obtained in the activity performed where we can highlight that: 1) the image in which first time the object
appears in the setting and the user is changing his position from zone A to zone B; 2) the image when
the object disappears of the base location; the image of the activity performed and the image obtained 2
seconds before that the object was removed of the patient; 3) the image when the user is recording the
activity; and the image where the object is removed to the base location.</p>
      <p>Once the activity has been inferred, the information related to the activity is stored and available in the
database. The activity is labeled and is possible to identify it as ”Blood pressure activity”. Also, this
activity is linked to an abstract level that belongs to the vital signs activity.</p>
      <p>These results allowed us to identify the interactions in the scene. Object recognition and classification
based on the states that arose according to objects’ behavior was accomplished. Finally, we presented an
activity representation whose aim is to provide a tool for checking the execution of the activities.
8</p>
    </sec>
    <sec id="sec-8">
      <title>Conclusions and future work</title>
      <p>
        We presented an approach for automatic activity estimation. Our approach was developed based on
a modified version of the general framework proposed by [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. We divided our development in four
phases that includes recognizing of the interactions in the scene; object recognition through using
composite correlation filter; activity classification; and two activity representations related to the healthcare
activities performed by caregivers in a visual fashion.
      </p>
      <p>Using slit-scanning technique we checked the interactions that happen in a scene. The highest values
are related to the activity performed when it starts or ends. Therefore, these values represent a cue of
interaction, namely something is happening in the scene. Based on these highest values, the indexes
related to the activity are saved in a database as a cue of when the interaction starts or ends.</p>
      <p>Grouping objects and linking these groups to specific activities allowed us to accomplish the early
steps of the challenges related to activity recognition. Such activities include concurrent activities,
interleaved activities and ambiguity of interpretation, using vision techniques.</p>
      <p>Our algorithms were able to recognize 51 activities from 55 performed, thus the rate of effectiveness
for activities inference was 92.72 percent.</p>
      <p>Indexes generation of a video sequences allowed us to obtain data related to activities such as time, date,
index of each frame and objects involved in a specific activity. Using slit-scanning technique, we
obtained a better and wider interpretation related to activity estimation through of a visual representation.
We know that with this representation we can omit some details related to activity performance; however,
our intention is to reduce the query time in videos so that this does not become too tedious. Furthermore,
this information can be configured to send alarms related to risk events.</p>
      <p>
        On the other hand, for each activity estimation allows us to create annotation in a data base linked to the
indexes of the video sequences in real time. Using these indexes, we provide an abstract meaning that it
is possible to extend showing their components involved. At the same form, using these indexes through
linking to other video sequences acquired by other cameras, which are installed in the setting. Therefore,
to give another meaning to the activities, where it is easy to see and identify actions. Additionally, it is
possible to see specific situations like activities without to check all video sequence. Allowing that other
applications as shown [
        <xref ref-type="bibr" rid="ref2 ref7">7, 2</xref>
        ] will be focused on images retrieval specifically on these video segments.
      </p>
      <p>This approach was implemented in Matlab, it is due to, allows matrix manipulations, plotting of
functions and data, implementation of algorithms, creation of user interfaces, and interfacing with
programs written in other languages, including java, C, C++, and Fortran. However our intention is develop
this system in another language as java, because it provides several advantages like portability,
generalpurpose, concurrent, and so forth. In addition, we must be aware of the amount of data that we can
obtain from handling video stream. It involves to the elderly, staff and relatives, but their perceptions and
implications have not been considered in this research. Other topics for us are the privacy and security
of data, but these are beyond our reach.</p>
      <p>The evaluation of this work was developed in our usability lab, however our intention is to implement
this approach in real scenarios and this implies having a better performance in our filters. Also, to extend
this approach we have planned to recognize the people that are performing the activities in the room.
Another ongoing work is related to knowing and monitoring the postures of immobile patients. Finally,
we intend to improve our activity representation related to the activities by responding user requirements.
9</p>
    </sec>
    <sec id="sec-9">
      <title>Acknowledgment</title>
      <p>We thank the personnel at Lourdes Nursing Home in Ensenada, Me´xico, especially to Argel Grisolle.
This work was funded through the scholarship provided to the first author (CONACYT 243422/217747).
This work has been developed within the program Maestr´ıa y Doctorado en Ciencias e Ingenier´ıa
(MyDCI) at UABC.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Javidi</given-names>
            <surname>Bahram</surname>
          </string-name>
          ,
          <string-name>
            <surname>Wang Wenlu</surname>
            ,
            <given-names>and Zhang</given-names>
          </string-name>
          <string-name>
            <surname>Guanshen</surname>
          </string-name>
          .
          <article-title>Composite fourier-plane nonlinear filter for distortioninvariant pattern recognition</article-title>
          .
          <source>Optical Engineering</source>
          ,
          <volume>36</volume>
          (
          <issue>10</issue>
          ):
          <fpage>2690</fpage>
          -
          <lpage>2696</lpage>
          ,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Amalia</given-names>
            <surname>Charisi</surname>
          </string-name>
          and
          <string-name>
            <given-names>V.</given-names>
            <surname>Megalooikonomou</surname>
          </string-name>
          .
          <article-title>Content-based medical image retrieval in peer-to-peer systems</article-title>
          .
          <source>In Proceedings of the 1st ACM International Health Informatics Symposium</source>
          , pages
          <fpage>724</fpage>
          -
          <lpage>733</lpage>
          . ACM,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Davidhazy</surname>
          </string-name>
          .
          <article-title>Slit-scan photography</article-title>
          .
          <source>In School of Photographic Arts and Sciences Rochester Institute of Technology</source>
          ,
          <year>2007</year>
          . http://people.rit.edu/andpph/text-slit-scan.html,
          <source>last viewed Nov</source>
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Kim</given-names>
            <surname>Eunju</surname>
          </string-name>
          , Helal Sumi, and
          <string-name>
            <given-names>Cook</given-names>
            <surname>Diane</surname>
          </string-name>
          .
          <article-title>Human activity recognition and pattern discovery</article-title>
          .
          <source>Pervasive Computing IEEE</source>
          ,
          <volume>9</volume>
          (
          <issue>1</issue>
          ):
          <fpage>48</fpage>
          -
          <lpage>53</lpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Bremond</surname>
            <given-names>F.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Nevatia R</surname>
          </string-name>
          .
          <article-title>Representation and optimal recognition of human activities</article-title>
          .
          <source>Proc IEEE Conf. on Computer Vision and Pattern Recognition CVPR</source>
          <year>2000</year>
          , pages
          <fpage>818</fpage>
          -
          <lpage>825</lpage>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Favela</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Tentori</surname>
          </string-name>
          , Luis A.
          <string-name>
            <surname>Castro</surname>
          </string-name>
          ,
          <string-name>
            <surname>Victor M. Gonzalez</surname>
          </string-name>
          , and
          <string-name>
            <surname>Elisa</surname>
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Moran</surname>
          </string-name>
          . Hospital Workers'
          <article-title>Activities and its use in Context-Aware Hospital Applications</article-title>
          . In Pervasive Healthcare, Austria,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Sergio</surname>
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Furuie</surname>
          </string-name>
          , Marina S. Rebelo,
          <article-title>Ramon a</article-title>
          . Moreno, Marcelo Santos, Nivaldo Bertozzo,
          <string-name>
            <surname>Gustavo H. M. B. Motta</surname>
          </string-name>
          ,
          <article-title>Fabio a</article-title>
          . Pires, and
          <article-title>Marco a</article-title>
          .
          <source>Gutierrez. Managing Medical Images and Clinical Information: InCor's Experience. IEEE Transactions on Information Technology in Biomedicine</source>
          ,
          <volume>11</volume>
          (
          <issue>1</issue>
          ):
          <fpage>17</fpage>
          -
          <lpage>24</lpage>
          ,
          <year>January 2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Hayit</given-names>
            <surname>Greenspan and Adi T Pinhas</surname>
          </string-name>
          .
          <article-title>Medical image categorization and retrieval for PACS using the GMMKL framework</article-title>
          .
          <article-title>IEEE transactions on information technology in biomedicine : a publication of the IEEE Engineering in Medicine</article-title>
          and
          <source>Biology Society</source>
          ,
          <volume>11</volume>
          (
          <issue>2</issue>
          ):
          <fpage>190</fpage>
          -
          <lpage>202</lpage>
          ,
          <year>March 2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>L.</given-names>
            <surname>Hollink</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Little</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Hunter</surname>
          </string-name>
          .
          <article-title>Evaluating the application of semantic inferencing rules to image annotation</article-title>
          .
          <source>In Proceedings of the 3rd international conference on Knowledge capture</source>
          , pages
          <fpage>91</fpage>
          -
          <lpage>98</lpage>
          , New York, New York, USA,
          <year>2005</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>W.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Wang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Maybank</surname>
          </string-name>
          .
          <article-title>A survey on visual surveillance of object motion and behaviors</article-title>
          .
          <source>IEEE Transact on Systems, Man, and Cybernetics</source>
          , Part C:
          <article-title>Applications</article-title>
          and Reviews,
          <volume>34</volume>
          (
          <issue>3</issue>
          ):
          <fpage>334</fpage>
          -
          <lpage>352</lpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Teddy</given-names>
            <surname>Ko</surname>
          </string-name>
          .
          <article-title>A survey on behavior analysis in video surveillance for homeland security applications</article-title>
          .
          <source>In Applied Imagery Pattern Recognition Workshop 2008 37th IEEE</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          . IEEE,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Golan</given-names>
            <surname>Levin</surname>
          </string-name>
          .
          <article-title>An informal catalogue of slit-scan video artworks</article-title>
          .
          <year>2006</year>
          . http://www.flong.com/ writings/lists/list_slit_scan.html,
          <source>last viewed Nov</source>
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>F. E.</given-names>
            <surname>Martinez-Perez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Gonzalez-Fraga</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Tentori</surname>
          </string-name>
          . Artifacts'
          <article-title>Roaming Beats Recognition for Estimating Care Activities in a Nursing Home</article-title>
          .
          <source>In 4th Internat. Conf. on Pervasive Computing Technologies for Healthcare</source>
          , Munchen Germany,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>H.</given-names>
            <surname>Muller</surname>
          </string-name>
          , Nicolas Michoux, David Bandon,
          <string-name>
            <given-names>and Antoine</given-names>
            <surname>Geissbuhler</surname>
          </string-name>
          .
          <article-title>A review of content-based image retrieval systems in medical applications clinical benefits and future directions</article-title>
          .
          <source>International journal of medical informatics</source>
          ,
          <volume>73</volume>
          (
          <issue>1</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>23</lpage>
          ,
          <year>February 2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M.</given-names>
            <surname>Nunes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Greenberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Carpendale</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Gutwin</surname>
          </string-name>
          .
          <article-title>What Did I Miss? Visualizing the Past through Video Traces</article-title>
          .
          <source>In ECSCW 2007 Proceedings of the 10th European Conference on Computer Cooperative Work Limerick Ireland 24-28 September</source>
          <year>2007</year>
          , number
          <year>2007</year>
          , page 1. Springer Verlag,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>Matthai</given-names>
            <surname>Philipose</surname>
          </string-name>
          ,
          <string-name>
            <surname>Kenneth</surname>
            <given-names>P</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mike</surname>
            <given-names>Perkowitz</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Donald J.</given-names>
            <surname>Patterson</surname>
          </string-name>
          , Dieter Fox, and
          <string-name>
            <given-names>Henry</given-names>
            <surname>Kautz</surname>
          </string-name>
          .
          <article-title>Inferring Activities from Interactions with Objects</article-title>
          .
          <source>Portal</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>Marcelo</given-names>
            <surname>Ponciano-Silva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.J.M.</given-names>
            <surname>Traina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.M.</given-names>
            <surname>Azevedo-Marques</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.C.</given-names>
            <surname>Felipe</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Caetano</given-names>
            <surname>Traina</surname>
          </string-name>
          .
          <article-title>Including the perceptual parameter to tune the retrieval ability of pulmonary CBIR systems</article-title>
          .
          <source>In Computer-Based Medical Systems</source>
          ,
          <year>2009</year>
          .
          <source>CBMS</source>
          <year>2009</year>
          . 22nd IEEE International Symposium on, pages
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          . IEEE,
          <year>August 2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>Monica</given-names>
            <surname>Tentori</surname>
          </string-name>
          and
          <string-name>
            <given-names>Jesus</given-names>
            <surname>Favela</surname>
          </string-name>
          .
          <article-title>Activity-aware computing for healthcare</article-title>
          .
          <source>IEEE Pervasive Computing</source>
          ,
          <volume>7</volume>
          (
          <issue>2</issue>
          ):
          <fpage>51</fpage>
          -
          <lpage>57</lpage>
          ,
          <year>April 2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Hu</surname>
            <given-names>Weiming</given-names>
          </string-name>
          , Xie Dan, Fu Zhouyu, Zeng Wenrong, and
          <string-name>
            <given-names>Maybank</given-names>
            <surname>Steve</surname>
          </string-name>
          .
          <article-title>Semantic-based surveillance video retrieval</article-title>
          .
          <source>IEEE transactions on image processing</source>
          ,
          <volume>16</volume>
          (
          <issue>4</issue>
          ):
          <fpage>1168</fpage>
          -
          <lpage>81</lpage>
          ,
          <year>April 2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Alper</surname>
            <given-names>Yilmaz</given-names>
          </string-name>
          , Omar Javed, and
          <string-name>
            <given-names>Mubarak</given-names>
            <surname>Shah</surname>
          </string-name>
          .
          <article-title>Object tracking A survey</article-title>
          .
          <source>ACM Computing Surveys</source>
          ,
          <volume>38</volume>
          (
          <issue>4</issue>
          ):
          <fpage>13</fpage>
          -es,
          <year>December 2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>Lei</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.W.</given-names>
            <surname>Wetzel</surname>
          </string-name>
          , John Gilbertson, and
          <string-name>
            <given-names>M.J.</given-names>
            <surname>Becich</surname>
          </string-name>
          .
          <article-title>Design and analysis of a content-based pathology image retrieval system</article-title>
          .
          <source>Information Technology in Biomedicine, IEEE Transactions on, 7</source>
          (
          <issue>4</issue>
          ):
          <fpage>249</fpage>
          -
          <lpage>255</lpage>
          ,
          <year>December 2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lao</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zheng</surname>
          </string-name>
          .
          <article-title>Object Tracking in Structured Environments for Video Surveillance Applications</article-title>
          .
          <source>IEEE Transact on Circuits and Systems for Video Technology</source>
          ,
          <volume>20</volume>
          (
          <issue>2</issue>
          ):
          <fpage>223</fpage>
          -
          <lpage>235</lpage>
          ,
          <year>February 2010</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>