<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>BIDAL@imageCLEFlifelog2019: The Role of Content and Context of Daily Activities in Insights from Lifelogs</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Minh-Son Dao</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anh-Khoa Vo</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Trong-Dat Phan</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Koji Zettsu</string-name>
          <email>zettsug@nict.go.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Big Data Analytics Laboratory National Institute of Information and Communications Technology</institution>
          ,
          <country country="JP">Japan</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Faculty of Information Technology University of Science</institution>
          ,
          <addr-line>VNU-HCMC</addr-line>
          ,
          <country country="VN">Vietnam</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>imageCLEFlifelog2019 introduces two exciting challenges of getting insights from lifelogs. Both challenges aim to have a memory assistant that can accurately bring a memory back to a human when necessary. In this paper, two new methods to tackle these challenges by leveraging the content and context of daily activities are introduced. Two backbone hypotheses are built based on observations: (1) under the same context (e.g., a particular activity), one image should have at least one associative image (e.g., same content, same concepts) taken from di erent moments. Thus, a given set of images can be rearranged chronologically by ordering their associative images whose orders are known precisely, and (2) a sequence of images taken during a speci c period can share the same context and content. Thus, if a set of images can be clustered into sequential atomic clusters, given an image, it is possible to automatically nd all images sharing the same content and context by rst nding the atomic cluster sharing the same content, then watershed reward and forward to nd other clusters sharing the same context. The proposed methods are evaluated on the imageCLEFlifelog 2019 dataset and compared to participants joined this event. The experimental results con rm the high productivity of the proposed method in both stable and accuracy aspects.</p>
      </abstract>
      <kwd-group>
        <kwd>lifelog</kwd>
        <kwd>content and context</kwd>
        <kwd>watershed</kwd>
        <kwd>image retrieval image rearrange</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Recently, research communities start getting used with new terminologies
"lifelogging" and "lifelog". The former represents the activity of continuously
recording people everyday experiences. The latter implies the dataset contained data
generated by lifelogging. The small sizes and a ordable prices of wearable
sensors, high bandwidth of the Internet, and exible cloud storages encourage
people to lifelogging more frequent than before. That leads to the fact that people
can have more opportunities to understand their lives thoroughly due to daily
recording data whose content conceal both cognitive and physiological
information. One of the most exciting topics when trying to get insights from lifelogs is
to understand human activities from the rst-person perspective[6][7]. Another
interesting topic is to augment human memory towards improving human
capacity to remember[10]. The former aims to understand how people act daily
towards having e ective and e cient support to improve the quali cation of
living both in social and physical activities. The latter tries to create a memory
assistant that can accurately and quickly bring a memory back to a human when
necessary.</p>
      <p>In order to encourage people to pay more attention to the topics above,
several events have been organized[8][9][2][4]. These events o er a large annotated
lifelog collected from various sensors such as physiology (e.g., heartbeat, step
counts), images (e.g., lifelog camera), location (e.g., GPS), users tags,
smartphone logs, and computer logs. A series of tasks introduced throughout these
events started attracting people leading to an increase in the number of
participants. Unfortunately, the results of proposed solutions from participants are far
from expectation. It means that lifelogging still a mystical land that needs to be
discovered more both in what kind of insights people can extract from lifelogs
and how the accuracy of these insights are.</p>
      <p>
        Along this direction, imageCLEFlifelog2019 - a session of imageCLEF2019[11]
- is organized with two tasks[3]: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) solve my puzzle: the new task introduced
this time, and (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) lifelog moment retrieval: the old one with some modi cation
to make it more di cult and enjoyable than before. In this paper, two new
methods are introduced to tackle two tasks mentioned. The basement of the proposed
methods is built by utilizing the association of contents and contexts of daily
activities. Two backbone hypotheses are built based on this basement:
1. Under the same context (e.g., a particular activity), one image should have
at least one associative image (e.g., same content, same concepts) taken from
di erent moments. Thus, a given set of images can be rearranged
chronologically by ordering their associative images whose orders are known precisely,
and
2. A sequence of images taken during a speci c period can share the same
context and content. Thus, if a set of images can be clustered into sequential
atomic clusters, given an image, it is possible to automatically nd all images
sharing the same content and context by rst nding the atomic cluster
sharing the same content, then watershed reward and forward to nd other
clusters sharing the same context.
      </p>
      <p>The paper is organized as follows: Section 2 and 3 introduce the solutions
for Task1 and Task 2, respectively. Section 4 describes the experimental results
gotten by applying two proposed methods as well as related discussions. The
last Section gives conclusions and future works.</p>
    </sec>
    <sec id="sec-2">
      <title>Task 1: Solve my puzzle</title>
      <p>In this section, we introduce the proposed method, as well as its detail
explanation, algorithms, and examples related to the Task 1.</p>
      <p>Task 1: solve my life puzzle is stated as: Given a set of lifelogging images
with associated metadata such as biometrics and location, but no timestamps,
these images are required to be rearranged in chronological order and predict the
correct day (e.g., Monday or Sunday) and part of the day (morning, afternoon,
or evening)[3].
2.1</p>
      <sec id="sec-2-1">
        <title>From Chaos to Order</title>
        <p>We build our solution based on one of the characteristics of lifelogs: activities
of daily living (ADLs)[5]. Some common contents and contexts people live in
every day can be utilized to associate unordered time images to ordered time
images. For example, one always prepares and has breakfast in a kitchen every
morning from 6 am to 8 am, except for a weekend. Hence, all record rid recorded
from 6 am to 8 am in the kitchen are shared the same context (i.e., in a kitchen,
chronological order of sequential concepts) and content (e.g., objects in a kitchen,
foods). If we can associate one-by-one each image of a set of unordered time
images riq to a subset of ordered time images rid, riq can totally be ordered by
the order of rid. This above observation brings the following useful hints:
{ If one record=(image, metadata) rq captured in the scope of ADLs, there
probably is a record rd sharing the same content and context.
{ If we can nd all associative pairs (rid; riq), we can rearrange riq by rearranging
rid utilizing metadata of rid, especially timestamps and locations.</p>
        <p>We call rq and rd a query record/image and associative record/image (of
the query record/image), respectively, if similary(rd; rq) &gt; (the prede ned
threshold). We call a pair of (rd; rq) an associative pair if it satis es the condition
similary(rd; rq) &gt; . Hence, we propose a solution as follows: Given sets of
q
unordered images Q = fri g and ordered images D = rid , we
1. Find all associative pairs (rd; rq) by using the similarity function (described
in subsection 3.2). The output of this stage is a set of associative pairs
ordering by timestamps of D, call P . Next, P is re ned to generate PT OU
to guarantee that there is only one associate pair remained with the highest
similarity score among consecutive associative pairs of the same query image.
Algorithm 1 models this task. Fig. 1, the rst and second rows, illustrates
how this task works.
2. Extract all patterns pat of associative pair to create PUT OU . The pattern is
de ned as the set of associative pairs so that the query images set is precisely
the same as the images of Q. Algorithm 2 models this task. Figure 1, the
third row 27 denotes the process of extracting patterns.
3. Find the best pattern PBEST that has the highest average similarity score
comparing to the rest. Algorithm 3 models this task. Fig. 1, the third row
denotes the process of selecting the best pattern.
4. The unordered images Q is ordered by order of associative images of PBEST .</p>
        <p>Fig. 1, the red rectangle at the left-bottom corner illustrates this task.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Similarity Functions</title>
        <p>The similarity function to measure the similarity between a query record and its
potentially associative record has two di erent versions built as follows:
1. Color histogram vector and Chi-square: The method introduced by
Adrian Rosebrock3 is utilized to build the histogram vector. As mentioned
above, the environment of images taken by lifelogging can be repeated daily,
both indoor and outdoor. Hence, the color information probably is the
primary cue to nd similar images captured under the same context.
2. Category vector and FAISS: The category vector is concerned as a deep
learning feature to overcome problems that handcraft features cannot do.
The ResNet18 of the pre-trained model PlaceCNN4 is utilized to create the
category vector. The 512-dimension vector extracted from the "avgpool"
layer of the ResNet18 is used. The reason such a vector is used is to search
images sharing the same context (spatial dimension). The FAISS (Facebook
AI Similarity Search)5 is utilized to push the speed of searching for similar
images due to its high productivity of similarity searching and clustering
dense vectors.
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Parameters De nitions</title>
        <p>Let Q = fq[k]gk=1::kQk denote the set of query images, where kQk is the total
number of images in the query set.</p>
        <p>Let D = fd[l]gl=1::kDk denote the set of images contained in the given data
set, where kDk is the total number of images in the data set.</p>
        <p>Let P = fp(iD; iQ)[m]g denote the set of associative indices that point to
related images contained in D and Q, respectively. In other words, using p[m]:iD
and p[m]:iQ we can access images d[p[m]:iD] 2 D and q[p[m]:iQ] 2 Q,
respectively.</p>
        <p>Let PT OU and PUT OU denote the sets of time-ordered-associative (TOU)
images and unique-time-ordered-associative (UTOU) images, respectively.</p>
        <p>Let PBEST = fp(iD; iQ)[k]g denote the best subset of PUT OU satisfy
{ 8m 6= n : p[m]:iQ 6= p[n]:iQ AND kPBEST k == kQk (i.e., unique)
3
https://www.pyimagesearch.com/2014/12/01/complete-guide-building-imagesearch-engine-python-opencv
4 https://github.com/CSAILVision/places365
5 https://code.fb.com/data-infrastructure/faiss-a-library-for-e
cient-similaritysearch
{ 8m : time(p[m]:iD) &lt; time(p[m + 1]:iD) (i.e., time-ordered)
{ k1 Pk</p>
        <p>i=1 (similarity(d[p[i]:iD]; q[p[i]:iQ])) ) M AX (i.e., associative)
Let , , and denote the similarity threshold, the number of days in the
data set, and the number of clusters within one day (i.e. parts of the day),
respectively.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Task 2: Lifelog Moment Retrieval</title>
      <p>In this section, we introduce the proposed method, as well as its detail
explanation, algorithms, and examples related to Task 2.</p>
      <p>Task 2: Lifelog moment retrieval is stated as[3]: Retrieve a number of speci c
prede ned activities in a lifelogger's life.
3.1</p>
      <p>The Interactive-Watershed-based for Lifelog Moment Retrieval
The proposed method is built based on the following observation: A sequence of
images taken during a speci c period can share the same context and content.
Thus, if a set of images can be clustered into sequential atomic clusters, given
an image, we can automatically nd all images sharing the same content and
context by rst nding the atomic cluster sharing the same content, then
watershed reward and forward to nd other clusters sharing the same context. The
terminology atomic cluster is understood that all images inside should share the
similarity higher than the prede ned threshold and must share the same context
(e.g., location, time).</p>
      <p>Based on the above discussion, the Algorithm 1 is built to nd a Lifelog
moment from a dataset de ned be a query events text. Following is the description
to explain the Algorithm 1.</p>
      <p>{ Stage 1 (O ine): This stage aims to cluster a dataset into the atomic
clusters. The category vector V c and attribute vector V a extracted from
images are utilized to build the similarity function. Besides, each image is
analyzed by using I2BoW function to extract its BoW contains concepts
and attributes reserved for later processing.
{ Stage 2 (Online): This stage targets to nd all clusters satis ed a given
query. Since the given query is described by text, a text-based query method
must be used for nding related images. Thus, we create Pos tag,
GoogleSearchAPI, and Filter functions to nd the best BoWs that represent the
taxonomy of the querys context and content. In order to prune the output of
these functions, we utilize the Interactive function to select the best
taxonomy. First, images queried by using BoW (i.e., seeds) are utilized for nding
all clusters, namely LMRT1, that contain these images. These clusters are
then hidden for the next step. Next, these seeds are used to query on the
rest clusters (i.e., unhidden clusters) to nd the second set of seeds. These
second set of seeds are used to nd all clusters, namely LRMT2, that contain
these seeds. The nal output is the union of C1 and C2. At this step, we
might apply Interactive function to re ne the output (e.g., select clusters
manually). Consequently, all lifelog moment satis ed the query are found.
3.2</p>
      <sec id="sec-3-1">
        <title>Functions</title>
        <p>In this subsection, the signi cant functions utilized in the proposed method are
introduced, as follows:
{ Pos tag: processes a sequence of words tokenized from a given set of
sentences, and attaches a part of speech tag (e.g., noun, verb, adjective, adverb)
to each word. The library NLTK6 is utilized to build this function.
{ featureExtractor: analyzes an image and return a pair of vectors vc
(category vector, 512 dimension) and va (attribute vector, 102 dimension). The
former is extracted from the "avgpool" layer of the RestNet18 of the
pretrained model PlaceCNN [13]. The latter is calculated by using the equation
va = WaT vc, introduced in [12], where Wa is the weight of Learnable
Transformation Matrix.
{ similarity: measures the similarity between two vectors using the cosine
similarity.
{ I2BoW: converts an image into a bag of words using the method introduced
in [1]. The detector developed by the authors return concepts, attribute, and
relation vocab. Nevertheless, only a pair of attribute and concept is used for
building I2BoW.
{ GoogleSearchAPI: enriches a given set of words by using Google Search
API 7. The output of this function is the set of words that could be probably
similar to the queried words under certain concepts.
{ Interactive: allows users to interfere with re ning the results generated by
related functions.
{ Query: nds all items of the searching dataset that are similar to queried
items. This function can adjust its similarity function depending on the type
of input data.
{ cluster: clusters images into sequential clusters so that all images of one
cluster must share the highest similarity comparing to its neighbors. All
images are sorted by time before being clustered. Algorithm 2 describes this
function in detail. Figure 3.1 illustrates how this function works.
3.3</p>
      </sec>
      <sec id="sec-3-2">
        <title>Parameter De nitions</title>
        <p>Let I = fIigi=1::N denote the set of given images (e.g., dataset).</p>
        <p>Let F = f(V ci; V ai)gi=1::N denote the set of feature vectors extracted from
Let C = fCkg denote a set of atomic clusters.
6 https://www.nltk.org/book/ch05.html
7 https://github.com/abenassi/Google-Search-API</p>
        <p>Let fBoWIi gi=1::N denote the set of BoWs; each of them is a BoW built by
using the I2BoW function.</p>
        <p>Let fBoWQT g ; BoWQNToun ; nBoWQATugo denote the Bag of Words extracted
from the query, the NOUN part of BoWQT , and the augmented part of BoWQT ,
respectively.</p>
        <p>Let Seedij and LM RT k denote a set of seeds and lifelog moments,
respectively.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experimental Results</title>
      <p>In this section, datasets and evaluation metrics used to evaluate the proposed
solution are introduced. Besides, the comparison of our results with others from
di erent participants is also discussed.
4.1</p>
      <sec id="sec-4-1">
        <title>Datasets and Evaluation Metrics</title>
        <p>We use the dataset released by the imageCLEFlifelog 2019 challenge[3]. The
challenge introduces an entirely new rich multimodal dataset which consists of
29 days of data from one lifeloggers. The dataset contains images (1,500-2,500
per day from wearable cameras), visual concepts (automatically extracted visual
concepts with varying rates of accuracy), semantic content (semantic locations,
semantic activities) based on sensor readings (via the Moves App) on mobile
devices, biometrics information (heart rate, galvanic skin response, calorie burn,
steps, continual blood glucose, etc.), music listening history, computer usage
(frequency of typed words via the keyboard and information consumed on the
computer via ASR of on-screen activity on a per-minute basis).</p>
        <p>Generally, the organizers of the challenge do not release the ground truth for
the test set. Instead, participants must send their arrangement to the organizers
and get back their evaluation.</p>
        <p>Task 1: Solve my puzzle Two training and testing query sets are given. Each
of them has a total of 10 sets of images. Each set has 25 unordered images need
to be rearranged. The training set has its ground truth to let participants can
evaluate their solution.</p>
        <p>For evaluating, the Kendall rank correlation coe cient is utilized to evaluate
the similarity between the arrangement and the ground truth. Then, the mean
of the accuracy of the prediction of which part of the day the image belongs to
and the Kendall's Tau coe cient is calculated to have the nal score.
Task 2: Lifelog Moment Retrieval The training and testing set stored with
the JSON format, have ten queries whose titles are listed in Tables 3 and 4,
respectively. Besides, more descriptions and constraints are also listed to guide
participants on how to understand queries precisely.</p>
        <p>The evaluation metrics are de ned by imageCLEFlifelog 2019 as follows:
{ Cluster Recall at X (CR@X): a metric that assesses how many di erent
clusters from the ground truth are represented among the top X results,
{ Precision at X (P@X): measures the number of relevant photos among the
top X results,
{ F1-measure at X (F1@X): the harmonic mean of the previous two.</p>
        <sec id="sec-4-1-1">
          <title>X is chosen as 10 for evaluation and comparison.</title>
          <p>4.2</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>Evaluation and Comparison</title>
        <p>
          Task 1: Solve my puzzle Although eleven runs were submitted to evaluate,
only the best ve runs are introduced and discussed in this paper: (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) only color
histogram vector, (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) color histogram vector and the constraint of timezone (e.g.,
only places in Dublin, Ireland) (
          <xref ref-type="bibr" rid="ref3">3</xref>
          ) object search (i.e., using visual concepts of
metadata), and category vector plus the constraint of timezone, (
          <xref ref-type="bibr" rid="ref4">4</xref>
          ) timezone
is strictly narrowed into only working days and inside Dublin, Ireland, and (
          <xref ref-type="bibr" rid="ref5">5</xref>
          )
only category vector, and the constraint of timezone (i.e., only places in Dublin,
Ireland).
        </p>
        <p>The reason we reluctantly integrate places information to the similarity
function due to the uncertainty of metadata given by the organizer. Records that
have location data (e.g., GPS, place names) occupy only 14:90%. Besides, most
of the location data give wrong information that can lead to similarity scores
coverage in a wrong optimal peak.</p>
        <p>As described in Table 1, the similarity function using only color histogram
gave the worst results compared to those who use the category vector. Since
users are living in Dublin, Ireland, the same daily activities should share the
same context and content. Nevertheless, if users go abroad or other cities, this
context and content will be changed totally. Hence, the timezone could be seen
as the majority factor that can improve the nal accuracy since it constraints the
context of searching scope. As mentioned above, the metadata does not usually
contain useful data that can be leveraged the accuracy of searching. Thus, when
utilizing object concepts as a factor for measuring the similarity, it can pull down
the accuracy comparing to not using them at all. Distinguish between working
day and weekend activities depending totally on types of activities. In this case,
the number of activities that happen every day is more signi cant than those
that happen only on working days. That explains the higher accuracy of ignoring
the working day/weekend factor. In general, the similarity funtion integrated
timezone (i.e., daily activities context) and category vector (i.e., rich content)
gives the best accuracy compared to others. It should be noted that the training
query sets are built by picking up exactly images from dataset while the testing
query sets not appear in the dataset at all. That explains why the accuracy of
the proposed method is perfect when running on training query sets.</p>
        <p>
          Four participants submitted their outputs for evaluating (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) DAMILAB, (
          <xref ref-type="bibr" rid="ref2">2</xref>
          )
HCMUS (Vietnam National University in HCM city), (
          <xref ref-type="bibr" rid="ref3">3</xref>
          ) BIDAL (ourselves),
and (
          <xref ref-type="bibr" rid="ref4">4</xref>
          ) DCU (Dublin City University, Ireland). Table 2 denotes the di erence
in primary scores between results generated by methods proposed by these
participants and the proposed method. Although the nal score of the proposed
method is lower than of HCMUS, the proposed method might be more stable
than others. In other words, the variance of accuracy scores when rearranging ten
di erent queries of the proposed method is less than others. Hence, the proposed
method can cope with under tting, over tting and bias problems.
Task 2: Lifelog Moment Retrieval Although three runs were submitted to
evaluate, two best runs are chosen to introduce and discuss in this paper: (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) run
1: interactive mode. In this run, interactive functions are activated to let users
interfere and manually get rid of those images that do not relevant to a query,
(
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) run 2: automatic mode. In this run, a program runs without any interfere
from users.
        </p>
        <p>Table 3 and 4 show the results running on the training and testing sets,
respectively. Opposite to the results of Task 1, the results of Task 2 do not
have much di erence between training and testing stages. That could lead to
the conclusion that the proposed method probably is stable and robust enough
to cope with di erent types of queries.
In all cases, the P@10 results are very high. It proves that the approach
used for querying seeds of the proposed method is useful and precise. Moreover,
the watershed-based stage after nding seeds can help not only to decrease the
complexity of querying related images but also to increase the accuracy of event
boundaries. Unfortunately, the CR@10 results are less accuracy comparing to
P@10. The reason could come from merging clusters. Currently, clusters gained
after running watershed are not merged and rearranged. That could lead to low
accuracy when evaluating CR@X. This issue is investigated thoroughly in the
future.</p>
        <p>Nevertheless, misunderstanding context and content of queries sometimes
lead to the worst results. For example, both runs failed in query 2 "driving
home" and query 9 "wearing a red plaid shirt." The former was understood as
"driving from o ce to home regardless of how many times stop at in-middle
places," and the latter was distracted by synonym words of "plaid shirt" when
leveraging GoogeSearchAPI to augmented the BoW. The rst case should be
understood that "driving home from the last stop before home," and the
second case should focus on only "plaid shirt" not "sweater" nor "fannel." After
xing these mistakes, both runs have higher scores on query 2 and query 9, as
described in Table 4. The second rows of event 2, 9, and the average score show
the results after correcting the mentioned misunderstanding. Hence, building a
exible mechanism to automatically build a useful taxonomy from a given query
to avoid these mistakes is built in the future.</p>
        <p>
          Nine teams participated to Task 2 included (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) HCMUS, (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) ZJUTCVR,
(
          <xref ref-type="bibr" rid="ref3">3</xref>
          ) BIDAL (ourselves), (
          <xref ref-type="bibr" rid="ref4">4</xref>
          ) DCU, (
          <xref ref-type="bibr" rid="ref5">5</xref>
          ) ATS, (
          <xref ref-type="bibr" rid="ref6">6</xref>
          ) REGIMLAB, (
          <xref ref-type="bibr" rid="ref7">7</xref>
          ) TUCMI, (
          <xref ref-type="bibr" rid="ref8">8</xref>
          )
UAPT, and (
          <xref ref-type="bibr" rid="ref9">9</xref>
          ) UPB. The detail information of these teams could be referred
to in [3]. Table 5 denotes the comparison among these teams. We are ranked
in the third position. In general, the proposed method can nd all events with
acceptance accuracy (i.e., no event with zero F1@10 scores comparing to others
those have at least one event with zero F1@10 scores). That con rms again the
stability and anti-bias of the proposed method.
5
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>
        We introduce two methods for tackling two tasks challenged by
imageCLEFlifelog 2019: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) solve my puzzle, and (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) lifelog moment retrieval. Although
each method is built di erently, the backbone of these methods is the same:
considering the association of content and context of daily activities. The rst
method is built for augmenting human memory when the chaos of memory
snapshots must be rearranged chronologically to give a whole picture of a users life
moment. The second method is constructed to visualize peoples memories from
their explanation: start from their pinpoints of memory and watershed to get
all image clusters around those pinpoints. The proposed method is thoroughly
evaluated by the benchmark dataset provided by the imageCLEFlifelog 2019,
and compared with other solutions coming from di erent teams. The nal
results show that the proposed method is developed in the right direction even
though it needs more improvement to reach the expected targets. Many issues
are raised during the experimental results and need to be investigated further
for better results. In general, two issues should be investigated more in the
future: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) Similarity functions and watershed boundaries, and (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) A taxonomy
generated from queries.
Algorithm 1 Create a set of time-ordered-associative (TOU) images
Input: Q = fq[k]g, D = fd[l]g, , (#days), (#clusters)
Output: PT OU [i][j]
1: PT OU [i][j] ( ;
2: for i=1.. do
3: for j=1.. do
4: Filter D so that only images taken on day i within cluster j is enable
5: Establish P = fp(iD; iQ)[m]g so that 8m &lt; n : time(d[p[m]:iD]) &lt;
time(d[p[n]:iD])^ similarity(d[p[m]:iD]; q[p[m]:iQ]) &gt; and 8m 6= n :
p(iD; iQ)[m] 6= p(iD; iQ)[n]
6: repeat
7: 8n kP k
8: if (p[n].iQ == p[n+1].iQ) then
9: if (similarity(d[p[n]:iD]; q[p[n]:iQ]) &gt;
      </p>
      <p>similarity(d[p[n + 1]:iD]; q[p[n]:iQ])) then
10: Delete p[n + 1] else Delete p[n]
11: end if
12: Rearrange the index of P
13: end if
14: until cannot delete any item of p
15: PT OU [i][j] ( P (P must satisfy 8n : p[n]:iQ 6= p[n + 1]:iQ)
16: end for
17: end for
18: return PT OU [i][j]
Input: Q = fq[k]g, D = fd[l]g, PT OU [i][j],
Output: PUT OU [i][j][z]
1: slidingW indow = kQk
2: for i = 1:: do
3: z = 1, N = kPT OU [i][j]k
4: for j = 1:: do
5: for n = 1::(N
6:
7:</p>
      <p>slidingW indow) do
substr ( PT OU [i][j][m]m=n::(n+slidingW indow)
if (8m 6= n : substr[m]:iQ 6= substr[n]:iQ and ksubstrk
slidingW indow) then
8: PUT OU [i][j][z] ( substr, z ( z + 1
9: end if
10: end for
11: end for
12: end for
13: return PUT OU [i][j][z]
==
Algorithm 2 Create a set of unique-time-ordered-associative (UTOU) images
(#days), (#clusters)</p>
      <sec id="sec-5-1">
        <title>Algorithm 3 Find the BEST set of UTOU images</title>
        <p>(#days), (#clusters)
Input: Q = fq[k]g, D = fd[l]g, PUT OU [i][j][z],
Output: PBEST
1: for j = 1:: do
2: CoverSet ( ;
3: for i = 1:: do
4: CoverSet[i] ( CoverSet [ PST OAU [i][j]
5: end for
6: Find the most common pattern pat from all patterns contained in CoverSet
7: Delete all subsets of CoverSet that do not match pat
8: For each remained subset of CoverSet, calculate the average similarity with Q
9: Find the subset CoverSetlargest that has the largest average similarity with Q
10: PBEST CLUST ER[j] ( CoverSetlargest
11: end for
12: Find the subset PBEST of PBEST CLUST ER[j] that has the largest average similarity
with Q
13: return PBEST
Input: QT , BoWConcepts, fIigi=1::N
Output: LM RT</p>
        <p>fOFFLINEg
1: fBoWIi gi=1::N ( ;
2: f(V ci; V ai)gi=1::N ( ;
3: 8i 2 [1::N ]; BoWIi ( I2BoW (Ii)
4: 8i 2 [1::N ]; (V ci; V ai) ( f eatureExtractor(Ii)
5: fCmg ( cluster(fIigi=1::N ) using f(V ci; V ai)g
6: fBOoNWLQNITNouEng( P os tag(QT ):N oun
7: BoWQATug ( GoogleSearchAP I(BoWQNToun)
8: BoWQNToun ( F ilter(BoWQATug \ BoWConcepts)
9: BoWQT ( Interactive(BoWQNToun; P os tag(QT ))
10: Seedj1 ( Interactive(Query(BoWQT ;</p>
        <p>fBoWIi gi=1::N ))
11: LM RT 1 ( Ckj8j 2 Seedj1 ; Seedj1 2 fCkg
12: fClrem LM RT 1
13: Seedj2g ((fCQmuegry( Seedj1 ; fClremg)
14: LM RT 2 ( Clj8j 2 Seedj2 ; Seedj2 2 fClremg
15: LM RT ( LM RT 1 [ LM RT 2
16: LM RT ( Intearactive(LM RT )
17: return LM RT</p>
        <sec id="sec-5-1-1">
          <title>Algorithm 5 cluster</title>
          <p>Input: I = fIigi=1::N ; F = f(V ci; V ai)gi=1::N ; a; b
Output: fCkg
1: k 0
2: C = fCkg ;
3: Itemp ( SORTbytime(I; F )
4: repeat
5: va Itemp:F:V a[0]
6: vc Itemp:F:V c[0]
7: ICt[e0m]p( [Itemp:I[0]
8: ( Itemp Itemp[0]
9: for i=1.. Itemp
10:</p>
          <p>do
if (similarity(va; Itemp:F:V a[i]) &gt; a and similarity(vc; Itemp:F:V c[i]) &gt; c
then
11:
12:
13: else
14: k
15: Break
16: end if
17: end for
18: until Itemp
19: return C</p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Anderson</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>He</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buehler</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Teney</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , Johnson,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Gould</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Zhang</surname>
          </string-name>
          , L.:
          <article-title>Bottom-up and top-down attention for image captioning and visual question answering</article-title>
          .
          <source>In: CVPR</source>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Dang-Nguyen</surname>
            ,
            <given-names>D.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Piras</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Riegler</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boato</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gurrin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Overview of imagecle ifelog 2017: Lifelog retrieval and summarization</article-title>
          .
          <source>In: CLEF (Working Notes)</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Dang-Nguyen</surname>
            ,
            <given-names>D.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Piras</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Riegler</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tran</surname>
            ,
            <given-names>M.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lux</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Le</surname>
            ,
            <given-names>T.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ninh</surname>
            ,
            <given-names>V.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gurrin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          : Overview of ImageCLEFlifelog 2019:
          <article-title>Solve my life puzzle and Lifelog Moment Retrieval</article-title>
          .
          <source>In: CLEF2019 Working Notes. CEUR Workshop Proceedings</source>
          , CEUR-WS.org &lt;http://ceur-ws.
          <source>org&gt;</source>
          , Lugano,
          <source>Switzerland (September</source>
          <volume>09</volume>
          -12
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Dang-Nguyen</surname>
            ,
            <given-names>D.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Piras</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Riegler</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lux</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gurrin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Overview of imagecle ifelog 2018: daily living understanding and lifelog moment retrieval</article-title>
          .
          <source>In: CLEF2018 Working Notes (CEUR Workshop Proceedings)</source>
          .
          <article-title>CEUR-WS</article-title>
          . org&lt; http://ceur-ws.
          <source>org&gt;</source>
          , Avignon, France (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Dao</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kasem</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nazmudeen</surname>
            ,
            <given-names>M.S.H.</given-names>
          </string-name>
          :
          <article-title>Leveraging content and context in understanding activities of daily living</article-title>
          .
          <source>In: Working Notes of CLEF 2018 - Conference and Labs of the Evaluation Forum</source>
          , Avignon, France,
          <source>September 10-14</source>
          ,
          <year>2018</year>
          . (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Dao</surname>
            ,
            <given-names>M.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nguyen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tien</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Riegler</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gurrin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Smart lifelogging: recognizing human activities using phasor (</article-title>
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Dimiccoli</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cartas</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Radeva</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Activity recognition from visual lifelogs: State of the art and future challenges</article-title>
          .
          <source>In: Multimodal Behavior Analysis in the Wild</source>
          , pp.
          <volume>121</volume>
          {
          <fpage>134</fpage>
          .
          <string-name>
            <surname>Elsevier</surname>
          </string-name>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Gurrin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joho</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hopfgartner</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Albatal</surname>
          </string-name>
          , R.:
          <article-title>Overview of ntcir-12 lifelog task</article-title>
          . In: Kando,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Kishida</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Kato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.P.</given-names>
            ,
            <surname>Yamamoto</surname>
          </string-name>
          , S. (eds.)
          <source>Proceedings of the 12th NTCIR Conference on Evaluation of Information Access Technologies</source>
          . pp.
          <volume>354</volume>
          {
          <issue>360</issue>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Gurrin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joho</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hopfgartner</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Albatal</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nguyen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tien</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Overview of ntcir-13 lifelog-2 task</article-title>
          .
          <source>NTCIR</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Harvey</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Langheinrich</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ward</surname>
          </string-name>
          , G.:
          <article-title>Remembering through lifelogging</article-title>
          .
          <source>Pervasive Mob. Comput. 27(C)</source>
          ,
          <volume>14</volume>
          { 26 (Apr
          <year>2016</year>
          ). https://doi.org/10.1016/j.pmcj.
          <year>2015</year>
          .
          <volume>12</volume>
          .002, http://dx.doi.org/10.1016/j.pmcj.
          <year>2015</year>
          .
          <volume>12</volume>
          .002
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Ionescu</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , Muller, H.,
          <string-name>
            <surname>Peteri</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cid</surname>
            ,
            <given-names>Y.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liauchuk</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kovalev</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klimuk</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tarasau</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Abacha</surname>
            ,
            <given-names>A.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hasan</surname>
            ,
            <given-names>S.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Datla</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Demner-Fushman</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dang-Nguyen</surname>
            ,
            <given-names>D.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Piras</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Riegler</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tran</surname>
            ,
            <given-names>M.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lux</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gurrin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pelka</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Friedrich</surname>
            ,
            <given-names>C.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>de Herrera</surname>
            ,
            <given-names>A.G.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garcia</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kavallieratou</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>del Blanco</surname>
            ,
            <given-names>C.R.</given-names>
          </string-name>
          , Rodr guez, C.C.,
          <string-name>
            <surname>Vasillopoulos</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Karampidis</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chamberlain</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Campello</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>ImageCLEF 2019: Multimedia retrieval in medicine, lifelogging, security and nature</article-title>
          . In:
          <article-title>Experimental IR Meets Multilinguality, Multimodality, and Interaction</article-title>
          .
          <source>Proceedings of the 10th International Conference of the CLEF Association (CLEF</source>
          <year>2019</year>
          ),
          <source>LNCS Lecture Notes in Computer Science</source>
          , Springer, Lugano,
          <source>Switzerland (September 9-12</source>
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Patterson</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Su</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hays</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>The sun attribute database: Beyond categories for deeper scene understanding</article-title>
          .
          <source>Int. J. Comput. Vision</source>
          <volume>108</volume>
          (
          <issue>1-2</issue>
          ),
          <volume>59</volume>
          {81 (May
          <year>2014</year>
          ). https://doi.org/10.1007/s11263-013-0695-z, http://dx.doi.org/10.1007/s11263-013-0695-z
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lapedriza</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khosla</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oliva</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Torralba</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Places: A 10 million image database for scene recognition</article-title>
          .
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>