<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>DCU team at The 2019 Insight for Wellbeing Task: Multimodal personal health lifelog data analysis</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Tu-Khiem Le</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Van-Tu Ninh</string-name>
          <email>tu.ninhvan@adaptcentre.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Liting Zhou</string-name>
          <email>zhou.liting2@mail.dcu.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Duc-Tien Dang-Nguyen</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Cathal Gurrin</string-name>
          <email>cathal.gurrin@dcu.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dublin City University</institution>
          ,
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Bergen</institution>
          ,
          <country country="NO">Norway</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <fpage>27</fpage>
      <lpage>29</lpage>
      <abstract>
        <p>In this paper, the authors described their proposed method in analyzing lifelog data in association with the environment. Tackling the problem of incomplete data, we proposed a replacement method using linear regression method which results in a normalized L2 distance score of 0.0153. Meanwhile, the authors solved the personal air quality subtask by inferring from lifeloggers' PM2.5 data, which achieves 1.0 in the arithmetic mean of absolute distance score between the predictions and the actual classes.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        Along with the development of engineering and technology, more
and more personal devices such as smartphones, video cameras
and wearable sensors have come to life which provide people the
ability to easily capture every aspect of their life. On top of that,
the term lifelogging is defined to be the process of recording a
detailed trace of life passively[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], which generates a large collection
of multimedia data. The huge amount of lifelog data leads to the
need to quickly retrieve and extract particular insight based on
the associations between data. In the MediaEval 2019 Insight for
Wellbeing Challenge, they defined a new approach to lifelog data
in relation with the environment. This is potential in analyzing the
efect of general pollution on the living quality on individual scale.
Beside the information recorded from the weather and air pollution
stations, lifelog data could add in the true nature of particular
regions where the stations are not set up.
      </p>
      <p>
        The organisers generated a novel dataset called SEPHLA [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]
which is collected by multiple lifeloggers who walk on several
selected routes in the city and record data through wearable sensors
and smartphones. The lifelog images, biometrics, weather, urban
perception tags, emotional tags and air pollution data are provided
within the dataset. To better understand the data and gain insights
for personal wellbeing, the organizers defines two subtasks:
Segment Replacement and Personal Air Quality prediction. In the first
subtask, the participants are asked to investigate the associations
among data and develop a solution to reconstruct the segments of
data which are removed by the organisers. Meanwhile, The second
subtask aims to estimate people wellbeing by predicting the AQI
(Air Quality Index) on particular positions in a specified time. More
details about the this challenge can be found in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>RELATED WORK</title>
      <p>
        In recent years, lifelogging has gained more and more attentions
and many research works have been proposed to provide better
understanding of personal digital collections. To support, many
international benchmarking eforts have been made and various
challenges on lifelogging data were hosted, the most recent of which
is NTCIR-14 Lifelog-3 Task [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], LSC 2018 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], and
ImageCLEF2019lifelog [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. While the purpose of these challenges is to mainly focus
on developing a solution to retrieve relevant moments based on a
set of given queries, each challenge has diferent subtasks to
further explore this multimodal data. In the Lifelog Search Challenge
(LSC), not only are the participants required to build an interactive
retrieval system, but they also need to compete with each other in
the competition with real-time on-screen query.
      </p>
      <p>The datasets, which were utilised in these challenges, are
collected by many lifeloggers who wear a passive-captured wearable
camera and other tracking sensors. Each lifelogger normally
generates around 1250 - 4500 images per day in association with other
biometrics (e.g. heart rate, calorie), locations (GPS), physical
movements and music. They share nearly the same structure with the
lifelog data in the MediaEval 2019 Insight for Wellbeing Challenge.
However, this challenge also considers additional information from
the environment, which makes the insight more general and enables
us to obtain an overview of the wellbeing among individuals.
3</p>
    </sec>
    <sec id="sec-3">
      <title>APPROACH</title>
      <p>From the dataset, we are provided air quality data gathered by
the stations and lifeloggers’ sensors. These are extremely useful
information to reconstruct missing segments of data and predict
air quality index for specific areas. Besides, we also got a collection
of image data recorded by the lifeloggers with corresponding
visual concepts extracted from the neural network, along with the
information on the checkpoints where they are asked to take
pictures. However, the images which are actively taken might vary
from the lifeloggers’ preferences. Therefore, it’s hard to capture and
generalize the context across individuals. Based on the observation
we gained, we proposed the solutions to both sub-tasks which are
described in the following subsections.
3.1</p>
    </sec>
    <sec id="sec-4">
      <title>Segment Replacement</title>
      <p>In this subtask, the sequence of missing PM2.5 data is specified
in each query with a starting and ending time. As the lifeloggers
walked in groups, the data from others could help regenerate the
missing segments. The data from the stations, however, is not quite
reliable since their distance is too far from the routes and they
might contribute noises to the result. Therefore, considering the
NO2, O3, temperature, humidity and heartbeat data from people
who share the same route with the targeted lifelogger, we build
a simple linear regression model to predict the removed PM2.5
data. Specifically, let x be the 5-dimensional L2-normed feature
vectors composed of five components mentioned above, and y be
the targeted PM2.5 value that needs to be predicted. We construct
a linear regression model y = wT x + b and apply gradient descent
to find the best parameters w and b to minimize root-mean-square
error, which aims to minimize the gap between model predictions
and ground-truth of train data. Then the trained model is used to
generate the missing PM2.5 of the targeted person in that group.
As NO2, and O3 values are not almost zero for most of the times,
temperature, humidity, and heart-beat are the main factors that
contribute most to our predictions.</p>
    </sec>
    <sec id="sec-5">
      <title>3.2 Personal Air Quality</title>
      <p>To obtain AQI for each day, we would need to first gather the air
quality data. From the checkpoints of each route, we could obtain a
list of GPS along the route. Then, we extracted all air quality data
where lifeloggers’ GPS is closed to the checkpoints. The distance
between two GPSs is calculated using the Haversine formula.</p>
      <p>As we observed from the air quality data of each route, NO2
and O3 values are mostly zeros while PM2.5 values have some
lfuctuations. Therefore, we choose PM2.5 to predict the ultimate
Air Quality Index (AQI). At first, we refine the data to get the right
PM2.5 data for each route by calculate the distance between the
route’s GPS and collectors’ current GPS. After this step, we obtain
data for 27 routes on 7 days from diferent groups of collectors.
For each data on a day collected by a user, we compute its average
PM2.5. Therefore, we receive many average PM2.5 values from
many collectors in one day. We consider the maximum value of
these average PM2.5 values as the criteria to evaluate AQI for that
route on that day. Then, we average the AQI value of 7 days and
re-evaluate again to infer the AQI level of the route.
4</p>
    </sec>
    <sec id="sec-6">
      <title>RESULTS AND ANALYSIS</title>
      <p>It can be seen from the table 1 that our team (DCU) manages to
achieve the 3r d highest score of approximately 0.0153 among the
best submission list in the Segment Replacement sub-task. It means
that our approach manages to generate relatively good prediction
with low error. However, there are other solutions could provide
more precise result with significantly low error.</p>
      <p>Meanwhile, in the Personal Air Quality sub-task, our approach
got the arithmetic mean absolute L1 distance score of 1.0. This
means that our approach to handle the data for this task is not good
and the operation that we apply to process PM2.5 data to infer AQI
level is not correct. Since the data recorded from the lifeloggers
walking through the route is not totally correct (as the values are
almost zeros for all) and the collected data is not enough (less than
24 hours during seven non-consecutive days), we can hardly infer
the right AQI level for the route.</p>
      <p>As we do not exploit all the provided materials such as the data
recorded from the stations, images and related metadata, we might
miss some important features that could be used to improve our
predictions. Moreover, as we rely on the users’ recorded data along
the route that they pass through, the recorded values such as PM2.5,
NO2, O3 are not reliable as the most of their values are zeros. These
are the main factors that afect our results in both sub-tasks. In order
to improve it in future work, we might need to consider additional
data on the internet, which is recorded from nearby stations, to
provide the missing PM2.5 values during the days to generate the
correct estimation of AQI score.</p>
    </sec>
    <sec id="sec-7">
      <title>ACKNOWLEDGMENTS</title>
      <p>This publication has emanated from research supported in party
by research grants from Irish Research Council (IRC) under Grant
Number GOIPG/2016/741 and Science Foundation Ireland under
grant numbers SFI/12/RC/2289 and 13/RC/2106.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <source>[1] 2018. LSC '18: Proceedings of the 2018 ACM Workshop on The Lifelog Search Challenge. ACM</source>
          , New York, NY, USA.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Duc-Tien</surname>
            Dang-Nguyen,
            <given-names>Luca</given-names>
          </string-name>
          <string-name>
            <surname>Piras</surname>
          </string-name>
          , Michael Riegler,
          <string-name>
            <surname>Minh-Triet</surname>
            <given-names>Tran</given-names>
          </string-name>
          , Liting Zhou, Mathias Lux,
          <string-name>
            <surname>Tu-Khiem Le</surname>
          </string-name>
          ,
          <string-name>
            <surname>Van-Tu Ninh</surname>
            , and
            <given-names>Cathal</given-names>
          </string-name>
          <string-name>
            <surname>Gurrin</surname>
          </string-name>
          .
          <year>2019</year>
          . Overview of ImageCLEFlifelog 2019:
          <article-title>Solve my life puzzle and Lifelog Moment Retrieval</article-title>
          .
          <source>In CLEF2019 Working Notes (CEUR Workshop Proceedings)</source>
          .
          <article-title>CEUR-WS</article-title>
          .org &lt;http://ceur-ws.
          <source>org&gt;</source>
          , Lugano, Switzerland.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Cathal</given-names>
            <surname>Gurrin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Joho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Frank</given-names>
            <surname>Hopfgartner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <surname>Van-Tu</surname>
            <given-names>Ninh</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tu-Khiem</surname>
            <given-names>Le</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Rami</given-names>
            <surname>Albatal</surname>
          </string-name>
          , D.-T
          <string-name>
            <surname>Dang-Nguyen</surname>
            , and
            <given-names>Graham</given-names>
          </string-name>
          <string-name>
            <surname>Healy</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Overview of the NTCIR-14 Lifelog-3 task.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Cathal</given-names>
            <surname>Gurrin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Alan F.</given-names>
            <surname>Smeaton</surname>
          </string-name>
          , and
          <string-name>
            <surname>Aiden</surname>
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Doherty</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>LifeLogging: Personal Big Data</article-title>
          .
          <source>Foundations and Trends® in Information Retrieval 8</source>
          ,
          <issue>1</issue>
          (
          <year>2014</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>125</lpage>
          . https://doi.org/10.1561/1500000033
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Tomohiro</given-names>
            <surname>Sato Koji Zettsu Duc-Tien Dang-Nguyen Cathal Gurrin Ngoc-Thanh Nguyen Minh-Son</surname>
          </string-name>
          <string-name>
            <surname>Dao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Peijiang</given-names>
            <surname>Zhao</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Overview of MediaEval 2019: Insights for Wellbeing Task: Multimodal Personal Health Lifelog Data Analysis</article-title>
          .
          <source>In MediaEval2019 Working Notes (CEUR Workshop Proceedings)</source>
          .
          <article-title>CEUR-WS</article-title>
          .org &lt;http://ceur-ws.
          <source>org&gt;</source>
          ,
          <string-name>
            <surname>Sophia</surname>
            <given-names>Antipolis</given-names>
          </string-name>
          , France.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Tomohiro</given-names>
            <surname>Sato</surname>
          </string-name>
          , Minh Dao, Kota Kuribayashi, and
          <string-name>
            <given-names>Koji</given-names>
            <surname>Zettsu</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>SEPHLA: Challenges and Opportunities within Environment-Personal Health Archives</article-title>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>