<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Automatic Text Searching For Personal Photos</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Neil O'Hare</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hyowon Lee</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Saman Cooray</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Cathal Gurrin</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gareth J.F. Jones</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jovanka Malobabic</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Noel E. O'Connor</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alan F. Smeaton</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bartlomiej Uscilowski</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>All authors are members of the Centre for Digital Video Processing, Dublin City University. Alan F. Smeaton and Noel E. O'Connor are members of the Adaptive Information Cluster, Dublin City University</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>- This demonstration presents the MediAssist prototype system for organisation of personal digital photo collections based on contextual information, such as time and location of image capture, and content-based analysis, such as face detection and recognition. This metadata is used directly for identification of photos which match specified attributes, and also to create text surrogates for photos, allowing for text-based queries of photo collections without relying on manual annotation. MediAssist illustrates our research into digital photo management, showing how a combination of automatically extracted context and content-based information, together with user annotation and traditional text indexing techniques, facilitates efficient searching of personal photo collections.</p>
      </abstract>
      <kwd-group>
        <kwd>Personal Photo Management</kwd>
        <kwd>Text Search</kwd>
        <kwd>Context</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        In recent years digital photography has become increasingly
popular, resulting in the accumulation of large numbers of
personal digital photos. The MediAssist project [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] at the
Centre for Digital Video Processing (CDVP) addresses this
situation by developing tools for the efficient searching of
photo archives. The system uses both automatically generated
contextual metadata (eg. time, location) and content-based
analysis tools (eg. face detection and recognition).
Semiautomatic annotation allows the user to interactively improve
the automatically generated annotations. Retrieval tools allow
for complex query formulation, in addition to the facility to
create simple text queries, based on these features. In previous
work using context for photo management, Davis et al [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]
utilised context to recommend recipients for sharing photos
taken with a context-aware phone, although their system does
not support retrieval. Naaman et al [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] use context-based
features for photo management, but they do not use
contentbased analysis tools, or facilitate semi-automatic annotation
or text-based searches. There is also a huge body of work
on content-based image retrieval [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], but it has been shown
that users to not find this facility useful for personal photo
management [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>II. CONTENT AND CONTEXT-AWARE PHOTO</title>
      <p>ORGANISATION</p>
      <p>
        The MediAssist photo archive contains over 17,000
location-stamped photos taken with a number of different
camera models, including camera phones. Over 11,000 of
these have been manually annotated for a number of concepts,
including buildings, indoor/outdoor and the presence and
Fig. 1. The MediAssist Photo Management System
identity of faces. This manually annotated dataset serves as a
ground truth for the evaluation of content-based analysis tools,
and also can be used to bootstrap semi-automatic tools (which
depend on a certain level of user annotation). All photos are
indexed using both context and content-based analysis. Time
and location of photo capture are used to derive additional
contextual information such as daylight status, weather and
indoor/outdoor classification [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. A face detection system is
used to detect the presence of frontal-view faces [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Other
content-based tools used include body patch (the area under
the face, modelling the clothes worn by the individual)
feature extraction [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], face recognition using ICA (Independent
Component Analysis) and building detection based on the
distribution of edges in the image [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. All of this information
can prove very useful for searching photo collections.
      </p>
    </sec>
    <sec id="sec-3">
      <title>III. THE MEDIASSIST WEB DEMONSTRATOR SYSTEM</title>
      <p>
        The MediAssist Web-based desktop interface allows users
to search through their personal photo collections using the
contextual and content-based features described above. The
MediAssist system interface is shown in Fig. 1. Our earlier
version of the MediAssist prototype supported filter-based
searching using the photo metadata features [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. The new
version presented here has been extended to include free-text
ranked information retrieval functionality.
      </p>
      <sec id="sec-3-1">
        <title>A. Filter-Based Search</title>
        <p>The system presents search options enabling a user to enter
details of desired locations, times, and advanced options such
as people present, weather, light status, indoor/outdoor and
building/non-building. Semi-automatic person identification
relies on a combination of automatic methods and manual
annotation as described below. Time filters enable powerful
timebased queries, for example all photos taken in the evening, at
the weekend, during the summer or within certain date ranges.</p>
      </sec>
      <sec id="sec-3-2">
        <title>B. Text Search Interface</title>
        <p>
          For text-based search the automatic context and
contentbased features are mined to construct text surrogates for all
photos, creating a textual equivalent of each feature (e.g. if the
date is October 21th 2006, the text ‘october autumn saturday
weekend 21 twenty-first 2006’ would form a surrogate textual
description). So an image might have the text ‘dublin ireland
september weekend afternoon person alan’ associated with it,
representing the features location, time, face detection and
person annotation. We index the text document associated with
an image using a conventional text search engine based on
the standard BM25 information retrieval model [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. We also
create text surrogates for ‘events’ (see below) to allow for
textbased searching of events in the ‘Event List’ view described
below. The system presents a text search box to allow for
the quick and easy formulation of text queries based on the
content and context features described above. We will conduct
an evaluation of this search interface in future work.
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>C. Collection Browsing</title>
        <p>
          Four different views are available to present the results of
searches. The default view, Event List, organises the filtered
photos into ‘events’ in which the photos are grouped together
based on time proximity, by detecting large temporal gaps
between consecutively captured photos, similar to [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. Each
event is summarized by a label (location and date/time) and
five representative thumbnail photos selected based on the
query. Event Detail is composed of the full set of photos in
an event, automatically organized into sub-events. Individual
Photo List is an optional view where the thumbnail size photos
are presented without any particular event grouping, but sorted
by date/time. Photo Detail is an enlarged single photo view
presented when the user selects one of the thumbnail size
photos in any of the above views. In all of the above presentation
options, each photo is presented with its associated automatic
annotation information.
        </p>
      </sec>
      <sec id="sec-3-4">
        <title>D. Semi-Automatic Annotation</title>
        <p>
          MediAssist allows users to manually change or update any
of the automatically tagged information for a single photo or
for a group of photos. In Photo Detail view, the user can
highlight all detected faces in the photo and tidy up the results
of the automatic detection by removing false detections or
adding missed faces. The system uses a body patch feature (i.e.
a feature modeling the clothes worn by a person) combined
with a face recognition feature to suggest names for detected
faces: the suggested name for an unknown face is the known
face with the most similar body patch and face [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. The user
can confirm the system choice or choose from a shortlist of
suggested names, again based on similarity. Other work has
shown effective methods of suggesting identities within photos
using context-based data [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]: in our ongoing research we are
exploring the combination of this type of approach with both
face recognition and body-patch matching.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>IV. CONCLUSIONS</title>
      <p>
        We have presented the MediAssist demonstrator system for
context-aware management of personal digital photo
collections. Automatically extracted features are supplemented with
semi-automatic annotation which allows the user to correct or
add to the automatically generated annotations. The system
allows the user to formulate precise queries using content and
context based features, or alternatively the user can formulate
simple text queries, which are enabled without the need for
manual annotation. We plan to leverage context metadata to
improve on the performance of content analysis tools [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], and
we will use combined context and content-based approaches
to identity annotation, based on face recognition, body-patch
matching and contextual information. We will also extend the
integration of person recognition to enable the user to query for
a given individual, and the system will (in addition to returning
the photos with confirmed annotations) return a ranked list of
candidate photos which should contain this person.
      </p>
    </sec>
    <sec id="sec-5">
      <title>ACKNOWLEDGMENTS The MediAssist project is supported by Enterprise Ireland under Grant No CFTD-03-216. This work is partly supported by Science Foundation Ireland under Grant No 03/IN.3/I361</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.</given-names>
            <surname>Ahern</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>King</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Davis</surname>
          </string-name>
          .
          <article-title>MMM2: mobile media metadata for photo sharing</article-title>
          .
          <source>In ACM Multimedia</source>
          , pages
          <fpage>267</fpage>
          -
          <lpage>268</lpage>
          , Singapore,
          <year>November 2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Cooray</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. O</given-names>
            <surname>'Connor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Gurrin</surname>
          </string-name>
          , G. Jones,
          <string-name>
            <surname>N.</surname>
          </string-name>
          <article-title>O'Hare, and</article-title>
          <string-name>
            <given-names>A. F.</given-names>
            <surname>Smeaton</surname>
          </string-name>
          .
          <article-title>Identifying person re-occurrences for personal photo management applications</article-title>
          .
          <source>In VIE 2006</source>
          , pages
          <fpage>144</fpage>
          -
          <lpage>149</lpage>
          , Bangalore, India,
          <year>September 2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Graham</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Garcia-Molina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Paepcke</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Winograd</surname>
          </string-name>
          .
          <article-title>Time as essence for photo browsing through personal digital libraries</article-title>
          .
          <source>In ACM Joint Conference on Digital Libraries</source>
          , pages
          <fpage>326</fpage>
          -
          <lpage>335</lpage>
          , Portland, USA,
          <year>July 2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Naaman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Harada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Garcia-Molina</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Paepcke</surname>
          </string-name>
          .
          <article-title>Context data in geo-referenced digital photo collections</article-title>
          .
          <source>In ACM Multimedia</source>
          , pages
          <fpage>196</fpage>
          -
          <lpage>203</lpage>
          , New York, USA,
          <year>October 2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Naaman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. B.</given-names>
            <surname>Yeh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Garcia-Molina</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Paepcke</surname>
          </string-name>
          .
          <article-title>Leveraging context to resolve identity in photo albums</article-title>
          .
          <source>In ACM Joint Conference on Digital Libraries</source>
          , pages
          <fpage>178</fpage>
          -
          <lpage>187</lpage>
          , Denver, CO, USA,
          <year>June 2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>N. O</given-names>
            <surname>'Hare</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Gurrin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Murphy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. F.</given-names>
            <surname>Smeaton</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Jones</surname>
          </string-name>
          .
          <article-title>Digital photos: Where and when?</article-title>
          <source>In ACM Multimedia</source>
          <year>2005</year>
          , pages
          <fpage>261</fpage>
          -
          <lpage>262</lpage>
          , Singapore,
          <year>November 2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>N. O</given-names>
            <surname>'Hare</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Cooray</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Gurrin</surname>
          </string-name>
          , G. Jones,
          <string-name>
            <given-names>J.</given-names>
            <surname>Malobabic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. O</given-names>
            <surname>'Connor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. F.</given-names>
            <surname>Smeaton</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Uscilowski</surname>
          </string-name>
          . Mediassist:
          <article-title>Using content-based analysis and context to manage personal photo collections</article-title>
          .
          <source>In CIVR 2006</source>
          , pages
          <fpage>529</fpage>
          -
          <lpage>532</lpage>
          , Tempe,
          <string-name>
            <surname>AZ</surname>
          </string-name>
          ,
          <year>July 2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S. E.</given-names>
            <surname>Robertson</surname>
          </string-name>
          , S. Walker, ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. M.</given-names>
            <surname>Hancock-Beaulieu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Gatford</surname>
          </string-name>
          .
          <article-title>Okapi at TREC-3</article-title>
          .
          <source>In Proceedings of the Third Text REtrieval Conference (TREC-3)</source>
          , pages
          <fpage>109</fpage>
          -
          <lpage>126</lpage>
          , NIST,
          <year>November 1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>K.</given-names>
            <surname>Rodden</surname>
          </string-name>
          and
          <string-name>
            <given-names>K. R.</given-names>
            <surname>Wood</surname>
          </string-name>
          .
          <article-title>How do people manage their digital photographs?</article-title>
          <source>In CHI 2003</source>
          , pages
          <fpage>409</fpage>
          -
          <lpage>416</lpage>
          , Florida, USA,
          <year>April 2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A.</given-names>
            <surname>Smeulders</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Worring</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Santini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Jain</surname>
          </string-name>
          .
          <article-title>Contentbased image retrieval at the end of the early years</article-title>
          .
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          ,
          <volume>22</volume>
          (
          <issue>12</issue>
          ):
          <fpage>1349</fpage>
          -
          <lpage>1380</lpage>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>