<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>ISWC 2013 Doctoral Consortium `Ontology Evolution for End-User Communities'</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Peter J. Goodall</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Peter Eklund</string-name>
          <email>peklund@uow.edu.au</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Centre for Digital Ecosystems University of Wollongong</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>This project supports application for a practice-based [4] PhD by Peter Goodall. The project will produce a system architecture and proof-of-concept laboratory implementation to model, instrument, prototype, and evaluate the e ectiveness of several alternate system designs which are intended to enable small end-user communities to evolve specialized ontologies and annotations for entities important to them. Depending on available resources, the laboratory may also be used to study bridging-subset ontologies for interchange, federation and seeding of community ontologies. There will be a discussion of the strengths and weaknesses of various approaches tried by others, and a re ection on the usefulness of the system resulting from innovative work of this project.</p>
      </abstract>
      <kwd-group>
        <kwd>information systems</kwd>
        <kwd>ontology</kwd>
        <kwd>taxonomy</kwd>
        <kwd>emergent systems</kwd>
        <kwd>digital ecosystems</kwd>
        <kwd>curation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Each living human community, small or large, cultural, scienti c or
commercial has its characteristic evolving shared nomenclature which enables discussion
and action involving entities and activities important to that community.
Further, communities that interact or exchange information with each-other require
intersecting terminologies.</p>
      <p>Ontology is the concept which encompasses terminology with its meanings
and mechanisms. Ontology has philosophical and computational streams of
interest to this project: Philosophical Ontology has been described as `that branch of
metaphysics concerned with the nature or essence of being or existence' [17]. Its
modern child - Computational Ontology can be described as `a conceptualization
of some universe of discourse, embodied as a declarative formal abstraction'[14].</p>
      <p>
        This project is motivated by observation of a contradiction between the
curatorial and cultural perspectives of an Australian Research Council Linkage Grant
(now inactive) which I project-managed - known as `The Virtual Museum of the
Paci c'[
        <xref ref-type="bibr" rid="ref10 ref9">10,9</xref>
        ] (no longer active).
      </p>
      <p>The Cultural Collections section of our project partner - The Australian
Museum, has over many decades iterated through developing or adopting controlled
vocabularies for cataloging their Paci c Collection. They currently use a concise
in-house taxonomy for very practical reasons.</p>
      <p>While preparing for the launch of the project web-site, we spent time working
with representatives of some Paci c communities. They were concerned that the
collection taxonomy had little relationship with their own way of describing their
artifacts within the collection.</p>
      <p>Both the cultural and curatorial ontologies struggled to be e ective in the
context of the collection. Both the curatorial sta and cultural owners were
experts in the conceptualization of their particular perspectives of the collection,
and both need aid in ongoing development of terms to suit their both separate
and overlapping needs.</p>
      <p>My project goal is to research and develop systems for exploring ontologies
computationally emergent from technically-informal community annotation.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Relevancy</title>
      <sec id="sec-2-1">
        <title>This project has both social and commercial relevancy:</title>
        <p>
          Social - many communities of commitment or interest [
          <xref ref-type="bibr" rid="ref11">21,11</xref>
          ] not equipped
or resourced to create formal ontologies for curating their physical and electronic
artifacts, even though their members are the primary domain experts of those
communities.
        </p>
        <p>Commercial - because nearly every modern large-scale electronic market or
business, has a shortage of technical experts to formally classify objects and data
for cataloging and recommendation.</p>
        <p>Both of these application areas could greatly bene t from further
development of categorization generated from community or customer based annotation.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Related</title>
    </sec>
    <sec id="sec-4">
      <title>Work</title>
      <p>Eleanor Rosch and others, in a number of seminal papers [18,19] developed
the notions of Basic Categories and Prototype Theory, prosecuting the idea
that there are a number of fundamental levels of category that have a degree
of coherence amongst end-users. This experimental work inspires some con
dence in community labeling of objects being used to produce derivable category
structure. More recently work on the Rational Model of Categorization and the
Feature Induction Model examined by Sanborn [20] show promise for inferring
categories from user tags.</p>
      <p>
        The work of Cimiano and colleagues [
        <xref ref-type="bibr" rid="ref6 ref7">6,7</xref>
        ] provides a theoretical and practical
resource for generating Formal Concept Lattices for adaption to work done in our
Virtual Museum of the Paci c project and its Formal Concept Analysis-based
navigation and tagging [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>Mika [15] provides an extended ontology tripartite hypergraph model,
splitting the hypergraph into three bipartite graphs associating actors with concepts,
concepts with instances and actors with instances. These bipartite graphs are
then folded using matrix operations to create various useful a liation networks,
whose properties are used to develop lightweight ontologies demonstrated via a
number of case-studies.</p>
      <p>
        A data-model for capture, archiving and processing of user tagging events
requires some thought. A user labels an object while engaged in some cognitive
context, the person chooses a symbol relevant to that context as a reminder
of the concept from that context. Some investigation of processes and schools
of Semiotics [
        <xref ref-type="bibr" rid="ref5">5,16</xref>
        ] leads to a semiotic triad augmented with a time-stamp and
some exible metadata.
      </p>
      <p>
        The time-trace of these tags bears an interesting a nity with Dynamic Topic
Modeling and Topic Hierarchies [
        <xref ref-type="bibr" rid="ref12 ref2 ref3">2,3,12</xref>
        ] - an approach for incremental discovery
of hierarchic categories which appears worth investigating.
      </p>
      <p>
        As noted in the Problem Description of this overview, development of
category systems by domain experts who are not ontologists is discouraged by the
complexity of traditional approaches. This complexity is re ected in the
software tools readily available. Protege [13] is an example of a quality open-source
ontology development system, many others are available in various states of
development usage and repair [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] .
4
      </p>
      <p>Research Questions
1. How to model simple hierarchies of community generated categories?
2. What is the relationship between community and classi cation?
3. Are time-based data-sets available which demonstrate category drift within
a community over time?
4. How much information external to the tagging discourse is required for an
ontology to emerge?
5. Is emergence as a computational model valid?
6. How to represent collections of tagging events in a way that is simple and
e ective, yet consistent with relevant semiotic theory?
7. How to model requirements for selecting suitable experimental datasets and
feeds
8. How to access and transform datasets/feeds for ingestion into the system,
and select appropriate tool-sets for this?
9. How to represent and derive ontologies from tagging event streams?
10. Which formal models of tags and ontologies are suitable for incremental and
bulk computation?
11. De ne a system architecture for processing, interaction, and visualization.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Hypotheses</title>
      <p>These hypotheses are still very much under development.
1. Basic-level ontologies can be computationally emergent from a suitably
annotated tagging event stream. The emergent ontologies will be judged useful
by the annotating community.
2. Ontologists or curators working in the domain of a community's emergent
ontologies (as derived above) will judge those ontologies to be useful to their
work.
3. A proof of concept technical work-bench/tool-chain can be constructed from
Free and Open Source Software (FOSS) and reasonably generic cloud
computing resources, which can be used to e ectively experiment with extraction
of latent categories in tagging discourses.
4. Cohesiveness of a subject domain's community of interest corresponds to the
coherence of the vocabulary used when tagging objects from that domain.
5. A cohesive community of interest tagging their domain of interest, will
produce a stream of tagging events whose latent categorizations will converge
into a set of basic and superordinate categories [19], which are recognized as
useful by that community.
6. A tagging event data-set with a long-enough time-base will demonstrate
temporal drift of its latent categories.
7. Scalable algorithms can be found to derive basic and superordinate categories
incrementally from suitable annotation event streams.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Approach</title>
      <sec id="sec-6-1">
        <title>1. Literature Survey - Research literature to understand</title>
        <p>{ the fundamental relationships between semiotics, classi cation,
community, crowd-sourcing, folksonomy.
{ computational topics - such as Formal Concept Analysis, ontology
extraction from text, functional programming, analysis packages such as
R, Incanter and Pandas.
2. Characterize and discover suitable data-sets. It is important to nd, where
possible, suitable existing data-sets for testing hypotheses before committing
to the expense and di culty of user-testing.</p>
        <p>{ There are a number of ad-hoc web-sites indicating the existence of
possibly suitable data-sets, many of these sites are out-of-date or informally
curated.
{ Formal experimental data-set registries are being actively implemented.</p>
        <p>These also need characterization and selection to nd examples suitable
for this project.
{ Stream based data-feeds of tagging events are particularly desirable,
especially if the same sources have archives of their data-feeds.
3. Identify end-user communities that have an interest in curation of their
subject areas. Although very disappointing, strong advice that the di culty of
progressing an ethnographic project through ethics committee process has
caused me to design experiments using contemporary western user domains.
I am still expecting to use members of some communities of interest - where
groups have overlapping domains but divergent interests.
4. Perform preliminary processing of data-sets to re ne technical approach and
use that to inform further development of my thesis. I expect to iterate
through this cycle several times.
5. Once algorithms for deriving categorizations latent in the test data-streams
are performing adequately it will be feasible to begin on user-interface
development and visualization. I expect to base the display and tagging interface
for end-users on the The Virtual Museum of the Paci c, and iteratively
evolve it from that base.
6. Once an end-user user-interface is working under informal testing,
nalization of user-testing strategy should be developed. Resources such as Amazon
Mechanical Turk will probably be used for scalable user-testing.
7. Finalize the system architecture for the experimental work-bench. At this
stage of the project I should have su cient knowledge of the characteristics
of data-sets, user-iteration and analysis tools to perform this task with some
con dence. Initially I am strongly motivated to use FOSS where practicable,
and to implement the system so that it can be scaled by moving from a
Linux-based workstation onto a cloud platform.
8. Complete the thesis and package the system for archiving.
7</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Re ections</title>
      <p>I don't regard my possible success as a sign of failure of others, rather I hope
to answer a variant requirement with my own synthesis. Most of the work done
with Ontology development has been very formal, which excludes most busy
domain-experts from their development.</p>
      <p>I hope to produce some interesting results which are a step on the way to
generating useful, evolvable user-ontologies, and that the 'workbench' of tools I
compose will be useful to others in the eld, or a good starting point for further
development.
8</p>
    </sec>
    <sec id="sec-8">
      <title>Evaluation plan</title>
      <p>There are three facets of evaluation for this project:
1. Subject domain usefulness - If a community of end-users tags objects from
their domain of interest, they should recognize the extracted categories and
rate them as useful
2. Ontologist usefulness - A working ontologist or curator should nd the
extracted categories useful in themselves, and in a standard format that they
can use of in their normal working environment. They should also recognize
features in the project's developed workbench that they would like to see as
an improvement in their professional working environment.
3. Workbench deployment e ectiveness - The running and deployment of the
work-bench should be accessible to a researcher in ontology extraction who
has a reasonable ability to maintain and con gure a Linux workstation.
4. Demonstrate that relationships between annotators can be convincingly
deduced by agent a liation analysis from the tripartite graph of agent, object,
and annotation. Agents would need some veri able, relevant relationship
whose description is independent of the tagging dataset.
5. Given a reasonably large tagging event stream from a particular community,
determine how well computationally emergent ontologies from that dataset
are received by that community.
6. Use temporal sliding windows to generate subsets of a large tagging event
stream as input to observe if temporal category drift is observable.
7. Demonstrate computation that generates emergent ontologies incrementally
in an environment where the whole data-set cannot be held in system RAM.
13. Gennari, J.H., Musen, M.A., Fergerson, R.W., Grosso, W.E., Crubzy, M.,
Eriksson, H., Noy, N.F., Tu, S.W.: The evolution of protg: an environment for
knowledge-based systems development. International Journal of
Human-computer studies 58(1), 89123 (2003),
http://www.sciencedirect.com/science/article/pii/S1071581902001271
14. Gruber, T.R., et al.: A translation approach to portable ontology speci cations.</p>
      <p>Knowledge acquisition 5, 199199 (1993)
15. Mika, P.: Ontologies are us: A uni ed model of social networks and semantics.</p>
      <p>Web Semantics: Science, Services and Agents on the World Wide Web 5(1), 5{15
(Mar 2007), http://www.sciencedirect.com/science/article/</p>
      <p>B758F-4MYF67P-1/2/56984a3ddf4632bb98b722551cdb1151
16. Nth, W.: Handbook of semiotics. Indiana University Press (1995), http:
//books.google.com/books?hl=en&amp;lr=&amp;id=rHA4KQcPeNgC&amp;oi=fnd&amp;pg=PR9&amp;dq=
Handbook+of+Semiotics&amp;ots=ddo1tWkS4f&amp;sig=9NNalr 6O7bREpvCPhwQXB9lBn4
17. Oxford English Dictionary: "ontology, n.". (2004),</p>
      <p>http://www.oed.com/view/Entry/131551?redirectedFrom=Ontology
18. Rosch, E., Mervis, C., Gray, W., Johnson, D., Boyes-Braem, P.: Basic objects in
natural categories. Cognitive psychology 8(3), 382439 (1976)
19. Rosch, E.: Principles of categorization. Concepts: core readings p. 189206 (1999),
http://books.google.com/books?hl=en&amp;lr=&amp;id=sj1gczQ-7K8C&amp;oi=fnd&amp;pg=
PA189&amp;dq=%0D%0A1%0D%0APrinciples+of+Categorization%0D%0AEleanor+
Rosch,+1978+&amp;ots=NoqbGizR2u&amp;sig=0WEAZVboo9yC0P8tnryxMC7DT3M
20. Sanborn, A.N., Gri ths, T.L., Navarro, D.J.: A more rational model of
categorization. In: Proceedings of the 28th annual conference of the cognitive
science society. p. 726731 (2006), http://citeseerx.ist.psu.edu/viewdoc/
download?doi=10.1.1.163.3800&amp;rep=rep1&amp;type=pdf
21. Wenger, E.: Communities of Practice: Learning, Meaning, and Identity.</p>
      <p>Cambridge University Press (Sep 1999)</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bergman</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>The sweet compendium of ontology building tools</article-title>
          (
          <year>Jan 2010</year>
          ), http://www.mkbergman.com/862/ the-sweet
          <article-title>-compendium-of-ontology-building-tools/</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Blei</surname>
            ,
            <given-names>D.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>La</surname>
            <given-names>erty</given-names>
          </string-name>
          , J.D.:
          <article-title>Dynamic topic models</article-title>
          .
          <source>In: Proceedings of the 23rd international conference on Machine learning</source>
          . p.
          <volume>113120</volume>
          (
          <year>2006</year>
          ), http://dl.acm.org/citation.cfm?id=
          <fpage>1143859</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Blei</surname>
            ,
            <given-names>D.M.</given-names>
          </string-name>
          , Gri ths, T.L.,
          <string-name>
            <surname>Jordan</surname>
            ,
            <given-names>M.I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tenenbaum</surname>
            ,
            <given-names>J.B.</given-names>
          </string-name>
          :
          <article-title>Hierarchical topic models and the nested chinese restaurant process</article-title>
          .
          <source>In: Advances in Neural Information Processing Systems</source>
          . p.
          <year>2003</year>
          . MIT Press (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Candy</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Practice based research: A guide</article-title>
          .
          <source>CCS Report: 2006-V1</source>
          . 0 November p.
          <volume>19</volume>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Chandler</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Semiotics the Basics</article-title>
          . Taylor &amp; Francis,
          <string-name>
            <surname>Hoboken</surname>
          </string-name>
          (
          <year>2007</year>
          ), http://public.eblib.com/EBLPublic/PublicView.do?ptiID=
          <fpage>308502</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Cimiano</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Staab</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tane</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Automatic acquisition of taxonomies from text: FCA meets NLP</article-title>
          . In: International Workshop &amp;
          <article-title>Tutorial on Adaptive Text Extraction and Mining held in conjunction with the 14th</article-title>
          <source>European Conference on Machine Learning and the 7th European Conference on Principles and Practice of</source>
          . p.
          <volume>10</volume>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Cimiano</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Ontology Learning and Population from Text: Algorithms, Evaluation and Applications</article-title>
          . Springer-Verlag New York, Inc., Secaucus, NJ, USA (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Eklund</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goodall</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wray</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Information retrieval and social tagging for digital libraries using formal concept analysis</article-title>
          .
          <source>In: Computing and Communication Technologies</source>
          , Research, Innovation, and
          <article-title>Vision for the Future (RIVF</article-title>
          ),
          <year>2010</year>
          IEEE RIVF International Conference on. p.
          <volume>16</volume>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Eklund</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goodall</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wray</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bunt</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lawson</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Christidis</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Daniels</surname>
          </string-name>
          , V.,
          <string-name>
            <surname>Van Ollfen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Designing the digital ecosystem of the virtual museum of the paci c</article-title>
          .
          <source>In: 3rd IEEE International Conference on Digital Ecosystems and Technologies</source>
          . IEEE Press (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Eklund</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goodall</surname>
            ,
            <given-names>P.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wray</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Daniel</surname>
          </string-name>
          , V., Van Ol en, M.:
          <article-title>Folksonomy with practical taxonomy, a design for social metadata of the virtual museum of the paci c</article-title>
          .
          <source>In: Proceedings of the 6th International Conference on Information Technology and Applications</source>
          . p.
          <volume>112117</volume>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Fischer</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>Communities of interest: Learning through the interaction of multiple knowledge systems</article-title>
          .
          <source>In: Proceedings of the 24th IRIS Conference</source>
          . p.
          <volume>114</volume>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Fu</surname>
          </string-name>
          , W.T.:
          <article-title>The microstructures of social tagging: a rational model</article-title>
          .
          <source>In: Proceedings of the 2008 ACM conference on Computer supported cooperative work</source>
          . p.
          <volume>229238</volume>
          (
          <year>2008</year>
          ), http://dl.acm.org/citation.cfm?id=
          <fpage>1460600</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>