<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>OUC's participation in the 2012 INEX Book and Linked-Data Tracks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Michael Preminger</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ragnar Nordlie</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>David Massey</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nils Pharo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Oslo and Akershus University College of Applied Science</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this article we describe the Oslo University College's participation in the INEX 2012 endeavor. This year we participate in the Book Track's "Prove it" (as in 2011) and Social search tasks, as well as the Linked Data track's Ad-hoc task. In 2011, the OUC submitted retrieval results for the "Prove It" task with traditional relevance detection combined with detection of confrmation based on speci city detected through the Wordnet concept hierarchy. In line with our belief that proving or refuting facts are di erent semantic aware actions of speech, we have this year attempted to incorporate some semantic support based on Named entity recognition. For the Social search task, we wish to examine the utility of the MARCdata (subject heading eld) in social searching for readings. For the Linked data task, we wish to explore the possibility of using links as a query expansion mechanism.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>{ what is the di erence between simply being relevant to a claim and
expressing support for a claim
{ how do we modify traditional retrieval to reveal support or refutal of a claim?</p>
      <p>We also made the claim that "Prove It" sorts within the (not very
wellde ned) category "semantic-aware retrieval", which, for the time being will be
de ned by us as retrieval that goes beyond simple string matching, and is aware
of the meaning (semantics) of text.</p>
      <p>Those question, being rhetorical in part, may be augmented by the questions
{ How can one detect the meaning of texts (words, sentences and passages) and
incorporate those in the retrieval process to attain semantic-aware retrieval
and consequently
{ can one exploit technologies developed within the semantic web to improve
semantic-aware retrieval</p>
      <p>The latter is not directly addressed in this paper, but we claim that the
techniques used here point in this direction.
1.1</p>
    </sec>
    <sec id="sec-2">
      <title>Task De nition and User Scenario</title>
      <p>The prove-it task is still at its infancy, and may be subject to some modi cations
in the future. Quoting the user scenario as formulated by the organizers
The scenario underlying this task is that of a user searching for speci c
information in a library of books that can provide evidence to con rm
or refute a given factual statement. Users expect to be pointed directly
at book pages that can help them to con rm or refute the claim of the
topic. Users are assumed to view the ranked list of retrieved book pages
starting from the top of the list and moving down, examining each result.
No browsing is considered (only the returned book pages are viewed by
users).</p>
      <p>This user scenario is a natural point of departure as it is in the tradition
of information retrieval and facilitates the development of the task by using
existing knowledge. As a future strategy, it may be argued that this user scenario
is gradually modi ed, as ranking in the context of proving is a highly complex
process, and, in the context where Prove-it algorithms are most likely to be used,
arguably super uous.
1.2</p>
    </sec>
    <sec id="sec-3">
      <title>What Is a Proof ?</title>
      <p>What constitutes a proof is well de ned in elds like mathematics and computer
science. In connection with a claim or a statement of fact, it is less obvious what
demands a passage of text should satisfy in order to be considered proof of the
claim. Obviously, we are looking for a passage which expresses a relevant truth
about the claim, but what are the characteristics which signal a su cient degree
of relevance and truthfulness? We might want to identify a trustworthy passage,
which in turn might be identi ed by considering the source of the passage, the
degree to which the passage agreed with other passages treating the same claim
or fact, or the centrality of the claim to the main content of the text. We might
want to identify a concentrated passage, a passage where the largest amount of
elements contained in the claim were represented or where they were by some
measure most heavily represented. We might look for a de nitional passage,
which typographically or linguistically showed the characteristics of a de nition.
Or we might try to identify a "proof" by linguistic characteristics, mostly
semantic, which might be of di erent kinds: certain typical words might be relatively
consistently used to speak about a fact or claim in a "proving" manner, writing
in a "proving" mode might entail using terms on a certain level of speci city, etc.
These latter aspects are orthogonal to the statement or claim itself in the sense
that they (at least ideally) apply equally to whatever claim being the subject of
proving / con rming.
1.3</p>
    </sec>
    <sec id="sec-4">
      <title>Semantic Approaches to Proof</title>
      <p>A statement considered as a "proof" (or con rmation) may be characterized
semantically by several indicators:
{ the phenomenon to be supported may be introduced or denoted by speci c
terms, for instance verbs indicating a de nition: "is", "constitutes",
"comprises" etc.
{ terms describing the phenomenon may belong to a speci c semantic category
{ nouns describing the phenomenon may be on a certain level of speci city
{ named entities of di erent kinds are heavily used
{ verbs describing the phenomenon may denote a certain type of action or
state
Deciding which speci city level or which semantic categories will depend on the
semantic content and the relationship between the terms of the original claim.
Without recourse to the necessary semantic analysis, we assume that in general,
terms indicating a proof / con rmation will be on a relatively high level of
speci city. It will in some way constitute a treatment of one or more aspects
of the claim at a certain level of detail, which we expect to be re ected in the
terminology which is applied.</p>
      <p>
        In 2011, we were investigating whether terms, in our case nouns, found on
a page indicated as a potential source of proof diverges in a signi cant way
from other text in terms of level of speci city. We determined the level of noun
speci city through their place in the WordNet([
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]) term hierarchies.
      </p>
      <p>In this year's experiments, we proceed along the same line, trying to detect
named entities. This years e ort represents a starting point in taking into use
named entity detection to assist in identifying con rming pages. Con rmation
or proofs will often be about subjects identi able by a name. Gradually, we rst
need to nd the limits of current detection of named entities, how easy it is to
adapt it to a relatively diverse text mass that the (English part of) our text
collection is, and then the approach's e ectiveness in detecting proving pages.
The two main possibilities in taking NED into use are:
{ Detecting of named entities in general: pages that mention many named
entities are candidates for being "con rming of something". Other methods
are used to nd the speci c subject of proof. this means we only detect
named entities in the book pages.
{ Detecting the named entity being the subject of the statement to be proved.</p>
      <p>This means detecting named entities in the query, and in the.</p>
      <p>Even though the latter possibility looks obvious it entails some problems, like
polymorphism in identi cation of entities, which must be approached. This is
the main rationale for starting out with the former possibility.
1.4</p>
    </sec>
    <sec id="sec-5">
      <title>Ranking According to "Proof E ciency"?</title>
      <p>In this paper we are still following the two-step strategy of rst nding pages
relevant to the claim, and from those pages trying to identify pages that are likely
to prove the claim1. The rst step is naturally done using current strategies for
ranked retrieval. The second stage identi es among relevant documents those
which prove / con rm the statement. Rank order is not necessarily preserved in
this process: if document A comprises a better string-wise match with the claim
than does document B, document B can still be more e cient at proving the
claim than document A is. Not all elements that make a document relevant also
make it a good prover</p>
      <p>Another issue is the context in which prove-it is used. One example is the
writing of a paper. A writer is (again, arguably) more likely to evaluate a greater
number of sources for proof of a claim than he or she would in a context of pure
fact nding. Additionally, di erent contexts would arguably invite di erent proof
emphases. All this advocates for use of other strategies of presenting proving
results than ranked lists.
1.5</p>
    </sec>
    <sec id="sec-6">
      <title>Indexing and Retrieval Strategies</title>
      <p>
        The point of departure of the strategies discussed here is that con rming or
refuting a statement is a simple action of speech that does not require from the
book (the context of the retrieved page) to be about the topic covering the fact.
In this way the "Prove It" task is di erent than e.g. the one referred to in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
This means that we do not need the index we build for search purposes to be
context-faithful (pages need not be indexed in a relevant book context). It is the
formulation of the statement in the book or page that matters.
1 We see refutal as a totally di erent type of task and will not address it in this paper.
      </p>
    </sec>
    <sec id="sec-7">
      <title>Indexing</title>
      <p>In line with the above, indexing should facilitate two main aspects at retrieval
time: identifying relevant pages and nding which of these is likely to prove a
claim. The rst aspect is catered for creating a simple index of all the words in
the corpus, page by page. The pages are treated as separate documents regardless
of the book in which they appear. The second aspect is catered for by
1.7</p>
    </sec>
    <sec id="sec-8">
      <title>Named entity discovery</title>
      <p>Named entity discovery is a natural language processing (NLP) activity. There
exist several tools that perform NED. The choice this time fell on the opennlp
package of the Apache project. The package was used with default settings (no
special training), with the assumption that the big diversity of the book
collection is not apt to any signi cant improvement with respect to the default
settings.
1.8
2
2.1</p>
    </sec>
    <sec id="sec-9">
      <title>Runs and Results</title>
      <sec id="sec-9-1">
        <title>Social Book search</title>
      </sec>
    </sec>
    <sec id="sec-10">
      <title>Introduction</title>
      <p>The social book search features two representations of books: the social data,
which is a mixture of "Amazon data" (descriptive and social data to facilitate
book sale via Amazon) and social encounters as recorded in the libraryThing
fora on one hand, and, on the other hand, traditional library data (MARC
records) entered by professional catalogers. The main purpose is to nd out the
relative utility of each of these representations when it comes to automatic book
recommendation.</p>
      <p>[?] has attempted to compare the utility of social data to that of DEWEY
classi cation data (which are also available in the Amazon records). In this
paper we try to build upon this, and look at subject headings extracted from
the MARC data.
2.2</p>
    </sec>
    <sec id="sec-11">
      <title>Indexing and retrieval strategies</title>
      <p>The collection has been loaded to a database where all types of data about
each book are associated with the book's ISBN. We create an Indri index that
includes all the tagged XML information that is extracted from both amazon
and the LT fora. To each Indri document (book representation) we also add
a section containing the subject headings extracted from the MARC record or
records of that book2.
2 Some of the books contain MARC records from both the Library of congress as well
as the British Library, with a greater diversity of subject headings.
&lt;DOC&gt;
&lt;DOCNO&gt;0525949283&lt;/DOCNO&gt;
&lt;TEXT&gt;
&lt;SH&gt;Balloon a s c e n s i o n s Women b a l l o o n i s t s &lt;/SH&gt;
&lt;bookdoc&gt;&lt;book&gt;&lt;isbn &gt;0525949283&lt;/ isbn &gt;&lt;t i t l e &gt;The L i t t l e
B a l l o o n i s t &lt;/ t i t l e &gt;&lt;ean &gt;9780525949282 &lt;/ ean&gt;&lt;binding &gt;
Hardcover &lt;/binding &gt;&lt;l a b e l &gt;Dutton Adult&lt;/l a b e l &gt;&lt;
l i s t p r i c e &gt;$21 .95 &lt;/ l i s t p r i c e &gt;&lt;manufacturer&gt;Dutton Adult
&lt;/manufacturer&gt;&lt;p u b l i s h e r &gt;Dutton Adult&lt;/p u b l i s h e r &gt;&lt;
r e a d i n g l e v e l/&gt;&lt;r e l e a s e d a t e/&gt;&lt;p u b l i c a t i o n d a t e
&gt;2006 01 19&lt;/ p u b l i c a t i o n d a t e &gt;&lt;s t u d i o &gt;Dutton Adult&lt;/
s t u d i o &gt;&lt;e d i t i o n/&gt;&lt;dewey &gt;813.6&lt;/dewey&gt;&lt;numberofpages
&gt;224&lt;/numberofpages&gt;&lt;dimensions &gt;&lt;height &gt;70&lt;/height &gt;&lt;
width &gt;580&lt;/width&gt;&lt;length &gt;850&lt;/length &gt;&lt;weight &gt;75&lt;/
weight &gt;&lt;/dimensions &gt;&lt;reviews &gt;&lt;review&gt;&lt;authorid &gt;
A1HA6KZZNDCME9&lt;/authorid &gt;&lt;date &gt;2007 02 25&lt;/date&gt;&lt;
summary&gt;More l i k e 2 1/2 s t a r s . . . &lt; / summary&gt;&lt;content &gt;
H i s t o r y c o l l i d e s with f i c t i o n i n THE LITTLE BALLOONIST
s e t . . . may f e e l that the o v e r a l l s t o r y i s a b i t
rushed and that the p o s s i b l e depths that could have
been conveyed j u s t never emerged . &amp;l t ; br /&amp;gt ; &amp;l t ;
br /&amp;gt ;COURTESY OF CK2S KWIPS AND KRITIQUES&lt;/content
&gt;&lt;r a t i n g &gt;2&lt;/r a t i n g &gt;&lt;t o t a l v o t e s &gt;2&lt;/t o t a l v o t e s &gt;&lt;
h e l p f u l v o t e s &gt;1&lt;/h e l p f u l v o t e s &gt;&lt;/review&gt;&lt;review&gt;&lt;
authorid &gt;A1UDDVTG2K1K72&lt;/authorid &gt;&lt;date &gt;2006 02 07&lt;/
date&gt;&lt;summary&gt;Ahh . . . A Wonderful Love Story f o r
Valentine ' s Day&lt;/summary&gt;&lt;content &gt;What a l o v e s t o r y ! A
p e r f e c t p r e s e n t f o r Valentine ' s Day i f you ' r e s t i l l
l o o k i n g f o r any l a s t minute g i f t s . I bought t h i s book
on a f r i e n d ' s recommendation and read i t from s t a r t to
f i n i s h i n one evening . A l o s t l o v e ? A r e k i n d l e d
romance ? A l l d e f i n i t e l y keep the pages t u r n i n g . And
Donn ' s v i v i d d e s c r i p t i o n s o f Napoleonic France a l l
come a l i v e . I h i g h l y recommend . &amp;l t ; br /&amp;gt ; &amp;l t ; br /&amp;
gt ;&lt;/ content &gt;&lt;r a t i n g &gt;5&lt;/r a t i n g &gt;&lt;t o t a l v o t e s &gt;0&lt;/
t o t a l v o t e s &gt;&lt;h e l p f u l v o t e s &gt;0&lt;/h e l p f u l v o t e s &gt;&lt;/review &gt;&lt;/
reviews &gt;&lt;/browseNode&gt;&lt;/browseNodes&gt;&lt;/book&gt;&lt;/bookdoc&gt;
&lt;/TEXT&gt;
&lt;/DOC&gt;</p>
      <p>Every possible element in this XML is made known to the indexing system, so
it can be used as an extent e.g. for retrieval time weighting. An obvious strategy
here is to weight &lt;SH&gt; at retrieval time when trying to nd the e ect of the
subject headings</p>
      <p>One problem we have is that only about two-thirds of our books have MARC
records. This means that full comparison of the utility uses less documents. Still,
with many enough topics we may hope that a good number of those do have
relevant and judged books among those with MARC records3.</p>
      <p>At retrieval time, weighting of subject headings can be done in the following
way:
&lt;query&gt;</p>
      <p>&lt;number&gt;530&lt;/number&gt;
&lt;t e x t &gt;
#combine ( #w e i g h t ( 1 . 0 #combine ( J e s u s . t e x t Why.
t e x t s c h o l a r l y . t e x t p e r s p e c t i v e . t e x t h i s t o r i c a l .
t e x t From . t e x t ) 2 . 0 #combine ( J e s u s . sh Why. sh
s c h o l a r l y . sh p e r s p e c t i v e . sh h i s t o r i c a l . sh From .
sh ) ) )
&lt;/t e x t &gt;
&lt;/query&gt;
3</p>
      <sec id="sec-11-1">
        <title>Preliminary runs and results</title>
        <p>Preliminary runs were performed in accordance with the description in
Section2.2. Table 1 summarizes the results. There seems to be a problem with the
basic setting that makes it di cult to assess the contribution of the subject
headings. This makes us refrain from further analysis at the present moment, apart
from the suspicion that the results may be due to insensitive use of elements
from the data collected from each book.</p>
      </sec>
      <sec id="sec-11-2">
        <title>The linked data track, the Ad-Hoc task</title>
      </sec>
    </sec>
    <sec id="sec-12">
      <title>4.1 Introduction</title>
      <p>The purpose of the linked data track is to nd out how techniques within the
semantic web / linked data can be used to improve and enhance retrieval of
Wikipedia articles. The data collection is an XML'i ed version of a Wikipedia
3 We do not know it at the time of writing.
subset (about 4.1M articles), where incoming and outgoing links are tagged in
terms of RDF-properties (DBpedia), and the article text is also included.</p>
      <p>Our experiment is based on a two-stage approach. The initial search is in
an index built from the entire corpus. Here each article is only represented by
heading and category texts. For each topic, the initial search result (1000 articles)
is enhanced by articles that form triples with the initial articles (as subjects or
objects). This process results in a set of typically several thousands articles,
including the initial set. Those are used to create a smaller "topic-wise index"
with more data on each of the retrieved articles. In addition links in the articles
are collected to enhance the results with the most popular articles that are not
captured in the initial search.</p>
    </sec>
    <sec id="sec-13">
      <title>4.2 Indexing and retrieval strategies</title>
      <p>Stage one Prior to building the main index, a lter removes from the corpus
les that are not considered to be articles. This included les that describe
images. Files that had titles with the pre xes 'File:', 'Wikipedia:', 'Category',
'Portal:' and 'Template:' are removed. The aim of this process is to reduce the
potential noise such les would create.</p>
      <p>Text contained within the following tags is extracted for indexing:
{ Title: tag-element with name-attribute 'title' within the metadata-template
(TITLE)
{ Heading 1: heading-element with level-attribute '2' (H2)
{ Heading 2: heading-element with level-attribute '3' (H3)
{ Category: property-element with name-attribute 'type' (CAT)</p>
      <p>Common headings such as 'References', 'External links' and 'See also' are
excluded from the index.</p>
      <p>An example of a document ready for indexing by Indri:
&lt;DOC&gt;
&lt;DOCNO&gt;1x6bx0ax212420&lt;/DOCNO&gt;
&lt;TEXT&gt;&lt;TITLE&gt;The Scream&lt;/TITLE&gt;&lt;H2&gt;Sources of inspiration&lt;/H2&gt;
&lt;H2&gt;Thefts&lt;/H2&gt;&lt;H2&gt;In popular culture&lt;/H2&gt;&lt;H2&gt;Gallery&lt;/H2&gt;
&lt;H3&gt;Depersonalization disorder&lt;/H3&gt;&lt;CAT&gt;1893 paintings&lt;/CAT&gt;
&lt;CAT&gt;Edvard Munch paintings&lt;/CAT&gt;&lt;CAT&gt;Expressionist paintings&lt;/CAT&gt;
&lt;CAT&gt;Modern paintings&lt;/CAT&gt;&lt;CAT&gt;Symbolist paintings&lt;/CAT&gt;&lt;/TEXT&gt;
&lt;/DOC&gt;</p>
      <p>Our aim is to use the link structure within the Wikipedia articles to enhance
the initial search. It was therefore necessary to build a hash that links the title
of the article to the le location. A Perl hash with this structure is used:
$fileHash{'La_Concorde'} = '1x6bx0ax265017';
$fileHash{'Embassy_of_Barbados_in_Washington,_D.C.'} = '1x6bx0ax31828384';
$fileHash{'Bosnia_and_Herzegovina_Hockey_League'} = '1x6bx0ax25617492';
$fileHash{'The_Scream'} = '1x6bx0ax212420';
$fileHash{'Belgium_at_the_1924_Summer_Olympics'} = '1x6bx0ax7521518';
$fileHash{'Lambrini'} = '1x6bx0ax5264971';</p>
      <p>The query against the large index is limited to the text of the
descriptionelement of the topic.</p>
      <p>An example of a search:
&lt;text&gt;
#weight(
10.0 #combine(learn.title major.title bicycle.title
races.title multi.title affairs.title tour.title
de.title france.title runs.title milan.title
san.title remo.title )
5.0 #combine(learn.cat major.cat bicycle.cat races.cat
multi.cat affairs.cat tour.cat de.cat france.cat
runs.cat milan.cat san.cat remo.cat )
2.5 #combine(learn.h2 major.h2 bicycle.h2 races.h2
multi.h2 affairs.h2 tour.h2 de.h2 france.h2
runs.h2 milan.h2 san.h2 remo.h2 )
1.0 #combine(learn.h3 major.h3 bicycle.h3 races.h3 multi.h3
affairs.h3 tour.h3 de.h3 france.h3 runs.h3
milan.h3 san.h3 remo.h3 ) )
&lt;/text&gt;</p>
      <p>The elds title, cat, h2 and h3 where weighted in falling importance. The top
1,000 results are returned.</p>
      <p>Stage two In-links to the articles had the following mark-up:
&lt;property name='http://dbpedia.org/ontology/wikiPageWikiLink'&gt;
&lt;subject name='http://dbpedia.org/resource/List_of_paintings_by_Edvard_Munch'&gt;
&lt;/subject&gt;
&lt;/property&gt;
&lt;property name='http://dbpedia.org/ontology/wikiPageWikiLink'&gt;</p>
      <p>&lt;subject name='http://dbpedia.org/resource/Culture_of_Norway'&gt;&lt;/subject&gt;
&lt;/property&gt;
&lt;property name='http://dbpedia.org/ontology/wikiPageWikiLink'&gt;</p>
      <p>&lt;subject name='http://dbpedia.org/resource/Silence_%28Doctor_Who%29'&gt;&lt;/subject&gt;
&lt;/property&gt;
While out-links:
&lt;property name='http://dbpedia.org/ontology/wikiPageWikiLink'&gt;</p>
      <p>&lt;object name='http://dbpedia.org/resource/Pastel'&gt;&lt;/object&gt;
&lt;/property&gt;
&lt;property name='http://dbpedia.org/ontology/wikiPageWikiLink'&gt;</p>
      <p>&lt;object name='http://dbpedia.org/resource/Tempera'&gt;&lt;/object&gt;
&lt;/property&gt;
&lt;property name='http://dbpedia.org/ontology/wikiPageWikiLink'&gt;</p>
      <p>&lt;object name='http://dbpedia.org/resource/Oil_painting'&gt;&lt;/object&gt;
&lt;/property&gt;</p>
      <p>All the in and out links from the 1,000 articles in the result for a topic are
stored in a single array together with their frequencies. The 500 most popular
links are added to the initial result set.</p>
      <p>The combination of the original articles and the most popular linked articles
are then indexed. This new index is based on the entire text of the article, i.e.
the tags are removed. An example of a document ready for indexing:
&lt;DOC&gt;
&lt;DOCNO&gt;1x6bx0ax212420 &lt;/DOCNO&gt;
&lt;TEXT&gt;
212420 The Scream The Scream d i s a m b i g u a t i o n The Scream
j p g 220 px The Scream Norwegian S k r i k Edvard Munch
1893 O i l p a i n t i n g O i l tempera and p a s t e l on cardboard
91 73 5 Oslo N a t i o n a l G a l l e r y o f Norway N a t i o n a l G a l l e r y
The Scream Norwegian S k r i k c r e a t e d i n 18931910 The</p>
      <p>Scream r e t u r n s damaged but younger News com au 2008
05 21 i s the t i t l e o f e x p r e s s i o n i s m e x p r e s s i o n i s t
p a i n t i n g s and p r i n t s i n a s e r i e s by Norway Norwegian
a r t i s t Edvard Munch showing an a g o n i z e d f i g u r e a g a i n s t
a blood red sky The l a n d s c a p e . . .
&lt;/TEXT&gt;
&lt;/DOC&gt;</p>
      <p>The new smaller index, approximately 1,400 articles, was searched to obtain
the nal ranking.
4.3</p>
    </sec>
    <sec id="sec-14">
      <title>Alternative strategies</title>
      <p>The chosen strategy is only one of many. Future experiments could study the
importance of these components in the retrieval algorithm:
{ Alternative weighting of the indexed elds in the initial (stage one) search
{ More elds or di erent elds in the initial index
{ Finding related articles using only in links or out links
{ A smaller or larger indexing of the articles in stage two. The index could
be enhanced with Google data on incoming link texts. Freebase and other
linked data sources could also be used to enhance the index.
{ Only the description was used to represent the query. Would additional elds
from the topic give improved results?</p>
      <sec id="sec-14-1">
        <title>Discussion, Limitation and Further Research</title>
        <p>At the same time that the book world becomes more and more digital,as old
books are being digitized and new books are increasingly published digitally,
information not published in book format becomes more and more "semantic"
in the sense that data pieces (as opposed to exclusively documents in the web's
rst years) are linked together and made available. These two parallel
development entail great opportunities in the exploitation of book material for di erent
purposes, of which the topic of this paper is one example.</p>
        <p>This paper provides an example of the possibilities and the challenges. Whereas
"WordNet speci city", here representing content independent linguistic
semantic, is one simple example of information that can be used to systematically
extract semantics from written content, other much larger and much more
complicated sources of semantics, the semantic web and linked data, are waiting to
be used in a similar (or related) way. To explore these possibilities we will need
to experiment with more modern texts than what our present test collection
contains.</p>
        <p>To judge by the results of the runs presented here, this path of research,
though promising, still requires a lot of modi cation and calibration.</p>
        <p>Exploring the semantics of a page in a basically statistical manner may be
seen as a superposition of independent components. Counting occurrences of
special words is one component on which we superimpose the detection of noun
speci city. The treatment using WordNet represents further progress from the
2010 experiments, but is still rudimentary. Nouns are currently the only
wordclass we are treating, using only level of speci city. trying to detect classes nouns
using the lateral structure of synsets may be another path to follow. It is also
conceivable that treating of other word classes, primarily verbs, might contribute
to the treatment. Verbs are more complicated than nouns in WordNet and such
treatment will be more demanding.</p>
        <p>Utilizing digital books poses new challenges on information retrieval. The
mere size of the book text poses both storage, performance and content related
challenges as compared to texts of more moderate size. But the challenges are
even greater if books are to be exploited not only for nding facts, but also to
support exploitation of knowledge, identifying and analyzing ideas, a.s.o.</p>
        <p>This article represents work in progress. We explore techniques gradually in
an increasing degree of complexity, trying to adapt and calibrate them.</p>
        <p>
          Even though such activities may be developed and re ned using techniques
from e.g. Question Answering[
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], we suspect that employing semantics-aware
retrieval [
          <xref ref-type="bibr" rid="ref5 ref6">5,6</xref>
          ], which is closely connected to the development of the Semantic
Web [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] would be a more viable (and powerful) path to follow.
        </p>
        <p>
          One obstacle particular to this research is the test collection. Modern
ontologies code facts that are closely connected to the modern world. For example
the Yago2 [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] ontology, that codes general facts automatically extracted from
Wikipedia, may be complicated to apply to an out-of-copyright book collection
emerging from academic specialized environments. But this is certainly a path
to follow.
6
        </p>
      </sec>
      <sec id="sec-14-2">
        <title>Conclusion</title>
        <p>This article is a further step in a discussion about semantics-aware retrieval in
the context of the INEX book track. Proving (or con rmation or support) of
factual statements is discussed in light of some rudimental retrieval experiments
incorporating semantics. We also discuss the task of proving statement, raising
the question whether it is classi able as a semantics-aware retrieval task. Results
are highly inconclusive.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Kazai</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koolen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kamps</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Doucet</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Landoni</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Overview of the inex 2010 book track: Scaling up the evaluation using crowdsourcing</article-title>
          .
          <source>In: Comparative Evaluation of Focused Retrieval. Volume 6932 of Lecture Notes in Computer Science</source>
          . Springer Berlin / Heidelberg (
          <year>2011</year>
          )
          <volume>98</volume>
          {
          <fpage>117</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Fellbaum</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>WordNet : an electronic lexical database</article-title>
          . MIT Press, Cambridge, Mass (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Cartright</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Feild</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Allan</surname>
          </string-name>
          , J.:
          <article-title>Evidence nding using a collection of books</article-title>
          .
          <source>In: BooksOnline '11 Proceedings of the 4th ACM workshop on Online books, complementary social media and crowdsourcing</source>
          , Amherst, MA (
          <year>2011</year>
          )
          <volume>11</volume>
          {
          <fpage>18</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>VOORHEES</surname>
            ,
            <given-names>E.M.:</given-names>
          </string-name>
          <article-title>The trec question answering track</article-title>
          .
          <source>Natural Language Engineering</source>
          <volume>7</volume>
          (
          <year>2001</year>
          )
          <volume>361</volume>
          {
          <fpage>378</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Finin</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          , May eld, J.,
          <string-name>
            <surname>Joshi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cost</surname>
            ,
            <given-names>R.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fink</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Information retrieval and the semantic web</article-title>
          .
          <source>In: Proc. 38th Int. Conf. on System Sciences, Digital Documents Track (The Semantic Web: The Goal of Web Intelligence)</source>
          . (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6. May eld, J.,
          <string-name>
            <surname>Finin</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Information retrieval on the semantic web: Integrating inference and retrieval</article-title>
          .
          <source>In: SIGIR Workshop on the Semantic Web</source>
          , Toronto. (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Berners-Lee</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hendler</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lassila</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>The semantic web</article-title>
          .
          <source>Scienti c American</source>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Ho</surname>
            <given-names>art</given-names>
          </string-name>
          , J.,
          <string-name>
            <surname>Suchanek</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berberich</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weikum</surname>
          </string-name>
          , G.:
          <article-title>Yago2: A spatially and temporally enhanced knowledge base from wikipedia</article-title>
          .
          <source>Special issue of the Arti cial Intelligence Journal</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>