<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>April</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>When is the Structural Context Effective?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Muhammad Ali Norozi</string-name>
          <email>mnorozi@idi.ntnu.no</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paavo Arvola</string-name>
          <email>paavo.arvola@uta</email>
          <email>paavo.arvola@uta.fi</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dept. of Computer and Information Science, Norwegian University of Science and Technology</institution>
          ,
          <addr-line>Trondheim</addr-line>
          ,
          <country country="NO">Norway</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>School of Information Sciences, University of Tampere</institution>
          ,
          <addr-line>Tampere</addr-line>
          ,
          <country country="FI">Finland</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2013</year>
      </pub-date>
      <volume>26</volume>
      <issue>2013</issue>
      <abstract>
        <p>Structural context surrounding the relevant information is intuitively and empirically considered important in information retrieval. Utilizing this context in scoring has improved the retrieval e ectiveness. In this study we will objectively look into the signi cance of the structural context in contextualization process, and try to answer the core question of under which circumstances do we need to deal with the such types of context?</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>Document parts, referred to as elements, have both a
hierarchical and a sequential relationship with each other. The
hierarchical relationship is a partial order of the elements,
which can be represented with a directed acyclic graph, or
more precisely, a tree. In the hierarchy of a document, the
upper elements form the context of the lower ones. In
addition to the hierarchical order, the sequential relationship
corresponds to the order of the running text. From this
perspective, the context covers the surroundings of an
element. An implicit chronological order of a document's text
is formed, when the document is read by a user.</p>
      <p>In focused retrieval, the use of context is a driving force
to alleviate or \un-bias" the retrieval of items with varying
length. Namely, information retrieval is based on evidence
of the retrievable units at hand, and longer text units have
indeed more textual evidence. This has led to a play-safe
strategy where the larger elements are favoured by retrieval
systems. How e ective the context is to neutralize the
sidee ects or bias because of size or length (smaller elements
with less textual evidence gets same opportunity to satisfy
the users need), has been reported experimentally in many
studies [1{3, 6, 9, 10, 8, 7]. The question asked here is:
why the structural context is important in the retrieval of
focused items? In addition, we also ask if the use of context,
under certain circumstances (worst-case), would harm the
retrieval. This means if the context is poor or even
misleading.</p>
    </sec>
    <sec id="sec-2">
      <title>CONTEXT</title>
      <p>
        In semi-structured documents, context of an element
covers everything in the document excluding the element
itself. The surrounding items or elements of the relevant
information is the context. The representation of the
semistructured documents aims to follow the established
structure of documents, i.e., an academic book is typically
composed of hchaptersi, hsectionsi, hsubsectionsi etc.,
structures. hchapter1i is followed by hchapter2i and within
hchapter1i, hsection1i is followed by hsection2i. Elements
hsection1i and hsection2i are siblings, and hence most
likely, semantically related. The following element takes
the concepts further from the preceding elements, and the
preceding elements provide the basics or foundation for the
following elements. Therefore, together in the document
order, the preceding and following elements form a strong
and connected perspective (the kinship structural context),
surrounding the relevant information. Two general types of
context can be distinguished based on the standard
relationships. Hierarchical context, for one, refers to the ancestors,
whereas horizontal refers to the preceding and following
elements [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. In existing studies, context has been referred to
externally as the hyperlink structure of the elements as well.
The context is internal when it is considered from within the
document, and it is external when it is considered outside
the document(s).
      </p>
      <p>
        Contextualization [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] is a re-scoring model, where the
basic score, usually obtained from a full-text retrieval model,
of a contextualized document or element is re-enforced by
the weighted scores of the contextualizing documents or
elements (elements in the sub-tree of interest or structural
context). In this section, we will formalize the context from
in and outside the document using contextualization model.
2.1
      </p>
    </sec>
    <sec id="sec-3">
      <title>Structural Context</title>
      <p>
        Structural context is the sub-tree of interest from the
hierarchical tree structure of the semi-structured document.
Internally, in hierarchical contextualization [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], the
intrinsic tree structure within the XML document is employed.
Structural context in hierarchical or vertical
contextualization is the context based on parent-child relationship in
document's hierarchical structure. An element's parent or
ancestors are accounted to be the structural context, while
contextualizing the element itself. The sub-tree of interest
is shown in Figure 1(a). Horizontal contextualization [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
takes into account the sibling elements in the document's
hierarchical structure as the structural context. If we
visualize the document's hierarchically tree structure, horizontal
structural context is horizontal, as it is based on one level
(the same level as the element to be contextualized) of the
tree at a time (see Figure 1(b)). The most recent form of
contextualization, the Kinship contextualization [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], is both
horizontal (siblings) and vertical (ancestors &amp; descendants
 
 
 
 
      </p>
      <p> 
header
&lt;1.1&gt;
body
&lt;1.2&gt;
sec
&lt;1.2.2&gt;
body
&lt;1.2&gt;
sec
&lt;1.2.2&gt;
&lt;1t.i1t.l1e&gt; &lt;1.i1d.2&gt; r&lt;e1v.i1s.i3o&gt;n ca&lt;t1e.g1o.r4i&gt;es
sec
&lt;1.2.1&gt;
sec
&lt;1.2.3&gt;
&lt;1t.i1t.l1e&gt; &lt;1.i1d.2&gt; r&lt;e1v.i1s.i3o&gt;n ca&lt;t1e.g1o.r4i&gt;es
sec
&lt;1.2.1&gt;
sec
&lt;1.2.3&gt;
&lt;1t.i1t.l1e&gt; &lt;1.i1d.2&gt; r&lt;e1v.i1s.i3o&gt;n ca&lt;t1e.g1o.r4i&gt;es
sec
&lt;1.2.1&gt;
sec
&lt;1.2.3&gt;
elements) but intrinsically non-hierarchical perspective of
the hierarchical information. Structural context is hence
both vertical and horizontal in the document's hierarchical
form, Figure 1(c).</p>
      <p>
        And externally, in citation contextualization [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], the
document's hyperlink structure is taken in to account. The
structural context here is based on the hyperlinks' graph of
documents hyper-linking (connecting) one another in form
of inlinks (indegree) and outlinks (outdegree). In this case,
the sub-graphs instead of tree of interest are the out-links
graph and the in-links graphs (see Figure 1(d)).
2.2
      </p>
    </sec>
    <sec id="sec-4">
      <title>Why Structural Context?</title>
      <p>
        Structural context is the essential component of the
Contextualization model [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. With contextualization model,
using the structural context, the aim is to rank higher an
element in a good context (strong evidence in the structural
context) than an identical element in a not so good context
(less or no evidence in the structural context) within the
document. And therefore, retrieve elements independent of
their sizes. A small element, in term of size, can be viewed
and hence scored in relation to its structural context, and
its smaller size (which means having less evidence in total)
doesn't stop it from being selected as one of the best results.
      </p>
      <p>In order to cope up with the \biasness" issue (described
earlier), in contextualization model, the weight of a relevant
element is adjusted by the basic weights of the elements in
the structural context (its contextualizing elements). In
addition to basic weights, each element in the structural
context of the contextualized element, should possess an impact
factor. An higher impact factor shows the importance of the
contextualizing element and vice versa. The role and
relation of elements in the structural context are operationalized
by giving the element a contextualizing weight. A
contextualization vector is de ned to capture the impact factor
of each contextualizing element, and this contextualization
vector is represented by a g function in Equation 1.
2.3</p>
    </sec>
    <sec id="sec-5">
      <title>Contextualization and Random Walks</title>
      <p>Random walk principle is employed, for contextualization,
to induce a similarity structure over the documents based
on the containment and reverse-containment relationships
(element, sub-element and vice versa). Hence, these
relationships a ect the weight each element, in the structural
context, has in contextualization.</p>
      <p>
        The premise is that good structural context (identi ed by
random walk and the contextualization model [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]) provides
evidence that an element in focused retrieval is a good
candidate to satisfy the user's need and therefore, the elements
should be contextualized by the elements in the sub-tree
of interest. Hence, the good structural context contains a
strong likelihood factor that should be used to deduce that
the contextualized element is a good candidate for the posed
query.
      </p>
      <p>The tree-structure of the XML document (Figure 1) is
assumed to be a graph. In order for the structural context to
take part in the contextualization process, each of the nodes
in the sub-tree of interest should possess an impact factor.
Conceptually, the impact factor is produced in the
following manner: Myriad of random surfers traverse the XML
graphs. In particular, at any time step a random surfer is
found at an element and either (a) makes a next move to the
sub-element of the existing element by traversing the
containment edge, or (b) makes a move to the parent-element
of the existing element, or (c) jumps randomly to another
element in the XML graph. As the time goes on (the
number iterations), the expected percentage of surfer at each
node converges to a limit, the dominant eigenvector of the
XML graph. This limit provides the impact or strength of
each element in the structural context of the element to be
contextualized, in the form of g function. All the elements,
in the structural context of the contextualized element, are
considered for contextualization; where the
contextualization vector g identi es the importance of each of the unit of
the structural context (Equation 1).
2.4</p>
    </sec>
    <sec id="sec-6">
      <title>Generalized Combination Functions</title>
      <p>The generalized re-ranking combination function based on
the random walk principle, which also captures the
structural context, can be formally de ned as follows:
CR(x; f; Cx; gk) = (1
f
y2Cx
f ) BS(x) +
X BS(y) gk(y)</p>
      <p>X gk(y)
y2Cx
(1)
where</p>
      <p>BS(x) is the basic score of contextualized element x
(text-based score, e.g., tf ief )
f is a parameter which determines the weight of the
context in the overall scoring.</p>
      <p>Cx is the kinship context surrounding the
contextualizing element x, i.e., Cx structural context(x),
, because only the structural context containing the
query terms are considered.
gk(y) is the generalized contextualization vector based
on random walk, which gives the authority weight (the
impact) of y, the contextualizing elements (elements in
structural context) of x in the sub-tree of interest.</p>
    </sec>
    <sec id="sec-7">
      <title>EFFECTS OF CONTEXTUALIZATION ON</title>
    </sec>
    <sec id="sec-8">
      <title>DIFFERENT TEST COLLECTIONS</title>
      <p>
        Structural context in the contextualization framework, is
independent of the basic weighting scheme of the elements
and it could be applied on the top of any query language,
retrieval systems and test collections. The e ects of
contextualization on di erent test collections have been observed
in the existing studies. Contextualization model has been
applied on the top of di erent and competitive baseline
systems using a diverse set of test collections, e.g., semantically
annotated Wikipedia collection from INEX 20091, IEEE
collection, and iSearch scienti c collection [
        <xref ref-type="bibr" rid="ref3 ref7 ref8">3, 7, 8</xref>
        ]. In order to
get the best possible baseline system, a data fusion was
performed based on sum of normalized scores (CombSUM) [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]
and Reciprocal Ranking [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] of INEX 2009 submitted runs.
      </p>
      <p>
        In the experimental evaluation the retrieval e ectiveness
at di erent granularity levels were observed. Mainly,
retrieval e ectiveness at paragraph, article and INEX's
focused retrieval level selection has been observed. The
approaches were evaluated using the evaluation framework
provided by TREC and INEX evaluation initiatives. The
reported results were shown to be promising using both TREC
and INEX evaluation framework [
        <xref ref-type="bibr" rid="ref3 ref7">3, 7</xref>
        ].
      </p>
      <p>The focused task in INEX ad-hoc track is to retrieve
most focused elements satisfying an information need
without overlapping elements. An overlapping result list means
that the elements in the result list may have a descendant
relationship with each other and share the same text content.
For instance, in Figure 1 the hentryi element h1.2.2.2.1i
and the hseci element h1.2.2i are overlapping. In the
existing studies, in the focused retrieval task, the INEXs'
focused approach is followed, considering a result list where
only one of the overlapping elements from each branch is
selected. This means that including the hseci element in
the results would mean excluding the entry element in the
results or vice versa.</p>
      <p>Contextualization and the fusion approach as scoring
methods, however, do not take any stand on which elements
should be selected from each branch. Thus a structural
fusion has been performed, where the element level selection
is taken from the baseline run and subsequently re-rank the
elements of the baseline run.
3.1</p>
    </sec>
    <sec id="sec-9">
      <title>Test Settings</title>
      <p>
        The hierarchical structure of XML documents in the
Wikipedia 2009 collection, are captured using the dewey encoding
scheme (as shown in Figure 1). This way each element in the
document possess a unique index within the document, and
together with document's unique id, this becomes unique for
the entire collection. The tree structure of XML documents
are converted into a matrix, and random walk is performed
on this matrix at indexing time, as it is described in detail, in
our earlier work [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. The contextualization vector gk from
Equation 1 is computed o -line for each and every XML
document in the Wikipedia collection. This suggests that
1Wikipedia collection containing 2:66 million semantically
annotated XML documents (50; 7 Gb) and 68 related topics
provided by the INEX 2009 ad-hoc track [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
computing gk vector is feasible for a reasonably large XML
document collections. At the query time, the scores from
gk vector and the basic scores are combined to produce an
overall ranking score, using Equation 1.
      </p>
      <p>
        In the generalized combination function given (Equation 1),
the contextualization force has to be parametrized. In our
earlier work [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], the contextualization force was tuned and
reported the values leading to best overall performance. In
the parametrization process it was found that the optimal
values of contextualization force f (from Equation 1) lies in
the range, (f 2 f.25,..., 2.50g). These optimal values for f
are obtained by using cross-validation technique. A 68-fold2
cross-validation (or complete cross-validation) technique has
been performed - by randomly partitioning the collection
into 68 training and test samples based on the number of
assessed topics. Of the 68 samples, a single sample is retained
as the validation set for testing, and remaining 67 samples
are used as training set. The cross-validation process is
repeated 68 times (for each fold), with each of 68 samples used
exactly only once as validation set. These 68 independent
or unseen samples are then combined to produce a single or
a set of estimations for parameter f .
3.2
      </p>
    </sec>
    <sec id="sec-10">
      <title>Query Term Probabilities</title>
      <p>If a relevant element does not contain any of the query
term(s), it does not match to the query. Hence, in order
to retrieve such elements, some expansive methods, such as
contextualization, ought to be used. It seems obvious that,
in a relevant small element, the probability of occurrence
of a query term is smaller than in a larger element. In
order to demonstrate this lack of evidence on small elements,
we calculated some posteriori probabilities for query term
occurrences in a relevant document (Rd) and in a relevant
paragraph (Rp, i.e., the relevant hpi elements from the XML
graph), based on INEX 2009, 68 topics (title eld) and their
relevance assessments. The probabilities are calculated as
the fraction of relevant elements containing any query term,
or all query terms over all relevant elements of same kind.
The probability of occurrence of any query term (from the
query Q) in a Rp and in a Rd respectively are:</p>
      <p>P
[ q Rp
q2Q
!
= 0:847;</p>
      <p>P
[ q Rd
q2Q
!
= 0:995
This means that the probability of occurrence of none of the
query terms in Rp and a Rd is 0.153 and 0.005 respectively3.
Accordingly, the probabilities of occurrence of all the query
terms in Rp and Rd, respectively are:</p>
      <p>Y P (qjRp) = 0:127;</p>
      <p>Y P (qjRd) = 0:469
q2Q</p>
      <p>q2Q
The di erence in the amount of evidence at di erent
granularity levels become even more obvious, when we draw the
frequencies of the query terms in this picture. A query term
occurs on average 3:4 times in a Rp and 45:4 times in a Rd.</p>
    </sec>
    <sec id="sec-11">
      <title>4. WORST CASE ANALYSIS</title>
      <p>Worst-case for a document d, in contextualization models,
means when structural context of element x is chosen such
that:
structural context(x) = elementsy(d)
2
(2)
(8 elements y in document d
where x and y 2 d )
268, because of the 68 topics from INEX 2009.
3Test is performed without stemming or stop-word removal
1.0
0.9
0.8
0.7
0.3
0.2
0.1
0.0
CRhwioerrsatr-cchaiscea-lp</p>
      <p>p
BaselineFusion
BaselineaFusion
CRhwioerrsatr-cchaiscea-la
0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.0
Recall
Figure 2: Precision - recall, worst-case scenario at
article (a) and paragraph (p) granulation and the
fusion baseline systems.</p>
      <p>The non-structural context (Equation 2), should
theoretically expose the worst-case e ects of the contextualization
model. Non-structural context is structural by de nition,
but physically not in the structural context of element x.
How should we interpret the non-structural context, in
order to experimentally visualize the worst-case scenario?
Instead of taking the actual and true structural context, we
randomly select the structural context from another
nonrelevant but retrieved document. Such a document
(retrieved but not relevant) would have misleading evidence
(false positive) and hence best suited for the worst-case
evaluation. Randomly selecting a document with zero basic
score would be trivial and not suitable for our purposes.</p>
      <p>
        By applying this simplistic approach on every element to
be contextualized, we can formulate the worst-case scenario.
We have used the reciprocal rank fusion approach (fusing
98 INEX 2009 runs) as the baseline system, for worst-case
analysis, which has been used before in our earlier work, nd
further details from [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]:
      </p>
      <p>RRScore(e; q) = X (3)
where</p>
      <p>R is the set of runs (rankings)
and rank(r; e; q) returns the rank of element e as a
result of query q in run r.</p>
      <p>If e is not in the ranking, rank(r; e; q) is not de ned
1
and the outcome of k+rank(r;e;q) is 0.</p>
      <p>The parameter k is for tuning.</p>
      <p>Figure 2 reveals the worst-case depiction of the
contextualization model. Not unexpectedly, the worst-case scenario is
as good as the baseline system, slightly better but not
signi cant enough to be visible statistically. We can claim here
that, when the structural context is chosen randomly
(haphazardly), in the worst-case, the contextualization method
will not be worse than the basic scoring method.</p>
    </sec>
    <sec id="sec-12">
      <title>CONCLUSIONS AND FUTURE WORK</title>
      <p>
        Structural context is the sub-tree of interest, utilized in
conjunction with contextualization model, improves the
retrieval e ectiveness. We have presented an exploratory and
theoretical study into the use of structural context from
elements in the hierarchical structure of information, to
improve retrieval performance. We looked into the structural
context from document's hierarchical structure internally,
and hyperlinks structure externally. We looked
theoretically into the hypothesis that structural context gathered
from within the document, \horizontally" and \vertically"
using the hierarchical tree structure of document, and from
outside, using the hyperlinks graph structure of documents
referencing each other, in uences the retrieval e ectiveness.
Worst-case experiments also support the theoretical
soundness of contextualization, i.e., if we apply contextualization
blindly, in the worst case, we would have as good result
as the basic scoring method. The results obtained in this
study are in-line with the earlier work on
contextualization [
        <xref ref-type="bibr" rid="ref1 ref10 ref3 ref6 ref7 ref9">1, 3, 6, 7, 9, 10</xref>
        ]. In this study we have experimented
with semi-arti cial data, in the sense that we muddled the
context for the worst-case analysis. However, in real data the
quality of context varies as well. For example in Wikipedia
there are di erent kinds of pages ranging from listings to
topically very coherent documents. In order to get the best
results in retrieval, analysing the quality and topical coherency
of context would be of great bene t. The analysis of
context may be topic dependent, since some queries may have
contextual parts. For instance a query: \Losses Belgium in
WW2", crave for answers about Belgium in the context of
WW2.
      </p>
    </sec>
    <sec id="sec-13">
      <title>6. REFERENCES</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>P.</given-names>
            <surname>Arvola</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Junkkari</surname>
          </string-name>
          , and
          <string-name>
            <surname>J. Keka</surname>
          </string-name>
          <article-title>lainen. Generalized Contextualization Method for XML Information Retrieval</article-title>
          .
          <source>In Proc. of the 14th ACM CIKM</source>
          , pages
          <volume>20</volume>
          {
          <fpage>27</fpage>
          . ACM,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>P.</given-names>
            <surname>Arvola</surname>
          </string-name>
          , J. Kekalainen, and
          <string-name>
            <given-names>M.</given-names>
            <surname>Junkkari</surname>
          </string-name>
          .
          <article-title>The E ect of Contextualization at Di erent Granularity Levels in Content-oriented XML Retrieval</article-title>
          .
          <source>In Proc. of the 17th ACM CIKM</source>
          , pages
          <volume>1491</volume>
          {
          <fpage>1492</fpage>
          . ACM,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>P.</given-names>
            <surname>Arvola</surname>
          </string-name>
          , J. Kekalainen, and
          <string-name>
            <given-names>M.</given-names>
            <surname>Junkkari</surname>
          </string-name>
          .
          <article-title>Contextualization Models for XML Retrieval</article-title>
          .
          <source>Info. Processing &amp; Management</source>
          , pages
          <volume>1</volume>
          {
          <fpage>15</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>G.</given-names>
            <surname>Cormack</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Clarke</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Buettcher</surname>
          </string-name>
          .
          <article-title>Reciprocal rank fusion outperforms condorcet and individual rank learning methods</article-title>
          .
          <source>In Proc. of the 32nd international ACM SIGIR conference on Research and development in information retrieval</source>
          , pages
          <volume>758</volume>
          {
          <fpage>759</fpage>
          . ACM,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Geva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kamps</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lethonen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Schenkel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Thom</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Trotman</surname>
          </string-name>
          .
          <article-title>Overview of the INEX 2009 ad-hoc track</article-title>
          .
          <source>Focused Ret. and Evaluation</source>
          , pages
          <volume>4</volume>
          {
          <fpage>25</fpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Mass</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Mandelbrod</surname>
          </string-name>
          .
          <article-title>Component Ranking and Automatic Query Re nement for XML Retrieval</article-title>
          .
          <article-title>Advances in XML IR</article-title>
          , pages
          <volume>1</volume>
          {
          <fpage>18</fpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Norozi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Arvola</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <surname>A. P.</surname>
          </string-name>
          de Vries.
          <article-title>Contextualization using hyperlinks and internal hierarchical structure of wikipedia documents</article-title>
          .
          <source>In Proc. of the 21st ACM CIKM</source>
          , pages
          <volume>734</volume>
          {
          <fpage>743</fpage>
          . ACM,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Norozi</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. P. de Vries</surname>
            , and
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Arvola</surname>
          </string-name>
          .
          <article-title>Contextualization from the Bibliographic Structure</article-title>
          .
          <source>In Proc. of the ECIR 2012 Workshop on Task-Based and Aggregated Search (TBAS2012), page 9</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>P.</given-names>
            <surname>Ogilvie</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Callan</surname>
          </string-name>
          .
          <article-title>Hierarchical Language Models for XML Component Retrieval</article-title>
          .
          <source>Advances in XML IR</source>
          , pages
          <volume>269</volume>
          {
          <fpage>285</fpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>G. Ramirez</given-names>
            <surname>Camps</surname>
          </string-name>
          .
          <article-title>Structural Features in XML Retrieval</article-title>
          .
          <source>PhD thesis</source>
          , SIKS,
          <source>the Dutch Research School for Information and Knowledge Systems</source>
          .,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Shaw</surname>
          </string-name>
          and
          <string-name>
            <given-names>E. A.</given-names>
            <surname>Fox</surname>
          </string-name>
          .
          <article-title>Combination of multiple searches</article-title>
          .
          <source>In The 2nd TREC. Citeseer</source>
          ,
          <year>1994</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>