<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Logistic Regression for Metadata: Cheshire takes on Adhoc-TEL</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Ray R. Larson School of Information University of California</institution>
          ,
          <addr-line>Berkeley</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper we will briefly describe the approaches taken by the Berkeley Cheshire Group for the Adhoc-TEL 2008 tasks (Mono and Bilingual retrieval). Since the AdhocTEL task is new for this year, we took the approach of using methods that have performed fairly well in other tasks. In particular, the approach this year used probabilistic text retrieval based on logistic regression and incorporating blind relevance feedback for all of the runs. All translation for bilingual tasks was performed using the LEC Power Translator PC-based MT system. This approach seems to be a fit good for the limited TEL records, since the overall results show Cheshire runs in the top five submitted runs for all languages and tasks except for Monolingual German.</p>
      </abstract>
      <kwd-group>
        <kwd>Cheshire II</kwd>
        <kwd>Logistic Regression</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>The CLEF Adhoc-TEL collections are different from most of the data used for testing in the various
CLEF tasks. The three sub-collections – British Library (BL), Biblioteque Nationale de France
(BNF), and the Austrian National Library (ONB) – each represent about 1 million bibliographic
records from The European Library union catalog (TEL). The records, we can assume, were
originally in some version of MARC (Machine Readable Cataloging) before they were converted
to a much more simplified bibliographic record based on the Dublin Core metadata schema. Each
of the subcollections use somewhat differing encoding of the (assumed) original MARC data, not
always including all of the fields that might be useful in retrieval.</p>
      <p>Although each the collections were considered to be “mainly” in a particular language (English
for BL, French for BNF, and German for ONB), according to the language codes of the records,
only about half of each collection was in that main language, with virtually all other languages
represented by one or more entries in one or another of the collections. German, French,
English, and Spanish records were available in all of collections. Although this overlap of languages
presents an interesting multilingual search (and evaluation) problem, it was not addressed in our
experiments this year.</p>
      <p>This paper concentrates on the retrieval algorithms and evaluation results for Berkeley’s official
submissions for the Adhoc-TEL 2008 track. All of the runs were automatic without manual
intervention in the queries (or translations). We submitted six Monolingual runs (two German,
two English, and two French) and nine Bilingual runs (each of the three main languages to both
of the other main languages (German, English and French). In addition we submitted three runs
from Spanish translations of the topics to the three main languages.</p>
      <p>This paper first describes the retrieval algorithms used for our submissions, followed by a
discussion of the processing used for the runs. We then examine the results obtained for our
official runs, and finally present conclusions and future directions for Adhoc-TEL participation.
2</p>
    </sec>
    <sec id="sec-2">
      <title>The Retrieval Algorithms</title>
      <p>
        Note that this section is virtually identical to one that appears in our papers from previous CLEF
participation[
        <xref ref-type="bibr" rid="ref7 ref8">8, 7</xref>
        ] The basic form and variables of the Logistic Regression (LR) algorithm used for
all of our submissions was originally developed by Cooper, et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. As originally formulated, the
LR model of probabilistic IR attempts to estimate the probability of relevance for each document
based on a set of statistics about a document collection and a set of queries in combination
with a set of weighting coefficients for those statistics. The statistics to be used and the values
of the coefficients are obtained from regression analysis of a sample of a collection (or similar
test collection) for some set of queries where relevance and non-relevance has been determined.
More formally, given a particular query and a particular document in a collection P (R | Q, D)
is calculated and the documents or components are presented to the user ranked in order of
decreasing values of that probability. To avoid invalid probability values, the usual calculation of
P (R | Q, D) uses the “log odds” of relevance given a set of S statistics, si, derived from the query
and database, such that:
|Qc| is the number of matching terms between a document component and a query,
qtfi is the within-query frequency of the ith matching term,
tfi is the within-document frequency of the ith matching term,
ctfi is the occurrence frequency in a collection of the ith matching term,
ql is query length (i.e., number of terms in a query like |Q| for non-feedback situations),
cl is component length (i.e., number of terms in a component), and
Nt is collection length (i.e., number of terms in a test collection).
ck are the k coefficients obtained though the regression analysis.
      </p>
      <p>If stopwords are removed from indexing, then ql, cl, and Nt are the query length, document
length, and collection length, respectively. If the query terms are re-weighted (in feedback, for
example), then qtfi is no longer the original term frequency, but the new weight, and ql is the
sum of the new weight values for the query terms. Note that, unlike the document and collection
lengths, query length is the “optimized” relative frequency without first taking the log over the
matching terms.</p>
      <p>
        The coefficients were determined by fitting the logistic regression model specified in log O(R|C, Q)
to TREC training data using a statistical software package. The coefficients, ck, used for our
official runs are the same as those described by Chen[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. These were: c0 = −3.51, c1 = 37.4,
c2 = 0.330, c3 = 0.1937 and c4 = 0.0929. Further details on the TREC2 version of the Logistic
Regression algorithm may be found in Cooper et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
2.2
      </p>
      <sec id="sec-2-1">
        <title>Blind Relevance Feedback</title>
        <p>
          In addition to the direct retrieval of documents using the TREC2 logistic regression algorithm
described above, we have implemented a form of “blind relevance feedback” as a supplement to the
basic algorithm. The algorithm used for blind feedback was originally developed and described by
Chen [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. Blind relevance feedback has become established in the information retrieval community
due to its consistent improvement of initial search results as seen in TREC, CLEF and other
retrieval evaluations [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. The blind feedback algorithm is based on the probabilistic term relevance
weighting formula developed by Robertson and Sparck Jones [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ].
        </p>
        <p>Blind relevance feedback is typically performed in two stages. First, an initial search using
the original topic statement is performed, after which a number of terms are selected from some
number of the top-ranked documents (which are presumed to be relevant). The selected terms
are then weighted and then merged with the initial query to formulate a new query. Finally the
reweighted and expanded query is submitted against the same collection to produce a final ranked
list of documents. Obviously there are important choices to be made regarding the number of
top-ranked documents to consider, and the number of terms to extract from those documents. For
ImageCLEF this year, having no prior data to guide us, we chose to use the top 10 terms from 10
top-ranked documents. The terms were chosen by extracting the document vectors for each of the
10 and computing the Robertson and Sparck Jones term relevance weight for each document. This
weight is based on a contingency table where the counts of 4 different conditions for combinations
of (assumed) relevance and whether or not the term is, or is not in a document. Table 1 shows
this contingency table.</p>
        <p>The relevance weight is calculated using the assumption that the first 10 documents are relevant
and all others are not. For each term in these documents the following weight is calculated:
wt = log</p>
        <p>Rt
R−Rt</p>
        <p>Nt−Rt
N−Nt−R+Rt
(4)</p>
        <p>The 10 terms (including those that appeared in the original query) with the highest wt are
selected and added to the original query terms. For the terms not in the original query, the new
“term frequency” (qtfi in main LR equation above) is set to 0.5. Terms that were in the original
query, but are not in the top 10 terms are left with their original qtfi. For terms in the top 10 and
in the original query the new qtfi is set to 1.5 times the original qtfi for the query. The new query
is then processed using the same LR algorithm as shown in Equation 4 and the ranked results
returned as the response for that topic.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Approaches for Adhoc-TEL</title>
      <p>In this section we describe the specific approaches taken for our submitted runs for the
AdhocTEL task. First we describe the indexing and term extraction methods used, and then the search
features we used for the submitted runs.
3.1</p>
      <sec id="sec-3-1">
        <title>Indexing and Term Extraction</title>
        <p>The Cheshire II system uses the XML structure of the documents to extract selected portions for
indexing and retrieval. Any combination of tags can be used to define the index contents.</p>
        <p>Name
recid
names
title
topic
anywhere
date
lang
subject</p>
        <p>Table 2 lists the indexes created by the Cheshire II system for the Adhoc-TEL database and the
document elements from which the contents of those indexes were extracted. The “Used” column
in Table 2 indicates whether or not a particular index was used in the submitted Adhoc-TEL runs.
As the table shows we used only the topic index, which contains most of the content-bearing parts
of records, for all of our submitted runs.
M-DE-TD-T2FB
M-DE-T-T2FB
M-EN-TD-T2FB
M-EN-T-T2FB
M-FR-TD-T2FB
M-FR-T-T2FB
B-ENDE-TD-T2FB
B-ESDE-TD-T2FB
B-FRDE-TD-T2FB
B-DEEN-TD-T2FB
B-ESEN-TD-T2FB
B-FREN-TD-T2FB
B-DEFR-TD-T2FB
B-ENFR-TD-T2FB
B-ESFR-TD-T2FB
0.1742
0.1980 *
0.3466 *
0.2773
0.2438 *
0.1931</p>
        <p>For all indexing we used language-specific stoplists to exclude function words and very common
words from the indexing and searching. The German language runs did not use decompounding
in the indexing and querying processes to generate simple word forms from compounds. The
Snowball stemmer was used by Cheshire for language-specific stemming.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Search Processing</title>
        <p>Searching the Adhoc-TEL collection using the Cheshire II system involved using TCL scripts
to parse the topics and submit the title and description or the title alone from the topics. For
monolingual search tasks we used the topics in the appropriate language (English, German, and
French), for bilingual tasks the topics were translated from the source language to the target
language using the LEC Power Translator PC-based machine translation system.</p>
        <p>The scripts for each run submitted the topic elements as they appeared in the topic to the
system for TREC2 logistic regression searching with blind feedback. When both the “title” and
“description” topic elements were used, they were combined into a single probabilistic query. Table
3 shows which element were used in the “Type” column, T for title only and TD for title and
description.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results for Submitted Runs</title>
      <p>The summary results (as Mean Average Precision) for the submitted bilingual and monolingual
runs for English German and French are shown in Table 3, the Recall-Precision curves for these
runs are also shown in Figures 1 (for monolingual) and 2 (for bilingual). In Figures 1 and 2 the
names for the individual runs represent the language codes, which can easily be compared with
full names and descriptions in Table 3 (since each language combination has only a single run).</p>
      <p>Table 3 indicates runs that had the highest overall MAP for the task by asterisks next to the
run name.</p>
      <p>Obviously the “weak man” in our current implementation remains monolingual German. This
may be due to decompounding issues, but the higher results for title-only monolingual seem
anomalous, since for each other language, the combination of title and description performed
better than title alone.</p>
      <p>In spite of this relatively poor performance in monolingual German, we had the rather
surprising results that for bilingual English to German our submitted run B-ENDE-TD-T2FB was
ranked third overall among the bilingual “to German” runs submitted, and our German to French
bilingual run B-DEFR-TD-T2FB was ranked first in the bilingual “to French” task well ahead of
our English to French run. This would seem to indicate that the our translation system works
quite well with the Adhoc-TEL topics.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>In looking at the overall results for the various Adhoc-TEL tasks, it would appear that the basic
logistic regression with blind relevance feedback approach, coupled with the LEC translation
system is a fairly good combination. Since Adhoc-TEL is a new task, we took a fairly conservative
approach using methods that have worked well in the past.</p>
      <p>In our experiments for other tracks (GeoCLEF for example) we reintroduced fusion approached
for retrieval that performed quite well and could be easily applied to this task as well. For future
work we intend to test these approaches as well as some other approaches that would incorporate
external supplementary topical indexing for the books (primarily) represented by Adhoc-TEL
records.</p>
      <p>1
0.9
0.8
0.7
ion 0.6
ics 0.5
reP 0.4
0.3
0.2
0.1
0</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Aitao</given-names>
            <surname>Chen</surname>
          </string-name>
          .
          <article-title>Multilingual information retrieval using english and chinese queries</article-title>
          . In Carol Peters, Martin Braschler, Julio Gonzalo, and Michael Kluck, editors,
          <source>Evaluation of CrossLanguage Information Retrieval Systems: Second Workshop of the Cross-Language Evaluation Forum</source>
          , CLEF-2001, Darmstadt, Germany,
          <year>September 2001</year>
          , pages
          <fpage>44</fpage>
          -
          <lpage>58</lpage>
          . Springer Computer Scinece Series LNCS 2406,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Aitao</given-names>
            <surname>Chen</surname>
          </string-name>
          .
          <source>Cross-Language Retrieval Experiments at CLEF</source>
          <year>2002</year>
          , pages
          <fpage>28</fpage>
          -
          <lpage>48</lpage>
          . Springer (LNCS #2785),
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Aitao</given-names>
            <surname>Chen</surname>
          </string-name>
          and
          <string-name>
            <given-names>Fredric C.</given-names>
            <surname>Gey</surname>
          </string-name>
          .
          <article-title>Multilingual information retrieval using machine translation, relevance feedback and decompounding</article-title>
          .
          <source>Information Retrieval</source>
          ,
          <volume>7</volume>
          :
          <fpage>149</fpage>
          -
          <lpage>182</lpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>W. S.</given-names>
            <surname>Cooper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Chen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F. C.</given-names>
            <surname>Gey</surname>
          </string-name>
          .
          <article-title>Full Text Retrieval based on Probabilistic Equations with Coefficients fitted by Logistic Regression</article-title>
          .
          <source>In Text REtrieval Conference (TREC-2)</source>
          , pages
          <fpage>57</fpage>
          -
          <lpage>66</lpage>
          ,
          <year>1994</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>William</surname>
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Cooper</surname>
          </string-name>
          ,
          <string-name>
            <surname>Fredric C. Gey</surname>
          </string-name>
          , and Daniel P. Dabney.
          <article-title>Probabilistic retrieval based on staged logistic regression</article-title>
          .
          <source>In 15th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          , Copenhagen, Denmark, June 21-24, pages
          <fpage>198</fpage>
          -
          <lpage>210</lpage>
          , New York,
          <year>1992</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Ray</surname>
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Larson</surname>
          </string-name>
          .
          <article-title>Probabilistic retrieval, component fusion and blind feedback for XML retrieval</article-title>
          .
          <source>In INEX 2005</source>
          , pages
          <fpage>225</fpage>
          -
          <lpage>239</lpage>
          .
          <source>Springer (Lecture Notes in Computer Science, LNCS 3977)</source>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Ray</surname>
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Larson</surname>
          </string-name>
          . Cheshire at geoclef 2007:
          <article-title>Retesting text retrieval baselines</article-title>
          .
          <source>In 8th Workshop of the Cross-Language Evaluation Forum</source>
          ,
          <string-name>
            <surname>CLEF</surname>
          </string-name>
          <year>2007</year>
          , Budapest, Hungary,
          <source>September 19- 21</source>
          ,
          <year>2007</year>
          ,
          <string-name>
            <given-names>Revised</given-names>
            <surname>Selected</surname>
          </string-name>
          <string-name>
            <surname>Papers</surname>
          </string-name>
          , LNCS
          <volume>5152</volume>
          , page to appear
          <source>2008</source>
          , Budapest, Hungary,
          <year>September 2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Ray</surname>
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Larson</surname>
          </string-name>
          .
          <article-title>Experiments in classification clustering and thesaurus expansion for domain specific cross-language retrieval</article-title>
          .
          <source>In 8th Workshop of the Cross-Language Evaluation Forum</source>
          ,
          <string-name>
            <surname>CLEF</surname>
          </string-name>
          <year>2007</year>
          , Budapest, Hungary,
          <source>September 19-21</source>
          ,
          <year>2007</year>
          ,
          <string-name>
            <given-names>Revised</given-names>
            <surname>Selected</surname>
          </string-name>
          <string-name>
            <surname>Papers</surname>
          </string-name>
          , LNCS
          <volume>5152</volume>
          , page to appear
          <source>2008</source>
          , Budapest, Hungary,
          <year>September 2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S. E.</given-names>
            <surname>Robertson</surname>
          </string-name>
          and
          <string-name>
            <given-names>K. Sparck</given-names>
            <surname>Jones</surname>
          </string-name>
          .
          <article-title>Relevance weighting of search terms</article-title>
          .
          <source>Journal of the American Society for Information Science</source>
          , pages
          <fpage>129</fpage>
          -
          <lpage>146</lpage>
          , May-June
          <year>1976</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>