<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The death of General Custer</article-title>
      </title-group>
      <pub-date>
        <year>1994</year>
      </pub-date>
      <volume>2313</volume>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>General Custer was Civil War
figures in United States Military history. Graduated last in his Possibly.
Union Major soldier. One of the mostfamous andcontroversial
West Point Class (June 1861). Spent first part of the Civil War
Brigadier General of Volunteers just prior to the Bat le of more
as a courier and staff of icer. Promoted from Captain to
Get ysburg, and was given command of the Michigan
relevant
"Wolverines" Cavalary brigade.</p>
      <p>He helped defeat General Stuart's at empt to make a cavalry
strike behind Union lines on the 3rd Day of the Bat le (July 3,
1863), thus markedlycontributing to the Armyof the Potomac's
victory(alarge monument tohis Brigade now stands in theEast
Cavalry Field in Getysburg). Participated in nearly every
cavalry action in Virginia from that point until the end of the
war, always performing boldly, most often bril iantly, and
always seeking publicity for himself and his actions. Ended the Same
war as a Major General of Volunteers and a Brevet Major
relevance
General in the Regular Army.</p>
      <p>Upon Army reorganization in 1886, he was appointed
Lieutenant Colonel of the soon to be renown 7th United States
Cavalry. Fought in the various actions against the Western
Indians, often with a singular brutality (exemplified by his
wiping out of a Cheyenne vilage on the Washita in November
1868). His exploits on the Plains were romanticized by Eastern
Unites States newspapermen, and hewas elevated to legendary
status in his time. The death of his friend, Lucareli
change his life.</p>
      <p>At Get ysburg heremainedwith GeneralGregg
east of town to face jeb Stuart's threat to the Union
rear, although he was previously ordered to the
south. The combined Union force defeated Stuart.</p>
      <p>Returning to the Army of the Potomac in early
1865, he fought at Five Forks; and in the
Appomat ox Campaign. His victories against the
rebel cavalry came at a time when that force was a
ghost of its former self Custer was brevet ed in the
regulars through grades to major general for
Get ysburg, Yelow Tavern, Winchester, Five Forks,
and the Appomat ox Campaign. In addition he was
brevet ed major general of volunteers for
Winchester.</p>
      <p>Remaining in the armyafter thewar, in 1866he
was appointed Lt. Col. of the newly authorized 7th
Cavalry, remaining its active commander until his
death. He took part in the1867 Sioux and Cheyenne
expedition, but was court-martialed and suspended
from duty one year for paying an unauthorized visit
to his wife.
retrieval terms, since they use the structure of the document itself. However, this model greatest
structural properties of the documents, such as sentences or paragraphs [2] in order to dene the
approaches.
and a query in accordance with the relevance of the passages each document is divided (see Figure</p>
    </sec>
    <sec id="sec-2">
      <title>2). This approach, called Passage Retrieval (PR), is not so aected b y the length of the documents</title>
    </sec>
    <sec id="sec-3">
      <title>This paper is structured as follows. The following section presents the basic features of IR</title>
      <p>n system. Third section describe the main improvements introduced for Clef-2002 Conference.
passages. The second one divides each document into semantic pieces according to the dieren t
models. Although each passage is made up by a xed n umber of sentences, we consider that our
Figure 2: Passage retrieval</p>
    </sec>
    <sec id="sec-4">
      <title>The passage extraction model that we propose (IR-n) allows us to benet from the adv antages</title>
      <p>models have to bear in mind the variable size of each passage. Nevertheless, discourse-based and
of discourse-based models since self-contained information units of text, such as sentences, are
models, is not based on the number of passage terms, but on a xed n umber of passage sentences.
semantic models have the main advantage that they return full information units of the document,</p>
    </sec>
    <sec id="sec-5">
      <title>A possible alternative to these models consists on computing the similarity between a document</title>
      <p>At rst glance, we could think that discourse-based models would be the most eectiv e, in
to determine passage boundaries [5].
of each document. On the other hand, window models have as main advantage that they are
simpler to accomplish, since the passages have a previously known size, whereas the remaining
guishes between discourse models, semantic models, and window models. The rst one uses the
which is quite important if these units are used as input by other applications.
used for building passages. Moreover, the relevance measure which, unlike other discourse-based</p>
    </sec>
    <sec id="sec-6">
      <title>This fact allows a simpler calculation of this measure unlike other discourse-based or semantic</title>
      <p>topics in the document [3]. The last one uses windows of a xed size (usually a n umber of terms)
proposal diers from the windo w models since our passages do not have a xed size (i.e. a xed</p>
    </sec>
    <sec id="sec-7">
      <title>Fourth section describes the dieren t runs performed for this campaign and discusses the results</title>
      <p>problem relies on detecting passage boundaries since it depends on the writing style of the author</p>
    </sec>
    <sec id="sec-8">
      <title>PR community generally agrees with the classication proposed in [1], where the author distin</title>
      <p>and besides, they add the concept of proximity to the similarity measure by analysing small pieces</p>
    </sec>
    <sec id="sec-9">
      <title>PR systems can be classied in accordance with the w ay of dividing documents into passages. of text instead of whole documents. Figures 1 and 2 show the main dierences between both obtained. Finally, last section extracts initial conclusions and opens directions for future work. number of words) since we use sentences with a variable size.</title>
      <p>The death of General
Custer occurs in June 25, 1876, at the
bat le of Lit le Big Horn, which resulted in the
extermination of his immediate command and a total
loss of some 266 officers and men. On June 28th, the
bodies were given a hasty burial on the field. The
fol owing year, what may have been Custer's
remains were disinterred and given a military funeral
at West Point. (Monaghan, Jan, r:</p>
    </sec>
    <sec id="sec-10">
      <title>As it can be observed, this formulation is similar to the cosine measure dened in [9 ]. The</title>
      <p>main dierence is that length normalisation is omitted. Instead, our proposal accomplishes
the number of sentences is constant.
In this case, the discourse unit selected is the sentence and a passage is dened as a xed
number N of sentences. This way, although the number of terms of each passage may vary,
length normalisation by dening passage size as a xed n umber of textual discourse units.
Inc
0.00
2.79%
2.92%
- The relevance measure was modied in order to increase the score of the passages when a
sentence contained more than a consecutive word of the question.
- A serie of experiments was performed to determine the suitable size of the passages (the
number N of sentences).
- We could not make any previous experiment for determining the optimum size of the passage
since it was the rst time this approac h was applied.
show these results for short and long questions respectively.</p>
    </sec>
    <sec id="sec-11">
      <title>Once we had determined the optimum length for passages, we designed a second experiment</title>
      <p>performance was measured using the standard average interpolated precision (AvgP).</p>
    </sec>
    <sec id="sec-12">
      <title>In both cases better results are obtained although, the dierence with baseline is more con</title>
      <p>were carried out on the same document collection (EFE agency), but using the 49 test questions
results were obtained and then, we proceeded to determine the optimum value for N. System</p>
    </sec>
    <sec id="sec-13">
      <title>As baseline system we selected the well-known document retrieval model based on the cosine</title>
      <p>passages to 7 sentences since this length achieved the best results for short questions and they
similarity measure [9]. The experiments were designed for detecting the best value for N (the
proposed in Clef-2001.</p>
    </sec>
    <sec id="sec-14">
      <title>For long questions, best results were achieved for passages of 6 or 7 sentences. Tables 1 and 2</title>
    </sec>
    <sec id="sec-15">
      <title>For short questions, best results were obtained when passages were 7 or 8 sentences length.</title>
      <p>also were nearly the best for long queries.
measure when more than one question term was found into a sentence and they presented the
for adapting the similarity measure described before in such a way that allowed increasing this
siderable when using long queries. After analysing these results, we determined to x the size of</p>
    </sec>
    <sec id="sec-16">
      <title>We developed a serie of experiments in order to optimize system performance. These experiments</title>
      <p>number of sentences that make up a passage). Initially, we detected the interval where the best</p>
    </sec>
    <sec id="sec-17">
      <title>IR-n 7 sentences</title>
    </sec>
    <sec id="sec-18">
      <title>Baseline</title>
    </sec>
    <sec id="sec-19">
      <title>IR-n 8 sentences</title>
      <p>the following changes were introduced:</p>
    </sec>
    <sec id="sec-20">
      <title>The main changes proposed for Clef-2002 were designed to solve these problems. Therefore,</title>
    </sec>
    <sec id="sec-21">
      <title>Baseline</title>
    </sec>
    <sec id="sec-22">
      <title>IR-n 7 sentences</title>
      <p>IR-n 6 sentences
evaluating this way, how passages respond to each of them. This approach is fully described in [7]</p>
    </sec>
    <sec id="sec-23">
      <title>This run is a little more complex. The question is divided into several queries. Each query</title>
      <p>contains an isolated idea appearing into the whole question. Then each query is posed for retrieval,
and basic steps are summed up as follows:
conicto. &lt;/ES-narr&gt;
&lt;/top&gt;</p>
    </sec>
    <sec id="sec-24">
      <title>Tambien pueden incluir informacion sobre propuestas o soluciones adoptadas para resolver este</title>
      <p>,description and a sentence of the narrative.</p>
    </sec>
    <sec id="sec-25">
      <title>2. The system generates as many queries as sentences are detected. Each query contains title</title>
    </sec>
    <sec id="sec-26">
      <title>This expansion consists on detecting the 10 more excellent terms of rst 5 reco vered documents,</title>
      <p>and adding them to the original question.</p>
    </sec>
    <sec id="sec-27">
      <title>This run is similar to IR-n1 but applies query expansion according to the model dened in [4].</title>
      <p>de intereses del primer ministro italiano, Silvio Berlusconi.</p>
    </sec>
    <sec id="sec-28">
      <title>Conicto de inter eses en Italia. Encontrar documentos que discutan el problema del conicto</title>
      <p>interes es del primer ministro italiano, Silvio Berlusconi. Los documentos relevantes se referiran</p>
    </sec>
    <sec id="sec-29">
      <title>Conicto de inter eses en Italia. Encontrar documentos que discutan el problema del conicto de</title>
      <p>de forma explcita al c onicto de inter eses entre el Berlusconi poltic o y cabeza del gobierno italiano,
y el Berlusconi hombre de negocios. Tambien pueden incluir informacion sobre propuestas o
soluciones adoptadas para resolver este conicto.</p>
    </sec>
    <sec id="sec-30">
      <title>In this case, from the example question described before the system generates the following</title>
      <p>two queries:</p>
    </sec>
    <sec id="sec-31">
      <title>Query 1. Conicto de inter eses en Italia. Encontrar documentos que discutan el problema del</title>
    </sec>
    <sec id="sec-32">
      <title>Query 2. Conicto de intereses en Italia. Encontrar documentos que discutan el problema</title>
      <p>se referiran de forma explcita al conicto de intereses entre el Berlusconi poltic o y cabeza del
conicto de inter es es del primer ministro italiano, Silvio Berlusconi. Los documentos relevantes
informacion sobre propuestas o soluciones adoptadas para resolver este conicto.
gobierno italiano, y el Berlusconi hombre de negocios.
del conicto de inter es es del primer ministro italiano, Silvio Berlusconi.Tambien pueden incluir
This run uses long questions formed by title, description and narrative. The example questions
was posed for retrieval as follows.
generated queries processed.</p>
    </sec>
    <sec id="sec-33">
      <title>4. Relevant documents are punctuated with the maximum similarity value obtained for all the</title>
      <p>follows:</p>
    </sec>
    <sec id="sec-34">
      <title>This run takes only short questions (title + description). The example question was processed as</title>
      <p>+12.85%
0.00
+4.32%
+10.91%
Inc
+10.82%</p>
    </sec>
    <sec id="sec-35">
      <title>General conclusions are positive. We have obtained considerably better results than in previous</title>
      <p>documents carried out. Second, the system has been correctly trained to obtain the optimum size
modications for the relev ance formula in order to improve the application of vicinity factors.
After this new experience, we are examining several lines of future work. We want to analyse
of passage. Third, the errors introduced by the Spanish lemmatizer have been avoided by using a
edition. This fact has been caused mainly by three aspects. First, the better preprocessing of
the possible improvements that could be obtained using another type of lemmatizer instead of the
simple stemmer that we have used this year. On the other hand we are going to continue studying
simple stemmer.
(IR-n1) improved around a 4% and the remaining runs performed better between 11 and 13%.
systems that participated at this conference. Table 5 shows the average precision for
monolin</p>
    </sec>
    <sec id="sec-36">
      <title>In this section the results achieved by our four runs are compared with the obtained by all the</title>
    </sec>
    <sec id="sec-37">
      <title>As it can be observed, our four runs performed better than median results. Our baseline</title>
      <p>calculated by taking as base the median average precision of all participant systems.
gual runs and computes the increment of precision achieved. This increment (or decrement) was</p>
    </sec>
    <sec id="sec-38">
      <title>4.2 Results</title>
    </sec>
    <sec id="sec-39">
      <title>5 Conclusions and Future Work</title>
    </sec>
    <sec id="sec-40">
      <title>References</title>
    </sec>
    <sec id="sec-41">
      <title>6 Acknowledgements</title>
    </sec>
    <sec id="sec-42">
      <title>Recall 5 10</title>
      <p>Table 5: Results comparison.</p>
    </sec>
    <sec id="sec-43">
      <title>This work has been supported by the Spanish Government (CICYT) with grant TIC2000-0664</title>
      <p>C02-02.</p>
    </sec>
    <sec id="sec-44">
      <title>Annual International ACM SIGIR Conference on Research and Development in Information</title>
    </sec>
    <sec id="sec-45">
      <title>Retrieval, Text Structures, pages 178{185, Philadelphia, PA, USA, 1997.</title>
      <p>[5] Marcin Kaszkiel and Justin Zobel. Passage Retrieval Revisited. In Proceedings of the 20th
[8] Fernando Llopis and Jose L. Vicedo. IR-n system, a passage retrieval systema at CLEF 2001.</p>
    </sec>
    <sec id="sec-46">
      <title>In Workshop of Cross-Language Evaluation Forum (CLEF 2001), Lecture notes in Computer</title>
    </sec>
    <sec id="sec-47">
      <title>Science, Darmstadt, Germany, 2001. Springer-Verlag.</title>
      <p>[6] Marcin Kaszkiel and Justin Zobel. Eectiv e Ranking with Arbitrary Passages. Journal of the</p>
    </sec>
    <sec id="sec-48">
      <title>American Society for Information Science (JASIS), 52(4):344{364, February 2001.</title>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>