<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Managing Personal Information by Automatic Titling of E-mails</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Cedric Lopez</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Violaine Prince</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mathieu Roche</string-name>
          <email>mrocheg@lirmm.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Univ. Montpellier 2, LIRMM</institution>
          ,
          <addr-line>Montpellier</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents an approach that enables automatic titling of e-mails relying on the morphosyntactic study of real titles. Automatic titling of e-mails has two interests: Titling mails 'no object' and managing personal information. The method is developed in three stages: Candidate sentences determination for titling, noun phrases extraction in the candidate sentences, and nally, selecting a particular noun phrase as a possible e-mail title. A human evaluation associated with ROC Curves are presented.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>A title de nition met in any dictionary is 'word, expression, sentence, etc.,
serving to indicate a paper, one of its parts [...], to give its subject.' So it seems that
a title role can be assumed by a well formed word group, an expression, a topic
or a simple word, related to the text content, in one way or another. It ensues
that some groups of well formed words can be convenient for a title, which means
that a text might get several possible titles. A title varies in length (i.e. number
of words), form and local focus. So, the human judgment on a title quality will
always be subjective and several di erent titles might be judged as relevant to
a given content.</p>
      <p>This paper deals with an automatic approach providing a title to an e-mail,
which meets the di erent characteristics of human issued titles. So, when a title
is absent (e-mails without subject), the described method enables the user to
save time by informing him/her in order to manage its personal data. Actually,
a relevant title is an important issue for the person who wants to correctly
classify its e-mails. Let us note that titling is not a task to be confused with
automatic summarization, text compression, and indexation, although it has
several common points with them. This will be detailed in the 'related work'
section.</p>
      <p>The originality of this method is that it relies on the morphosyntactic
characteristics of existing titles to automatically generate a document heading. So
the rst step is to determine the nature of the morphosyntactic structure in
e-mail titles. A basic hunch is that a key term of a text can be used as its
title. But studies have shown that very few titles are restricted to a single term.
Besides, the reformulation of a text relevant elements is still a quite di cult
task, which will not be addressed in the present work. The state-of-the art in
automatic titling (section 2) and our own corpus study have stressed out the
following hypothesis: It seems that the rst sentences of a document tend to
contain the relevant information for a possible title. Our approach (section 3)
extracts crucial knowledge in these selected sentences and provide a title. An
evaluation obtained on real data is presented in section 4.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>It seems that no scienti c study leading to an automatic titling application was
published. However, the title issue is studied in numerous works.</p>
      <p>
        Titling is a process aiming at relevantly representing the contents of
documents. It might use metaphors, humor or emphasis, thus separating a titling
task from a summarization process, proving the importance of rhetorical status
in both tasks [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Titles have been studied as textual objects focusing on fonts,
sizes, colors, . . . [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Also, since a title suggests an outline of the associated
document topic, it is endowed with a semantic contents that has three functions:
Interest and captivate the reader, inform the reader, introduce the topic of the
text.
      </p>
      <p>
        It was noticed that elements appearing in the title are often present in the
body of the text [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] has showed that the rst and last sentences of
paragraphs are considered important. The recent work of [
        <xref ref-type="bibr" rid="ref19 ref2 ref7">2, 7, 19</xref>
        ] supports this idea
and shows that the covering rate of those words present in titles, is very high in
the rst sentences of a text. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] notices that very often, a de nition is given
in the rst sentences following the title, especially in informative or academic
texts, meaning that relevant words tend to appear in the beginning since de
nitions introduce the text subject while exhibiting its complex terms. The latter
indicate relevant semantic entities and constitute a better representation of the
semantic document contents [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <p>
        A title is not exactly the smallest possible abstract. While a summary, the
most condensed form of a text, has to give an outline of the text contents that
respects the text structure, a title indicates the treated subject in the text
without revealing all the content [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Summarization might rely on titles, such as
in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] where titles are systematically used to create the summary. This method
stresses out the title role, but also the necessity to know the title to obtain a good
summary. Text compression could be interesting for titling if a strong
compression could be undertaken, resulting in a single relevant word group. Compression
texts methods (e.g. [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]) could be used to choose a word group obeying to titles
constraints. However, one has to largely prune compression results to select the
relevant group [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
      <p>A title is not an index : A title does not necessarily contain key words (and
indexes are key words), and might present a partial or total reformulation of the
text (what an index is not).</p>
      <p>Finally, a title is a full entity, has its own functions, and titling has to be
sharply distinguished from summarizing and indexing.</p>
      <p>A rapid survey of existing documents helps to fathom some of title
characteristics such as length, and nature of part-of-speech items often used. Next section
is devoted to our automatic titling approach.
3</p>
    </sec>
    <sec id="sec-3">
      <title>The Automatic Titling Approach</title>
      <p>
        By leaning on the previous work (section 2) and our previous study [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], we
propose an automatic titling approach in order to title e-mails.
      </p>
      <p>
        The rst elementary step consists in determining the textual data from which
we will build a title. These data have to contain the information necessary for
the titling of the document. As said before, [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] has concluded that the maximal
covering of the words of the title in the text, was obtained by extracting the rst
seven sentences and both last ones.
      </p>
      <p>
        The following sections present our methods. The main idea consists in
selecting the most relevant Noun Phrase (NP) for its use as title [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
3.1
      </p>
      <sec id="sec-3-1">
        <title>Extracting of the Noun Phrases (NP)</title>
        <p>Corpus analysis showed that the titles of e-mails contain few verbs and are short
(between approximately two and six words) (Table 1). Our aim is to extract the
most relevant noun phrases in order to provide a title.</p>
        <p>Nature
E-mails
% Noun
73
% Named entity
53
% Verb
6</p>
        <p>Number of Words
5</p>
        <p>
          For that purpose, e-mails are tagged with TreeTagger [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. Our NP extraction
method is inspired from [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] who determined syntactical patterns allowing noun
phrase extraction, e.g. N oun1 Adjective1, N oun1 Det1 N oun2, N oun1
N oun2, and so forth. We set up syntactical lters, adapted to French, allowing
the extraction of NP having a maximal size of 6 words (For example 'noun
prep - det - noun - prep - det'). This limit of size is inspired from the maximal
title length for e-mails.
        </p>
        <p>Next section consist in selecting the most relevant NP extracted, for its use
as title. In the following section, we shall use the TF-IDF measure to calculate
the score of every NP. This score can be the maximal TF-IDF obtained for a
word of the NP (TMAX ) either the sum of the TF-IDF of every word of the NP
(TSUM ). Finally, the TALL method is presented.</p>
      </sec>
      <sec id="sec-3-2">
        <title>Selection of NP with statistical criteria</title>
        <p>
          We shall use the TF-IDF measure [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] to calculate the score of every NP
extracted from the e-mail text.
        </p>
        <p>The TF-IDF mesure is a weight often used in information retrieval and text
mining. This weight is a statistical measure used to evaluate how important a
word is to a document in the corpus.</p>
        <p>tfi;j =</p>
        <p>ni;j
Pk(nk;j )
(1)
ni;j is the number of occurrences of the considered term ti in document
dj , and the denominator is the sum of number of occurrences of all terms in
document dj .</p>
        <p>idfi = log</p>
        <p>jDj
jdj : ti 2 dj j
(2)
jDj : total number of documents in the corpus.
jdj : ti 2 dj j : number of documents where the term ti appears.</p>
        <p>Let us note that if new emails arrive in the corpus, the TF-IDF will be
recalculed. The NP score can be the maximal TF-IDF obtained for a word of
the NP (TMAX ) either the sum of the TF-IDF of every word of the NP (TSUM ).
Finally, an improvement of these methods is presented (TALL).</p>
        <p>
          TMAX . The TMAX method consists in calculating a score for each NP in the
rst sentences [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. For each word of the candidate NP, the TF-IDF is computed.
The score for each candidate NP is the maximum TF-IDF of the words of the
NP. With this method, discriminant terms are highlighted. For example, in the
noun phrase 'contribution recherche' (research contribution) (N P 1) and
'nouvelle relecture' (new review ) (N P 2), N P 1 will be retained, the term contribution
being more discriminant than 'recherche' (research), 'nouvelle' (new ), and
'relecture' (review ) in our e-mail corpus.
        </p>
        <p>Contrarly to TMAX , another method consists in extracting the NP containing
the most information: TSUM .</p>
        <p>TSUM . For each word of the NP candidate (extracted from rst sentences of the
e-mail), the TF-IDF is calculated. The score of each NP candidate is the sum of
each TF-IDF. This method favors long noun phrases. For example, let both NP
'soucis de vibration' (vibration nuisance) (N P 3) and 'soucis de vibration avec
Saxo' (vibration nuisance with Saxo) (N P 4). N P 4 will be privileged because it
is a superset of N P 3. However, this method still allows to distinguish between
noun phrases of the same size: N P 2 obtains a better score than N P 1 because
the sum of the TF-IDF for the terms 'nouvelle' (new ) and 'relecture' (review ) is
higher than the sum for 'contribution' (contribution) and 'recherche' (research).</p>
        <p>With these methods (TMAX and TSUM ), we only worked on the rst
sentences (two sentences) of the e-mails. In the next section, we propose an approach
using all the texts.</p>
        <p>TALL. Generally, it is advisable that relevant terms for titling are present in the
rst and last sentences of the text (see Section 2). However, as regards e-mails,
our statistic study shows that terms appearing in real title are rarely at the end
of the text (Fig. 1).</p>
        <p>In the Figure 1, the Y axis represents the number of words that appears both
in the title and in the text. The X axis represents the parts of the text. Actually
in order to identify the parts of the text where the terms of the title appear,
the text was divided in eight parts. For instance, in the Figure 1, four words
are both in the title and on the sixth part of the text. Of course, determiners,
prepositions, articles, and so forth, are not considered in this study. We note
that the dispersal of relevant terms in the text takes an hyperbolic form.</p>
        <p>Let us note that if the NP score is based only on the TF-IDF 1, the results
indicate that NP candidates for a title could be extracted wherever in the text
1 Score calculated in the same way as TSUM , but on the complete text and not only
on the rst sentences
(Fig. 2). We will see that this method, called TF REQ, does not obtain good
results (see Section 4).</p>
        <p>N Pnumber 1, 43 for N Pnumber 43). We use
apply di erent values to .</p>
        <p>Our objective is to use this information during the calculation of the NP
score. We propose a method combining the NP position in the text and its
semantic contents.</p>
        <p>The ScoreP enables to give more importance to the NP extracted at the
beginning (section 3.2) of the text. P is the position of the NP (e.g., 1 for
= 12 . In a future work, we plan to
ScoreP =
1
P
(3)</p>
        <p>The ScoreT F IDF is calculated in the same way as TSUM , but on the
complete text and not only on the rst two sentences. Finally, the score of the NP
(ScoreTALL ) is the sum of ScoreP and ScoreT F IDF .</p>
        <p>ScoreT F IDF =
n</p>
        <p>X (T F IDF )term
term=1
ScoreTALL = ScoreP + ScoreT F IDF
(4)
(5)</p>
        <p>With the example given in the Fig. 3, the fourth extracted NP is chosen:
1. Dans un soucis (In a concern)
2. Soucis d'amelioration (Concerns of improvements)
3. Amelioration de la Journee (Improvement of the Day)
4. Amelioration de la Journee Scienti que (Improvement of the Scienti c Day)
5. La Journee Scienti que du LIRMM (The Scienti c Day of the LIRMM)
6. Scienti que du LIRMM (Scienti c of the LIRMM)
7. Du LIRMM (Of the LIRMM)
8. LIRMM
9. La frequence d'une fois (Frequency of one time)
10. ...</p>
        <p>The Figure 4 shows that the ScoreP gives an important weight to the rst
noun phrases. Moreover, the second and fourth NP have an important value of
ScoreT F IDF . Finally, the ScoreTALL favors the fourth NP as a relevant title.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experiments</title>
      <p>The corpora consists of French personal e-mails from di erent persons and
registers ; they are more or less well written. Our three methods studied in this
paper are evaluated. First of all, we have studied the behavior of our methods
by using ROC Curves.
4.1</p>
      <sec id="sec-4-1">
        <title>ROC Curves</title>
        <p>
          ROC Curves measure the quality of the obtained ranking. Initially the ROC
Curves (Receiver Operating Characteristic), detailed in [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], come from the eld
of signal processing. ROC Curves are often used in medicine to evaluate the
validity of diagnosis tests. ROC Curves show in X-coordinate the rate of false
positives (in our case, not relevant title) and in Y-coordinate the rate of true
positives (relevant titles). The surface under the ROC Curve (AUC - Area Under
the Curve), can be seen as the e ectiveness of a measurement of interest. The
criterion related to the surface under the curve is equivalent to the statistical
test of Wilcoxon-Mann-Whitney (see [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]).
        </p>
        <p>In the case of the noun phrase extracting, a perfect ROC Curve corresponds to
obtaining all relevant NP at the beginning of the list and all irrelevant NP at
the end of the list. This situation corresponds to AU C = 1.</p>
        <p>The diagonal corresponds to the performance of a random system, progress
of the rate of true positives being accompanied by an equivalent degradation of
the rate of false positives. This situation corresponds to AU C = 0:5.</p>
        <p>A human expert have manually evaluated the list of extracted NP, from 7
e-mails (i.e. approximately 210 NP).</p>
        <p>ROC curves indicate that the favorable titling methods are TALL (0.77) and
TSUM (0.69) (see Table 2). The score of TALL (i.e. NP extracted on the whole
text) seems to give better results than TSUM . With TMAX , the choice of the
title among the NP candidate is irrelevant for e-mails.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Human evaluation</title>
        <p>The experiments have been run on personal e-mails. Twenty e-mails were
selected. Texts are variable in size (i.e. number of words), topics, technicality, and
e ort of writing. Evaluation results are presented in Table 3. The expert had to
tag " " or "+" all the titles proposed with our system. The + symbol indicates
that the title given by the method (i.e. TMAX , TF REQ, TSUM , TALL) is relevant,
and indicates a title as irrelevant.</p>
        <p>Titling with TMAX does not o er good results (9/20) perhaps because of the
rarity/speci city of the terms of the title. Moreover, it could be interesting to
evaluate this method on speci c e-mails, for example on e-mails sent between
specialists of a same domain.</p>
        <p>Titles determinated by TSUM are relevant (12/20). However, the results show
that any titles are irrelevant, and thus that it is possible that the titles were not
found in the rst two sentences.</p>
        <p>Finally, TALL obtains a high score (16/20), that indicates a real interest to
extract the NP in the whole text, with the condition of use their position. In order
to see if this condition is really necessary, we have evaluated the TF REQ method.
This one is identical in TALL, but without the consideration of ScoreT F IDF in
the nal NP score. TF REQ obtains a bad result (8/20). This result justi es the
use of the position score called ScoreP (see Section 3.2).</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>We set up a method that enables to combine the NP position importance in
e-mails and its semantic content.</p>
      <p>Statistic study shows that it is necessary to use all the sentences of the e-mail
in order to propose a relevant title. The method TALL seems to be adapted to
e-mails titling.</p>
      <p>The quality of automatically computed titles strongly depends on the care
brought to the text writing. Nevertheless, the TALL method2 proposes relevant
titles for e-mails. The results show all the same that improvements can be brought.
Even if a part of the performance of this approach depends on Tree Tagger, it
seems possible to improve results. In particular, it could be interesting to give
more importance to Named Entities using TALL approach.</p>
      <p>The evaluation tends to indicate a possible bene t of an automatic method.
This one enables a time saving procedure for an e-mail writer... Then, the
proposed title makes possible a relevant indexing process of personal data as e-mails.
2 Available on the address
http : ==www:lirmm:f r=</p>
      <p>lopez=T itrage general=T iM ail:php</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Baxendale</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Man-made index for technical literature - an experiment</article-title>
          .
          <source>IBM Journal of Research</source>
          and Development pp.
          <volume>354</volume>
          {
          <issue>361</issue>
          (
          <year>1958</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Belhaoues</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          : Titrage automatique de pages web.
          <source>Master Thesis</source>
          , University Montpellier II, France (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Daille</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Study and implementation of combined techniques for automatic extraction of terminology</article-title>
          .
          <source>The Balancing Act : Combining Symbolic and Statistical</source>
          Approaches to language pp.
          <volume>29</volume>
          {
          <issue>36</issue>
          (
          <year>1996</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Ferri</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Flach</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hernandez-Orallo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Learning decision trees using the area under the ROC curve</article-title>
          .
          <source>In: Proceedings of ICML'02</source>
          . pp.
          <volume>139</volume>
          {
          <issue>146</issue>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Goldsteiny</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kantrowitz</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mittal</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carbonelly</surname>
          </string-name>
          , J.:
          <article-title>Summarizing text documents: Sentence selection and evaluation metrics</article-title>
          . pp.
          <volume>121</volume>
          {
          <issue>128</issue>
          (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Ho-Dac</surname>
            ,
            <given-names>L.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jacques</surname>
            ,
            <given-names>M.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rebeyrolle</surname>
          </string-name>
          , J.:
          <article-title>Sur la fonction discursive des titres</article-title>
          . S. Porhiel and
          <string-name>
            <given-names>D.</given-names>
            <surname>Klingler</surname>
          </string-name>
          (Eds). L'unit texte, Pleyben, Perspectives. pp.
          <volume>125</volume>
          {
          <issue>152</issue>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Jacques</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rebeyrolle</surname>
          </string-name>
          , J.:
          <article-title>Titres et structuration des documents</article-title>
          .
          <source>Actes International Symposium: Discourse and Document</source>
          pp.
          <volume>125</volume>
          {
          <issue>152</issue>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Lopez</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prince</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roche</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Text titling application (demonstration session</article-title>
          , to appear).
          <source>In: Proceedings of EKAW'10</source>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Lopez</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prince</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roche</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Titrage automatique de documents electroniques par extraction de syntagmes nominaux</article-title>
          .
          <source>In: Acte des 21emes Journees Francophones d'Ingenierie des Connaissances</source>
          . pp.
          <volume>17</volume>
          {
          <issue>28</issue>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Mitra</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buckley</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Singhal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cardi</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>An analysis of statistical and syntactic phrases</article-title>
          .
          <source>In: RIAO'</source>
          <year>1997</year>
          (
          <year>1997</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Salton</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buckley</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Term-weighting approaches in automatic text retrieval</article-title>
          .
          <source>Information Processing and Management</source>
          <volume>24</volume>
          p.
          <fpage>513</fpage>
          <lpage>523</lpage>
          (
          <year>1988</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Schmid</surname>
          </string-name>
          , H.:
          <article-title>Probabilistic part-of-speech tagging using decision trees</article-title>
          .
          <source>In: International Conference on New Methods in Language Processing</source>
          . pp.
          <volume>44</volume>
          {
          <issue>49</issue>
          (
          <year>1994</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Teufel</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moens</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Sentence extraction and rhetorical classi cation for exible abstracts</article-title>
          .
          <source>In: AAAI Spring Symposium on Intelligent Text Summarisation</source>
          . pp.
          <volume>16</volume>
          {
          <issue>25</issue>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Vinet</surname>
          </string-name>
          , M.T.:
          <article-title>L'aspet et la copule vide dans la grammaire des titres</article-title>
          .
          <source>Persee</source>
          <volume>100</volume>
          ,
          <issue>83</issue>
          {
          <fpage>101</fpage>
          (
          <year>1993</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gong</surname>
          </string-name>
          , Y.:
          <article-title>Multi-document summarization using sentence-based topic models</article-title>
          .
          <source>In: ACL-IJCNLP '09: Joint conference of the 47th Annual Meeting of the Association for Computational Linguistics and the 4th International Joint Conference on Natural Language Processing of the Asian Federation of Natural Language Processing</source>
          . pp.
          <volume>297</volume>
          {
          <issue>300</issue>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Yan</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dodier</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , Mozer,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Wolniewicz</surname>
          </string-name>
          , R.:
          <article-title>Optimizing classi er performance via an approximation to the Wilcoxon-Mann-Whitney statistic</article-title>
          .
          <source>In: Proceedings of ICML'03</source>
          . pp.
          <volume>848</volume>
          {
          <issue>855</issue>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Yous -Monod</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prince</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Sentence compression as a step in summarization or an alternative path in text shortening</article-title>
          .
          <source>In: Coling'08: International Conference on Computational Linguistics</source>
          , Manchester, UK. pp.
          <volume>139</volume>
          {
          <issue>142</issue>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Zajic</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Door</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schwarz</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          :
          <article-title>Automatic headline generation for newspaper stories</article-title>
          . Workshop on Text Summarization (
          <article-title>ACL 2002</article-title>
          and
          <article-title>DUC 2002 meeting on Text Summarization)</article-title>
          . Philadelphia. (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hovy</surname>
          </string-name>
          , E.:
          <article-title>Headline summarization at isi</article-title>
          .
          <source>In: Document Understanding Conference (DUC-2003)</source>
          , Edmonton, Alberta, Canada. (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>