<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>SINAI at INFILE 2009: Experiments with Google News</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Arturo Montejo-Raez, Jose M. Perea-Ortega, Manuel Carlos D az-Galiano, L. Alfonso Uren~a-Lopez SINAI research group. Computer Science Department. University of Jaen Campus Las Lagunillas</institution>
          ,
          <addr-line>Ed. A3, E-23071, Jaen</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes the SINAI team participation in the INFILE routing and ltering track of the CLEF campaign. This is the rst participation of the SINAI research group in the INFILE task. We have participated in the batch ltering subtask and submitted two experiments: one using the topics' text as learning data to train a classi er, and another one where training data has been constructed from Google News pages. Our results show that our use of Google News did not improved the classi cation obtained using only topics description.</p>
      </abstract>
      <kwd-group>
        <kwd>Text classi cation</kwd>
        <kwd>Information Retrieval</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        only topics texts were used a learning data. The learning algorithm was Support Vector Machines
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] on both cases.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Experiment Description</title>
      <p>A traditional supervised-learning scheme has been followed. The di erence between the runs
submitted rely on the training corpus used. One was on Google News entries and another on the
descriptions of the topics provided.</p>
      <p>The Google News corpus was generated by querying Google News on each of the topics
keywords. For example, topic 101 contains the keywords doping, legislation doping, athletes, doping
substances and ght against doping. For each of these keywords, 50 links were retrieved,
downloaded and their HTML cleaned out. In this way, about 200 documents existed per topic.</p>
      <p>Once each corpus was generated, a SVM model was trained on it. This is a binary classi er
turned into a multi-class classi er by training a di erent SVM model per topic. The topic with the
highest con dence was selected as label for the incoming document and, therefore, the document
was routed to that topic. It is important to note here that a label was proposed for all of the
incoming documents, that is, no document was left without one of the 50 labels (topics).
3</p>
    </sec>
    <sec id="sec-3">
      <title>Results and Discussion</title>
      <p>Overall results are displayed in Table 1. The results obtained are discouraging: few relevant
assignations are made. In fact, the use of Google News as learning source leads to very poor
results.</p>
      <p>Topics descriptions
Retrieved
Relevant
Relevant Retrieved</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusions and Further Work</title>
      <p>Google News as a source of information for generating learning corpus has shown quite bad results.
After inspecting these problematic corpus, we found that huge amount of useless text was not
ltered. Therefore, we plan to improve the quality of the data extracted from the web in order to
avoid undesirable side e ects due to noisy content.</p>
      <p>Although the results obtained in this task are really very low in terms of performance, it
represents a challenge in text mining, as real data has been used, compared to previous too
controlled corpora. We expect to continue our research on this data, and analyze in depth the
e ect of incorporating web content in ltering tasks.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgements</title>
      <p>This work has been supported by the Andalusian Regional Government (Spain) under excellence
project GeOasis (P08-41999), under project on Tourism (FFIEXP06-TU2301-2007/000024), the
Spanish Government under project Text-Mess TIMOM (TIN2006-15265-C06-03) and the local
project RFC/PP2008/UJA-08-16-14.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Romaric</given-names>
            <surname>Besan</surname>
          </string-name>
          <string-name>
            <surname>A</surname>
          </string-name>
          ~xon, Stephane Chaudiron, Djamel Mostefa, Ismail Timimi, and
          <string-name>
            <given-names>Khalid</given-names>
            <surname>Choukri</surname>
          </string-name>
          .
          <article-title>The in le project: a crosslingual ltering systems evaluation campaign</article-title>
          .
          <source>In Proceedings of the Sixth International Language Resources and Evaluation (LREC'08)</source>
          , Marrakech, Morocco, may
          <year>2008</year>
          .
          <article-title>European Language Resources Association (ELRA)</article-title>
          . http://www.lrecconf.org/proceedings/lrec2008/.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Francisco</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Couto</surname>
            , Bruno Martins, and
            <given-names>Mario J.</given-names>
          </string-name>
          <string-name>
            <surname>Silva</surname>
          </string-name>
          .
          <article-title>Classifying biological articles using web resources</article-title>
          .
          <source>In SAC '04: Proceedings of the 2004 ACM symposium on Applied computing</source>
          , pages
          <volume>111</volume>
          {
          <fpage>115</fpage>
          , New York, NY, USA,
          <year>2004</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>M.C.</surname>
          </string-name>
          <article-title>D az-</article-title>
          <string-name>
            <surname>Galiano</surname>
            ,
            <given-names>J.M.</given-names>
          </string-name>
          <string-name>
            <surname>Perea-Ortega</surname>
            ,
            <given-names>M.T.</given-names>
          </string-name>
          <string-name>
            <surname>Mart</surname>
            n-Valdivia,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Montejo-Raez</surname>
            , and
            <given-names>L.A.</given-names>
          </string-name>
          <article-title>Ure na Lopez. Sinai at trecvid 2007</article-title>
          . In Paul Over, editor,
          <source>Proceedings of the TREC Video Retrieval Evaluation 2007 (TRECVID'07)</source>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Risto</given-names>
            <surname>Gligorov</surname>
          </string-name>
          , Warner ten Kate, Zharko Aleksovski, and Frank van Harmelen.
          <article-title>Using google distance to weight approximate ontology matches</article-title>
          .
          <source>In WWW '07: Proceedings of the 16th international conference on World Wide Web</source>
          , pages
          <volume>767</volume>
          {
          <fpage>776</fpage>
          , New York, NY, USA,
          <year>2007</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Thorsten</given-names>
            <surname>Joachims</surname>
          </string-name>
          .
          <article-title>Text categorization with support vector machines: learning with many relevant features</article-title>
          .
          <source>In Claire Nedellec and Celine Rouveirol</source>
          , editors,
          <source>Proceedings of ECML-98, 10th European Conference on Machine Learning, number 1398</source>
          , pages
          <fpage>137</fpage>
          {
          <fpage>142</fpage>
          ,
          <string-name>
            <surname>Chemnitz</surname>
          </string-name>
          , DE,
          <year>1998</year>
          . Springer Verlag, Heidelberg, DE.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>