<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Hierarchical Text Classi cation for Supporting Educational Programs</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Qi Ju</string-name>
          <email>qi@disi.unitn.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chiara Ravagni?y</string-name>
          <email>chiara.ravagni@erickson.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alessandro Moschitti</string-name>
          <email>moschitti@disi.unitn.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giampiero Vaschetto?</string-name>
          <email>giampiero.vaschetto@erickson.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>?Centro Studi Erickson</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>DISI, University of Trento</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>More than two decades have passed since the rst design of the CONSTRUE system [2], a powerful rule-based model for the categorization of Reuters news. Nowadays, statistical approaches are well assessed and they allow for an easy design of text classi cation (TC) systems. Additionally, the Web has emphasized the need of approaches for digesting large amount of textual information and making it more easily accessible, e.g., thorough hierarchical taxonomies like Dmoz or Yahoo! categories. Surprisingly, automated approaches have not proved yet to be indispensable for such categorization processes. This suggests that the role of TC might be di erent from simply routing documents to di erent topical categories. In this paper, we provide evidence of the promising use of TC as a support for an interesting and high level human activity in the educational context. The latter refers to the selection and de nition of educational programs tailored on speci c needs of pupils, who sometime require particular attention and actions to solve their learning problems. TC in this context is exploited to automatically extract several aspects and properties from learning objects, i.e., didactic material, in terms of semantic labels. These can be used to organized the di erent pieces of material in speci c didactic program, which can address speci c de ciencies of pupils. The TC experiments, carried out with state-of-the-art algorithms and a small set of training data, show that automatic classi ers can easily derive labels like, didactic context, school matter, pupil di culties and educative solution type.</p>
      </abstract>
      <kwd-group>
        <kwd>hierarchical text classi cation</kwd>
        <kwd>information management applications</kwd>
        <kwd>e-learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The last two decades have seen an impressive development of methods for automated
text categorization (TC) [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. This has been mainly due to the combination of two
important factors: (i) the exponential development of the Web, requiring for e ective
methods of information access and management; and (ii) the enhancement in theory
and practice of machine learning methods, which constitute the bases of TC.
      </p>
      <p>Despite the success of the TC research, it is still not clear if such technology
should be devoted to the design of topical categorization systems as very famous
Web hierarchical categorization systems are currently manually maintained, e.g.,
Dmoz or Yahoo! categories. On the other hand, TC also regards the association of
semantic labels that go beyond the simple routing of information to the most
appropriate user feeds. Indeed, this kind of task inevitably su ers from errors in Recall
and/or in Precision. Di erent would be the approach and results, if the outcome
of the TC system were cooperatively used as a tool to organize the information in
di erent and creative ways. In this respect, TC would be seen as a tool similarly
to search engines, rather than an end-to-end system forced to demonstrate a very
high accuracy.</p>
      <p>In this paper, we report on our experience with the e-Value project, whose
aims are the reorganization or combination of educational materials in di erent
pedagogical contexts. The Erickson Research Centre has been cataloging a large set
of published educational materials in smaller units, according to the SCORM (2004)
standards, Shareable Content Object Reference Model1. These documents are used
for the creation of novel and speci c didactic product as follows: (i) school classes
are evaluated about target cognitive processes; (ii) processes in which pupils have
di culties are detected and recorded in a huge database (DB) of normative data
along with the results of its elaboration; (iii) The Decision Support System (DSS)
chooses the proper didactic material for the class according to the DB content.</p>
      <p>The above steps require: (a) to identify cognitive processes involved in pupils'
learning; (b) to divide the didactic materials in smaller parts (learning objects); and
(c) classify such objects according to their bibliographic characteristics and to the
cognitive processes involved, which depends on the user context (e.g., age, class,
special situations). An automatic classi er can be used for easing and speeding up
the last step. It can provide a rough classi cation, which can constitute the starting
point for the work of expert catalogers.</p>
      <p>The use of the classi er would reduce the cataloging costs, both in terms of
time and human resources. Indeed, any educational material, being part of a book,
article or best practice, needs to be read and evaluated by experts, before being
assigned to the proper categories; this process takes a huge amount of time. As an
alternative model, the classi er can perform a rst approximate categorization and
after, the experts can re ne it. The clear advantage is that materials pertaining
to a certain subject can be directly assigned to its experts (working in that eld),
thus improving the accuracy of classi cation and avoiding the burden to exchange
materials among the di erent experts.</p>
      <p>However, the above scenario could be realized only if the adopted multi-class
classi er (MCC) performed accurate hierarchical categorization. Given the novelty
of the intended taxonomy, it is not simple to predict if MCC can deploy the needed
accuracy. For this purpose, we have:
{ designed a new taxonomy that meets the organization needs of e-Value;
{ de ned an annotation procedure and produced an initial datasets of 122
documents, organized in 112 categories (of course the documents are repeated in the
hierarchy); and
{ implemented an MCC, which exploits state-of-the-art TC models such as,
Support Vector Machines, structured in binary at categorizers.</p>
      <p>The preliminary experiments on the overall hierarchy of 112 nodes show promising
results, ranging from a Micro-F1 of above 95% for the rst level to about 70% on
the whole hierarchy. This outcome is rather promising and enables future research
in the use of TC for the e cient implementation of educational programs.</p>
      <p>In the reminder of this paper, Section 2 describes the tackled task in more detail,
Section 3 reports on our results and Section 4 derives the nal conclusions.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Automatic Support to the e-Value project</title>
      <p>The main objective of the e-Value project is to design, develop and test a multimedia
platform (consisting of a set of web applications), which integrates the evaluation
of various learning abilities and the application of didactic processes. These can
bene t from automatic methods for classifying the didactic material used in such
1 http://www.adlnet.gov/capabilities/scorm/scorm-2004-4th
C1</p>
      <p>C2</p>
      <p>C3</p>
      <p>C4</p>
      <p>C11 C12 C13 C21 C22 C23 C24 C31 C32 C41 C42 C43 C44 C45
C121 C122 C123 C124 C231 C232 C233 C311 C312 C321 C322</p>
      <p>C2321 C2322 C2323 C2324
processes. The next sections describe the problem in more detail and suggest how
a TC system can be used in such context.
2.1</p>
      <p>e-Value Framework
The framework includes di erent interconnected processes:
{ standard evaluation procedures and dynamic assessment of learning abilities of
pupils;
{ collection of normative data, e.g., educational material and pupils' evaluations;
{ continuous data ow, i.e., the related database is continuously updated and the
normative data currently available is integrated and compared with the new
arriving data; and
{ qualitative and quantitative evaluation of the collected data.</p>
      <p>The educational material is used for de ning didactic products, which address
speci c action (intervention). It consists of books, CD-ROMs, collections of articles,
etc. The e-Value project aims at both using independently and jointly the materials
above.</p>
      <p>Designing an intervention often requires the use of units taken from several
books or CD-ROMs but including the entire sources is very ine ective, considering
that only some small parts will be used. To enable more exibility in the creation
of training programs, the material collections are divided into basic training units,
called learning objects, which can be reassembled in a exible way. This requires
to analyze the materials to be used in the interventions and selecting the portion
involved in the target cognitive processes.
2.2</p>
      <sec id="sec-2-1">
        <title>A framework use-case</title>
        <p>A use of the framework is illustrated by the following example. In a school context
some classes are evaluated with respect to targeted cognitive processes. The tests
may reveal that some of the pupils have di culties in certain processes. Thus,
the test results are recorded (building a large database of normative data) along
with some elaboration of them, i.e., basic data statistics. Then the DSS chooses
the proper didactic material for the class by proposing di erent material to pupils
requiring attention and quick intervention. For this purpose the educational team
need to:
{ identify every cognitive process that can be involved in learning. At the moment,
this has been restricted to mathematics and reading-writing (with linguistic
skills and metaphonetics);
!"#
!""#
!"$#
!"$"#
!"$$#
!"$0#
!"$/#
!"$1#
!"0#
!$#
!$"#
!$""#
!$"""#
!$""$#
!$""0#
!$"$#
!$"$"#
!$"$$#
!$"$0#
!$"$/#
!$"$1#
!$"0#
!$"0"#
!$"0$#
!$"00#
!$"/#
!$"/"#
!$"/$#
!$"/0#
!$"//#
!$"1#
!$"1"#
!$"1$#
!$"10#
!$"&gt;#
!$"&gt;"#
!$"&gt;$#
{ divide the didactic materials in smaller parts (learning objects). This because
the use of the entire books or CD-Rom would be unfeasible, considering that
just a few exercises need to be applied. Thus the whole material has to be
checked by experts to be subdivided in learning objects. The latter are then
used to design the formative o er, in place of the entire material, obtaining a
more personalized and individualized learning.
{ Categorize the materials according to their bibliographic characteristics and,
most importantly for the fruition of the materials, to features of the involved
cognitive processes, e.g., the age, class and special situations of the target pupils
etc.
{ Porting the material from paper or optical media to an electronic format (pdf
or swf) so that it can be reassembled online and o ine.</p>
        <p>In the last phase the application of an automatic classi er can provide signi cant
bene ts to the whole process as explained in the following section.
2.3</p>
      </sec>
      <sec id="sec-2-2">
        <title>Classi cation Task</title>
        <p>To meet the need of the e-Value project, we have de ned a new taxonomy as well
as the annotation procedure and initial datasets. Our hierarchical categorization
scheme is shown in Figure 1, whose more descriptive labels are reported in Table
1. The materials have to be classi ed according to four macro-categories, and then
divided into a structure of sub-categories of 4 levels. Each category is meaningful
for a correct description of the materials, from both administrative perspective
(e.g., in which educational context should be applied) and subject/cognitive process
viewpoint (e.g. Mathematics { Number { Lexical and semantic processes instead of
Mathematics { Basic processes of calculus { Numerical facts). The Macro-categories
are: C1 { School and class (referring to the ages 5 { 14); C2 { Subject/cognitive
process (referring to the subjects of mathematics, linguistics, phonetics,
readingwriting abilities); C3 { Pupils' situation (for the cases of special needs or particular
situations); and C4 { Type of material (or the normal didactic usage in the class,
or for pupils with special situation or greater di culties in the subject).</p>
        <p>Such automatic classi cation could improve the manual categorization costs, in
terms of both time and human resource. Each piece of educational material, being
part of a book, article or best practice, needs to be read and evaluated by experts,
before being assigned to the proper categories, and this process takes a huge amount
of time. Therefore, the use of an automatic classi er could signi cantly reduce the
time required to read and evaluate the materials. Of course, experts will need to read
part of the material in any case to re ne and validate the output of the classi er.
However, the materials pertaining to a certain subject can be directly routed to the
experts of such eld, thus improving the categorization accuracy.
The aim of our evaluation is to demonstrate that state-of-the-art TC methods can be
applied to learn hierarchical classi ers for our e-Value taxonomy. This task is made
complex by two di erent aspects: (i) in addition to topic labels such as, Euclidean
Geometry, Problem Solving or Geometric Transformation, the taxonomy also
contains semantic characterization such as Story Development or Story Understanding,
whose characterization using simple terms seems harder; and (ii) given the novelty
of the taxonomy, we could only produce a small dataset, which makes the learning
of classi cation functions more di cult. To deal with and analyze such problems,
we experimented with hierarchy subsets, de ned according to the hierarchy's levels,
ranging from 1 to 4 (the maximum depth of our hierarchy). The deeper the level,
the more di cult TC is.
One major drawback of machine learning and thus of TC based on it is the need
of training data, i.e., a set of documents manually classi ed into the referring
taxonomy. This data is di cult to nd and/or to produce as it requires human labor.
Given the novelty of our taxonomy de ned in Figure 1, no previous data was
available. Thus, we set an annotation procedure (with only one annotator) of the didactic
material available in the Erickson's database. We randomly selected 60 documents
and we classi ed each of them according to all the 112 nodes of the taxonomy. This
led to a dataset of 122 documents (repetitions are considered).</p>
        <p>
          We randomly divided the above data in training and test set by taking care that
for each document all its repetitions were all put either in the training or in the test
set. The training data was used to learn the set of 112 binary classi ers, one for each
category, following the one-vs-all schema. The output of the multi-class classi er is
the merged set of the individual binary classi er decisions. Although simple, this is
considered a state-of-the-art approach [
          <xref ref-type="bibr" rid="ref3 ref5">5, 3</xref>
          ]. We used default SVM parameters as the
small training data prevented to apply any reasonable parameterization approach.
We used a bag-of-term representation (string separated by space and punctuation)
without applying any feature selection, stop list or lemmatization. Although, we are
        </p>
        <p>?
78;=&lt;=
786;;&amp;
(a) third level
con dent that the latter may relevantly improves our models. We used the classical
log(T F ) IDF weighting scheme and normalized vectors.</p>
        <p>The performance is provided by means of Micro- and Macro-Average F1,
evaluated from our test data over all 112 categories. Additionally, the F1s of the binary
classi ers are reported. For measuring the performance of di erent hierarchical
levels, only the nodes up to the target level are considered, e.g., for the rst level, we
only measure the Micro/Macro F1 of C1, C2, C3 and C4.
3.2</p>
      </sec>
      <sec id="sec-2-3">
        <title>Results and Discussion</title>
        <p>Table 2 reports the performance on the rst level. We note that for each category
there are about 40 documents for training. These seem to be enough as the accuracy
of the individual categories as well as the overall Micro/Macro F1 is exceptionally
high. This is not completely surprising as most documents are repeated in the above
four categories.</p>
        <p>Table 3 illustrates the results for the second level. We note that when the training
documents are more than 20, very good results can be achieved. Low performance is
shown for C11 and C13, which are trained with less than 7 documents. Additionally,
they have only one test document, this means that their accuracy cannot really be
estimated. The situation of C31 is even worse as it has no test documents. In this
case, we do not report any accuracy in the related row. It should also be noted
that, since we use one-vs-all schema, the accuracy of C1,..,C4 is the same as before.
Thus, from now on, we will not report the accuracy of previously reported binary
classi ers.</p>
        <p>Table 4 shows the performance on levels 3 and 4. Again the few training
documents available for the classi ers prevent to achieve a reasonable F1. There are
some good cases such as C124 and C322 but also bad cases such as C122 and C123.
The latter two refer to Primaria Classe II and Primaria Classe III, respectively,
which have large overlap with the other classes, i.e., I, IV and V. For separating
such categories, the simple bag-of-words may not be enough.
4</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Conclusions</title>
      <p>In this paper, we have described an interesting and new semantic classi cation
problem in the context of the educational framework of the e-Value project. We
have de ned a new hierarchical taxonomy, which is promising for improving the
production cycle of educational systems. To test the feasibility of the approach, we
have also built a corpus annotated according to the above taxonomy. Such data
was used for training an MCC based on SVMs. The results show that when there
is a reasonable amount of training documents the classi ers can deploy remarkably
high accuracy. On the other hand, the F1 of lower level categories is highly a ected
by data scarceness. Some categories would probably require the de nition of more
expressive features to better model their separation.</p>
      <p>
        Possible solutions are also provided by previous work, which shows more
advanced TC models, e.g., [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], in which global dependencies between hierarchical nodes
are encoded in a gradient descendent learning approach. They experimented with
Reuters Volume 1 (RCV1) 2 on a subhierarchy only containing 34 nodes. Other
relevant work such as [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] uses a rather di erent datasets and a di erent idea of
dependencies based on the feature distributions over the linked categories. Finally,
[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] experiment with models similar to ours achieving state-of-the-art on RCV1.
      </p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgements</title>
      <p>The research described in this paper has been partially supported by the Italian
Project e-Value (PAT) and by the European Community's Seventh Framework
Programme (FP7/2007-2013) under the grants #231126: LivingKnowledge {
Facts, Opinions and Bias in Time, #247758: EternalS { Trustworthy Eternal
Systems via Evolving Software, Data and Knowledge, and #288024: LiMoSINe {
Linguistically Motivated Semantic aggregation engiNes.
2 trec.nist.gov/data/reuters/reuters.html</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Dumais</surname>
            ,
            <given-names>S.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
          </string-name>
          , H.:
          <article-title>Hierarchical classi cation of web content</article-title>
          . In: Belkin,
          <string-name>
            <given-names>N.J.</given-names>
            ,
            <surname>Ingwersen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Leong</surname>
          </string-name>
          , M.K. (eds.)
          <source>Proceedings of SIGIR-00, 23rd ACM International Conference on Research and Development in Information Retrieval</source>
          . pp.
          <volume>256</volume>
          {
          <fpage>263</fpage>
          . ACM Press, New York, US, Athens, GR (
          <year>2000</year>
          ), http://research.microsoft.com/ ~sdumais/sigir00.pdf
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Hayes</surname>
            ,
            <given-names>P.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weinstein</surname>
            ,
            <given-names>S.P.</given-names>
          </string-name>
          :
          <article-title>Construe/Tis: a system for content-based indexing of a database of news stories</article-title>
          . In: Rappaport,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <surname>R</surname>
          </string-name>
          . (eds.)
          <source>Proceedings of IAAI-90, 2nd Conference on Innovative Applications of Arti cial Intelligence</source>
          . pp.
          <volume>49</volume>
          {
          <fpage>66</fpage>
          . AAAI Press, Menlo Park, US (
          <year>1990</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Lewis</surname>
            ,
            <given-names>D.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rose</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Rcv1: A new benchmark collection for text categorization research</article-title>
          .
          <source>The Journal of Machine Learning Research (5)</source>
          ,
          <volume>361</volume>
          {
          <fpage>397</fpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>McCallum</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosenfeld</surname>
            , R., Mitchell,
            <given-names>T.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ng</surname>
            ,
            <given-names>A.Y.</given-names>
          </string-name>
          :
          <article-title>Improving text classi cation by shrinkage in a hierarchy of classes</article-title>
          . In: ICML. pp.
          <volume>359</volume>
          {
          <issue>367</issue>
          (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Rifkin</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klautau</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>In defense of one-vs-all classi cation</article-title>
          .
          <source>J. Mach. Learn. Res</source>
          .
          <volume>5</volume>
          ,
          <issue>101</issue>
          {
          <issue>141</issue>
          (
          <year>December 2004</year>
          ), http://dl.acm.org/citation.cfm?id=
          <volume>1005332</volume>
          .
          <fpage>1005336</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Rousu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Saunders</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Szedmak</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shawe-Taylor</surname>
          </string-name>
          , J.:
          <article-title>Kernel-based learning of hierarchical multilabel classi cation models</article-title>
          .
          <source>The Journal of Machine Learning Research (7)</source>
          ,
          <volume>1601</volume>
          {
          <fpage>1626</fpage>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Sebastiani</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <article-title>: Machine learning in automated text categorization</article-title>
          .
          <source>ACM Computing Surveys</source>
          <volume>34</volume>
          (
          <issue>1</issue>
          ),
          <volume>1</volume>
          {
          <fpage>47</fpage>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>