<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>On the Use of PU Learning for Quality Flaw Prediction in Wikipedia</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Edgardo Ferretti</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Donato Hernandez Fusilier</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rafael Guzman Cabrera</string-name>
          <email>guzmancg@ugto.mx</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Manuel Montes y Gomez</string-name>
          <email>mmontesg@inaoep.mx</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marcelo Errecalde</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paolo Rosso</string-name>
          <email>prosso@dsic.upv.es</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Departamento de Ciencias Computacionales, Instituto Nacional de Astrof sica</institution>
          ,
          <addr-line>Optica y Electronica (INAOE). Puebla</addr-line>
          ,
          <country country="MX">Mexico</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Departamento de Informatica, Universidad Nacional de San Luis (UNSL).</institution>
          <addr-line>San Luis</addr-line>
          ,
          <country country="AR">Argentina</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Division de Ingenier as, Campus Irapuato-Salamanca, Universidad de Guanajuato. Salamanca</institution>
          ,
          <addr-line>Guanajuato</addr-line>
          ,
          <country country="MX">Mexico</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>NLE Lab - ELiRF, Universidad Politecnica de Valencia (UPV).</institution>
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this article we describe a new approach to assess Quality Flaw Prediction in Wikipedia. The partially supervised method studied, called PU Learning, has been successfully applied in classi cations tasks with traditional corpora like Reuters-21578 or 20-Newsgroups. To the best of our knowledge, this is the rst time that it is applied in this domain. Throughout this paper, we describe how the original PU Learning approach was evaluated for assessing quality aws and the modi cations introduced to get a quality aws predictor which obtained the best F1 scores in the task \Quality Flaw Prediction in Wikipedia" of the PAN challenge.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Given the daily increase in the amount of data on the Web, machine-based
assessment of Information Quality (IQ) is becoming a topic of enormous interest. This
fact is rooted, among others, in the increasing popularity of user-generated Web
content and the unavoidable divergence of the delivered content's quality [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. In
this respect, Wikipedia is a paradigmatic undertaking. This free-access
encyclopedia generated from among the content contributed by millions of users, has
this characteristic as main strength regarding its increased popularity.
Nonetheless, this feature is probably, the main challenge that Wikipedia faces on how to
systematically improve the quality of its articles.
      </p>
      <p>
        According to our literature review, there are three main research lines
related to IQ in Wikipedia, namely: (a) Featured articles identi cation [
        <xref ref-type="bibr" rid="ref10 ref12">10, 12</xref>
        ];
(b) Development of quality measurement metrics [
        <xref ref-type="bibr" rid="ref11 ref16">11, 16</xref>
        ]; and (c) Quality aws
detection [2{4]. It is clear that all the e orts made in improving IQ in Wikipedia
should be enhanced, nevertheless, as indicated by Anderka et al. in [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ], a rst
step towards automatic quality assurance in Wikipedia is detecting quality aws.
      </p>
      <p>
        In [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], it has been presented the rst complete breakdown of Wikipedia's
quality aw structure, which reveals the quality aws that actually exist, the
distribution of aws in Wikipedia, and the extent of awed content. It is
important to notice that the majority of quality aws are not caused due to malicious
intentions but stem from edits by inexperienced authors.
      </p>
      <p>
        In previous editions of the PAN challenge, assessing quality issues in
Wikipedia has been addressed in the form of vandalism detection. Given the context
above, in PAN@CLEF 2012,5 the vandalism detection task has been generalized
in focussing on the prediction of quality aws in Wikipedia articles. In particular,
the quality aws to be predicted are the ten most frequent quality aws of the
English Wikipedia articles, namely: Advert, Empty Section (Empty), No footnotes
(No-foot), Notability (Notab), Original research (OR), Orphan (Orph), Primary
sources (PS), Re mprove (Ref), Unreferenced (Unref) and Wikify (Wiki).
Besides, the task is formally de ned as follows: \Given a set of Wikipedia articles
that are tagged with a particular quality aw, decide whether an untagged
article su ers from this aw". That is to say, that detection of text quality aws is
cast as a one-class classi cation, as proposed in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        In our view, the most notable proposals to quality aw predictions in
Wikipedia have been made by Anderka et al. [2{4]. In [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], it is reported on the
exploratory analysis performed to target IQ aws, and also a one-class classi cation
technology for their identi cation is devised. The proposed method combines density
estimation with class probability estimation. The experimental results show that
certain aws can be detected with a nearly perfect precision, while for others
precision deteriorates signi cantly. In [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] it is performed a more in-depth
experimental analysis, where two settings are considered in deriving outlier examples:
an optimistic setting which uses featured articles 6 as outliers, and a pessimistic
setting that uses a random sample of documents not tagged as containing the
aw. Finally, in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], this idea is pushed further and previous work is extended
with: (a) a comprehensive breakdown of prior work on quality assessment, (b) an
in-depth discussion of the clean-up tag mining approach, (c) a description of the
quality aw model, and (d) a detailed analysis of the one-class problem.
      </p>
      <p>
        As mentioned above, all the work done in literature with respect to
quality aw prediction in Wikipedia, has been carried out following supervised
approaches. Despite the fact that very good results have been achieved in [
        <xref ref-type="bibr" rid="ref2 ref4">2, 4</xref>
        ] in
the so-called optimistic setting, when using untagged articles as outliers
(pessimistic setting ) the e ectiveness of aws predictions notably decrease. In this
way, motivated by [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], where several partially supervised learning techniques
are discussed and it is also shown their good performances in Web mining
applications, we decided to assess this task by means of a semi-supervised method.
After considering several alternatives we came to the decision of following the
5http://pan.webis.de/
6http://en.wikipedia.org/wiki/Wikipedia:Featured_article_criteria.
approach proposed by Liu et al. [
        <xref ref-type="bibr" rid="ref14 ref15">14, 15</xref>
        ]. This approach, called PU Learning is
explained next in Sect. 2. The key feature of this method is that it uses as input
a small labelled set of the positive class to be predicted and a large unlabelled
set to help learning. To the best of our knowledge, this is the rst time that this
method is used to predict information quality aws in Wikipedia.
      </p>
      <p>In Sect. 3, it is described in detail the research questions which guided the
development of our proposal to participate in PAN@CLEF 2012. Besides, in this
section it is also described a more intuitive rule-based approach to assess certain
quality aws. Then, Sect. 4 reports and discusses the results obtained in the
competition with our PU Learning approach and with the rule-based approach
as well. Finally, in Sect. 5 some general conclusions are drawn.
2</p>
    </sec>
    <sec id="sec-2">
      <title>PU Learning</title>
      <p>Text classi cation is an important problem which has numerous applications. As
pointed out by Liu et al.:7 \Although this classic model is important,8 in practice
one also encounters another problem. That is, one has a set of documents of a
particular topic or class P (positive class), and is given a large set U of mixed
(unlabelled) documents that contains documents from class P and also other
types of documents (negative documents). One wants to classify the documents
in U into documents from P and documents not from P. The key feature of
this problem is that there is no labeled negative training data, which makes the
traditional text classi cation techniques inapplicable. This problem is termed,
partially supervised classi cation (PSC). We also call it PU-learning (Learning
from Positive and Unlabelled examples)."</p>
      <p>
        In particular, given its simplicity and robust performance we decided to
implement the two-step strategy proposed in [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], which addresses the problem of
building two-class classi ers with only positive and unlabelled examples, but no
negative examples. This strategy is brie y described below and for extra details,
the interested reader should refer to [
        <xref ref-type="bibr" rid="ref14 ref15">14, 15</xref>
        ].
      </p>
      <p>Step 1: Identifying a set of reliable negative documents from the unlabelled set.
Step 2: Building a set of classi ers by iteratively applying a classi cation
algorithm and then selecting a good classi er from the set.</p>
      <sec id="sec-2-1">
        <title>7http://www.cs.uic.edu/~liub/NSF/PSC-IIS-0307239.html</title>
        <p>8Here, \this" refers to the supervised approach.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Experimental Setting and Preliminary Results</title>
      <p>It is well-known in Machine Learning research community that documents'
representation is a key issue. However, given that the team expertise is stronger in
the research eld of algorithms, we decided to use features already proposed in
the literature for modelling the documents and focussing our e orts in exploiting
as much as possible the characteristics of the PU learning approach described in
Sect. 2. There are four research questions which guided our experiments, namely:</p>
      <sec id="sec-3-1">
        <title>1. What is the best classi er in each stage?</title>
        <p>2. How to determine the sets of untagged documents for training the rst stage
classi er to improve its performance in selecting RNs?
3. After determining the RNs set, what documents should be used for training
the second stage classi er?
4. Which parameters setting of the second classi er improves its performance?</p>
        <p>
          These four questions are discussed next in below subsections. Regarding the
documents' representation, we used as a guide the work performed in [
          <xref ref-type="bibr" rid="ref4 ref8">4, 8</xref>
          ],
where they explore a signi cant number of quality features to assess the quality
of Wikipedia articles. These features are detailed in the Sect. 3.1. Given the
characteristics of some aws like Empty, No-foot, Ref and Unref, there is no need
to generate a complex document model to predict them. This is why, we also
devised a simpler rule-based approach based on parsing the articles' wikitexts
to nd particular patterns indicating the presence of these aws. This approach
is brie y described in Sect. 3.6.
3.1
        </p>
        <p>
          Documents Model
In [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], several features are conceptually grouped in three classes: Text Features
(those extracted from articles' textual content), Review features (those extracted
from articles' review history) and Network features (those extracted from the
social network inherent to the collection). Similarly, in [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], a four dimension
classi cation of features is devised. Our document model is composed by the
Text Length: character count, information-to-noise ratio, sentence count,
syllaFeatures bles count, one-syllable word count, word count; Structure: average
sentence length, average word length, average section length, average
subsection length, average subsubsection length, average sections nesting, average
subsections nesting, category count, external link count, le count,
heading count, image count, longest section length, longest subsection length,
longest subsubsection length, mandatory sections count, section count,
subsection count, subsubsection count, tables count, templates count, trivia
sections count, passive sentences rate, citation count, reference sections
count, shortest section length, shortest sentence length, shortest subsection
length, shortest subsubsection length; Style: Complex word rate,
Conjunction rate, Di cult word rate, Doubt word rate, Easy word rate, stop
words rate, longest sentence length, long sentences rate, long words rate,
average word syllables, one-syllable word rate, short sentences rate,
Peacock words rate, prepositions rate, pronouns rate, questions rate, \To be"
verb rate, Auxiliary verb rate, Weasel word rate, rate of sentences
beginning with: article coordinating conjunction, interrogative pronoun,
preposition, pronoun, subordinating conjunction; Readability: ARI, Bormuth,
Coleman-Liau, Dale-Chall, Flesch, Gunning-Fog, Kincaid, Lix, Miyazaki,
SMOG-Grading
Network In-link count, Internal link count, Inter-language link count
Features
features mentioned in Table 1, which are a subset of the ones used in [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. The
results reported in [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] show that textual features perform best and this is why
almost all of our features belong to this category. It is worth noticing that all the
features shown in Table 1 have been proposed by di erent authors [
          <xref ref-type="bibr" rid="ref12 ref16 ref4 ref6 ref8">4, 6, 8, 12, 16</xref>
          ]
and for a better understanding they have been organized as suggested in [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ].
3.2
        </p>
        <p>
          What Classi er in Each Stage?
As mentioned above, in [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] a benchmark system is proposed where a
comprehensive evaluation of sixteen combinations of classi ers for both steps is performed.
From this study it is shown that Support Vector Machines (SVM) variants
perform best as classi ers for the second step. Also it is corroborated that Nave
Bayes (NB) performs very well as rst stage classi er. In this way, based on this
evidence we decided to evaluate this combination rst. Moreover, given the
results reported in [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] where KNN is proposed as rst stage classi er, we decided
to study this technique as well. Consequently, several experiments were carried
out using NB, SVM and KNN as rst and second stage classi ers, respectively.
These experiments involved di erent corpora created from the PAN training
release. Similarly to the ndings in [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ], using NB and SVM as rst and second
stage classi ers, respectively, achieved very good results. Besides, NB + SVM
also presented a very good trade-o between running times and good results.
Thus, this combination was used in the remaining experimental setting.
As indicated in Sect. 3.2, several corpora were built from the PAN training
release. The main concern in building these corpora was studying how the sampling
strategy of untagged documents (U) could in uence the results obtained by our
approach. Instead of using a tenfold cross-validation approach as usual, splitting
the documents in U by our own gave us the possibility of having much more
control of the experimental setting, mainly on the issues related with determining
the proportions of positive vs. untagged documents in the training sets.
        </p>
        <p>To avoid a bias in how U ( U) was determined, 40 di erent samples were
selected to cover all the 50000 untagged documents in U. Figure 2 shows how these
40 di erent samples were obtained. Originally, U was split in 10 sub-sets Ui such
that jUij = 5000, for i = 1 : : : 10. Then, subsets Ui:1, were built such that: Ui:1 =
Ui + U(i mod 10)+1. Hence, jUi:1j = 10000 for all i = 1 : : : 10. Similarly, subsets
Ui:2, were obtained as: Ui:2 = Ui:1 +U((i+1) mod 10)+1. Thus, jUi:2j = 15000 for all
i = 1 : : : 10. Finally, subsets Ui:3, were built as Ui:3 = Ui:2 + U((i+2) mod 10)+1. In
this way, for all i = 1 : : : 10, jUi:3j = 20000. The idea of building these untagged
sets in an incremental way aims at analysing the e ect of increasing the
proportion of untagged documents versus positive documents in the training sets.
Larger proportions up to jU j = 50000 were tried but no improvements in the
results were achieved. Moreover, increasing the size of U also increases running
times, so 20000 was set as the upper amount of untagged samples to be used.</p>
        <p>Given that the number of positive sample documents for each aw was highly
unbalanced, for each aw it was also determined the minimum amount of positive
documents required to get the best results. For eight of the ten aws, the
number of positive documents in the training sets was set to 1000. Besides, several
proportions for the respective test sets were analysed. From these experiments
we decided to use for testing only 110 positive documents, since having more
of them in average resulted in similar performance rates. Flaws Advert and OR
contain 1109 and 507 documents, respectively. Hence, in order to have a uni ed
experimental test setting, it was also set to 110 the number of positive
documents to be used in the test sets of these aws. Hence, the number of positive
documents for training were 999 for Advert and 397 for OR, respectively.</p>
        <p>In this way, for each aw f it was xed a positive set Pf which was combined
with each of the 40 di erent subsets of U depicted in Fig. 2, thus yielding in
40 di erent training sets for the rst stage classi er. As explained in Sect. 2,
set Pf also comprise the positive sample of the training set of the second stage
classi er. In the following section, the di erent approaches used to determine
the negative training set of the second stage classi er, are explained.
3.4</p>
        <p>Strategies for Selecting Reliable Negatives
Four strategies were used for selecting the reliable negative documents (RNs) to
compose the training set of the second stage classi er, namely:</p>
      </sec>
      <sec id="sec-3-2">
        <title>1. Selecting all RNs as negative set.</title>
        <p>2. Selecting jPf j documents by random from RNs set.
3. Selecting the jPf j best RNs (those assigned the highest con dence prediction
values by the rst stage classi er).
4. Selecting the jPf j worst RNs (those assigned the lowest con dence prediction
values by the rst stage classi er).</p>
        <p>
          Strategy 1 is the original one proposed in [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] and was the rst one used in
our experiments. Testing our approach with positive samples only, we realised
that this strategy produces in average more false negatives (f n) than strategies
2 4. Table 2 reports the average, median, minimum and maximum f n values
for these strategies. Since the performance of the second stage classi er can only
be measured by considering its recall values, the average recall values over the
ten aws are also presented. As it can be observed, the maximum number of
f n for strategy 1 is 110, the actual number of positive samples in the test set.
Besides, the average number of f n for strategy 1 is close to the maximum f n
predictions for strategies 2 and 4, and it is higher than the maximum f n value
for strategy 3. This shows that having a highly unbalanced training set for the
second stage classi er a ects the performance of our approach.
        </p>
        <p>A statistical analysis (a non-parametric ANOVA) showed that the existing
di erences in the false negative rates of strategies 2 and 4, were signi cant when
compared against strategies 1 and 3, respectively. For strategy 3, the di erences
with strategy 1 were found not signi cant. Similarly, taking into account the
median recall values calculated on the ten aws when trained with the 40
different training sets described in Sect. 3.3, the statistical analysis determined as
signi cant the existing di erences between strategies 2 and 4, against 1 and 3.
Moreover, the mean rank di erences found between strategies 2 vs. 4, and 1 vs.
3, were determined not signi cant, respectively. Table 3 presents the average
recall values obtained per each aw on the 40 training sets. As it can be observed,
strategy 1 for ve out of the ten aws performs very poorly. Furthermore, when
considering the running times for each strategy, strategy 1 was found at least
three times slower than the other ones. Based on this evidence, we decided to
continue working with strategies 2 4.</p>
        <p>
          When compared against strategies 3 and 4, strategy 2 is conceptually the
simplest one, since it just selects at random jPf j RNs documents to make a
balanced training set for the second stage classi er. Conversely, strategy 3 selects
those documents assigned the highest con dence prediction values by the rst
classi er, on the grounds that they are better candidates in representing the
real negative documents' features. Finally, strategy 4 aims at selecting those
documents that in spite of being predicted as negatives, are still quite similar to
the positive ones. The underlying idea of this last strategy, is that selecting these
documents could help to build a much more ne-grained borderline between both
sets of documents. As shown in Table 2, strategies 2 and 4 perform best.
In Sect. 3.2, it was mentioned that using a SVM variant as second stage
classier reported the best results in [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] and in our experiments as well.
Following the suggestions of Chang and Lin [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ],9 the authors of the SVM
implementation we used, we tried all the di erent combinations of the parameters
2 f2 15; 2 13; 2 11; : : : ; 21; 23g and C 2 f2 5; 2 3; 2 1; : : : ; 213; 215g of the
RBF kernel which is used by default in this software package. We found that
combinations reporting the best results were those having a high penalty value
(C) for the error term and very low values, which allow reproducing in a
highsensitive way decision boundaries. In particular, in our experiments with the
training sets described in Sect. 3.3, C = 215 was found as the best penalty value
        </p>
        <sec id="sec-3-2-1">
          <title>9http://www.csie.ntu.edu.tw/~cjlin/papers/guide/guide.pdf</title>
          <p>for all the aws, while the best values are indicated in Table 4. It is worth
mentioning, that values presented in Tables 2 and 3 were obtained with the RBF
kernel, accordingly set with the parameters values reported in Table 4.</p>
          <p>During the training stage, we studied the performance of our algorithm using
the default kernel, i.e., RBF. When the PAN test set was released, the only clue
we had about it, was that it was balanced. That is to say, that for all the aws,
there were the same amount of positives and untagged documents. In this way,
approximately 50% of the documents was expected to be predicted as positives
by our algorithm. When we ran the di erent variants of our algorithm, in average,
for all the aws, they predicted as positive nearly 75% of the test documents.</p>
          <p>With the aim of improving the classi ers expected performance, the same
experimental setting carried out for the training set was also run for the test set.
Instead of studying the classi ers performance based on recall and f n measures,
they were studied with respect to their prediction rates. For each document in the
test sets, statistics were gathered considering if a particular classi er predicted
it or not as positive. For the aws Advert, Empty, No-foot, OR and Ref, most
of the classi ers agree on their predictions, while for the remaining aws the
classi ers shown di erent predicting behaviours. As this phenomenon could be
caused by an over- tting in the models learned from the training sets, therefore,
it was tried a more simple approach like a linear kernel instead of RBF. With this
kernel, in average, the number of documents predicted as positive was 62.55%,
a more balanced percentage than the one obtained for RBF.</p>
          <p>
            Based on these studies on the PAN test set, the linear SVM was also studied
as second stage classi er in our PU learning approach for the PAN training set.
For this particular kernel, we used the default parameters provided by WEKA [
            <xref ref-type="bibr" rid="ref9">9</xref>
            ].
Table 5 reports recall and f n values for RNs selection strategies for the linear
kernel. Conversely to the results presented in Table 2, strategy 3 is the best
performing for this kernel. We also evaluated the RBF and linear SVM classi ers
with positives plus untagged samples to reproduce the experimental conditions
of the PAN test set. From these experiments we noticed that strategy 3 tends
to predict as positives many untagged documents, while strategies 2 and 4 tend
to maintained their positive predictions rates.
          </p>
          <p>In this way, based on all the experiments performed with both kernels and
also considering the prediction statistics gathered for each document for the PAN
test set, we decided to use strategy 2 for RNs selection in nine out of the ten
aws. Flaw Orph, was the only one where strategy 4 was used. Regarding the
kernel selection for the second stage classi er, the linear kernel was used in eight
out of the ten aws. Flaws Advert and OR were the only ones where the RBF
kernel was used. In the Sect. 4 the results obtained with the PAN test set are
presented.
As mentioned above, for some aws like Empty, No-foot, Ref and Unref,
generating a complex document model to predict them is not necessary. In this
way, according to our experimental study, a document is predicted as having
the Empty aw when one of the following three conditions hold: the number of
empty sections is greater than zero; the number of subsections without content
is greater than zero or when the pattern \&lt; - - - : : : - - -&gt;" is present. Similarly,
the No-foot aw is predicted when at least one of the following conditions hold:
there are no external links (\==External link==" = 0); there are no in-text
citations (\http" = 0) or when expression \ref" is found less than 80 times.
Moreover, the Ref aw is predicted when: the number of external references is
less than 22; the regular expression (\ref &gt;") is found less than 65 times or when
reference section is empty. Finally, for Unref aw three conditions are used:
expression \ref" is found less than 45 times; the reference section does not exist or
there are no in-text citations (\http" = 0).</p>
          <p>
            In [
            <xref ref-type="bibr" rid="ref4">4</xref>
            ], a rule-based approach has also been proposed to predict aws Empty,
Orphan and Unref, in what they have called the \intensional modeling". Their
results are very accurate. It is worth noticing that their rules are applied on
particular features belonging to the document model also used by the supervised
approach they work with. In our case, our rule-based approach works with a
di erent document model than the one used by our PU Learning approach.
4
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>PAN Results</title>
      <p>
        Throughout this paper, we have mainly described the PU Learning approach
implemented to participate in the PAN challenge. Despite the fact that we
competed with this approach, as indicated in Sect. 3.6, also a rule-based approach
was developed to assess the prediction of some quality aws. As it can be
observed in Table 6, the results obtained with our rule-based approach are not very
encouraging. Nonetheless, we believe that we have failed in capturing the gist of
the wikitext patterns which characterize best these aws, and as suggested in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ],
a rule-based approach can be very useful in detecting these particular aws.
Regarding the results obtained with our PU Learning approach, they show a better
performance. In fact, as shown in Table 6, we got an average F1 score of 0.81.
Table 7 shows in row np, the amount of positive documents predicted per aw
to get these performance values in the challenge. Similarly, in row tn, the total
amounts of documents composing the test sets, are shown.
The use of a partially supervised method to predict quality aws in Wikipedia
has proven to be e ective. In this domain, our PU Learning approach which
implements several strategies for selecting RNs has outperformed the original
proposal which uses all the RNs found by the rst stage classi er. Strategies
2 and 4 achieved the best results. We expected that strategy 4 would perform
well. This is due to the fact that its underlying idea consists of building a
negrained borderline between both classes by selecting those documents that in
spite of being predicted as negatives, are still quite similar to the positive ones.
Likewise, in our view, strategy 2 achieved also very good results since by choosing
the RNs randomly it captures in a better way the implicit heterogeneity of the
documents not containing the aw.
      </p>
      <p>Also, it has been described the exhaustive experimental setting carried out to
set up as best as possible all the features of this approach, in order to participate
in the PAN challenge, where our proposal obtained the best F1 scores in the task
\Quality Flaw Prediction in Wikipedia". As future work, we think that exploring
other di erent semi-supervised techniques is a promising direction to improve
quality aw predictions in this free-access encyclopedia available to the entire
world.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgements</title>
      <p>Edgardo Ferretti and Marcelo Errecalde thank Universidad Nacional de San
Luis (PROICO 30310). The collaboration of UNSL, INAOE and UPV has been
funded by the European Commission as part of the WIQ-EI project (project
no. 269180) within the FP7 People Programme. Manuel Montes is partially
supported by CONACYT, No. 134186. The work of Paolo Rosso was carried out
also in the framework of the MICINN Text-Enterprise (TIN2009-13391-C04-03)
research project and the Microcluster VLC/Campus (International Campus of
Excellence) on Multimodal Intelligent Systems.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Anderka</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>A breakdown of quality aws in Wikipedia</article-title>
          . In: 2nd Joint WICOW/AIRWeb Workshop on Web quality. pp.
          <volume>11</volume>
          {
          <fpage>18</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Anderka</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lipka</surname>
          </string-name>
          , N.:
          <article-title>Detection of text quality aws as a one-class classi cation problem</article-title>
          .
          <source>In: 20th ACM International Conference on Information and Knowledge Management (CIKM'11)</source>
          . pp.
          <volume>2313</volume>
          {
          <fpage>2316</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Anderka</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lipka</surname>
          </string-name>
          , N.:
          <article-title>Towards Automatic Quality Assurance in Wikipedia</article-title>
          .
          <source>In: 20th International Conference on World Wide Web. ACM</source>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Anderka</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lipka</surname>
          </string-name>
          , N.:
          <article-title>Predicting Quality Flaws in User-generated Content: The Case of Wikipedia</article-title>
          .
          <source>In: 35rd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Baeza-Yates</surname>
            ,
            <given-names>R.:</given-names>
          </string-name>
          <article-title>User generated content: how good is it?</article-title>
          <source>In: 3rd Workshop on Information Credibility on the Web (WICOW'09)</source>
          . pp.
          <volume>1</volume>
          {
          <issue>2</issue>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Blumenstock</surname>
            ,
            <given-names>J.E.</given-names>
          </string-name>
          :
          <article-title>Size matters: word count as a measure of quality on Wikipedia</article-title>
          .
          <source>In: 17th Int'l Conference on World Wide Web. ACM</source>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <issue>7</issue>
          .
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>C.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>C.J.:</given-names>
          </string-name>
          <article-title>LIBSVM: A library for support vector machines</article-title>
          .
          <source>ACM Transactions on Intelligent Systems and Technology</source>
          <volume>2</volume>
          ,
          <issue>27</issue>
          :1{
          <fpage>27</fpage>
          :
          <fpage>27</fpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Dalip</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goncalves</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cristo</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Calado</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Automatic quality assessment of content created collaboratively by Web communities: a case study of Wikipedia</article-title>
          .
          <source>In: 9th ACM/IEEE-CS Joint Conference on Digital Libraries. ACM</source>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Hall</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Frank</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Holmes</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pfahringer</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reutemann</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Witten</surname>
            ,
            <given-names>I.H.</given-names>
          </string-name>
          :
          <article-title>The weka data mining software: An update</article-title>
          .
          <source>SIGKDD Explorations</source>
          <volume>11</volume>
          (
          <issue>1</issue>
          ) (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Lex</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          , Volske,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Errecalde</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Ferretti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Cagnina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Horn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Granitzer</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          :
          <article-title>Measuring the quality of web content using factual information</article-title>
          . In: 2nd joint WICOW/AIRWeb Workshop on Web quality.
          <source>ACM</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Lih</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Wikipedia as participatory journalism: reliable sources? Metrics for evaluating collaborative media as a news resource</article-title>
          .
          <source>In: 5th International Symposium on Online Journalism</source>
          . pp.
          <volume>16</volume>
          {
          <issue>17</issue>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Lipka</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Identifying featured articles in Wikipedia: writing style matters</article-title>
          .
          <source>In: 19th international Conference on World Wide Web. ACM</source>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Web Data Mining: Exploring Hyperlinks, Contents, and Usage Data</article-title>
          .
          <source>Data-Centric Systems and Applications</source>
          , Springer-Verlag,
          <year>2nd</year>
          edn. (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dai</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>W.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Building text classi ers using positive and unlabeled examples</article-title>
          .
          <source>In: 3rd IEEE Int'l Conference on Data Mining</source>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>W.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          :
          <article-title>Partially supervised classi cation of text documents</article-title>
          .
          <source>In: 19th International Conference on Machine Learning</source>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Stvilia</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Twidale</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gasser</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Assessing information quality of a community-based encyclopedia</article-title>
          .
          <source>In: 10th International Conference on Information Quality (ICIQ'05)</source>
          . pp.
          <volume>442</volume>
          {
          <fpage>454</fpage>
          .
          <string-name>
            <surname>MIT</surname>
          </string-name>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zuo</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          :
          <article-title>Reliable Negative Extracting Based on kNN for Learning from Positive and Unlabeled Examples</article-title>
          .
          <source>Journal of Computers</source>
          <volume>4</volume>
          (
          <issue>1</issue>
          ),
          <volume>94</volume>
          {
          <fpage>101</fpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>