<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview of the International Sexual Predator Identification Competition at PAN-2012</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Giacomo Inches</string-name>
          <email>giacomo.inches@usi.ch</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fabio Crestani</string-name>
          <email>fabio.crestani@usi.ch</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Faculty of Informatics University of Lugano (USI)</institution>
          ,
          <country country="CH">Switzerland</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This contribution presents the evaluation methodology for the identification of potential “sexual predators” in online conversations as part of PAN 2012. We provide details of the realized collection and analyse the submissions of the participants, who had to solve two problems: identify the predators among all the users in the different conversations and identify the part (the lines) of the predator conversations which are the most distinctive of the predator bad behaviour. The methods proposed by the 16 teams participating in the contest made possible the recognition of common pattern for predator identification (e.g. no preprocessing of the conversations, lexical and behavioral analysis, blacklisting of predator terms) as well as possible extension to existing systems (e.g. victimpredator distinction, pre-filtering of not relevant conversations).</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>“Chat messages” or “online/IRC conversations” are part of almost everybody’s
everyday life with services like Skype, Yahoo Messanger, MSN Messanger, ICQ but also
IRC networks like Freenode or Quakenet. Although these services facilitate the
establishment of new connections between persons or reinforce existing ones, they also allow
for misbehaviours or cybercriminal acts. The Sexual Predator Identification competition
ran for the first time in 2012 within PAN1 and aimed at providing researches with a
common framework to test methods for identifying such misbehaviours or cybercryminal
activities. For simplicity, in the competition we only concentrate on the identification of
“sexual predator” inside a chat, not dealing with other kind of misbehaviour or media.
A “sexual predator” is defined in the New Oxford American Dictionary as “a person
or group that ruthlessly exploits others” while Wikipedia noticed how the definition
“is used pejoratively to describe a person seen as obtaining or trying to obtain sexual
contact with another person in a metaphorically “predatory” manner”. We refer to these
interpretations of the term “sexual predator’ for the competition.</p>
      <p>
        In defining the tasks for the competition, we were also inspired by some previous
works [
        <xref ref-type="bibr" rid="ref12 ref16 ref9">12,9,16</xref>
        ] that addressed similar problem, even if none of them aimed at being an
evaluation laboratory or containing a challenging collection to be used as a reference. In
fact, we were the firsts to propose the following two kind of problems: given a collection
containing chat logs involving two (or more) persons the participants had to:
1 A benchmarking activity on uncovering plagiarism, authorship and social software misuse
http://pan.webis.de
1. identify the predators among all users in the different conversations (problem 1)
2. identify the part (the lines) of the conversations which are the most distinctive of
the predator behaviour (problem 2).
      </p>
      <p>We are presenting in Section 2 the details of the collection used in the competition
and in Section 3 the analysis of the methods employed by the participants to the task.
Finally, in Section 4 we are presenting the results of the competition, concluding with
Section 5.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Description of the Evaluation Framework</title>
      <p>In this section we present the corpus realized specifically for this competition and
highlight its properties and novelty with respect to existing collections. We also describe the
measures of performance used for evaluating the submission of the participants the two
problems of the task.
2.1</p>
      <sec id="sec-2-1">
        <title>Corpus</title>
        <p>
          In creating our collection we were animated by the same spirit of TREC and, more
recently, CLEF: we wanted to build a large collection that could serve as common
reference point for researchers of different fields (from Information Retrieval to Natural
Language Processing, from Text Mining to Machine Learning) and where they could
compare the performances of their different approaches. The realistic (large) size of the
collection is very important and is one of the central aspect of TREC tracks [20] and
PAN laboratories [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. It serves to fill the gap between the research and the industrial
application of the technologies developed in lab. For this reason we created a large
collection (hundred of thousands of conversations) with realistic properties: few number
of true positives (conversations with a potential “sexual predator”), large number of
false positives (people talking about sex or shared topic with the “sexual predator”)
and large number of false negatives (general conversations between users on different
topics). We believe that in a realistic scenario the percentage of “predator” conversations
with respect to the “regular” ones should be very low. In a different field (paedophile
queries in peer-to-peer system) the number of “predator” queries was found to be 0.25%
of the total [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. In our collection we therefore tried to respect that number but, in
order to make the identification of the predator a doable investigation, we increased the
percentage of one order of magnitude and set this to less than 4%.
        </p>
        <p>
          When looking for the “predator” collections, we found a common source for the
different datasets that were already used in the literature [
          <xref ref-type="bibr" rid="ref12 ref16 ref9">12,9,16</xref>
          ]: the http://www.
perverted-justice.com/ (PJ) website. This is a website where logs of online
conversations between convicted sexual predators and volunteers posing as underage
teenagers are published. The controversial creation and preliminary usage of these data
has been already discussed in [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], where the authors also give a detailed overview of
other collections [
          <xref ref-type="bibr" rid="ref16">16,21</xref>
          ], tools and approaches to cybercrime and online deception
detection. We therefore started with the PJ data for building our collection and kept in
mind the observations present in [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ], where two kinds (and different subkinds) of
suspicious interactions were identified: I) Predator/Other interaction, subdivided into: (Ia)
Predator/Victim (victim is underage); (Ib) Predator/Pseudo-Victim (volunteer posing as
child); (Ic) Predator/Pseudo-Victim (law enforcement officer posing as child) and II)
Adult/Adult (consensual relationship). Data of type (Ia) and (Ic) are difficult to obtain,
since it involves the police or law enforcement agencies in the process of data
acquisition. To our personal experience, police and law enforcement agencies are reluctant and
not very enthusiastic in collaborating on this sensible topic, therefore we ignored this
approach to data acquisition and focused on (Ib), which corresponds to the PJ data. PJ
data constitutes therefore our true positive set.
        </p>
        <p>Regarding interaction of type II) we found at first several online sources2 that could
had been of help but we later discarded them, because they were based on a single
person experience or were not of sufficiently large size (only some hundreds
conversations) to be successful employed in our collection. The documents present in the
Omegle repository3, to the contrary, served exactly our purpose. The original service
Omegle (where the documents come from) is a website that allows two strangers,
connected at the same time to the website, to have an anonymous online conversation. The
repository presents a random sample of more than 1 million original Omegle
conversations and by admission of the provider contains “abusive language and general silliness
online” and sometimes users “engage in cybersex” 4. The quantity of conversations as
well their nature and characteristics made this repository perfect to augment the level
of false positives in our collection, thus to make it more challenging and somehow real.</p>
        <p>The major difficulty that we encountered was in crawling “regular” online
conversations to complete the false negative set of documents and add a variety of topic of
discussion, to possibly hide an eventual general topicality of our true positive
conversations. We already mitigated the fact that the “predator” conversations are
between just two users by introducing the conversations extracted from Omegle, so now
we just needed to focus on topics about general discussions. To our surprise, the
Internet lacks of this kind of conversations: few people share their (private)
conversations online and the massive crawling of the public channels of the major IRC
networks5 is neither trivial nor encouraged6. We decided then to rely on those IRC logs
that included thousand of conversations and that were already made available on the
website of the IRC channel managers, namely http://www.irclog.org/ and
http://krijnhoetmer.nl/irc-logs/. Having a large volume of
conversations allowed us to increase the probability of having general discussions, interactions
between just few users and a variety of messages in length and duration, despite the
topical similarity between these conversations.</p>
        <p>A few other issues had to be solved in merging so different collections together. A
first problem occurred when deciding about the semantic definition of conversation. In
fact, we downloaded files from different sources of different formats, containing from
continuous logs of conversations on a daily basis to transcripts of unique conversations
2 See for example: http://www.oocities.org/urgrl21f/, http://www.fugly.</p>
        <p>com/victims/ or http://chatdump.com/
3 http://omegle.inportb.com/
4 See: http://inportb.com/2010/02/21/the-omeglean-society/
5 See: http://irc.netsplit.de/
6 See: http://wiki.vorratsdatenspeicherung.de/IRSeeK-en
of few lines, and we needed to combine them together in a single collection. To make
the conversations contained in the different files comparable, we decided to segment
all the messages exchange between the users in threads, where the cut was a break in
the messages exchange of more than 25 minutes. We empirically observed that this was
a reasonable threshold for a topic change in the conversation or the starting of a total
new one. After this step we obtained a consistent collection of hundred of thousand
conversations. We then noticed, by studying the length of the conversations, that the
wast majority (from 77% to 99% depending on the source) were below 150 messages
exchange. We therefore decided to include in the collection all the conversations that
were less or equal to 150 message exchange. Finally, we decided to generate an arbitrary
unique id for each conversation and also for each user and to replace nicknames within
each message with the corresponding user ids. Where possible we also substituted real
email addresses with arbitrary tags, in order to avoid the identification of real users.</p>
        <p>To the purpose of the competition we divided the collection into two parts, a training
one and a testing one. Given the fact that the training part is intended as “practicing”
rather than “training” as in Machine Learning, we decided to release 30% of the
collection as training set. In Table 1 we report the main properties of the whole collection.</p>
        <p>#conversations
#conv. length ≤150</p>
        <p>(% all )
#conv. length≤150
” and exactly 2 user</p>
        <p>(% training)
unique (perverted) users
#conv. length≤150
” and exactly 2 user</p>
        <p>(% testing)
unique (perverted) users
For the evaluation of the performance of the participants of the two problems, we
referred to the standard Information Retrieval measure of Precision (P), Recall (R) and F
(weighted harmonic mean between Precision and Recall):</p>
        <p>Precision (P) =
#(relevant items retrieved)
#(retrieved items)
(1)
Recall (R) =
#(relevant items retrieved)</p>
        <p>#(relevant items)
F =</p>
        <p>1
α P1 + (1 − α) R1 =
(β2 + 1)PR
β2P + R
where β2 =
1 − α
α
(2)
(3)</p>
        <p>
          The “items” retrieved are in one case (problem 1, identify the predators) the ids of
the authors considered perverted and in the second one (problem 2, identify the
predators’ lines) the line numbers considered indicative of a bad behaviour within a
conversation. We also noticed that, while the standard F measure equally weighted P and R with
β equal to 1, this is not always desired. In our case, in fact, for the first problem, despite
we observed that retrieving lot of relevant authors is important (Recall), to facilitate the
work of a police agent who would like to receive the largest number of suspect, what
is more important is the fact that the retrieved authors are relevant (Precision). This to
optimize the time of the police agent towards the “right” suspect rather than “all” the
possible suspects. For this reason we used a measure of F with the β factor equal to 0.5,
in order to emphasize Precision [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. For the second problem, instead, we observed that
retrieving lot of relevant lines (Recall) is more important than finding only the relevant
ones (Precision). Having lot of relevant lines, in fact, augments the possibility of finding
good evidences towards a suspect and for this reason we used a measure of F with the
β factor equal to 3, for emphasizing Recall [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ].
        </p>
        <p>It is also to be said that, while for the first problem the evaluation was quite
straightforward, having an a-priori indication of convicted perverted from the PJ website, for
the second it was harder (and more discussed) to define the ground truth. We decided
to adopt a TREC-like methodology for the evaluation and manually evaluated all
submitted lines by at least one participants (this accounted for 91% of all the predators’
lines). Given the particular nature of the task, that requires a particular training for the
evaluator in order to be able to distinguish between a predator chat and a regular chat,
this could not be done in a distributed way (e.g. mechanical turk). Moreover given the
limited time for the evaluation, we could not train other experts than us, thus relying
on the evaluation of a single expert in our group. For this reason, evaluations contain a
certain grade of subjectivity that we could not avoid. This is certainly a weak point in
this year competition that we will try to address better next year.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Overview of the Participants’ Approaches</title>
      <p>We received 16 submissions for the first problem (identifying the predators) and 14 for
the second problem (identifying the distinctive chat lines of the predator behaviour)
of the Sexual Predator Identification competition. Few users decide not to submit a
notebook paper to explain their used methods, therefore we are presenting an analysis
based on the 12 notebook paper received.
3.1</p>
      <sec id="sec-3-1">
        <title>Problem 1: identify predators</title>
        <p>
          Pre-filtering For the first problem, where the participants had to return a list of
potential predator, different pre-filtering techniques as well as classification methods have
been applied. The collection given to the participants was by design very unbalanced
(as most of them noticed) having few true positive authors (1% or less) in both training
and testing dataset and containing lot of false negatives that needed to be filtered out. A
common approach to overcome this problem was the use of a two stage classifier, where
in the first stage the classifier had to distinguish between conversation involving a
predator (true positive) and conversation without a predator (false negatives) [
          <xref ref-type="bibr" rid="ref13 ref15 ref6">19,13,15,6</xref>
          ]. In
addition to this, one of the most successful approaches [19] decides for the pre-filtering
of all the conversations that manifested some particular patterns: presence of 1
participants only, those with less then 6 interventions per user or those that contained 3 long
sequences of unrecognised characters. Similar attempts were done by other participants
but with a rule-based approach and on different features for different approaches [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ].
Features Apart from one case [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ], where participants used machine learning
approaches that work at character level (kernel with character 5-gram presence bit), in
all the others submissions we can divide the used features into two main categories:
“lexical” features and “behavioural” features. Lexical features are those that can be
derived from the raw text of the conversation: example of these features are unigram or
bigram [
          <xref ref-type="bibr" rid="ref13 ref14 ref4">19,13,4,14</xref>
          ], their weighting using TF-IDF or the cosine similarity and emoticons
counting. Other examples are the name recognition of the participants in the
conversation (self, other, group) [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] but also features obtained by the LIWC tool7 that calculates
the degree to which people use different categories of words across a wide array of
texts [
          <xref ref-type="bibr" rid="ref14">14,18</xref>
          ]. It is to be noted that, in general, lexical features have been used
without any stemming or stopword removal, to preserve each author own style, including
misspelling and grammatical errors.
        </p>
        <p>
          Behavioural are all those features that captures the “actions” of a user within a
conversation [
          <xref ref-type="bibr" rid="ref6">6,18</xref>
          ]: the number of times a user starts a dialogue, the response time after
a message of the partner in the conversation, the number of questions asked, the
frequency of turn-taking, intention (grooming, hooking, ...), etc. One of the most common
approach was the creation of a single set of features for each author, to be able to profile
him and exploit his predator potential. Some participants decided to build up not just the
Language Model (LM) of a single author, but also a LM as a combination of the LMs
of the two participants in the chat [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. Some other approaches were working, instead,
at a conversation or at a line level, therefore participants that used this strategies had
to aggregate the partial scores relative to all the lines or conversations of an author to
obtain a unique set of features for each author [
          <xref ref-type="bibr" rid="ref1 ref10 ref15 ref6 ref8">1,10,6,15,8</xref>
          ].
        </p>
        <p>
          Classification approaches In the classification step we could observe different
proposed method, but Support Vector Machines (SVM) were the most used [
          <xref ref-type="bibr" rid="ref13 ref14 ref15">13,14,15,19</xref>
          ].
In general, they were used in most cases for the first (predator-vs-all), then also for
the second step of the classification (predator-vs-victim). Sometime participants found
out that other solutions worked better than SVM, for example when they used a
Neural Network classifier [19]. Other classifier applied were based on Maximum-Entropy
[
          <xref ref-type="bibr" rid="ref4 ref8">4,8</xref>
          ], decision trees[
          <xref ref-type="bibr" rid="ref10">10</xref>
          ], k-NN [
          <xref ref-type="bibr" rid="ref17 ref7">7,17</xref>
          ] and/or random forest [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] as well as Naïve Bayes
[
          <xref ref-type="bibr" rid="ref1 ref6">6,1</xref>
          ]. In combination with the classifier sometimes we observed a filtering approach
based on a self-compiled dictionary of predatory terms.
7 http://www.liwc.net/
        </p>
        <p>To conclude, we should noticed that for this first problem we release a training set,
that allowed for supervised algorithms to be easily used. The situation was different for
the second problem, were no training data was available.
3.2
For this second problem, no training data was available for the participants. This was
intentionally done, mostly to test how participants approached the problem without
apriori relevance.</p>
        <p>
          The difficulty of the problem reduced the number of submissions (from 16 to 14)
and obliged the participants to use different approaches, compared with the supervised
ones of problem 1. The straightforward solution was to return as relevant all the
conversations lines of all the identified predators from the first problem [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. One of the
most used method was a filtering of all the predator conversations through a dictionary
of “perverted” terms or with a particular score (e.g. TF-IDF weighting) [
          <xref ref-type="bibr" rid="ref13 ref14 ref15 ref4">15,13,14,4</xref>
          ].
Similar to this approach, another first computed the LMs of the part of the conversation
considered predatory and then computed the differences between the actual
conversation and the LMs [19]. To conclude, the last approach was simply to return those lines
already labelled as predatory in the proposed algorithm by the default method for
problem 1 (working at line level) [
          <xref ref-type="bibr" rid="ref10 ref6 ref8">10,6,8</xref>
          ].
4
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Evaluation Results of the Participants’ Approaches and</title>
    </sec>
    <sec id="sec-5">
      <title>Discussion</title>
      <p>As reported in Table 1, participants received a training and a testing set, the first
containing 142 users labelled as predator, the second containing 254 predators to be discovered.
This was useful for the first problem (identify the predators), while for the second
problem (identify the lines manifesting the predators bad behaviour) we did not release any
training set. We wanted, in fact, to test how such a problem could be addressed without
any evidence. We later evaluated manually all the 113888 lines submitted by the
participants and identified 6478 that we considered expression of a predator bad behaviour.
In Table 2 and Table 3 we present the results for the first and second problem, with the
measures of evaluation explained in Section 2.2.
4.1</p>
      <sec id="sec-5-1">
        <title>Problem 1: identify predators</title>
        <p>If we analyse in details the results for the first problem, in particular the ranking in the
case of the two different metrics F with β = 1 and F with β = 0.5, we can notice that
only two positions swaps (1st and 5th) in case we consider one or the other measure of F.
This is due to the fact that we emphasised Precision with the F with β = 0.5. This choice
did not encountered the favour of all the participants, in fact some manifested their
disagreement and suggested giving more weight to Recall (thus, having a F measure
with β ≥ 2). In a real scenario, the proposed idea is to let the police agent decide
who is a predator and “manually” filter the results automatically obtained. Another
suggestion into this directions is the creation of a ranked list of suspects, that could
serve to prioritize the investigations.</p>
        <p>
          Besides this issues, from an operational point of view, it is interesting to notice
how important was the pre-filtering of unrelated conversations (at the cost of few true
positive) [19] and the similar use of lexical features in all the first ranked approaches:
bag-of-words with boolean weighting scheme [
          <xref ref-type="bibr" rid="ref13">13,19</xref>
          ], unigrams with TF-IDF
weighting scheme [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ], unigram and bigram [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. Participants also created a unique profile
for each author, by computing the features on an author-based file that collects all the
posts/messages of that author [
          <xref ref-type="bibr" rid="ref13 ref14 ref4">4,14,13</xref>
          ]. Behavioural/conversational features were, on
the other hand, used by all [
          <xref ref-type="bibr" rid="ref13 ref14 ref4">4,14,13</xref>
          ] except one [19] of the top-5 participants. These
last one [19] also choose to use a Neural Network classifier instead of SVM (in both
cases, two step classifiers) that were instead used by two others [
          <xref ref-type="bibr" rid="ref13 ref14">14,13</xref>
          ], while others
employed a Maximum-Entropy Classifier.
        </p>
        <p>Despite the similar features used and the relatively closeness of the performance
measures, the different classification strategies are a signal of still possible improvement
possibilities in the problem.
4.2</p>
      </sec>
      <sec id="sec-5-2">
        <title>Problem 2: identify predators’ lines</title>
        <p>As mentioned before, problem 2 was more difficult than problem 1 and presented more
open-issues than problem 1 too. Despite the suggestion of giving more weight to
Precision than to Recall, we should mention at least two issues that touched this part of the
competition. The first one is a certain dependency from the first problem: identifying
lines of the predator conversation requires at the beginning the correct identification of a
good number of predators. This might disadvantage participants that performed poorly
in the first part of the task. A solution to this problem might be having two stages for the
competition that corresponds to the two problems. The best result set of the first
problem could be used as a starting point for the second task. It has to be noticed, however,
that in the best-performers list (first-half of the ranking) we find also participants that
were not in the top-5 of the first problem. A preliminary explanation for this is that few
conversations of relatively few predators contributes to generate the ground truth for
the predators’ lines, therefore it is enough to identify such predators to obtain a good
score for problem 2. This fact leads to a second issue for problem 2, the creation of the
ground truth for the predators’ lines. At the beginning of the competition, there was no
ground truth for this second problem and we generated it on the basis of the received
submissions. We could have generated the ground truth by analysing all the predators’
conversation but by labelling only the submitted lines we spared 10% of all the
conversations and approximatively 1 week of work time. The real issue was determined by the
fact that one expert only labelled the lines of the conversation, leading to exclusion of
possibly relevant lines or the over-consideration of some others. We would have liked to
have more experts (at least 2 or more) for labelling the relevant lines in all the predator
conversations, but due to time and resource constraints that was not possible this year.
For a future edition of the Sexual Predator Identification task, we should plan more time
and resources for generating the ground truth and maybe we should consider the release
of a training set for this part of the problem as well.</p>
        <p>P</p>
        <p>R</p>
        <p>Official
Fβ=1 Fβ=0.5 run rank</p>
        <p>P</p>
        <p>R</p>
        <p>Official
Fβ=1 Fβ=3 run rank
We presented in this document the results of the first International Sexual Predator
Identification Competition at PAN-2012 within CLEF 2012. Given a realistic and
challenging collection containing chat logs involving two (or more) persons, the 16 participants
to the competition had to identify the predators among all the users in the different
conversations and identify the part (the lines) of the predator conversations which were the
most distinctive of the predator bad behaviour.</p>
        <p>For the first problem we can conclude that lexical and behavioural features should
be used when dealing with this kind of task. However, there is no unique method to
identify predators but different approaches could be used, from SVM to
MaximumEntropy algorithm. Having a pre-filtering step to prune irrelevant conversations seems
an important addition to the systems. For the second problem the most effective methods
appeared to be those based on filtering on a dictionary or LM basis, partly due to the
lack of ground truth for this specific problem (if we exclude the one based on 5-gram
characters presence bit). The identification of common set of features and a group of
effective strategies to identify predators is an achievement for this first part of the task.</p>
        <p>During the competition some issues were raised about the measurement of
performances for the two problems, whether we should emphasise Precision or Recall and
about the degree of subjectivity in the creation of the ground truth for problem 2. This
is an achievement, too: with this competition we wanted to give researchers a unique
place for comparing their methods but also for discussing and debating about future
directions on this research area.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>We want to thank all the PAN 2012 and CLEF 2012 organisers for their hard work and
support, as well as the participants of the competition for their patience and suggestions.</p>
      <p>
        Development of this competition was funded in part by the Swiss National Science
Foundation (SNF) project “Mining Conversational Content for Topic Modelling and
Author Identification (ChatMiner)” under grant number 200021_130208.
18. Vartapetiance, A., Gillam, L.: Quite simple approaches for authorship attribution, intrinsic
plagiarism detection and sexual predator identification - notebook for pan at clef 2012. In:
Forner et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]
19. Villatoro-Tello, E., Juárez-González, A., Escalante, H.J., Montes-Y-Gómez, M.,
Villaseñor-Pineda, L.: A two-step approach for effective detection of misbehaving users in
chats - notebook for pan at clef 2012. In: Forner et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]
20. Voorhees, E.M., Harman, D.K.: TREC: Experiment and Evaluation in Information
      </p>
      <p>Retrieval. Digital Libraries and Electronic Publishing, MIT Press (2005)
21. Yin, D., Xue, Z., Hong, L., Davison, B.D., Kontostathis, A., Edwards, L.: Detection of
Harassment on Web 2.0. In: CAW 2.0 ’09. Madrid, Spain (2009)</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Ayala</surname>
            ,
            <given-names>D.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Castillo</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pinto</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Olmos</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>León</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Information retrieval and classification based approaches for the sexual predator identification - notebook for pan at clef 2012</article-title>
          . In: Forner et al. [
          <volume>5</volume>
          ]
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Christopher</surname>
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Manning</surname>
            ,
            <given-names>P.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schütze</surname>
          </string-name>
          , H.: Introduction to Information Retrieval. Cambridge University Press (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Clough</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferro</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Forner</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gonzalo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huurnink</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kekäälå</surname>
            <given-names>inen</given-names>
          </string-name>
          , J.,
          <string-name>
            <surname>Lalmas</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Petras</surname>
          </string-name>
          , V.,
          <string-name>
            <surname>de Rijke</surname>
            ,
            <given-names>M.: CLEF</given-names>
          </string-name>
          <year>2011</year>
          .
          <source>ACM SIGIR Forum</source>
          <volume>45</volume>
          (
          <issue>2</issue>
          ),
          <volume>32</volume>
          (Jan
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Eriksson</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Karlgren</surname>
          </string-name>
          , J.:
          <article-title>Features for modelling characteristics of conversations- notebook for pan at clef 2012</article-title>
          . In: Forner et al. [
          <volume>5</volume>
          ]
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Forner</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Karlgren</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Womser-Hacker</surname>
            ,
            <given-names>C</given-names>
          </string-name>
          . (eds.):
          <article-title>CLEF 2012 Evaluation Labs</article-title>
          and Workshop - Working Notes Papers,
          <volume>17</volume>
          -20
          <source>September</source>
          <year>2012</year>
          , Rome, Italy (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Hidalgo</surname>
            ,
            <given-names>J.M.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Díaz</surname>
            ,
            <given-names>A.A.C.</given-names>
          </string-name>
          :
          <article-title>Combining predation heuristics and chat-like features in sexual predator identification - notebook for pan at clef 2012</article-title>
          . In: Forner et al. [
          <volume>5</volume>
          ]
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Kang</surname>
            ,
            <given-names>I.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>C.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kang</surname>
            ,
            <given-names>S.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Na</surname>
            ,
            <given-names>S.H.</given-names>
          </string-name>
          :
          <article-title>Ir-based k-nearest neighbor approach for identifying abnormal chat users - notebook for pan at clef 2012</article-title>
          . In: Forner et al. [
          <volume>5</volume>
          ]
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Kern</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klampfl</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zechner</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Vote/veto classification, ensemble clustering and sequence classification for author identification - notebook for pan at clef 2012</article-title>
          . In: Forner et al. [
          <volume>5</volume>
          ]
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Kontostathis</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Edwards</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leatherman</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Text mining and cybercrime</article-title>
          .
          <source>In: Text Mining</source>
          , pp.
          <fpage>149</fpage>
          -
          <lpage>164</lpage>
          . Wiley Online Library (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Kontostathis</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>West</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garron</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reynolds</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Edwards</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Identify predators using chatcoder 2.0 - notebook for pan at clef 2012</article-title>
          . In: Forner et al. [
          <volume>5</volume>
          ]
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Latapy</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Quantifying Paedophile Queries in a Large P2P System</article-title>
          . System pp.
          <fpage>401</fpage>
          -
          <lpage>405</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>McGhee</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bayzick</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kontostathis</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Edwards</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McBride</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jakubowski</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          :
          <article-title>Learning to Identify Internet Sexual Predation</article-title>
          .
          <source>International Journal of Electronic Commerce</source>
          <volume>15</volume>
          (
          <issue>3</issue>
          ),
          <fpage>103</fpage>
          -
          <lpage>122</lpage>
          (
          <year>Apr 2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Morris</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hirst</surname>
          </string-name>
          , G.:
          <article-title>Identifying sexual predators by svm classification with lexical and behavioral features - notebook for pan at clef 2012</article-title>
          . In: Forner et al. [
          <volume>5</volume>
          ]
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Parapar</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Losada</surname>
            ,
            <given-names>D.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barreiro</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>A learning-based approach for the identification of sexual predators in chat logs - notebook for pan at clef 2012</article-title>
          . In: Forner et al. [
          <volume>5</volume>
          ]
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Peersman</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vaassen</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Asch</surname>
            ,
            <given-names>V.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Daelemans</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          :
          <article-title>Conversation level constraints on pedophile detection in chat rooms - notebook for pan at clef 2012</article-title>
          . In: Forner et al. [
          <volume>5</volume>
          ]
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Pendar</surname>
          </string-name>
          , N.:
          <article-title>Toward Spotting the Pedophile Telling victim from predator in text chats</article-title>
          .
          <source>In: International Conference on Semantic Computing (ICSC</source>
          <year>2007</year>
          ). pp.
          <fpage>235</fpage>
          -
          <lpage>241</lpage>
          . No. c, IEEE (Sep
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Popescu</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grozea</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Kernel methods and string kernels for authorship analysis - notebook for pan at clef 2012</article-title>
          . In: Forner et al. [
          <volume>5</volume>
          ]
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>