<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Detecting Automatically Generated Sentences with Grammatical Structure Similarity</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nguyen Minh Tien</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Cyril Labbe</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Univ. Grenoble Alpes</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Grenoble INP ??</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Grenoble</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <fpage>73</fpage>
      <lpage>84</lpage>
      <abstract>
        <p>Detection of automatically generated papers has been a new eld of research. However, all current approaches are working at the document level and are unable to detect a small amount of generated text inside a large body of genuine written text. This paper will present the Grammatical Structure Similarity (GSS) measurement to detect sentences or short fragments from known generators. The proposed approach is tested against common machine learning methods, the ability to detect a modi ed generator is also tested.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The rest of the paper is organized as follows: Section 2 shows some current approaches to detect
automatically generated papers and for using parse trees with di erent goals, while Section 3 gives
a deeper understanding of PCFG and how a parse tree might be a good approach for our need.
Section 4 shows our method and the results of di erent tests. From the test results previously
obtained, Section 5 describes our system and in Section 6 a test is performed to compare di erent
machine learning approaches and our proposal to detect automatically generated sentences as well
as the possibility of detecting a modi ed generator.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Detecting Automatically Generated Paper and Parse Tree Usage on</title>
    </sec>
    <sec id="sec-3">
      <title>Sentence Similarity</title>
      <p>In this paper, we are interested in detecting and measuring the similarity between sentences as a
means to identify speci c ones as being automatically generated using PCFG.
2.1</p>
      <sec id="sec-3-1">
        <title>Detecting Automatically Generated Paper</title>
        <p>
          Over the years, some questionable events have surfaced such as [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]'s nonsense paper was accepted
to more than 150 journals or when Ike [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] became one of the most highly cited author on Google
Scholar despite the fact that all his \research" was automatically generated. Even within well
know publishers such as IEEE and Springer, more than 120 nonsense automatic-generated articles
have been found and retracted [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. These nonsense papers were generated using Natural language
generation (NLG) more speci cally Probabilistic Context Free Grammar (PCFG). Even though
this type of automatically generated paper is easy to be detected by an experienced human reader,
to the eyes of the general public they appear to have proper sentences and structure comparable
to a normal scienti c paper. There are multiple approaches to detect such type of papers using
di erent characteristics. [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] makes use of references to make a decision based on whether or not
those references are properly cached. [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] uses an ad-hoc similarity measure with custom weight
for sections, keywords, and references. Also [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ], with textual distance, achieves a very good result
comparable to other current approaches.
        </p>
        <p>
          Recently, [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] demonstrates the possibility of using similarity search to detect automatically
generated papers on a dataset of 43k genuine and 110 SCIgen1 papers where 10 SCIgen papers
are used as search seeds. They propose a pseudo-relevance feedback method where the returns of
a search query are reused as new search seeds. This results in a very promising accomplishment
with 0.96 precision and 0.99 recall. Also, [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] makes use of complex networks to obtain a Scigen
discrimination rate of at least 89%. They model the texts as complex networks with edges and
vertexes. These networks are then used with di erent machine learning methods to show that there
are hidden patterns in SCIgen papers that di er from real texts.
        </p>
        <p>However, all of these methods are working at the document level and are unable to detect a
small quantity of generated text inside a large body of genuine written text, thus making them
easier to be deceived. That was the reason the parse trees are investigated as a means to determine
the similarity between sentences.
2.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Using Dependency/Parse Tree To Measure Sentence Similarity</title>
        <p>A dependency tree or parse tree is a tree that represents the syntactic structure of a sentence
or a phrase (Example 21 and 22). Each sentence can be separated into Verb Phrase (VP), Noun
Phrase (NP) and then deeper level as Noun, Verb, Adjective, etc.... This might make one think
that sentences with a similar structure and word pairs might be related to each other.</p>
        <p>
          Over the years, there have been multiple proposals using parse/dependency tree to discover
the similarity between sentences. For example [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] uses their sentence similarity measurement in
Exactus [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] to detect plagiarism. This measurement uses di erent characteristics from the sentence
including TF-IDF, IDF overlap and a syntactic similarity measurement where the syntactic links
between pairs of words from di erent sentences are measured. Using this method, they were able
to obtain the second highest score for the plagiarism detection track at the PAN workshop 2014.
        </p>
        <p>
          [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] proposes a method to estimate the similarity between sentences using the number of
common tree segments in their augmented parse tree. This augmented parse tree is created using a
normal parse tree structure, but each node is represented as a feature vector instead of an entity
in the sentence. The similarity score between two augmented trees is then de ned using the tree
kernels function which uses both the matching subsequence of the children of each node and the
compatibility between two feature vectors.
        </p>
        <sec id="sec-3-2-1">
          <title>1 http://pdos.csail.mit.edu/scigen/</title>
          <p>Example 21. A parse tree in di erent forms for the
phrase: \a novel framework for the development of
scater/gather I/O".
(ROOT(NP(NP(DT a) (NN novel)) (NP(NP(NN
system)) (PP(IN for) (NP(NP(DT the) (NN analysis))
(PP(IN of) (NP(NNP scater/gather) (NNP I/O))))))))
Example 22. A parse tree in di erent forms for the
phrase: \a novel system for the analysis of gigabit switches".
(ROOT(NP(NP(DT a) (NN novel)) (NP(NP(NN
system)) (PP(IN for) (NP(NP(DT the) (NN analysis))
(PP(IN of) (NP(JJ gigabit) (NNS switches))))))))
.</p>
          <p>
            The DLSITE-2 system [
            <xref ref-type="bibr" rid="ref18">18</xref>
            ] aims at determining the truth of a text fragment from another
text (textual entailment) using syntactic parse tree. In this work, sentences are selected based on
words with signi cant grammatical value likes nouns, verbs, adjectives, etc. Sentences with the same
or similar signi cant grammatically value words are parsed to syntactic trees and these trees are
compared to detect if one is contained by another. This information is then used to demonstrate
that it is possible to deduce textual entailment from a hypothesis using syntactic trees.
          </p>
          <p>
            Parse tree along with common words are also used by [
            <xref ref-type="bibr" rid="ref7">7</xref>
            ] to determine the similarity between
sentences. They propose a method to search for semantic relation based on the exploding of Parse
tree starting from pairs of similar words. For each pair of sentences, the syntactic dependency trees
are obtained using Stanford Parser[
            <xref ref-type="bibr" rid="ref10">10</xref>
            ], then the most signi cant terms of each sentence, like nouns
or verbs, are discovered. From those terms, the tree is explored by going up to the ancestors step
by step, until a connection is formed; this connection is used as a common subtree between the two
trees. The similarity between them is then calculated with the number of common nodes along with
custom weight for each tree.
          </p>
          <p>However, these methods might not t our particular need since they either focus on pairs of
common word or are too expensive when it is required to parse every single sentence. Thus section
4.1 shows our framework to detect automatically generated sentences.
3</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Probabilistic Context Free Grammar (PCFG)</title>
      <p>At the moment, sentences generated with RNN or with Markov chain might appear to be more
diverse but are not always guarantee to be grammatically correct. That is why the quality of text
generated with PCFG are of a higher quality and are less easily detectable by unskilled human.
5
This section gives a deeper explanation on how probabilistic context free grammar function and why
parse tree structure might be an e ective method to detect sentence generated from such method.
3.1</p>
      <sec id="sec-4-1">
        <title>PCFG</title>
        <p>
          The seminal generator SCIgen was the rst realization of a family of scienti c oriented text
generators: SCIgen-Physic2 focuses on physics (It has been built using the structure rules of SCIgen and
modifying a subset of the vocabulary), Mathgen3 deals with mathematics, and the Automatic SBIR
Proposal Generator 4 (Propgen in the following) focuses on grant proposal generation. These four
generators were originally developed as hoaxes whose aim was to expose \bogus" conferences or
meetings by submitting meaningless, automatically generated papers. These generators make use
of PCFG which is a set of rules for the arrangement of the whole paper as well as for individual
sections and sentences (example 31). The richness of generated texts depends on the generator but
is quite limited when compared to a real human written text in both structure and vocabulary [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ].
Example 31. simple rules to generate a sentence S:
        </p>
        <p>S ! The implications of SCI BUZZWORD ADJ SCI BUZZWORD NOUN have been far-reaching and pervasive.
SCI BUZZWORD ADJ ! relationalj compactj ubiquitousj linear-timej fuzzyj embeddedj etc...</p>
        <p>SCI BUZZWORD NOUN ! technologyj communicationj algorithmsj theoryj methodologiesj informationj etc...
Using the previous rule, the owing sentences can be generated:
- The implications of relational epistemologies have been far-reaching and pervasive.
- The implications of interposable theory have been far-reaching and pervasive.</p>
        <p>For example, 31 shows a simple rule which can be used to generate several simple sentences.
However, a sentence usually comes from a more complicated rule where not only the noun and
adjective were randomized but at a deeper level as shown in example 32.It is also give the possibility
to modify a generator to a di erent eld by changing the terminal words such as the case of
SCIgen-physic. So it might be hard to produce a fully comprehensive list of all possible sentences.
Nevertheless, these sentences should have somewhat similar parse tree structures because they were
generated using the same rule as seen in example 21 and 22. Thus we propose to use a sample
corpus of PCFG automatically generated sentences hoping to gather most, if not all of the available
parse tree.</p>
        <p>Example 32. parts of a more complicated rule for phrase P generation:</p>
        <p>P ! a novel SCI SYSTEM for the SCI ACT.</p>
        <p>SCI ACT ! SCI ACT A SCI THINGj SCI ACT A SCI THING that SCI EFFECT
SCI ACT A ! understanding ofj SCI ADJ uni cation of SCI THING andj SCI VERBION of
SCI SYSTEM ! algorithmj systemj frameworkj heuristicj applicationj methodologyj SCI APPROACH
SCI THING ! IPv4j IPv6 j telephonyj multi-processorsj compilers j semaphoresj RPCsj virtual machinesj etc...
SCI VERBION ! explorationj developmentj re nementj investigationj analysis j improvementj etc...
Using this rule, these phrases can be generated:
- a novel heuristic for the understanding of randomized algorithms
- a novel system for the typical uni cation of massive multiplayer online role-playing games and symmetric encryption
- a novel framework for the development of scatter/gather I/O
- a novel system for the analysis of gigabit switches
3.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Corpora</title>
        <p>Multiple text corpora are used for testing purpose and this section will give a detailed description
about them.</p>
        <sec id="sec-4-2-1">
          <title>2 https://bitbucket.org/birkenfeld/scigen-physics</title>
        </sec>
        <sec id="sec-4-2-2">
          <title>3 http://thatsmathematics.com/mathgen/</title>
        </sec>
        <sec id="sec-4-2-3">
          <title>4 http://www.nadovich.com/chris/randprop/</title>
          <p>PCFG Corpus: PCFG corpora of di erent sizes were used as samples and tried to learn the di erent
parse tree structures from the generators. Table 1 shows the correlation between the size of the
corpus (evenly distributed between four generators) with the number of sentence and the number
of distinct parse tree that can be obtained from those sentences.</p>
          <p>Even though the number of distinct sentences and trees increase steadily,it is possible that there
are only small variations between them and this will be tested in section 4.2 later on.
Test Corpus: This corpus is composed of 4 smaller corpora each contains 100 texts from an
automatic generator and a real corpus of 100 genuine human written texts which were selected at
random from di erent elds. This resulted to about 110k sentences was used as the test corpus.
This section shows the testing process, how data is handled and how the similarity is de ned.</p>
          <p>
            First pdf les are converted to plain text, then normalized (de-capitalize, remove numbers,
symbols, non conventional characters, etc.). Later these texts are separated into sentences, and
each sentence is parsed using Stanford Parser[
            <xref ref-type="bibr" rid="ref10">10</xref>
            ] to obtain a parse tree. Since in our case, the
keywords have little to no value in deciding the similarity between sentences, they are removed
from the structure, and only the nodes are kept (Example 41). These parse trees are then compared
to the PCFG corpus of known parse trees from pre-processing generated sentences using a recursive
loop to nd the biggest possible subtree match of the tree structure.
          </p>
          <p>Example 41. the parse tree in example 21 would be considered only as.</p>
          <p>(ROOT(NP(NP(DT) (NN)) (NP(NP(NN)) (PP(IN) (NP(NP(DT) (NN)) (PP (IN) (NP(NNP)
(NNP))))))))</p>
          <p>Once a similar structure is found, the similarity between them need to be quanti ed. So the
grammatical structure similarity is de ned as follows:
De nition 1. Grammatical structure similarity (GSS):</p>
          <p>Let NA be the number of node in the parse tree TA of sentence A, NB be the number of node in
the parse tree TB of sentence B, and NAB be the number of node in the biggest common subtree of
TA and TB. Then the Grammatical Structure Similarity between A and B is de ned as:</p>
          <p>GSS(A;B) = N2AN+ANBB</p>
        </sec>
      </sec>
      <sec id="sec-4-3">
        <title>Example 42. Grammatical Structure Similarity between Example 21 and Example 22</title>
        <p>GSS(E21=E22) = 129+1179 = 0:89</p>
        <p>In our proposal, the computation is quite expensive since each and every sentence needs to go
through the parser and then compared with all samples in the PCFG corpus. This will be explored
further in the following section.
Our hypothesis is that even though the number of distinct sentences as well as parse trees seem quite
numerous, most of them should also be somewhat similar to each other. Only a small proportion of
the sentences are di erent. To verify this, the maximum GSS (MGSS) of all sentences in the test
corpus for three di erent size PCFG corpora that were computed and presented in section 3 and
the results are shown in gure 2.</p>
        <p>De nition 2. Maximum Grammatical Structure Similarity (MGSS): For a sentence A in the test
corpus (CT ), The MGSS between A and the PCFG corpus (CP CF G) is:</p>
        <p>M GSS(A;CP CF G) = M ax(B2CP CF G)(GSSA;B)</p>
        <p>It is understandable when comparing the PCFG corpus of size 80 with the others. With more
sample in the PCFG corpus, it is possible to nd much more high GSS match for generated sentences.
However comparing the PCFG of size 160 and 320, it is di cult to see any signi cant di erence. This
suggests that the previous hypothesis is true. Even though the size of the PCFG sample corpus was
doubled, it did not double the match rate because most of the additional parse tree structure have
very little di erences with what has already been obtained in the smaller size corpus. Subsequently,
from now on, only the PCFG corpus of size 160 is used.</p>
        <p>Figure 2 also shows that for SCIgen, physgen and mathgen it is possible to nd more than 50%
of really high match (GSS higher than 0.9) as compared to less than 2% for genuine written paper.
Even though there is no clear separation for the score, it is easy to see that there are di erent bell
curves for genuine written and generated ones. The curves for generated sentences lean very heavily
toward the end of the histogram thus making them stand out.</p>
        <p>Furthermore, table 2 shows some examples of genuine written sentence with high GSS to other
sentences in the PCFG corpus. It can be seen that most of them are just common sentences that
are also appearing in the PCFG corpus. To deal with such problem, the context of the sentence is
taken into account, and this will be presented in section 5 .</p>
        <p>Genuine written sentence Sentence in PCFG corpus
our main contributions are as follows our main contributions are as follows
it is easy to see that it is easy to see that
the states of this network are the contributions of this work are as follows
the rest of the paper is organized as follows the rest of this paper is organized as follows
the remainder of the paper is organized as the rest of this paper is organized as follows
follows
the proof of the claim can be found in ap- useful survey of the subject can be found
pendix in
the interpretation of the walk is as follows the rest of the paper proceeds as follows</p>
        <p>Table 2: Some typical mistakes from using only GSS
Jaccard similarty GSS
1 1
1 1
0.4 0.89
0.88 1
0.8 1
0.58
PCFG corpus with 80 samples</p>
        <p>PCFG corpus with 160 samples</p>
        <p>PCFG corpus with 320 samples
As mentioned before, the cost for parsing is quite expensive, this raises the need to implement
a lter to reduce the number of sentences that need to be parsed. To do such task, the Jaccard
similarity were used (the number of common word over the total number of distinct word in two
sentences). Figure 3 shows the Jaccard similarity between sentences in the test corpus with the
sentences in the PCFG corpus that have highest GSS to them.</p>
        <p>As shown in Figure 3, the majority (more than 90%) of genuine written sentence have Jaccard
similarity less than 0.3 to sentences in the PCFG corpus, while it was only about 20% for other
types of generator (except about 40% for propgen). Subsequently, this would make 0.3 a good
candidate for a threshold to be used in the lter since it is possible to keep a large number of
"suspected generated" sentence while greatly reducing the number of "irrelevant" sentences.Even
though Jaccard similarity is able to lter around 90% of the sentences but at the same time 20% of
genuine written sentences were also marked, this would result in a large number of false positive.
However this lter signi cant reduces the computational cost since it is no longer required to parse
and compare each sentence with the whole PCFG corpus, only those that are similar to a generated
sentence.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>GSS System with Jaccard Filter and Sentence's Context</title>
      <p>Since the aim is to detect a small portion of automatically generated text, each sentence is considered
along with its context, which includes the direct previous and next sentence to balance out special
cases. Thus, for each sentence in the test that is longer than 5 words and less than 35 words, the
PCFG corpus is used to nd other sentences that have Jaccard similarity higher than 0.3. Then
the GSS between them are calculated to obtain the maximum result; the same process is repeated
for the previous and the next sentences. The nal GSS with context for the sentence is the average
GSS of itself along with its direct neighbours.</p>
      <p>The result for the Jaccard lter is shown in table 3. It can be seen that the lter seems to
serve its purpose. Even though on average there are more sentences in a genuine written paper,
only 20% of them pass the lter and need to be parsed as compared to 70% to more than 90%
of sentences in automatically generated papers. This greatly reduced the processing time required,
since, in reality, one would assume that an overwhelming number of sentence are genuinely written.
The result of the GSS system using the Jaccard lter and GSS with context is shown in Figure 4.
It shows that a majority (96.7%) of genuine written sentences which pass through the lter have
less than 0.5 GSS. On the other hand, for automatically generated sets of sentences if a threshold
is set at 0.5, it is possible to detect more than 90% for SCIgen, physgen and about 75% mathgen
and propgen.</p>
    </sec>
    <sec id="sec-6">
      <title>Comparison With Other Methods</title>
      <p>
        To evaluate the approach, R and Rtexttools package[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] are used; this package is used for supervised
learning and includes di erent learning algorithms. The PCFG corpora were converted to texts and
separated into sentences. Since it is also required to have a sample corpus of genuine written paper,
other real corpora of the same size to the counterpart PCFG corpora were chosen at random to be
used as genuine-sample-corpora. The test corpus was split into sentences, they are classi ed with
di erent methods using a document-term-matrix to obtain each a label as \generated" or \real".
It is understandable that for a truly impartial comparison, the information from the parse trees
should also be given to the classi ers. However, transforming a parse tree to a feature vector is
not a straightforward task and it may even be impossible in practice. If trees would be added as a
particular feature for learning algorithms, then one would have to de ne how to compute similarity
for this particular feature and this exactly what GSS is de ning.
      </p>
      <p>The results of these classi cations are shown in table 4, considering precision or recall can be
easily manipulated by using di erent size corpora so we decided to use false positive rate which is the
probability of a genuine written sentence marked as generated and vice versa for false negative rate
to fairly represent the results. This table shows that conventional machine learning methods might
not be appropriate to our need since the results vary from very bad (Glmnet, SLDA, Tree) where
most of the sentence were marked as \generated" to mediocre (Max entropy, boosting, bagging
random forest) where they marked about half of the genuine written sentences as \generated". For
the GSS only and GSS system, 0.5 was used as a threshold to determine neither or not if a sentence
is automatically generated. As seen in gure 2 with GSS only, this is not a good threshold for single
sentence however if the context is taken into account as in GSS system ( gure 4), a very promising
result are obtained with very few genuine sentences marked as \generated" (for corpus size 160
there were less than 200 sentences marked as automatically generated out of more than 22k genuine
written sentences) but still catch a good number of automatically generated ones.</p>
      <p>Furthermore, to verify the possibility of detecting a modi ed generator where only the terminal
terms or keywords were changed. Physgen which is only a version of SCIgen with all the \hot
keywords" switched from computer science to physic ones is used. For this test, a corpus of 40
SCIgen papers is used trying to detect physgen sentences along 100 physgen papers and 100 genuine
papers. As before, a genuine-sample-corpus of 40 genuine papers is also used to aid machine learning
techniques. The results are shown in the \40 SCIgen" column of table 4. As suspected, using GSS
only was able to catch most of the sentences from physgen with only samples from SCIgen (0.015
false negative rate); even if the context and Jaccard lter are used, GSS system is still able to nd
80% of them. This suggests that the GSS system would also be e ective against cases of newly
modi ed version of existing generators.
7</p>
    </sec>
    <sec id="sec-7">
      <title>Conclusion</title>
      <p>There is a need for automatic detection of automatically generated texts and even though current
approaches have reasonably good results, they all focus on the document level. So, in this paper
we have shown our GSS system which is capable of detecting sentences from known generators
with su cient sample, which has 80% positive detection rate and less than 1% false detection rate.
Furthermore, the system has been tested against some well-known machine learning techniques
to demonstrate that it is able to provide the best results. The possibility of detecting a modi ed
version of current generators is also veri ed with great success.</p>
      <p>However, against new automatic generators without samples or generators which use other
techniques such as Markov chains or RNN, the system is impractical. This calls for more research
algorithm
Glmnet
Maxentropy
SLDA
Boosting
Bagging
Random Forest
Tree
GSS only
GSS system
corpus size</p>
      <p>False positive rate</p>
      <p>False Negative rate
80</p>
      <p>160 320 40 SCIgen 80 160 320 40 SCIgen
in di erent aspects of the text, for instance, checking the meaning of the words based on their
context or styles of generated texts.</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgments</title>
      <p>This research was funded by Springer Nature. We would like to thank our colleagues in PCM
department of Springer Nature who provided us with valuable insights, expertise as well as test
data that greatly assisted our research.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Amancio</surname>
            ,
            <given-names>D.R.</given-names>
          </string-name>
          :
          <article-title>Comparing the topological properties of real and arti cially generated scienti c manuscripts</article-title>
          .
          <source>Scientometrics</source>
          <volume>105</volume>
          (
          <issue>3</issue>
          ),
          <volume>1763</volume>
          {1779 (Dec
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Beel</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gipp</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Academic search engine spam and google scholars resilience against it</article-title>
          .
          <source>Journal of Electronic Publishing (December</source>
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bohannon</surname>
          </string-name>
          , J.:
          <source>Who's afraid of peer review? Science</source>
          <volume>342</volume>
          (
          <issue>6154</issue>
          ),
          <volume>60</volume>
          {5 (Oct
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Collingwood</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jurka</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boydstun</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grossman</surname>
            , E., van Atteveldt,
            <given-names>W.</given-names>
          </string-name>
          :
          <article-title>Rtexttools: A supervised learning package for text classi cation</article-title>
          .
          <source>The R Journal</source>
          <volume>5</volume>
          (
          <issue>1</issue>
          ),
          <volume>6</volume>
          {
          <fpage>13</fpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Culotta</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sorensen</surname>
          </string-name>
          , J.:
          <article-title>Dependency tree kernels for relation extraction</article-title>
          .
          <source>In: Proceedings of the 42Nd Annual Meeting on Association for Computational Linguistics. ACL '04</source>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computational Linguistics, Stroudsburg, PA, USA (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Dalkilic</surname>
            ,
            <given-names>M.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clark</surname>
          </string-name>
          , W.T.,
          <string-name>
            <surname>Costello</surname>
            ,
            <given-names>J.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Radivojac</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Using compression to identify classes of inauthentic texts</article-title>
          .
          <source>In: Proc. of the 2006 SIAM Conf. on Data Mining</source>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Duran</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rodr</surname>
            <given-names>guez</given-names>
          </string-name>
          , J.,
          <string-name>
            <surname>Bravo</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Similarity of sentences through comparison of syntactic trees with pairs of similar words</article-title>
          .
          <source>In: Electrical Engineering, Computing Science and Automatic Control (CCE)</source>
          ,
          <year>2014</year>
          11th International Conference on. pp.
          <volume>1</volume>
          {
          <issue>6</issue>
          (Sept
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Fahrenberg</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Biondi</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corre</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jegourel</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kongsh</surname>
          </string-name>
          j, S.,
          <string-name>
            <surname>Legay</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Measuring global similarity between texts</article-title>
          . In: Second International Conference, SLSP. pp.
          <volume>220</volume>
          {
          <issue>232</issue>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Ginsparg</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <source>Automated screening: ArXiv screens spot fake papers Nature</source>
          -
          <volume>508</volume>
          (-
          <fpage>7494</fpage>
          ),
          <volume>44</volume>
          (Mar
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Klein</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manning</surname>
          </string-name>
          , C.D.:
          <article-title>Fast exact inference with a factored model for natural language parsing</article-title>
          .
          <source>In: In Advances in Neural Information Processing Systems 15 (NIPS</source>
          . pp.
          <volume>3</volume>
          {
          <fpage>10</fpage>
          . MIT Press (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Labbe</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Ike Antkare one of the great stars in the scienti c rmament</article-title>
          .
          <source>ISSI Newsletter</source>
          <volume>6</volume>
          (
          <issue>2</issue>
          ),
          <volume>48</volume>
          {
          <fpage>52</fpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Labbe</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Labbe</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Duplicate and fake publications in the scienti c literature: How many scigen papers in computer science</article-title>
          ?
          <source>Scientometrics</source>
          <volume>94</volume>
          (
          <issue>1</issue>
          ),
          <volume>379</volume>
          {396 (Jan
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Labbe</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Labbe</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Portet</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          : Detection of Computer-Generated Papers in Scienti c Literature, pp.
          <volume>123</volume>
          {
          <fpage>141</fpage>
          . Springer International Publishing (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Lavoie</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krishnamoorthy</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Algorithmic detection of computer generated text</article-title>
          .
          <source>arXiv preprint arXiv:1008.0706</source>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Nguyen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Labbe</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Engineering a tool to detect automatically generated papers</article-title>
          .
          <source>In: Proceedings of the Third Workshop on Bibliometric-enhanced Information Retrieval co-located with the 38th European Conference on Information Retrieval (ECIR</source>
          <year>2016</year>
          ). pp.
          <volume>54</volume>
          {
          <issue>62</issue>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Noorden</surname>
            ,
            <given-names>R.V.</given-names>
          </string-name>
          :
          <article-title>Publishers withdraw more than 120 gibberish papers</article-title>
          .
          <source>Nature News (Feb</source>
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Sochenkov</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zubarev</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tikhomirov</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smirnov</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shelmanov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suvorov</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Osipov</surname>
          </string-name>
          , G.:
          <article-title>Exactus like: Plagiarism detection in scienti c texts</article-title>
          .
          <source>In: European Conference on Information Retrieval</source>
          . pp.
          <volume>837</volume>
          {
          <issue>840</issue>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neumann</surname>
          </string-name>
          , G.:
          <article-title>Recognizing textual entailment using sentence similarity based on dependency tree skeletons</article-title>
          .
          <source>In: Proceedings of the ACL-PASCAL Workshop on Textual Entailment and Paraphrasing</source>
          . pp.
          <volume>36</volume>
          {
          <fpage>41</fpage>
          . RTE '
          <volume>07</volume>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computational Linguistics, Stroudsburg, PA, USA (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giles</surname>
            ,
            <given-names>C.L.</given-names>
          </string-name>
          :
          <article-title>On the use of similarity search to detect fake scienti c papers</article-title>
          .
          <source>In: Similarity Search and Applications - 8th International Conference</source>
          ,
          <string-name>
            <surname>SISAP</surname>
          </string-name>
          <year>2015</year>
          . pp.
          <volume>332</volume>
          {
          <issue>338</issue>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Xiong</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>An e ective method to identify machine automatically generated paper</article-title>
          .
          <source>In: Knowledge Engineering and Software Engineering</source>
          . pp.
          <volume>101</volume>
          {
          <issue>102</issue>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Zubarev</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sochenkov</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Using sentence similarity measure for plagiarism source retrieval</article-title>
          .
          <source>In: CLEF (Working Notes)</source>
          . pp.
          <volume>1027</volume>
          {
          <issue>1034</issue>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>