<?xml version="1.0" encoding="UTF-8"?>
<TEI xml:space="preserve" xmlns="http://www.tei-c.org/ns/1.0" 
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" 
xsi:schemaLocation="http://www.tei-c.org/ns/1.0 https://raw.githubusercontent.com/kermitt2/grobid/master/grobid-home/schemas/xsd/Grobid.xsd"
 xmlns:xlink="http://www.w3.org/1999/xlink">
	<teiHeader xml:lang="en">
		<fileDesc>
			<titleStmt>
				<title level="a" type="main">Common Criteria for Genre Classification: Annotation and Granularity 1st Author</title>
			</titleStmt>
			<publicationStmt>
				<publisher/>
				<availability status="unknown"><licence/></availability>
			</publicationStmt>
			<sourceDesc>
				<biblStruct>
					<analytic>
						<author>
							<persName><forename type="first">Marina</forename><surname>Santini</surname></persName>
						</author>
						<title level="a" type="main">Common Criteria for Genre Classification: Annotation and Granularity 1st Author</title>
					</analytic>
					<monogr>
						<imprint>
							<date/>
						</imprint>
					</monogr>
					<idno type="MD5">E9F48C889312EF8D209C31AD9A531DD4</idno>
				</biblStruct>
			</sourceDesc>
		</fileDesc>
		<encodingDesc>
			<appInfo>
				<application version="0.7.2" ident="GROBID" when="2023-03-25T02:15+0000">
					<desc>GROBID - A machine learning software for extracting information from scholarly documents</desc>
					<ref target="https://github.com/kermitt2/grobid"/>
				</application>
			</appInfo>
		</encodingDesc>
		<profileDesc>
			<abstract>
<div xmlns="http://www.tei-c.org/ns/1.0"><p>In this paper, 1 we present two experiments that use machine learning for automatically classifying web pages by genre. These experiments highlight the influence that genre annotation and genre granularity can have on the accuracy of the classification. From a practical point of view these experiments show that a collection annotated with the criteria of 'objective sources' and consistent genre granularity ensures a very good classification accuracy (Experiment 1). Additionally, the classification model built out of such a collection can be exported more profitably for predictive tasks on an unclassified web page collection (Experiment 2). These experiments represent a starting point for a discussion about the need of common criteria for building a genre collection in the absence of an official genre-annotated benchmark.</p></div>
			</abstract>
		</profileDesc>
	</teiHeader>
	<text xml:lang="en">
		<body>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1">INTRODUCTION</head><p>In this paper, we present two experiments that use machine learning for automatically classifying web pages by genre.</p><p>Many definitions of genre have been proposed so far in literary studies (e.g. <ref type="bibr" target="#b19">[20]</ref>), academic writing (e.g. <ref type="bibr" target="#b22">[23]</ref>), professional settings (e.g. <ref type="bibr" target="#b1">[2]</ref> and <ref type="bibr" target="#b23">[24]</ref>), organizational environment (e.g. <ref type="bibr" target="#b25">[26]</ref>), and so on. More specifically, in automatic genre classification studies, genres have often been seen as non-topical categories that could help reduce information overload (e.g. <ref type="bibr" target="#b15">[16]</ref> or <ref type="bibr" target="#b14">[15]</ref>). In this area, not only text categories such as 'article', 'FAQs', 'home page', etc. have been considered to be genres, but also polarities, such as subjective-objective and positive-negative ( <ref type="bibr" target="#b6">[7]</ref>), and style ( <ref type="bibr" target="#b0">[1]</ref>, <ref type="bibr" target="#b8">[9]</ref> and <ref type="bibr" target="#b4">[5]</ref>). Regardless the different definitions and connotations, a classification by genre has been acknowledged to be useful in information retrieval (e.g. <ref type="bibr" target="#b8">[9]</ref>, <ref type="bibr" target="#b11">[12]</ref>, etc.), information filtering ( <ref type="bibr" target="#b6">[7]</ref>), digital libraries ( <ref type="bibr" target="#b18">[19]</ref>) and other practical applications.</p><p>In this paper we present two experiments of genre classification of web pages based on a simplified and intuitive definition of genre, which is suitable for all kind of genresincluding genres on the web -and for an automatic approach. In our view, genres can be defined as named socio-cultural communication artefacts, linked to a society or a community, bearing standardized traits, leaving space for the creativity of the text producer, and raising expectations in the text receiver. For example, the personal home page (cf. also <ref type="bibr" target="#b5">[6]</ref>) has standard traits, such as self-narration, personal interests, contact details, and often pictures related to one's life. However, these conventions do not hinder the creativity of the producer, and as receivers, we expect a blend of standardized information and personal touch. Though unsophisticated, this definition of genre allows us to suggest a practical solution to the main shortcoming in genre classification, i.e. the lack of a genre-annotated benchmark. Because of this lack, the main tendency has always been to build one's own collection according to subjective criteria as for genre annotation and genre granularity. This is especially true for genre studies based on collections of web pages. Although building a genre-annotated benchmark of web pages is difficult and maybe not feasible, because annotating a web page by genre is both hard and controversial (cf. <ref type="bibr" target="#b20">[21]</ref>), a few criteria should be discussed and agreed upon. Without some kind of commonality, any comparison becomes unfeasible. For instance, can we state that the 92% accuracy achieved by <ref type="bibr" target="#b2">[3]</ref> is better than the accuracy (about 70%) achieved by <ref type="bibr" target="#b16">[17]</ref>? The solution we suggest for building more comparable genre collections is to exploit the socio-cultural aspect of the concept of genre. As pointed out earlier, genres have a function in a society, culture or community, i.e. they have a social or public role that implies a number of conventions and raises predictable expectations. This means that the role or the function of different genres is recognized and correctly used in the communication interaction. Leveraging on this public and collective acknowledgement it is possible to create a genreannotated collection without involving human annotators. The key is to download documents from genre-specific archives or portals and use their membership in these containers as an automatic membership in a specific genre. For example, eshops can be randomly downloaded from the portal http://www.eshops.co.uk/ and considered to be eshops without any further manual annotation or inter-rater agreement assessment. We include in the public acknowledgement also genres used as title of documents (for example, "Insects Hotlist"). The idea behind selecting documents with a genre in the title or picking them up randomly from public resources, such as an archives or a portals, is the following: if there is an archive, a portal or a website specialized in, say, pointing to or collecting genres such as eshops, blogs or search engines, this means that the documents pointed to or collected there are considered to belong to these genres by the collectivity of web users. We call this criterion 'annotation by objective sources'. A genre collection annotated by objective sources tends to be more representative as for intra-genre variation than a collection annotated relying on the genre stereotypicality that two, three, or more annotators have in mind. We suggest that annotating a collection using objective sources is faster and closer to real-world conditions.</p><p>Genre granularity is also important when building a collection for genre classification. In fact, genre palettes often show different levels of granularity. For instance, <ref type="bibr" target="#b8">[9]</ref> includes in his genre palette both FAQs (genre) and journalistic materials (super-genre). We suggest the use of the prototype theory (cf. <ref type="bibr" target="#b17">[18]</ref> and <ref type="bibr" target="#b12">[13]</ref>) to achieve a consistent level of genre granularity. A prototype is the most typical instance of a more encompassing or fuzzy category. Categories that can be dealt with the prototype theory can be ordered into a three-tiered hierarchy: superordinate level, basic level and subordinate level. For example, the genre 'advertisement' represents the basic level (genre) of the superordinate level 'advertising' (super-genre), while a 'web ad' represents the subordinate level (subgenre) of the basic level. The basic level embodies the information level at which concepts are most easily recognized, remembered and learned with respect to their function. The basic level included in the prototype theory should not be mixed up with document stereotypicality or exemplarity. Building a genre collection choosing exemplars, i.e. only stereotypical documents, to unambiguously represent a genre can return biased results. According to the prototype theory, instead, instances of a genre may vary in their prototypicality, thus allowing intra-genre variation.</p><p>The two experiments presented in this paper highlight the influence that genre annotation and genre granularity can have on the accuracy of genre classification of web pages. They were designed to point out several issues (some already covered in <ref type="bibr" target="#b21">[22]</ref>). In this paper, these two experiments allow us to emphasize two general aspects of genre classification, one practical and one theoretical. From a practical point of view these experiments show that a collection annotated with the criteria of objective sources and consistent genre granularity ensures a very good classification accuracy (Experiment 1). Additionally, the classification model built out of such a collection can be exported more profitably for predictive tasks on an unclassified web page collection (Experiment 2). From a theoretical point of view, they represent a starting point for a discussion about the need of common criteria in the absence of an official genre-annotated benchmark</p><p>In order to ensure replicability, all the materials used for these experiments, including web page collections, feature sets and the manual evaluation of Experiment 2, are available at http://www.nltg.brighton.ac.uk/home/Marina.Santini/, bottom of the page.</p><p>The paper is organized as follows: Section 2 provides an overview of recent work in genre classification of web pages; Section 3 presents the web page collections and the two experiments; conclusions are drawn in Section 4.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2">PREVIOUS WORK</head><p>Several experiments have been recently carried out with genres and web pages. Here we list the latest studies in order to show how difficult is to compare their results in the absence of common criteria as for corpus building and genre palettes. <ref type="bibr" target="#b6">[7]</ref>: Number of web pages: 2150; Annotation: single rater; Categories: subjectivity, positive-ness. They tried to discriminate among texts coming from different domains in terms of two polarities: subjective vs. objective and positive vs. negative. Their aim was to see how a classification model tuned on one domain performed in another domain. According to their results, in single domain classification the best accuracy is achieved with Multi-View-Ensemble (MVE) (see <ref type="bibr" target="#b6">[7]</ref> for details) for subjectivity, and with bag-of-words (BOW) features for positive-ness. In domain transfer classification, the best accuracy is achieved with Parts-of-Speech (POS) tags for subjectivity and MVE for positive-ness. Although it is true that genres can be divided into more subjective genres (e.g. editorials), or more objective genres (e.g. surveys), and that the opposition positive-negative can suggest specific genres (such as reviews), these two polarities can hardly be considered as "genres" in themselves. Nonetheless, <ref type="bibr" target="#b6">[7]</ref>'s contribution is extremely valuable because they shed some light on the performance of different feature sets across several domains, providing insight into the extent of feature exportability. <ref type="bibr" target="#b4">[5]</ref>: Number of web pages: 2700; Annotation: one or more raters; Categories: functional styles. They carried out an experiment on style-dependent document ranking. Their research explored the possibility of incorporating style-dependent ranking into ranking schemata for searching the web and digital libraries. Their basic idea was to reduce styles (more specifically, the five functional styles theorized by the School of Prague) to a single continuous parameter. Regardless the promising preliminary results, they could see little improvement in relevance ranking when stylistic parameters were included. <ref type="bibr" target="#b2">[3]</ref>: Number of web pages: 343; Genre annotation: the author plus at least one or more raters; Genres: abstract, call for papers, FAQs, hub/sitemap, job description, resume/C.V., statistics, syllabus, technical paper. She tried out the efficiency of several feature sets and automatic feature selection techniques on a small corpus of 10 genres, using a number of classification algorithms. Although her results can be considered only indicative given the reduced number of pages per genre (an average of 20 web pages per genre class), she made interesting remarks about discrimination across similar genres, and the influence of the genre palette and document exemplarity on discrimination tasks. Her best accuracy (92.1%) was achieved by one of the feature combinations resulting from an automatic feature selection technique.</p><p>[10]: Number of web pages: 321; Genre annotation: do not say; Genres: personal, corporate, organizational home pages, including also non-home pages, as noise. They tried the hard task of home page genre discrimination. The best accuracy (71.4%) is achieved on personal home pages with a single classifier, manual feature selection, and without noisy pages. <ref type="bibr" target="#b15">[16]</ref>: Number of web pages: 1224; Genre annotation: two graduate students; Genres: personal home page, public home page, commercial home page, bulletin collection, link collection, image collection, simple table/lists, input pages, journalistic material, research report, official materials, FAQs, discussions, product specification, informal texts (poem, fiction, etc.). They investigated the efficiency of several feature sets to discriminate across these 16 genres. They also tested the classification efficiency on different parts of the web page space (title and metacontent, body, and anchors). The best accuracy (75.7%) was achieved with one of their features sets when applied only to the body and anchors. <ref type="bibr" target="#b16">[17]</ref>: Number of web pages: 800; Genre annotation: three raters; Genres: help, article, discussion, shop, portrayal (non-private), portrayal (private), link collection, download. They worked out a genre palette of eight genres following the outcome of a study on genre usefulness. As they aimed at a classification performed on the fly, they assessed features according to the computational effort they required, giving preference to those requiring low or medium effort. They achieved around 70% accuracy with discriminant analysis on the palette of eight genres. Other results relate to groups of genres tailored for web user profiles. <ref type="bibr" target="#b13">[14]</ref> and the follow up <ref type="bibr" target="#b14">[15]</ref>: Number of web pages: 321; Genre annotation: at least two raters; Genres: reportage-editorial, research article, review, home page, Q&amp;A, specification. They aimed at selecting genre-revealing terms from the training document set using collection of web pages annotated both at topic level and at genre level. Their formula (the deviation formula) makes use of both genre-classified documents and subject-classified documents and eliminate terms that are more subject-related than genre-related. They report a micro-average of precision and recall of about 90%.</p><p>As already stressed, the absence of common criteria or evaluation ground makes most of these experiments (see Table <ref type="table" target="#tab_0">1</ref> for a summary) difficult to compare, however fruitful each study can be in itself. A cross-evaluation of these experiments remains virtually unfeasible because genre palettes are mostly disparate. Also in the case of 'home page', which is probably one of the few genres in common in several experiments, any comparison appear to be difficult, because selection criteria and level of exemplarity are not declared. The two criteria of annotation by objective sources and consistent level of granularity are suggested to overcome this un-comparability. The KI-04 corpus was collected using bookmarks from about five people. Some genres were extended to get a better balance. The corpus was sorted by three people, one of them wrote a bachelor thesis (in German) on the corpus building process. One of the author of <ref type="bibr" target="#b16">[17]</ref> checked many of the pages, and most of the sorting complied with his understanding of the genre categories. The download date was January 26th, 2004.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.3">SPIRIT collection</head><p>The SPIRIT collection is a random crawl carried out in 2001 (see <ref type="bibr" target="#b7">[8]</ref>). It contains single web pages and not full websites. The size of the whole collection is about one terabyte, and the number of web pages (mostly HTML files) is about 95 millions. It is multilingual and without any meta-information, apart from a short header including the original URL, the date and time when the pages were crawled from the web, and few other details. It represents a genuine slice of the real web. In Experiment 2, we used only 1,000 English web pages (available online at the URL reported in the Introduction) from this random, multilingual and unclassified collection.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.4">Experiment 1</head><p>The practical aim of Experiment 1 was to build two single-label discrete classification models, one out of the 7-web-genre collection, the other from KI-04 corpus, and compare their accuracy results. Both collections were submitted to the same preprocessing. The unit of analysis was a single static web page in HTML format. The feature set, called 1_set, used in Experiment 1 includes:</p><p>• the 50 most common words in English;</p><p>• 24 Part-of-Speech (POS) tags;</p><p>• 8 punctuation marks: full stop (.), colon (:), semi-colon (;), comma (,), exclamation mark (!), question mark (?), apostrophe ('), and quotes ("); • genre-specific words 3  As you can see in Table <ref type="table" target="#tab_1">2</ref>, the accuracy of the model built with the 7-web-genre collection is much higher than the model built with KI-04 corpus, namely +21.7%. In order to see whether the feature set was too tailored or biased towards the 7-web-genre collection, we compared the accuracy of this feature set on KI-04 corpus with the accuracy rates reported in <ref type="bibr" target="#b16">[17]</ref>. To make this comparison possible, we ran discriminant analysis using our feature set on KI-04 corpus. As <ref type="bibr" target="#b16">[17]</ref> ran their discriminant analysis only on 800 web pages while we used 1,205 3 Genre-specific words were selected through a cursory manual analysis.</p><p>A total of 13 sets of genre-specific words were built. 13 and not 15 because two sets were shared across the two collections, namely those related to home-page/portrayal (priv) and eshop/shop. It is worth saying that genre-specific words (available online at the URL reported in the Introduction) are not numerous. For example, genre-specific words for the search web genre are only: search, crawl, directories, engine, find, and see.</p><p>web pages, we converted all the results into percentages. A breakdown of the different accuracy rates achieved with discriminant analysis and two different feature set is shown in Table <ref type="table" target="#tab_2">3</ref>. Our feature set performs better than <ref type="bibr" target="#b16">[17]</ref>'s feature set. Although the difference is rather small (+2.1%), it is statistically significant (chi-square test). This means that our feature set is not biased toward the 7-web-genre collection, but it performs significantly better than <ref type="bibr" target="#b16">[17]</ref>'s feature set on KI-04 corpus with discriminant analysis, i.e. the same algorithm used in <ref type="bibr" target="#b16">[17]</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.4.1">Discussion</head><p>Experiment 1 compares the accuracies of two models built with the same classification algorithm, the same feature set but different web page collections, the 7-web-genre collection and KI-04 corpus. The accuracy on the 7-web-genre collection (1,400 web pages) is above 90% while the accuracy on KI-04 corpus is definitely lower. A first thought was that our feature set did not represent the genre palette of KI-04 corpus adequately. However, after having compared the performance of our feature set with <ref type="bibr" target="#b16">[17]</ref>'s feature set using the same algorithm (discriminant analysis) on the same collection, we saw that the accuracy achieved by our feature set was slightly higher than the accuracy stated in <ref type="bibr" target="#b16">[17]</ref>.</p><p>Although KI-04 corpus contains eight genres, i.e. one genre more than the 7-web-genre collection (error rate usually increases with the number of categories), this does not justify such a wide the gap in the classification accuracy. Also, it is important to stress that genre-specific words are tailored to the genre palette. This means, the genre-specific words used for the 7-web-genre collection account for blogs, search, front page, etc., while those employed for KI-04 corpus include words relate to articles, discussion, download, etc. Since these two genre palettes have two web genres in common, i.e. home page/portrayal (priv) and eshop/shop, in these two cases the same set of genre-specific words was used for both web genre collections. That the feature set used in the KI-04 corpus is not biased towards the 7-web genre collection is confirmed by the results shown in Table <ref type="table" target="#tab_2">3</ref>, where the performance of our features set is higher than <ref type="bibr" target="#b16">[17]</ref>'s feature set.</p><p>In conclusion, if neither the feature set nor the classification algorithm is the cause of this large discrepancy in accuracy, then the suspicion is that the selection of the web pages representing genres in KI-04 corpus might be responsible for the lower performance. Although the issue of subjectivity of the assignment of genre to web pages needs further investigation (cf. also <ref type="bibr" target="#b3">[4]</ref>), for the time being we interpret the higher performance on the 7-webgenre collection as a result of the application of the two criteria of annotation by objective sources annotation and consistent genre granularity.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.5">Experiment 2</head><p>The goal of Experiment 2 was to see whether the classification model built with the collection complying to the criteria of annotation by objective source and consistent genre granularity is more effective also for predictive tasks. In other words, predictions are used here as a kind of evaluation metrics of the efficiency of classification models.</p><p>In this experiment we used the two classification models built in the previous experiment together with additional models. The practical aim was to make predictions on unclassified and non-annotated web pages, i.e. 1,000 random English web pages from the SPIRIT collection. The relevance of the agreed upon web pages (see Tables <ref type="table" target="#tab_5">5 and 6</ref>) to a genre was manually assessed by the author of this paper (the breakdown of this manual evaluation is available online at the URL reported in the Introduction).</p><p>When making a prediction, the classifier returns a probability score to be interpreted in terms of classification confidence. This confidence score can be exploited when assessing the value of a prediction and for setting a threshold for reliable guesses. In order to get predictions on genre labels which were as reliable as possible, we devised an approach inspired by co-training. The basic idea was to exploit three different views (i.e. three different feature sets) on the same data. When the three models built with the three feature sets agreed on the same genre label (3-out-of-3 agreement) at very high confidence score, namely &gt;=0.9, this was for us an indication of a good prediction. Additionally, as we have two web page collections with two different genre palettes, we can have multi-label predictions. Ideally, a web page might get a prediction of "personal home page", following the palette adopted in the 7-web-genre collection, and "portrayal (private)", following the genre palette adopted in KI-04 corpus. Also, as the two palettes are mostly not overlapping, it is interesting to see which palette is more suitable for the classification of this SPIRIT random sample. From the previous experiment we had two models built with a single feature set (1_set). To these models, we add four additional models (two per collection) in order to get the three simultaneous views on each collection. The additional two models were built using the feature sets called 2_set and 3_set (these feature sets, together with a description, are available online at the URL reported in the Introduction). 2_set contains the following features:</p><p>• POS trigrams; • 8 punctuation symbols (as above);</p><p>• genre-specific words (as above);</p><p>• 28 HTML tags (as above); • 1 nominal attribute representing the length of the web page (as above).</p><p>3_set contains the following features:</p><p>• 86 linguistic facets 4 ; • genre-specific words; • 6 HTML facets; • 1 nominal attribute representing the length of the web page (as above). 4 Linguistic facets and HTML facets are groups of features highlighting an aspect in the communicative context that is reflected in the use of language. They are listed in the URL reported in the Introduction.</p><p>Table <ref type="table" target="#tab_3">4</ref> shows the performance of the three feature sets on the two web genre collections. From the summary shown in Table <ref type="table" target="#tab_4">5</ref>, we can see that a very low number of pages were agreed upon by the three classification models (second column) built on the 7-web-page collection. This is not necessarily bad when aiming at high precision (future work will explore the possibility of increasing precision). However, predictions are even sparer with the models built using KI-04 corpus (Table <ref type="table" target="#tab_5">6</ref>). As there was no 3-out-of-3 agreement for discussion, download, help, and portrayal (non-private), these genres were evaluated with 2-out-of-3 agreement. No correct guesses were returned for article, discussion, download, and help. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.5.1">Discussion</head><p>Experiment 2 shows that the classification models built with the 7-web-genre collection return a higher number of predictions. This seems to confirm the interpretation that using the two criteria of objective source annotation and consistent level of granularity ensures better classification models and consequently a higher number of correct predictions. Also, this experiment shows a useful methodology to follow for multi-genre classification of web pages, which can be refined and further investigated in future.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4">CONCLUSIONS</head><p>In this paper we pointed out how classification models learned from a web collection annotated by genre using the two criteria of annotation by objective source and consistent level of granularity can return higher accuracy and a higher number of correct predictions.</p><p>The annotation by objective source is not only less subjective and closer to real-world conditions, but also much faster than annotation by human raters, which is usually time-consuming, controversial, and expensive. Further, a collection built with a consistent level of genre granularity seems to be learned more profitably by the classifier. Together, these two criteria enhance the performance of classification algorithms. However, a full comparison between the results achieved with the two web page collections built with different criteria is not entirely feasible because the two genre palettes are mostly different. Nonetheless, these findings are indicative of a tendency that can be further investigated in future. It is also worth pointing out that objective sources may still contain biases. Biases in web collections relate to the well-known issue of 'corpus representativeness', dating back to Chomsky's aversion to the use of corpora. However, in the present days and with the web available, biases can be alleviated by randomly picking up web pages from several genre-specific web archives or portals.</p><p>Although the two criteria of annotation by objective source and consistent level of granularity represent a practical solution that can help genre classification, the concept of genre remains hard to capture computationally and statistically in its entirety.</p><p>First, it would be interesting to investigate more about the ideal proportion among corpus size, number of features and number of classes and its influence on classification results. Also, up to now only single-label discrete classification has been tried out in genre classification studies. Experiment 2 implicitly shows an easy method that can be exploited for multi-label classification: the use of concurrent genre palettes over the same unclassified collection. Ideally, the use of several classification models built with different collections annotated by external sources and a consistent granularity, and including different genre palettes can suggest several genre labels for the same web page. Multi-genre documents and genre hybridism are particularly acute when dealing with web pages, which appear much more unpredictable and individualized than paper documents. Using concurrent genre palettes might represent an alternative to the multi-faceted approach by <ref type="bibr" target="#b10">[11]</ref>. What is less reassuring is the absence of a proper evaluation metrics for multi-label problems. We leave these problems open to further investigations and invite the genre classification community to make use of the three collections employed in these experiments and now available online.</p></div><figure xmlns="http://www.tei-c.org/ns/1.0" type="table" xml:id="tab_0"><head>Table 1 .</head><label>1</label><figDesc>Summary TableThe 7-web-genre collection includes 200 English web pages per genre, amounting to a total of 1,400 web pages (available online at the URL reported in the Introduction). These web pages were collected by the author of this paper in early spring 2005. This collection was built with genres belonging to a consistent level of granularity and applying the annotation by objective source. The seven web genres included in the collection are the following: Personal home page' is the basic level of the superordinate level 'home page' and has 'academic personal home page', 'administrative personal home page', etc. as subordinate level.The web pages included in the 7-web-genre collection were randomly downloaded from the following public archives or portals (download date: Feb-March 2005):</figDesc><table><row><cell></cell><cell></cell><cell></cell><cell></cell><cell>•</cell><cell>Blogs:</cell></row><row><cell></cell><cell></cell><cell></cell><cell></cell><cell></cell><cell>o</cell><cell>http://www.britblog.com/</cell></row><row><cell></cell><cell></cell><cell></cell><cell></cell><cell></cell><cell>o</cell><cell>http://www.nataliedarbeloff.com/augustinearchive.html.</cell></row><row><cell></cell><cell></cell><cell></cell><cell></cell><cell>•</cell><cell>Eshops:</cell></row><row><cell></cell><cell></cell><cell></cell><cell></cell><cell></cell><cell>o</cell><cell>http://www.shops.co.uk/</cell></row><row><cell></cell><cell></cell><cell></cell><cell></cell><cell></cell><cell>o</cell><cell>http://www.eshops.co.uk/</cell></row><row><cell></cell><cell></cell><cell></cell><cell></cell><cell>•</cell><cell>FAQs:</cell></row><row><cell></cell><cell></cell><cell></cell><cell></cell><cell></cell><cell>o</cell><cell>http://www.cybernothing.org/faqs/net-abuse-faq.html</cell></row><row><cell></cell><cell></cell><cell></cell><cell></cell><cell></cell><cell>o</cell><cell>http://www.irs.gov/faqs/</cell></row><row><cell></cell><cell></cell><cell></cell><cell></cell><cell></cell><cell>o</cell><cell>http://www.copyright.gov/help/faq/</cell></row><row><cell>Studies [7] [5]</cell><cell cols="3">No. of web pages 2,150 single rater Subjectivity vs. objectivity, positive Annotation Labels vs. negative 2,700 One or more raters public affairs style, everyday communication style, scientific style, journalistic style, literary style</cell><cell>• •</cell><cell>o Newspaper front pages belong to a number of different http://www.aoml.noaa.gov/hrd/tcfaq/tcfaqHED.html online newspaper and are available at Internet Archive: o www.archive.org Personal home pages are heterogeneous, and include academic and administrative personal home pages, as well as more informal personal home pages. They were downloaded from:</cell></row><row><cell>[3]</cell><cell>343</cell><cell>Two or more</cell><cell>abstract, call for papers, FAQs,</cell><cell></cell></row><row><cell></cell><cell></cell><cell>raters</cell><cell>hub/sitemap, job description,</cell><cell></cell></row><row><cell></cell><cell></cell><cell></cell><cell>resume/C.V., statistics, syllabus,</cell><cell></cell></row><row><cell></cell><cell></cell><cell></cell><cell>technical paper</cell><cell></cell></row><row><cell>[10]</cell><cell>321</cell><cell>do not say</cell><cell>home pages (personal, corporate,</cell><cell></cell></row><row><cell></cell><cell></cell><cell></cell><cell>organizational)</cell><cell></cell></row><row><cell>[16]</cell><cell cols="2">1,224 two graduate</cell><cell>personal home page, public home</cell><cell></cell></row><row><cell></cell><cell></cell><cell>students</cell><cell>page, commercial home page,</cell><cell></cell></row><row><cell></cell><cell></cell><cell></cell><cell>bulletin collection, link collection,</cell><cell></cell></row><row><cell></cell><cell></cell><cell></cell><cell>image collection, simple table/lists,</cell><cell></cell></row><row><cell></cell><cell></cell><cell></cell><cell>input pages, journalistic material,</cell><cell></cell></row><row><cell></cell><cell></cell><cell></cell><cell>research report, official materials,</cell><cell></cell></row><row><cell></cell><cell></cell><cell></cell><cell>FAQs, discussions, product</cell><cell></cell></row><row><cell></cell><cell></cell><cell></cell><cell>specification, informal texts</cell><cell></cell></row><row><cell>[17]</cell><cell>800</cell><cell>3 raters</cell><cell>article, discussion, shop, portrayal</cell><cell></cell></row><row><cell></cell><cell></cell><cell></cell><cell>(non-private), portrayal (private),</cell><cell></cell></row><row><cell></cell><cell></cell><cell></cell><cell>link collection, download</cell><cell></cell></row><row><cell>[14] and [15]</cell><cell>321</cell><cell>at least two raters</cell><cell>reportage-editorial, research article, review, home page, Q&amp;A,</cell><cell cols="2">3.2 KI-04 corpus</cell></row><row><cell></cell><cell></cell><cell></cell><cell>specification</cell><cell cols="2">KI-04 corpus was built following a palette of eight genres</cell></row><row><cell></cell><cell></cell><cell></cell><cell></cell><cell cols="2">suggested by a user study on genre usefulness ([17]). It includes</cell></row><row><cell cols="3">3 EXPERIMENTS</cell><cell></cell><cell cols="2">1,295 English web pages (HTML documents), but only 800 web</cell></row><row><cell></cell><cell></cell><cell></cell><cell></cell><cell cols="2">pages (100 per genre) were used in the experiment described in</cell></row><row><cell cols="4">3.1 7-Web-Genre Collection</cell><cell cols="2">[17]. In Experiment 1, we used 1,205 web pages because some</cell></row><row><cell></cell><cell></cell><cell></cell><cell></cell><cell cols="2">web pages were empty (both original version, 1,295 web pages,</cell></row><row><cell></cell><cell></cell><cell></cell><cell></cell><cell cols="2">and working version, 1,205 web pages, are available online at the</cell></row><row><cell></cell><cell></cell><cell></cell><cell></cell><cell cols="2">URL reported in the Introduction). KI-04 corpus includes:</cell></row><row><cell></cell><cell></cell><cell></cell><cell></cell><cell cols="2">1. article (127 web pages)</cell><cell>5. discussion (127 w. p)</cell></row><row><cell></cell><cell></cell><cell></cell><cell></cell><cell cols="2">2. download (151 w. p)</cell><cell>6. help (139 w. p)</cell></row><row><cell></cell><cell></cell><cell></cell><cell></cell><cell cols="2">3. link collection (205 w. p)</cell><cell>7. portrayal (non-priv) (163 w. p.)</cell></row><row><cell></cell><cell></cell><cell></cell><cell></cell><cell cols="2">4. portrayal (priv.) (126 w. p)</cell><cell>8. shop (167 w. p)</cell></row><row><cell>1. blog</cell><cell></cell><cell></cell><cell>5. list</cell><cell></cell></row><row><cell>2. eshop</cell><cell></cell><cell></cell><cell>6. personal home page 2</cell><cell></cell></row><row><cell>3. FAQs</cell><cell></cell><cell cols="2">7. search page</cell><cell></cell></row><row><cell cols="3">4. online newspaper front page</cell><cell></cell><cell></cell></row></table><note>2 'o http://dmoz.org/Society/People/Personal_Homepages/ o http://www.math.unl.edu/~mbritten/ldt/homepage.html o http://www.bradley.edu/people/fac-staff.html o http://www.daimi.au.dk/local/map/PeopleandLocationsPe opleFrame.html o http://www.mit.edu/Home-byUser.html o http://dir.yahoo.com/Society_and_Culture/People/Person al_Home_Pages o http://hpsearch.uni-trier.de/hp/a-tree/ o Search pages comes from: o http://www.searchenginecolossus.com/ The web pages included in the genre 'list', were selected searching keywords in Google and selecting relevant web pages from the results. All the lists include one of the following keywords (and orthographic variants) in the heading: checklist, hot list, table of content, and sitemap (see, for example, Insect Hotlist at http://www.fi.edu/tfi/hotlists/insects.html).</note></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" type="table" xml:id="tab_1"><head>Table 2 .</head><label>2</label><figDesc>; • 28 HTML tags; • 1 nominal attribute representing the length of the web page (SHORT, MEDIUM and LONG). Averaged Accuracies with SMO</figDesc><table><row><cell cols="2">(This feature set, together with a description, is available online at</cell></row><row><cell cols="2">the URL reported in the Introduction). The classification</cell></row><row><cell cols="2">algorithm used both in Experiments 1 and 2 is SMO (which</cell></row><row><cell cols="2">implements the Sequential Minimal Optimisation (SMO) for</cell></row><row><cell cols="2">training support vectors) with default parameters and logistic</cell></row><row><cell cols="2">regression model, from Weka machine learning workbench ([25]).</cell></row><row><cell cols="2">Accuracy results, shown in Table 2, are averaged over stratified</cell></row><row><cell cols="2">10-fold crossvalidations repeated 10 times.</cell></row><row><cell>Averaged Accuracy on the 7-</cell><cell>Averaged Accuracy on KI-04</cell></row><row><cell>web-genre collection</cell><cell>corpus</cell></row><row><cell>90.6%</cell><cell>68.9%</cell></row></table></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" type="table" xml:id="tab_2"><head>Table 3 .</head><label>3</label><figDesc>Accuracy rates with discriminant analysis</figDesc><table><row><cell>KI-04 corpus</cell><cell>Our feature set</cell><cell>[17]'s feature set</cell></row><row><cell>Article</cell><cell>80.3%</cell><cell>81.3%</cell></row><row><cell>Discussion</cell><cell>76.4%</cell><cell>68.5%</cell></row><row><cell>Download</cell><cell>74.2%</cell><cell>79.6%</cell></row><row><cell>Help</cell><cell>59.7%</cell><cell>55.1%</cell></row><row><cell>Link Collection</cell><cell>69.3%</cell><cell>67.6%</cell></row><row><cell>Portrayal (non-priv)</cell><cell>59.5%</cell><cell>57.9%</cell></row><row><cell>Portrayal (priv)</cell><cell>73.8%</cell><cell>67.7%</cell></row><row><cell>Shop</cell><cell>68.3%</cell><cell>66.9%</cell></row><row><cell>Accuracy</cell><cell>70.2%</cell><cell>68.1%</cell></row></table></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" type="table" xml:id="tab_3"><head>Table 4 .</head><label>4</label><figDesc>Accuracies of three feature sets on two collections</figDesc><table><row><cell>Classification</cell><cell>Averaged accuracy on the</cell><cell>Averaged accuracy on</cell></row><row><cell>algorithm: Weka</cell><cell>7-web-genre collection</cell><cell>KI-04 corpus</cell></row><row><cell>SMO</cell><cell></cell><cell></cell></row><row><cell>1_set</cell><cell>90.6%</cell><cell>68.9%</cell></row><row><cell>2_set</cell><cell>89.4%</cell><cell>64.1%</cell></row><row><cell>3_set</cell><cell>88.8%</cell><cell>65.9%</cell></row></table></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" type="table" xml:id="tab_4"><head>Table 5 .</head><label>5</label><figDesc>Correct predictions with the 7-web-genre palette</figDesc><table><row><cell>7 WEB GENRE</cell><cell># OF AGREED</cell><cell>CORRECT</cell><cell>INCORRECT</cell><cell>ERROR</cell></row><row><cell>PALETTE</cell><cell>UPON WEB PAGES</cell><cell>GUESSES</cell><cell>GUESSES AND</cell><cell>RATE</cell></row><row><cell></cell><cell>(OUT OF 1,000)</cell><cell></cell><cell>UNCERTAIN</cell><cell></cell></row><row><cell>BLOG</cell><cell>17</cell><cell>1</cell><cell>16</cell><cell>0.94</cell></row><row><cell>ESHOP</cell><cell>11</cell><cell>3</cell><cell>8</cell><cell>0.73</cell></row><row><cell>FAQs</cell><cell>8</cell><cell>1</cell><cell>7</cell><cell>0.88</cell></row><row><cell>FRONTPAGE</cell><cell>7</cell><cell>0</cell><cell>7</cell><cell>1.00</cell></row><row><cell>LISTING</cell><cell>18</cell><cell>7</cell><cell>11</cell><cell>0.61</cell></row><row><cell>PHP</cell><cell>44</cell><cell>10</cell><cell>34</cell><cell>0.77</cell></row><row><cell>SPAGE</cell><cell>12</cell><cell>6</cell><cell>6</cell><cell>0.50</cell></row><row><cell>TOTAL</cell><cell>117</cell><cell>28</cell><cell>89</cell><cell></cell></row><row><cell>PERCENTAGE</cell><cell>11.7%</cell><cell>2.8%</cell><cell>8.9%</cell><cell></cell></row></table></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" type="table" xml:id="tab_5"><head>Table 6 .</head><label>6</label><figDesc>Correct predictions with KI-04 corpus</figDesc><table><row><cell>KI-04 CORPUS</cell><cell># OF AGREED</cell><cell>CORRECT</cell><cell>INCORRECT</cell><cell>ERROR</cell></row><row><cell></cell><cell>UPON WEB</cell><cell>GUESSES</cell><cell>GUESSES AND</cell><cell>RATE</cell></row><row><cell></cell><cell>PAGES (OUT OF</cell><cell></cell><cell>UNCERTAIN</cell><cell></cell></row><row><cell></cell><cell>1,000)</cell><cell></cell><cell></cell><cell></cell></row><row><cell>ARTICLE</cell><cell>4</cell><cell>0</cell><cell>4</cell><cell>1.00</cell></row><row><cell>DISCUSSION</cell><cell>8</cell><cell>0</cell><cell>8</cell><cell>1.00</cell></row><row><cell>DOWNLOAD</cell><cell>4</cell><cell>0</cell><cell>4</cell><cell>1.00</cell></row><row><cell>HELP</cell><cell>3</cell><cell>0</cell><cell>3</cell><cell>1.00</cell></row><row><cell>LINK</cell><cell>3</cell><cell>3</cell><cell>0</cell><cell>0.00</cell></row><row><cell>PORTRAYAL (NON-</cell><cell>5</cell><cell>1</cell><cell>4</cell><cell>0.80</cell></row><row><cell>PRIVATE)</cell><cell></cell><cell></cell><cell></cell><cell></cell></row><row><cell>PORTRAYAL</cell><cell>7</cell><cell>3</cell><cell>4</cell><cell>0.57</cell></row><row><cell>(PRIVATE)</cell><cell></cell><cell></cell><cell></cell><cell></cell></row><row><cell>SHOP</cell><cell>6</cell><cell>3</cell><cell>3</cell><cell>0.50</cell></row><row><cell>TOTAL</cell><cell>36</cell><cell>10</cell><cell>26</cell><cell></cell></row><row><cell>PERCENTAGE</cell><cell>3.6%</cell><cell>1%</cell><cell>2.6%</cell><cell></cell></row></table></figure>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="1" xml:id="foot_0">University of Brighton (UK); M.Santini@brighton.ac.uk</note>
		</body>
		<back>
			<div type="references">

				<listBibl>

<biblStruct xml:id="b0">
	<analytic>
		<title level="a" type="main">Routing documents according to style</title>
		<author>
			<persName><forename type="first">S</forename><surname>Argamon</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Koppel</surname></persName>
		</author>
		<author>
			<persName><forename type="first">G</forename><surname>Avneri</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proc. First International Workshop on Innovative Internet Information Systems</title>
				<meeting>First International Workshop on Innovative Internet Information Systems</meeting>
		<imprint>
			<date type="published" when="1998">1998</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b1">
	<monogr>
		<title level="m" type="main">Analysing Genre. Language Use in Professional Settings</title>
		<author>
			<persName><forename type="first">V</forename><surname>Bathia</surname></persName>
		</author>
		<imprint>
			<date type="published" when="1993">1993</date>
			<publisher>Longman</publisher>
			<pubPlace>London and New York</pubPlace>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b2">
	<monogr>
		<title level="m" type="main">Stereotyping the Web: Genre Classification of Web Documents</title>
		<author>
			<persName><forename type="first">E</forename><surname>Boese</surname></persName>
		</author>
		<imprint>
			<date type="published" when="2005">2005</date>
		</imprint>
		<respStmt>
			<orgName>Colorado State Univ.</orgName>
		</respStmt>
	</monogr>
	<note type="report_type">M.S. Thesis</note>
</biblStruct>

<biblStruct xml:id="b3">
	<analytic>
		<title level="a" type="main">Effects of Web Document Evolution on Genre Classification</title>
		<author>
			<persName><forename type="first">E</forename><surname>Boese</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><surname>Howe</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">CIKM&apos;05</title>
				<imprint>
			<date type="published" when="2005">2005</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b4">
	<analytic>
		<title level="a" type="main">Experiment on Style-Dependent Document Ranking</title>
		<author>
			<persName><forename type="first">P</forename><surname>Bravslavski</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><surname>Tselischev</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proc. of the 7th Russian Conference on Digital Libraries</title>
				<meeting>of the 7th Russian Conference on Digital Libraries</meeting>
		<imprint>
			<date type="published" when="2005">2005</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b5">
	<analytic>
		<title level="a" type="main">Genres and the Web: is the personal home page the first uniquely digital genre?</title>
		<author>
			<persName><forename type="first">A</forename><surname>Dillon</surname></persName>
		</author>
		<author>
			<persName><forename type="first">B</forename><surname>Gushrowski</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">JASIS</title>
		<imprint>
			<biblScope unit="volume">51</biblScope>
			<biblScope unit="issue">2</biblScope>
			<date type="published" when="2000">2000</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b6">
	<analytic>
		<title level="a" type="main">Learning to classify documents according to genre</title>
		<author>
			<persName><forename type="first">A</forename><surname>Finn</surname></persName>
		</author>
		<author>
			<persName><forename type="first">N</forename><surname>Kushmerick</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">JASIST, Special Issue</title>
		<imprint>
			<biblScope unit="volume">7</biblScope>
			<biblScope unit="issue">5</biblScope>
			<date type="published" when="2006">2006</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b7">
	<analytic>
		<title level="a" type="main">The SPIRIT collection: an overview of a large web collection</title>
		<author>
			<persName><forename type="first">H</forename><surname>Joho</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Sanderson</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">SIGIR Forum</title>
		<imprint>
			<biblScope unit="volume">38</biblScope>
			<biblScope unit="issue">2</biblScope>
			<date type="published" when="2004">2004</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b8">
	<monogr>
		<title level="m" type="main">Stylistic Experiments for Information Retrieval</title>
		<author>
			<persName><forename type="first">J</forename><surname>Karlgren</surname></persName>
		</author>
		<imprint>
			<date type="published" when="2000">2000</date>
			<pubPlace>Sweden</pubPlace>
		</imprint>
		<respStmt>
			<orgName>degree of Doctor of Philosophy ; Stockholm University</orgName>
		</respStmt>
	</monogr>
	<note>Thesis submitted for the</note>
</biblStruct>

<biblStruct xml:id="b9">
	<analytic>
		<title level="a" type="main">Automatic Identification of Home Pages on the Web</title>
		<author>
			<persName><forename type="first">A</forename><surname>Kennedy</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Shepherd</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proc. 38 HICSS</title>
				<meeting>38 HICSS</meeting>
		<imprint>
			<date type="published" when="2005">2005</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b10">
	<analytic>
		<title level="a" type="main">Automatic Detection of Text Genre</title>
		<author>
			<persName><forename type="first">B</forename><surname>Kessler</surname></persName>
		</author>
		<author>
			<persName><forename type="first">G</forename><surname>Numberg</surname></persName>
		</author>
		<author>
			<persName><forename type="first">H</forename><surname>Shütze</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proc. 35 Annual Meeting of the ACL and 8th Conference of the EACL</title>
				<meeting>35 Annual Meeting of the ACL and 8th Conference of the EACL</meeting>
		<imprint>
			<date type="published" when="1997">1997</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b11">
	<analytic>
		<title level="a" type="main">Identifying document genre to improve web search effectiveness</title>
		<author>
			<persName><forename type="first">B</forename><surname>Kwasnik</surname></persName>
		</author>
		<author>
			<persName><forename type="first">K</forename><surname>Crowston</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Nilan</surname></persName>
		</author>
		<author>
			<persName><forename type="first">D</forename><surname>Roussinov</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">The Bulletin of the American Society for Information Science and Technology</title>
		<imprint>
			<biblScope unit="volume">27</biblScope>
			<biblScope unit="issue">2</biblScope>
			<biblScope unit="page" from="23" to="26" />
			<date type="published" when="2000">2000</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b12">
	<analytic>
		<title level="a" type="main">Registers, Text types, Domains, and Styles: Clarifying the concepts and navigating a path through the BNC Jungle</title>
		<author>
			<persName><forename type="first">D</forename><surname>Lee</surname></persName>
		</author>
		<author>
			<persName><surname>Genres</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Language Learning and Technology</title>
		<imprint>
			<biblScope unit="volume">5</biblScope>
			<biblScope unit="issue">3</biblScope>
			<biblScope unit="page" from="37" to="72" />
			<date type="published" when="2001">2001</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b13">
	<analytic>
		<title level="a" type="main">Automatic Identification of Text Genres and Their Roles in Subject-Based Categorization</title>
		<author>
			<persName><forename type="first">Y</forename><surname>Lee</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Myaeng</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proc. 37 HICSS</title>
				<meeting>37 HICSS</meeting>
		<imprint>
			<date type="published" when="2004">2004</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b14">
	<analytic>
		<title level="a" type="main">Text Genre Classification with Genre-Revealing and Subject-Revealing Features</title>
		<author>
			<persName><forename type="first">Y</forename><surname>Lee</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Myaeng</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proc. 25 Annual International ACM SIGIR</title>
				<meeting>25 Annual International ACM SIGIR</meeting>
		<imprint>
			<date type="published" when="2002">2002</date>
			<biblScope unit="page" from="145" to="150" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b15">
	<analytic>
		<title level="a" type="main">Automatic Genre Detection of Web Documents</title>
		<author>
			<persName><forename type="first">C</forename><surname>Lim</surname></persName>
		</author>
		<author>
			<persName><forename type="first">K</forename><surname>Lee</surname></persName>
		</author>
		<author>
			<persName><forename type="first">G</forename><surname>Kim</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Natural Language Processing</title>
				<editor>
			<persName><forename type="first">K</forename><surname>Su</surname></persName>
		</editor>
		<editor>
			<persName><forename type="first">J</forename><surname>Tsujii</surname></persName>
		</editor>
		<editor>
			<persName><forename type="first">J</forename><surname>Lee</surname></persName>
		</editor>
		<editor>
			<persName><forename type="first">O</forename><forename type="middle">Y</forename><surname>Kwong</surname></persName>
		</editor>
		<meeting><address><addrLine>Berlin</addrLine></address></meeting>
		<imprint>
			<publisher>Springer</publisher>
			<date type="published" when="2005">2005</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b16">
	<analytic>
		<title level="a" type="main">Genre Classification of Web Pages: User Study and Feasibility Analysis</title>
		<author>
			<persName><forename type="first">S</forename><surname>Meyer Zu Eissen</surname></persName>
		</author>
		<author>
			<persName><forename type="first">B</forename><surname>Stein</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Advances in Artificial Intelligence</title>
				<editor>
			<persName><forename type="first">S</forename><surname>Biundo</surname></persName>
		</editor>
		<editor>
			<persName><forename type="first">T</forename><surname>Fruhwirth</surname></persName>
		</editor>
		<editor>
			<persName><forename type="first">G</forename><surname>Palm</surname></persName>
		</editor>
		<meeting><address><addrLine>Berlin</addrLine></address></meeting>
		<imprint>
			<publisher>Springer</publisher>
			<date type="published" when="2004">2004</date>
			<biblScope unit="page" from="256" to="269" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b17">
	<analytic>
		<title level="a" type="main">Working with genre: A pragmatic perspective</title>
		<author>
			<persName><forename type="first">B</forename><surname>Paltridge</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Journal of Pragmatics</title>
		<imprint>
			<biblScope unit="volume">24</biblScope>
			<biblScope unit="page" from="393" to="406" />
			<date type="published" when="1995">1995</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b18">
	<analytic>
		<title level="a" type="main">Integrating Automatic Genre Analysis into Digital Libraries</title>
		<author>
			<persName><forename type="first">A</forename><surname>Rauber</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><surname>Müller-Kögler</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">ACM/IEEE joint Conference on Digital Libraries</title>
				<meeting><address><addrLine>Roanoke, USA</addrLine></address></meeting>
		<imprint>
			<date type="published" when="2001">2001</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b19">
	<monogr>
		<title level="m" type="main">The Power of Genre</title>
		<author>
			<persName><forename type="first">A</forename><surname>Rosmarin</surname></persName>
		</author>
		<imprint>
			<date type="published" when="1985">1985</date>
			<publisher>University of Minnesota Press</publisher>
			<pubPlace>Minneapolis</pubPlace>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b20">
	<analytic>
		<title level="a" type="main">Genres In Formation? An Exploratory Study of Web Pages using Cluster Analysis</title>
		<author>
			<persName><forename type="first">M</forename><surname>Santini</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Proc. CLUK</title>
		<imprint>
			<biblScope unit="volume">05</biblScope>
			<date type="published" when="2005">2005</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b21">
	<analytic>
		<title level="a" type="main">Some Issues in Automatic Genre Classification of Web Pages</title>
		<author>
			<persName><forename type="first">M</forename><surname>Santini</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proc. of the JADT 2006</title>
				<meeting>of the JADT 2006<address><addrLine>Besançon</addrLine></address></meeting>
		<imprint>
			<date type="published" when="2006">2006</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b22">
	<monogr>
		<title level="m" type="main">Genre Analysis</title>
		<author>
			<persName><forename type="first">J</forename><surname>Swales</surname></persName>
		</author>
		<imprint>
			<date type="published" when="1990">1990</date>
			<publisher>Cambridge University Press</publisher>
			<pubPlace>Cambridge</pubPlace>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b23">
	<monogr>
		<title level="m" type="main">Analysing Professional Genres</title>
		<editor>Trosborg, A.</editor>
		<imprint>
			<date type="published" when="2000">2000</date>
			<publisher>J. Benjamins Publishing Company</publisher>
			<pubPlace>Amsterdam</pubPlace>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b24">
	<monogr>
		<title level="m" type="main">Data Mining: Practical Machine Learning Tools and Techniques</title>
		<author>
			<persName><forename type="first">I</forename><surname>Witten</surname></persName>
		</author>
		<author>
			<persName><forename type="first">E</forename><surname>Frank</surname></persName>
		</author>
		<imprint>
			<date type="published" when="2005">2005</date>
			<publisher>Morgan Kaufmann Publishers</publisher>
			<pubPlace>Amsterdam</pubPlace>
		</imprint>
	</monogr>
	<note>second edition</note>
</biblStruct>

<biblStruct xml:id="b25">
	<analytic>
		<title level="a" type="main">Genres of organizational communication: A structural approach to studying communications and media</title>
		<author>
			<persName><forename type="first">J</forename><surname>Yates</surname></persName>
		</author>
		<author>
			<persName><forename type="first">W</forename><surname>Orlikowski</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Academy of Management Review</title>
		<imprint>
			<biblScope unit="volume">17</biblScope>
			<biblScope unit="issue">2</biblScope>
			<biblScope unit="page" from="229" to="326" />
			<date type="published" when="1992">1992</date>
		</imprint>
	</monogr>
</biblStruct>

				</listBibl>
			</div>
		</back>
	</text>
</TEI>
