ICWE 2008 Workshops, 7th Int. Workshop on Web-Oriented Software Technologies – IWWOST 2008 JobOlize Headhunting by Information Extraction in the Era of Web 2.0 Christina Buttinger1, Birgit Pröll1, Jürgen Palkoska1, Werner Retschitzegger2, Manfred Schauer3, Reinhold Immler3 1 Institute for Application Oriented Knowledge Processing (FAW) 2 Information Systems Group (IFS) Johannes Kepler University Linz, Austria 3 JoinVision E-Services GmbH, Austria {christina.buttinger, birgit.proell, juergen.palkoska, werner.retschitzegger}@jku.at {manfred.schauer, reinhold.immler}@joinvision.com Abstract manage the large amount of job offers on the Web and to adequately support, both, headhunters and job seekers, the E-recruitment is one of the most successful e-business number of job portals is increasing rapidly (cf., e.g., Mon- applications supporting both, headhunters and job seek- ster [18], Stepstone [24] or Jobscout24 [12]), some of ers. The explosive growth of online job offers makes the them even allowing matchmaking between offers and pro- usage of information extraction techniques to build up, files of job seekers. e.g., job portals in a semi-automatic way a necessity. Ex- A crucial prerequisite for these systems is the usage of isting approaches, however, hardly cope with the hetero- proper techniques for automating the extraction of rele- geneous and semi-structured nature of job offers nor do vant information from available job offers. Many existing they consider potentials offered by Web 2.0 technologies. information extraction techniques have in common that This paper proposes an information extraction system a target Web page is automatically atomized into its con- called “JobOlize”1, realized for arbitrarily structured IT stituents called tokens, which further might be annotated job offers. To improve extraction quality, a hybrid ap- on basis of a vocabulary defined by a common schema. proach is employed, combining existing NLP-techniques Some of these techniques focus more on Web pages hav- with a new form of context-driven extraction, incorporat- ing a more or less common underlying structure and per- ing layout, structure and content information. To allow form screen scraping, meaning that information is ex- users a proper adaptation of the extraction results while tracted based on, e.g., the HTML tags used, whereas oth- preserving the look and feel of the original Web pages, ers emphasize more on completely unstructured text and a rich client interface is provided. The improvements in employ natural language processing (NLP), analyzing, extraction quality are justified on basis of a case study e.g., the grammar of the content’s language (for a survey and the experiences gained are generalized and critically of extraction techniques, cf., e.g., [14]). reflected by discussing lessons learned. Considering the area of e-recruitment, job offers fall typically under both categories – although a common 1. Introduction structure between them can seldom be found, there are at least structured parts, which could be used for improving The Internet has already become the most important the quality of the extraction process. Existing techniques medium for recruitment processes. Over the past years, the hardly cope with this heterogeneous and semi-structured e-recruitment market grew steady. In the EU in 2007, 70% nature, thus decreasing the quality of extraction results. In of all job offers were published online, and more than addition, the environments provided by existing ap- 56% of employments were results of online offers [6]. To proaches for improving the extraction results manually, e.g., by manipulating the annotations of the extracted in- formation, is most often realized as a heavy-weight desk- 1 This prototype, which has been partly funded by the Austrian Research top application, taking no advantage of upcoming Web 2.0 Promotion Agency FFG under grant 813202, has been used for further techniques [25]. development of the Austrian online job portal www.joinvision.com. 26 ICWE 2008 Workshops, 7th Int. Workshop on Web-Oriented Software Technologies – IWWOST 2008 This paper proposes an information extraction system management, backed up by our experience in developing called “JobOlize”, which is realized for dealing with arbi- ontologies [20], [22], [23]. Most notably, from the ontol- trarily structured IT job offers. To improve extraction ogy of Mochol et al. [17] which is based on the compre- quality, a hybrid approach is proposed which combines hensive HR-XML standard [10], we have adopted the con- existing NLP-techniques to deal with completely unstruc- cept of skill levels, whereas the work of Garcia et al. [7] tured information, with a new form of context-driven ex- was influential for our notion of offer characteristics. Our traction, incorporating layout, structure and content infor- ontology also considers specialisation hierarchies of IT mation. As demonstrated in a case study, this hybrid ap- skills (e.g., “Oracle“ and “MySQL” are specializations of proach considerably improves the quality of the extracted “DBS”) and equivalence relationships (e.g., “DB” is se- information. To allow users a proper adaptation of the mantically equivalent to “DBS”), which are of particular extraction results while preserving the look and feel of the interest for the extraction process. Finally, to consider also original Web page, a rich client interface is provided using job offers in different languages, each concept contains Mozilla’s Firefox extensions, and the XML-based GUI a language property. language XUL [26]. The results and experiences gained during the development of JobOlize have been used for further development of the Austrian online IT job portal JoinVision (www.joinvision.com). The paper is structured as follows. Section 2 presents the architecture of the JobOlize prototype, especially fo- cusing on context-driven extraction together with the rich client interface. Section 3 discusses a case study evaluat- ing the improvements in extraction quality and provides a comparison to related approaches. Section 4 concludes the paper with a critical reflection in terms of lessons learned. 2. The JobOlize Prototype Figure 1. Overall Architecture of JobOlize. This section presents the architecture of JobOlize, es- Extraction Rule Base. The second part of our knowl- pecially focusing on the context-driven extraction process edge base contains about 50 extraction rules ranging from and describes functionality and implementation aspects of very simple ones, responsible, e.g., for matching input the system’s rich client interface realized on basis of Web tokens with ontology concepts to rather complex ones for, 2.0 techniques. e.g., job title detection. The rules are in fact condition- action tuples formulated using the regular-expression- 2.1 Context-Driven Information Extraction based language JAPE (Java Annotation Pattern Engine) provided by GATE and callouts to Java code for realizing The architecture of JobOlize, depicted in Figure 1, is more complex tasks. divided into two core components, a knowledge base pro- Pipeline of Extraction Components. On basis of the viding a domain ontology as well as an extraction rule second core part of our architecture, the pipeline, annota- base and a pipeline consisting of different extraction com- tions of web pages are incrementally built up by streaming ponents. Parts of the system are realized on basis of the input pages through the different components of the pipe- existing information extraction infrastructure GATE (Gen- line, each of them being responsible for a certain annota- eral Architecture for Text Engineering) [3] which has tion task, thereby adding new or manipulating already been favored mainly because it is a mature open source existing annotations. Parts of this pipeline are realized on implementation providing a highly extensible pluggable basis of GATE which provides not only predefined, plug- architecture. gable components for realizing basic extraction tasks but e-Recruitment Domain Ontology. For representing the also allows to plug-in self-developed, customized ones. To annotation vocabulary used by our extraction system, be more specific, JobOlize reuses existing NLP-based a light-weight domain ontology has been developed con- components for tokenizing and stemming and furthermore taining 15 core concepts of job offers which should be realizes customized components on basis of Java for pre- extracted (e.g., job title, IT-skill, language skill, operation processing the input Web pages (e.g., eliminating area and graduation). A design goal in this respect was to JavaScript code), for post-processing (e.g., exporting an- heavily build on reasonable concepts of existing ontolo- notated Web pages as XML documents) and for the core gies in the area of e-recruitment and human resource (HR) task of identifying those tokens of the input Web page, 27 ICWE 2008 Workshops, 7th Int. Workshop on Web-Oriented Software Technologies – IWWOST 2008 which are relevant for our purposes, e.g., IT-skill with an part of the bottom part, it gets a relevance of 0,25, only. associated skill-level like “basic knowledge in Java pro- During post-processing, annotated IT-skills with a rele- gramming”. The identification of relevant tokens is per- vance lower than a certain threshold, are eliminated from formed on basis of four different customized components, the result. as described in the following, which are making use of the above mentioned knowledge base. 2.2 Rich Client Interface for Annotations Initial Annotation. The first component is responsible for an initial annotation of the tokens of the input Web For ensuring an acceptable quality of the extraction re- page with appropriate concepts defined by our domain sults, information extraction out of arbitrarily structured ontology. For most of the tokens, this task is straightfor- and heterogeneous job offers is reasonable in a semi- ward, but considering the domain concepts IT-skill and automatic way only. Thus, the results generated automati- language skill, context-driven processing is required in cally by the extraction system, which are in our case, an- order to determine their corresponding levels. In particu- notations of the target Web pages containing the job of- lar, not only the position of the skill level with respect to fers, have to be assessed by a user and potentially cor- the skill type itself is taken into account by means of ap- rected accordingly. For this task, we provide a rich client propriate JAPE rules, but also, e.g., if it is located within interface which allows the user not only to initiate the ex- the same sentence or not. traction process by entering the URL’s to the desired tar- Page Segmentation. The remaining three components get Web pages, but especially to visualize the results of promote the basic idea of context-driven extraction even the extraction process in terms of the annotated Web further, with the ultimate goal to improve the quality of pages and the possibility to manipulate existing annota- the extraction results. For this, the Web page is first seg- tions or to add additional ones directly in a Web browser2. mented into three parts, a top part, a content part and The screenshot in Figure 2 depicts on the left hand side a bottom part, simply taking the content itself into ac- a menu sidebar and on the right hand side the original count. This is done by identifying common text fragments Web page. between two job offers of the same Web site on basis of the well-known Longest-Common-Subsequence algorithm (LCS) [11]. These common text fragments represent the top and bottom parts of a page and contain tokens which are most probably irrelevant for further processing. For example, the occurrence of “powered by Typo3” within the bottom page would lead to an IT-skill annotation of “Typo3” during the initial annotation phase, being refined by page segmentation, i.e., classified as irrelevant. Block Identification. The content part of the Web page identified before, is further divided into so-called blocks, representing a couple of tokens which “visually” belong together and are normally grouped under a header title. For the identification of such blocks, first, context in terms of layout information (e.g., a -tag) and structural in- formation (e.g., a
  • -tag) is considered by appropriate JAPE rules. In a second step, context in form of content information is used to categorize the identified blocks. In particular, on basis of their header title and the corre- Figure 2. Rich Client Interface of JobOlize. sponding domain concepts defined by our ontology, blocks are categorized and annotated as requirements, Annotation Highlighting. Within the sidebar, the user responsibilities, offer characteristics and contact details, can choose from an annotation list (cf. c), representing i.e., those chunks of information most commonly found in the concepts of our domain ontology, which of the auto- job offers. matically generated annotations should be indicated within Relevance Assignment. The final component of our ex- the Web page on the right hand side. To preserve the look traction process assigns pre-defined relevance values, be- and feel of the original Web page, annotations are indi- tween 0 and 1, to the initial annotations, depending on the cated by marking the concerned parts of the Web page block category, the annotated token is contained in. E.g., in case that a token annotated as IT-skill is part of the re- 2 It has to be noted that the GUI provided by GATE is realized as quirements block it gets a relevance of 1, whereas if it is a heavy-weight desktop application only, which among others, lack a proper visualization of the extraction results. 28 ICWE 2008 Workshops, 7th Int. Workshop on Web-Oriented Software Technologies – IWWOST 2008 with corresponding colors and highlighting those, being tion extraction, comprising three different measures [5]: currently focused on, by means of a flash effect. First, precision, which reflects the share of extracted con- Annotation Details. The details of annotations (e.g., the tent which is relevant among all extracted content. Sec- relevance values assigned) are shown for each annotation, ond, recall, which depicts the share of extracted content both, within the sidebar (cf. d) and directly within the which is relevant among all content. Third, F-measure, Web page, activated by a mouse over effect. This func- which is a balanced measure, combining recall and preci- tionality also serves as a simple form of explanation com- sion. These measures were computed using the Annota- ponent, which allows tracing back the genesis of a certain tionDiff tool, provided by GATE [8]. annotation. Interpretation of Results. Figure 3 shows the results of Annotation Manipulation. Finally, the details of each of the case study. Overall, it can be seen that our context- the automatically generated annotations can be manipu- driven extraction process outperforms a non-context- lated in accordance with the underlying job ontology, the driven extraction, as expected. The values for recall were annotation can be completely deleted or new ones can be equal in all three cases, simply due to the fact that the ini- added. Thus, e.g., a language skill can be modified from tial annotation determines all potentially relevant token German to English (cf. e). Manipulations are immedi- candidates, while the subsequent components of our ex- ately propagated to the Web page at the right hand side. traction pipeline do not further contribute to the recall Implementation Aspects. Concerning implementation values. aspects, the rich client interface is realized as a Firefox The most significant gain in extraction quality could be extension using XUL (XML User Interface Language observed regarding the ontology concept operation area. [26]), allowing for portability of the application, Precision and F-measure showed an increase of 64% and JavaScript and CSS. The primary reason to favor the Fire- 38%, respectively. The reason for this improvement is fox extension mechanisms instead of alternatives like that, in several Web pages of the test set, operation areas Apache’s MyFaces [19], was the rather easy realization of also describe the business focus of the company, not only a rich client interface experience as in native desktop ap- the profile required from a job seeker. This completely plications, considerably reducing programming effort. different intention of using operation areas can be detected Annotation highlighting as well as manipulation is based by our context-driven extraction process, since operation on DOM. As a pre-requisite for asynchronously initiating areas of a company are naturally positioned in the top part the extraction process (executed on the server) directly of a Web page, whereas operation areas required from from within the rich client interface, the underlying ex- a job seeker occur within a dedicate block of the content traction system had to be ported to the used Tomcat Web part. server. 1,2000 3. Evaluation of JobOlize 1,0000 0,8000 0,6000 This section is dedicated to an evaluation of JobOlize, 0,4000 first on basis of a case study and second by a comparison 0,2000 0,0000 to related work. Precision Precision Precision F-Meas. F-Meas. F-Meas. Recall Recall Recall 3.1 Case Study Operation IT-Skill Language Context-driven Extraction The goal of this case study is to evaluate the amount of Area Skill Non Context-driven Extraction improvement in extraction quality when employing our context-driven extraction process, compared to a simple ontology- and rule-based annotation (called “initial anno- Figure 3. Results of the Evaluation. tation” in Section 2.1). Set-up of the Case Study. The test set of our case study consisted of 32 randomly chosen online IT job offers from Precision and F-measure for IT-skills increased by 30% 16 different Austrian recruitment Web sites, partly differ- and 20%, respectively. The reason for this improvement is ing considerably in content, structure, and layout. The similar to that given for operation areas – there are job extraction quality has been tested for three different con- offers where pretended IT-skills show up at the bottom cepts of our job ontology, namely IT skill, operation area, part of the Web page as already argued in Section 2.1. Our and language skill, since each of them poses unique chal- context-driven extraction process allows again eliminating lenges to the extraction process. Finally, we employed the these false positives, although there seems to be not that common metric used for evaluating the quality of informa- many false positive IT-skills than operations areas out 29 ICWE 2008 Workshops, 7th Int. Workshop on Web-Oriented Software Technologies – IWWOST 2008 there, as indicated by the lower improvement in extraction tributions of this paper, the lessons learned are grouped quality. into three categories. Finally, the least improvement could be observed for language skills, being nearly the same for precision (8%) 4.1 Context-Driven Information Extraction and F-measure (6%). This rather marginal advance is due to the fact that language skills are more or less unambigu- Combining different kinds of context counteracts intri- ously used in job offers, seldom leading to a false positive cacies of HTML encodings. As one would expect, the lib- content extraction. erty of users to ignore any accessibility or well- formedness criteria when designing HTML pages, using 3.2 Comparison to Related Work e.g., proprietary enumerations instead of explicit
  • - tags, different font sizes instead of explicit -tags or There are already numerous information extraction HTML tables instead of CSS, considerably challenges the techniques available [14]. Considering our rich client in- extraction process. It has been shown that the combination terface first, most existing systems build, as already men- of different kinds of context as done in JobOlize eases the tioned, on desktop-based interfaces. The remaining ones, task to deal also with these intricacies of HTML encoding e.g., KIM [21] provide restricted functionality only, in that in a proper way. they visualize the annotated Web page only, but do not Eliminating irrelevant parts of a web page considered allow arbitrary manipulations. harmful. One has to be aware that eliminating parts of Second, regarding the extraction itself, most existing a Web page which are presumably considered as irrelevant techniques represent the targeted Web page in form of as done in JobOlize on basis of the LCS algorithm can a DOM tree. On this basis, some of them just separate also be counterproductive, e.g., in case that the eliminated content from non-content parts (cf. e.g., [4]), often focus- part indeed contains links to further skills, self refreshing ing on certain kinds of structure only, such as tables (cf., date items, or RSS-based news-tickers. In addition, the e.g.,[16]). Other approaches add additional context to im- accurateness of the used algorithm completely depends on prove extraction quality, either in terms of content infor- a proper choice of the two Web pages used for initial com- mation (cf. e.g., the “Pagelet” approach of [1]) or using parison. When choosing two job offers with, e.g., similar structure (cf. the WISDOM approach of [13] and the tem- skill demands, these blocks could, because of their over- plate detection approach of [2]), or layout information (cf. lap, be falsely classified as irrelevant. e.g., the “Styletree” approach of [27] and the visual fea- ture approach of [9]). 4.2 Rich Client User Interface for Annotations Compared to these systems, we approach the extraction process from a different angle, since we do not represent Firefox extensions and XUL proved to be adequate for the input Web page as a DOM tree but let our extraction rich client interfaces. Our experience of using Firefox system directly work with the original Web page, thus extension mechanisms and XUL confirms that the devel- preventing to be dependent on the decreased fidelity of opment of a rich client interface is possible without much DOM representations. Another major difference is that training and programming effort, being additionally facili- JobOlize does not focus on certain structures of the Web tated by a comprehensive documentation and a steadily pages but rather deals with arbitrary structures, which is enlarging developer community. Nevertheless, as common made possible by the combination of NLP techniques of in the Web Engineering area, one must expect permanent the underlying GATE infrastructure and our context- changes of the underlying SW-components which re- driven extraction process. And finally, we do not consider quired, in our case, due to fundamental changes to the a certain kind of context in isolation, but rather consecu- Firefox API from version 2 to 3, considerable re-coding. tively apply all kinds of context available, i.e., content, Annotations partly conflicting with original HTML- structure and presentation information to produce the re- tags. To visualize also the annotated Web page within the sulting annotations and calculate corresponding relevance rich client interface, JobOlize currently merges the values. HTML-tags of the original page with the tags representing the annotations. There are, however, cases where this 4. Concluding Lessons Learned leads to irresolvable HTML validity conflicts, for exam- ple, if a title element overlaps an end-of-list HTML-tag. The experiences gained in the course of developing our Currently, our rich client interface is not able to visualize system JobOlize for the domain of e-recruitment are now annotations causing such conflicts, requiring more ad- critically reflected and generalized in order to provide vanced visualization techniques such as an overlay con- valuable lessons learned also for other domains relying on cept. information extraction techniques. According to the con- 30 ICWE 2008 Workshops, 7th Int. Workshop on Web-Oriented Software Technologies – IWWOST 2008 4.3 Application of Extraction Systems [7] Garcia-Sanchez, F. et al. “An ontology-based intelligent system for recruitment”. In Expert Systems with Applications, 31(2), Aug. 2006, pp. 248-263. Guidelines how to apply information extraction systems [8] GATE User Guide, http://gate.ac.uk/sale (accessed May 2008) required. Not least due to the high degree of extensibility provided by GATE, guidelines how to apply an extraction [9] Gatterbauer, W., et al. “Towards domain-independent informa- tion extraction from web tables”. In Proc. of the 16th Int. World system depending on certain characteristics of the domain Wide Web Conf. (WWW), Alberta, May 2007. are urgently needed. Especially the specification of extrac- [10] HR XML, http://www.hr-xml.org (accessed May 2008) tion rules for JobOlize turned out as a trial and error proc- [11] Hunt, J. W. and Szymanski, T. G. “A fast algorithm for com- ess, getting completely lost in several rule dependencies puting longest common subsequences”. CACM 20(5), May which might be alleviated by a visual (pattern-based) rule- 1977. editor or by some kind of decision support system, apply- [12] JobScout24, http://www.jobscout24.de (accessed May 2008) ing further AI expertise to the area of information extrac- [13] Kao, H., Ho, J. and Chen, M. “WISDOM: Web Intrapage In- tion. formative Structure Mining Based on Document Object Measures for evaluating the quality of information ex- Model”. IEEE Transactions on Knowledge and Data Engineer- traction needed. During the case study, it emerged that, ing, 17(5), May, 2005, pp. 614-627. e.g., for evaluating the skill levels, a simple Boolean [14] Kayed, M. and Shaalan, K. F. “A Survey of Web Information measure in terms of true positives and false positives is Extraction Systems”. IEEE Trans. on Knowledge and Data En- not adequate. For example, instead of “perfect” for a skill gineering 18(10), Oct. 2006. level the slightly different value “very good” might have [15] Lavelli, A., et al. “A critical survey of the methodology for IE been extracted or instead of “Oracle” for a skill the more Evaluation.” Int. Conf. on Lang. Resources & Evaluation, 2004. generalize concept “DBS” might have been extracted. [16] Lin, S. and Ho, J. “Discovering informative content blocks from These examples argue for a measure, which takes a simi- Web documents”. In Proc. of the 8th ACM SIGKDD Int. Conf. larity aspect into account, e.g., derived from relations de- on Knowledge Discovery and Data Mining (KDD), Alberta, July, 2002, pp. 588-593. fined within the ontology. These explorations confirm the [17] Mochol, M., Paslaru, E. and Simperl, B. “Practical Guidelines argumentation of Lavelli et al. [15], that the field still for Building Semantic eRecruitment Applications”. Proc. Of the lacks standard datasets, evaluation procedures and appro- Int. Conf. on Knowledge Mgmt. (iKnow)”, Austria, Sept. 2006. priate measures. [18] Monster, http://www.monster.com (accessed May 2008) [19] Myfaces, http://www.myfaces.org (accessed May 2008) Acknowledgements: We kindly thank our colleague Simon [20] Palkoska, J., et al. “On Cooperatively Creating Dynamic On- Thalbauer for valuable contributions to the realization of tologies”. In Proc. of Int. Conf. on Hypertext, Salzburg, 2005. Jobolize. [21] Popov, B., et al. ”KIM - A Semantic Platform For Information Extraction and Retrieval”. In Journal of Natural Language En- References gineering, 10(3-4), Sep 2004. [22] Baumgartner, N., Retschitzegger, W. and Schwinger, W. ”Lift- [1] Bar-Yossef, Z. and Rajagopalan, S. “Template detection via ing Metamodels to Ontologies – a Step to the Semantic Integra- data mining and its applications”. In Proc. of the Int. World tion of Modeling Languages”. In Proc. of the 9th Int. Conf. on Wide Web Conf. (WWW), Honolulu, May 2002. Model Driven Engineering Languages and Systems (MODELS). [2] Chakrabarti, D., Kumar, R., and Punera, K. “Page-level tem- Springer, 2006. plate detection via isotonic smoothing.” In Proc. of the 16th Int. [23] Baumgartner, N., Retschitzegger, W. and Schwinger, W. World Wide Web Conf. (WWW), Alberta, 2007. “A Software Architecture for Ontology-Driven Situation [3] Cunningham, H., et al. “GATE: A Framework and Graphical Awareness”. In Proc. of the 23rd ACM Symp. on Applied Development Environment for Robust NLP Tools and Applica- Comp. (ACM SAC), Fortaleza, March 2008. tions”. In Proc. of the 40th Anniversary Meeting of the Associa- [24] Stepstone, http://www.stepstone.com (accessed May 2008) tion for Comput. Linguistics (ACL), Philadelphia, 2002. [25] Uren, V, et al. “Semantic Annotation for Knowledge Manage- [4] Debnath, S., Mitra, P. Pal, N. and Lee Giles, C. “Automatic ment: Requirements and a Survey of the State of the Art”. Web Identification of Informative Sections of Web Pages”, IEEE Semantics: Science, Services and Agents on the World Wide Transactions on Knowledge and Data Enginering, 17(09), Sept. Web 4(1), 2006, pp. 14-28. 2005, pp. 1233-1246. [26] XUL, http://developer.mozilla.org (accessed May 2008). [5] Douthat, A., “The message understanding conference scoring [27] Yi, L., Liu, B., and Li, X. “Eliminating noisy information in software user’s manual”. In Proc. of the 7th message Under- Web pages for data mining”. In Proc. of the Int. Conf. on standing Conference (MUC7), 1998 Knowledge Discovery & Data Mining (KDD), Aug., 2003. [6] Eckhardt, A., König, W., Weitzel, T. and Deininger, K. “Re- cruiting Trends 2007 – European Union – An empirical survey with the top 1.000-enterp. in the EU”, No. 2007-798. 31