<?xml version="1.0" encoding="UTF-8"?>
<TEI xml:space="preserve" xmlns="http://www.tei-c.org/ns/1.0" 
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" 
xsi:schemaLocation="http://www.tei-c.org/ns/1.0 https://raw.githubusercontent.com/kermitt2/grobid/master/grobid-home/schemas/xsd/Grobid.xsd"
 xmlns:xlink="http://www.w3.org/1999/xlink">
	<teiHeader xml:lang="en">
		<fileDesc>
			<titleStmt>
				<title level="a" type="main">Towards efficient scoring of student-generated long-form analogies in STEM</title>
			</titleStmt>
			<publicationStmt>
				<publisher/>
				<availability status="unknown"><licence/></availability>
			</publicationStmt>
			<sourceDesc>
				<biblStruct>
					<analytic>
						<author>
							<persName><forename type="first">Thilini</forename><surname>Wijesiriwardene</surname></persName>
							<email>thilini@sc.edu</email>
							<affiliation key="aff0">
								<orgName type="department">AI Institute</orgName>
								<orgName type="institution">University of South Carolina</orgName>
								<address>
									<settlement>Columbia</settlement>
									<region>SC</region>
									<country key="US">USA</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Ruwan</forename><surname>Wickramarachchi</surname></persName>
							<affiliation key="aff0">
								<orgName type="department">AI Institute</orgName>
								<orgName type="institution">University of South Carolina</orgName>
								<address>
									<settlement>Columbia</settlement>
									<region>SC</region>
									<country key="US">USA</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Valerie</forename><forename type="middle">L</forename><surname>Shalin</surname></persName>
							<email>valerie.shalin@wright.edu</email>
							<affiliation key="aff0">
								<orgName type="department">AI Institute</orgName>
								<orgName type="institution">University of South Carolina</orgName>
								<address>
									<settlement>Columbia</settlement>
									<region>SC</region>
									<country key="US">USA</country>
								</address>
							</affiliation>
							<affiliation key="aff1">
								<orgName type="department">Department of Psychology</orgName>
								<orgName type="institution">Wright State University</orgName>
								<address>
									<settlement>Dayton</settlement>
									<region>OH</region>
									<country key="US">USA</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Amit</forename><forename type="middle">P</forename><surname>Sheth</surname></persName>
							<affiliation key="aff0">
								<orgName type="department">AI Institute</orgName>
								<orgName type="institution">University of South Carolina</orgName>
								<address>
									<settlement>Columbia</settlement>
									<region>SC</region>
									<country key="US">USA</country>
								</address>
							</affiliation>
						</author>
						<title level="a" type="main">Towards efficient scoring of student-generated long-form analogies in STEM</title>
					</analytic>
					<monogr>
						<idno type="ISSN">1613-0073</idno>
					</monogr>
					<idno type="MD5">4B25EF15F07FB7AD2498B60365F107D3</idno>
				</biblStruct>
			</sourceDesc>
		</fileDesc>
		<encodingDesc>
			<appInfo>
				<application version="0.7.2" ident="GROBID" when="2023-06-19T14:46+0000">
					<desc>GROBID - A machine learning software for extracting information from scholarly documents</desc>
					<ref target="https://github.com/kermitt2/grobid"/>
				</application>
			</appInfo>
		</encodingDesc>
		<profileDesc>
			<textClass>
				<keywords>
					<term>Descriptive analogies</term>
					<term>Analogical features</term>
					<term>Analogy scoring</term>
					<term>Long-form analogies</term>
				</keywords>
			</textClass>
			<abstract>
<div xmlns="http://www.tei-c.org/ns/1.0"><p>Switching from an analogy pedagogy based on comprehension to analogy pedagogy based on production raises an impractical manual analogy scoring problem. Conventional symbol-matching approaches to computational analogy evaluation focus on positive cases, and challenge computational feasibility. This work presents the Discriminative Analogy Features (DAF) pipeline to identify the discriminative features of strong and weak long-form text analogies. We introduce four feature categories (semantic, syntactic, sentiment, and statistical) used with supervised vector-based learning methods to discriminate between strong and weak analogies. Using a modestly sized vector of engineered features with SVM attains a 0.67 macro F1 score. While a semantic feature is the most discriminative, out of the top 15 discriminative features, most are syntactic. Combining this engineered features with an ELMo-generated embedding still improves classification relative to an embedding alone. While an unsupervised K-Means clustering-based approach falls short, similar hints of improvement appear when inputs include the engineered features used in supervised learning.</p></div>
			</abstract>
		</profileDesc>
	</teiHeader>
	<text xml:lang="en">
		<body>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1.">Introduction</head><p>Analogical reasoning relies on the ability to draw on the relational similarities between two systems of objects in different contexts <ref type="bibr" target="#b0">[1,</ref><ref type="bibr" target="#b1">2,</ref><ref type="bibr" target="#b2">3]</ref>. Analogies appear in several disciplines such as engineering design, scientific reasoning, and often in STEM education. However, the dominant pedagogical paradigm requires students to comprehend curated analogies. In this work, we are focusing on the evaluation of student-generated analogies in their first undergraduate biochemistry course.</p><p>Problem sets, specifically created to explore the underlying mechanisms of analogical reasoning, consist of visual and verbal analogies <ref type="bibr" target="#b3">[4,</ref><ref type="bibr" target="#b4">5]</ref>. Verbal analogies have two primary forms; analogical proportions and long-form analogies. Analogical proportions follow a four-term format such as "A to B as to C to D" or A:B::C:D <ref type="bibr" target="#b5">[6]</ref>. Recent work on computational analogy making focuses on analogical proportions <ref type="bibr" target="#b6">[7,</ref><ref type="bibr" target="#b7">8,</ref><ref type="bibr" target="#b8">9]</ref>. Our interest here lies in long-form analogies consisting of a narrative/ description of a target unfamiliar situation/ system (for the context or lesson to be learned) using several sentences and a familiar source (base) <ref type="bibr" target="#b9">[10,</ref><ref type="bibr" target="#b2">3]</ref>. While the objects across the two descriptions differ, they employ similar relations between these objects. An example of a well-known long-form analogy is between the solar system (source) and the Rutherford-Bohr model of the atom (target) <ref type="bibr" target="#b10">[11]</ref> where small objects revolving around a large central object provide relational similarity with the target. The solar system and the atom can each be described using several sentences. Parallels between these two systems can then be drawn, making the two descriptions analogous.</p><p>The atom-solar system analogy exemplifies the curated analogies in STEM textbooks. <ref type="bibr" target="#b11">[12]</ref> has developed algorithms for evaluating correct or slightly incorrect long-form analogies. In <ref type="bibr" target="#b12">[13]</ref> we solicited analogies from STEM students, with the expectation, relative to a comprehension exercise, that analogy production is both more engaging and allows students to employ existing familiar knowledge to scaffold the acquisition of new knowledge. No matter how pedagogically successful, manual scoring is impractical for an analogy production pedagogy. Production pedagogy elevates an analogy scoring problem for computational solution.</p><p>We aim to identify discriminative features between strong and weak, long-form student generated verbal analogies collected in a college biochemistry class (see Section 2.1) to support efficient computational scoring. To this end, we develop the Discriminative Analogy Features (DAF) pipeline.</p><p>We use a long-form analogy dataset, instructor-graded as strong or weak. We explore both supervised and unsupervised learning classifiers using vectors based on embeddings, engineered features and both. Given manually annotated data for a supervised learning classifier (i.e. SVM), we identify the discriminative features of strong and weak analogies.</p><p>We introduce DAF, a pipeline to identify the discriminative features of strong and weak analogies. We also introduce four feature categories -semantic, syntactic, sentiment, and statistical used in supervised learning to discriminate between strong and weak analogies. We show that "unique attribute count", a semantic feature, is the most discriminative when identifying between strong and weak analogies. Out of the top 15 discriminative features, most are syntactic. Unsupervised learning is unable to obtain comparable success, though it slightly improves with features corresponding to the above categories.</p><p>The rest of this paper is organized as follows: Section 2 introduces and describes the DAF pipeline and identifies the discriminative features. Section 3 presents the discussion with findings, insights, limitations, and future work subsections. Section 4 concludes the paper.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.">Discriminative Analogy Features (DAF) Pipeline</head><p>To identify discriminative features, we introduce the pipeline illustrated in Figure <ref type="figure" target="#fig_0">1</ref>. In the subsequent subsections, we describe each pipeline component. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.1.">Dataset and Input</head><p>The dataset used in this work was drawn from 500 student-created analogies, collected in a college classroom. An instructor explained a process in the domain of biochemistry, e.g., Glycolysis (source analogy), and requested students to construct a scenario analogous to the explained process from a domain of their choice (target analogy). The instructor then evaluated 31 studentgenerated analogous scenarios as a strong or weak analogy based on its correspondence to the Biochemistry concept. A strong analogy corresponds well with the target analogy, and a weak analogy minimally corresponds with the target analogy. To increase the size of the 31 exemplar data set from the original we split each analogy into its constituent sentences, generating a data set of 526 strong exemplars and 140 weak exemplars. Each constituent sentence of an analogy falls into the same annotation category as the original analogy. Ergo, the initial input to the DAF pipeline is a sentence. This work does not distinguish between the analogy's target domains (Enzyme Kinetics and Glycolysis). Table <ref type="table" target="#tab_0">1</ref> presents the summarized statistics of the dataset. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.2.">Input Processing</head><p>Sentences were processed and used as inputs to a Support Vector Machine (SVM) classifier (supervised learning) and K-means clustering (unsupervised learning) separately. In the following section we briefly review the background of input processing techniques, learning methods and implementation details.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.2.1.">Background</head><p>SVM is a supervised learning technique that creates functions to map inputs to pre-existing annotations <ref type="bibr" target="#b13">[14]</ref>. SVM is an easy-to-interpret classifier providing competitive performance in classification, regression, and outlier detection tasks <ref type="bibr" target="#b14">[15]</ref>. The following paragraphs detail the background of four feature groups of interest here.</p><p>The obviously relevant features are semantic. Abstract Meaning Representation (AMR) is a semantic representation language that expresses a sentence's logical meaning by converting it to a rooted, directed, acyclic, edge-labeled, and leaf-labeled graph. <ref type="bibr" target="#b15">[16]</ref>. To abstract away from syntactic idiosyncrasies, AMR assigns the same AMR graph to sentences with the same meaning. Nodes of an AMR graph are labeled as concepts, edges as relations, and concept properties as attributes. Concepts are either English words, PropBank framesets <ref type="bibr" target="#b16">[17]</ref> or special keywords. There are approximately 100 relations <ref type="bibr" target="#b15">[16]</ref>. AMR is used as a semantic representation of text in several NLP tasks such as summarization <ref type="bibr" target="#b17">[18]</ref>, machine comprehension <ref type="bibr" target="#b18">[19]</ref>, and event extraction <ref type="bibr" target="#b19">[20,</ref><ref type="bibr" target="#b20">21]</ref>. In this work we use AMR representations to extract concepts, relations and attributes present in sentence analogies. Figure <ref type="figure" target="#fig_1">2</ref> illustrates the AMR for a sentence from the dataset.</p><p>Sentiment-based features potentially reveal student engagement. Sentiment analysis aims to identify emotional or affective tendencies in user-generated content such as tweets, product reviews, and feedback <ref type="bibr" target="#b21">[22]</ref>. Subjectivity detection and polarity determination are two common tasks in sentiment analysis <ref type="bibr" target="#b22">[23]</ref>. Subjectivity quantifies the personal opinions versus factual information contained in the text. High subjectivity indicates the text contains more personal opinions compared to factual information <ref type="bibr" target="#b22">[23]</ref>. Polarity describes the sentiment of a piece of text as positive, negative, or neutral <ref type="bibr" target="#b21">[22]</ref>.</p><p>We extract three groups of syntactic features. The first feature group is Part of Speech (POS), a grammatical classification of the word types in a sentence. These POS tags commonly include nouns, verbs, adjectives, etc. <ref type="bibr" target="#b23">[24]</ref>. Named Entities Recognition (NER), the second feature group, is used to identify occurrences of named entities such as people, organizations, times, and locations in a sentence <ref type="bibr" target="#b24">[25]</ref>. The third feature group is sentence type. Sentences in the dataset are identified as complex or compound sentences and simple sentences. In linguistics, complex sentences are sentences with two or more clauses connected with a subordinate conjunction. Simple sentences contain one independent clause <ref type="bibr" target="#b25">[26]</ref>.</p><p>We use four routine and straightforward statistical features, word count, character count, the average word length of a sentence (character count/ word count), and the number of unique words in a sentence.</p><p>K-means is a non-deterministic, iterative, and unsupervised machine learning technique to produce clusters from data <ref type="bibr" target="#b26">[27]</ref>. Unsupervised learning here serves as both a baseline for comparison with supervised learning results, and as a long term goal in itself, independent of any manual annotation. In the simplest test, we converted the input sentences to embeddings and cluster them using K-means clustering. Sentence embeddings were created using two techniques, context-based Embeddings from Language Models (ELMo) and knowledge-graphbased ConceptNet Numberbatch (CNNB). The following two paragraphs give a brief overview of these two embedding techniques.</p><p>ELMo embeddings are deep, contextualized representations of words computed using a twolayer bidirectional language model (biLM), which is pretrained on a large text corpus <ref type="bibr" target="#b27">[28]</ref>. ELMo is robust in creating embeddings for out-of-vocabulary (OOV) words because it incorporates subword and character-level information when creating embeddings. Handling OOV words is particularly important in this work as most of the sentences often contain domain-specific keywords such as "Glucose-6-p", "DHP", "GAP" which can fall into the OOV category. ELMo sentence vectors are 1024-dimensional.</p><p>ConceptNet is a semantic network of knowledge about word meanings <ref type="bibr" target="#b28">[29]</ref>. CNNB embeddings <ref type="bibr" target="#b29">[30]</ref> are semantic word vectors created by encoding the knowledge in ConceptNet <ref type="bibr" target="#b28">[29]</ref>. ConceptNet Numberbatch sentence embeddings are produced by taking the mean of single word embeddings in a sentence. CNNB sentence vectors are 300-dimensional. CNNB embeddings are not as robust as ELMo embeddings when handling OOV words, yet the percentage of OOV words in the current dataset is rather small (6%). Hence we use CNNB as the second embedding technique to create sentence embeddings.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.2.2.">Implementation Details</head><p>We use Pandas DataFrames <ref type="bibr" target="#b30">[31]</ref> to process and manipulate the sentence features. We also used other external libraries used in the extraction of sentence features as follows. To extract semantic features, the sentences are sent through a transition-based AMR parser named CAMR <ref type="bibr" target="#b31">[32]</ref>. Textblob <ref type="bibr" target="#b32">[33]</ref> is used to assess the subjectivity and polarity scores of the sentences. POS tag and NER-related features (in syntactic features category) are extracted using spaCy <ref type="foot" target="#foot_0">1</ref> . Matplotlib<ref type="foot" target="#foot_1">2</ref> and seaborn<ref type="foot" target="#foot_2">3</ref> are used for the visualizations.</p><p>ConceptNet Numberbatch embeddings are static representations for words available publicly <ref type="foot" target="#foot_3">4</ref> . ELMo sentence embeddings were created using the model available at Tensorflow Hub<ref type="foot" target="#foot_4">5</ref> .</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.3.">Analysis</head><p>In the following sections, we look at semantic, sentiment-based, syntactic, and statistical feature distributions for strong and weak analogies. We then compare the performances of an SVM classifier and K-means clustering.</p><p>Figure <ref type="figure" target="#fig_2">3</ref> illustrates the distributions of counts of concepts, relations, attributes, unique concepts, unique relations, and unique attributes of strong and weak analogies. Figure <ref type="figure" target="#fig_3">4</ref> presents the polarity and subjectivity distribution of strong and weak analogies. As shown in the plots, both strong and weak analogies contain sentences with neutral polarity and less subjectivity. Seventeen POS tags are present in the dataset. Distributions of the three most prevalent POS tags in strong and weak analogies are depicted in Figure <ref type="figure" target="#fig_4">5</ref> to utilize space effectively. Nevertheless, we used all 17 POS tags in the SVM classifier as features. Out of the fourteen named entities in the dataset (that are used in the SVM classifier), the distributions of the top three (ORG, CARDINAL, and PERSON) are plotted in Figure <ref type="figure" target="#fig_5">6</ref>. We use the spaCy English pipeline <ref type="foot" target="#foot_5">6</ref> for NER tagging. Analogies written by students contain several references to biochemicals. These are misidentified as organizations (ORG) by spaCy, resulting in the ORG tag being the top named entity in the dataset. Figure <ref type="figure" target="#fig_6">7</ref> shows the distribution of simple and complex/ compound sentences. Weak analogies tend to have a slightly higher number of complex/ compound sentences, and strong analogies have slightly more simple sentences. Figure <ref type="figure" target="#fig_7">8</ref> presents the distributions of word counts, character counts, average word lengths, and unique word counts of strong and weak analogies. Modest discrepancies between distributions suggest the potential for such features to distinguish between strong and weak analogies. Therefore a feature vector combining the abovementioned features (engineered features) was then used in an SVM classifier to classify strong and weak analogies. We use five variants of sentence vectors as inputs to the SVM classifier and K-means clustering. The first variant is the ELMo embeddings vector (ELMo). The second variant is the CNNB embeddings vector (CNNB), and the third variant is the engineered features vector with the vector dimension of 44. The fourth variant is a simple concatenation between ELMo embeddings and the engineered feature vector (ELMo composite). The fifth variant is a simple concatenation between CNNB embeddings and the engineered feature vector (CNNB composite).  We opted to train an SVM classifier with stratified K-fold cross validation due to the limited size of our dataset (less than 1000 data points). Due to the imbalanced nature of the dataset and Thilini Wijesiriwardene et al.   identifying strong and weak analogies were equally important in this initial analysis, we used macro-F1 as the performance metric <ref type="bibr" target="#b33">[34]</ref>. Performance of the SVM classifier with five variants of sentence vectors are listed in table 2.3. We further inspect the contributions of the engineered features from the four feature categories mentioned in section 2.2 when discriminating between strong and weak analogies. We observe (see Figure <ref type="figure" target="#fig_8">9</ref>) that most of the top 15 discriminating features belong to the syntactic feature category, but a semantic feature contributes the most to discriminate between strong and weak analogies. We use K-means to cluster the five variants of sentence vectors mentioned above with cluster centers randomly selected and K set to two. Based on the Rand index 7 , the clusters are not well Thilini Wijesiriwardene et al.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>ICCBR'22 Workshop Proceedings</head><formula xml:id="formula_0">ICCBR'22 Workshop Proceedings (a)<label>(b) (c) (d)</label></formula></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>ICCBR'22 Workshop Proceedings</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Table 2</head><p>Different sentence vector variants and their performance on SVM classifier measured by macro averaged Precision/Recall/F1-score along with their K-means cluster qualities given by Rand Index. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.">Discussion</head><p>This section presents our findings and insights, followed by the limitations and future work.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.1.">Findings and Insights</head><p>We introduce the DAF pipeline to identify discriminative features of strong and weak analogies.</p><p>We show that just a few engineered features does a surprisingly good job as input to SVM. To be sure, the ELMo composite sent through the SVM classifier performs better than the rest of the sentence vector variants. Nevertheless, the ELMo composite score is slightly higher (∼0.03) than the ELMo. This increase highlights that the engineered features encode some aspects of the analogies not well-captured by the ELMo embeddings. Although the SVM's performance with the engineered feature vector is 26% lower than that of the ELMo embedding, its embedding size is ∼23 times smaller than the ELMo. This phenomenon hints that considerable performance gains can be achieved with a much smaller number of better hand-crafted features, and most importantly, the better performance is explainable. We also note that the CNNB composite vector's performance in SVM is slightly poorer than that of the CNNB itself (∼0.02). Although further exploration is required to explain this phenomenon clearly, we suspect this may be the result of feature multicollinearity specific to the manner in which CNNB creates its embeddings, combined with idiosyncracies of the subsets constructed in cross-validation. We show that, among the features passed to the SVM classifier, the most discriminative feature for classification is a semantic feature (unique attribute count) and three out of the four semantic features (unique relations count, unique concepts count, concepts count) fall in the list of top 15 discriminative features. Also, among the top 15 discriminative features, syntactic features have the most representation. Overall, the engineered features are few in number, meaningful, and relatively cheap to calculate. Given the range of content in the data set -anything from marbles to cake-the modest success reported here is impressive. These features will contribute to our future efforts based on more computationally intensive semantic analysis. A successful unsupervised learning method would liberate classifier training from dependence on manual annotation. Unsupervised learning results remain largely unimpressive. Nevertheless, there are some hints of promise. Engineered features improve clustering results relative to embeddings alone or embeddings and engineered features. This reinforces our claim that such features are identifying discriminators that are not captured by embeddings.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.2.">Limitations and Future Work.</head><p>The dataset used in this work is imbalanced, with more strong analogy data points than weak ones. This may cause the "uniform effect" where K-means produces clusters of the same size, even when the "true" cluster sizes of the dataset are varied <ref type="bibr" target="#b34">[35]</ref>. To overcome such issues we plan to improve class imbalance through SMOTE <ref type="bibr" target="#b35">[36]</ref>, GANS <ref type="bibr" target="#b36">[37]</ref>, and the expansion of the manually-annotated corpus.</p><p>The natural language processing techniques employed in this work do not handle the particular nature of the dataset. For example, the spaCy model we use is trained on a generic English corpus <ref type="foot" target="#foot_7">8</ref> . However, we plan to use models/ techniques trained on subject-specific corpora to overcome issues like misidentifying biochemical terms as organizations in NER. Also, students use the term "like" in their target analogies to signify the similarity between their analogy and the source domain (biochemistry) concept. These are wrongly picked up by the sentiment analysis tool when evaluating polarity. Modified corpora will allow us to better manage these issues.</p><p>We classified analogy strength using individual sentences, which is both a benefit and a limitation. As a result, we identified very simple discriminators. However, some sentences in the dataset might not contribute when creating strong/ weak analogies. Constraining analysis to the sentence level requires annotation to eliminate this potential source of noise. However, the long-term goal is to evaluate analogies at the document level, for their epistemic quality. Though still vector based, our ongoing work in this area employs referent knowledge bases for both the target and variable student sources, to guide semantic interpretation.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.">Conclusion</head><p>This work introduces the DAF pipeline to identify discriminative features between strong and weak long-form analogies. We show that an SVM-based supervised-learning approach can successfully discriminate component sentences drawn from strong and weak analogies. Semantic and several syntactic features are the main contributors to discrimination, helping us to realize our goal of efficient evaluation of student generated long-form analogies.</p></div><figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_0"><head>-Figure 1 :</head><label>1</label><figDesc>Figure 1: Illustration of the DAF Pipeline. The analogies are split into sentences to create the analogy dataset. Single sentences are sent through Analogy Processing. Features of the sentences are extracted and used in SVM-based classification and K-means clustering. An analysis is then conducted on SVMbased classification and K-means clustering results.</figDesc></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_1"><head>Figure 2 :</head><label>2</label><figDesc>Figure 2: AMR representation of the sentence "Adding more marbles to the box will not increase the amount of product produced since it relies heavily on red marbles being oriented into the grooves. " from the dataset.</figDesc><graphic coords="5,89.29,106.65,416.71,249.55" type="bitmap" /></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_2"><head>Figure 3 :</head><label>3</label><figDesc>Figure 3: Plots illustrating the distributions of (a) Concepts counts, (b) Relations counts, (c) Attributes counts, (d) Unique concepts counts, (e) Unique relations counts, and (f) Unique attributes counts of sentence analogies.</figDesc></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_3"><head>Figure 4 :</head><label>4</label><figDesc>Figure 4: Plots illustrating the distributions of (a) Polarity, (b) Subjectivity across sentence analogies.</figDesc></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_4"><head>Figure 5 :</head><label>5</label><figDesc>Figure 5: (a) VERB, (b) DETERMINER, (c) NOUN POS tags counts distribution across sentence analogies.</figDesc></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_5"><head>Figure 6 :</head><label>6</label><figDesc>Figure 6: (a) ORG, (b) CARDINAL, and (c) PERSON NER tag counts distribution across sentence analogies.</figDesc></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_6"><head>Figure 7 :</head><label>7</label><figDesc>Figure 7: Sentence type statistics of strong and weak analogies.</figDesc><graphic coords="8,192.22,403.72,208.35,136.49" type="bitmap" /></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_7"><head>Figure 8 :</head><label>8</label><figDesc>Figure 8: Plots illustrating the distributions of (a) Word count, (b) Character count, (c) Avg. word length, and (d) Number of unique words across sentence analogies.</figDesc></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_8"><head>Figure 9 :</head><label>9</label><figDesc>Figure 9: Top 15 discriminative features between strong and weak analogies</figDesc><graphic coords="9,108.88,426.89,375.03,164.99" type="bitmap" /></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" type="table" xml:id="tab_0"><head>Table 1</head><label>1</label><figDesc></figDesc><table><row><cell>Dataset statistics</cell><cell></cell><cell></cell></row><row><cell></cell><cell cols="2">Strong Weak</cell></row><row><cell>Num. of analogies</cell><cell>25</cell><cell>6</cell></row><row><cell cols="2">Num. of analogies (sentences) 586</cell><cell>140</cell></row></table></figure>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="1" xml:id="foot_0">https://spacy.io/</note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="2" xml:id="foot_1">https://matplotlib.org/</note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="3" xml:id="foot_2">https://seaborn.pydata.org/</note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="4" xml:id="foot_3">https://github.com/commonsense/conceptnet-numberbatch</note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="5" xml:id="foot_4">https://tfhub.dev/google/elmo/3</note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="6" xml:id="foot_5">https://spacy.io/models/en#en_core_web_md</note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="7" xml:id="foot_6">https://scikit-learn.org/stable/modules/clustering.html#rand-index</note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="8" xml:id="foot_7">https://spacy.io/models/en#en_core_web_md</note>
		</body>
		<back>

			<div type="acknowledgement">
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Acknowledgments</head><p>We thank Dr. Biplav Srivastava for his valuable feedback and Dr. Nitin Jain for providing the data used in this work. We also thank the reviewers for their constructive comments.</p></div>
			</div>

			<div type="references">

				<listBibl>

<biblStruct xml:id="b0">
	<analytic>
		<title level="a" type="main">Structure-mapping: A theoretical framework for analogy</title>
		<author>
			<persName><forename type="first">D</forename><surname>Gentner</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Cognitive science</title>
		<imprint>
			<biblScope unit="volume">7</biblScope>
			<biblScope unit="page" from="155" to="170" />
			<date type="published" when="1983">1983</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b1">
	<analytic>
		<title level="a" type="main">High-level perception, representation, and analogy: A critique of artificial intelligence methodology</title>
		<author>
			<persName><forename type="first">D</forename><forename type="middle">J</forename><surname>Chalmers</surname></persName>
		</author>
		<author>
			<persName><forename type="first">R</forename><forename type="middle">M</forename><surname>French</surname></persName>
		</author>
		<author>
			<persName><forename type="first">D</forename><forename type="middle">R</forename><surname>Hofstadter</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Journal of Experimental &amp; Theoretical Artificial Intelligence</title>
		<imprint>
			<biblScope unit="volume">4</biblScope>
			<biblScope unit="page" from="185" to="211" />
			<date type="published" when="1992">1992</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b2">
	<monogr>
		<title level="m" type="main">Mental leaps: Analogy in creative thought</title>
		<author>
			<persName><forename type="first">K</forename><forename type="middle">J</forename><surname>Holyoak</surname></persName>
		</author>
		<author>
			<persName><forename type="first">P</forename><surname>Thagard</surname></persName>
		</author>
		<imprint>
			<date type="published" when="1996">1996</date>
			<publisher>MIT press</publisher>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b3">
	<analytic>
		<title level="a" type="main">Distraction during relational reasoning: The role of prefrontal cortex in interference control</title>
		<author>
			<persName><forename type="first">D</forename><forename type="middle">C</forename><surname>Krawczyk</surname></persName>
		</author>
		<author>
			<persName><forename type="first">R</forename><forename type="middle">G</forename><surname>Morrison</surname></persName>
		</author>
		<author>
			<persName><forename type="first">I</forename><surname>Viskontas</surname></persName>
		</author>
		<author>
			<persName><forename type="first">K</forename><forename type="middle">J</forename><surname>Holyoak</surname></persName>
		</author>
		<author>
			<persName><forename type="first">T</forename><forename type="middle">W</forename><surname>Chow</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><forename type="middle">F</forename><surname>Mendez</surname></persName>
		</author>
		<author>
			<persName><forename type="first">B</forename><forename type="middle">L</forename><surname>Miller</surname></persName>
		</author>
		<author>
			<persName><forename type="first">B</forename><forename type="middle">J</forename><surname>Knowlton</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Neuropsychologia</title>
		<imprint>
			<biblScope unit="volume">46</biblScope>
			<biblScope unit="page" from="2020" to="2032" />
			<date type="published" when="2008">2008</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b4">
	<analytic>
		<title level="a" type="main">A neurocomputational model of analogical reasoning and its breakdown in frontotemporal lobar degeneration</title>
		<author>
			<persName><forename type="first">R</forename><forename type="middle">G</forename><surname>Morrison</surname></persName>
		</author>
		<author>
			<persName><forename type="first">D</forename><forename type="middle">C</forename><surname>Krawczyk</surname></persName>
		</author>
		<author>
			<persName><forename type="first">K</forename><forename type="middle">J</forename><surname>Holyoak</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><forename type="middle">E</forename><surname>Hummel</surname></persName>
		</author>
		<author>
			<persName><forename type="first">T</forename><forename type="middle">W</forename><surname>Chow</surname></persName>
		</author>
		<author>
			<persName><forename type="first">B</forename><forename type="middle">L</forename><surname>Miller</surname></persName>
		</author>
		<author>
			<persName><forename type="first">B</forename><forename type="middle">J</forename><surname>Knowlton</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Journal of cognitive neuroscience</title>
		<imprint>
			<biblScope unit="volume">16</biblScope>
			<biblScope unit="page" from="260" to="271" />
			<date type="published" when="2004">2004</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b5">
	<analytic>
		<title level="a" type="main">Verbal analogy problem sets: An inventory of testing materials</title>
		<author>
			<persName><forename type="first">N</forename><surname>Ichien</surname></persName>
		</author>
		<author>
			<persName><forename type="first">H</forename><surname>Lu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">K</forename><forename type="middle">J</forename><surname>Holyoak</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Behavior research methods</title>
		<imprint>
			<biblScope unit="volume">52</biblScope>
			<biblScope unit="page" from="1803" to="1816" />
			<date type="published" when="2020">2020</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b6">
	<analytic>
		<title level="a" type="main">Analogical proportions: Why they are useful in ai</title>
		<author>
			<persName><forename type="first">H</forename><surname>Prade</surname></persName>
		</author>
		<author>
			<persName><forename type="first">G</forename><surname>Richard</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">IJCAI</title>
				<imprint>
			<date type="published" when="2021">2021</date>
			<biblScope unit="page" from="4568" to="4576" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b7">
	<analytic>
		<title level="a" type="main">Classifying and completing word analogies by machine learning</title>
		<author>
			<persName><forename type="first">S</forename><surname>Lim</surname></persName>
		</author>
		<author>
			<persName><forename type="first">H</forename><surname>Prade</surname></persName>
		</author>
		<author>
			<persName><forename type="first">G</forename><surname>Richard</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">International Journal of Approximate Reasoning</title>
		<imprint>
			<biblScope unit="volume">132</biblScope>
			<biblScope unit="page" from="1" to="25" />
			<date type="published" when="2021">2021</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b8">
	<analytic>
		<title level="a" type="main">BERT is to NLP what alexnet is to CV: Can pre-trained language models identify analogies?</title>
		<author>
			<persName><forename type="first">A</forename><surname>Ushio</surname></persName>
		</author>
		<author>
			<persName><forename type="first">L</forename><surname>Espinosa-Anke</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Schockaert</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Camacho-Collados</surname></persName>
		</author>
		<ptr target="https://openreview.net/forum?id=BdWgrMFxdW9" />
	</analytic>
	<monogr>
		<title level="m">ACL 2022 Workshop on Commonsense Representation and Reasoning</title>
				<imprint>
			<date type="published" when="2022">2022</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b9">
	<analytic>
		<title level="a" type="main">Pragmatics in analogical mapping</title>
		<author>
			<persName><forename type="first">B</forename><forename type="middle">A</forename><surname>Spellman</surname></persName>
		</author>
		<author>
			<persName><forename type="first">K</forename><forename type="middle">J</forename><surname>Holyoak</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Cognitive psychology</title>
		<imprint>
			<biblScope unit="volume">31</biblScope>
			<biblScope unit="page" from="307" to="346" />
			<date type="published" when="1996">1996</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b10">
	<analytic>
		<title level="a" type="main">The structure-mapping engine: Algorithm and examples</title>
		<author>
			<persName><forename type="first">B</forename><surname>Falkenhainer</surname></persName>
		</author>
		<author>
			<persName><forename type="first">K</forename><forename type="middle">D</forename><surname>Forbus</surname></persName>
		</author>
		<author>
			<persName><forename type="first">D</forename><surname>Gentner</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Artificial intelligence</title>
		<imprint>
			<biblScope unit="volume">41</biblScope>
			<biblScope unit="page" from="1" to="63" />
			<date type="published" when="1989">1989</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b11">
	<analytic>
		<title level="a" type="main">Extending analogical generalization with near-misses</title>
		<author>
			<persName><forename type="first">M</forename><surname>Mclure</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Friedman</surname></persName>
		</author>
		<author>
			<persName><forename type="first">K</forename><surname>Forbus</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceedings of the AAAI Conference on Artificial Intelligence</title>
				<meeting>the AAAI Conference on Artificial Intelligence</meeting>
		<imprint>
			<date type="published" when="2015">2015</date>
			<biblScope unit="volume">29</biblScope>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b12">
	<monogr>
		<title level="m" type="main">Use of student-generated process analogies to enhance student engagement</title>
		<author>
			<persName><forename type="first">R</forename><surname>Shrem</surname></persName>
		</author>
		<author>
			<persName><forename type="first">T</forename><surname>Vonderhaar</surname></persName>
		</author>
		<author>
			<persName><forename type="first">V</forename><surname>Shalin</surname></persName>
		</author>
		<author>
			<persName><forename type="first">N</forename><surname>Jain</surname></persName>
		</author>
		<imprint>
			<date type="published" when="2022">2022</date>
		</imprint>
	</monogr>
	<note>in preparation</note>
</biblStruct>

<biblStruct xml:id="b13">
	<analytic>
		<title level="a" type="main">Support-vector networks</title>
		<author>
			<persName><forename type="first">C</forename><surname>Cortes</surname></persName>
		</author>
		<author>
			<persName><forename type="first">V</forename><surname>Vapnik</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Machine learning</title>
		<imprint>
			<biblScope unit="volume">20</biblScope>
			<biblScope unit="page" from="273" to="297" />
			<date type="published" when="1995">1995</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b14">
	<analytic>
		<title level="a" type="main">A comprehensive survey on support vector machine classification: Applications, challenges and trends</title>
		<author>
			<persName><forename type="first">J</forename><surname>Cervantes</surname></persName>
		</author>
		<author>
			<persName><forename type="first">F</forename><surname>Garcia-Lamont</surname></persName>
		</author>
		<author>
			<persName><forename type="first">L</forename><surname>Rodríguez-Mazahua</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><surname>Lopez</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Neurocomputing</title>
		<imprint>
			<biblScope unit="volume">408</biblScope>
			<biblScope unit="page" from="189" to="215" />
			<date type="published" when="2020">2020</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b15">
	<analytic>
		<title level="a" type="main">Abstract Meaning Representation for sembanking</title>
		<author>
			<persName><forename type="first">L</forename><surname>Banarescu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">C</forename><surname>Bonial</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Cai</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Georgescu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">K</forename><surname>Griffitt</surname></persName>
		</author>
		<author>
			<persName><forename type="first">U</forename><surname>Hermjakob</surname></persName>
		</author>
		<author>
			<persName><forename type="first">K</forename><surname>Knight</surname></persName>
		</author>
		<author>
			<persName><forename type="first">P</forename><surname>Koehn</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Palmer</surname></persName>
		</author>
		<author>
			<persName><forename type="first">N</forename><surname>Schneider</surname></persName>
		</author>
		<ptr target="https://aclanthology.org/W13-2322" />
	</analytic>
	<monogr>
		<title level="m">Proceedings of the 7th Linguistic Annotation Workshop and Interoperability with Discourse, Association for Computational Linguistics</title>
				<meeting>the 7th Linguistic Annotation Workshop and Interoperability with Discourse, Association for Computational Linguistics<address><addrLine>Sofia, Bulgaria</addrLine></address></meeting>
		<imprint>
			<date type="published" when="2013">2013</date>
			<biblScope unit="page" from="178" to="186" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b16">
	<analytic>
		<title level="a" type="main">From TreeBank to PropBank</title>
		<author>
			<persName><forename type="first">P</forename><surname>Kingsbury</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Palmer</surname></persName>
		</author>
		<ptr target="http://www.lrec-conf.org/proceedings/lrec2002/pdf/283.pdf" />
	</analytic>
	<monogr>
		<title level="m">Proceedings of the Third International Conference on Language Resources and Evaluation (LREC&apos;02), European Language Resources Association (ELRA)</title>
				<meeting>the Third International Conference on Language Resources and Evaluation (LREC&apos;02), European Language Resources Association (ELRA)<address><addrLine>Las Palmas, Canary Islands -Spain</addrLine></address></meeting>
		<imprint>
			<date type="published" when="2002">2002</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b17">
	<analytic>
		<title level="a" type="main">Toward abstractive summarization using semantic representations</title>
		<author>
			<persName><forename type="first">F</forename><surname>Liu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Flanigan</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Thomson</surname></persName>
		</author>
		<author>
			<persName><forename type="first">N</forename><surname>Sadeh</surname></persName>
		</author>
		<author>
			<persName><forename type="first">N</forename><forename type="middle">A</forename><surname>Smith</surname></persName>
		</author>
		<idno type="DOI">10.3115/v1/N15-1114</idno>
		<ptr target="https://aclanthology.org/N15-1114.doi:10.3115/v1/N15-1114" />
	</analytic>
	<monogr>
		<title level="m">Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Association for Computational Linguistics</title>
				<meeting>the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Association for Computational Linguistics<address><addrLine>Denver, Colorado</addrLine></address></meeting>
		<imprint>
			<date type="published" when="2015">2015</date>
			<biblScope unit="page" from="1077" to="1086" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b18">
	<analytic>
		<title level="a" type="main">Machine comprehension using rich semantic representations</title>
		<author>
			<persName><forename type="first">M</forename><surname>Sachan</surname></persName>
		</author>
		<author>
			<persName><forename type="first">E</forename><surname>Xing</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics</title>
		<title level="s">Short Papers</title>
		<meeting>the 54th Annual Meeting of the Association for Computational Linguistics</meeting>
		<imprint>
			<date type="published" when="2016">2016</date>
			<biblScope unit="volume">2</biblScope>
			<biblScope unit="page" from="486" to="492" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b19">
	<analytic>
		<title level="a" type="main">Biomedical event extraction using Abstract Meaning Representation</title>
		<author>
			<persName><forename type="first">S</forename><surname>Rao</surname></persName>
		</author>
		<author>
			<persName><forename type="first">D</forename><surname>Marcu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">K</forename><surname>Knight</surname></persName>
		</author>
		<author>
			<persName><forename type="first">H</forename><surname>Daumé</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Iii</forename></persName>
		</author>
		<idno type="DOI">10.18653/v1/W17-2315</idno>
		<ptr target="https://aclanthology.org/W17-2315.doi:10.18653/v1/W17-2315" />
	</analytic>
	<monogr>
		<title level="m">BioNLP 2017, Association for Computational Linguistics</title>
				<meeting><address><addrLine>Vancouver, Canada</addrLine></address></meeting>
		<imprint>
			<date type="published" when="2017">2017</date>
			<biblScope unit="page" from="126" to="135" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b20">
	<analytic>
		<title level="a" type="main">Liberal event extraction and event schema induction</title>
		<author>
			<persName><forename type="first">L</forename><surname>Huang</surname></persName>
		</author>
		<author>
			<persName><forename type="first">T</forename><surname>Cassidy</surname></persName>
		</author>
		<author>
			<persName><forename type="first">X</forename><surname>Feng</surname></persName>
		</author>
		<author>
			<persName><forename type="first">H</forename><surname>Ji</surname></persName>
		</author>
		<author>
			<persName><forename type="first">C</forename><surname>Voss</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Han</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><surname>Sil</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics</title>
		<title level="s">Long Papers</title>
		<meeting>the 54th Annual Meeting of the Association for Computational Linguistics</meeting>
		<imprint>
			<date type="published" when="2016">2016</date>
			<biblScope unit="volume">1</biblScope>
			<biblScope unit="page" from="258" to="268" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b21">
	<analytic>
		<title level="a" type="main">Turning words into consumer preferences: How sentiment analysis is framed in research and the news media</title>
		<author>
			<persName><forename type="first">C</forename><surname>Puschmann</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><surname>Powell</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Social Media+ Society</title>
		<imprint>
			<biblScope unit="volume">4</biblScope>
			<biblScope unit="page">2056305118797724</biblScope>
			<date type="published" when="2018">2018</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b22">
	<analytic>
		<title level="a" type="main">Subjectivity analysis in opinion mining-a systematic literature review</title>
		<author>
			<persName><forename type="first">E</forename><surname>Kasmuri</surname></persName>
		</author>
		<author>
			<persName><forename type="first">H</forename><surname>Basiron</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Int J Adv Soft Comput Appl</title>
		<imprint>
			<biblScope unit="volume">9</biblScope>
			<biblScope unit="page" from="132" to="159" />
			<date type="published" when="2017">2017</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b23">
	<analytic>
		<title level="a" type="main">Part of speech tagging: a systematic review of deep learning and machine learning approaches</title>
		<author>
			<persName><forename type="first">A</forename><surname>Chiche</surname></persName>
		</author>
		<author>
			<persName><forename type="first">B</forename><surname>Yitagesu</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Journal of Big Data</title>
		<imprint>
			<biblScope unit="volume">9</biblScope>
			<biblScope unit="page" from="1" to="25" />
			<date type="published" when="2022">2022</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b24">
	<analytic>
		<title level="a" type="main">Named entity recognition without gazetteers</title>
		<author>
			<persName><forename type="first">A</forename><surname>Mikheev</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Moens</surname></persName>
		</author>
		<author>
			<persName><forename type="first">C</forename><surname>Grover</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Ninth Conference of the European Chapter of the Association for Computational Linguistics</title>
				<imprint>
			<date type="published" when="1999">1999</date>
			<biblScope unit="page" from="1" to="8" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b25">
	<analytic>
		<title level="a" type="main">A novel system for generating simple sentences from complex and compound sentences</title>
		<author>
			<persName><forename type="first">B</forename><surname>Das</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Majumder</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Phadikar</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">International Journal of Modern Education and Computer Science</title>
		<imprint>
			<biblScope unit="volume">11</biblScope>
			<biblScope unit="page">57</biblScope>
			<date type="published" when="2018">2018</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b26">
	<analytic>
		<title level="a" type="main">Classification and analysis of multivariate observations</title>
		<author>
			<persName><forename type="first">J</forename><surname>Macqueen</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">5th Berkeley Symp. Math. Statist. Probability</title>
				<imprint>
			<date type="published" when="1967">1967</date>
			<biblScope unit="page" from="281" to="297" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b27">
	<analytic>
		<title level="a" type="main">Deep contextualized word representations</title>
		<author>
			<persName><forename type="first">M</forename><forename type="middle">E</forename><surname>Peters</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Neumann</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Iyyer</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Gardner</surname></persName>
		</author>
		<author>
			<persName><forename type="first">C</forename><surname>Clark</surname></persName>
		</author>
		<author>
			<persName><forename type="first">K</forename><surname>Lee</surname></persName>
		</author>
		<author>
			<persName><forename type="first">L</forename><surname>Zettlemoyer</surname></persName>
		</author>
		<idno type="DOI">10.18653/v1/N18-1202</idno>
		<ptr target="https://aclanthology.org/N18-1202.doi:10.18653/v1/N18-1202" />
	</analytic>
	<monogr>
		<title level="m">Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</title>
		<title level="s">Long Papers</title>
		<meeting>the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies<address><addrLine>New Orleans, Louisiana</addrLine></address></meeting>
		<imprint>
			<date type="published" when="2018">2018</date>
			<biblScope unit="volume">1</biblScope>
			<biblScope unit="page" from="2227" to="2237" />
		</imprint>
	</monogr>
	<note>Association for Computational Linguistics</note>
</biblStruct>

<biblStruct xml:id="b28">
	<analytic>
		<title level="a" type="main">Conceptnet 5.5: An open multilingual graph of general knowledge</title>
		<author>
			<persName><forename type="first">R</forename><surname>Speer</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Chin</surname></persName>
		</author>
		<author>
			<persName><forename type="first">C</forename><surname>Havasi</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Thirty-first AAAI conference on artificial intelligence</title>
				<imprint>
			<date type="published" when="2017">2017</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b29">
	<monogr>
		<author>
			<persName><forename type="first">R</forename><surname>Speer</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Chin</surname></persName>
		</author>
		<idno type="arXiv">arXiv:1604.01692</idno>
		<title level="m">An ensemble method to produce high-quality word embeddings</title>
				<imprint>
			<date type="published" when="2016">2016</date>
		</imprint>
	</monogr>
	<note type="report_type">arXiv preprint</note>
</biblStruct>

<biblStruct xml:id="b30">
	<monogr>
		<title level="m" type="main">pandas: a foundational python library for data analysis and statistics, Python for high performance and scientific computing</title>
		<author>
			<persName><forename type="first">W</forename><surname>Mckinney</surname></persName>
		</author>
		<imprint>
			<date type="published" when="2011">2011</date>
			<biblScope unit="volume">14</biblScope>
			<biblScope unit="page" from="1" to="9" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b31">
	<analytic>
		<title level="a" type="main">Camr at semeval-2016 task 8: An extended transition-based amr parser</title>
		<author>
			<persName><forename type="first">C</forename><surname>Wang</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Pradhan</surname></persName>
		</author>
		<author>
			<persName><forename type="first">X</forename><surname>Pan</surname></persName>
		</author>
		<author>
			<persName><forename type="first">H</forename><surname>Ji</surname></persName>
		</author>
		<author>
			<persName><forename type="first">N</forename><surname>Xue</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceedings of the 10th international workshop on semantic evaluation (semeval-2016)</title>
				<meeting>the 10th international workshop on semantic evaluation (semeval-2016)</meeting>
		<imprint>
			<date type="published" when="2016">2016</date>
			<biblScope unit="page" from="1173" to="1178" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b32">
	<analytic>
		<title level="a" type="main">textblob documentation</title>
		<author>
			<persName><forename type="first">S</forename><surname>Loria</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Release</title>
		<imprint>
			<biblScope unit="volume">0</biblScope>
			<biblScope unit="issue">15</biblScope>
			<date type="published" when="2018">2018</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b33">
	<analytic>
		<title level="a" type="main">A systematic analysis of performance measures for classification tasks</title>
		<author>
			<persName><forename type="first">M</forename><surname>Sokolova</surname></persName>
		</author>
		<author>
			<persName><forename type="first">G</forename><surname>Lapalme</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Information processing &amp; management</title>
		<imprint>
			<biblScope unit="volume">45</biblScope>
			<biblScope unit="page" from="427" to="437" />
			<date type="published" when="2009">2009</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b34">
	<analytic>
		<title level="a" type="main">K-means clustering versus validation measures: A datadistribution perspective</title>
		<author>
			<persName><forename type="first">H</forename><surname>Xiong</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Wu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Chen</surname></persName>
		</author>
		<idno type="DOI">10.1109/TSMCB.2008.2004559</idno>
	</analytic>
	<monogr>
		<title level="j">IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics)</title>
		<imprint>
			<biblScope unit="volume">39</biblScope>
			<biblScope unit="page" from="318" to="331" />
			<date type="published" when="2009">2009</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b35">
	<analytic>
		<title level="a" type="main">Smote: synthetic minority over-sampling technique</title>
		<author>
			<persName><forename type="first">N</forename><forename type="middle">V</forename><surname>Chawla</surname></persName>
		</author>
		<author>
			<persName><forename type="first">K</forename><forename type="middle">W</forename><surname>Bowyer</surname></persName>
		</author>
		<author>
			<persName><forename type="first">L</forename><forename type="middle">O</forename><surname>Hall</surname></persName>
		</author>
		<author>
			<persName><forename type="first">W</forename><forename type="middle">P</forename><surname>Kegelmeyer</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Journal of artificial intelligence research</title>
		<imprint>
			<biblScope unit="volume">16</biblScope>
			<biblScope unit="page" from="321" to="357" />
			<date type="published" when="2002">2002</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b36">
	<analytic>
		<title level="a" type="main">Generative adversarial nets</title>
		<author>
			<persName><forename type="first">I</forename><surname>Goodfellow</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Pouget-Abadie</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Mirza</surname></persName>
		</author>
		<author>
			<persName><forename type="first">B</forename><surname>Xu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">D</forename><surname>Warde-Farley</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Ozair</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><surname>Courville</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Y</forename><surname>Bengio</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Advances in neural information processing systems</title>
		<imprint>
			<biblScope unit="volume">27</biblScope>
			<date type="published" when="2014">2014</date>
		</imprint>
	</monogr>
</biblStruct>

				</listBibl>
			</div>
		</back>
	</text>
</TEI>
