<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Uncovering Business Relationships: Context-sensitive Relationship Extraction for Difficult Relationship Types</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Zhe Zuo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michael Loster</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ralf Krestel</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Felix Naumann</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Hasso Plattner Institute Prof.</institution>
          <addr-line>-Dr.-Helmert-Straße 2-3, 14482 Potsdam</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <abstract>
        <p>This paper establishes a semi-supervised strategy for extracting various types of complex business relationships from textual data by using only a few manually provided company seed pairs that exemplify the target relationship. Additionally, we offer a solution for determining the direction of asymmetric relationships, such as “ownership of”. We improve the reliability of the extraction process by using a holistic pattern identification method that classifies the generated extraction patterns. Our experiments show that we can accurately and reliably extract new entity pairs occurring in the target relationship by using as few as five labeled seed pairs.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Extracting structured data from text, and thus harnessing the valuable information on
the web and hidden in the vast amounts of other textual data, is a well-known and
wellstudied research area. As the text corpora and the kind of information to be extracted
from them can vary greatly, many research works have focused on specific types of
information, on specific corpora, on specific application domains, on specific languages,
or any combination of the above. In this paper, we regard the problem of extracting
relationships of several specific types among companies from news articles.</p>
      <p>Many tasks, such as building business networks, predicting risks, or valuating
companies, can significantly benefit from accurately extracting relationships between
companies. Imagine a scenario in which Dell wants to acquire EMC. Dell plans to finance
the deal by taking out a loan. The chosen bank has to decide whether to award the loan
based on the careful assessment of the risk associated with this transaction. With the
explosive growth of the textual data on the web, it becomes possible to discover not
only the information of Dell and EMC but also the dependencies by extracting business
relationships and building up a company network. In the same example, by analyzing
the network structure of both companies, the bank might reach the conclusion that the
risk of granting a loan is too high, because many of EMC’s subsidiaries, as given by
the relationship network, are struggling. With this knowledge the bank might award a
smaller or no loan at all or propose a higher interest rate.</p>
      <p>
        To build up a business network between companies, it is critical to reliably extract
business relationships. Companies often connect to each other via the activities in which
they participate. Business relationships represent a subset of these activities; examples
include ownership of, partnership with, supplier of, and so on. Only very few of them
can be found in structured knowledge base like Freebase [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] or semi-structured data like
Wikipedia infoboxes – a substantial amount of relationships is hidden in unstructured
data sources. Aggravating this situation, both Freebase and infoboxes contain only the
major subsidiaries of some companies (i.e., ownership of relationship). Other
relationships, such as partnership with or supplier of, are not covered.
      </p>
      <p>Given a corpus of unstructured textual data, we aim to (1) discover whether two
co-occurring companies participate in a business relationship, (2) identify the type of
the relationship, and (3) in the case of an asymmetric business relationship, determine
its direction.</p>
      <p>The task of business relationship extraction is challenging due to the complex
nature of the relationships between companies. First, multiple types of relationships can
exist between two companies. Samsung as one of the biggest competitors of Apple is
also the supplier of displays for Apple’s products. Moreover, as an example of
resolving the direction of asymmetric relationships, such as the ownership of relationship,
consider that Walt Disney owns ABC Studios but not the other way around. Being able
to successfully derive the direction of the relationships is of vital importance for many
subsequent tasks.</p>
      <p>
        The Snowball system addresses the general problem of relationship extraction [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ],
and our work is based in parts on its general idea. It takes a small set of entity pairs as
a seed set and generates candidate patterns that are based on the context of these pairs.
Subsequently, the most prominent patterns are selected according to a scoring function
and used to extract new entity pairs that participate in the target relationship. In the end,
the newly selected pairs are added to the seed set and the process repeats to generate
more patterns. However, Snowball functions only correctly if there is a one-to-many
relationship between the participating entities, e.g., in the headquarter of(Microsoft,
Redmond) relationship, Microsoft has exactly one headquarter. Business relationships do
not adhere to this characteristic, which is the reason Snowball is unable to solve the
problem at hand.
      </p>
      <p>We extend the Snowball idea by introducing a key-phrase extraction strategy, which
allows us to remove irrelevant parts of the context surrounding the company pairs. To
determine the direction of asymmetric relationships, we propose a process that
leverages information contained in the seed set. Since Snowball cannot deal with
many-tomany business relationships, we propose a generalization of their tuple- and
patternevaluation strategy by specifying a new selection method to select patterns and new
seeds. We further define a holistic pattern identification strategy, which enables us to
extract multiple relationship types simultaneously.</p>
      <p>In summary, we propose a system to perform (directed) relationship extraction
(RE) between companies from textual data. Addressing this problem, we present a
novel, semi-supervised relationship extraction method, which requires only a minimum
amount of manually specified company pairs to efficiently extract new ones that belong
to the same target relationship. Additionally, we provide a straightforward solution to
reliably identify the direction of asymmetric relationships. We show that our approach
is superior to more advanced distant learning approaches for the particularly difficult
case of many-to-many relationships.</p>
    </sec>
    <sec id="sec-2">
      <title>Background and Related Work</title>
      <p>
        The most related work is the Snowball system [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], which we have already introduced
in Section 1. Despite the fact that there is a large body of work that focuses on the
topic of relationship extraction, the subject of extracting business relationships between
companies from unstructured data has not been sufficiently addressed by research.
      </p>
      <p>
        One way to approach the general relationship extraction problem is to use
supervised learning techniques by classifying whether two entities participate in a specific
relationship. Kambhatla [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] employed Maximum Entropy models to solve the
relationship extraction task. Zhou et al. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] also applied a feature-based relationship extraction
strategy that uses Support Vector Machines (SVM) [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Further, kernel methods with
string-kernels have successfully been applied to deal with the RE problem [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. The
major drawback of these techniques is that a large amount of labeled data is required
for training. As a representative example, Kambhatla [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] uses a training set that contains
around 9,752 instances of relationships to generate their results. Moreover, relabeling
and retraining of the model becomes necessary, as soon as either the underlying
characteristics of the data sources or the target relationship change substantially.
      </p>
      <p>
        Mintz et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] introduced a distant supervision approach, which avoids the
expensive labeling process. The idea is to automatically label the training data according to
the relationships included in knowledge bases, i.e., Freebase [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. One of the limitations
is that it is highly rely on the given knowledge base, only the types of relationships that
are included can be extracted, while most of the business relationships are not covered
at all, such as partnership with, competitor of and supplier of .
      </p>
      <p>
        Another way to address the problem was presented by Banko et al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. They
introduced an unsupervised approach called TextRunner to extract all possible relationships
in a given corpus without requiring any labeled data. This task is known as the open
information extraction (Open IE) task. Wu and Weld [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] proposed the WOE system,
which enhances TextRunner by including additional information from Wikipedia
infoboxes to construct a training dataset. Although the Open IE approaches can
automatically extract all possible relationships from a given corpus, their results cannot directly
be used in further applications. They can neither disambiguate mentions nor provide
semantic information about the extracted relationships automatically.
      </p>
      <p>
        We avoid labeling large amounts of training data and predefining a specific type
of relationship by using a few examples of a target relationship for bootstrapping.
This idea was first introduced by Brin [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] in the context of the DIPRE system, which
focused on extracting relationships between authors and their corresponding book
titles. Some other approaches were developed based on this bootstrapping strategy, e.g.,
Snowball [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and StatSnowball [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
      </p>
      <p>We focus on reliably extracting business relationships between companies. By
applying the semi-supervised algorithm, we can extract more complicated many-to-many
relationships from large amounts of unlabeled data without requiring the expensive
initial labeled data. A user only has to supply a very small number (3–5) of seeds to
achieve good results, which makes our approach flexible to be applied to variant target
relationships or data sources by simply provide another small seed set. Furthermore, our
approach is able to determine the direction of asymmetric relationships. This enables
us to directly use the generated results in subsequent applications.</p>
    </sec>
    <sec id="sec-3">
      <title>Overview of our Approach</title>
      <p>
        Relationship Extraction
Figure 1 gives a high-level overview of our relationship extraction approach: Given
some textual data and a seed set of multiple company pairs that occur as members of a
particular relationship, our system outputs new company pairs participating in the same
relationship type. As a preprocessing step we simplified the algorithm introduced by
Zuo et al. [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] to recognize and link the mentions of companies to their corresponding
Wikipedia pages.
      </p>
      <p>Given these disambiguated
company mentions we generate patterns
from their contexts. To this end, we
follow the intuition that if a company Textual Named Entity
pair from the seed set co-occurs in Data Recognition
the same sentence, it is likely that
the context characterizes the
relationship specified by the seed (see
Section 4.1). Therefore, the sentences that SSeeetd GePnaettreartnion
contain two or more distinct compa- Selected Seeds
nies are selected as the input for the re- Candidate Patterns
lationship extraction phase. An exam- New Seed Pattern
ple tagged sentence is “. . . [[Verizon Selection Selection
CommunicationsjVerizon]]’s acquisi- Selected Patterns
ttihoen oroifgin[a[MlmCeIntIinocn.sjM“CVIe]r]i”z,onw”haenrde CoPmapirasny New Pairs ENxetrwacPtaioirn
“MCI” are separately linked to
Verizon Communications and MCI Inc.</p>
      <p>From the contexts surrounding
company pairs, we generate possible can- Fig. 1. Processing pipeline of our approach
didate extraction patterns that are
likely to represent the target relationship (see Section 4.1). Suppose we are interested in
the ownership of relationship and the company pair (Verizon Communication, MCI Inc.)
is contained in the seed set, then a candidate pattern pattern = hCOMP1, COMP2,
acquisition of, !i can be generated. The last element in pattern describes the direction
of the ownership of relationship (see Section 4.4). After generating a list of candidate
patterns, we select the most promising ones according to the measurements to be
introduced in Section 4.2.</p>
      <p>We then use the selected patterns to discover new company pairs from the input. If
the previous pattern pattern is selected, we can extract a new company pair (The Walt
Disney Company, Pixar) from a sentence like “. . . after Disney’s acquisition of Pixar
Animation Studios”. Afterwards, we select the most prominent newly extracted pairs
to extend the seed set (see Section 4.3). We then iterate the procedure to extract new
patterns using the extended seed set until no more new company pairs can be selected
as seeds or the iteration number reaches a predefined limit. The company pairs that are
extracted based on the current set of patterns are considered to participate in the same
type of relationship as the target one. Our evaluation shows that this is indeed almost
always the case regardless of the initial choice of seed pairs.</p>
      <p>Company
Disambiguation</p>
      <p>Preprocessing
Tagged Sentences</p>
    </sec>
    <sec id="sec-4">
      <title>Extraction of Business Relationships</title>
      <p>This section introduces our semi-supervised relationship extraction strategy, which
iteratively extracts new company pairs that participate in a given target relationship.
4.1</p>
      <sec id="sec-4-1">
        <title>Pattern generation</title>
        <p>Generating the extraction patterns represents a crucial step in our approach. The
context surrounding a company pair represents the main source to identify relationships
occurring in textual data. To capture the key information that represents the relationship
between two companies we extract the most determining phrases from the context as a
key-phrase. This key-phrase is then used to generate a pattern.</p>
        <p>Candidate pattern An extracted pattern includes two company variables COMP1 and
COMP2, the key-phrase extracted from the context in between those companies, and a
direction. We explain each of these parts in the following. From an example sentence
“. . . YouTube, the video-sharing Web site owned by Google . . . ” we can generate the
pattern hCOMP1, COMP2, owned by, i. By applying this pattern to this sentence we
obtain the following instantiation of the pattern hYouTube, Google, owned by, i,
indicating that Google owns YouTube.</p>
        <p>
          Key-phrase extraction The quality of a pattern depends on the key-phrase it contains.
A good pattern should satisfy two criteria: characterize a single type of relationship
(which in turn improves the precision of the extraction result) and be as general as
possible (to extract many new company pairs). For this reason, it is beneficial to
generalize the context and don’t keep idiosyncratic key-phrases. The key-phrase should be
as compact as possible, while maintaining the information in the context. Extracting
patterns for business relationships in the news is particularly challenging since
journalists are used to introduce the same type of business relationships using different
writing styles spanning a relatively large context. This can be shown using the excerpt
“. . . News Corporation, which owns a minority interest in DirecTV”. In this sentence,
we can easily figure out that News Corporation is one of the owners of DirecTV by
finding the verb “owns” in the intermediary context. If we now use the entire context
(i.e., “, which owns a minority interest in”) between the two companies as a pattern
to extract additional company pairs, we would find only very few since the pattern is
not general enough. The problem can be solved by extracting the key-phrase “owns”
that defines the ownership relationship. Thus we can conceptually simplify the original
sentence to “New Corporation owns DirecTV”. To this end, we developed a key-phrase
extraction strategy to automatically extract the most determining phrases from the
intermediary context. Intuitively, relationships in sentences are often conveyed by verbs
or nouns. In [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] most of the binary relationships are indicated by four types of phrases,
which cover over 86% of the cases. These key-phrase types are “Verb”, “Noun+Prep.”,
“Verb+Prep.”, and “Infinitive” located in between two entities in English text. To
extract key-phrases from contexts, we apply the Stanford Part-Of-Speech(POS) Tagger 1.
1 http://nlp.stanford.edu/software/
Based on the POS tags, we keep the phrases that match any of the four types above.
We abandon the context if the containing verb is “to be”, because it usually does not
indicate any business relationship, or if the context contains multiple key-phrases.
4.2
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>Pattern selection</title>
        <p>In each iteration, we generate candidate patterns based on the (extended) seed set.
However, patterns that do not represent the target relationship might also end up in the
candidate list. Therefore it is important to keep only the most representative patterns while
filtering out unfavorable patterns. In the following, we introduce two strategies to select
the best patterns.</p>
        <p>Hit score Building on the intuition that patterns that frequently match company pairs
in the seed set are likely to be representative ones, we introduce a Hit score for each
pattern as follows,</p>
        <p>Hit(patternjPairseed; S) =</p>
        <p>X</p>
        <p>X [match(pairi; p; sj )]
(1)
pairi2Pairseed sj2S</p>
        <p>Thus Hit is defined as the summation of how frequently a pattern matches a
company pairi 2 Pairseed in the set of input sentences S. A pattern with a high Hit score
denotes that the corresponding key-phrase is more likely to represent the target
relationship. Given a list of candidate patterns that are sorted in descending order by their
respective Hit score, we select the top-k ranked patterns to extend the set of the current
extraction patterns.</p>
        <p>Coverage score A good pattern should frequently be used in the context between
different company pairs to describe a particular relationship. If the pattern can be extracted
by using only one of the seed pairs, it is either too specific or it describes some other
type of relationship between the corresponding companies. We introduce a Coverage
(Cov) score, which represents the percentage of company pairs from the seed set that
are able to generate this pattern.</p>
        <p>Cov(patternjPairseed; S) =</p>
        <p>P
pairi2Pairseed
[Psj2S[match(pairi; p; sj )] &gt; 0]
jPairseedj
(2)</p>
        <p>The Cov score of a pattern equals 1:0 when all seed pairs match the pattern at least
once. All patterns that have a Cov score greater than a threshold are selected.
4.3</p>
      </sec>
      <sec id="sec-4-3">
        <title>New seeds selection</title>
        <p>We introduce a similar strategy for selecting newly extracted company pairs to extend
the seed set. Using the selected patterns, we compute the Hit score for each of the
extracted company pairs. We select the top-k pairs by their Hit scores. We can also
compute the Cov score of an extracted company pair, which is the percentage of selected
patterns that match the company pair in the text. In a similar fashion to the pattern
selection, we extend the seed set by selecting the company pairs that have a greater Cov
score than the same given threshold for pattern selection.
4.4</p>
      </sec>
      <sec id="sec-4-4">
        <title>Direction of relationship</title>
        <p>In Section 1 we introduced the challenge of determining the direction of asymmetric
business relationships. Compared to the extraction of symmetric relationships,
extracting asymmetric ones, such as supplier of, ownership of, and sued by require not only
the extraction of a new company pair occurring in the target relationship, but also the
detection of its correct semantic direction.</p>
        <p>
          Previous work, such as Snowball [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], naturally avoids this direction problem, since
they focus on relationships that relate two objects of different entity types (i.e.,
organization, location). However, in our case, the entities are of the same type (i.e., company).
Zhu et al. [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] present a similar challenge, e.g. an entity of type person e1 is the husband
of e2. They solve this problem by manually adding new rules, such as IsHusband(e1; e2)
) IsWife(e2; e1), during their iterations.
        </p>
        <p>We introduce an elegant strategy to automatically classify the direction of newly
extracted relationships. The idea is to include the direction information already in the
seed set. When the target relationship is asymmetric, the company pairs in the initial
seed set must be specified by also providing the direction of the relationship. E.g., in
the case of the ownership of relationship, we specify a forward direction, denoting that
the first company is the owner of the second.</p>
        <p>Given this directed seed set, we can identify the direction of the generated patterns
as follows: When two companies are mentioned in the same order as in the seed pair,
the pattern is annotated with the same direction as the seed pair. Finally, the direction of
a pattern is derived by assigning the direction that is more frequently marked. Table 1
in the evaluation section shows some examples of determined directions of patterns.
4.5</p>
      </sec>
      <sec id="sec-4-5">
        <title>Multiple types of relationships</title>
        <p>With our semi-supervised business relationship extraction approach, we can
independently extract different relationship types by providing multiple initial seed sets each
characterizing one type of relationship.</p>
        <p>As mentioned in Section 1, different types of business relationships can exist
between two companies at the same time. Therefore, the patterns generated from the seed
set do not always represent the desired relationship type. Even worse, once a pattern
that represents an undesired relationship type is selected, the following iterations can be
negatively influenced in a way that they yield more and more irrelevant patterns, which
leads to incorrect extraction results comparable to a topic drift in pseudo-relevance
feedback methods. We can avoid this problem by assigning each pattern that is generated
for multiple relationship types exclusively to one single type.</p>
        <p>We followed the intuition that each pattern characterizes one kind of relationship
and implemented a holistic pattern identification strategy by using the Cov score. In
case the same pattern is generated for multiple relationship types, we exclusively assign
the pattern to the type that yields the highest Cov score.</p>
        <p>As a preliminary experiment to show the effect of this holistic strategy, we
applied our approach to extract the ownership of and partnership with relationships at the
same time. The selection of patterns and new seeds was made using the Hit score.
By applying the holistic pattern identification strategy, most of the patterns, especially
the top-ranked ones, characterize the partnership with relationship. Without our
holistic strategy, the top-ranked patterns (i.e., “stake in”, “deal with”, and “buy”) mainly
represent the ownership of relationship. This problem was caused by a falsely selected
pattern (i.e., “owned by”), which led to more and more patterns that characterize the
ownership of relationship.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Experiments</title>
      <p>In our evaluation we focus on the extraction of an asymmetric relationships (i.e.,
ownership of) from articles of the New York Times corpus.
5.1</p>
      <sec id="sec-5-1">
        <title>NYTimes corpus and seeds</title>
        <p>The full New York Times corpus contains 1,855,658 news articles, spanning a period
of 20 years from Jan. 1987 to Jun. 2007. We observed that about 74% of all company
pairs within a sentence occurred in the “Technology” and “Business” categories. Thus,
we reduced our corpus to articles with at least one of those two labels. Our final corpus
(called NYTimes from now on) consists of 359,459 articles.</p>
        <p>An initial seed set serves as the input for our approach and predefines the
relationship type we would like to extract. We investigated two different seed sets to evaluate
their influence on the results. To this end, we generated a list of distinct company pairs
that co-occur in the NYTimes corpus and sorted it in descending order by co-occurrence
frequency. We manually labeled the relationship type for the first 100 pairs and then
randomly selected five company pairs (FreqSeed) that share the ownership of relationship
from the top-100 list entries. Following this random selection strategy, we also
generated a seed set called InfreqSeed from the top-1000 company pairs. Keep in mind that
seed selection and our evaluation is based on a corpus dating from 1987 to 2007,
resulting in relationships that might not hold today. FreqSeed contains company pairs, such
as (AOL, Netscape), (Viacom, Viacom Media Networks), (Ford, Jaguar), (Time Warner,
TBS), and (GE, NBC Sports), while InfreqSeed contains less frequently mentioned pairs,
such as (Disney, ESPN), (IPC, Campbell Mithun), (GM, Saturn), (Chrysler, American
Motors), and (Investcorp, Saks Fifth Avenue).
5.2</p>
      </sec>
      <sec id="sec-5-2">
        <title>Experimental results</title>
        <p>We first show which patterns were generated and then evaluate the quality of the actual
business relationships we extracted.</p>
        <p>Results of pattern generation Based on the two randomly generated seed sets we
applied our approach to extract new company pairs that are also members of the
ownership of relationship. Table 1 shows the key-phrases of the selected patterns that are
automatically generated by using FreqSeed and InfreqSeed. In this experiment, we
applied the Hit score in each iteration for selecting the top-10 candidate patterns and the
respective company pairs. The first column shows the key-phrases of the selected
patterns. By using either FreqSeed or InfreqSeed, the extraction process terminates after
Extracted Patterns FRraenqkS(eIetderation)
(key-phrase) 1 2 3 I1nfre2qSeed3 Direction
unit of
parent of
owned by
part of
division of
owns
company of
acquisition of
subsidiary of
owner of
including
include
bought
acquired
buy
bought by
three iterations resulting in 14 selected patterns. These patterns are sorted in descending
order by their Hit score. We also include the ranks of the patterns per iteration to show
the changes that occur from iteration to iteration.</p>
        <p>Further, Table 1 shows that most of the automatically generated key-phrases are
typical phrases frequently used to describe an ownership of relationship. Already in the
first iteration, our approach can generate representative patterns. Differences between
the two sets of generated patterns can be observed mainly in the tail. Thus, our approach
is not particularly sensitive towards the chosen seed set (we made similar observations
for various other seed sets, both in terms of size and co-occurrence frequency).</p>
        <p>The last column in Table 1 contains the extracted direction of the patterns as
determined by the strategy introduced in Section 4.4. Only 2 out of the 16 directions are
incorrect Although the direction of these two patterns is classified incorrectly, most
directions of the newly extracted company pairs, are identified correctly as the statistics
in Section 5.2 show. This is because the direction of newly extracted company pairs is
determined by multiple patterns.</p>
        <p>
          Quality of extraction results We applied our approach using different settings for both
pattern and seed selection to verify the extraction result. We conducted experiments
with the Hit and Cov scores strategies introduced in Section 4. To show the effect of
our key-phrase extraction strategy, we also executed our algorithm without using this
strategy. In other words, we employed the original context to generated patterns, which
is similar to previous work, e.g., [
          <xref ref-type="bibr" rid="ref1 ref5">1, 5</xref>
          ]. As a baseline, we select the most frequently
cooccurred company pairs to check how many of them are in an ownership of relationship.
        </p>
        <p>We had to manually check relationships between company pairs, because no gold
standard with known business relationships is available. The design of our approach is
30.0% 36.0% 30.5%
mainly concerned with achieving a high precision value because we aim to use it in
the context of risk-analysis, which has only a small tolerance for incorrectly extracted
information. Therefore, we mainly focus on evaluating the precision performance of
our approach. We manually examined the top-200 most frequently extracted company
pairs from each result set produced by our algorithm with different parameterizations.</p>
        <p>Table 2 presents the evaluation results using the FreqSeed seed set to extract the
ownership of relationships. As this table shows, by applying Cov ( = 0:7) score, 90%
of the top-200 extracted company pairs indeed participate in the ownership of
relationship. The performance of our approach, excluding the key-phrase extraction strategy,
also shows the significant effect of including it. In comparison to the baseline, our
approach can produce much better results.</p>
        <p>Apart from the precision measure, we also present a detailed error analysis based
on the top-200 extracted company pairs: The first error type is that company pairs that
do not participate in an ownership of relationship are extracted (Rel.). Another error
case is that our approach extracted the correct company pair, but failed to identify the
correct direction (Dir.). A third error case is caused by recognition or disambiguation
errors made by the preprocessing steps (Pre.). An incorrect result can also be due to
misinterpretation of the semantics (Sem.). E.g., one company finally canceled the plan
of acquiring another one, such as the abandoned merger between EMI and Time Warner.
Such events are covered by a series of New York Times articles, but our approach was
unable to successfully capture the final cancellation of the deal. As the result shows,
only around half of the incorrectly extracted relationships (i.e., Rel. and Dir.) are caused
by our RE strategy.</p>
        <p>Furthermore, according to the mechanism of our approach, when a relationship is
mentioned in the given corpus more frequently, the probability that our approach can
extract that relationship is higher. Thus, by including more documents the recall of our
approach increases. We iteratively applied our approach (with the setting Cov ( =
0:7)) to an NYTimes corpus of increasing size, starting from 10 years of data up to 21
years. In Figure 2, the red line denotes the total number of tagged sentences after our
preprocessing step. The blue bars show the accumulated count of extracted company
pairs. As the figure shows, by enlarging the size of the dataset more unique aimed
relationships can be extracted.</p>
        <p>
          We compared our approach
with a state-of-the-art distant
learning approach developed by Zeng
et al. [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. They applied the
piecewise convolutional neural
networks with multi-instance learning
for relationship extraction, which
we refer as PCNNs2. In their
experiments, the dataset3 which
contains the New York Times
articles labeled with Freebase
relationships was used. The sentences Fig. 2. The accumulated count of extracted company
from 2005 2006 were used for pairs from subsets of NYTimes corpus
training, while the ones from 2007
were used for testing. To compare the performance between PCNNs and our approach,
we apply our approach on this dataset. As we have introduced in Section 1, Freebase
only contains the major acquisitions of companies, which can be considered as the
ownership of relationship. However, all of the instances of the ownership of relationship
were mislabeled to be negative in the original dataset. Therefore, to compare PCNNs
with our approach for extracting the ownership of relationship, we had to relabel the
training set according to the corresponding Freebase relationships (99 pairs are matched
in the training set). Since only 14 Freebase relationships can be matched in the test set,
we randomly picked and manually validated 100 company pairs (including 50 positives
and 50 negatives) from the articles in 2007. For this specific type of relationship,
PCNNs labels 5 pairs as positive, which are all correct. Our approach extracts 19 pairs,
where 18 of them are correct. In this experiment, our approach outperforms PCNNs in
both recall and F-measure.
        </p>
        <p>
          Regarding efficiency, our approach can extract business relationships at a rate of
about 650 documents per minute on a standard consumer PC, with most of the time
spent on preprocessing. The efficiency can be further improved by implementing a
distributed system to apply our approach as the strategy introduced in [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ].
        </p>
        <p>More detailed statistics as well as the annotated data are available online4.
6</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusion and Future Work</title>
      <p>The focus of this work was to efficiently extract complex business relationships from
news articles. We are the first to focus on the class of many-to-many relationships. To
this end, we proposed a relationship extraction approach that not only extracts new
relationships from text but also indicates their direction in case of non-symmetric ones, such
2 The original code is available online: http://www.nlpr.ia.ac.cn/cip/
˜liukang/publications.html
3 http://iesl.cs.umass.edu/riedel/ecml/
4 https://hpi.de/naumann/projects/knowledge-discovery-and-mining/
business-relationship-extraction.html
as the ownership of relationship. Another contribution is the holistic pattern
identification strategy, which is used to avoid the semantic drift of generated extraction patterns
while dealing with multiple business relationships simultaneously.</p>
      <p>Further, we would like to include the duration and domain information of
relationships. Moreover, the performance of our approach can be further improved by
understanding the semantics of the underlying sentences to avoid incorrect extractions caused
by misinterpretations.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Agichtein</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gravano</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Snowball: Extracting relations from large plain-text collections</article-title>
          .
          <source>In: Proceedings of the International Conference on Digital Libraries (DL)</source>
          . pp.
          <fpage>85</fpage>
          -
          <lpage>94</lpage>
          (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Banko</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cafarella</surname>
            ,
            <given-names>M.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soderland</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Broadhead</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Etzioni</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Open information extraction for the web</article-title>
          .
          <source>In: Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI)</source>
          . pp.
          <fpage>2670</fpage>
          -
          <lpage>2676</lpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Banko</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Etzioni</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>The tradeoffs between open and traditional relation extraction</article-title>
          .
          <source>In: Proceedings of the Meeting of the Association for Computational Linguistics (ACL)</source>
          . pp.
          <fpage>28</fpage>
          -
          <lpage>36</lpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Bollacker</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Evans</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paritosh</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sturge</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Taylor</surname>
          </string-name>
          , J.:
          <article-title>Freebase: a collaboratively created graph database for structuring human knowledge</article-title>
          .
          <source>In: Proceedings of the International Conference on Management of Data (SIGMOD)</source>
          . pp.
          <fpage>1247</fpage>
          -
          <lpage>1250</lpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Brin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Extracting patterns and relations from the world wide web</article-title>
          .
          <source>In: The World Wide Web and Databases</source>
          , pp.
          <fpage>172</fpage>
          -
          <lpage>183</lpage>
          . Springer (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Kambhatla</surname>
          </string-name>
          , N.:
          <article-title>Combining lexical, syntactic, and semantic features with maximum entropy models for extracting relations</article-title>
          .
          <source>In: Proceedings of the ACL 2004 on Interactive poster and demonstration sessions</source>
          . p.
          <volume>22</volume>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Mintz</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bills</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Snow</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jurafsky</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Distant supervision for relation extraction without labeled data</article-title>
          .
          <source>In: Proceedings of the Joint Conference of the Meeting of the Association for Computational Linguistics (ACL) and the International Joint Conference on Natural Language Processing of the AFNLP</source>
          . pp.
          <fpage>1003</fpage>
          -
          <lpage>1011</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Vapnik</surname>
            ,
            <given-names>V.N.</given-names>
          </string-name>
          :
          <article-title>Statistical learning theory</article-title>
          , vol.
          <volume>1</volume>
          . Wiley, New York (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Min</surname>
          </string-name>
          , H.:
          <article-title>Entity relation mining in large-scale data</article-title>
          .
          <source>In: Database Systems for Advanced Applications:</source>
          DASFAA 2015 International Workshops, SeCoP, BDMS, and Posters. p.
          <volume>109</volume>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weld</surname>
            ,
            <given-names>D.S.</given-names>
          </string-name>
          :
          <article-title>Open information extraction using Wikipedia. In: Proceedings of the Meeting of the Association for Computational Linguistics (ACL)</article-title>
          . pp.
          <fpage>118</fpage>
          -
          <lpage>127</lpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Zelenko</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aone</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Richardella</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Kernel methods for relation extraction</article-title>
          .
          <source>The Journal of Machine Learning Research</source>
          <volume>3</volume>
          ,
          <fpage>1083</fpage>
          -
          <lpage>1106</lpage>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Zeng</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Distant supervision for relation extraction via piecewise convolutional neural networks</article-title>
          .
          <source>In: Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP)</source>
          . pp.
          <fpage>1753</fpage>
          -
          <lpage>1762</lpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Su</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , Zhang,
          <string-name>
            <surname>J.</surname>
          </string-name>
          , Zhang, M.:
          <article-title>Exploring various knowledge in relation extraction</article-title>
          .
          <source>In: Proceedings of the Meeting of the Association for Computational Linguistics (ACL)</source>
          . pp.
          <fpage>427</fpage>
          -
          <lpage>434</lpage>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nie</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wen</surname>
            ,
            <given-names>J.R.:</given-names>
          </string-name>
          <article-title>StatSnowball: a statistical approach to extracting entity relationships</article-title>
          .
          <source>In: Proceedings of the International World Wide Web Conference (WWW)</source>
          . pp.
          <fpage>101</fpage>
          -
          <lpage>110</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Zuo</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kasneci</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gruetze</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Naumann</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>BEL: Bagging for entity linking</article-title>
          .
          <source>In: Proceedings of the International Conference on Computational Linguistics (COLING)</source>
          . pp.
          <fpage>2075</fpage>
          -
          <lpage>2086</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>