Propensity Score Matching for Causal Inference with Relational Data David Arbour Katerina Marazopoulou Dan Garant David Jensen School of Computer Science University of Massachusetts Amherst Amherst, MA 01003 {darbour, kmarazo, dgarant, jensen}@cs.umass.edu Abstract Hometown Propensity score matching (PSM) is a widely From used method for performing causal inference Checks-In Has with observational data. PSM requires fully Places User Friends specifying the set of confounding variables of treatment and outcome. In the case of Figure 1: Example of relational data: users are friends relational data, this set may include non- with other users, each user comes from a hometown, intuitive relational variables, i.e., variables and users check-in at places. derived from the relational structure of the data. In this work, we provide an automated method to derive these relational variables To illustrate this, consider the example domain shown based on the relational structure and a set of in Figure 1, depicting a plausible relational domain. naive confounders. This automatic construc- Foursquare is an example of a real system that could tion includes two unusual classes of variables: produce this sort of data. Suppose a researcher is inter- relational degree and entity identifiers. We ested in using data from this domain to assess whether provide experimental evidence that demon- smoking causes a user to gain weight. One approach strates the utility of these variables in ac- would be to construct a propensity score model with counting for certain latent confounders. Fi- user attributes that the researcher believes could be nally, through a set of synthetic experiments, causes of whether a user smokes and the user’s weight, we show that our method improves the per- such as alcohol consumption and ethnicity: formance of PSM for causal inference with [User ].Smokes ∼ [User ].Drinks + [User ].Ethnicity. relational data. While this accounts for attributes associated with the user, it fails to account for possible confounders de- 1 INTRODUCTION rived from relational variables. For example, it is plau- sible that the alcohol consumption of a user’s friends is Propensity score matching (PSM) [Rosenbaum and a common cause of [User ].Weight and [User ].Smokes. Rubin, 1983] is a widely used tool for determining To account for these effects, the corresponding rela- causal effects from observational data. Propensity tional variables should be included in the propensity scores summarize the effects of a potentially large num- score model. ber of confounding variables by creating a predictive It is not difficult to envision more complicated rela- model of treatment. The computation of a propensity tional variables having an effect. In fact, as previous score requires specifying a set of potentially confound- work has shown [Maier et al., 2013b], the number of ing variables. This task is relatively straightforward relational variables can be arbitrarily large depending for propositional (i.i.d.) data. However, many causal on how many entity and relationship types exist in the analyses consider data in which treatment, outcome, network, the size of the network, and the length of the and potential confounders can arise from the inter- longest path (the largest degree of separation) in the actions among multiple types of interrelated entities. network where direct dependence exists. Propensity score matching becomes substantially more challenging in such relational data. An additional level of complexity introduced by rela- tional data is that relational structures may result in can be used to pair similar samples from the treated multiple instances of a given variable. For example, a and the control groups. Matching can be generally de- user with multiple friends could be influenced by the fined as a method that aims to approximate random drinking behaviour of each of those friends. Typically, assignment by equating the distribution of covariates an aggregation function, such as mean, is used to com- in the treated and control group [Stuart, 2010]. bine this set of values into a single value. Properly Matching requires a measure quantifying how similar conditioning on a relational variable entails choosing two individuals are. This is achieved by (1) selecting the correct set of aggregation functions to represent a set of features to be used in the computation of sim- the distribution of values contained in the set. For ex- ilarity, and (2) choosing a similarity function to apply ample, in order to condition on a relational variable, it on those features (for example Mahalanobis distance, may be necessary to condition on multiple aspects of propensity score, etc.). Once a similarity measure has the distribution of those values, such as the mean and been chosen, individuals are matched based on this the standard deviation (stdev). measure. There are multiple methods for performing To address these issues, we introduce relational matching (see Stuart [2010] and Ho et al. [2007] for a propensity score matching (RPSM), a method that survey of matching methods). In this paper, we em- applies propensity score matching to relational do- ploy full matching [Hansen and Klopfer, 2006], which mains. RPSM leverages the framework of relational creates a collection of matched sets (the size of the models [Getoor and Taskar, 2007, Maier et al., 2013b] collection is chosen automatically). Each matched set to automatically construct the set of possible relational contains at least one treated and one control unit. Full confounders given a simpler specification of the as- matching has been shown to be optimal with respect sumed dependency structure. RPSM also identifies op- to similarity within matched sets [Rosenbaum, 1991]. portunities to use relational degree variables and entity Matching methods make the assumption of ignorable identifiers, which, as we show empirically, can reduce treatment assignment, i.e., treatment assignment is in- the bias arising from latent relational confounders. We dependent of the outcome given the observed covari- evaluate RPSM via a set of synthetic experiments us- ates. This assumption guides the selection of appro- ing the relational structure of a real-world relational priate covariates for the computation of similarity. domain, Foursquare. 2.2 PROPENSITY SCORE 2 BACKGROUND The propensity score [Rosenbaum and Rubin, 1983] is In this section we provide a brief overview of matching the probability of receiving treatment, given the ob- methods and propensity scores. We then introduce the served covariates Xi relational concepts necessary to formalize RPSM. ei (Xi ) = P (Ti = 1|Xi ). 2.1 MATCHING Propensity scores are a form of dimensionality re- duction that projects the original covariates down to In the framework of potential outcomes [Rubin, 1974], a single value which preserves distance with respect estimating the causal effect of treatment T on variable to the likelihood of treatment. Matching can then Y is formalized as a comparison of potential outcomes. be performed on the propensity score, as opposed to More formally, let Ti be a binary treatment variable the covariates directly. The prevailing explanation for unit i and let Yi be the outcome variable for unit i, for why propensity scores are appropriate for match- where i ∈ {1, . . . , n}. Yi (Ti = 0) denotes the value of ing is that they are balancing scores (given the value Yi that would be observed if no treatment was applied of the propensity score, the treatment and control to unit i. Similarly, Yi (Ti = 1) is the value of Yi that groups have the same distribution of covariates), and would be observed if unit i had received treatment. they preserve ignorability of treatment assignment (if The causal effect of T on Y is estimated by comparing treatment assignment is ignorable given the covariates, the difference Yi (Ti = 1)−Yi (Ti = 0) across all units i. then treatment assignment is also ignorable given the propensity score) [Stuart, 2010]. In practice, a specific unit either receives treatment or not. Therefore, for a given value of i we never know Any method that models the conditional probability both Yi (T = 1) and Yi (T = 0). Experimental studies of a binary variable given a set of predictors can be often randomly assign units to treatment and control used to estimate a propensity score. In this work, we groups, so that the expected distribution of the co- employ logistic regression, a widely used method for variates in these groups is identical. In observational obtaining a propensity score. However, other mod- studies, where randomization is not possible, matching els (such as boosted trees, support vector machines, FRIENDS Smokes Weight CHECKS-IN Price Range OM Drinks Hillside Diner FR Smokes Weight Alice C H EC S FROM Price Range School Quality ND KS FRIE School Quality CHECKS-IN -IN FR Drinks PLACE OM HOMETOWN Springfield USER Smokes Weight CHECKS-IN Price Range [User , Friends, User ].Drinks ! [User ].Weight Drinks Corner Cafe Bob Figure 2: Relational model for the Foursquare do- main. The underlying relational schema (ER diagram) Figure 3: Example relational skeleton for the is shown in black. The attributes on the entities are Foursquare domain. This could be a small fragment fictional. The relational dependency is shown in gray. of a (potentially) larger skeleton. and neural networks) have been explored in the liter- domain. It includes a cardinality function that im- ature [Westreich et al., 2010, McCaffrey et al., 2004, poses constraints on the number of times an entity Lee et al., 2010]. instance can participate in a relationship. A rela- A key advantage of propensity scores is their robust- tional schema can be graphically represented with an ness to model misspecification [Drake, 1993], i.e., in- Entity-Relationship (ER) diagram. Figure 2 shows cluding irrelevant variables1 in the calculation of the the ER diagram for the Foursquare domain. In this propensity score. Because the propensity score model example, there are three entity classes (User , Place, is built upon a predictive rather than causal model of Hometown), and three relationship classes, (Friends, treatment, many of the issues that arise with tradi- ChecksIn, From). The entity class User has three at- tional regression modeling, such as multicollinearity, tributes: Smokes, Weight, and Drinks. The cardinal- are no longer a threat to validity. Further, in con- ity constraints are depicted using crow’s feet notation. trast to matching directly on the covariates, propen- For example, the cardinality of the From relationship sity scores can down-weight or disregard variables that is one-to-many, indicating that one user has one home- are not associated with treatment and have been er- town, but many users can be from the same hometown. roneously included in the propensity model. How- A relational skeleton is a partial instantiation of a re- ever, as Pearl [2009] has observed, common effects of lational schema that specifies the set of entity and re- the treatment and outcome must not be included in lationship instances that exist in the domain. Fig- the propensity score model. In general, the set of ure 3 depicts an example relational skeleton for the d-connecting paths between treatment and outcome Foursquare domain. The network consists of two User needs to be considered. The propensity score model instances, Alice and Bob, who are friends with each must include a (not necessarily minimal) separating other and come from the same hometown. There are set of treatment and outcome. One approach to elim- two Place instances, Hillside Diner and Corner Cafe. inating variables that are potential common effects of treatment and outcome is the injunction of Rosenbaum Given a relational schema, one can specify rela- and Rubin [1983] to restrict the set of covariates to pre- tional paths, which intuitively correspond to pos- treatment variables (variables whose values are mea- sible ways of traversing the schema (see Maier sured prior to treatment). et al. [2013b] for a formal definition). For the schema shown in Figure 2, possible paths in- clude [User , Friends, User ] (a person’s friends), and 2.3 RELATIONAL CONCEPTS [User , Friends, User , From, Hometown] (the home- Propositional representations, such as Bayesian net- towns of a person’s friends). Relational variables con- works, describe domains with a single entity type. sist of a relational path and an attribute that can be However, many real-world systems involve multiple reached through that path. For example, the relational types of entities that interact with each other. Data variable [User , Friends, User ].Drinks corresponds to produced by such systems are called relational or net- the alcohol consumption of a person’s friends. Prob- work data. In this section, we introduce the basic rela- abilistic dependencies can be defined between rela- tional concepts, following the notation and terminol- tional variables. In this work, we consider dependen- ogy of Maier et al. [2013b]. cies where the path of the outcome relational variable is a single item. In this case, the path of the treat- A relational schema S = (E, R, A, card ) specifies the ment relational variable describes how dependence is set of entity, relationship, and attribute classes of a induced. For example, the relational dependency 1 Variables that are marginally independent of treat- ment or outcome. [User , Friends, User ].Drinks ! [User ].Weight states that the alcohol consumption of a user’s friends affects that user’s weight. FROM IN PERSON TOWN STATE A relational model M = (S, D, Θ) is a collection of relational dependencies D defined over a relational Figure 4: Relational schema that depicts a hierarchy. schema along with their parameterizations Θ (a condi- A state has many towns, but each town is in one state, tional probability distribution for each attribute given and many people are from the same town, but each its parents). The structure of a relational model can person is from one town. be depicted by superimposing the dependencies on the ER diagram of the relational schema, as shown in Figure 2, and labeling each arrow with the depen- summarize the values of related instances into a single dency it corresponds to. If labels are omitted, the value, representative of the distribution. Common ag- resulting graphical representation is known as a class- gregation functions include mean, stdev, mode, count, dependency graph. sum, min, max, and median. Researchers have also de- fined more complex aggregation methods [Perlich and Recent work by Maier et al. [2013b] provides a frame- Provost, 2006]. work that enables reasoning about d-separation in re- lational models. Toward that end, they introduce ab- stract ground graphs (AGGs), a graphical structure 2.4.2 Degree Variables that captures relational dependencies and can be used Other work has pointed out that variation in the size of to answer relational d-separation queries. Abstract the set of values for a relational variable can strongly ground graphs are defined from a given perspective, affect the distribution of the observed values of many the base item of the analysis, and include nodes that aggregation functions [Jensen et al., 2003]. Jensen et correspond to relational variables. For practical appli- al. call the size of this set the “degree” and it is equiv- cations, the size of abstract ground graphs is limited alent, in the terminology of Maier et al. [2013b], to by a (domain dependent) hop-threshold, which con- the size of the terminal set of a relational path. To strains the length of relational paths that will be con- account for the effects of degree on aggregated values, sidered. Intuitively, the hop-threshold corresponds to RPSM includes degree variables in the calculation of the relational “distance” of a cause from its effect. propensity scores. 2.4 NEW TYPES OF VARIABLES 2.4.3 Entity Identifiers In this section we present the new types of variables Blocking designs are widely used in experimental stud- that are enabled by relational domains: (1) Relational ies to account for latent confounders [Fisher, 1935]. variables (a way of defining a larger number of poten- Rattigan et al. [2011] formalized relational blocking as tial confounders) and aggregation; (2) Degree variables an operator that can be used to infer causal depen- (a type of confounder not available without relational dence in observational data expressed in a relational data); (3) Entity identifiers (which enable blocking, a representation. By blocking on the identifier of an en- way to account for latent confounders only available tity, relational blocking accounts for the effect of la- within relational data). Those types of variables are tent variables associated with that entity. Blocking used in the calculation of relational propensity scores is uniquely available for relational data. Moreover, and are referred to as relational covariates. since blocking on an entity appears to avoid inducing dependence due to colliders on that entity, blocking 2.4.1 Aggregation Functions may partially alleviate a key threat to validity noted by Pearl [2009]. A fundamental characteristic of relational data is the heterogeneity of the underlying relational structure. In this work, we incorporate relational blocking with For example, a person can have many friends, differ- propensity scores by including entity identifiers as co- ent people have different sets of friends, and those sets variates in the calculation of propensity scores. We can overlap to varying degrees. This implies that when restrict the use of blocking to hierarchies, i.e., parts constructing relational variables for a specific individ- of the relational schema that are connected through a ual, the construction process will often return a set of series of many-to-one relationships. An example hier- values rather than a single value. For instance, the re- archy is shown in Figure 4. In this case, blocking on lational variable “friends’ age” for a person consists of the identifier for towns (i.e., grouping users based on a set of values containing the age of each one of that their hometown) accounts for the effect of latent vari- person’s friends. In the field of statistical relational ables associated with Hometown, and for the effect of learning, aggregation functions are commonly used to latent variables associated with the State within which each town is located. More generally, blocking on the Algorithm 1: RPSM (X , schema, treatment, outcome, identifier of an entity in a hierarchy accounts for the confoundingAttrs, aggrFunctions, h) effect of latent confounders that reside in that entity 1 perspective ← item class of treatment, outcome and in entities that appear higher up in the hierarchy. 2 covariates ← GetRelationalCovariates (schema, perspective, confoundingAttrs, aggrFunctions, h) 3 propensityScore ← Calculate propensity score for 3 RELATIONAL PROPENSITY treatment ∼ covariates using X SCORE MATCHING 4 matches ← Match (propensityScore, treatment, X ) 5 return matches We consider the following problem: given an entity E and two attributes on that entity, treatment [E].T and Algorithm 2: GetRelationalCovariates (schema, outcome [E].O, we seek to decide between [E].T ! perspective, confoundingAttrs, aggrFunctions, h) [E].O and [E].T 6! [E].O. For notational convenience, 1 relationalConfounders ← relational variables with we restrict our attention to cases where the treatment attributes in confoundingAttrs from perspective and outcome are on the same entity. In practice, perspective up to hop-threshold h RPSM can be applied to any treatment and outcome 2 relCovariates ← ∅ lying on entities that are connected through one-to-one 3 for P.X in relationalConfounders do relationships. We assume that the relational skeleton 4 if P == [perspective] then 5 relCovariates ← relCovariates ∪ P .X has been given a priori, i.e., all entity and relationship 6 else instances have been fully and correctly specified. Ad- 7 for agg in aggrFunctions do ditionally, we assume that the effects of all latent vari- 8 relCovariates ← relCovariates∪agg(P .X ) ables can be accounted for by using relational blocking 9 relCovariates ← relCovariates ∪ degree(P ) (in other words, latent variables exists only on paths 10 for P.X in relationalConfounders do that can be blocked on). 11 if P is valid blocking choice for perspective then 12 controlled ← relational variables that P Relational propensity score matching (RPSM) pro- controls for vides an automatic method for constructing the set 13 relCovariates ← relCovariates \ controlled 14 relCovariates ← relCovariates ∪ P .id of aggregated relational variables, degree variables 15 return relCovariates and entity identifiers (i.e., the relational covariates) to perform propensity score matching on relational data. The procedure for RPSM is described in Al- gorithm 1. RPSM takes as input a data-set X , a rela- are removed from the list of covariates (line 13). tional schema, the treatment and outcome attributes, Example 3.1. Consider our earlier scenario of assess- a set of possible confounding attributes, a set of ag- ing the effect of smoking on a user’s weight. The treat- gregation functions, and a hop-threshold h. The algo- ment is User .Smokes and the outcome is User .Weight rithm constructs the set of relational covariates based (the perspective of the analysis is the User entity on the confounding attributes, the aggregation func- class). If Drinks is given as a possible confounding tions, and hop-threshold (line 2, discussed below in attribute and the hop-threshold is 4, the algorithm detail). The propensity score of the treatment given will add the following relational variables to the set of the covariates is then computed (line 3) and matching relational confounders: is performed based on the propensity score (line 4). [User ].Drinks The construction of relational covariates is presented in Algorithm 2. The algorithm first constructs all [User , Friends, User ].Drinks potential relational variables for the confounding at- [User , Friends, User , Friends, User ].Drinks tributes from the given perspective, up to the speci- [User , ChecksIn, Place, ChecksIn, User ].Drinks fied hop-threshold (line 1).2 This is the set of relational [User , From, Hometown, From, User ].Drinks confounders. Then, for each relational confounder, it creates the appropriate relational covariates by apply- ing the given aggregation functions (lines 7-8). A de- The next step is to create relational covariates based gree variable is then added for the paths of the re- on the above relational variables. First, relational vari- lational confounders (line 9). Finally, the algorithm ables that only involve the User entity, in this case identifies parts of the schema that form a hierarchy and [User ].Drinks, are added to the set of relational co- adds identifier variables for the schema item lowest in 2 The algorithm can be trivially extended to exclude cer- the hierarchy to perform blocking (lines 10-14). Rela- tain relational paths. For example, if the user has domain tional covariates that were constructed from relational knowledge that would exclude specific relational paths or variables that are now determined by the blocking path relational variables from the list of potential confounders. variates. Because these covariates are propositional, Table 1: Descriptive statistics for the Foursquare rela- aggregation functions are not applied. tional skeleton used in the synthetic experiments. The aggregation functions are then applied to rela- Aggregate Friends Check-Ins tional variables that cross the boundaries of the User mean 9.45 120.09 entity. If the set of aggregation functions is {mean}, median 5 73 the algorithm will add the following to the set of rela- min 1 1 tional covariates: max 3674 2477  mean [User , Friends, User ].Drinks ,  mean [User , Friends, User , Friends, User ].Drinks ,  mean [User , ChecksIn, Place, ChecksIn, User ].Drinks et al., 2012], augmented with synthetic attributes on  , the entities. This allows for controlling the dependen- mean [User , From, Hometown, From, User ].Drinks cies between attributes as well as the marginal and The set of relational covariates is augmented by in- conditional distributions, while leveraging relation- cluding the degree of the relational paths that involve ships from a real-network. The relational schema for more than one entity classes: the Foursquare network is shown in Figure 2. The re-  degree [User , Friends, User ] , lational skeleton consists of 9,599 users, 47,164 friend-  ships, 182,968 locations where users “checked-in” via degree [User , Friends, User , Friends, User ] ,  the mobile application, 1,360,123 check-ins, and the degree [User , ChecksIn, Place, ChecksIn, User ] , users’ hometowns. Aggregate statistics for the net- degree [User , From, Hometown, From, User ] work are shown in Table 1. Finally, id variables are added to the relational paths. For our experiments we generated data from multi- In this case, there exists a hierarchy expressed by the ple models to test each hypothesis individually. In all relational path [User , From, Hometown]. Therefore, experiments, the treatment is [User ].Smokes and the the algorithm adds the following relational covariate: outcome [User ].Weight. Each model was parameter- [User , From, Hometown].id ized as follows: The value of the treatment was drawn from a logistic model parametrized using coefficients In practice, the hop-threshold should be chosen on drawn from U (−2, 2) and interaction terms increasing a case by case basis, using expert knowledge of the in degree from 1 (no interaction) to 10 (up to 10 inter- application domain. The choice of aggregation func- acting covariates, not necessarily distinct, per term). tions can be guided by an analysis of each variable’s We refer to this varying degree as “covariate complex- marginal distribution from the perspective of the treat- ity”. The value of outcome was drawn from a linear ment and outcome. model with coefficients drawn from U (−2, 2) and an error distribution drawn from N (0, 1). Marginal dis- tributions for each variable were drawn from N (µ, σ), 4 SYNTHETIC EXPERIMENTS with µ and σ sampled for each variable individually from U (0, 5) and U (1, 3), respectively. To evaluate the performance of RPSM we examine the following hypotheses: We used logistic regression to calculate the propen- sity score and then performed full matching using the 1. Propensity score matching models that are lim- optmatch package [Hansen and Klopfer, 2006]. A lin- ited to simplistic relational attributes (h = 2) fail ear model was applied using treatment and matching to fully account for confounding network effects assignment as covariates and outcome as the response (h = 4) (Section 4.1). variable to assess statistical significance, with an α 2. Traditional aggregates for relational data, such as value of 0.01 for determining dependence. In this set- mean, when used in isolation do not sufficiently ting, we would expect a low error rate for linear log- condition on the distribution of confounding rela- odds functions (covariate complexity is 1), given the tional variables (Section 4.2). perfect correspondence between the generating mod- 3. The inclusion of identifiers for entities that lie els and the estimation methods when the set of co- along valid blocking paths accounts for latent con- variates is correctly specified (no interaction terms). founders on those entities as well on entities con- Adding interaction terms renders the models progres- nected to them. That is, including entity identi- sively less appropriate. We report Type I and Type fiers in the propensity model performs an implicit II errors. Type I error corresponds to cases where a causal blocking design (Section 4.3). valid causal dependence exists between treatment and For all experiments we used the structure derived from outcome and RPSM incorrectly concludes that there a sample of a real-world network, Foursquare [Gao exists no such dependence. Type II error corresponds FRIENDS Model Used in RPSM RPSM2: all aggregations, h=2 RPSM4: all aggregations, h=4 Smokes Weight 1.00 School Quality FROM CHECKS-IN Price Range Drinks PLACE Error Type HOMETOWN USER 0.75 Type I Type II World 2 Figure 5: Class-dependency graph for the models used 0.50 to evaluate the effect of using relational variables with 0.25 longer hop-thresholds as covariates. True Model Error Percentage 0.00 1.00 to cases where RPSM incorrectly concludes that there exists a dependency between treatment and outcome. 0.75 World 4 0.50 4.1 SIMPLE NETWORK DEPENDENCIES 0.25 We examine the first hypothesis, whether propensity score models limited to simplistic relational attributes 0.00 fail to fully account for confounding network effects, by 1 2 3 4 5 6 7 8 9 10 1 2 3 4 5 6 7 8 9 10 generating data from two models. Both models have Covariate Complexity the same class-dependency graph, shown in Figure 5, but differ in the length of the longest true dependency. Figure 6: Percentage of Type I and II error when RPSM2 and RPSM4 are applied to data generated by In the first model (World2), the true relational con- World2 and World4 models with increasing covariate founders are at most two hops away from the treat- complexity, averaged over 100 trials. ment and outcome entity. This corresponds to depen- dencies that can be read directly from the class depen- dency graph, e.g., the places a user checks in to. The 2. RPSM using mean, stdev, max, min as aggregation set of true relational confounders for the model is: functions and h = 4 without blocking or degree [User ].Drinks variables (RPSM4) [User , From, Hometown].SchoolQuality The results are shown in Figure 6. Along the diagonal [User , ChecksIn, Place].PriceRange the RPSM model is consistent with the world config- uration. When models are over-specified, for instance In the second model (World4), the set of true con- RPSM4 in the World2 configuration, RPSM4 achieves founders is extended to include relational variables up comparable performance to RPSM2. However, when to four hops away, e.g., other users that check in to the models are underspecified, for instance RPSM2 in the same places as a user. The set of confounders includes World 4 configuration, a spurious effect is inferred be- all of the confounders of the first model as well as: tween treatment and outcome in the conditionally in- dependent case. These results also demonstrate a case [User , Friends, User ].Drinks in which RPSM can successfully tolerate large num- [User , ChecksIn, Place, ChecksIn, User ].Drinks bers of irrelevant covariates. [User , From, Hometown, From, User ].Drinks [User , Friends, User , Friends, User ].Drinks 4.2 COMPLEX NETWORK [User , Friends, User , ChecksIn, Place].PriceRange DEPENDENCIES [User , Friends, User , From, Hometown].SchoolQuality In this section, we examine the second hypothesis re- garding the effect of using complex aggregation func- Using the above procedure we ran 100 trials. For each tion in the construction of relational covariates. We trial we considered two cases, one in which treatment generated data from models with the same class- and outcome are conditionally independent and one dependency graph as in Section 4.1. We used World2 in which there is a direct effect between them. We and World4, as before, and two simplified models then compared two methods for creating the relational which consider only mean as an aggregate, with hop- covariates for propensity score matching: thresholds of 2 (World2-) and 4 (World4-). We then 1. RPSM using mean, stdev, max, min as aggregation used the RPSM2 and RPSM4 methods for construct- functions and h = 2 without blocking or degree ing relational covariates and two simpler propensity variables (RPSM2) score models that only include mean as an aggregate Model Used in RPSM RPSM2−: mean only, h=2 RPSM2: all aggregations, h=2 RPSM4−: mean only, h=4 RPSM4: all aggregations, h=4 1.00 0.75 Error Type World 2− Type I 0.50 Type II 0.25 0.00 1.00 0.75 World 2 0.50 Error Percentage 0.25 True Model 0.00 1.00 0.75 World 4− 0.50 0.25 0.00 1.00 0.75 World 4 0.50 0.25 0.00 1 2 3 4 5 6 7 8 9 10 1 2 3 4 5 6 7 8 9 10 1 2 3 4 5 6 7 8 9 10 1 2 3 4 5 6 7 8 9 10 Covariate Complexity Figure 7: Type I and Type II error when RPSM2-, RPSM2, RPSM4-, and RPSM4 are applied to data generated by World2-, World2, World4-, and World4 models with increasing covariate complexity, averaged over 100 trials. (RPSM2- with h = 2 and RPSM4- with h = 4). well when true relational dependencies are limited to h = 2. RPSM4+ performs well in all cases. This The results are shown in Figure 7. Along the main is an indication that including the entity identifiers diagonal, the assumptions of the RPSM model are in the propensity model performs blocking, producing consistent with the true world configuration. In cases effects similar to the explicit conditioning performed where the employed model uses mean as the sole ag- by Rattigan et al. [2011]. This also strengthens the gregation function but distributional dependencies are connection between relational blocking and a conjec- more complex, RPSM commits a large number of Type ture made by Perlich and Provost [2006] that the in- I errors. However, the over-specified models (e.g., clusion of identifier variables in a non-causal setting RPSM4 in World2) maintain accuracy levels that are can be used to create a relational fixed or random ef- consistent with the most efficient RPSM configuration. fects model. Given these results, the ability to auto- matically identify and utilize entity identifiers provides 4.3 ENTITY IDENTIFIERS a strong argument for using RPSM as opposed to a propositional approach. While blocking accounts for The final experiment examines the third hypothesis re- a relatively small subset of all possible confounders, it garding the effect of including entity identifiers in the provides a substantial improvement over the alterna- relational propensity score model. We generated data tive of assuming no latent confounders. from a model similar to that of Figure 5, with an addi- tional latent confounder on the Hometown entity. We then created relational covariates using four strategies: 5 RELATED WORK 1. Use all observed variables and hop-threshold of 2 Multi-level propensity score models [Hong and Rau- (RPSM2) and 4 (RPSM4). denbush, 2006, Li et al., 2013] provide a method for 2. Use degree variables and entity-identifiers for all accounting for group or cluster level effects. This eligible blocking paths with either h = 2 or h = 4 corresponds to a one-to-many relationship in a rela- (RPSM2+ and RPSM4+ respectively). tional schema. RPSM can be seen as an extension The results are shown in Figure 8. RPSM2 and of the multi-level setting, capturing not only one-to- RPSM4 perform poorly, because of the bias induced many group level effects, but also many-to-many ef- by unconditioned confounders. RPSM2+ performs fects. There has also been significant progress in un- Model Used in RPSM RPSM2: all aggregations, h=2 RPSM2+: all aggregations, h=2, blocking & degree RPSM4: all aggregations, h=4 RPSM4+: all aggregations, h=4, blocking & degree 1.00 0.75 Error Type World 2 Type I 0.50 Type II 0.25 0.00 1.00 0.75 World 2+ 0.50 Error Percentage 0.25 True Model 0.00 1.00 0.75 World 4 0.50 0.25 0.00 1.00 0.75 World 4+ 0.50 0.25 0.00 1 2 3 4 5 6 7 8 9 10 1 2 3 4 5 6 7 8 9 10 1 2 3 4 5 6 7 8 9 10 1 2 3 4 5 6 7 8 9 10 Covariate Complexity Figure 8: Type I and Type II error when RPSM2, RPSM2+, RPSM4, and RPSM4+ are applied to data generated by World2, World2+, World4, and World4+ with increasing covariate complexity, averaged over 100 trials. derstanding the consequences of the stable unit treat- 6 FUTURE WORK ment value assumption (SUTVA)3 for matching and propensity models in the fields of statistics, epidemi- We plan on examining RPSM further, using more com- ology and econometrics [Hudgens and Halloran, 2008, plex synthetic data and real-world data. An interest- Tchetgen and VanderWeele, 2012, Manski, 2013]. This ing avenue for future research is extending RPSM to work does not address SUTVA violations, but exten- the case where the treatment or outcome lies along a sions to that setting are a focus of future work. one-to-many relational path (e.g., the effect of a treat- ment performed on an individual on an aggregate at- Perlich and Provost [2006] introduced relational fixed tribute of the individual’s friends). There are also a and random effects models using identifier attributes number of methods for performing matching without a as features in the ACORA framework. RPSM differs propensity score, such as matching on the full set of co- in two important aspects. First, the aim of the afore- variates [Stuart, 2010], coarsened exact matching [Ia- mentioned work is predictive, rather than causal. Sec- cus et al., 2012], and entropy balancing [Hainmueller, ond, RPSM incorporates degree variables and provides 2012]. Extending these methods to the relational set- an algorithm for deciding which relational variables ting would allow practitioners flexibility in terms of the should be included, rather than assuming the correct set of assumptions required for a given causal analysis. set of relational variables and aggregating. In the area of relational causal discovery, Maier et al. 7 CONCLUSIONS [2013a] introduced a constraint-based algorithm, RCD, that leverages relational d-separation [Maier et al., Propensity score matching provides a powerful and ro- 2013b] to learn causal models from relational data. bust method for causal inference on propositional data. RCD learns a joint causal model of a relational domain However, naively applying PSM to relational data ig- and abstracts away the mechanics of performing indi- nores both new challenges and opportunities presented vidual tests of conditional independence, while RPSM by this richer type of data. RPSM automatically con- focuses on evaluating a single causal dependence and structs the set of relational covariates to be used in the conditioning mechanism. the propensity score model given a set of confounding attributes, a set of aggregation functions, and a hop threshold. Further, it exploits the relational structure by identifying degree variables and entity identifiers, which can account for latent relational confounders. 3 SUTVA states that the outcome of an individual is We evaluate its efficacy via synthetic experiments that independent of the treatment status of other individuals. leverage a real-world relational skeleton. References M. Maier, K. Marazopoulou, and D. Jensen. Rea- soning about independence in probabilistic models C. Drake. Effects of misspecification of the propensity of relational data. arXiv preprint arXiv:1302.4381, score on estimators of treatment effect. Biometrics, 2013b. 49(4):1231–1236, 1993. C. F. Manski. Identification of treatment response R. A. Fisher. The Design of Experiments. Oliver and with social interactions. The Econometrics Journal, Boyd, Edinburgh, 1935. 16(1):S1–S23, 2013. H. Gao, J. Tang, and H. Liu. gSCorr: Modeling geo- social correlations for new check-ins on location- D. F. McCaffrey, G. Ridgeway, and A. R. Morral. based social networks. In Proceedings of the 21st Propensity score estimation with boosted regression ACM International Conference on Information and for evaluating causal effects in observational studies. Knowledge Management, pages 1582–1586. ACM, Psychological Methods, 9(4):403, 2004. 2012. J. Pearl. Remarks on the method of propensity score. L. Getoor and B. Taskar. Introduction to statistical Statistics in Medicine, 28(9):1415–1416, 2009. relational learning. MIT press, 2007. C. Perlich and F. Provost. Distribution-based aggrega- J. Hainmueller. Entropy balancing for causal effects: tion for relational learning with identifier attributes. A multivariate reweighting method to produce bal- Machine Learning, 62(1-2):65–105, February 2006. anced samples in observational studies. Political M. J. Rattigan, M. Maier, and D. Jensen. Relational Analysis, 20(1):25–46, 2012. blocking for causal discovery. In Proceedings of the B. B. Hansen and S. O. Klopfer. Optimal full matching Twenty-Fifth AAAI Conference on Artificial Intel- and related designs via network flows. Journal of ligence, pages 145–151, 2011. Computational and Graphical Statistics, 15(3):609– P. R. Rosenbaum. A characterization of optimal de- 627, 2006. signs for observational studies. Journal of the Royal D. E. Ho, K. Imai, G. King, and E. A. Stuart. Match- Statistical Society. Series B (Methodological), 53(3): ing as nonparametric preprocessing for reducing 597–610, 1991. model dependence in parametric causal inference. P. R. Rosenbaum and D. B. Rubin. The central role Political Analysis, 15(3):199–236, 2007. of the propensity score in observational studies for G. Hong and S. W. Raudenbush. Evaluating kinder- causal effects. Biometrika, 70(1):41–55, 1983. garten retention policy. Journal of the American D. B. Rubin. Estimating causal effects of treatments Statistical Association, 101(475), 2006. in randomized and nonrandomized studies. Journal M. G. Hudgens and M. E. Halloran. Toward causal of Educational Psychology, 66(5):688–701, 1974. inference with interference. Journal of the American E. A. Stuart. Matching methods for causal inference: Statistical Association, 103(482), 2008. A review and a look forward. Statistical Science, 25 S. M. Iacus, G. King, and G. Porro. Causal inference (1):1–21, 2010. without balance checking: Coarsened exact match- E. J. T. Tchetgen and T. J. VanderWeele. On causal ing. Political Analysis, 20(1):1–24, 2012. inference in the presence of interference. Statistical D. D. Jensen, J. Neville, and M. Hay. Avoiding bias Methods in Medical Research, 21(1):55–75, 2012. when aggregating relational data with degree dispar- D. Westreich, J. Lessler, and M. J. Funk. Propensity ity. In Proceedings of the Twentieth International score estimation: Neural networks, support vector Conference on Machine Learning, pages 274–281. machines, decision trees (cart), and meta-classifiers AAAI Press, 2003. as alternatives to logistic regression. Journal of Clin- B. K. Lee, J. Lessler, and E. A. Stuart. Improving ical Epidemiology, 63(8):826–833, 2010. propensity score weighting using machine learning. Statistics in Medicine, 29(3):337–346, 2010. F. Li, A. M. Zaslavsky, and M. B. Landrum. Propen- sity score weighting with multilevel data. Statistics in Medicine, 32(19):3373–3387, 2013. M. Maier, K. Marazopoulou, D. Arbour, and D. Jensen. A sound and complete algorithm for learning causal models from relational data. In Pro- ceedings of the Twenty-Ninth Conference on Un- certainty in Artificial Intelligence, pages 371–380, 2013a.