Submitted:
13 August 2025
Posted:
14 August 2025
You are already at the latest version
Abstract
Difference-in-Differences (DiD) is a useful statistical technique employed by researchers to estimate the effects of exogenous events on the outcome of some response variables in random samples of treated units (i.e. units exposed to the event) ideally drawn from an infinite population. The term effect should be intended as the difference between the actual post-event realization of the response and the (non-existing and therefore unobservable) hypothetical realization of that same response for the same treated units, were the event absent. To circumvent the implicit missing variables problem, DiD methods use the realizations of the response variable observed in comparable random samples of untreated units. The latter are samples of units drawn from the same infinite population, but they are not exposed to the event. They serve as control or comparison groups. They provide the “substitutes” for the non-existing untreated realizations of the responses in treated units during post-treatment periods. In short, DiD assumes that without treatment, and under certain circumstances, treated units would behave exactly as the control or untreated units during post treatment periods. Then for the estimation purposes, the method adopts a combination of before-after and treatment-control group comparisons. The event that affects the response variables was termed “treatment”, but it could be equally termed “causal factor” to emphasise that with DiD we are not estimating a mere statistical association among variates. With DiD we cultivate the ambition of evaluating whether a precise causative link between causes and effects –defined according to a model based on a proper identification of the relationship among variables– is actually consistent with the data, and estimate how intensive and statistically robust the causal-effect link actually is. DiD analysis has been widely employed in economics, public policy, health research, management, environment analysis, and other fields. There is a discussion about the true “fatherhood” of the method and, not surprisingly, there are clear pioneering antecedents of DiD applications outside economics. Examples include medicine (the study of the causes of London’ worst cholera epidemics of 1849 with 14,137 victims) and agriculture (studies of changes of soil productivity enhanced by new cultivation techniques in Africans’ neighbour areas in the 1980s conducted by revolutionary governments after the victory of their anti-colonialist movements in the second half of the 1970s). A recognised common methodological basis is R. A. Fisher’s analysis of variance (ANOVA). This Review is an introduction to the DiD techniques. It starts from the very basic methods used to estimate the so-called Average Treatment Effect upon Treated (ATET) in a 2–period and 2–group case and proceeds by covering many of the issues that emerge in a multi-unit and multi-period context. Particular attention will be devoted to the statistical assumptions needed for a correct definition of the identification process of the causal-effect relationship in the multi-period case, namely to the parallel trend hypothesis, to the no anticipation assumption, and to the SUTVA assumption. In the multi-period case, both the Homogeneous case (when treated units start being treated in the same periods) and the Heterogeneous case (when treated units start being treated in different periods) will be considered. Some space will be devoted to the developments associated to the DiD techniques employable in the presence of data clustering or spatial-temporal dependence. The Review includes brief presentations of some policy-oriented applications of DiD. Areas covered are income taxation, migration, regulation and environment management.
Keywords:
1. Introduction to DiD
1.1. A 2 × 2 (Two Groups and Two Periods) Homogenous DiD with No Cofactors
- C is the expected value of y for the treated group conditional upon the application of the treatment on that group
- D is the expected value of y for the untreated group conditional upon the absence of the treatment for that group
- A is the expected value of y for the treated group conditional upon the absence of the treatment
- B is the expected value of y for the untreated group conditional upon the absence of the treatment
- yht is the value of the response variable for a unit in the population under study. Its value is measured in each group and each t, i.e. before and after the introduction of the treatment. It will correspond either to the i-th or to j-th observation at time t depending on the group (treated or untreated) of the unit.
- β0 is the intercept of the regression model, common to treated and untreated units.
- D1 is the Time Period Dummy which is a dummy variable that takes the value 0 or 1 depending on whether the h.th observation of the response variable refers to the pre (D1 = 0) or post treatment period (D1 = 1) independently on the group (treated or control) the observation belongs to. It simply indicates if that t is a period in which the treatment existed or not.
- D2 is the Treatment Indicator Dummy which is a dummy variable that takes the value 0 or 1 depending on whether the h.th measurement refers to an individual in the control group (untreated) or in the treatment group respectively, independently on the time period. Therefore, D2 = 0 when the observation belongs to an untreated unit and D2 = 1 when the observation belongs to a treated unit (independently upon when the treatment was introduced). Clearly, in the simplified example of this section with only two periods, D2 = 0 means that the unit is never treated. Other settings are discussed in other sections.
- D1× D2 is the interaction term between the time dummy and the treatment dummy. It s the most important coefficient to estimate. It measures the average effect of the treatment on treated units the estimated average differential impact of the treatment.
- (i)
- corresponds to point B, as above, and must be interpreted as the model baseline average (constant) and
- (ii)
- , which corresponds to segment AB, is the constant difference between the two groups before the treatment.
1.2. Violations of the Parallel Trend Assumption
1.3. The Stable Unit Treatment Value Assumption (SUTVA) (Rubin, 1978, I980, 1990)
1.4. Exogeneity and Identification. DiD and Traditional Econometrics
2. The OLS Version of the Two-Way Fixed Effects Regression (TWFE)
- is the response variable
- is a time effect
- is a unit (not group) fixed effect
- is the dummy (indicator) for whether or not unit h is affected by the treatment in period t (the term D1 × D2 of the last column of Table 2)
- are idiosyncratic, time-varying unobservable factors.


| 1 | 1 | 0 | .5 |
| 1 | 2 | 0 | .5 |
| 1 | 3 | 0 | .5 |
| 1 | 4 | 0 | .5 |
| 1 | 5 | 0 | .5 |
| 1 | 6 | 0 | .5 |
| 1 | 7 | 0 | .5 |
| 1 | 8 | 0 | .5 |
| 1 | 9 | 0 | .5 |
| 1 | 10 | 0 | .5 |
| 2 | 1 | 0 | 1 |
| 2 | 2 | 0 | 1 |
| 2 | 3 | 0 | 1 |
| 2 | 4 | 0 | 1 |
| 2 | 5 | 1 | 2 |
| 2 | 6 | 1 | 2 |
| 2 | 7 | 1 | 2 |
| 2 | 8 | 1 | 2 |
| 2 | 9 | 1 | 2 |
| 2 | 10 | 1 | 2 |
| 3 | 1 | 0 | 2 |
| 3 | 2 | 0 | 2 |
| 3 | 3 | 0 | 2 |
| 3 | 4 | 0 | 2 |
| 3 | 5 | 1 | 4 |
| 3 | 6 | 1 | 4 |
| 3 | 7 | 1 | 4 |
| 3 | 8 | 1 | 4 |
| 3 | 9 | 1 | 4 |
| 3 | 10 | 1 | 4 |
| Consumers’ Id | Time | Consumption € | D1 | D2 |
|---|---|---|---|---|
| 1 | 2010 | 12 | 0 | 1 |
| 2 | 2010 | 9 | 0 | 1 |
| 3 | 2010 | 13 | 0 | 1 |
| 4 | 2010 | 14 | 0 | 1 |
| 5 | 2010 | 15 | 0 | 1 |
| 6 | 2010 | 13 | 0 | 0 |
| 7 | 2010 | 14 | 0 | 0 |
| 8 | 2010 | 13 | 0 | 0 |
| 9 | 2010 | 16 | 0 | 0 |
| 10 | 2010 | 15 | 0 | 0 |
| 1 | 2011 | 15 | 1 | 1 |
| 2 | 2011 | 17 | 1 | 1 |
| 3 | 2011 | 19 | 1 | 1 |
| 4 | 2011 | 18 | 1 | 1 |
| 5 | 2011 | 22 | 1 | 1 |
| 6 | 2011 | 13.5 | 1 | 0 |
| 7 | 2011 | 14 | 1 | 0 |
| 8 | 2011 | 15 | 1 | 0 |
| 9 | 2011 | 15.5 | 1 | 0 |
| 10 | 2011 | 14.4 | 1 | 0 |
- t* and t* – 1 the two periods that for simplicity correspond to two years
- Dh the treatment indicator D1 × D2 of Table 2 so that
2.1. Testing for the Parallel Trends and Anticipation Effects Assumptions in the TWFE Model
2.2. More on the Parallel Trend Assumption
- Selection bias relates to the fixed characteristics of the units
- Time trend is the same for treated and untreated units.
3. Simple Worked Examples
3.1. Example n.1
- The mean Consumption in the Control group before the treatment is
- The mean Consumption in the Treated group before treatment is
- The mean Consumption in Control group after the treatment is
- The mean Consumption in Treated group after the treatment is
| Control | Treated | |
| Pre-Treatment | 14.2 | 12.6 |
| Post-Treatment | 14.48 | 18.2 |

- The estimated Constant = 14.2 (with a p-value smaller than 0.05) is the mean value of the Consumption in the control group in 2010 (i.e. before the treatment). We can compare it with the result obtained from the numerical calculation reported above. The two figures coincide.
- If we sum the coefficient Constant and the d2 coefficient, i.e. if we calculate 14.2 + (– 1.6), we obtain 12.6. This is the expected Consumption of the control group in 2011, i.e. during the year of treatment.
- If we sum the coefficient Constant and the d1 coefficient, i.e. if we calculate 14.2 + 0.28 = 14.48, we obtain the mean value of the Consumption in the treatment group in 2010, i.e. before the treatment.
- The estimated TRET = 5.32 is the (statistically significant) treatment effect. Treated units increase their average consumption by 5.32 euros with respect to untreated individuals.
3.2. Example n.2
4. ATET vs ATE
5. The Confounding Factors
- (1)
- the covariate is associated with treatment
- (2)
- there is a time-varying relationship between the covariate and outcomes
- (3)
- there is differential time evolution in covariate distributions between the treatment and control populations (the covariate must have an effect on the outcome).
6. More Than Two Periods with Homogeneity
7. More Than Two Periods with Heterogeneity
- Irreversibility of the treatment or Staggered treatment (This assumption posits that once units receive treatment, they remain treated throughout the observation period.
-
Parallel Trends Assumption with respect to Never-Treated Units: When we examine groups and periods where treatment isn’t applied (C=1), we assume the average potential outcomes for the group initially treated at time g. The group that never received treatment would have followed similar trends in all post-treatment periods t ≥ g. Then, if we have T = (1, …, S) and g = (2, …, S) with t ≥ g. However, this assumption relies on two important conditions:
- There must be a sufficiently large group of units that have never received treatment in our data.
- These never-treated units must be similar enough to the units that eventually receive treatment so that we can validly compare their outcomes.
- 3.
- Parallel Trends Assumption with respect to Not-Yet Treated Units: When we’re studying groups treated first at time g, we assume that we can use the units that are not-yet treated by time s (where s ≥ t) as valid comparison groups for the group initially treated at time g.
- ▪
- DEPENDENT VARIABLE: CONSUMPTION
- ▪
- COACTOR: INCOME
- ▪
- HETEROGENOUS TREATMENT: A Consumption Credit (for instance a policy measure that supports consumption (for instance a consumption local credit card with public warrant.
| DiD DUMMIES |
|
D1 = 0 if the consumer was never treated D1 = 1 if the consumer was treated, sooner or later D2 = 0 if the treatment did not exist in that year for that consumer D2 = 1 if the treatment exists in that year for that consumer |
|
ID CONSUMERS 1 to 5 are Treated from 2011 6 to 10 are Never Treated 11 to 12 are Treated from 2012 13 is Treated from 2013 to 2014 |
|
TREATMENT TIMING |
|
From 2009 to 2010 No Treatment existed From 2011 to 2012 there was a treatment on individuals 1, 2, 3, 4, and 5 In 2012 a Treatment was extended to individuals 11 and 12 In 2013 a Treatment further extended to individuals of unit 13 |
|
UNITS AND COHORTS | |||
| Cohorts | Units and Observations | ||
| Never Treated Units | 5 units | 30 Observations | |
| First Cohort | Units Treated from 2011 | 5 units | 30 Observations |
| Second Cohort | Units Treated from 2012 | 2 units | 12 Observations |
| Third Cohort | Units Treated from 2013 | 1 unit | 6 Observations |
| ATET(SE in parenthesis) | ||||||
| Cohorts | YEARS | TWFE | RA | IPW | AIPW | |
| 2010 | // | .18 (.20) | .18 (.20) | .18 (.2) | ||
| 2011 | 5,20*** (.99) | 5.32*** (.93) | 5.32*** (.93) | 5.32*** (.93) | ||
| 2011 | 2012 | 5.22*** (1.1) | 5.26*** (1.01) | 5.26 *** (1.01) | 5.26*** (1.01) | |
| 2013 | 6.0*** (1.06) | 6*** (.97) | 6*** (.97) | 6*** (.97) | ||
| 2014 | 5.92*** (1.03) | 5.92*** (.96) | 5.92*** (.96) | 5.92*** (.96) | ||
| 2010 | // | .05 (.24) | .05 (.24) | .05 (.24) | ||
| 2011 | // | ,82 (.47) | .82 (.47) | .82 (.47) | ||
| 2012 | 2012 | 1.23** (.31) | .72*** (.10) | .72*** (.10) | .72 *** (.10) | |
| 2013 | 1.17** (.44) | .62* (.23) | .62* (.23) | .62* (.23) | ||
| 2014 | .97 (1.27) | .42 (.94) | .42 (.94) | .42 (.94) | ||
| 2010 | // | .2 (.16) | .2 (.16) | .2 (.16) | ||
| 2011 | // | -0.08 | -0.08 (.42) | -0.08 (.42) | ||
| 2013 | 2012 | // | .32*** (.07) | .32** (.08) | .32** (.08) | |
| 2013 | 1.5*** (.22) | 1*** (.09) | 1*** (.09) | 1*** (.09) | ||
| 2014 | 1.35** (.43) | 1.1 ** (.24) | 1.1 ( (.25) | 1.1 ( (.25) | ||
| Overall ATET | 4.32** (1.11) | 4.22*** (1.003) | 4.22*** (1.00) | 4.22*** (1.00) | ||
| Average ATET by years | ||||||
| 2011 | 5.2*** (.99) | 5.32*** (.93) | 5.32*** (.93) | 5.32*** (.93) | ||
| 2012 | 4.08** (1.16) | 3.96 ** (1.05) | 3.96 ** (1.05) | 3.96 ** (1.05) | ||
| 2013 | 4.2** (1.21) | 4.03** (1.08) | 4.03** (1.08) | 4.03** (1.08) | ||
| 2014 | 4.11** (1.28) | 3.94** (1.13) | 3.94** (1.13) | 3.94** (1.13) | ||
7.1. The Extended TWFE Method (Wooldridge, 2021)
7.2. The Regression Adjusted Method (Callaway and Sant’Anna, 2021)
7.3. The Inverse Probability Weighting Method, IPW, (Callaway and Sant’Anna, 2021) and the Augmented IPW (Callaway and Sant’Anna, 2021)
8. DiD with Complex Data Structure: Clustering and Spatial-Temporal Dependence
- data showing a grouping or clustering structure
- data exhibiting complex dependence generated by spatial and temporal relationships.
8.1. Clustering
8.2. Serial Correlation
- yhgt is the status of the response variable of individual h in group g in time t;
- βg is the time invariant group effect;
- λt is the group invariant time effect;
- is the interaction dummy representing the treatment state in post-treatment period;
- εhgt reflects the idiosyncratic variation of the response variable across individuals, groups and time.
8.3. Spatial Dependence
9. The Most Relevant Issues Discussed in This Review
- As in many causal inference procedures, DiD relies on strong assumptions that are difficult to test. The key assumption (parallel trends) is that the outcomes of the treated and comparison groups would have evolved similarly in the absence of treatment (the vis inertiae appearing in the title). Yet, even in simple 2 units and 3 periods case the optical (graphical) observation of similar trends in both groups prior to intervention is generally insufficient to establish the existence of post-treatment parallel trends. The issue become more complicated in the multi-unit and multi-period cases and makes it questionable the use of untreated observations as the appropriate counterfactuals for the (non-existing) untreated observations of treated units in the treatment periods. The search for the existence of parallel trends might became a search for the Arabian Phoenix since it requires elaborated statistical tests. The simple graphical appearance of a commune time path of mean realizations in the pre-treatment period might be a misleading suggestion of the perpetuation of a similar potential parallel trend path in the post treatment periods (when counterfactuals cannot be observed).
- Therefore, without a true randomized experiment, tools like DiD do not broaden the range of “natural experiments” we can use to identify causal effects.
- Even in the case of true randomization, SUTVA problems (so called spill-over effects across treated and untreated unites) might plague estimations and make it difficult to identify a DiD model that consistently estimate ATET (which requires unique potential outcome for each individual under each exposure condition).
- Often the interpretation of the role of covariates in DiD estimates is difficult and, sometimes, even what a covariate is might be controversial. In fact, DiD does not require the treated and comparison groups to be balanced on covariates, unlike in cross-sectional OLS studies. Thus, a covariate that differs by treatment group and is associated with the outcome is not necessarily a confounder in DiD. Only covariates that differ by treatment group and are associated with outcome trends are confounders in DiD as these can be the ones that violate the identification assumptions.
- Importantly, it can matter whether we believe the “correct” model is a linear probability model, probit or logit, since they assume different counterfactuals. Determining that two groups would have experienced parallel trends requires, first of all, a justification of the chosen functional forms for the adopted model.
10. Some Examples of DiD Applications
10.1. The Elasticity of Taxable Income (Feldstein, 1995)
- 4.
- The structural approach (closer to the “old” theoretical analysis of labour responses to income taxation) which separately account for each of the potential responses to taxation (intensive and extensive) and then aggregate.
- 5.
- The DiD approach first proposed by Feldstein (1995) which aims at estimating the elasticity of taxable income with respect to the net-of-tax rate and claims that this elasticity is a sufficient statistic for calculating the possible deadweight loss of income taxation.
- TI = the Taxable Income (defined as an aggregate measure of income from various sources)
- τ = proportional income tax rate

The use of tax return data rather than of a household survey permits analysing the response of taxable income as a whole and not just of labour force participation and working hours. A panel, in which each individual is observed both before and after the change in tax rates, permits a "differences-in-differences" form of estimator that identifies the tax effect in a way that is not available with a single year’s cross section.
- The income growth rate is the same for all income earners (medium, high and highest tax brackets) absent the treatment (“parallel trend assumption”).
- The taxpayers cannot adjust their income in 1985 (last year before reform) as to “choose” their change in tax rate through TRA1986 (“no selection into treatment” and no anticipation effect).
- The comparison of taxpayers that vary in the intensity of treatment (instead of comparing taxed to untaxed taxpayers) is legitimate. Implicitly, he needs to assume that the elasticity of taxable income is constant in income, i.e., the same across all income groups. This last assumption will reappear in other papers.

- Estimates of the elasticities are estimates high, ranging from 1 to 3.
- The so-called Laffer rate i.e. the rate that maximises the tax revenue, changes with the elasticity and corresponds to 1/(1+ϵ)
- The USA are on the wrong side of the Laffer curve (excessive levels of income tax rates)?
- No proper untreated control group is present in the study. Treatment and control groups differ in the intensity of treatment.
- An equal elasticity of taxable income across the income distribution is assumed. Elasticity of taxable income is likely higher for high-income taxpayers (with more adjustment opportunities).
- Small and unstratified sample: very few high-income taxpayers are included.
- The presence of increasing earnings inequality in the US determined by for non-tax reasons should be considered.
- Results may be affected by a regression-to-the-mean bias due to classification of treatment groups by pre-treatment income: Rich people in year t may tend to revert to the mean in year t+1.
- Panel analysis introduces a downward bias in the estimated elasticity if marginal tax rate for rich people decreases.
- It is unclear whether the common trend assumption really holds. Not even the simplest tests are conducted (parallel trends, anticipation effects, etc.).
- Estimated elasticity overestimates welfare loss if behavioural response involves transfers between individuals.
- The study really provides some shaky indication about the effects of changes of MTR on the aggregate income tax yield, but it is silent about taxpayers behavioural reactions to income taxation in spite of the claim that “The Tax Reform Act of 1986 is a particularly useful natural experiment for studying the responsiveness of taxpayers to changes in marginal tax rates” (Feldstein, 1995 p. 552). The potential role that confounders (likely affected by the treatment) may play in this estimation is completely ignored.
10.2. Top Income Taxation and the Migration Decisions of Rich Taxpayers (Kleven, Landais, and Saez, 2013)
- pnd = total of domestic players in country n
- pnf = total of foreign players in country n

- In the graphical analysis, the elasticities of the Average Tax Rate are not presented for the pre-Bosman period and the Danish case studies because of lack of individual earnings data before 1996. Similarly, the average tax rate elasticity for Spain is based on the 1996-2003 versus 2004-2008 comparison. It is therefore difficult to conduct a complete comparison study (not even graphical).
- The sample used is limited to a very special category of privileged migrants (the well-paid football players whose behaviour is affected by several treatment-related confounding factors). Out of sample projections seems problematic.
- Bosman ruling could have had differential impacts on low-tax and high-tax countries for nontax reasons. Tax rates may correlate with country size and thus league quality. Better leagues may have benefited more from Bosman ruling.
- Football players contracts are generally signed in advance with respect of the year of the actual transfer and then anticipation effects of the Borman ruling might be present.
- Other factors could have changed from the pre-Bosman to the post-Bosman era that impacted low-tax and high-tax countries differentially.
10.3. Toxic Emission and the Environment (Zhou, Zhang, Song, and Wang, 2019; Dong, Li, Qin, Zhang, Chen, Zhao, and Wang, 2022)

10.4. Regulation, Privatization, Management (Galiani, Gertler, and Schargrodsky, 2005; Gertler et al. 2016)

Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
Appendix A. An Example with an Easy Visualization of the Data Set
References
- Angrist J. D. and J.-S. Pischke (2009), Mostly Harmless Econometrics: An Empiricist’s Companion. Princeton University Press.
- Angrist J. D. and J.-S. Pischke (2010), The Credibility Revolution in Empirical Economics: How Better Research Design is Taking the Con out of Econometrics. Journal of Economic Perspectives, 24, 2, p. 3–30.
- Angrist, J. D., G. W. Imbens, & D. B Rubin, (1996). Identification of Causal Effects Using Instrumental Variables. Journal of the American Statistical Association, 91(434), 444-455. [CrossRef]
- Ashenfelter O. & D. Card (1985) Using the Longitudinal Structure of Earnings to Estimate the Effect of Training Programs The Review of Economics and Statistics, 67, 4, pp. 648-660.
- Bertrand, M., E. Duflo, & S. Mullainathan (2004), How much should we trust difference-in-differences estimates? Quarterly Journal of Economics 119: 249–275. [CrossRef]
- Bester, C. A., T. G. Conley, and C. B. Hansen (2011), Inference with dependent data using cluster covariance estimators. Journal of Econometrics 165: 137–151. [CrossRef]
- Bilinski A. and L. Hatfield (2020), Nothing to See Here? Non-Inferiority Approaches to Parallel Trends and Other Model Assumptions, JSM 2020 Virtual Conference, august, https://ww2.amstat.org/meetings/jsm/2020/onlineprogram/AbstractDetails.cfm?abstractid=312323.
- Borusyak, Kirill, Jaravel, Xavier, 2018. Revisiting Event Study Designs. SSRN Scholarly Paper ID 2826228, Social Science Research Network, Rochester, NY.
- Callaway B. (2022), Difference-in-Differences for Policy Evaluation, in K. F. Zimmermann (ed.), Handbook of Labor, Human Resources and Population, pp. 1–61 . [CrossRef]
- Callaway B. and P. Sant’Anna (2021), Difference-in-differences with multiple time periods. Journal of Econometrics 225, pp. 200–230. [CrossRef]
- Cameron A. C. and P. K. Trivedi (2005) Microeconometrics: Methods and Applications, Cambridge University Press.
- Card D. and A. Krueger (1994), Minimum Wages and Employment: A Case Study of the Fast-Food Industry in New Jersey and Pennsylvania, American Economic Review, 84, issue 4, p. 772–93.
- Cerqua, A., Letta, M., and Menchetti, F. (2022). Losing control (group)? The Machine Learning Control Method for counterfactual forecasting.
- Cerqua, A., Letta, M., and Menchetti, F. (2023). The Machine Learning Control Method for Counterfactual Forecasting.
- Cerulli G. (2015), Econometric evaluation of socio-economic programs. Theory and applications. Springer-Verlag GmbH.
- Chesnaye N., Stel S., Tripepi G., Dekker F. W., Fu E. L., Zoccali G., and Jager K. J. (2022), An introduction to inverse probability of treatment weighting in observational research, Clinical Kidney Journal, 15, 1, pp. 14–20 doi: 10.1093/ckj/sfab158.
- Cole, S. R., and Frangakis, C. E. (2009). The Consistency Statement in Causal Inference: A Definition or an Assumption? Epidemiology, 20(1), 3-5. [CrossRef]
- Cox, D. R. (1958). Planning of experiments, Wiley Series in Probability and Statistics - Applied Probability and Statistics Section.
- de Chaisemartin C. and X. D’Haultfoeuille (2020), Two-Way Fixed Effects Estimators with Heterogeneous Treatment Effects, American Economic Review vol. 110, 9, pp. 2964–96.
- de Chaisemartin C. and X. D’Haultfoeuille (2023), Two-way fixed effects estimators with heterogeneous treatment effects: a survey, Econometrics Journal, 26, pp. C1–C30 . [CrossRef]
- Delgado, M. S., and Florax, R. J. G. M. (2015). Difference-in-differences techniques for spatial data: Local autocorrelation and spatial interaction. Economics Letters, 137, 123-126. [CrossRef]
- Elhorst, J. P. (2010). Applied Spatial Econometrics: Raising the Bar. Spatial Economic Analysis, 5(1), 9-28. [CrossRef]
- Fisher, R. A. (1935), Design of Experiments, Oliver and Boyd.
- Freyaldenhoven, S., C. Hansen, and Shapiro, J. (2019), Pre-event Trends in the Panel Event-study Design”, American Economic Review, 109, pp. 3307–3338.
- Goodman-Bacon A. (2021), Difference-in-differences with variation in treatment timing, Journal of Econometrics, 225, 2, pp. 254-277.
- Huber, M., and Steinmayr, A. (2021). A Framework for Separating Individual-Level Treatment Effects From Spillover Effects. Journal of Business and Economic Statistics, 39(2), 422-436. [CrossRef]
- Imbens G. W. and D. B. Rubin (2015), Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction, Cambridge University Press.
- Kahn-Lang, A. and Lang, K. (2020), The Promise and Pitfalls of Differences-in-Differences: Reflections on 16 and Pregnant and Other Applications, Journal of Business and Economic Statistics, 38, pp. 613–620.
- Laffers, L., and Mellace, G. (2020). Identification of the average treatment effect when SUTVA is violated. Discussion Papers on Business and Economics, University of Southern Denmark, 3.
- Myint L. (2024), Controlling time-varying confounding in difference-in-differences studies using the time-varying treatments framework. Health Services and Outcomes Research Methodology, 24, pp.95–111 . [CrossRef]
- Ogburn, E. L., Shpitser, I., and Lee, Y. (2020). Causal Inference, Social Networks and Chain Graphs. Journal of the Royal Statistical Society Series A: Statistics in Society, 183(4), 1659-1676. [CrossRef]
- Ogburn, E. L., Sofrygin, O., Díaz, I., and van der Laan, M. J. (2024). Causal Inference for Social Network Data. Journal of the American Statistical Association, 119(545), 597-611. [CrossRef]
- Qiu, F., and Tong, Q. (2021). A spatial difference-in-differences approach to evaluate the impact of light rail transit on property values. Economic Modelling, 99, 105496. [CrossRef]
- Rambachan A. and J. Roth (2023), A More Credible Approach to Parallel Trends, Review of Economic Studies 90, pp. 2555–2591.
- Rosenbaum P. R. (1984), The Consequences of Adjustment for a Concomitant Variable That Has Been Affected by the Treatment. Journal of the Royal Statistical Society. Series A (General), 147, 5, pp. 656-666.
- Roth J. (2022), Pretest with Caution: Event-Study Estimates after Testing for Parallel Trends, American Economic Review: Insights, 4, 3, pp. 305–22.
- Roth J., P. Sant’Anna, A. Bilinski, and J. Poe (2023), What’s trending in difference-in-differences? A synthesis of the recent econometrics literature, Journal of Econometrics 235, pp. 2218–2244.
- Rubin, D. B. (1978), Bayesian Inference for Causal Effects: The Role of Randomization, Annals of Statistics, Vol. 6: pp. 34–58.
- Rubin, D. B. (1980), Randomization Analysis of Experimental Data: The Fisher Randomization Test Comment, Journal of the American Statistical Association, Vol. 75, pp. 591-593. Rubin, D. B. (1990), Formal Modes of Statistical Inference for Causal Effects, Journal of Statistical Planning and Inference, Vol. 25: pp. 279–292.
- Rubin, D. B. (1980). Randomization Analysis of Experimental Data: The Fisher Randomization Test Comment. Journal of the American Statistical Association, 75(371), 591-593. [CrossRef]
- Rubin, D. B. (1990). [On the Application of Probability Theory to Agricultural Experiments. Essay on Principles. Section 9.] Comment: Neyman (1923) and Causal Inference in Experiments and observational Studies. Statistical Science, 5(4), 472-480, 479. [CrossRef]
- Schwartz, S., Gatto, N. M., and Campbell, U. B. (2012). Extending the sufficient component cause model to describe the Stable Unit Treatment Value Assumption (SUTVA). Epidemiologic Perspectives and Innovations, 9(1), 3. [CrossRef]
- Sobel, M. E. (2006). What Do Randomized Studies of Housing Mobility Demonstrate? Journal of the American Statistical Association, 101(476), 1398-1407. [CrossRef]
- Sun, S., and Delgado, M. S. (2024). Local spatial difference-in-differences models: treatment correlations, response interactions, and expanded local models. Empirical Economics. [CrossRef]
- VanderWeele, T. J. (2009). Concerning the Consistency Assumption in Causal Inference. Epidemiology, 20(6), 880-883. [CrossRef]
- VanderWeele, T. J. (2010). Direct and Indirect Effects for Neighborhood-Based Clustered and Longitudinal Data. Sociological Methods and Research, 38(4), 515-544. [CrossRef]
- VanderWeele, T. J., Tchetgen, E. J. T., and Halloran, M. E. (2015). Interference and sensitivity analysis. Statistical science: a review journal of the Institute of Mathematical Statistics, 29(4), 687.
- Wang, Y. (2021). Causal Inference with Panel Data under Temporal and Spatial Interference. arXiv preprint arXiv:2106.15074.
- Wang, Y., Samii, C., Chang, H., and Aronow, P. (2020). Design-based inference for spatial experiments under unknown interference. arXiv preprint arXiv:2010.13599.
- Wooldridge J. M. (2010), Econometric Analysis of Cross Section and Panel Data, Second Edition, The MIT Press.
- Wooldridge, J. M. (2021), Two-Way Fixed Effects, the Two-Way Mundlak Regression, and Difference-in-Differences Estimators, Available at SSRN: https://ssrn.com/abstract=3906345. [CrossRef]
- Wooldridge, J. M. (2023), Simple approaches to nonlinear difference-in-differences with panel data. Econometrics Journal, 26, pp. C31–C66 . [CrossRef]
- Xu, Y. (2024). Causal Inference with Time-Series Cross-Sectional Data: A Reflection. In J. M. Box-Steffensmeier, D. P. Christenson, and V. Sinclair-Chapman (Eds.), Oxford Handbook of Engaged Methodological Pluralism in Political Science (pp. 0). Oxford University Press. [CrossRef]
- Zeldow B. and L.A. Hatfield (2021), Confounding and regression adjustment in difference-in-differences studies, Health Services Research 56, pp. 932–941 . [CrossRef]
References for Section 10 (Applications) and Further Readings
- Abadie A., A. Diamond, A., & J. Hainmueller (2015). Comparative Politics and the Synthetic Control Method. American Journal of Political Science. 59 (2): pp. 495–510 https://doi.org/10.1111/ajps.12116
- Bosco B.P., C.F. Bosco, & P. Maranzano (2025). Labour responsiveness to income tax changes: empirical evidence from a DID analysis of an income tax treatment in Italy. Empirical Economics https://doi.org/10.1007/s00181-025-02748-7
- Courtemanche, C. J., and D. Zapata (2014). Does universal coverage improve health? The Massachusetts experience. Journal of Policy Analysis and Management, 33, pp. 36–69.
- Dong F., Y. Li, C. Qin, X. Zhang, Y. Chen, X. Zhao, and C. Wang (2022), Information infrastructure and greenhouse gas emission performance in urban China: A difference-in-differences analysis, Journal of Environmental Management 316, 115252 https://doi.org/10.1016/j.jenvman.2022.115252
- Feldstein M. (1995), The Effect of Marginal Tax Rates on Taxable Income: A Panel Study of the 1986 Tax Reform Act, Journal of Political Economy, 103, pp. 551–572.
- Feldstein M. (1999), Tax Avoidance And The Deadweight Loss Of The Income Tax, The Review of Economics and Statistics, 81(4): pp. 674-680
- Fredriksson A., and G. Magalhães de Oliveira (2019), Impact evaluation using Difference-in-Differences, RAUSP Management Journal, 54, 4, pp. 519-532
- Galiani, Gertler, and Schargrodsky (2005), Water for life: The impact of the privatization of water services on child mortality. Journal of Political Economy, 113, pp. 83–120.
- Gertler, P. J., Martinez, S., Premand, P., Rawlings, L. B., and Vermeersch, C. M. (2016). Impact evaluation in practice, Washington, DC: The World Bank.
- Goolsbee A. (2000), What Happens When You Tax the Rich? Evidence from Executive Compensation, Journal of Political Economy, 108, pp. 352–378.
- Jakobsen K., K. Jakobsen, H. Kleven, and G. Zucman (2020), Wealth Taxation and Wealth Accumulation: Theory and Evidence from Denmark, The Quarterly Journal of Economics, 135(1), pp. 329–388.
- Johannesen N. and G. Zucman (2014), The End of Bank Secrecy? An Evaluation of the G20 Tax Haven Crackdown, American Economic Journal: Economic Policy, 6(1), pp. 65–91.
- Kleven H., J. M. Knudsen, C. Kreiner, S. Pedersen, and E. Saez (2011), Unwilling or Unable to Cheat? Evidence from a Tax Audit Experiment in Denmark, Econometrica, 79(3), pp. 651–692.
- Kleven k. J., C. Landais, E. Saez (2013), Taxation and International Migration of Superstars: Evidence from the European Football Market, American Economic Review, 103(5), pp. 1892–1924.
- Moretti E. and D. J. Wilson (2017), The Effect of State Taxes on the Geographical Location of Top Earners: Evidence from Star Scientists, American Economic Review 107, 7, pp. 1858-1903
- Naritomi J. (2019), Consumers as Tax Auditors, American Economic Review, 109(9), pp. 3031–3072.
- Schnabl, P. (2012). The international transmission of bank liquidity shocks: Evidence from an emerging market. The Journal of Finance, 67, pp. 897–932
- Tørsløv T., L. Wierand, and G. Zucman (2023), Externalities in International Tax Enforcement, American Economic Journal: Economic Policy, vol. 15, 2, pp. 497–525.
- Zhou B., C. Zhang, H. Song, and Q. Wang (2019), How does emission trading reduce China’s carbon intensity? An exploration using a decomposition and difference-indifferences approach, Science of the Total Environment 676, pp. 514–523. https://doi.org/10.1016/j.scitotenv.2019.04.303
| 1 | Callaway (2022) discusses an ampler set of estimation strategies. According to Callaway (2022, 4) all of them explicitly make, in a first step, the same good comparisons that show up in the TWFE regression (i.e., the comparisons that use units that become treated relative to units that are not-yet-treated) while explicitly avoiding the “bad comparisons” that show up in the TWFE regression (i.e., the comparisons that use already-treated units as the comparison group). Then, in a second step, they combine these underlying treatment effect parameters into target parameters of interest such as an overall average treatment effect on the treated. See Section Alternative Approaches in Callaway (2022, 20). |
| 2 | A number of different ICC statistics have been proposed, not all of which estimate the same population parameter. There has been considerable debate about which ICC statistics are appropriate for a given use, since they may produce markedly different results for the same data |
| 3 | A list of bias correction procedures is provided by Angrist and Pischke (2009, p. 320-2). |
| 4 | In a later paper (Feldstein, 1999) he also argues that traditional analyses of the income tax greatly underestimate deadweight losses by ignoring its effect on forms of compensation and patterns of consumption. He calculated the full deadweight loss using the compensated elasticity of taxable income to changes in tax rates because leisure, excludable income, and deductible consumption are assumed (by Feldstein) to be a Hicksian composite good. According to his estimations a deadweight loss of as much as 30% of revenue or more than ten times Harberger’s classic 1964 estimate. The relative deadweight loss caused by increasing existing tax rates is substantially greater and, according to Feldstein’s results, may exceed $2 per $1 of revenue. Some enormous measure, one should say! |
| 5 | The treatment incorporated in the Feldstein’s analysis was the 1986 US tax reform that lowered marginal tax rates, and simultaneously broadened tax bases. The two elements were designed to net out. Approximately no revenue and distributional effects absent behavioural responses means that approximately there are no income effects. Important as the aim is to estimate the compensated elasticity of taxable income. |
| 6 | This Review does not discuss synthetic controls. One should see Abadie et al (2015). A synthetic control can be constructed as a weighted average of several units combined to recreate the trajectory that the response variable of a treated unit would have followed in the absence of the treatment. |
| 7 | This is in contrast with the view that the appropriate tax rate for decisions on the intensive margin is the marginal tax rate (MTR = tax rate on the last euro earned). In the paper ATR is not exact but approximated (for a subsample of football players). Since these taxpayers earn very high salaries, authors approximate the ATR by the top marginal tax rate (MTR). An alternative, and possible more reliable procedure is followed by Moretti and Wilson (2017). By focusing on the locational outcomes of star scientists, defined as scientists with patent counts in the top 5 percent of the distribution, their paper quantifies how sensitive is migration by these stars to changes in personal and business tax differentials across states in the USA. The study uncovers large, stable, and precisely estimated effects of personal and corporate taxes on star scientists’ migration patterns. The long-run elasticity of mobility relative to taxes is 1.8 for personal income taxes, 1.9 for state corporate income tax, and -1.7 for the investment tax credit. |




| D1 = 0 | D1 = 1 | |
|---|---|---|
| D2 = 0 | ||
| D2 = 1 |
| Consumers’ Identity | Time | Consumption € | D1 | D2 |
| 1 | 2009 | 11 | 0 | 1 |
| 1 | 2010 | 12 | 0 | 1 |
| 1 | 2011 | 15 | 1 | 1 |
| 2 | 2009 | 8.6 | 0 | 1 |
| 2 | 2010 | 9 | 0 | 1 |
| 2 | 2011 | 17 | 1 | 1 |
| 3 | 2009 | 12.5 | 0 | 1 |
| 3 | 2010 | 13 | 0 | 1 |
| 3 | 2011 | 19 | 1 | 1 |
| 4 | 2009 | 13 | 0 | 1 |
| 4 | 2010 | 14 | 0 | 1 |
| 4 | 2011 | 18 | 1 | 1 |
| 5 | 2009 | 14 | 0 | 1 |
| 5 | 2010 | 15 | 0 | 1 |
| 5 | 2011 | 22 | 1 | 1 |
| 6 | 2009 | 12 | 0 | 1 |
| 6 | 2010 | 13 | 0 | 0 |
| 6 | 2011 | 13.5 | 1 | 0 |
| 7 | 2009 | 13.7 | 0 | 0 |
| 7 | 2010 | 14 | 0 | 0 |
| 7 | 2011 | 14 | 1 | 0 |
| 8 | 2009 | 12.7 | 0 | 0 |
| 8 | 2010 | 13 | 0 | 0 |
| 8 | 2011 | 15 | 1 | 0 |
| 9 | 2009 | 14.9 | 0 | 0 |
| 9 | 2010 | 16 | 0 | 0 |
| 9 | 2011 | 15.5 | 1 | 0 |
| 10 | 2009 | 14.7 | 0 | 0 |
| 10 | 2010 | 15 | 0 | 0 |
| 10 | 2011 | 14.4 | 1 | 0 |
| 2000 | 2001 | 2002 | 2003 | 2004 | 2005 | 2006 | 2007 | |
|
UNITS 1 2 |
Period 1 (no treatment Dit = 0) | Period 2 (treatment Dit = 0 ˄ Djt = 1) | ||||||
|
Never treated | ||||||||
| 3 | Not yet treated | Treated since 2003 until 2007 | ||||||
| 4 | Not yet treated | Treated since 2003 until 2007 | ||||||
| 5 | Not yet treated | Treated since 2003 until 2007 | ||||||
| 6 | Not yet treated | Treated since 2003 until 2007 | ||||||
| YEARS UNITS |
2000 | 2001 | 2002 | 2003 | 2004 | 2005 | 2006 | 2007 |
|---|---|---|---|---|---|---|---|---|
| 1 2 |
Never treated |
|||||||
| 3 | Not yet treated | Treated | ||||||
| 4 | Not yet treated | Treated | ||||||
| 5 | Not yet treated | Treated | ||||||
| 6 | Not yet treated | Treated | ||||||
| ID | Year | Consumption | D1 | D2 | TRET | First Year of Treatment |
|---|---|---|---|---|---|---|
| 1 | 2009 | 11 | 1 | 0 | 0 | 2011 |
| 1 | 2010 | 12 | 1 | 0 | 0 | 2011 |
| 1 | 2011 | 15 | 1 | 1 | 1 | 2011 |
| 1 | 2012 | 14.8 | 1 | 1 | 1 | 2011 |
| 1 | 2013 | 15.8 | 1 | 1 | 1 | 2011 |
| 1 | 2014 | 17 | 1 | 1 | 1 | 2011 |
| 2 | 2009 | 8.6 | 1 | 0 | 0 | 2011 |
| 2 | 2010 | 9 | 1 | 0 | 0 | 2011 |
| 2 | 2011 | 17 | 1 | 1 | 1 | 2011 |
| 2 | 2012 | 18 | 1 | 1 | 1 | 2011 |
| 2 | 2013 | 18.8 | 1 | 1 | 1 | 2011 |
| 2 | 2014 | 19.1 | 1 | 1 | 1 | 2011 |
| 3 | 2009 | 12.5 | 1 | 0 | 0 | 2011 |
| 3 | 2010 | 13 | 1 | 0 | 0 | 2011 |
| 3 | 2011 | 19 | 1 | 1 | 1 | 2011 |
| 3 | 2012 | 19.8 | 1 | 1 | 1 | 2011 |
| 3 | 2013 | 21 | 1 | 1 | 1 | 2011 |
| 3 | 2014 | 22 | 1 | 1 | 1 | 2011 |
| 4 | 2009 | 13 | 1 | 0 | 0 | 2011 |
| 4 | 2010 | 14 | 1 | 0 | 0 | 2011 |
| 4 | 2011 | 18 | 1 | 1 | 1 | 2011 |
| 4 | 2012 | 19.1 | 1 | 1 | 1 | 2011 |
| 4 | 2013 | 22 | 1 | 1 | 1 | 2011 |
| 4 | 2014 | 21.8 | 1 | 1 | 1 | 2011 |
| 5 | 2009 | 14 | 1 | 0 | 0 | 2011 |
| 5 | 2010 | 15 | 1 | 0 | 0 | 2011 |
| 5 | 2011 | 22 | 1 | 1 | 1 | 2011 |
| 5 | 2012 | 21.9 | 1 | 1 | 1 | 2011 |
| 5 | 2013 | 22.2 | 1 | 1 | 1 | 2011 |
| 5 | 2014 | 22 | 1 | 1 | 1 | 2011 |
| 6 | 2009 | 12 | 0 | 0 | 0 | Never treated |
| 6 | 2010 | 13 | 0 | 0 | 0 | Never treated |
| 6 | 2011 | 13.5 | 0 | 1 | 0 | Never treated |
| 6 | 2012 | 13.9 | 0 | 1 | 0 | Never treated |
| 6 | 2013 | 14.2 | 0 | 1 | 0 | Never treated |
| 6 | 2014 | 15.1 | 0 | 1 | 0 | Never treated |
| 7 | 2009 | 13.7 | 0 | 0 | 0 | Never treated |
| 7 | 2010 | 14 | 0 | 0 | 0 | Never treated |
| 7 | 2011 | 14 | 0 | 1 | 0 | Never treated |
| 7 | 2012 | 14.9 | 0 | 1 | 0 | Never treated |
| 7 | 2013 | 15.1 | 0 | 1 | 0 | Never treated |
| 7 | 2014 | 14.9 | 0 | 1 | 0 | Never treated |
| 8 | 2009 | 12.7 | 0 | 0 | 0 | Never treated |
| 8 | 2010 | 13 | 0 | 0 | 0 | Never treated |
| 8 | 2011 | 15 | 0 | 1 | 0 | Never treated |
| 8 | 2012 | 15.5 | 0 | 1 | 0 | Never treated |
| 8 | 2013 | 16.1 | 0 | 1 | 0 | Never treated |
| 8 | 2014 | 17.2 | 0 | 1 | 0 | Never treated |
| 9 | 2009 | 14.9 | 0 | 0 | 0 | Never treated |
| 9 | 2010 | 16 | 0 | 0 | 0 | Never treated |
| 9 | 2011 | 15.5 | 0 | 1 | 0 | Never treated |
| 9 | 2012 | 16 | 0 | 1 | 0 | Never treated |
| 9 | 2013 | 16.7 | 0 | 1 | 0 | Never treated |
| 9 | 2014 | 17 | 0 | 1 | 0 | Never treated |
| 10 | 2009 | 14.7 | 0 | 0 | 0 | Never treated |
| 10 | 2010 | 15 | 0 | 0 | 0 | Never treated |
| 10 | 2011 | 14.4 | 0 | 1 | 0 | Never treated |
| 10 | 2012 | 15 | 0 | 1 | 0 | Never treated |
| 10 | 2013 | 15.7 | 0 | 1 | 0 | Never treated |
| 10 | 2014 | 16.1 | 0 | 1 | 0 | Never treated |
| 11 | 2009 | 13.1 | 1 | 0 | 0 | 2012 |
| 11 | 2010 | 14 | 1 | 0 | 0 | 2012 |
| 11 | 2011 | 14.8 | 1 | 0 | 0 | 2012 |
| 11 | 2012 | 16 | 1 | 1 | 1 | 2012 |
| 11 | 2013 | 16.2 | 1 | 1 | 1 | 2012 |
| 11 | 2014 | 15.5 | 1 | 1 | 1 | 2012 |
| 12 | 2009 | 12.9 | 1 | 0 | 0 | 2012 |
| 12 | 2010 | 13.3 | 1 | 0 | 0 | 2012 |
| 12 | 2011 | 14.7 | 1 | 0 | 0 | 2012 |
| 12 | 2012 | 16.1 | 1 | 1 | 1 | 2012 |
| 12 | 2013 | 16.7 | 1 | 1 | 1 | 2012 |
| 12 | 2014 | 18 | 1 | 1 | 1 | 2012 |
| 13 | 2009 | 12 | 1 | 0 | 0 | 2013 |
| 13 | 2010 | 12.8 | 1 | 0 | 0 | 2013 |
| 13 | 2011 | 13 | 1 | 0 | 0 | 2013 |
| 13 | 2012 | 13.9 | 1 | 0 | 0 | 2013 |
| 13 | 2013 | 15.4 | 1 | 1 | 1 | 2013 |
| 13 | 2014 | 16 | 1 | 1 | 1 | 2013 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).