Difference-in-differences
Estimates causal effects by comparing changes over time between a treated and a control group. Identifies the ATT under the parallel trends assumption, with no randomization required.
THE ESTIMAND
DiD targets the average treatment effect on the treated (ATT),[3]the average effect among units that actually received treatment. It does not recover the ATE without additional assumptions. Morgan & Winship situate repeated-observation designs of this kind within the counterfactual framework.[5]
THE LOGIC
Take the change in outcomes for treated units over time, subtract the change for control units over the same period. What remains is the treatment effect, provided both groups would have followed the same trend absent treatment.
THE 2×2 CASE
| Pre | Post | |
|---|---|---|
| Treated | Y¹₀ | Y¹₁ |
| Control | Y⁰₀ | Y⁰₁ |
| DiD estimate | (Y¹₁ − Y¹₀) − (Y⁰₁ − Y⁰₀) | |
CAUSAL STRUCTURE
In the absence of treatment, the average outcome for treated and control units would have followed the same trend over time. This is the core, and untestable in post-periods, assumption of DiD.
HOW TO TEST
Pre-period event study. Inspect coefficients on relative time dummies before treatment, they should be statistically and economically close to zero.
Units do not change behavior before treatment begins in anticipation of receiving it. Violated if firms or individuals respond to announced policies before they take effect.
HOW TO TEST
Pre-trend test at t−1, t−2. Statistically significant pre-period coefficients often signal anticipation effects.
Potential outcomes for unit i depend only on unit i's treatment status, no spillovers to other units, and only one version of treatment exists.
HOW TO TEST
Theoretical argument. Check for geographic spillovers by testing outcomes in border regions.
Both treated and control groups exist throughout the panel. Pure time-series units (always treated) cannot contribute to identification.
HOW TO TEST
Inspect treatment timing distribution. Ensure sufficient control units across all time periods.
Panel structure
Repeated observations of the same units over time. Minimum: two periods (pre and post). Unbalanced panels work but require care.
Treatment variation
Some units treated, others not, or units treated at different times. Cross-sectional variation in treatment timing is the source of identification.
Outcome variable
Observed for all units in all periods. Should be the same measure pre and post. Level or log depending on the estimand and interpretation.
library(tidyverse)
# Panel must have: unit id, time period, treatment indicator
panel <- read_csv("panel_data.csv") |>
mutate(
# Binary treatment: 1 if unit i is treated at time t
treated = as.integer(state_treated & year >= treat_year),
# Relative time: periods since/before treatment
rel_time = year - treat_year,
# Cap endpoints so sparse far-out bins don't drive the event study
rel_time_binned = case_when(
rel_time <= -4 ~ -4L,
rel_time >= 4 ~ 4L,
TRUE ~ rel_time
)
)
# Check panel balance
panel |>
count(unit_id, year) |>
filter(n > 1) # should be empty
# Check support: observations per relative-time bin
count(panel, rel_time_binned)The canonical DiD estimator adds unit and time fixed effects to a regression of the outcome on a treatment dummy. Unit FEs control for all time-invariant confounders; time FEs absorb common shocks. The coefficient on treated is the ATT, under parallel trends and homogeneous effects across cohorts.[2]
library(fixest)
# Two-way fixed effects DiD
# Unit FE absorbs time-invariant differences
# Time FE absorbs common trends
fit <- feols(
outcome ~ treated | unit_id + year,
data = panel,
cluster = ~unit_id # cluster SEs at treatment level
)
summary(fit)
# coefplot(fit) # visual coefficient plotNot a separate method, but a richer DiD specification. It replaces the single treatment dummy with a vector of dummies for each period relative to treatment, typically −k to +k, with t = −1 omitted as the baseline.
The static TWFE regression gives one average post-treatment effect. The event study gives a separate estimate for each horizon, revealing anticipation, dynamics, and fadeout, plus pre-period estimates that probe parallel trends.
Always, alongside the static ATT. It is both the validity check and the fuller characterization of the effect trajectory. A flat pre-period and a rising post-period is the canonical clean result.
STYLIZED EVENT STUDY PLOT
library(fixest)
# Event-study specification: replace the single treatment dummy
# with one dummy per relative-time period.
# i() builds the relative-time x treatment interactions;
# ref = -1 omits the period before treatment as the baseline.
fit_es <- feols(
outcome ~ i(rel_time_binned, treated, ref = -1) | unit_id + year,
data = panel,
cluster = ~unit_id
)
# Coefficient plot: the event-study figure itself
iplot(
fit_es,
main = "Event study",
xlab = "Periods relative to treatment",
pt.join = TRUE,
ci.lwd = 1.5
)
abline(h = 0, lty = 2, col = "grey60") # zero line
abline(v = -0.5, lty = 3, col = "grey60") # treatment onsetChoosing the baseline
The omitted period, usually t = −1, must be a clean pre-treatment period. If effects begin before the nominal start date (announcement effects), the baseline is contaminated. Re-estimate omitting t = −2 or t = −3; large shifts mean the choice is driving your results.
Binning the endpoints
Far-out relative-time bins get sparse when adoption timing varies widely. Cap the ends (e.g. at ±4) so a handful of units cannot drive the tails. Collapse any bin holding fewer than roughly 20 units.
Pre-treatment coefficients from the event study should sit near zero. Divergence suggests parallel trends fails before treatment begins. Note that a non-rejection is weak evidence, these tests have low power with few pre-periods, so sensitivity analysis matters more than the p-value.
Visual inspection
Plot every pre-period coefficient with its 95% CI. A clearly sloping pre-period is a red flag regardless of p-values, the eye test often decides the design's credibility.
Pre-trend F-test
Joint significance test on all pre-period coefficients. Rejection suggests pre-existing divergence, not just noise. Evaluate jointly rather than period by period.
Sensitivity to bounded violations
The HonestDiD package tests how robust the ATT is to bounded violations of parallel trends. Specify the maximum allowable deviation M and report sensitivity across values of M.
Placebo treatment
Assign a fake treatment date in the pre-period, or randomly reassign treatment status, and re-estimate. Effects should be near zero, any pattern reveals pre-existing dynamics.
Alternate control group
Re-estimate using a different, arguably comparable, control group. Estimates should be similar if parallel trends holds broadly.
Enough pre-periods
Two is the minimum, three or four is much stronger. With a single pre-period you cannot distinguish a true parallel trend from a brief convergence.
# Joint F-test on the pre-period coefficients
# Null: all pre-period effects are zero (no detectable pre-trend)
pre_coefs <- names(coef(fit_es)) |>
(\(x) x[grepl("rel_time.*:-[1-9]", x)])()
wald(fit_es, pre_coefs)
# Sensitivity to bounded violations of parallel trends
# install.packages("HonestDiD")
library(HonestDiD)
honest <- createSensitivityResults(
betahat = coef(fit_es),
sigma = vcov(fit_es),
numPrePeriods = 3,
numPostPeriods = 4,
Mvec = seq(0, 0.05, by = 0.01) # allowable trend deviation
)
createSensitivityPlot(honest)When units adopt treatment at different times, TWFE conflates ATTs across cohorts and periods, and can produce sign-reversed estimates when treatment effects are heterogeneous.[2] This applies to the event-study specification above just as it does to the static ATT, so with staggered timing both should be re-estimated with a cohort-robust method: either the Callaway–Sant'Anna dynamic aggregation[1] or the Sun–Abraham interaction-weighted estimator.[4]
What changes
Instead of one β, you get ATT(g,t): the average effect for cohort g (first treated in year g) at calendar time t.[1]
Aggregation options
Aggregate to a single ATT, a dynamic event-study plot, or group-specific effects, all from the same underlying estimates. The dynamic aggregation is the cohort-robust event study.[1]
Control group
Use never-treated units if available. If not, not-yet-treated units can serve as controls with additional assumptions.[1]
Why not TWFE
Goodman-Bacon (2021) decomposition shows that TWFE weights cohort-time ATTs negatively when effect sizes vary across cohorts.[2]
Sun & Abraham (2021)
An interaction-weighted estimator available directly in fixest and pyfixest via sunab(). Produces cohort-robust event-study plots without leaving the TWFE workflow, so it is the faster option to add to an existing specification.[4]
library(did)
# Callaway & Sant'Anna (2021)
# Robust to heterogeneous treatment effects across cohorts
cs <- att_gt(
yname = "outcome",
tname = "year",
idname = "unit_id",
gname = "first_treat_year", # 0 if never treated
control_group = "nevertreated",
data = panel
)
# Aggregate to a single overall ATT
aggte(cs, type = "simple")
# Or aggregate dynamically: this IS the cohort-robust event study
es_agg <- aggte(cs, type = "dynamic", min_e = -4, max_e = 4)
ggdid(es_agg, title = "CS event study: dynamic ATT")
# Sun & Abraham (2021) alternative, inside the familiar fixest workflow
fit_sa <- feols(
outcome ~ sunab(first_treat_year, year) | unit_id + year,
data = panel,
cluster = ~unit_id
)
iplot(fit_sa, main = "Sun & Abraham event study")What does β = 0.04 mean?
A 4 percentage point increase in the outcome for treated units relative to controls, after accounting for unit and time fixed effects. Interpret in the units of your outcome variable.
My pre-period coefficients are non-zero, now what?
First check magnitude, not just significance. Small deviations with wide CIs may be noise. Large deviations suggest the parallel trends assumption fails, consider a different control group, covariates-adjusted DiD, or a different research design. Where the treated unit is a single aggregate, a state or a hospital system, synthetic control replaces the parallel trends assumption with a weighted donor pool matched on the pre-treatment path.
One pre-period coefficient is significant but the rest are flat, should I worry?
One marginally significant pre-period coefficient out of several is often noise. Evaluate them jointly rather than individually, the joint F-test matters more than any single period. Check magnitude too: a statistically significant pre-trend can still be substantively negligible.
My post-period effects are increasing, is the treatment strengthening?
Possibly, but rule out compositional change first. Rising effects can reflect selection if the composition of the treated group shifts as marginal units are added over time. Also check for an Ashenfelter dip in the pre-period, which can mechanically inflate post-period estimates.
TWFE ATT differs from CS ATT: which do I report?
Report both and explain the difference. If treatment effects are homogeneous across cohorts, they should be close. Divergence is itself informative, it signals treatment effect heterogeneity across adoption cohorts.[1][2]
Should I use clustered standard errors?
Yes, cluster at the level of treatment assignment (typically the state or firm). With fewer than ~30 clusters, consider wild cluster bootstrap or aggregation to the cluster level before estimation.
Difference-in-differences with multiple time periods
Callaway, B. & Sant'Anna, P. H. C. (2021)
Journal of Econometrics, 225(2), 200–230
Difference-in-differences with variation in treatment timing
Goodman-Bacon, A. (2021)
Journal of Econometrics, 225(2), 254–277
Causal Inference: The Mixtape
Cunningham, S. (2021)
Yale University Press · free online edition, updated as The Remix
Estimating dynamic treatment effects in event studies with heterogeneous treatment effects
Sun, L. & Abraham, S. (2021)
Journal of Econometrics, 225(2), 175–199
Counterfactuals and Causal Inference: Methods and Principles for Social Research
Morgan, S. L. & Winship, C. (2014)
Cambridge University Press, 2nd edition
Built as a learning tool and working reference for developing knowledge and coding practices across causal inference and causal ML. The author is actively building expertise, and therefore this should not be taken as an authoritative source. Drafted with assistance from Claude (Anthropic) for prose, code scaffolding, and reference checking; method selection, technical claims, and every citation above were reviewed by the author.