Difference-in-differences
SkillMediaDesign, estimate, validate, and write up a difference-in-differences analysis, using the post-2018 heterogeneity-robust toolkit with an explicit statement of which parallel-trends assumption is imposed. TRIGGER on "difference-in-differences", "DiD", "TWFE", "event study", "staggered adoption", "parallel trends", "pre-trends", "Callaway-Sant'Anna", "HonestDiD", "triple differences", "Poisson DiD", "nonlinear DiD", or any panel or repeated-cross-section setting where units become treated over time (policy rollout, staggered feature launch, state law changes). One or a few treated aggregate units: synthetic-control. No design chosen yet: causal-design.
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the Difference-in-differences skill
What this skill tells your AI
The instructions your AI receives, as published by ericluo04/claude-academic-workflow in skills/did/SKILL.md and read by ahel’s review.
An opinionated DiD workflow grounded in a read canon (references/canon.md, current as of 2026-08-26). Deliverable: the recommendation with its citation, the R estimation and diagnostics code, and a methods paragraph. Where the literature is unsettled the skill names a default and the condition that moves you off it, and where the canon genuinely disagrees it says so instead of faking consensus.
Refresh path: run litreview on "difference-in-differences" since the canon date, then propose additions to references/canon.md as flagged addenda.
Exemplar designs
Seven shapes worth recognizing on sight. Find the one your data looks like before writing a specification; what each one teaches is in references/details.md.
| Design shape | Canonical case | Marketing analogue | What kills it |
|---|---|---|---|
| Never-treated comparison with the full evidence battery | Miller, Johnson, and Wherry (2021) | a feature launched in some markets and withheld in others, with take-up data | a control group quietly exposed to the treatment |
| Staggered rollout across institutions, adoption dates rebuilt from an archive | Braghieri, Levy, and Makarin (2022) | a platform reaching accounts, regions, or partners on its own schedule | adoption dates that are wrong, or dated to enforcement when behavior moved at announcement |
| Forward engineering from estimand to estimator | Baker et al. (2026) Medicaid | any panel whose units differ enormously in size | reporting the weighted and the unweighted ATT as a robustness pair |
| Estimand first on a heavy-tailed outcome | Winkler et al. (2026) UMG-TikTok | streams, revenue, views, sessions | a log taken for convenience |
| Compositional change in repeated cross-sections | Hong (2013) Napster | brand trackers and refreshed survey panels | who is sampled moves with treatment timing |
| Triple differences | Gruber (1994) | an ineligible segment inside the same treated markets | reading the placebo DiD as a test that has to return zero |
| The idea, not the inference | Card and Krueger (1994) | a two-region holdout test | G = 2 with one treated cluster |
Triage: three questions before anything else
From Roth, Sant'Anna, Bilinski, and Poe (2023), the most condensed decision object in the literature:
- Is everyone treated at the same time? Yes: TWFE is fine, static or dynamic, without covariates; with covariates see Covariates below. For two groups and two periods the estimator is interaction OLS or the long difference, clustered at the level of assignment. The identity of interaction OLS, TWFE, and the long difference holds exactly only in that 2x2 with no covariates, and it is what makes people reach for TWFE-with-controls, where it no longer holds. Everything else in this skill still applies to a 2x2. No (staggered): default to a heterogeneity-robust estimator; TWFE only if you will defend effect homogeneity.
- Are you sure about parallel trends, and on which scale? Justify levels vs logs (PT is functional-form dependent and generally cannot hold in both) and name the estimand the scale targets (Functional form below). If PT is plausible only conditional on covariates, use regression adjustment, IPW, or doubly robust, never bare TWFE-with-controls. Always pair the event study with a Rambachan-Roth sensitivity analysis, a universal mandate this skill family hardens beyond the canon's best-practice advice because the analysis is cheap and the pretest is low-powered.
- Do you have many treated and untreated clusters? Yes: cluster at the level at which treatment is independently assigned. No: pick a few-clusters method by which homogeneity assumption you believe (references/details.md has the map), or fall back to a cluster-level Fisher randomization test. Before picking, run the MacKinnon, Nielsen, and Webb (2023) battery (CV1, CV3 jackknife, and a restricted wild cluster bootstrap side by side): agreement means stop, disagreement means escalate to the map.
Two exits from DiD entirely, both routed to the synthetic-control skill: only one or a handful of treated units, or selection on lagged outcomes with autocorrelated errors, where DiD is inconsistent even as pre-periods grow while SC is consistent (Arkhangelsky-Hirshberg, via the panel survey of Arkhangelsky and Imbens 2024). Synthetic DiD also lives there. Few treated aggregate UNITS with failed pretrends exit to synthetic-control; few treated CLUSTERS of micro units with plausible parallel trends stay here on the few-clusters inference map (references/details.md). When parallel trends fails and the synthetic-control exit is also infeasible (no credible donors, short pre-period), decline the design and route back to causal-design; a failed gate is a verdict. If treatment timing is quasi-random, DiD is valid but inefficient; use the efficient random-timing estimators (R package staggered, Roth-Sant'Anna).
Assumptions and treatment dating
Parallel trends is the assumption this skill argues about. Two others get stated and forgotten, and both fail silently.
No anticipation: the treated group's outcome is Y(0) until treatment happens. An already-treated baseline attenuates the estimate and nothing in the diagnostics battery flags it. In the Mixtape's simulation (Cunningham, Causal Inference: The Remix, ch. 9) a contaminated baseline returns 0.017 against a true ATT of 10 under constant effects and 15.017 against a true 20 under dynamic effects. Make the baseline clean. When announcement precedes enforcement and behavior can respond (pre-announced price changes, regulatory effective dates published months ahead), date treatment to the announcement and state the two prices: the ATT now averages over a longer post window including announced-but-unenforced periods, and parallel trends is now stated against a different baseline.
SUTVA: no interference between units, and the treatment is the same thing for every treated unit. Panel threats are spillover to adjacent control markets, substitution across one firm's own units, and the announcement-versus-enforcement gap, which makes "announced" and "announced and enforced" two treatments. Interference contaminates the control group's Y(0), so the comparison understates the effect or reverses its sign. Buffer or drop adjacent controls, or model exposure and make the exposure the treatment.
Estimand before estimator
Weights define the target parameter, not the specification (Baker et al. 2026). An unweighted ATT answers "effect on the average treated county"; a population-weighted ATT answers "effect on the average treated person." In the Baker et al. Medicaid 2x2 these are +0.1 and -2.6 deaths per 100,000: different questions, not a robustness check of each other. In marketing panels where units differ enormously in size (DMAs, stores, channels), decide by the business or policy question and, if you report both, report them as different estimands.
Write the target in potential-outcomes notation before touching code: which ATT(g,t) cells, and which aggregation (event-time, calendar-time, overall; cohort-share or population weights). Callaway-Sant'Anna aggregates the same building blocks four ways: simple, one average over all feasible group-times; group, ATT(g), the effect on each cohort; calendar, ATT(t), the effect in each period, which is the target when the question is about a season, a platform change, or a macro shock; and dynamic, ATT(l), the event study. Those are four parameters, not four robustness checks, and the feasible (g,t) set shrinks at long horizons, so each covers a different slice of the data.
Functional form and weights are estimand choices too (Winkler, Hotz-Behofsits, Wlömert, Papies, and Liaukonytė 2026). Three parameters get reported as if they were one: the typical-unit proportional effect ΔΔE[log Y] (log OLS), the population-total proportional effect ΔΔ log E[Y] (PPML), and the level effect ΔΔE[Y] (levels OLS). Under heavy-tailed outcomes they differ in magnitude and can differ in sign with no staggered-timing problem anywhere: in the UMG-TikTok withdrawal, a clean two-group single-date design with 53,753 matched song pairs, unweighted log OLS gives +0.0063 and PPML gives -0.0310 on the same panel, and reweighting the log OLS toward the head gives -0.0286 without touching the transformation. Pick the estimand from the question ("did per-user usage rise 5%?" is typical-unit, "did total revenue rise 5%?" is population-total, "did this add $2M?" is level), then the estimator. Levels-vs-logs is never a robustness check: an appendix that reports both without naming the estimand each targets is reporting two parameters as one. The heterogeneity-robust estimator question and the estimand-scale question are orthogonal, and both get answered.
The parallel-trends menu and the estimator it implies
State explicitly which PT assumption you impose (Baker et al. 2026 make this a requirement). Three variants under staggered adoption, with the estimator crosswalk:
| PT variant | Comparison group | Pre-trends restricted? | Estimators | R |
|---|---|---|---|---|
| PT-Nev | never-treated | no | Callaway-Sant'Anna (never), Sun-Abraham | did::att_gt(control_group="nevertreated"), fixest::sunab() |
| PT-NYT | not-yet-treated | no | Callaway-Sant'Anna (NYT), dCDH instantaneous | did::att_gt(control_group="notyettreated") |
| PT-all | all groups, all periods | yes (testable, and baked in) | BJS/Gardner/LWX imputation, Wooldridge ETWFE | didimputation, did2s, etwfe |
Sun-Abraham sits under PT-Nev for a reason worth naming at the point of use: its cohort 2x2s use
the last-treated cohort or the never-treated, never the not-yet-treated. Against a CS-NYT default
the sunab cross-check in the template therefore compares two assumptions, and neither agreement
nor divergence is a robustness result.
Default: Callaway-Sant'Anna with not-yet-treated controls, Baker et al.'s own choice for their application. That default is licensed by no anticipation among the not-yet-treated: if their behavior already responds to the treatment they are about to get, they are not clean controls. Move to never-treated when the not-yet-treateds' timing plausibly responded to recent outcomes; move to imputation (PT-all) when you will defend parallel pre-trends over the whole panel and errors are near-serially-uncorrelated, where it buys real efficiency (Roth et al. 2023). A long covariate list is a further reason to move to imputation or Wooldridge ETWFE, since the CS propensity score is estimated per cohort and with fifty covariates common support gets hard to assess and harder to defend. The price is that imputation event-study plots are not comparable to CS or TWFE plots, because fitting the counterfactual on the whole pre-period puts a mechanical kink at t = -1, so do not overlay them. If Y(0) is close to a random walk, the CS last-pre-period baseline is the efficient choice and imputation's pre-period averaging buys nothing (Harmon's caveat: averaging is not guaranteed more precise). If PT is implausible for one specific cohort, drop that cohort rather than average over it.
"Never-treated" operationally means "not treated by the end of the sample." If all units are eventually treated, drop periods from when the last cohort adopts and use that cohort as the comparison, and drop units treated in the first period.
The TWFE question, stated honestly
The same regression is two designs. With an adoption date and units that are never or not-yet
treated, y ~ treat | unit + period is a DiD and the assumption is parallel trends. With a
treatment that varies within unit over time and no comparison group at all, the identical
regression is the within estimator and the assumption is strict exogeneity conditional on the
unit effect, non-nested with parallel trends and failing differently (feedback from past
outcomes to current treatment). The Mixtape (Cunningham, Causal Inference: The Remix, ch. 8)
teaches that they are the same thing, true as algebra and false as design. Route the second case
to causal-design's plain-panel-fixed-effects section; nothing below applies to it.
For the first case the canon disagrees, and the skill's position is a default with named dissent:
- Baker et al. (2026): TWFE under staggered adoption has "well-understood, potentially serious, and easily remedied problems, and we do not recommend using it."
- Arkhangelsky and Imbens (2024): "we recommend against the current routine use of the standard TWFE estimator or related estimators," though under block assignment TWFE still estimates the ATT, and they think negative-weight concerns "have perhaps been exaggerated."
- Abadie, Angrist, Frandsen, and Pischke (2025): the pathologies "are unlikely to derail DD or event-study designs in practice"; in the divorce data BJS and TWFE match. Their prescription is TWFE event studies with BJS as a check.
Default here: estimate the robust estimator as the headline number and report TWFE alongside it. Agreement is affirmative evidence the simple model suffices (AAFP); divergence means heterogeneity is doing real work and the robust estimate stands. The only way to know TWFE would have been fine is to run the robust estimator anyway, at which point you report it (Baker et al.). Run the building-block heterogeneity scan (2x2 DiDs by cohort, gap, and time since adoption, Arkhangelsky-Imbens) rather than trusting any single robust estimator blindly.
The weights do something the negative-weight framing hides. Goodman-Bacon weights combine a sample share and the treatment-timing variance Dbar(1 - Dbar), which peaks at 0.25 for a cohort treated at the panel midpoint, so TWFE upweights mid-panel cohorts and extending or truncating the panel moves the estimate through the weights alone. Under constant effects TWFE is unbiased for the variance-weighted ATT and not for the simple ATT. That also reconciles bacondecomp's all-positive weights, which sit on 2x2 comparisons, with TwoWayFEWeights' negative ones, which sit on unit-level treatment effects: both are right, and all-positive Bacon weights do not license TWFE.
The Mixtape (ch. 10) rejects stacking estimators on one event-study plot: "These estimators all have slightly different assumptions and should not be considered robustness checks for one another." Two practices are in play. Reporting TWFE beside one robust estimator is asymmetric, because TWFE is the estimator whose bias the exercise bounds, so agreement or divergence is information about heterogeneity. Stacking four heterogeneity-robust estimators is symmetric, they impose different PT variants and use different comparison groups, and the objection lands there: choose ex ante by design, and if you genuinely cannot, pre-commit to reporting all.
Event-study mechanics
Hard rules, mostly from Abadie, Angrist, Frandsen, and Pischke (2025):
- Feasible horizons: with panel end T and cohorts c(s), longest lag q = T - min c(s), longest lead m = max c(s) - 1. Here s indexes treatment cohorts and c(s) is cohort s's adoption period. The formulas assume periods renumbered 1..T, so calendar years must be reindexed first.
- Always omit event time -1. The Mixtape (ch. 9) gives the reason, that no anticipation requires an untreated baseline, and stops there; with never-treated units that single normalization identifies everything, and without them a second lead or lag must also be omitted, the linear component of the effect path is unidentified, and different second choices rotate the whole path around -1. Choose deliberately, never let the software's default drop decide, and show the path under at least two normalizations before interpreting dynamics.
- Short gaps versus long differences, a hard rule. OLS event studies mechanically use a
universal baseline, so every coefficient is a long difference against t = -1.
Callaway-Sant'Anna and dCDH can produce either, and a rolling baseline gives short gaps, which
estimate a different quantity (Roth 2026). Stata's
csdiddefaults to short gaps and needslong2,csdid2defaults to long differences, R'sdidneedsbase_period = "universal". Reader-side tell in someone else's paper: a confidence interval at t = -1 means a rolling baseline, since t = -1 cannot be its own baseline, and the leads top out one period earlier. - Reading the coefficients out loud. A lead of +1.5 means the treated group's change from that period to the omitted baseline ran 1.5 outcome units above the control group's. A lag is the ATT for that period, valid only under PT from the baseline to it, no anticipation, and an untreated comparison group. Plot disconnected points with whiskers: connecting the bands makes the intervals appear to narrow toward the omitted period, where nothing was estimated.
- Bin leads and lags beyond the horizon where only a few units identify the coefficient (AAFP use +/-15 with 41 periods). Clustered SEs fail at long horizons through leverage: a lone late adopter can identify the longest leads, and coverage collapses.
- Report simultaneous sup-t uniform bands, computed separately for leads and lags, not only pointwise bands (recipe in references/details.md; the did package produces them by multiplier bootstrap). Never use a clustered joint F over many leads.
- Under staggered timing, build the event study from a robust estimator, never from dynamic TWFE: cross-lag contamination means TWFE lead coefficients can be nonzero under valid PT and zero under violations (Sun-Abraham, via Roth et al. 2023).
- Check composition: cohorts entering and leaving event times can manufacture dynamics. Use the balanced-in-event-time aggregation or a fixed cohort set when in doubt.
Pre-trends and honest sensitivity
Pre-trend tests are underpowered: in simulations calibrated to three top journals, linear violations detected only 50% of the time produce bias as large as the estimated effect and a spurious significant effect about half the time (Roth 2022). Quote from Roth et al. (2023): "the lack of a significant pre-trend does not necessarily imply the validity of the parallel trends assumption." Also the converse (Kahn-Lang and Lang's bar mitzvah example): parallel pre-trends do not imply parallel post-trends.
The canon splits on pretesting itself. Roth (2022) shows conditioning on passing adds selection bias; AAFP's cost-benefit analysis concludes "the bias-mitigation benefits of pretesting are likely to outweigh the risks" and prescribes a sup-t joint test of the leads. The Mixtape (ch. 10) is a third position, nearer Roth than AAFP: "event studies were always only falsifications. They weren't true tests", parallel trends is untestable, and honest DiD is not a test of it either. Default here: run the sup-t pretest and report it, but never let a pass substitute for the sensitivity analysis, and report the test's power against economically relevant violations (R package pretrends). That pretest is not a claim that parallel trends is testable. It is a claim that a joint test with a reported power curve is a better-calibrated version of the eyeball heuristic everyone runs anyway.
Mandatory companion to every event study: Rambachan-Roth honest inference (HonestDiD), in one or both of its restrictions, a universality that is this family's own hardening of the canon's best-practice endorsement (cheap to run, against a low-powered pretest). Relative magnitudes bounds post-treatment violations by M-bar times the largest pre-treatment violation; use it when the worry is shocks like those already seen pre-treatment. Smoothness bounds deviations from a linear extrapolation of the pre-trend by M; use it when the worry is a smoothly evolving confound. Report the identified set, the robust CI, and the breakdown value at which the conclusion dies, then read it economically: robustness to M-bar = 2 is strong in a calm period and weak if treatment coincided with a shock larger than anything pre-treatment. Worked template numbers (Baker et al.): largest one-period pre-trend 4, identified set -2.6 +/- 4 = [-6.6, 1.4], robust CI [-11.1, 5.1].
Why these units were treated
Answer that before reading the pre-trend picture. The assignment mechanism is learned outside the dataset, from institutional detail, and it decides how the picture should be read (Ghanem, Sant'Anna, and Wüthrich; Marx, Tamer, and Tang 2024, via the Mixtape ch. 9). Five mechanisms are compatible with parallel trends: a common constant trend in Y(0), under which PT cannot be violated however units were selected; selection on baseline Y(0); selection on fixed effects; selection on observables, which is conditional PT and sends you to Covariates below; and selection under imperfect foresight. One breaks PT: selection on realized gains, where units take the treatment because they correctly infer it will help them. The table in references/details.md separates what each does to pre-trends from what it does to PT.
Selection on baseline Y(0) is the case that misleads. Enrolling everyone below a threshold on baseline Y does not violate PT and does break pre-trends mechanically, because the baseline is both the selection point and the omitted category, which manufactures a dip for the treated group at t = -1. There is no fix because there is no problem, and re-basing to t = -2 to make the picture look right breaks PT, which held from the original baseline.
That collides with the HonestDiD mandate, and the resolution here is the skill's own judgment. Relative magnitudes bounds post-treatment violations by M-bar times the largest pre-treatment violation, so under a documented selection-on-baseline-outcome mechanism the anchor is inflated by an artifact of the assignment rule and the skill would otherwise drive you to call a valid design fragile. Prefer the smoothness restriction there, or compute the anchor from the pre-treatment periods excluding the selection period, and say which you did. The reverse case is worse: under selection on realized gains, clean pre-trends are no comfort at all.
Covariates
Never bare TWFE-with-controls: even in the 2x2 it identifies the ATT only under constant effects across covariate strata, weights strata non-convexly, and adds three misspecification bias terms (Caetano-Callaway, via Baker et al. 2026). Choose covariates from theory (determinants of untreated trends or of selection), and ask domain experts what ordinarily drives untreated trends in this outcome, since that is the arm of the DAG nobody guesses well. If you select covariates from the data, run the selection on untreated units only, so only Y(0) informs it (Borgschulte and Vogler 2020). State the covariate list before you see results: the choice can swing the estimate and a DAG does not protect against specification search. Check the covariates are unaffected by treatment (a time-varying covariate is fine only if its whole path is unaffected), then use:
- Doubly robust (default):
did::att_gt(est_method="dr"), Sant'Anna-Zhao. Consistent if either the outcome-change model or the propensity model is right. - Regression adjustment when overlap is weak (extrapolates the outcome model; say so).
- IPW when you understand selection better than outcome dynamics; noisy as control propensity scores approach 1. Histograms and kernel densities do not reveal explosive weights, so count the control units with a propensity score near 1 directly. In the Mixtape's CAPS data (ch. 10) one control municipality scores 0.999971, which is a weight of p/(1-p) = 34,481, and eleven control units sit above 0.995. Trim at 0.995 and keep the trim. R's did trims automatically and not every package does, so check what yours does before trusting the estimate.
Shortened here. Read the whole file on GitHub.
Signals
- GitHub stars
- 23
- Forks
- 3
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Gateway key
did- Source
- github.com/ericluo04/claude-academic-workflow