Identification Strategy
SkillMediaDesign or review identification strategy for the sewage-house-prices project. Produces strategy memos with estimand, assumptions, pseudo-code, robustness plan, falsification tests, and referee objection anticipation. This skill should be used when asked to "design the strategy", "identify the effect", "write a strategy memo", or "think through identification".
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; ahel provides instructions and does not run this skill.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
Then ask your AI: use the Identification Strategy skill
What this skill tells your AI
The instructions your AI receives, as published by brycewang-stanford/auto-empirical-research-skills in skills/41-sticerd-eee-sewage-econometrics-check/skills/identify/SKILL.md and read by ahel’s review.
Design or review an identification strategy for the sewage-house-prices project.
Input: $ARGUMENTS — a research question, approach name (e.g. "hedonic", "dry spills"), or "review existing" to audit all current strategies.
Project-Specific Context
Existing Strategies
- Hedonic — Cross-sectional:
log(price) ~ spill_metrics + controls | lsoa + year_quarter. Assumption: spill exposure is conditionally exogenous given LSOA FE. - Repeat sales — Within-property:
Δlog(price) ~ Δspill_metrics | house_id. Eliminates time-invariant unobservables. - Long difference — Grid-level: changes in average prices within 250m grids. Eliminates level differences.
- News/media DiD — Treatment = post-media-coverage × exposure. Tests whether information matters for capitalisation.
- Upstream/downstream — River network topology via PostGIS. Downstream sites receive upstream pollution. Tests directionality.
- Dry spills — Spills without rainfall. If dry spills affect prices, suggests awareness/stigma channel over physical damage.
- Hydraulic capacity instrument — Planned IV using sewer capacity as instrument for spill frequency.
Key Data Features
- EDM data: 2021-2024+, high-frequency (event-level)
- Land Registry: universe of transactions
- Zoopla: rental listings
- Met Office: daily rainfall at LSOA level
- River networks: PostGIS topology
- Treatment radii: 250m, 500m, 1000m, 2000m, 5000m, 10000m
Workflow
Step 1: Context Gathering
- Read existing manuscript sections in
docs/overleaf/for how strategies are currently described - Read relevant analysis scripts in
scripts/R/09_analysis/ - Read
scripts/R/utils/spill_aggregation_utils.Rfor treatment construction - Check
docs/overleaf/refs.bibfor methodological references
Step 2: Strategy Development
For a new or revised strategy, produce:
- Strategy memo — Design choice, estimand (ATT/ATE/LATE), key assumptions, comparison group
- Estimating equation — LaTeX-formatted with clear variable definitions
- Pseudo-code — Implementation sketch (what the R code will do)
- Robustness plan — Ordered list with rationale:
- Radius sensitivity (250m → 10km)
- Time period variation (prior period vs full period)
- Alternative treatment measures (count vs hours vs binary)
- Subsample analysis (sales vs rentals, urban vs rural)
- Falsification tests — What SHOULD NOT show effects and why
- Referee objection anticipation — Top 5 objections with pre-emptive responses
Step 3: Strategy Review
If reviewing an existing strategy:
Phase 1: Claim Identification
- What is the claimed design?
- What is the estimand?
- What is the treatment / comparison?
Phase 2: Core Design Validity
- Are identifying assumptions stated and defensible?
- Are the biggest threats acknowledged?
- Does the specification match the stated design?
Phase 3: Robustness Assessment
- Does the robustness plan address the right concerns?
- Are falsification tests well-chosen?
- Is there radius sensitivity analysis?
Step 4: Present Results
# Identification Strategy: [Approach]
**Date:** YYYY-MM-DD
**Design:** [Hedonic / Repeat Sales / Long Diff / DiD / IV / etc.]
**Estimand:** [ATT / ATE / LATE]
## Strategy Summary
[2-3 sentence description]
## Estimating Equation
$$\log(p_{it}) = \alpha + \beta \cdot \text{SpillMetric}_{it} + \gamma X_{it} + \mu_i + \delta_t + \varepsilon_{it}$$
## Key Assumptions
1. [Assumption 1] — [defense]
2. [Assumption 2] — [defense]
## Assessment: [SOUND / CONCERNS / CRITICAL ISSUES]
## Robustness Plan (ordered)
1. [Most important check]
2. [Second check]
...
## Falsification Tests
1. [Test 1] — [expected null and why]
## Anticipated Referee Objections
1. [Objection] — [Response]
## Next Steps
- [ ] Implement main specification
- [ ] Run falsification tests
- [ ] Generate pre-trend evidence
Save to output/log/strategy_memo_[approach].md.
Principles
- Catch problems before coding. A flawed strategy caught now saves weeks of wasted analysis.
- Multiple strategies are strength. This paper uses 6+ approaches — consistency across them is the key argument.
- Cross-reference approaches. Each strategy should address threats the others cannot.
- The user decides. Present trade-offs, don't make choices unilaterally.
- Strategy memo is the contract. Once approved, analysis scripts implement it faithfully.
Signals
- GitHub stars
- 4k
- Forks
- 531
- Last commit
- Sep 2026
Advanced
- Item type
- skill
- Key
identify- Source
- github.com/brycewang-stanford/auto-empirical-research-skills
github.com/brycewang-stanford/auto-empirical-research-skills
Related picks
Skill · agricidaniel
The pick for Markdownmarkdown-formatter
Skill · nvidia
The pick for Markdownlatex-posters
Skill · k-dense-ai
The pick for LaTeXlatex-drawing-guide
Skill · brycewang-stanford
The pick for LaTeXsupabase-postgres-best-practices
Skill · supabase
The pick for Postgresprototype
Skill · mattpocock
More in Media