Data Provenance

SkillDatabases & data

Track dataset lineage, transformation steps, merge logic, and reproducibility risks in Stata workflows. Use when the user needs to explain where data came from, how it changed, or why a pipeline can be trusted.

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the Data Provenance skill

What this skill tells your AI

The instructions your AI receives, as published by brycewang-stanford/auto-empirical-research-skills in skills/64-tmonk-mcp-stata/skills/stata-data-provenance/SKILL.md and read by ahel’s review.

Use this skill when lineage and reproducibility matter.

  1. Map the sequence of source files and transformations.
  2. Flag untracked merges, overwrites, and silent sample restrictions.
  3. Produce a concise provenance narrative a coauthor can audit.

Read references/lineage.md for the provenance checklist.

Signals

GitHub stars
4k
Forks
531
Last commit
Sep 2026
Advanced
Item type
skill
Key
stata-data-provenance
Source
github.com/brycewang-stanford/auto-empirical-research-skills