Diagnosing experiment health
SkillDev toolsRuns diagnostics and a health check on one PostHog experiment: its setup, exposures, results and changes during the run. Covers 0 exposures, sample ratio mismatch, lost or uneven exposures, users in multiple variants, identity faults, who the experiment counts, significance traps (peeking, A/A, 'target reached', Bayesian vs Frequentist, sequential, CUPED), mid-run edits, pause and freeze, and survey follow-up.\nTRIGGER when: user asks 'is my experiment healthy / set up right / biased?' or 'why 0 exposures?', asks why it counts different people than their insight, mentions the bias or mismatch warning, an uneven split, significance that flips, numbers that disagree with their SQL, results that changed after an edit, or wants user feedback on an experiment.\nDO NOT TRIGGER when: creating an experiment (use creating-experiments), only configuring rollout (use configuring-experiment-rollout) or metrics (use configuring-experiment-analytics), or only asking lifecycle questions (use managing-experiment-lifecycle).
Use Diagnosing experiment health in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add Diagnosing experiment health and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the Diagnosing experiment health skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; ahel provides instructions and does not run this skill.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
What this skill tells your AI
The instructions your AI receives, as published by posthog/posthog-foss in products/experiments/skills/diagnosing-experiment-health/SKILL.md and read by ahel’s review.
This skill answers: My PostHog experiment results look wrong, biased, or empty — what's going on? It is the diagnostics and health check for one experiment: its setup, its exposures, its results and what changed during the run.
Match the user's complaint in the dispatch table, then read the matching reference file for the diagnostic.
Each diagnostic in the reference files is tagged [HIGH], [MEDIUM], or [LOW] based on how
strongly it's verified — [HIGH] is verified directly in PostHog code, [MEDIUM] is partially or
team-source verified, [LOW] describes SDK/external behavior that wasn't verified here. Treat [LOW]
items as hypotheses to test, not facts to assert.
Step 1 — Resolve the experiment
If the user refers to an experiment by name or description, load the finding-experiments skill first to
resolve it to a concrete ID.
Call experiment-get and pull these fields. They are inputs for almost every diagnostic:
status(draft/running/paused/exposure_frozen/stopped),start_date,end_date,is_legacyfeature_flag.active, andfeature_flag.filters.multivariate.variants[]— the variant keys and the split (rollout_percentage). This is the flag as it is now. After a ship or an edit, the split of the run is in the change history (Step 1.5).feature_flag.filters.aggregation_group_type_index— when set, the experiment's unit is a group, not a personfeature_flag.filters.groups[]— for each group readvariant,properties, androllout_percentage(that group's rollout, % of the matched bucketing units — persons, or groups when the flag or the group is aggregated by a group type — that enter the experiment; each group has its own, so there is no single overall rollout unless the flag has one unconditional group). A non-nullvariantthat is one of the flag's variant keys is a forced-variant override on the matched cohort (release-condition assignment, not randomized) — surfaces A7. Watch for the severe shape (A7b): a variant-pinned group with broad/emptypropertiesat high rollout, or no group left randomized (variant: null) / no release path to one arm — that starves the other variant (one arm gets ~0 analyzable exposures). Seereferences/bias-and-skew.md.feature_flag.bucketing_identifierandfeature_flag.ensure_experience_continuity— what the flag hashes (A3, A6, A8, A10)holdoutandexcluded_variants— people and arms that are outside the analysis on purposeexposure_criteria.multiple_variant_handling— defaults to"exclude"if absentexposure_criteria.exposure_config— the exposure event and its property filters. Unset, or naming$feature_flag_called, means the default exposure event: read which one fromresolved_exposure_event. A config that names$experiment_exposurecounts that event, whateverresolved_exposure_eventsays. Any other event, or an action, is a custom exposure event. Property filters in the config apply to either (B12). The table inreferences/diagnostic-snapshot.mdsays which property carries the variant.exposure_criteria.activation_config— an activation event, or an action, on top of the default exposure. When it is set, a person counts as exposed only from their first activation event at or after their first flag exposure, and the time of that event is the exposure time. It does not combine with a custom exposure event.exposure_criteria.filterTestAccounts— defaults totruemetricsandmetrics_secondary, with shared metrics — for each: the name, the type and the definition. Describe a metric from its definition, not from its name (D13).stats_config—method(Bayesian or Frequentist), the confidence level (bayesian.ci_level,frequentist.alpha),frequentist.sequential_testing_enabled,cuped,baseline_variant_key. A missingmethodreads as Bayesian and a missing confidence level as 95%. Sequential testing and CUPED fall back to the project default (C13).running_time_calculation— the target sample (a total across all variants), the days still missing to reach it, and the minimum detectable effect. When the estimate comes from numbers the user typed, or was written at creation and never refreshed by the page, the days are the whole estimated length (C12). A duplicated experiment starts with the values of its source until the page refreshes them.only_count_matured_users— E11
Step 1.5 — Pull a diagnostic snapshot (verify before asking)
Before asking the user clarifying questions, pull the diagnostic snapshot in references/diagnostic-snapshot.md. Most diagnostics in this skill can be confirmed or ruled out from that data without an interview.
Read the cheap sources first, in every diagnosis: the stored results, then the change history. When the user reports a surprising change, read the change history before the stored results. Then the exposure data, when the diagnosis needs what only it holds. Query raw events only for what those cannot show.
Step 2 — Match symptom to diagnostic
| User says... | Diagnostic group |
|---|---|
| "Smaller variant looks biased" / banner says bias | A — bias & skew |
| "Variant ratio doesn't match my split" / mismatch warning / "is 58/42 normal?" | A — bias & skew |
| "Why isn't it 50/50?" / "one variant gets fewer users every day" | A — bias & skew |
"Users in both control and test" / high $multiple % | A — bias & skew |
| "Our own analytics show an even split and PostHog doesn't" | A — bias & skew |
| "Most responses on the flag are false" / "users saw the new UI for a second, then lost it" | A — bias & skew |
| Multi-variant exposure on a server-rendered app | A — bias & skew |
| Banner about feature-flag/experiment state mismatch | A — bias & skew |
| "Migrating distinct_id" / "switching from anonymous to user_id" mid-run | A — bias & skew |
| Metric count is much smaller than exposures (e.g. 10× or 100× gap) | A — bias & skew (route here before D) |
| "Experiment shows 0 / not enough data" / empty | B — empty experiment |
| "Variant always undefined / false" | B — empty experiment |
| "$feature_flag_called fires but no exposures show up" | B — empty experiment |
| "Exposures dropped to zero after I added a filter" / "the copied experiment is empty" | B — empty experiment |
| "Plenty of exposures, zero conversions" | B — empty experiment (then D11) |
| "Experiment says running but exposures haven't moved in weeks/months" | B — empty experiment |
| "Significance keeps flipping as we run longer" | C — interpretation traps |
| "Significance was declared, then it wasn't significant anymore" | C — interpretation traps |
| "It's been running one day, can we read it?" | C — interpretation traps |
| "30/16 split at 46 exposures, is this broken?" | C — interpretation traps |
| "It says complete / target reached and nothing is significant" | C — interpretation traps |
| "A/A test is showing significant results" | C — interpretation traps |
| "Many metrics — some significant, some not" | C — interpretation traps |
| "The test variant does worse, is it set up correctly?" / "can you check this experiment?" | Step 3 — the two checks, then C and D12 |
| "Bayesian says 96% chance to win — should we ship?" | C — interpretation traps |
| "Confidence intervals overlap — does that mean not significant?" | C — interpretation traps |
| "The p-value is exactly 1" / "control shows a lift" | C — interpretation traps |
| "An external tool (significance calculator or AI agent) disagrees with PostHog" | C — interpretation traps |
| "Should I ship? Primary is up but a secondary is down" | C — interpretation traps |
| "PostHog numbers ≠ my SQL count" | D — numbers vs SQL |
| "Funnel says X% but my raw event count says Y" | D — numbers vs SQL |
| "The experiment says flat and our dashboard shows a drop" | D — numbers vs SQL |
| "Sum of revenue looks wrong" / "breakdown shows 'none'" | D — numbers vs SQL |
| "The metric's name says one thing and the number says another" | D — numbers vs SQL |
| "Recordings don't match the stats" / "no recordings for exposed users" | D — numbers vs SQL |
| "I applied a filter but the user count didn't change" | D — numbers vs SQL |
| "I want to slice results by current person properties (as of now, not as of exposure)" | D — numbers vs SQL |
| "The results haven't updated" / "refresh changes nothing" | D — numbers vs SQL |
| "Changed split / rollout / metric / criteria mid-run, now odd" | E — mid-run changes |
| "The numbers changed overnight and nobody touched it" | E — mid-run changes, then D8 in D |
| "We paused it for a few days, can we still trust it?" | E — mid-run changes |
| "Froze the experiment" / "after the freeze users lost the variant" / "freeze was refused" | E — mid-run changes |
| "Ended/shipped — flag now flipped to 0/100 unexpectedly" | E — mid-run changes |
| "We shipped the variant and some users still don't get it" | E — mid-run changes |
| "Long-term metric moves opposite from primary" | E — mid-run changes |
| "Retention metric counts users I didn't expect" | E — mid-run changes |
| "Can't convert the feature flag back to a simple (boolean) flag after the experiment ends" | E — mid-run changes |
| "How do I restart an experiment with new variants?" | E — mid-run changes |
| "It says this is a legacy experiment" / metrics can't be edited | E — mid-run changes (E13 legacy experiments) |
"Results won't load" / many metric rows show data: null (not a legacy experiment) | Step 1.5 — diagnostic snapshot (null rows) |
| "What do users think of the new flow?" / wants qualitative feedback on an experiment | F — qualitative feedback |
| "Why did users prefer control?" / "what did they dislike about the test variant?" | F — qualitative feedback |
If the symptom is unclear, ask one clarifying question before picking. Most diagnostics have different fixes — do not guess. A complaint that matches rows of several groups gets each of them read (Step 3).
A report from a scout or another automated check names its finding: route by the finding, and reuse the report's numbers before pulling them again.
Step 3 — Surface every diagnostic the evidence supports
After matching the symptom in Step 2 and reading the relevant reference file(s), list each diagnostic that applies before recommending an action.
Surface co-occurring mechanisms independently — even when one is more salient, don't collapse them into a single "wait" or "fix" recommendation. Different mechanisms have different fixes: a systematic bias (e.g. uneven-split + Exclude) doesn't resolve by waiting; a statistical pattern (e.g. small-sample variance) does. Bundling them leaves the bias in place after the user follows the bundled advice.
Only list mechanisms that have a path to verification in the project state — config (from
experiment-get), snapshot data, activity log, or repo source (when the agent has the repository). Config-derived mechanisms count: an
80/20 split with default multiple_variant_handling="exclude" is visible in experiment-get and is
therefore enumerable. Naming a mechanism with no source (e.g. SRM when the snapshot shows a clean
variant ratio) is not.
Two checks belong to every diagnosis, whatever the complaint:
- How young is the experiment? State its age (
start_date) and its share of the target sample (the exposed total againstrunning_time_calculation.recommended_sample_size) before reading a result (C1 inreferences/interpretation.md). The page takes the exposed total from the first primary metric: the sum of itsnumber_of_samplesover all variants in the stored results, notexposures.total_exposures.recommended_sample_sizeis the target the page saved when it last computed the estimate. When it is absent, no target was saved: say so. - Who is exposed? Compare the exposed population (
exposure_criteriaand the flag's release conditions) with the people the decision is about before calling the setup sound (D12 inreferences/numbers-vs-sql.md). Default exposure settings are not evidence that the right people are in the experiment. When nothing says who the decision is about, ask.
Diagnostic groups
A — Bias & skew
Variants don't look balanced, one variant looks biased, the in-app warning banner appeared, or users are
showing up under multiple variants. Covers the uneven-split + Exclude interaction, the sample ratio mismatch and its causes, exposure events that are lost or sent unevenly between the arms, flags that return no variant for part of the audience, identity
fragmentation, bootstrap × /flags mismatch, mid-run flag edits, forced variants, and flag/experiment state inconsistency.
→ See references/bias-and-skew.md
B — Empty experiment / 0 exposures / "not enough data"
A frequent pain point. Covers SDK call (wrong evaluation method, identify() timing, dedup),
exposure capture (custom event missing variant property, required properties, ad-blockers), and
exposure-criteria match (test-account filter, filters that cannot match, eligibility ordering, events firing before exposure).
→ See references/empty-experiment.md
C — Significance / interpretation traps
Significance flipping, A/A test showing significance, Bayesian vs Frequentist confusion, multiple comparisons, low-volume variance, peeking / early stopping. Includes the running-time estimate ("target reached" is not a result) and the statistics settings that change how a result reads: sequential testing, CUPED, the confidence level and the baseline variant.
→ See references/interpretation.md
D — Numbers don't match (PostHog vs the user's SQL / raw count)
The experiment page applies an exposure scope, $multiple exclusion, test-account filter, and date range
that ad-hoc SQL almost never replicates. Covers funnel attribution (only first→last step counts for stats),
breakdowns (read from the exposure event, not the metric event), the "sum of revenue" mean-of-per-user
confusion, stale stored results, the Recordings tab against the stats, an insight that counts a narrower population than the experiment, and a metric whose title does not match its definition.
→ See references/numbers-vs-sql.md
E — Surprises after mid-run changes (incl. lifecycle and retention quirks)
Increasing rollout is safe; decreasing is caution; changing the variant split is an anti-pattern; adding metrics mid-run is p-hacking; ship-variant rewrites the flag and ending alone does not; reset clears results not the flag (apart from removing an exposure freeze); a pause, a freeze and an unfreeze can each bias the result; a changed start date moves the whole analysis. Also covers retention-metric quirks, "matured users" filtering, and long-term vs short-term metric divergence.
→ See references/mid-run-changes.md
F — Qualitative feedback: how the change landed, not how far the number moved
Groups A–E find mechanisms. F is for the question they can't reach: what the people in the experiment made of the change. A short survey, shown when users finish the experimented flow, adds that qualitative half — a rating and an optional comment, readable per variant — and works over MCP today.
It suits some experiments and not others. The gate is whether a user could describe the change without being shown both versions: a reworked flow, layout, or process work; a threshold or ranking tweak don't, however large its measured effect. Don't offer it while a mechanical diagnostic is still open — a survey on top of a broken flag gate collects opinions about a feature half the audience never received — and check whether an existing or recent survey already covers the window before proposing a new one.
Two things to get right before creating one: a survey shown to a single variant is itself a difference between the variants, and a running survey linked to the flag can generate exposure events.
→ See references/qualitative-feedback.md
Step 4 — Calibrate recommendations to experiment state
Surface diagnostics first (Step 3). Then recommend — but scope what you recommend to what the experiment's current state permits.
- Draft — config changes are free; recommend and apply.
The flag of a draft can still serve users: a draft made by a reset, or created on an existing flag, can have
feature_flag.activetrue, and an edit of that flag reaches those users at once. - Running (and
pausedorexposure_frozen, which are launched and not ended) — every change has a tradeoff. Explain the mid-run impact (anti-pattern? safe? user-visible?) before recommending. Seeconfiguring-experiment-rolloutand its reference filereferences/changing-distribution-after-launch.mdfor the mid-run rules. A pause, a freeze, an unfreeze and a reset are changes too: each one interrupts the run and can bias the result (E8, E9 and E16 inreferences/mid-run-changes.md). - Stopped / archived — the experiment AND its feature flag represent the documented outcome of the run. Recommendations are scoped to (a) interpretation of the existing data, (b) what to do for the next experiment, or (c) explaining what happened.
A diagnosis ends with findings and a recommendation. A change to a running experiment or to its flag (a new metric, new exposure criteria, a reset, a relaunch, a flag edit) is the user's decision: say what the change does to the data collected so far, and make it only on an explicit instruction.
On a stopped or archived experiment, don't preemptively offer reversal of a state mutation (ship-variant flag rewrite, manual flag edit, reset, archive). If the user asks "why did X happen?", explain X — don't append a "here's how to undo it" coda. That pattern assumes intent the user didn't signal. Conditional offers like "if this wasn't intended, you could…" or "want me to revert it?" count as preemptive too — only the user explicitly naming the reversal action ("how do I undo this?", "can I roll back ship-variant?", "how do I get the 50/50 split back?") is a request to surface reversal mechanics.
Use consistent terminology: variant split (between variants) is distinct from rollout (the share of a release condition's audience that
enters); the default exposure event (resolved_exposure_event) is distinct from a custom exposure event; the
Exclude from analysis / Use first seen variant options of Multiple variant handling (exposure_criteria.multiple_variant_handling) decide how people exposed to several variants are counted, not which event is the exposure.
Related skills
configuring-experiment-analytics— fix the exposure or metric configuration the diagnosis points atconfiguring-experiment-rollout— split-change anti-patterns and safe rollout adjustmentsmanaging-experiment-lifecycle— reset, end, or restart when the experiment can't be salvaged in placeanalyzing-experiment-session-replays— when the numbers are fine but you need to see the behavior behind themdebugging-experiments— when the question arrives as a support request about someone else's experiment and the answer is a written reply
Signals
- GitHub stars
- 721
- Forks
- 120
- Last commit
- Oct 2026
ahel recommends instead
Advanced
- Item type
- skill
- Key
diagnosing-experiment-health-posthog- Source
- github.com/posthog/posthog-foss
github.com/posthog/posthog-foss