A/B Test Planner
SkillMonitoring & opsDesign rigorous A/B test plans with hypothesis, sample size calculation, Minimum Detectable Effect (MDE), randomization strategy, and decision rules. Includes guardrail metrics and rollout playbook. Use when planning product experiments, conversion optimization, or data-driven feature decisions.
Use A/B Test Planner in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add A/B Test Planner and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the A/B Test Planner skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; Ahel provides instructions and does not run this skill.
No other account needed.
Add Ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
What this skill tells your AI
The instructions your AI receives, as published by hoavdc/codexkit in skills/codexkit-a-b-test-planner/SKILL.md and read by Ahel’s review.
When to Use
- Before running any product experiment or feature test
- When optimizing conversion funnels or UX flows
- When leadership requires statistical rigor for feature decisions
- When planning multi-variant tests or sequential experiments
Procedure
Step 1 — Hypothesis
Write a clear, falsifiable hypothesis:
- If [change we're making]
- Then [metric we expect to change]
- Because [reasoning / user insight]
Step 2 — Metrics
| Type | Metric | Current Baseline |
|---|---|---|
| Primary | The metric that determines success | [value] |
| Secondary | Supporting metrics that provide context | [value] |
| Guardrail | Metrics that must NOT degrade | [value] |
Step 3 — Sample Size & Duration
Calculate required sample size using:
- Baseline conversion rate (p₁)
- Minimum Detectable Effect (MDE) — smallest meaningful change
- Statistical significance level (α) — typically 0.05
- Statistical power (1−β) — typically 0.80
n = f(p₁, MDE, α, β) → use standard sample size calculator
Duration = n / (daily traffic × allocation %)
Step 4 — Randomization Plan
- Randomization unit: user, session, device, or account
- Allocation: 50/50, or asymmetric with justification
- Stratification: any segments to balance (geography, plan, device)
- Exclusion: users to exclude (employees, bots, existing tests)
Step 5 — Decision Rules
| Outcome | Criteria | Action |
|---|---|---|
| Winner | Primary metric ↑ ≥ MDE, p < 0.05, guardrails stable | Ship to 100% |
| Neutral | No significant difference | Keep control, iterate hypothesis |
| Loser | Primary metric ↓ significantly | Revert, analyze why |
| Guardrail breach | Any guardrail metric degrades > threshold | Stop test immediately |
Step 6 — Rollout Playbook
- Ramp: 5% → 25% → 50% → 100% over [days]
- Monitoring: check metrics daily during ramp
- Rollback trigger: guardrail breach or unexpected anomaly
Inputs
| Input | Required | Format |
|---|---|---|
| Change description | Yes | What is being tested |
| Baseline metric | Yes | Current value |
| MDE target | Yes | Percentage or absolute |
| Daily traffic | Yes | Number of users/events |
| Test duration budget | Recommended | Max days willing to run |
Output
## A/B Test Plan — [Test Name]
### Hypothesis
If we simplify the checkout form from 5 fields to 3 fields,
then checkout completion rate will increase by ≥ 5%,
because user research shows 40% abandon at the address step.
### Metrics
| Type | Metric | Baseline | Target |
|------|--------|----------|--------|
| Primary | Checkout completion rate | 45% | ≥ 50% |
| Secondary | Average order value | $65 | Stable |
| Guardrail | Revenue per user | $12 | No decrease |
| Guardrail | Error rate | 0.5% | No increase |
### Sample Size
| Parameter | Value |
|-----------|-------|
| Baseline rate | 45% |
| MDE | 5% (absolute) |
| Significance (α) | 0.05 |
| Power (1−β) | 0.80 |
| Required n per variant | ~1,600 |
| Daily traffic | 800 users |
| Allocation | 50/50 |
| **Estimated duration** | **4 days** |
### Randomization
- Unit: User (cookie-based)
- Allocation: 50% control / 50% variant
- Exclusion: Internal users, users in other active tests
### Decision Rules
[As defined in procedure]
### Rollout Playbook
Day 1–2: 10% ramp → monitor → Day 3–4: 50% → Day 5: 100%
Definition of Done
- Hypothesis is clear and falsifiable
- Primary, secondary, and guardrail metrics defined
- Sample size calculated with stated parameters
- Duration estimated based on traffic
- Randomization plan documented
- Decision rules and rollout playbook included
Examples
Prompt
We want to test if a new onboarding flow increases activation rate.
Current activation: 32%. MDE: 3 percentage points. Daily new users: 500.
We have max 14 days for the test.
Design a complete A/B test plan with sample size and decision rules.
Quality Criteria
- Every finding is tied to a specific evidence source (log, test, metric)
- Pass/fail criteria are binary and measurable — no subjective judgments
- Severity levels are assigned with clear thresholds
- Remediation steps are provided for all critical and high findings
Verification (4C)
| Check | Question |
|---|---|
| Correctness | Are all pass/fail criteria applied against the correct standard or rule? |
| Completeness | Were all required dimensions or checklist items evaluated? |
| Context-fit | Does the verification scope match the actual risk level of the deliverable? |
| Consequence | If this passed verification but had a hidden flaw, what is the worst-case impact? |
Edge Cases
- Incomplete data for full assessment — Document which checks were limited and flag for re-verification when data becomes available.
- Ambiguous pass/fail criteria — Request clarification from the standard owner before scoring. Mark as 'Needs Review'.
- Multiple overlapping standards — Identify the governing standard and note where others diverge.
Changelog
- v1.0.0 — Initial release
Signals
- GitHub stars
- 25
- Forks
- 13
- Last commit
- Oct 2026
Advanced
- Item type
- skill
- Key
codexkit-a-b-test-planner- Source
- github.com/hoavdc/codexkit
Related picks
Skill · larksuite
The pick for Markdownmarkdown-formatter
Skill · nvidia
The pick for Markdowninternal-comms
Skill · anthropics
More in Monitoring & opsagent-eval
Skill · affaan-m
More in Monitoring & opspricing
Skill · coreyhaines31
More in Monitoring & opslark-okr
Skill · larksuite
More in Monitoring & ops