Trustworthy Experiment Insights

SkillMonitoring & ops

Assess whether experiment results are credible enough to influence product decisions. Use when checking false positive or false negative risk, underpowered metrics, suspiciously large lifts, replication needs, meta-analysis, stratified sampling, covariate adjustment, or whether A/B test insights should be trusted.

Use Trustworthy Experiment Insights in Claude, ChatGPT or Ahel Desktop

Free. Sign in, add Trustworthy Experiment Insights and connect your AI. About a minute.

Also: Claude Code · Cursor · Codex

Then ask your AI: use the Trustworthy Experiment Insights skill

Details

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Trustworthy Experiment InsightsStart free

What this skill tells your AI

The instructions your AI receives, as published by hashgraph-online/awesome-codex-plugins in plugins/LVTD-LLC/skills/skills/trustworthy-experiment-insights/SKILL.md and read by ahel’s review.

Use this skill to decide whether an experiment result is believable enough to shape a product or engineering decision. It focuses on false positives, false negatives, power, replication, meta-analysis, stratified sampling, covariate adjustment, and suspicious result review.

Source Traceability

Primary source: Next-Level A/B Testing by Leemay Nassery. Guidance is transformed and paraphrased from Chapter 6 on false positives and negatives, meta-analysis, metric sensitivity, stratified random sampling, covariate adjustments, replication, longer runs, and statistical power.

Related skills:

  • ab-test-results-readout for standard experiment reporting.
  • experiment-sensitivity-optimization for improving precision before or during experiment design.
  • experiment-verification-monitoring for operational validity checks.

Reference Routing

NeedRead
Insight-quality conceptsreferences/core/knowledge.md
Credibility and follow-up rulesreferences/core/rules.md
Result-review scenariosreferences/core/examples.md
Step-by-step credibility reviewworkflows/review-experiment-credibility.md

Workflow

  1. Confirm the experiment was operationally valid enough to interpret.
  2. Check power, practical significance, and whether metrics were underpowered.
  3. Look for false positive risk: suspicious lift, many comparisons, early stop, weak prior, or contradiction with prior experiments.
  4. Look for false negative risk: noisy metrics, small sample, low sensitivity, or over-broad metric choice.
  5. Compare with similar experiments or run meta-analysis when available.
  6. Recommend launch, replicate, extend, investigate, or reject the result.

Output Format

# Experiment Insight Credibility Review

## Result Under Review
[Experiment, metric, observed result, and proposed decision.]

## Credibility Assessment
[Trust | Trust with caveats | Replicate | Extend | Investigate | Do not trust]

## Evidence
| Check | Finding | Risk |
|-------|---------|------|

## Follow-Up
- Replication needed:
- Longer run needed:
- Meta-analysis/comparison:
- Variance reduction opportunity:

## Decision Guidance
[What decision can be made now, and what should wait.]

Quality Bar

  • Do not celebrate a result before checking whether it could be a false positive.
  • Do not dismiss a flat result before checking power and sensitivity.
  • Do not compare against prior experiments without noting differences in population, metric, design, and timing.
  • Do not use statistical checks to hide operational failures; verify experiment health first.

Signals

GitHub stars
1k
Forks
316
Last commit
Oct 2026
Advanced
Item type
skill
Key
trustworthy-experiment-insights
Source
github.com/hashgraph-online/awesome-codex-plugins