absolute-deflake

SkillDev tools

Flaky test fixes: detect nondeterministic tests empirically (repeat/shuffle/parallel runs), diagnose the root cause, fix it — never retry/skip/sleep — and verify across many randomized runs. Triggers on "absolute deflake", "fix flaky tests", "CI is flaky", "this test fails randomly/intermittently".

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the absolute-deflake skill

What this skill tells your AI

The instructions your AI receives, as published by maddhruv/absolute in skills/absolute-deflake/SKILL.md and read by ahel’s review.

Start your first response with the 🧪 emoji.

Absolute Deflake

Find tests that pass and fail nondeterministically, diagnose the root cause of each, and fix it — not by retrying or skipping, but by removing the source of nondeterminism. Output is evidence (failure rate per test) → cause → fix, verified by repeated runs.

Runs the shared engine in references/health-engine.md — read it for the DETECT → SCAN → TRIAGE → FIX → VERIFY → REPORT loop and the safety contract. This file covers only what's specific to flaky tests.


When to use

  • "Our CI is flaky", "this test fails randomly", "fix the intermittent failures".
  • A test passes locally but fails in CI (or vice versa), or fails ~1 in N runs.
  • Burning down a backlog of retry/skip-marked tests that mask real flakiness.

Not for tests that fail deterministically — that's a real bug or a real regression (/absolute work for a fix, or just fix it). deflake targets nondeterministic failures.


What it scans

Establish flakiness empirically — a test isn't flaky because someone said so. Use preferences.health.deflakeRuns from config as the default N for repeat-runs (else 20):

EcosystemRepeat-run / detect
Jest/Vitestrun suite N× (--run loop), randomize order (--shuffle / testSequencer)
pytestpytest-randomly + pytest --count=N (pytest-repeat); -p no:randomly to A/B
Gogo test -count=N -shuffle=on ./..., -race

Also mine signals: existing retry/flaky/skip annotations, CI history if reachable, and run the suite both in isolation and in full/parallel — order- and concurrency- dependent failures only show one way. Record a failure rate per suspect test.


Common root causes (diagnose, don't guess)

CauseTellFix
Test-order / shared statepasses alone, fails in suite (or vice versa)isolate state; reset/teardown between tests
Time / clockfails near midnight, DST, or under loadfake timers / inject clock; no real sleep
Async race / missing awaitfails under parallelism or slow CIawait the actual condition; no fixed timeouts
Randomnessfails ~X% with no patternseed the RNG; fix the seed in tests
Network / external I/Ofails offline or on slow linksmock/stub the boundary
Unordered collectionsfails on map/set iteration ordersort before asserting
Resource leak / port reusefails on repeat or parallel runsunique resources; clean up

Risk ranking (TRIAGE)

WaveClassDefault
1clear, isolated cause (seed, await, fake clock, sort)fix now
2shared-state / ordering — needs fixture refactorfix this pass, per test
3flakiness pointing at a real product race, not just the testgated — surface; may be a genuine bug to fix in code

A flaky test sometimes means the code has a race, not the test. Don't "stabilize" the test into hiding a real concurrency bug — flag wave-3 cases for a real fix.


Fix & verify

  • Fix the cause. Then prove it: re-run the test many times (and shuffled / parallel / with -race) — green once is not deflaked; green across N randomized runs is.
  • Remove the retry/skip/flaky annotation that was masking it once the cause is fixed.
  • Never "fix" by adding retries, raising timeouts blindly, sleep, or skipping the test — that hides flakiness, doesn't remove it.
  • Re-run the full suite to confirm the fix didn't destabilize neighbors.

Gotchas

  1. Retry/skip as a fix. Masks the flake, ships the nondeterminism. Forbidden here.
  2. sleep to dodge a race. Slows the suite and still flakes under load. Await the condition.
  3. One green run = done. Flakes are probabilistic — verify with many randomized runs.
  4. Stabilizing a real product race. If the code races, fix the code, not just the assertion.
  5. Ignoring order/parallel dimension. Run isolated and in-suite; the bug hides in whichever you skip.

Companion commands

  • /absolute upgrade — a flaky suite makes upgrade verification unreliable; deflake first.
  • /absolute debt — flaky-test annotations are test debt; this clears them at the root.
  • /absolute work — when the flake is a genuine product-code race needing real design.

Signals

GitHub stars
211
Forks
30
Last commit
Jul 2026
Advanced
Catalog kind
skill
Gateway key
absolute-deflake
Source
github.com/maddhruv/absolute