QA Resilience

SkillAI & models

Designs and tests distributed-system resilience. Use when adding retries, deadlines, hedging, circuit breakers, overload protection, chaos experiments, or SLO reliability gates.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the QA Resilience skill

What this skill tells your AI

The instructions your AI receives, as published by vasilyu1983/ai-agents-public in frameworks/shared-skills/skills/qa-resilience/SKILL.md and read by ahel’s review.

Use this skill when reliability work is about failure behavior, overload protection, degraded mode, or resilience testing. The goal is not "add retries everywhere." The goal is predictable failure handling, clear ownership, and testable recovery behavior.

Quick Reference

SymptomStart With
slow or hanging dependencydeadline and timeout budget
transient dependency failurebounded retry with jitter and retry budget
sustained dependency failurecircuit breaker and fallback
rate-limited dependencyhonor Retry-After, expose degraded behavior, and test quota paths intentionally
one bad host in a healthy pooloutlier detection or endpoint ejection
queue or pool saturationbulkheads, concurrency limits, load shedding
non-critical feature outagegraceful degradation or feature flag fallback
resilience validationdeterministic fault injection before chaos

When to Use This Skill

  • retries, deadlines, hedging, breakers, bulkheads, and overload protection
  • degraded-mode UX or API behavior
  • service-mesh or gateway resilience policy
  • chaos engineering, game days, DR drills, and fault injection
  • release gates based on failure behavior, not only happy-path load tests

Route Elsewhere


Workflow

  1. Identify the critical user journeys and the dependencies that can break them.
  2. Define the contract per dependency:
    • timeout or deadline budget
    • retry ownership
    • breaker or outlier policy
    • concurrency and queue limits
    • degraded behavior if the dependency is unavailable
  3. Decide where the policy lives:
    • app code
    • client library
    • mesh or gateway
  4. Test in stages:
    • deterministic fault injection
    • staged chaos in non-production
    • narrow prod canary or game day only with guardrails
  5. Define pass or fail signals:
    • burn rate
    • p95 or p99
    • fallback rate
    • breaker transitions
    • shed volume
    • recovery time

Pattern Rules

  • retries happen at one layer only
  • deadlines come before retries
  • hedging is only for idempotent or cancellation-safe reads
  • overload handling must shed early instead of collapsing late
  • readiness and liveness must stay bounded and shallow
  • resilience behaviors belong in targeted checks; do not let rate-limit or degraded-mode coverage leak into unrelated happy-path suites
  • prod experiments require blast-radius limits, abort criteria, dashboards, and owners

Failure Modes to Validate

  • timeouts and deadline propagation
  • retry storms and duplicate side effects
  • partial dependency outages
  • slow downstreams and long-tail latency
  • queue buildup and connection-pool exhaustion
  • one-bad-host behavior inside a pool
  • degraded-mode responses and stale-data fallbacks
  • rate-limit handling, Retry-After, and client backoff expectations
  • visible state convergence after recovery or backend resets
  • failover and failback behavior
  • metastable failure: a self-sustaining feedback loop (retry amplification, cache-miss stampede, queue backlog, connection-pool churn) that does not self-resolve after the original trigger clears — see references/cascading-failure-prevention.md

Testing Ladder

Deterministic first

  • inject latency, errors, timeouts, malformed payloads, and unavailable endpoints in controlled tests
  • verify the intended control activates and the wrong controls do not

Chaos second

  • start in non-production
  • use a small blast radius and a fixed time window
  • stop immediately on error-budget or customer-impact breach

DR and game days last

  • validate RTO and RPO claims explicitly
  • rehearse recovery ownership, not only technical failover

Operational Guardrails

  • every experiment needs a stated hypothesis and steady-state metric
  • every run should capture timestamps, targets, blast radius, and dashboard links
  • telemetry fields for retries, breaker transitions, hedging, shedding, and fallback are part of the resilience contract
  • releases should be gated on reliability behavior, not only resource usage

Anti-Patterns

  • no timeouts
  • retries at every hop
  • fixed-interval or unbounded retries
  • fixed per-call retry count with no system-wide retry budget (does not tighten as error rate rises)
  • retries without idempotency
  • hedging unsafe writes
  • no bulkheads or queue bounds
  • deep readiness or liveness checks
  • silent degraded mode
  • untested failover plans
  • happy-path-only load testing

Scripts

ScriptPurpose
scripts/resilience_checker.pyScores resilience pattern coverage and reports gaps

Typical usage:

python scripts/resilience_checker.py assess --input data/sample-service-profile.json
python scripts/resilience_checker.py gaps --input data/sample-service-profile.json
python scripts/resilience_checker.py report --input data/sample-service-profile.json --output resilience-report.md

See scripts/README.md for the input format and scoring logic.

ASCII Flow

Resilience request
  -> Identify failure mode: slow, transient, sustained, overloaded, or degraded
  -> Set deadlines, retry budgets, isolation, and fallback ownership
  -> Add telemetry for saturation, errors, latency, and degraded behavior
  -> Validate with deterministic fault injection before chaos experiments
  -> Gate release on recovery evidence, SLO impact, and rollback path
  -> Document runbook actions and residual risk

Navigation

Foundation applied recipes

Core references

Operational resources

Related Skills

Fact-Checking

  • Known bugs, regressions, framework/compiler/runtime footguns, and version-specific crash or workaround guidance must be verified against current primary web sources before being treated as current fact.
  • Verify current platform features, mesh capabilities, and vendor-specific behavior before final answers when the recommendation depends on a live product.
  • Prefer primary docs for runtime or tooling specifics.
  • If web access is unavailable, keep external product guidance marked as unverified.

Learnings Loop

Before applying this skill on a non-trivial task, read learnings.consolidated.md in this directory (and learnings.md if present).

After applying it, if you encountered a pattern worth remembering, a mistake worth preventing, or a domain fact that surprised you, append one dated bullet to learnings.md via agents-skills-feedback-loop/scripts/append_learning.py. Do not modify SKILL.md itself.

Signals

GitHub stars
87
Forks
19
Last commit
Sep 2026

ahel review

  • K1binfo
    installs-packages (in references/chaos-tooling-recipes.md)

Automated review, not a security audit. Ruleset v1+k2.

Advanced
Catalog kind
skill
Gateway key
qa-resilience
Source
github.com/vasilyu1983/ai-agents-public