production-deployment

SkillCloud & infra

Zero-downtime deployments with pre-flight checks, staged rollouts, and rollback plans. Never ship to production without a verified rollback strategy.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the production-deployment skill

What this skill tells your AI

The instructions your AI receives, as published by developersglobal/ai-agent-skills in skills/production-deployment/SKILL.md and read by ahel’s review.

Overview

Production is not a test environment. Every deployment is a live operation with real consequences — user impact, data integrity risks, and potential outages. This skill encodes the discipline senior engineers apply before, during, and after every production deployment.

The core rule: never deploy without a rollback plan you've verified can execute in under 5 minutes.

When to Use

  • Before any deployment to a production or production-equivalent environment
  • When reviewing deployment scripts or CI/CD pipelines
  • When adding new services or infrastructure changes

Process

Step 1: Pre-Deployment Checklist

  1. All tests pass — CI is green on the exact commit being deployed. Not "mostly green."
  2. Migrations are backward-compatible — The old code must work with the new schema (for zero-downtime). New columns are nullable; columns aren't dropped until after full rollout.
  3. Feature flags configured — New features are behind flags, off by default.
  4. Rollback plan written — Document exactly how to rollback: which commands, which configs, estimated time.
  5. Deployment window confirmed — Low-traffic period? On-call engineer available?
  6. Stakeholders notified — Anyone affected by downtime or behavior change knows.

Verify: All 6 checklist items confirmed. Do not proceed if any is blocked.

Step 2: Staged Rollout

  1. Never deploy to 100% of traffic immediately. Use a staged rollout:
    • Canary: 1–5% of traffic
    • Staged: 10% → 25% → 50% → 100%
  2. Monitor key metrics at each stage for at least 15 minutes before expanding:
    • Error rate (baseline vs. current)
    • Latency p50, p95, p99
    • Business metrics (conversion, orders, etc.)
  3. Define your abort threshold before starting: "If error rate exceeds X% or latency p99 exceeds Y ms, rollback immediately."

Verify: Rollout stages and abort thresholds are documented before deployment begins.

Step 3: Deploy

  1. Execute the deployment using your CI/CD pipeline (not manual commands).
  2. Monitor dashboards in real-time during the rollout.
  3. Keep communication channel open with on-call engineer.
  4. Do not perform any other changes during a deployment (no "quick fixes").

Verify: Deployment running via CI/CD, dashboards being monitored actively.

Step 4: Post-Deployment Verification

  1. Smoke tests pass on production.
  2. Key user journeys manually verified.
  3. Error rate within normal range (15 minutes post-deploy).
  4. No unexpected alerts triggered.
  5. Run post-deploy integration tests if available.

Verify: All post-deploy checks confirmed green. Deployment marked successful.

Step 5: Rollback (if needed)

  1. If any abort threshold is hit: rollback immediately, without debate.
  2. Execute the pre-written rollback plan.
  3. Verify rollback complete: service restored, error rate normalized.
  4. Write an incident report — even for near-misses.

Verify: Rollback completes in under 5 minutes. Service restored.

Common Rationalizations (and Rebuttals)

ExcuseRebuttal
"It works in staging"Staging is not production. Different data, traffic, and configuration.
"It's just a small change"Small changes cause the majority of outages.
"We don't have time for staged rollout"You have even less time for an incident.
"I'll watch it for a few minutes"15 minutes minimum. Most production failures take time to materialize under load.
"We can rollback if needed"Do you have a written, tested rollback plan? No? Then you can't.

Red Flags

  • Deploying directly to 100% without a staged rollout
  • No rollback plan documented before deployment
  • Deploying breaking schema changes without backward compatibility
  • Running deployment from a local machine, not CI/CD
  • Deploying during high-traffic periods without approval
  • "I'll fix any issues after we deploy"

Verification

  • All tests passing on exact commit being deployed
  • Migrations are backward-compatible
  • Rollback plan written and executable in <5 minutes
  • Staged rollout plan with abort thresholds defined
  • Post-deploy smoke tests passed
  • Dashboards clean for 15 minutes post-deploy

References

Signals

GitHub stars
66
Forks
9
Last commit
May 2026
Advanced
Catalog kind
skill
Gateway key
production-deployment
Source
github.com/developersglobal/ai-agent-skills