ECS Operations Review Skill

SkillCloud & infra

Lets your agent audit your Amazon ECS clusters for reliability, security, and performance issues and suggest fixes.

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the ECS Operations Review Skill skill

About this skill

Performs a comprehensive Amazon ECS operations review across the 6 review pillars (Resiliency & HA, Observability, Security, Operations, Performance, Additional Analysis) using read-only AWS APIs, with a 7-day CloudWatch metrics baseline, recommended alarm thresholds for IDR onboarding, per-pillar P

What this skill tells your AI

The instructions your AI receives, as published by aws/tools-for-devops-agent in skills/ecs-operation-review/SKILL.md and read by ahel’s review.

Overview

Execute a comprehensive Amazon ECS operations review across the 6 review pillars with ✓/✗/N/A observations, 7-day CloudWatch baseline metrics, and alarm threshold recommendations for IDR onboarding.

Usage

  • User mentions "ECS operations review", "ECS assessment", "ECS review"
  • User provides ECS service ARN(s) for operations review
  • User asks to review ECS services for a workload
  • User requests ECS reliability or security review
  • Delegated ECS service assessment from UOPS

Core Concepts

  • Review Pillars: Six assessment dimensions — Resiliency & HA, Observability, Security, Operations, Performance, Additional Analysis
  • ECS Service/Cluster: The primary resources assessed — includes tasks, services, and cluster configuration
  • AWS API: Public AWS service APIs used for all data collection, called read-only (describe/list/get) via the AWS CLI, an AWS SDK, or an AWS API MCP
  • Baseline Metrics: 7-day CloudWatch metric history used to establish normal operating patterns

Prerequisites

  • Read-only AWS API access (describe/list/get) for ECS, CloudWatch, IAM, Application Auto Scaling, ELB, ECR, EC2 APIs — via AWS CLI, an AWS SDK, or an AWS API MCP
  • AWS Knowledge MCP access (search_documentation, read_documentation, recommend, list_regions, get_regional_availability)
  • AWS account ID and region
  • CloudWatch metrics access (7-day minimum for baseline)
  • ECS service ARN — format: arn:aws*:ecs:::service//

Skill Files

  • references/checks.md — Checks index: pillar→file map, check-ID ranges, counts, and the access-limitation / minimum-baseline rules. Read this FIRST.
  • references/pillars/resiliency.md — Resiliency & HA checks (REL1-REL14). Read only when running the Resiliency pillar.
  • references/pillars/observability.md — Observability checks (OBS1-OBS9). Read only when running the Observability pillar.
  • references/pillars/security.md — Security checks (SEC1-SEC20). Read only when running the Security pillar.
  • references/pillars/operations.md — Operations checks (OPS1-OPS9). Read only when running the Operations pillar.
  • references/pillars/performance.md — Performance checks (PERF1-PERF11). Read only when running the Performance pillar.
  • references/pillars/additional-analysis.md — Additional Analysis checks (ADD1-ADD7). Read only when running the Additional pillar.
  • references/alarm-thresholds.md — Recommended CloudWatch alarm thresholds for IDR onboarding.
  • references/common-checks-coverage.md — Crosswalk proving the review covers the shared review-common baseline (tagging, encryption, IAM least-privilege, alarms, logging, cost) via existing ECS check IDs. Read for any full review / CWR.
  • references/report-format.md — Required report/artifact structure, severity model, finding-block format, and the coverage gate. Read before generating the report.

Context Management for Checks

Do NOT read all 6 pillar files at once. Read references/checks.md (the index) first, then read each references/pillars/<pillar>.md file ONE AT A TIME as you run that pillar's checks. Agent may add checks beyond the baseline using the next sequential ID in the pillar.

Assessment Workflow

This skill is self-contained — the procedural workflow is embedded below (this copy lives in aws-operations-review and does not depend on any external agent-sops/ SOP). Execute the steps in order; this skill's references/ files supply the check definitions and alarm thresholds.

  1. Setup — Create the output directory and a scratchpad for raw API responses. Record account ID, region, and timestamp.
  2. Parse ARN — Validate the ECS service ARN (arn:aws*:ecs:*:*:service/*/*); extract cluster name and service name. Halt if the ARN region does not match the region parameter.
  3. Collect configuration data — Call the AWS APIs following the tier order in the API Tier Dependency Chain below. Tier 1 (ecs.describeServices) MUST succeed before any other call. Save each response verbatim to the scratchpad. On non-Tier-1 access/API errors, mark dependent checks N/A and continue.
  4. Resolve compute platform — Before grading any pillar, classify the service's compute platform from launchType + capacityProviderStrategy + ecs.describeCapacityProviders: Fargate (FARGATE/FARGATE_SPOT), EC2 ASG capacity provider (autoScalingGroupProvider), Managed Instances (managedInstancesProvider), launchType-only EC2/Fargate (no strategy), or ECS Anywhere (EXTERNAL). Mixed strategies are valid — record every platform present. This decision drives check applicability in every pillar (full rules in references/checks.md). Record the platform in the report header.
  5. Run pillar checks — Read references/checks.md (the index) first, then for EACH pillar read its references/pillars/<pillar>.md file one at a time and apply its checks against the collected data, recording ✓ / ✗ / N/A with an observation. Use the exact check IDs and severities; mark platform-specific checks N/A where they don't apply (per the resolved compute platform). May add checks beyond the baseline.
  6. Generate report — Produce the per-service review artifact following references/report-format.md exactly (Workload Details, per-pillar ✓/✗/N/A scorecards for all 6 pillars, Prioritized Action Plan, Detailed Findings, Recommended CloudWatch Alarms from references/alarm-thresholds.md, Access Limitations, Review Summary). Before finalizing, run the report-format Coverage gate — every check ID across all 6 pillars must appear in a scorecard as ✓/✗/N/A (including passes; no pillar dropped or truncated), the alarms table must be present, every ✗ needs a detailed finding block, and the shared review-common baseline must be accounted for per references/common-checks-coverage.md (all eight common checks covered via their ECS equivalents or ⚪ N/A with a reason). Default to a Markdown artifact; render DOCX only if asked (build from the same content). Strip internal check IDs from the customer-facing report. Return the Review Summary with verified counts, then delete the scratchpad.

Review Pillars:

  • Resiliency and High Availability (REL1-REL14) — Multi-AZ, desired count, deployment config, circuit breaker, deployment alarms, health checks, subnet AZ spread, capacity-provider managed termination protection, target-group deregistration delay, capacity provider infrastructure multi-AZ
  • Observability (OBS1-OBS9) — Container Insights, CloudWatch alarms, logging, log retention, distributed tracing, metrics monitoring
  • Security (SEC1-SEC20) — IAM least privilege, network mode, secrets management, ECR image scanning, security groups, VPC endpoints, private connectivity, encryption at rest, encryption in transit (TLS), VPC Flow Logs, GuardDuty Runtime Monitoring
  • Operations (OPS1-OPS9) — Deployment controller, resource tagging, IaC-managed, platform version, ECS agent version, ECS Exec posture and session audit logging
  • Performance (PERF1-PERF11) — Auto scaling, CPU/memory rightsizing, capacity provider strategy, managed scaling / targetCapacity headroom, CapacityProviderReservation 7-day baseline analysis, base/weight strategy design, Compute Optimizer recommendations
  • Additional Analysis & Recommendations (ADD1-ADD7) — Graviton/ARM64, Fargate Spot, Service Connect, cost optimization, CloudWatch Logs Insights queries, ECS Managed Instances evaluation

AWS API Summary

All calls below are public AWS API operations. Use read-only (describe/list/get) operations only, via the AWS CLI, an AWS SDK (e.g. boto3), or an AWS API MCP with least-privilege read-only credentials.

ECS APIs (Tier 1, 2, 3)

APITierPurpose
ecs.describeServices1Foundation — service config, task def, LB, deployment, network
ecs.describeTaskDefinition2Container defs, CPU/memory, roles, log config, network mode
ecs.describeClusters2Cluster settings, Container Insights, capacity providers, Exec logging config (include=["CONFIGURATIONS"] for OPS9)
ecs.listTasks2Running task ARNs for the service
ecs.describeTasks3Task health, AZ spread, connectivity status
ecs.listContainerInstances2Container instance ARNs for EC2 launch type clusters
ecs.describeContainerInstances3Agent version, AMI ID, instance status (EC2 only)
ecs.describeCapacityProviders2Compute platform classification (ASG vs Managed Instances vs Fargate), managed termination protection (REL12), managed scaling status/targetCapacity (PERF9), MI network config (REL14)

Application Auto Scaling APIs (Tier 2)

APITierPurpose
applicationautoscaling.describeScalingPolicies2Auto scaling policies for the service
applicationautoscaling.describeScalableTargets2Min/max capacity configuration

ELB APIs (Tier 2)

APITierPurpose
alb.describeTargetHealth2Target health for service tasks behind ALB/NLB (skip if no LB configured)
alb.describeTargetGroups2Target group details including LoadBalancerArns — used to determine LB type (ALB vs NLB) by ARN path segment: /app/ = ALB, /net/ = NLB, for correct alarm recommendations (skip if no LB configured)
elbv2.describeListeners2Listener protocol/port for the LB fronting the service — HTTPS/TLS vs plaintext HTTP/TCP for encryption-in-transit (SEC20); uses LoadBalancerArns from describeTargetGroups (skip if no LB configured)

IAM APIs (Tier 4a, 4b)

APITierPurpose
iam.listAttachedRolePolicies4aManaged policies on execution/task roles
iam.listRolePolicies4aInline policy names on execution/task roles
iam.getRolePolicy4bInline policy document for execution/task roles

ECR APIs (Tier 3)

APITierPurpose
ecr.describeRepositories3Image scanning config, tag immutability for container image repos

EC2/VPC APIs (Tier 2, 3, 4)

APITierPurpose
ec2.describeSecurityGroups2Security group rules for service ENIs (awsvpc mode)
ec2.describeSubnets2Subnet AZ distribution for service network config
ec2.describeVpcEndpoints3VPC endpoints for ECR, CloudWatch Logs, Secrets Manager (uses VPC ID from describeSubnets)
ec2.describeRouteTables3Route table entries for NAT/internet access assessment
ec2.describeNatGateways3NAT Gateway availability for private subnets
ec2.describeImages4AMI creation date for container instance AMI currency check (EC2 only, uses imageId from describeContainerInstances)
ec2.describeVolumes3EBS volume encryption status for task-attached / container-instance volumes (SEC17)
ec2.describeFlowLogs3VPC Flow Logs enablement for the service VPC (SEC18, uses VPC ID from describeSubnets)

CloudWatch APIs (Tier 2, 5)

APITierPurpose
cloudwatch.describeAlarms2Existing alarms for ECS service
cloudwatch.getMetricStatistics57-day baseline: CPU, memory, task count; plus CapacityProviderReservation (AWS/ECS/ManagedScaling) for EC2 ASG capacity providers (PERF10)

CloudWatch Logs APIs (Tier 3)

APITierPurpose
logs.describeLogGroups3Log retention setting and Logs Insights query targeting for the awslogs group (OBS8, ADD6 — uses awslogs-group from task definition)

GuardDuty APIs (Tier 2)

APITierPurpose
guardduty.listDetectors2Detector presence in region (SEC19)
guardduty.getDetector2Runtime Monitoring feature status for ECS (SEC19 — uses detector ID from listDetectors)

Compute Optimizer APIs (Tier 2)

APITierPurpose
computeoptimizer.getECSServiceRecommendations2ECS service task CPU/memory rightsizing recommendations (PERF8)

AWS Knowledge MCP

ToolPurpose
aws___search_documentationSearch across all AWS documentation with optional topic-based filtering
aws___read_documentationRetrieve and convert AWS documentation pages to markdown
aws___recommendGet content recommendations for AWS documentation pages
aws___list_regionsRetrieve a list of all AWS regions
aws___get_regional_availabilityRetrieve AWS regional availability information

API Tier Dependency Chain

Tier 1: ecs.describeServices (FOUNDATION — must complete first)
  ├─ extracts: taskDefinition ARN, clusterArn, loadBalancers,
  │            desiredCount, launchType, networkConfiguration, tags
  │
  ├─► Tier 2 (parallel): ecs.describeTaskDefinition, ecs.describeClusters,
  │     ecs.listTasks, cloudwatch.describeAlarms,
  │     applicationautoscaling.describeScalingPolicies,
  │     applicationautoscaling.describeScalableTargets,
  │     ecs.describeCapacityProviders (compute platform classification;
  │       ASG providers — REL12/PERF9; Managed Instances providers — REL14),
  │     guardduty.listDetectors ─► guardduty.getDetector (SEC19),
  │     computeoptimizer.getECSServiceRecommendations (PERF8),
  │     alb.describeTargetHealth (if LB configured),
  │     alb.describeTargetGroups (if LB configured — resolves ALB vs NLB type from LoadBalancerArns: /app/ = ALB, /net/ = NLB; also deregistration delay for REL13),
  │     elbv2.describeListeners (if LB configured — listener protocol for encryption-in-transit SEC20),
  │     ec2.describeSecurityGroups (from networkConfiguration.securityGroups),
  │     ec2.describeSubnets (from networkConfiguration.subnets),
  │     ecs.listContainerInstances (EC2 launch type only)
  │     │
  │     ├─► Tier 3: ecs.describeTasks (using task ARNs from listTasks)
  │     │            ecs.describeContainerInstances (EC2 only, using instance ARNs from listContainerInstances)
  │     │            ecr.describeRepositories (using repo name from task definition image URI)
  │     │            ec2.describeVpcEndpoints (using VPC ID from describeSubnets)
  │     │            ec2.describeRouteTables (using subnet IDs from describeSubnets)
  │     │            ec2.describeNatGateways (using VPC ID from describeSubnets)
  │     │            ec2.describeFlowLogs (using VPC ID from describeSubnets — SEC18)
  │     │            ec2.describeVolumes (task-attached / container-instance EBS encryption — SEC17)
  │     │            logs.describeLogGroups (using awslogs-group from task definition — OBS8, ADD6)
  │     │
  │     ├─► Tier 4 (EC2 only): ec2.describeImages (using imageId from describeContainerInstances)
  │     │
  │     └─► Tier 4a (parallel): iam.listAttachedRolePolicies (execution + task role),
  │           iam.listRolePolicies (execution + task role)
  │           │
  │           └─► Tier 4b: iam.getRolePolicy (execution + task role)
  │                 (uses policy names from listRolePolicies)
  │
  └─► Tier 5 (per-metric loop): cloudwatch.getMetricStatistics

Access Limitation Handling

When AWS API calls return access denied or authorization errors:

  • Mark dependent checks as N/A with observation: "Unable to assess — access denied on {{api_name}}. Manual verification recommended."
  • Include a dedicated Access Limitations section in the report listing all checks that could not be evaluated due to permissions
  • Continue with remaining assessable checks — do NOT halt the entire assessment for non-Tier-1 access errors
  • In the Review Summary, note how many checks could not be evaluated due to access limitations

Example Output

The agent produces a per-service review artifact (Markdown by default; DOCX if asked) containing: service configuration summary, all 6 review pillar scorecards (✓/✗/N/A with observations, every check including passes), 7-day baseline metrics, the recommended-alarms table with clickable doc links, and priority action items. A Review Summary with verified counts is returned to the orchestrator.

What Not to Do

  • Don't skip any pillar — all 6 must be assessed
  • Don't use write or mutating API calls — this is a read-only assessment; use describe/list/get operations only
  • Don't hardcode doc URLs for ✗ check findings — use AWS Knowledge MCP (aws___search_documentation) to supplement the doc links provided in checks.md
  • Don't hardcode doc URLs for alarm recommendation hyperlinks — use the doc_url column from references/alarm-thresholds.md as the canonical link target
  • Don't skip alarm recommendations — this is a core IDR deliverable
  • Don't omit the metrics baseline section — if 7-day data is unavailable, note the limitation in the report rather than skipping it
  • Don't hallucinate findings — only report what AWS API data confirms
  • Don't silently skip checks when access is denied — always mark as N/A with explicit access limitation note

Failure Recovery

  • If ecs.describeServices fails after retries: HALT workflow — delete {{scratchpad_dir}}/ and return error to orchestrator
  • If Tier 2-5 APIs fail: mark dependent checks as N/A, continue assessment with available data
  • If report generation fails: default to the Markdown artifact (no external dependency); only fall back to DOCX (python-docx, retry with --user on install failure) when a DOCX deliverable was explicitly requested
  • If artifact write fails: capture the error, delete scratchpad, return error to orchestrator

Success Criteria

  • All 6 review pillars assessed with ✓/✗/N/A observations — coverage gate passed: every check ID across all 6 pillars appears in a scorecard (including passes), no pillar dropped or truncated
  • Review artifact generated at {{output_directory}}/ (Markdown by default; DOCX only if requested)
  • Recommended-alarms table populated from alarm-thresholds.md, marking exist-vs-missing
  • 7-day baseline metrics included (or limitation noted)
  • Every ✗ has a detailed finding block; every recommendation includes a resolvable AWS documentation link
  • Access limitations explicitly documented; unobtainable checks are N/A-with-reason, never omitted or guessed
  • Review Summary with verified counts returned to orchestrator
  • {{scratchpad_dir}}/ deleted after report verified

Signals

GitHub stars
82
Forks
60
Last commit
Sep 2026
Advanced
Item type
skill
Key
ecs-operation-review
Source
github.com/aws/tools-for-devops-agent