CAST AI Failure Triage

SkillAI & models

Diagnose CAST AI connection, agent, node autoscaling, and workload autoscaling failures without making speculative changes. Use when a cluster is disconnected, recommendations are absent, pods are not optimized, or capacity does not scale as expected. Trigger with: "debug CAST AI", "CAST AI is not scaling", "why is CAST AI disconnected".

Use CAST AI Failure Triage in Claude, ChatGPT or Ahel Desktop

Free. Sign in, add CAST AI Failure Triage and connect your AI. About a minute.

Also: Claude Code · Cursor · Codex

Then ask your AI: use the CAST AI Failure Triage skill

Details

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

CAST AI Failure TriageStart free

What this skill tells your AI

The instructions your AI receives, as published by jeremylongshore/tons-of-skills-marketplace in skills/.curated/castai-common-errors/SKILL.md and read by ahel’s review.

Overview

Separate observation, connectivity, policy, capacity, and disruption failures before proposing a change. Preserve the failing state, use current component topology, and stop when the evidence requires cloud-provider or CAST AI support access.

Prerequisites

  • The exact kube context, cluster, region, time window, and observed symptom
  • Read-only access to the castai-agent namespace
  • The declared installation owner: castctl, Terraform, GitOps, or console

Instructions

Step 1: Freeze the symptom

Record expected versus actual behavior, timestamps, workload identity, pending-pod reason, and recent configuration changes. Use Read and Grep on runbooks and IaC to determine whether Cost Monitoring, Node Autoscaling, or Workload Autoscaling is actually enabled.

Step 2: Check installation health

Use Bash(castctl:) for version or non-mutating status commands supported by the installed client. Use Bash(helm:) to inspect releases and values, then Bash(kubectl:*) to inspect workloads, readiness, events, and bounded logs in castai-agent. Do not restart components before collecting evidence.

Step 3: Classify the failure plane

PlaneEvidenceLikely boundary
ConnectionAgent readiness, outbound failures, console disconnectIdentity, network, or cloud permissions
Node scalingPending pods, policy bounds, node-template fitUnsatisfied constraints or maximum CPU boundary
Workload scalingMissing recommendations, policy assignment, metricsMetrics server, confidence, policy, or unsupported workload
DisruptionEviction denial, PDB events, deferred changesPDB or selected apply mode
ReportingMissing cost or savings windowIngestion, baseline, adoption, or pricing configuration

Step 4: Test one hypothesis

Choose the smallest reversible check. Confirm regional endpoint alignment, effective scaling-policy assignment, metrics availability, supported workload type, node-template constraints, and cloud quota. Treat the deprecated cluster minimum CPU setting as migration debt, not a current control to add.

Step 5: Decide the owner and remedy

Map the evidence to the owning layer. Change repository-managed values only through their source of truth; do not mix console edits into Terraform or GitOps ownership. Escalate with a redacted bundle when the failure is inside the hosted control plane or an undocumented provider response.

Tool Discipline

Use Read and Grep for configuration and runbook evidence. Use Bash(kubectl:), Bash(helm:), and Bash(castctl:*) only for bounded inspection commands. Do not apply, upgrade, restart, connect, disconnect, or expose Secret objects during diagnosis.

Output

  • A timestamped symptom and environment summary
  • Evidence grouped by failure plane
  • One supported root-cause hypothesis with confidence
  • A reversible remedy, rollback condition, and escalation owner

Examples

Recommendations are absent because metrics-server is missing, so the remedy belongs to cluster observability. A node remains pending because every approved node template conflicts with its constraints; increasing a global limit without reviewing the workload is not the remedy.

Error Handling

FailureResponse
Kube context is ambiguousStop before any cluster command and resolve it
Logs include credentials or inventoryRedact locally and do not attach raw output
A PDB blocks Immediate modePreserve the PDB and evaluate Deferred mode with the workload owner
Evidence points to cloud quotaEscalate to the cloud owner with the exact denied dimension

Resources

Signals

GitHub stars
3k
Forks
415
Last commit
Oct 2026
Advanced
Item type
skill
Key
castai-common-errors
Source
github.com/jeremylongshore/tons-of-skills-marketplace