rca:kind
SkillDev toolsRoot cause analysis on local Kind cluster - fast local debugging with full access
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the rca:kind skill
What this skill tells your AI
The instructions your AI receives, as published by rossoctl/rossoctl in .claude/skills/rca:kind/SKILL.md and read by ahel’s review.
Root cause analysis workflow for failures on local Kind clusters.
Context-Safe Execution (MANDATORY)
All diagnostic commands MUST redirect output to files.
export LOG_DIR="${LOG_DIR:-${WORKSPACE_DIR:-/tmp}/rossoctl-rca}"
mkdir -p "$LOG_DIR"
Rules:
- ALL kubectl/oc commands redirect to
$LOG_DIR/<name>.log - ALL analysis in subagents:
Task(subagent_type='Explore')withGrep(not full file reads) - Main context only sees: OK/FAIL status and subagent summaries
When to Use
- Kind E2E tests failed locally or in CI
- Need to reproduce and debug CI Kind failures
- Ollama/agent issues specific to the Kind environment
Auto-approved: All read and debug operations on Kind clusters are auto-approved.
Cluster Concurrency Guard
Only one Kind cluster at a time. Before any cluster operation, check:
kind get clusters 2>/dev/null
- No clusters → proceed normally (create cluster)
- Cluster exists AND this session owns it → reuse it (skip creation, inspect directly)
- Cluster exists AND another session owns it → STOP. Do not proceed. Inform the user:
A Kind cluster is already running (likely from another session). Options: (a) wait for that session to finish, (b) switch to
rca:cifor log-only analysis, (c) explicitly destroy the existing cluster first withkind delete cluster --name rossoctl.
To determine ownership: if the current task list or conversation created this cluster, it's yours. Otherwise assume another session owns it.
flowchart TD
START(["/rca:kind"]) --> GUARD{"Kind cluster?"}
GUARD -->|No clusters| CREATE["Deploy Kind cluster"]:::rca
GUARD -->|Owned| REUSE[Reuse existing]:::rca
GUARD -->|Another session| STOP([Stop - cluster busy])
CREATE --> P1["Phase 1: Reproduce"]:::rca
REUSE --> P2["Phase 2: Inspect"]:::rca
P1 --> P2
P2 --> P3["Phase 3: Diagnose"]:::rca
P3 --> P4["Phase 4: Fix and Verify"]:::rca
P4 --> RESULT{"Fixed?"}
RESULT -->|Yes| TDD["tdd:kind"]:::tdd
RESULT -->|No| P2
classDef rca fill:#FF5722,stroke:#333,color:white
classDef tdd fill:#4CAF50,stroke:#333,color:white
Follow this diagram as the workflow.
Workflow
Failure → Reproduce locally → Inspect cluster → Identify root cause → Fix → Verify
Phase 1: Reproduce
Deploy the cluster if not running:
./.github/scripts/local-setup/kind-full-test.sh --skip-cluster-destroy > $LOG_DIR/kind-deploy.log 2>&1; echo "EXIT:$?"
Phase 2: Inspect
All kubectl output redirected to files:
kubectl get pods -n rossoctl-system > $LOG_DIR/pods-system.log 2>&1 && echo "OK" || echo "FAIL"
kubectl get pods -n rossoctl-system --field-selector=status.phase!=Running > $LOG_DIR/failed-pods.log 2>&1 && echo "OK" || echo "FAIL"
kubectl get events -n rossoctl-system --sort-by='.lastTimestamp' > $LOG_DIR/events.log 2>&1 && echo "OK" || echo "FAIL"
Use Task(subagent_type='Explore') with Grep to analyze the logs for errors.
Phase 3: Diagnose
kubectl logs -n rossoctl-system deployment/rossoctl-ui --tail=50 > $LOG_DIR/ui.log 2>&1 && echo "OK" || echo "FAIL"
kubectl logs -n rossoctl-system deployment/ollama --tail=50 > $LOG_DIR/ollama.log 2>&1 && echo "OK" || echo "FAIL"
kubectl get pods -n team1 > $LOG_DIR/pods-team1.log 2>&1 && echo "OK" || echo "FAIL"
kubectl logs -n team1 deployment/weather-service --tail=50 > $LOG_DIR/agent.log 2>&1 && echo "OK" || echo "FAIL"
Use Task(subagent_type='Explore') with Grep to find errors in the log files.
Phase 4: Fix and Verify
After fixing, re-run the specific failing test:
uv run pytest rossoctl/tests/e2e/ -v -k "test_name" > $LOG_DIR/retest.log 2>&1; echo "EXIT:$?"
CVE Check Before Publishing Findings
Before posting RCA findings to any public destination:
If the root cause involves a dependency bug or version issue:
- Invoke
cve:scanto check if this is a known CVE - If a CVE is found → invoke
cve:brainstormBEFORE documenting publicly - Use neutral language in all public documentation
Kind-Specific Issues
| Issue | Cause | Fix |
|---|---|---|
| Ollama OOM | Model too large for Kind | Use smaller model or increase Docker memory |
| DNS resolution | CoreDNS not ready | Wait or restart CoreDNS pod |
| Port conflicts | 8080 already in use | lsof -i :8080 and kill process |
| Image pull errors | Local registry not configured | Check kind-registry container is running |
Escalation
If the issue can't be reproduced locally, escalate:
- CI-only failure → Use
rca:cito analyze from logs - Need real OpenShift → Use
rca:hypershiftwith a live cluster
Related Skills
rca:ci- RCA from CI logs onlyrca:hypershift- RCA with live HyperShift clustertdd:kind- TDD workflow on Kindkind:cluster- Create/destroy Kind clustersk8s:pods- Debug pod issuesrossoctl:ui-debug- Debug UI issues (502, API, proxy)cve:scan- CVE scanning (check if root cause is a known CVE)cve:brainstorm- Disclosure planning (if CVE found during RCA)
Signals
- GitHub stars
- 300
- Forks
- 107
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Gateway key
rca-kind- Source
- github.com/rossoctl/rossoctl