k8s Cluster Operations — joelclaw on Talos
SkillCloud & infraOperate the joelclaw Kubernetes cluster, Talos Linux on Colima (Mac Mini). Deploy services, check health, debug pods, recover from restarts, add ports, manage Helm releases, inspect logs, fix networking. Triggers on: 'kubectl', 'pods', 'deploy to k8s', 'cluster health', 'restart pod', 'helm install', 'talosctl', 'colima', 'nodeport', 'flannel', 'port mapping', 'k8s down', 'cluster not working', 'add a port', 'PVC', 'storage', any k8s/Talos/Colima infrastructure task. Also triggers on service-specific deploy: 'deploy redis', 'redeploy inngest', 'livekit helm', 'pds not responding'.
Available today. Use it from your connected AI after setup.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
Then ask your AI: use the k8s Cluster Operations skill
What this skill tells your AI
The instructions your AI receives, as published by joelhooks/joelclaw in skills/k8s/SKILL.md and read by ahel’s review.
Architecture
Mac Mini (localhost ports)
└─ Colima/Lima port forwarding (grpc for host-published service ports; avoid separate persistent autossh tunnels)
└─ Colima VM (8 CPU, 24 GiB, 100 GiB, VZ framework, aarch64)
└─ Docker 29.x + buildx (joelclaw-builder, docker-container driver)
└─ Talos v1.12.4 container (joelclaw-controlplane-1, 18 GiB cap)
└─ k8s v1.35.0 (single node, Flannel CNI)
└─ joelclaw namespace (privileged PSA)
⚠️ Talos has NO shell. No bash, no /bin/sh, nothing. You cannot docker exec into the Talos container. Use talosctl for node operations and the Colima VM (ssh lima-colima) for host-level operations like modprobe.
Colima Stability Rules (2026-03-17 incident)
| Setting | Value | Reason |
|---|---|---|
| CPU | 8 | Match k8s workload requests (~2.8 CPU, 72%) |
| Memory | 24 GiB | Current post-reboot profile; 16 GiB left too little headroom once the Talos container cap is raised. Re-evaluate if macOS memory pressure returns. |
| nestedVirtualization | OFF by default | Crashes VM under load (image builds, heavy scheduling). Toggle ON only for Firecracker testing |
| vmType | vz | Required for Apple Silicon |
| mountType | virtiofs | Fastest option with VZ |
nestedVirtualization: true is unstable on M4 Pro under load. It causes the Colima VM to silently crash during Docker builds/pushes. Each crash:
- Kills the Talos container mid-operation
- Corrupts Redis AOF (if caught mid-write) → crash-loop on restart
- Breaks Lima socket forwarding →
dockerCLI on macOS disconnects - Creates stale k8s pods that re-pull images → amplifies pressure
Recovery from Colima crash-loop:
colima stop && colima start— basic restart- If Redis crash-loops:
redis-check-aof --fix(see Redis AOF Recovery below) - If Restate has stuck invocations: purge PVC or kill via admin API
- If native Docker socket dead: use SSH tunnel
ssh -L /tmp/docker.sock:/var/run/docker.sock
Docker image builds should use the buildx container builder (docker buildx build --builder joelclaw-builder) to isolate build IO from k8s workloads.
Redis AOF Recovery
If Redis crash-loops after a VM restart with Bad file format reading the append only file:
# 1. Scale down Redis (or use a temp pod if StatefulSet can't mount PVC concurrently)
kubectl -n joelclaw apply -f - <<'EOF'
apiVersion: v1
kind: Pod
metadata:
name: redis-fix
namespace: joelclaw
spec:
tolerations:
- key: node-role.kubernetes.io/control-plane
operator: Exists
effect: NoSchedule
containers:
- name: fix
image: redis:7-alpine
command: ["sh", "-c", "cd /data/appendonlydir && echo y | redis-check-aof --fix *.incr.aof && redis-check-aof *.incr.aof"]
volumeMounts:
- name: data
mountPath: /data
restartPolicy: Never
volumes:
- name: data
persistentVolumeClaim:
claimName: data-redis-0
EOF
# 2. Wait, check logs, then clean up
kubectl -n joelclaw logs redis-fix
kubectl -n joelclaw delete pod redis-fix --force
# 3. Restart Redis
kubectl -n joelclaw delete pod redis-0
For port mappings, recovery procedures, and cluster recreation steps, read references/operations.md.
Reboot-heal persistence rule (2026-04-15 incident)
infra/k8s-reboot-heal.sh runs under launchd as a fresh process every interval. Any recovery marker that only lives in shell memory dies at the end of that tick.
That means flannel/event-healing state must be persisted on disk. Canonical path:
~/.local/state/k8s-reboot-heal.env
Persist at least:
COLIMA_START_EPOCHRECOVERY_START_EPOCHLAST_FLANNEL_RESTART_EPOCHCOLIMA_UNHEALTHY_STREAKLAST_COLIMA_UNHEALTHY_EPOCHLAST_COLIMA_FORCE_CYCLE_EPOCHLAST_COLIMA_FAILED_RECOVERY_EPOCH
Why this matters: kubelet FailedCreatePodSandBox events mentioning missing subnet.env can stay recent for minutes after the first repair. If the healer forgets that it already restarted flannel, the next launchd tick can bounce flannel again and knock healthy services like Typesense back into 503 warmup for no good reason. The extra failed-recovery marker also stops the system from counting a one-minute green flash as success and then force-cycling Colima again when the control path collapses.
Kubeconfig / Talos Endpoint Contract (2026-05-30 incident)
Current operator access uses Colima/Lima's directly published TCP ports:
- Kubernetes API:
https://127.0.0.1:6443 - Talos API:
127.0.0.1:50000
Do not rewrite kubeconfig/Talos back to the older manual tunnel ports (16443 / 15000) unless explicitly testing KUBE_OPERATOR_MODE=ssh. After the 2026-05-30 reboot, grpc publishing kept service ports open but could leave Colima SSH and the Docker socket dead after load; a live ssh port-forwarder test did the opposite and dropped the service ports. The current availability-first default remains Colima portForwarder: grpc, but completion audits must verify both service ports and Docker/limactl instead of trusting either mode by vibes.
Fix/verify:
talosctl config endpoint 127.0.0.1:50000
talosctl config node 10.5.0.2
kubectl config set-cluster joelclaw --server=https://127.0.0.1:6443 --insecure-skip-tls-verify=true
kubectl config use-context admin@joelclaw
kubectl get nodes --request-timeout=10s
talosctl -e 127.0.0.1:50000 -n 10.5.0.2 version --client=false
infra/kube-operator-access.sh now defaults to KUBE_OPERATOR_MODE=direct and runs as a boring launchd monitor that keeps configs pointed at 6443 / 50000. It still has an ssh mode for deliberate fallback testing.
Durable recovery rule (ADR-0244)
A Colima restart is not recovery.
After any colima start / force-cycle, the system only counts recovery as real if a post-restart stability window stays healthy across repeated passes for:
- Colima SSH
- Docker socket
- Kubernetes API
- Typesense localhost health
- Inngest localhost health
If those regress during the verification window, classify the event as a failed recovery, capture proof artifacts, and stop repeated force-cycles for the configured hold period. The point is durability, not healer theatre.
Quick Health Check
kubectl get pods -n joelclaw # all pods
curl -s localhost:3111/api/inngest # system-bus-worker → 200
curl -s localhost:7880/ # LiveKit → "OK"
curl -s localhost:8108/health # Typesense → {"ok":true}
curl -s localhost:8288/health # Inngest → {"status":200}
curl -s localhost:9070/deployments # Restate admin → deployments list
curl -s localhost:9627/xrpc/_health # PDS → {"version":"..."}
kubectl exec -n joelclaw redis-0 -- redis-cli ping # → PONG
joelclaw restate cron status # Dkron scheduler → healthy via temporary CLI tunnel
Services
| Service | Type | Pod | Ports (Mac→NodePort) | Helm? |
|---|---|---|---|---|
| Redis | StatefulSet | redis-0 | 6379→6379 | No |
| Typesense | StatefulSet | typesense-0 | 8108→8108 | No |
| Inngest | StatefulSet | inngest-0 | 8288→8288, 8289→8289 | No |
| Restate | StatefulSet | restate-0 | 8080→8080, 9070→9070, 9071→9071 | No |
| system-bus-worker | Deployment | system-bus-worker-* | 3111→3111 | No |
| restate-worker | Deployment | restate-worker-* | in-cluster only (restate-worker:9080) | No |
| docs-api | Deployment | docs-api-* | 3838→3838 | No |
| LiveKit | Deployment | livekit-server-* | 7880→7880, 7881→7881 | Yes (livekit/livekit-server 1.9.0) |
| PDS | Deployment | bluesky-pds-* | 9627→3000 | Yes (nerkho/bluesky-pds 0.4.2) |
| MinIO | StatefulSet | minio-0 | 30900→30900, 30901→30901 | No |
| Dkron | StatefulSet | dkron-0 | in-cluster only (dkron-svc:8080) | No |
AIStor Operator (aistor ns) | Deployments | adminjob-operator, object-store-operator | n/a | Yes (minio/aistor-operator) |
AIStor ObjectStore (aistor ns) | StatefulSet | aistor-s3-pool-0-0 | 31000 (S3 TLS), 31001 (console) | Yes (minio/aistor-objectstore) |
Restate / Firecracker runtime notes
deployment/restate-workeris intentionally privileged and mounts/dev/kvm(hostPath type""— optional).- PVC
firecracker-imagesat/tmp/firecracker-teststores kernel, rootfs, and snapshot artifacts. - When
nestedVirtualizationis OFF:/dev/kvmabsent,microvmDAG handler fails, butshell/infer/noophandlers work normally. - When
nestedVirtualizationis ON: Firecracker one-shot exec works (create workspace ext4 → write command → boot VM → guest executes → poweroff → read results). - Restate retry caps: dagWorker maxAttempts=5, dagOrchestrator maxAttempts=3. Prevents journal poisoning.
- Stuck Restate invocations: Inspect and target the affected invocation through supported APIs. Journal or PVC destruction requires separate explicit data-loss authorization and a verified backup; never use it as a routine stuck-job fix.
- Re-register worker:
curl -X POST http://localhost:9070/deployments -H 'content-type: application/json' -d '{"uri":"http://restate-worker:9080"}'
⚠️ PDS port trap: Docker maps 9627→3000 (host→container). NodePort must be 3000 to match the container-side port. If set to 9627, traffic won't route.
Rule: NodePort value = Docker's container-side port, not host-side.
Agent Runner (Cold k8s Jobs)
Status: local sandbox remains the default/live path; the k8s backend is now code-landed and opt-in, but still needs supervised rollout before calling it earned runtime.
The agent runner executes sandboxed story runs as isolated k8s Jobs. Jobs are created dynamically via @joelclaw/agent-execution/job-spec — no static manifests.
Runtime Image Contract
See k8s/agent-runner.yaml for the full specification.
Required components:
- Git (checkout, diff, commit)
- Bun runtime
- runner-installed agent tooling (currently
claudeand/or other installed CLIs) /workspaceworking directory- runtime entrypoint at
/app/packages/agent-execution/src/job-runner.ts
Configuration via environment variables:
- Request metadata:
WORKFLOW_ID,REQUEST_ID,STORY_ID,SANDBOX_PROFILE,BASE_SHA,EXECUTION_BACKEND,JOB_NAME,JOB_NAMESPACE - Repo materialization:
REPO_URL,REPO_BRANCH, optionalHOST_REQUESTED_CWD - Agent identity:
AGENT_NAME,AGENT_MODEL,AGENT_VARIANT,AGENT_PROGRAM - Execution config:
SESSION_ID,TIMEOUT_SECONDS - Task prompt:
TASK_PROMPT_B64(base64-encoded) - Verification:
VERIFICATION_COMMANDS_B64(base64-encoded JSON array) - Callback path:
RESULT_CALLBACK_URL,RESULT_CALLBACK_TOKEN
Expected behavior:
- Decode task from
TASK_PROMPT_B64 - Materialize repo from
REPO_URL/REPO_BRANCHatBASE_SHA - Execute the requested
AGENT_PROGRAM - Run verification commands (if set)
- Print
SandboxExecutionResultmarkers to stdout and POST the same result to/internal/agent-result - Exit 0 (success) or non-zero (failure)
Current truthful limit:
piremains local-backend only for now; do not pretend the pod runner can execute pi story runs yet.
Job Lifecycle
import { generateJobSpec, generateJobDeletion } from "@joelclaw/agent-execution";
// 1. Generate Job spec
const spec = generateJobSpec(request, {
runtime: {
image: "ghcr.io/joelhooks/agent-runner:latest",
imagePullPolicy: "Always",
command: ["bun", "run", "/app/packages/agent-execution/src/job-runner.ts"],
},
namespace: "joelclaw",
imagePullSecret: "ghcr-pull",
resultCallbackUrl: "http://host.docker.internal:3111/internal/agent-result",
resultCallbackToken: process.env.OTEL_EMIT_TOKEN,
});
// 2. Apply to cluster (via kubectl or k8s client library)
// 3. Job runs → Pod materializes repo, executes agent, posts SandboxExecutionResult callback
// 4. Host worker can recover the same terminal result from log markers if callback delivery fails
// 5. Job auto-deletes after TTL (default: 5 minutes)
// Cancel a running Job
const deletion = generateJobDeletion("req-xyz");
// kubectl delete job ${deletion.name} -n ${deletion.namespace}
Resource Defaults
- CPU:
500mrequest,2limit - Memory:
1Girequest,4Gilimit - Active deadline:
1 hour - TTL after completion:
5 minutes - Backoff limit:
0(no retries)
Security
- Non-root execution (UID 1000, GID 1000)
- No privilege escalation
- All capabilities dropped
- RuntimeDefault seccomp profile
- Control plane toleration for single-node cluster
Verification Commands
# List agent runner Jobs
kubectl get jobs -n joelclaw -l app.kubernetes.io/name=agent-runner
# Check Job status
kubectl describe job <job-name> -n joelclaw
# View logs
kubectl logs job/<job-name> -n joelclaw
# Check for stale Jobs (should be auto-deleted by TTL)
kubectl get jobs -n joelclaw --show-all
Current State
- ✅ Job spec generator (
packages/agent-execution/src/job-spec.ts) - ✅ Runtime contract (
k8s/agent-runner.yaml) - ✅ Tests (
packages/agent-execution/__tests__/job-spec.test.ts) - ⏳ Runtime image not yet built (Story 3)
- ⏳ Hot-image CronJob not yet implemented (Story 4)
- ⏳ Warm-pool scheduler not yet implemented (Story 5)
- ⏳ Restate integration not yet wired (Story 6)
NAS NFS Access from k8s (ADR-0088 Phase 2.5)
k8s pods can mount NAS storage over NFS via a LAN route through the Colima bridge.
How it works
k8s pod → Talos container (10.5.0.x) → Docker NAT → Colima VM
→ ip route 192.168.1.0/24 via 192.168.64.1 dev col0
→ macOS host (IP forwarding enabled) → LAN → NAS (192.168.1.163)
Root cause of prior failures: VZ framework's shared networking on eth0 doesn't properly forward LAN-bound traffic. The fix routes LAN traffic through col0 (Colima bridge → macOS host) instead.
Route persistence
The LAN route is set in two places for reliability:
- Colima provision script (
~/.colima/default/colima.yaml) — runs oncolima start(cold boot) - k8s-reboot-heal (
~/Code/joelhooks/joelclaw/infra/k8s-reboot-heal.sh) — reasserts the route during reboot recovery ticks
Both execute: ip route replace 192.168.1.0/24 via 192.168.64.1 dev col0
Duplicate tunnel ownership is a bug (2026-04-16)
com.joel.colima-tunnel is deprecated. Colima/Lima already forwards the docker-published host ports for joelclaw-controlplane-1, so a second autossh daemon on those same ports is not redundancy — it's interference.
Rules:
com.joel.colimais the only boot/start helper for the VM; it must not keep a periodicStartIntervalcom.joel.colima-tunnelshould be absent from/Library/LaunchDaemons/;install-critical-launchdaemons.shremoves it instead of reinstalling it- do not run a second autossh daemon on ports Colima/Lima already publishes for
joelclaw-controlplane-1(3838,6379,7880,7881,8108,8288,8289,9627,64784) - do not kill generic
sshlisteners on those host ports; that can kill Lima's own forwarders infra/colima-tunnel.shis now only a deprecated compatibility stub so stale launchd installs exit cleanly instead of fighting Limacom.joel.kube-operator-accessis now a direct-mode monitor, not a persistent tunnel. It keeps kubectl/talosctl aimed at Colima/Lima's published loopback ports:6443for kube-apiserver and50000for Talos- the old SSH tunnel mode (
16443 -> 10.5.0.2:6443,15000 -> 10.5.0.2:50000) is fallback-only viaKUBE_OPERATOR_MODE=ssh; do not leave it crash-looping under launchd - once the daemon is installed, kubectl should use
https://127.0.0.1:6443and talosctl should use127.0.0.1:50000 com.joel.k8s-reboot-healmust use the same JSON status check; a plaincolima statusfalse-negative can force-cycle the VM and retrigger the flannel/NAS failure cascade during reboot recovery- do not trust status output alone when deciding to cycle Colima; if the Docker socket or Colima SSH path is still healthy, treat the VM as alive and keep your hands off it
- a Colima force-cycle now requires confirmed evidence; one ugly observation is not enough to panic-cycle the VM
- confirmation can come from consecutive launchd ticks or from a short rapid-confirmation window when both the Docker socket and Colima SSH path stay down long enough to prove a severe collapse
- after any Colima force-cycle, honor the persisted cooldown in
~/.local/state/k8s-reboot-heal.envso Talos and workload warmup can finish before another escalation is even considered - if the host path is still down but escalation is not yet earned, bail out early and mark the tick failed; do not pretend downstream kube/NAS repair steps are actionable without Colima host access
- reboot recovery is not healthy until the NAS route
192.168.1.0/24 via 192.168.64.1 dev col0exists again and NFS is reachable from the Colima VM - flannel can be "Running" while kubelet still reports
failed to load flannel 'subnet.env' file; treat recentFailedCreatePodSandBoxevents with that message as a restart signal for the flannel pod
Available PVs
| PV | NFS Path | Capacity | Access | Use |
|---|---|---|---|---|
nas-nvme | 192.168.1.163:/volume2/data | 1.5TB | RWX | NVMe RAID1: backups, snapshots, models, sessions |
nas-hdd | 192.168.1.163:/volume1/joelclaw | 50TB | RWX | HDD RAID5: books, docs-artifacts, archives, otel |
minio-nfs-pv | 192.168.1.163:/volume1/joelclaw | 1TB | RWO | HDD tier: MinIO object storage (same export) |
Mounting NAS in a pod
volumes:
- name: nas
persistentVolumeClaim:
claimName: nas-nvme
containers:
- volumeMounts:
- name: nas
mountPath: /nas
# Optional: subPath for specific dir
subPath: typesense
Rules
- Always use IP (192.168.1.163), never hostname (three-body). DNS doesn't resolve from inside k8s.
- Always use
nfsvers=3,tcp,resvport,noatimemount options. NFSv4 has issues with Asustor ADM. - NAS unavailability degrades gracefully with
softmount option — returns errors, doesn't hang pods. - NFS write performance: ~660 MiB/s over 10GbE with jumbo frames. Good for sequential I/O (backups, snapshots). Latency-sensitive workloads (Redis, active Typesense indexes) stay on local SSD.
- If NFS mount fails after Colima restart: verify the route exists:
colima ssh -- ip route | grep 192.168.1.0
Verify connectivity
# From Colima VM
colima ssh -- timeout 2 bash -c "echo > /dev/tcp/192.168.1.163/2049" && echo "NFS OK"
# From k8s pod
kubectl run nfs-test --image=busybox --restart=Never -n joelclaw \
--overrides='{"spec":{"tolerations":[{"key":"node-role.kubernetes.io/control-plane","operator":"Exists","effect":"NoSchedule"}],"containers":[{"name":"t","image":"busybox","command":["sh","-c","ls /nas && echo OK"],"volumeMounts":[{"name":"n","mountPath":"/nas"}]}],"volumes":[{"name":"n","persistentVolumeClaim":{"claimName":"nas-nvme"}}]}}'
kubectl logs nfs-test -n joelclaw && kubectl delete pod nfs-test -n joelclaw --force
Deploy Commands
# Manifests (redis, typesense, inngest, dkron)
kubectl apply -f ~/Code/joelhooks/joelclaw/k8s/
# Restate runtime
kubectl apply -f ~/Code/joelhooks/joelclaw/k8s/restate.yaml
kubectl apply -f ~/Code/joelhooks/joelclaw/k8s/firecracker-pvc.yaml
kubectl rollout status statefulset/restate -n joelclaw
~/Code/joelhooks/joelclaw/k8s/publish-restate-worker.sh
curl -fsS http://localhost:9070/deployments
# Dkron phase-1 scheduler (ClusterIP API + CLI-managed short-lived tunnel access)
kubectl apply -f ~/Code/joelhooks/joelclaw/k8s/dkron.yaml
kubectl rollout status statefulset/dkron -n joelclaw
joelclaw restate cron status
joelclaw restate cron sync-tier1 # seed/update ADR-0216 tier-1 jobs
# system-bus worker (build + push GHCR + apply + rollout wait)
~/Code/joelhooks/joelclaw/k8s/publish-system-bus-worker.sh
# LiveKit (Helm + reconcile patches)
~/Code/joelhooks/joelclaw/k8s/reconcile-livekit.sh joelclaw
# AIStor (Helm operator + objectstore)
# Defaults to isolated `aistor` namespace to avoid service-name collisions with legacy `joelclaw/minio`.
# Cutover override (explicit only): AISTOR_OBJECTSTORE_NAMESPACE=joelclaw AISTOR_ALLOW_JOELCLAW_NAMESPACE=true
~/Code/joelhooks/joelclaw/k8s/reconcile-aistor.sh
# PDS (Helm) — always patch NodePort to 3000
# (export current values first if the release already exists)
helm get values bluesky-pds -n joelclaw > /tmp/pds-values-live.yaml 2>/dev/null || true
helm upgrade --install bluesky-pds nerkho/bluesky-pds \
-n joelclaw -f /tmp/pds-values-live.yaml
kubectl patch svc bluesky-pds -n joelclaw --type='json' \
-p='[{"op":"replace","path":"/spec/ports/0/nodePort","value":3000}]'
Auto Deploy (GitHub Actions)
- Workflow:
.github/workflows/system-bus-worker-deploy.yml - Trigger: push to
maintouchingpackages/system-bus/**or worker deploy files - Behavior:
- builds/pushes
ghcr.io/joelhooks/system-bus-worker:${GITHUB_SHA}+:latest - runs deploy job on
self-hostedrunner - updates k8s deployment image + waits for rollout + probes worker health
- builds/pushes
- If deploy job is queued forever, check that a
self-hostedrunner is online on the Mac Mini.
GHCR push 403 Forbidden
Cause: GITHUB_TOKEN (default Actions token) does not have packages:write scope for this repo. A dedicated PAT is required.
Fix already applied: Workflow uses secrets.GHCR_PAT (not secrets.GITHUB_TOKEN) for the GHCR login step. The PAT is stored in:
- GitHub repo secrets as
GHCR_PAT(set via GitHub UI) - agent-secrets as
ghcr_pat(secrets lease ghcr_pat)
If this breaks again: PAT may have expired. Regenerate at github.com → Settings → Developer settings → PATs, update both stores.
Local fallback (bypass GHA entirely):
DOCKER_CONFIG_DIR=$(mktemp -d)
echo '{"credsStore":""}' > "$DOCKER_CONFIG_DIR/config.json"
export DOCKER_CONFIG="$DOCKER_CONFIG_DIR"
secrets lease ghcr_pat | docker login ghcr.io -u joelhooks --password-stdin
~/Code/joelhooks/joelclaw/k8s/publish-system-bus-worker.sh
Note: publish-system-bus-worker.sh uses gh auth token internally — if gh auth is stale, use the Docker login above before running the script, or patch it to use secrets lease ghcr_pat directly.
Resilience Rules (ADR-0148)
Shortened here. Read the whole file on GitHub.
Signals
- GitHub stars
- 64
- Forks
- 2
- Last commit
- Sep 2026
Advanced
- Item type
- skill
- Key
k8s-joelhooks- Source
- github.com/joelhooks/joelclaw