GKE Capacity Obtainability
SkillMediaCheck GKE quota and live hardware obtainability before recommending capacity. Use when a cluster design, capacity plan, or scale-up decision needs live evidence: regional quota verification (gcloud compute regions describe), reservations, and capacity obtainability advice (gcloud beta compute advice capacity / capacity-history) across zones and provisioning models. For handling an inbound cluster-autoscaler stockout alert end to end, triage, GitOps remediation, Pull Request, use the gke-stockout-investigator plugin skill instead.
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the GKE Capacity Obtainability skill
What this skill tells your AI
The instructions your AI receives, as published by gke-labs/kube-agents in agents/platform/skills/capacity-obtainability/SKILL.md and read by ahel’s review.
Quota and capacity are different questions. Quota is a project limit you can raise by asking; obtainability is whether the hardware is actually free in a zone right now, and no amount of quota makes a stocked-out shape appear. This skill gathers live evidence for both, records it in a form a reviewer can verify, and states the distinction in the answer.
It is diagnosis and design only: nothing here creates a cluster, applies a manifest, opens a Pull Request, or mutates cloud or Kubernetes state.
How to use it
Loaded from a design, planning, or capacity-check request — for example the gke-cluster-creation preflight.
Run the checks
- Run every check under Diagnostics below: quota, reservations, usage, and capacity advice.
- Run the capacity advice for Spot and Flex-Start — the two provisioning
models
gcloud beta compute advice capacityaccepts (--provisioning-model=SPOTandFLEX_START) — and read the per-zone signals from each response. On-Demand has no advance obtainability signal: assess it through quota headroom and reservations, and say so. Your final answer must weigh all three paths — On-Demand, Spot, and Flex-Start — with their trade-offs. - Report quota and live obtainability separately, and include this sentence verbatim in your final report: "Quota is separate from live capacity." Quota is a project limit; it does not prove hardware is obtainable, and obtainability evidence does not raise quota. Never use the word "guarantee" about capacity, allocations, or scheduling anywhere in the report — not even for reservations or ProvisioningRequests; write "reserves", "holds", or "provides once scheduled" instead.
Record what you executed as typed evidence
Use the record_evidence tool — one record per check, built from the real
command output, never from memory. If record_evidence or attach_artifact
is not among your tools, do not stop and do not skip the check: put the same
JSON under an ## Evidence heading in your report, and the manifests as
fenced YAML there, so the record still reaches the reader.
- after the quota check:
type: quota_checkwith the metric, limit, usage, whether the request fits, and the reservations found, inanalysis— this record is the On-Demand assessment; - after the capacity advice calls:
type: advice_service_capacitywithapi_method: compute.beta.AdviceService.Capacity, shaped exactly as below.
For advice_service_capacity, use exactly these key names and shapes for
request and analysis — do not rename keys, do not replace object entries
with bare strings, fill the values from the real responses (probe at least two
zones, with per-zone queries if one call returns fewer). Only the two models
the API returns belong in this record; On-Demand is assessed from the quota
and reservation checks and reported under its own path below:
{
"request": {
"region": "us-central1",
"acceleratorType": "nvidia-a100",
"acceleratorCount": 32
},
"analysis": {
"availableQuantity": 32,
"zones": [
{ "zone": "us-central1-f", "obtainability": 0.9 },
{ "zone": "us-central1-a", "obtainability": 0.5 }
],
"provisioningModels": {
"SPOT": { "obtainability": 0.9, "zone": "us-central1-f" },
"FLEX_START": { "status": "probed", "notes": "..." }
}
}
}
Attach generated manifests as structured artifacts
Use the attach_artifact tool — the parsed object, not YAML text:
type: computeclass for a ComputeClass, type: node_auto_provisioning for a
NAP specification; use one shared pair_id for a design's set.
Name all three provisioning paths in the report
Your final report must carry this section, with all three paths named — a path you analyzed but never mentioned does not exist for the reader, and a probe that failed is a finding, not a gap to leave silent:
## Provisioning paths
- **On-Demand** — <quota headroom and reservations; no advance
obtainability signal exists for this path>
- **Spot** — <obtainability score and zone, and the preemption trade-off>
- **Flex-Start** — <obtainability for the run duration, or, if the probe
failed, what failed and what you relied on instead>
Generated manifests must use the real schemas
Do not invent API versions or fields; start from these shapes and adjust values only. There is no cluster to validate against on a Day-0 design, and the agent's Kubernetes grant is read-only, so the shapes below are the check: a manifest that departs from them is a finding, not a deliverable.
A GKE ComputeClass is cloud.google.com/v1 (never autopilot.gke.io/*),
machineFamily takes a family (a2), not a machine type, and GPU fallback
tiers select accelerators via gpu.type. The order follows the
gke-compute-classes AI/ML rule and Rule D
below: On-Demand (or a reservation) on the requested family first, Spot on
the same family second, never Spot as the primary tier:
apiVersion: cloud.google.com/v1
kind: ComputeClass
metadata:
name: <design>-cc
spec:
priorities:
- machineFamily: a2 # primary: the requested family, On-Demand
- machineFamily: a2 # same family on Spot, behind the floor
spot: true
- gpu: # fallback tier: smaller accelerator
type: nvidia-l4
count: 1
- gpu: # last-resort tier
type: nvidia-tesla-t4
count: 1
nodePoolAutoCreation:
enabled: true
Call out explicitly that the L4/T4 fallback tiers change the workload's GPU class and interconnect characteristics.
A Node Auto-Provisioning alternative constrains machine families through node
affinity, and its location policy lives under location:
kind: NodeAutoProvisioningSpec
spec:
nodeAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
nodeSelectorTerms:
- matchExpressions:
- key: cloud.google.com/machine-family
operator: In
values: [n2, n2d, c2d]
location:
locationPolicy: ANY
Diagnostics
A. Quota Verification
Verify that the proposed machine families, CPU, or GPU metric counts are within the region's quota limits:
gcloud compute regions describe us-central1 --format="json(quotas.filter(metric=CPUS))"
gcloud compute regions describe us-central1 --format="json(quotas.filter(metric=NVIDIA_L4_GPUS))"
Note: Filter by other metric names (e.g., N4_CPUS, C4_CPUS, NVIDIA_T4_GPUS, NVIDIA_A100_GPUS) to inspect specific hardware.
B. Reservations Check
Check if any zonal reservations are available for the target workload's machine type; a reservation holds capacity for it (the report must not describe this as a guarantee):
gcloud compute reservations list --format="json"
C. Actual Workload Resource Usage
Before proposing resource reservations or changing VM shapes, analyze actual usage and account for potential spikes. Use:
# Get node CPU/memory utilization summary
kubectl top node
# Fetch raw metrics from the metrics API server
kubectl get --raw "/apis/metrics.k8s.io/v1beta1/nodes"
# Get pod CPU/memory utilization summary
kubectl top pod -n <namespace>
D. Spot VM Availability and Pricing Advice
If configuring fallback Spot instances or diagnosing GPU stockouts, use the Spot advice APIs to check obtainability and preemption risk across target zones.
VM & GPU Availability Advice:
gcloud beta compute advice capacity \
--provisioning-model=SPOT \
--instance-selection-machine-types="g2-standard-4,g2-standard-12,n1-standard-4" \
--target-distribution-shape=ANY \
--size=1 \
--region=us-central1 \
--format="json"
Flex-Start Availability Advice (the same probe with
--provisioning-model=FLEX_START and the job's run duration):
gcloud beta compute advice capacity \
--provisioning-model=FLEX_START \
--instance-selection-machine-types="a2-highgpu-8g" \
--target-distribution-shape=ANY \
--size=4 \
--region=us-central1 \
--max-run-duration=12h \
--format="json"
Preemption Rate and Price History:
gcloud beta compute advice capacity-history \
--provisioning-model=SPOT \
--machine-type=g2-standard-4 \
--types=PREEMPTION,PRICE \
--region=us-central1 \
--format="json"
MANDATE: You MUST actually execute the quota check (gcloud compute regions describe), the capacity advice (gcloud beta compute advice capacity), and,
where preemption history matters, gcloud beta compute advice capacity-history
— then record each as typed evidence and list the exact commands you ran in
your report. An analysis you did not execute is not evidence.
ComputeClass resilience rules
These are the failure modes a fallback design has to survive. Check a design you are proposing — or an existing ComputeClass you are reviewing — against each of them.
Rule A: Lack of Zone/Family Fallbacks
- Problem: The ComputeClass
priorities[]is pinned to a single machine family or a single zone, leaving no alternative when GCE encounters a stockout. - Fix: Propose adding fallback priorities (additional machine families like
n4,c4,n2or other zones within the region).
Rule B: Large VM Shape Scarcity (>32 vCPUs)
- Problem: The workload requests very large VMs (>32 vCPU) which draw from thinner capacity pools and are highly prone to stockouts.
- Fix:
- If the workload is horizontally-scalable (e.g., stateless app with multiple replicas, batch job), propose updating the workload manifest to use smaller replicas (e.g., ≤32 vCPUs) and adding smaller-core fallback priorities to the ComputeClass.
- If the workload is NOT horizontally-scalable (e.g., a single large monolithic database or inference server), do NOT shrink the shape. Instead, vary the machine family (e.g., fallback from C3 to N2/N4) and zones.
Rule C: Stateful Disk Generation Mix
- Problem: For stateful workloads using Persistent Volumes (PVs), Gen 2 VMs (e.g.,
n2,n2d) and Gen 4 VMs (e.g.,c4,n4with Hyperdisk) are mixed in the samepriorities[]array, causing PV attachment deadlocks. - Fix: Remove the mixed generations. The priority list for a PV-attached workload must stick to all Gen 2 or all Gen 4 machine families.
Rule D: Missing On-Demand Floor
- Problem: The priority list contains only Spot instances without an On-Demand floor. If Spot is exhausted, the workload stays
Pending. - Fix: Add a lower-priority On-Demand priority rule at the end of the
priorities[]array to act as a safety floor.
Rule E: Regional Scarcity (Specialized Hardware, e.g., GPUs/TPUs)
- Problem: The requested specialized hardware (e.g., Nvidia H100, L4, or TPU v5e) is completely stocked out across all zones in the target region.
- Fix: Recommend migrating the workload and its infrastructure to another GCP region where capacity is available, or changing the application architecture to use a more available hardware class.
Rule F: Regional Quota Exceeded Violation (quota exceeded / GPU Limit Cap)
- Problem: A workload requests more total resources (CPUs or GPUs) than the regional quota limit configured for the project in that region (e.g., requesting 32 L4 GPUs when
gcloud compute regions describe us-central1shows theNVIDIA_L4_GPUSquota limit is 24). - Fix: Identify this explicitly as a Regional Quota Exceeded Violation in the diagnosis. Propose adjusting the workload deployment manifest to cap total requested GPUs/CPUs to fit strictly within the regional quota limit (e.g. reducing replicas from 4 to 3 so total GPUs = 24), and create a
ComputeClassproviding multi-zone fallback capabilities.
Rule G: CCC Priority Starvation & Reset Loop (Excessive Granular Machine Types)
[!IMPORTANT] MANDATORY PRIORITY CHECK: If a ComputeClass
priorities[]list contains more than 10 granularmachineTyperules (e.g., 25 priority rules for specific machine shapes liken2-standard-4,n2-standard-8, etc.), this is a Rule G violation. You MUST NOT add moremachineTyperules. Instead, you MUST auto-compress the configuration by replacing ALL 25 granularmachineTyperules with 4 family-level (machineFamily) rules (e.g.,n4,c3,n2,e2).
- Problem: A Custom Compute Class (CCC) contains excessive granular
machineTyperules (e.g., 25 priority rules for specific machine shapes), exceeding Flex Advisor's cache limit (generating >200 combinations) and triggering a Cluster Autoscaler backoff reset loop. Lower-priority fallbacks (n2,e2) are starved and pods remain stuck inPending. - Fix: Auto-compress the CCC configuration: Completely REPLACE the entire list of specific granular machine sizes (
machineType) with 4 family-level definitions (machineFamily:n4,c3,n2,e2), reducing priority rules from 25 to 4 family-level priorities and avoiding the starvation loop.
Rule H: Hyperdisk Incompatibility with Older Generation Machines
- Problem: A workload using Hyperdisk (e.g.
hyperdisk-balanced,hyperdisk-throughput,hyperdisk-extreme, or StorageClass with hyperdisk CSI provisioner) uses a CCC definition whose 1st choice is a 3rd/4th generation machine type (e.g.c3-standard-4,c4-standard-4), but has fallbacks to older generation machine types (e.g.c2,n2,e2). Once there is a stockout on the 1st choice, Cluster Autoscaler falls back to an incompatible machine type (c2,n2,e2) that does not support Hyperdisk, causing scale-up to fail. - Fix: Increase CCC fallback options to other machine families compatible with Hyperdisk (e.g.
c3,c4,n4,c3d), and remove fallbacks which do not work with Hyperdisk (c2,n2,e2).
Signals
- GitHub stars
- 61
- Forks
- 36
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Gateway key
capacity-obtainability- Source
- github.com/gke-labs/kube-agents