Splunk Observability Cisco AI Pod Integration (Umbrella)
SkillCloud & infra"Use when deploying Splunk Observability Cloud for a Cisco AI Pod with UCS, Nexus, NVIDIA GPUs, NIM/vLLM
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the Splunk Observability Cisco AI Pod Integration (Umbrella) skill
What this skill tells your AI
The instructions your AI receives, as published by chambear2809/splunk-cisco-skills in skills/splunk-observability-cisco-ai-pod-integration/SKILL.md and read by ahel’s review.
Prerequisites
| Tool or access | Purpose | Verify |
|---|---|---|
| Bash and Python 3 | Run bundled setup and validation helpers | bash --version && python3 --version |
| Required product/platform access | Inspect or configure the selected target | Complete the documented preflight |
| Credential files for live modes | Keep secrets out of chat | Verify paths only |
Workflow Overview
┌───────────┐ ┌───────────────┐ ┌───────────────┐ ┌─────────────────┐
│ Preflight │ → │ Render/review │ → │ Apply/handoff │ → │ Validate evidence │
└───────────┘ └───────────────┘ └───────────────┘ └─────────────────┘
When to Activate
- Deploying Splunk Observability Cloud for a Cisco AI Pod with UCS, Nexus, NVIDIA GPUs, NIM/vLLM inference, and storage telemetry. Hand off base collector, HEC, dashboards, and detectors to the owning skills.
- Preview and review the splunk observability cisco ai pod integration workflow before any live apply phase.
- Diagnose failed prerequisites, generated assets, configuration, or validation evidence.
Scope
Follow the documented read-only or render-first path whenever it is available. This skill does not imply permission to mutate live systems. Require explicit apply flags, protected credentials, and operator review for state changes.
Examples
Inspect the supported setup modes before selecting one:
bash skills/splunk-observability-cisco-ai-pod-integration/scripts/setup.sh --help
Expected output: usage, supported modes, and required arguments are displayed without changing the target environment.
Inspect validation modes before running completion checks:
bash skills/splunk-observability-cisco-ai-pod-integration/scripts/validate.sh --help
Expected output: offline, live, and completion options are displayed when the skill supports them; help exits without mutation.
Troubleshooting
| Issue | Cause | Resolution |
|---|---|---|
| Preflight fails | A required tool or access path is missing | Resolve it before rendering or applying |
| Rendered assets are incomplete | Required non-secret inputs are absent | Complete intake and render again |
| Apply is blocked | Review, credentials, or explicit acceptance is missing | Use the documented handoff |
| Validation is incomplete | Live evidence is unavailable | Record the gap and keep completion open |
This is the AI Pod umbrella that ties together every component skill needed for end-to-end Cisco AI Pod observability in Splunk Observability Cloud. It composes:
- splunk-observability-cisco-nexus-integration for Cisco Nexus 9000 fabric metrics (cisco_os receiver).
- splunk-observability-cisco-intersight-integration for Cisco UCS metrics via Intersight OTel deployment.
- splunk-observability-nvidia-gpu-integration for NVIDIA GPU telemetry via DCGM Exporter.
And adds AI-Pod-specific bits documented in the configuration guide and production-validated by the atl-ocp2 OpenShift cluster:
- NIM scrapes (multi-job: llm/embedqa/rerankqa, port 8000
/v1/metrics). - vLLM scrape (port 8000
/metrics). - Milvus vector DB scrape (port 9091).
- NetApp Trident storage scrape (port 8001
/metrics). - Pure Portworx storage scrape (ports 17001 + 17018).
- Redfish exporter (user-supplied) on port 9210.
- Cisco AI PODs Splunk Observability dashboard pipeline (
metrics/cisco-ai-pods, unfiltered). - NIM dashboard pipeline (
metrics/nvidianim-metrics, unfiltered). k8s_attributes/nimprocessor forapp -> model_nameextraction.- OpenShift SCC helper,
workshop/multi-tenant.sh, and dual-pipeline filtering pattern.
Critical production lessons encoded
These are the silent failure traps the umbrella prevents:
- RBAC gap: base chart's ClusterRole grants only
podsandservices. Anykubernetes_sd_configs.role: endpointsscrape (e.g. NIM in endpoint mode) silently fails withendpoints is forbidden. The umbrella emits therbac.customRulesblock withendpoints+discovery.k8s.io/endpointslicesget/list/watch when needed. - receiver_creator naming:
receiver_creator/dcgm-cisco, NOTreceiver_creator/nvidia(collision with chart autodetect). Inherited from the GPU child skill. - DCGM dual-label discovery: matches both
appandapp.kubernetes.io/name. Inherited from the GPU child skill. - Dual-pipeline filtering: filtered standard pipeline + unfiltered specialized pipelines for AI Pod dashboards. Smarter than the canonical single-pipeline pattern.
- OpenShift defaults:
kubeletstats.insecure_skip_verify: true(REQUIRED),certmanager.enabled: false,cloudProvider: "". - Existing collector apply: use
--apply-existing-collectorwhen a Splunk OTel Collector is already running. This path renders the overlay, reads current Helm values without persisting the token, removes stalereceiver_creator/nvidia, wiresotlpinto the metrics pipeline for Intersight, applies via Helm, restarts the existing collector agent, restarts Intersight, and runs live validation. - Helm token pattern: apply scripts use a file-backed token (
--set-file splunkObservability.accessToken=...) so the token is never written to a tracked values file or temporary values file.
Composition model
When you run --render, the umbrella:
- Invokes each child skill's renderer to produce its overlay under a sub-directory.
- Merges the child overlays into a unified
splunk-otel-overlay/values.overlay.yamlwith the renderer's deterministic Python deep-merge. The rendered base-collector handoff usesyqlater to merge that reviewed overlay with base collector values. - Adds AI-Pod-specific blocks on top of the merged overlay.
- Renders unified handoff scripts.
When you run --apply-existing-collector, the umbrella applies its rendered overlay to the already running Splunk OTel Collector Helm release instead of standing up a second collector.
What it renders (composed + AI-Pod-specific)
splunk-otel-overlay/values.overlay.yaml— composed overlay (Nexus + Intersight + GPU children + AI-Pod additions).child-renders/<skill>/— each child skill's full rendered output (preserved for debugging the merge).intersight-integration/— from the Intersight child.secrets/cisco-nexus-ssh-secret.yaml— from the Nexus child.dcgm-pod-labels-patch/— from the GPU child when--enable-dcgm-pod-labels.- NIM, vLLM, Milvus, Trident, Portworx, and Redfish scrape configuration embedded in
splunk-otel-overlay/values.overlay.yaml. openshift/scc.sh— OpenShift SCC helper script.workshop/multi-tenant.sh— Workshop multi-tenant deploy script (when--workshop-mode).dashboards/— AI-Pod-specific dashboards for NIM/vLLM inference, Milvus, and Trident/Portworx storage.detectors/— AI-Pod-specific detectors (vLLM error rate, NIM TTFT regression, Milvus query latency, Portworx node offline, Trident volume allocation).scripts/handoff-base-collector.sh— emits the base collector + merge command with--distribution openshift(default).scripts/handoff-hec-token.sh— for K8s container log shipping to Splunk Platform.scripts/handoff-dashboards.sh,handoff-detectors.sh— emit reviewed dashboard and detector commands across all four skills.scripts/explain-composition.sh— prints the per-child contribution summary.metadata.json.
Safety Rules
- File-backed token flags only:
--o11y-token-file(Splunk Observability Org access token; passed through to all child skills + base collector).--platform-hec-token-file(optional; for K8s container logs to Splunk Platform).--intersight-key-id-fileand--intersight-key-file(passed through to the Intersight child).
- Reject every direct token / key flag.
- Token files must be
chmod 600;--allow-loose-token-permsoverrides with WARN. - Cisco Nexus SSH credentials handled by the Nexus child (K8s Secret stub; user creates the Secret out-of-band).
Primary Workflow
-
Confirm prerequisites are installed: NVIDIA GPU Operator (or standalone DCGM Exporter), NIM/vLLM with the standard pod labels, Milvus, NetApp Trident, Pure Portworx, Redfish exporter, Cisco Intersight account + API key.
-
Render the composed overlay:
bash skills/splunk-observability-cisco-ai-pod-integration/scripts/setup.sh \ --render --validate \ --realm us0 \ --cluster-name atl-ai-pod \ --distribution openshift \ --nim-scrape-mode endpoints \ --enable-dcgm-pod-labels \ --output-dir splunk-observability-cisco-ai-pod-rendered -
If a Splunk OTel Collector is already running, apply the overlay in place and run live validation:
bash skills/splunk-observability-cisco-ai-pod-integration/scripts/setup.sh \ --render --apply-existing-collector --validate --live \ --realm us0 \ --cluster-name atl-ai-pod \ --distribution openshift \ --collector-release splunk-otel-collector \ --collector-namespace splunk-otel \ --o11y-token-file /path/to/o11y-token \ --output-dir splunk-observability-cisco-ai-pod-rendered -
For greenfield installs, apply child manifests (Intersight, optional DCGM patch) + merge overlay + apply via base collector:
bash splunk-observability-cisco-ai-pod-rendered/scripts/handoff-base-collector.sh bash splunk-observability-cisco-ai-pod-rendered/scripts/handoff-dashboards.sh bash splunk-observability-cisco-ai-pod-rendered/scripts/handoff-detectors.sh
Hand-offs
- Splunk OTel Collector base install: splunk-observability-otel-collector-setup with
--distribution openshift(default; configurable). - HEC for K8s container logs: splunk-hec-service-setup.
- Dashboards: splunk-observability-dashboard-builder.
- Detectors: splunk-observability-native-ops.
- Component skills (composed): Nexus / Intersight / GPU child skills.
Out of scope
- All children's out-of-scope items (NVIDIA GPU Operator install, DCGM Exporter install, NIM/vLLM/Milvus/Trident/Portworx/Redfish exporter deployment, OpenShift cluster bootstrap, Cisco Intersight account creation).
Validation
bash skills/splunk-observability-cisco-ai-pod-integration/scripts/validate.sh
Runs each child skill's validate.sh recursively, then checks the composed overlay, endpoint-discovery RBAC, OpenShift kubelet settings, and rendered secret safety. With --live, it probes collector and Intersight resources and logs plus the live collector ConfigMap. The umbrella validator does not make direct SignalFlow API probes for NIM, Milvus, or vLLM metrics.
With --live, validation prefers oc, falls back to kubectl, passes --live through to child validators, and fails on Intersight OTLP export errors such as unknown service opentelemetry.proto.collector.metrics.v1.MetricsService.
See reference.md and references/composition-and-overlay-merge.md, nim-vllm-scrape-catalog.md, milvus-storage-redfish.md, openshift-scc.md, workshop-multi-tenant.md, ai-pod-dashboards-catalog.md, endpoints-rbac-patch.md, dual-pipeline-filtering.md, nim-scrape-modes.md, production-troubleshooting-atl-ocp2.md, troubleshooting.md for the full annexes.
Signals
- GitHub stars
- 37
- Forks
- 8
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Gateway key
splunk-observability-cisco-ai-pod-integration- Source
- github.com/chambear2809/splunk-cisco-skills