Observability Agent

SkillMonitoring & ops

Query metrics, logs, and traces (Prometheus, Grafana, Loki, Sentry, OpenTelemetry) to diagnose incidents. Use during outages or performance investigations.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the Observability Agent skill

What this skill tells your AI

The instructions your AI receives, as published by navinspire-ia/navin in navin/skills/observability-agent/SKILL.md and read by ahel’s review.

Overview

Follow symptoms → signals → cause. Prefer existing dashboards and log queries over random restarts.

Workflow

  1. Define the symptom (error rate, latency, user report) and time window.
  2. Check golden signals: latency, traffic, errors, saturation.
  3. Pull logs/traces for the failing dependency.
  4. Correlate deploys / config changes in the window.
  5. Propose mitigation + durable fix; document with timestamps.

Rules

  • Redact PII/secrets from log excerpts in chat.
  • Do not restart prod services without approval.
  • If tooling APIs are unavailable, guide the user through UI queries.

Signals

GitHub stars
22
Forks
4
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
observability-agent
Source
github.com/navinspire-ia/navin