Performance Operations

SkillAI & models

Performance profiling and optimization orchestrator - diagnoses symptoms, dispatches skill-preloaded profiling agents, manages before/after comparisons. Triggers on: performance, profiling, flamegraph, pprof, py-spy, clinic.js, memray, heaptrack, bundle size, webpack analyzer, load testing, k6, artillery, vegeta, locust, benchmark, hyperfine, criterion, slow query, EXPLAIN ANALYZE, N+1, caching, optimization, latency, throughput, p99, memory leak, CPU spike, bottleneck.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the Performance Operations skill

What this skill tells your AI

The instructions your AI receives, as published by 0xdarkmatter/claude-mods in skills/perf-ops/SKILL.md and read by ahel’s review.

Orchestrator for cross-language performance profiling and optimization. Classifies symptoms inline, dispatches profiling to general-purpose agents preloaded with the relevant language -ops skill (background), and manages optimization with confirmation.

Architecture

User describes performance issue or requests profiling
    |
    +---> T1: Diagnose (inline, fast)
    |       +---> Classify symptom (decision tree)
    |       +---> Detect language/runtime from project
    |       +---> Check installed profiling tools
    |       +---> Determine production vs development
    |       +---> Gather system baseline (CPU/mem/disk)
    |       +---> Present: diagnosis + recommended profiling approach
    |
    +---> T2: Profile (dispatch general-purpose + skill preload, background)
    |       +---> Select skill preload from routing table
    |       +---> Build perf-focused dispatch prompt
    |       +---> Agent runs profiler, collects data, interprets results
    |       |       +---> Fallback: tool commands inlined (no skill preload)
    |       +---> Returns: findings + bottleneck identification + suggestions
    |       |
    |       +---> [Optional parallel dispatch]:
    |             +---> CPU profiling agent  ---+
    |             +---> Memory profiling agent --+--> Consolidate findings
    |             +---> Baseline benchmark ------+
    |
    +---> T3: Optimize (dispatch general-purpose + skill preload, foreground + confirm)
            +---> Agent proposes specific code changes
            +---> Preflight: what changes, expected impact, risks
            +---> User confirms
            +---> Apply changes
            +---> Re-benchmark for before/after delta

Safety Tiers

T1: Diagnose - Run Inline

No agent needed. Execute directly via Bash for instant results.

OperationCommand / Method
Detect Python profilerswhich py-spy && which memray && which scalene
Detect Go profilerswhich go && go tool pprof -h 2>/dev/null
Detect Rust profilerswhich cargo-flamegraph && which samply
Detect Node profilerswhich clinic && which 0x
Detect benchmarking toolswhich hyperfine && which k6 && which vegeta
System CPU baselinetop -bn1 -o %CPU | head -20 (Linux) or wmic cpu get loadpercentage (Win)
System memory baselinefree -h (Linux) or wmic OS get FreePhysicalMemory (Win)
Disk I/O checkiostat -x 1 3 (Linux)
Identify languageCheck for package.json, go.mod, Cargo.toml, pyproject.toml, requirements.txt
Production vs devAsk user or detect from environment (NODE_ENV, FLASK_ENV, etc.)
Read existing profilesParse .prof, .svg, .bin files in project

Production safety rule: In production environments, only recommend sampling profilers (py-spy, pprof HTTP endpoint, perf). Never suggest attaching debuggers, tracing profilers, or tools that require process restart.

T2: Profile - Dispatch to Profiling Agent

Gather context from T1 diagnosis, then dispatch a general-purpose agent preloaded with the relevant language -ops skill plus the perf-ops references below.

Language Routing:

Detected LanguageDispatchPreloadKey Profiling Tools
Python (.py, pyproject.toml, requirements.txt)general-purposerelevant skills/python-*/SKILL.md + perf-ops referencespy-spy, memray, scalene, tracemalloc
Go (go.mod, .go files)general-purposeskills/go-ops/SKILL.md + perf-ops referencespprof (CPU/heap/goroutine/mutex), benchstat
Rust (Cargo.toml, .rs files)general-purposeskills/rust-ops/SKILL.md + perf-ops referencescargo-flamegraph, samply, DHAT, criterion
TypeScript/JavaScript (backend, package.json + server)general-purposeskills/javascript-ops/SKILL.md + perf-ops referencesclinic flame/doctor/bubbleprof, 0x
TypeScript/JavaScript (frontend, bundle issues)general-purposeskills/typescript-ops/SKILL.md + perf-ops referenceswebpack-bundle-analyzer, Lighthouse, source-map-explorer
SQL / PostgreSQLgeneral-purposeskills/postgres-ops/SKILL.md + perf-ops referencesEXPLAIN ANALYZE, pg_stat_statements, pgbench
SQL / SQLite, Cloudflare D1, libSQL/Turso (*.db, *.sqlite, wrangler.toml with a d1_databases binding)general-purposeskills/sqlite-ops/SKILL.md + perf-ops referencesEXPLAIN QUERY PLAN, sqlite-ops/scripts/eqp-triage.py, sqlite3 .timer/.stats, wrangler d1 insights, sql_duration_ms + rows_read
General / unknown / CLI benchmarkinggeneral-purposeperf-ops referenceshyperfine, perf, strace

Dispatch template (T2):

You are handling a performance profiling task dispatched by the perf-ops orchestrator.

## Diagnosis (from T1)
- Symptom: {classified symptom from decision tree}
- Language/Runtime: {detected language}
- Environment: {production | development}
- Installed tools: {list from tool detection}
- System baseline: {CPU/memory/disk metrics}

## Profiling Task
{specific profiling request - e.g., "CPU profile the API server under load"}

## Target
- Process/file: {target application or endpoint}
- Expected workload: {how to generate representative load if needed}

## Domain Knowledge
Before starting, read the relevant profiling reference for this language:
- Read: skills/perf-ops/references/cpu-memory-profiling.md

For load testing tasks, also read:
- Read: skills/perf-ops/references/load-testing.md

For database profiling, also read:
- Read: skills/postgres-ops/SKILL.md (if PostgreSQL)

## Instructions
1. Run the appropriate profiler for this language and symptom
2. Collect sufficient samples (minimum 30 seconds for CPU, multiple snapshots for memory)
3. Interpret the results - identify the top 3-5 bottlenecks
4. For each bottleneck: explain what it is, why it's slow, and suggest a fix
5. Report findings in structured format with metrics

Execution mode:

ScenarioModeWhy
User waiting for resultsrun_in_background=FalseThey need findings before continuing
User continuing other workrun_in_background=TrueDon't block the main session
Quick benchmark (hyperfine)run_in_background=FalseFast enough to wait
Load test (k6, artillery)run_in_background=TrueTakes minutes

T3: Optimize - Preflight Required

Dispatch a general-purpose agent (preloaded per the language routing table) with explicit instruction to produce a preflight report before any code changes.

Dispatch template (T3 preflight):

You are handling a performance optimization dispatched by the perf-ops orchestrator.

## Profiling Results (from T2)
{bottleneck findings, metrics, flamegraph interpretation}

## Optimization Request
{specific optimization - e.g., "Fix the N+1 query in UserController.list"}

IMPORTANT: Do NOT apply changes yet. Produce a Preflight Report:
1. Exactly what code/config changes you will make
2. Expected performance improvement (with reasoning)
3. Risks (correctness, side effects, edge cases)
4. How to verify the improvement (specific benchmark or test)
5. How to revert if the optimization causes issues

After user confirms: Re-dispatch with execute authority plus the before/after protocol.

Dispatch template (T3 execute + before/after):

User confirmed the optimization. Proceed with execution.

## Approved Changes
{exact changes from preflight report}

## Before/After Protocol
1. Record the current benchmark baseline: {specific command from T2}
2. Apply the approved changes
3. Run the same benchmark again
4. Report comparison:
   - Metric: before value -> after value (% change)
   - Include statistical confidence if tool supports it
5. If regression detected: revert and report

Parallel Profiling

When multiple independent symptoms are detected, or the user requests comprehensive profiling, dispatch parallel agents.

Parallelizable combinations:

Agent 1Agent 2Why Independent
CPU profilerMemory profilerDifferent tools, different data
CPU profilerBaseline benchmarkRead vs measurement
Backend profilerFrontend bundle analysisDifferent runtimes
Service A profilerService B profilerDifferent processes

NOT parallelizable:

Operation AOperation BWhy Sequential
ProfileInterpret resultsDependency
Before benchmarkAfter benchmarkRequires code change between
Load testCPU profile same processTool interference

Dispatch pattern for parallel profiling:

# Example: CPU + memory profiling in parallel
Agent(
    subagent_type="general-purpose",
    model="sonnet",
    run_in_background=True,
    prompt="First read skills/perf-ops/references/cpu-memory-profiling.md "
           "and the relevant language skill (e.g. skills/python-pytest-ops/SKILL.md). "
           "Then: CPU profiling task: {cpu_prompt}"
)
Agent(
    subagent_type="general-purpose",
    model="sonnet",
    run_in_background=True,
    prompt="First read skills/perf-ops/references/cpu-memory-profiling.md "
           "and the relevant language skill (e.g. skills/python-pytest-ops/SKILL.md). "
           "Then: Memory profiling task: {memory_prompt}"
)
# Both run simultaneously, consolidate findings when both complete

Fallback: When No Language Skill Matches

If no language -ops skill covers the target, dispatch general-purpose with profiling commands inlined instead of a skill preload.

Agent(
    subagent_type="general-purpose",
    model="sonnet",
    run_in_background=True,
    prompt="""You are acting as a performance profiling agent for {language}.

Use these specific tools and commands:
{tool commands from diagnosis-quickref.md for the detected language}

{original dispatch prompt}
"""
)

For simple benchmarks (hyperfine, single command timing), skip agent dispatch entirely and run inline via Bash.

Decision Logic

When a performance-related request arrives:

1. Classify the request:
   - Symptom description? -> Start at T1 (diagnose)
   - "Profile my app"? -> T1 (detect language + tools) then T2 (profile)
   - "Benchmark X vs Y"? -> T2 directly (hyperfine or language benchmark)
   - "Optimize this"? -> T2 (profile first) then T3 (optimize)
   - "Why is X slow"? -> T1 (diagnose) then T2 (targeted profile)

2. T1 Diagnose (always runs first for new issues):
   - Detect language/runtime
   - Check installed profiling tools
   - Classify symptom using decision tree (see diagnosis-quickref.md)
   - Determine production vs development
   - Present findings + recommend next step

3. T2 Profile (when diagnosis points to a specific bottleneck):
   - Route per the language routing table (general-purpose + skill preload)
   - Decide foreground vs background
   - Consider parallel dispatch if multiple symptoms
   - Consolidate findings from all agents

4. T3 Optimize (only when user wants changes applied):
   - Always produce preflight report first
   - Wait for explicit user confirmation
   - Execute with before/after comparison
   - Report delta with statistical confidence

Quick Reference

TaskTierExecution
Detect toolsT1Inline
Check system metricsT1Inline
Classify symptomT1Inline
Identify languageT1Inline
Run CPU profilerT2Agent (bg)
Run memory profilerT2Agent (bg)
Run load testT2Agent (bg)
Run benchmarkT2Agent (bg or inline for hyperfine)
Bundle analysisT2Agent (bg)
EXPLAIN ANALYZET2Agent (fg)
Before/after comparisonT2Agent (fg)
Apply optimizationT3Agent + confirm
Add indexT3Agent + confirm
Refactor hot pathT3Agent + confirm

Reference Files

FileContents
references/diagnosis-quickref.mdDecision tree, tool selection matrix, quick references for all profiling domains, common gotchas
references/cpu-memory-profiling.mdDeep flamegraph interpretation, language-specific CPU/memory profiling guides
references/load-testing.mdk6, Artillery, vegeta, wrk, Locust methodology and CI integration
references/optimization-patterns.mdCaching, database, frontend, API, concurrency, memory optimization strategies
references/ci-integration.mdPerformance budgets, regression detection, CI pipeline patterns, benchmark baselines

Load reference files when deeper tool-specific guidance is needed beyond what the dispatch prompt provides.

See Also

SkillWhen to Combine
debug-opsRoot cause analysis for performance regressions
monitoring-opsProduction metrics, alerting on latency/throughput
testing-opsPerformance regression tests in CI, benchmark suites
code-statsIdentify complex code that may be performance-sensitive
postgres-opsPostgreSQL-specific query optimization, indexing, EXPLAIN
sqlite-opsSQLite/D1/libSQL query plans, covering indexes, rows-read economics, eqp-triage.py
container-orchestrationResource limits, pod scaling, container performance

Signals

GitHub stars
36
Forks
5
Last commit
Aug 2026
Advanced
Catalog kind
skill
Gateway key
perf-ops
Source
github.com/0xdarkmatter/claude-mods