Interpret profiler output and optimize hot paths
SkillDev toolsAnalyze cProfile or py-spy tables to optimize real bottlenecks without code guessing.
Use Interpret profiler output and optimize hot paths in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add Interpret profiler output and optimize hot paths and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the Interpret profiler output and optimize hot paths skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; ahel provides instructions and does not run this skill.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
What this skill tells your AI
The instructions your AI receives, as published by 00200200/maintainer-skills-lab in skills/mkl-profile-performance/SKILL.md and read by ahel’s review.
Do not guess performance bottlenecks by reading code. Human intuition about execution time is notoriously unreliable in dynamic languages like Python. Never rewrite algorithms, add complex caches, or micro-optimize functions without empirical profiler evidence establishing that the target function is a dominant contributor to overall runtime.
Capture empirical profiler data
Establish a reproducible, deterministic workload that exercises the suspected performance regression or slow operation. Capture profile data using Python's built-in cProfile or sampling profilers (py-spy, scalene).
To capture the top CPU consumers sorted by internal function time:
python3 -m cProfile -s tottime script.py | head -n 25
To capture functions by cumulative time (useful for locating high-level subsystems):
python3 -m cProfile -s cumtime script.py | head -n 25
Record the complete environment: Python version, OS, CPU architecture, library versions, dataset size, and iteration count.
Interpret the profiler table columns
A standard cProfile output presents five critical numeric columns:
ncalls: The number of times the function was called. If two numbers appear (e.g.1000/10), the function is recursive; the first is total calls, the second is primitive calls.tottime: Total time spent in the given function's own body, excluding all time spent in calls to subfunctions. This is the primary indicator of CPU bottlenecks.percall(first): Average time spent per call inside the function body (tottime / ncalls).cumtime: Cumulative time spent in this function and all called subfunctions. Top-level runners (such asmainorrun_pipeline) will naturally have largecumtime.percall(second): Average cumulative time per call (cumtime / ncalls).filename:lineno(function): The exact code site. Distinguish application code from third-party libraries and Python built-ins.
Distinguish hot bodies from orchestrators
- High
tottime, moderate/lowcumtime: The function itself executes slow Python bytecode (e.g., tight loops, string concatenations, quadratic lookups). This is an immediate optimization candidate. - Low
tottime, highcumtime: The function is merely an orchestrator or wrapper delegating to subroutines. Do not optimize this function body; drill down into its subcalls to locate where the cumulative time is actually consumed. - High
ncallswith tinypercall: Function overhead multiplied across millions of iterations. Common causes include unmemoized helper calculations, repeated regex compilation, redundant object instantiation, or repeated attribute lookups inside nested loops.
Focus strictly on the top-3 bottlenecks
Follow Amdahl's Law: optimizing a function that accounts for 2% of total runtime can yield at most a 2% overall improvement even if optimized to zero seconds.
- Rank all application functions by
tottime. - Select only the top 1 to 3 functions that account for the majority of execution time.
- Ignore functions below the top tier. Reject PRs and changes that add caching, cythonization, or algorithmic complexity to non-bottleneck functions.
Select targeted remedies
Match the diagnosed symptom with the appropriate optimization:
- Repeated idempotent calculation with high
ncalls: Applyfunctools.lru_cacheor precompute a lookup table outside the loop. - Repeated membership tests in lists (
item in sequence): Convert lists tosetordictto reduce lookup from $O(N)$ to $O(1)$. - Repeated string building (
s += chunk): Replace with list collection and"".join(chunks). - Heavy element-wise math in nested Python loops: Replace with vectorized NumPy, PyTorch, or Polars operations to move computation into compiled C/CUDA kernels.
- Repeated regex compilation inside loops: Hoist pattern compilation to module level via
re.compile(). - JSON serialization/deserialization hot paths: Switch to faster native parsers (
orjson) or defer serialization until transmission.
Worked Example: Hot path diagnosis and verification
Baseline Profiler Output
ncalls tottime percall cumtime percall filename:lineno(function)
1 0.002 0.002 4.821 4.821 pipeline.py:1(run)
1000000 3.210 0.000 3.210 0.000 processor.py:42(is_valid_token)
1000000 1.120 0.000 1.120 0.000 {method 'match' of 're.Pattern'}
5000 0.410 0.000 0.480 0.000 processor.py:88(lookup_metadata)
1 0.079 0.079 4.821 4.821 main.py:1(<module>)
Diagnosis:
is_valid_token consumes 3.21s of 4.82s (66.6% of runtime) with 1,000,000 invocations. Inspecting processor.py:42 reveals an internal check: token in self.blacklist where self.blacklist was initialized as a list containing 500 items. Searching a 500-element list $10^6$ times performs $5 \times 10^8$ equality checks.
Remedy:
Convert self.blacklist from list to set at class initialization.
Verified Profiler Output
ncalls tottime percall cumtime percall filename:lineno(function)
1 0.002 0.002 1.682 1.682 pipeline.py:1(run)
1000000 1.115 0.000 1.115 0.000 {method 'match' of 're.Pattern'}
5000 0.408 0.000 0.478 0.000 processor.py:88(lookup_metadata)
1000000 0.082 0.000 0.082 0.000 processor.py:42(is_valid_token)
1 0.075 0.075 1.682 1.682 main.py:1(<module>)
Outcome:
is_valid_token dropped from 3.210s to 0.082s (39x function speedup). Overall execution dropped from 4.821s to 1.682s (2.86x end-to-end speedup).
Verification and Pull Request checklist
Before committing performance improvements:
- Re-run the profiler command under the exact same dataset and hardware conditions.
- Confirm that total elapsed time and target function
tottimedecreased measurably. - Run the full unit and regression test suite to ensure functional correctness was not compromised.
- Include both before and after profiler summaries in the pull request description.
Signals
- GitHub stars
- 49
- Forks
- 2
- Last commit
- Sep 2026
Advanced
- Item type
- skill
- Key
mkl-profile-performance- Source
- github.com/00200200/maintainer-skills-lab
github.com/00200200/maintainer-skills-lab
Related picks
Skill · wshobson
The pick for Pythonpython-pro
Skill · jeffallan
The pick for Pythonhandsontable-playwright-e2e
Skill · handsontable
The pick for End-to-end testingmstar-e2e
Skill · btspoony
The pick for End-to-end testingteach
Skill · mattpocock
More in Dev toolsimplement
Skill · mattpocock
More in Dev tools