Smart Contract Security Audit

SkillSecurity

Audit your Solidity smart contracts for security bugs while you develop.

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the Smart Contract Security Audit skill

About this skill

Security audit of Solidity code while you develop. Trigger on "audit", "check this contract", "review for security", "loop mode", "run the auditor in loop mode", "run 3 passes". Modes - default (full repo) or a specific filename. Loop mode runs several passes in one scan, each pass told what the ear

What this skill tells your AI

The instructions your AI receives, as published by pashov/skills in solidity-auditor/SKILL.md and read by ahel’s review.

You are the orchestrator of a parallelized smart contract security audit.

Mode Selection

Exclude pattern: skip dependency and build directories — node_modules/, lib/, artifacts/, cache/, out/, broadcast/, coverage/, typechain*/ — and non-production code — interfaces/, mocks/, test/ and files matching *.t.sol, *Test*.sol or *Mock*.sol.

Deploy scripts stay in scope. script/, deploy/ and *.s.sol are audited like any other code: a deploy script sets constructor arguments, hands over ownership and seeds state, and it carries real bugs. Excluding a dependency is an argument about code you did not write; a deploy script is code you did write.

  • Default (no arguments): scan all .sol files using the exclude pattern. Use Bash find (not Glob), exactly this command:

    find . -type f -name '*.sol' \
      -not -path '*/node_modules/*' -not -path '*/lib/*' -not -path '*/artifacts/*' \
      -not -path '*/cache/*' -not -path '*/out/*' -not -path '*/broadcast/*' \
      -not -path '*/coverage/*' -not -path '*/typechain*/*' \
      -not -path '*/interfaces/*' -not -path '*/mocks/*' -not -path '*/test/*' \
      -not -name '*.t.sol' -not -name '*Test*.sol' -not -name '*Mock*.sol'
    

    -type f is required, not tidiness: Hardhat writes each build artifact into a directory named GatewayCrossChain.sol/, so a find without it matches directories as if they were source files and hands them to cat. Do not drop it, and do not re-derive this command — paste it.

  • $filename ...: scan the specified file(s) only. The exclude pattern does not apply here. A file named on the command line is always scanned, wherever it lives — naming node_modules/@openzeppelin/contracts/token/ERC20/ERC20.sol scans that file. The pattern chooses what a default scan discovers; it never overrides an explicit request.

Flags:

  • --file-output (off by default): copy the assembled report into the working directory (name per {resolved_path}/report-formatting.md). The flag never causes a report to be produced — every scan assembles .solidity-auditor/runs/{stamp}/full-report.md whether it is passed or not. It only decides whether a copy lands where the runner can see it.

    Line 33 used to read "Never write a report file unless explicitly passed", and that is now false. The rule was written to protect the runner's working directory, and that protection is unchanged: without the flag, nothing is written outside .solidity-auditor/. What changed is that the report is assembled by shell on every scan, so the flag can no longer collapse the report — nothing regenerates, it copies. A later editor must not read this as a mistake and revert it.

  • --memory (off by default): remember findings between scans in a ledger at .solidity-auditor/memory.tsv in the audited repo. Any pass count above 1 turns memory on by itself, whether or not the flag was passed.

  • --loop [N] (off by default): run N passes in one scan, each pass told what the earlier ones found, ending in one combined report. --loop with no number is 3 passes. --loop 1 is a 1-pass scan. The flag exists for runners who prefer flags; when it is passed the picker in Turn 1b does not ask, it obeys.

Vocabulary (used throughout this file):

  • run — one pass of the 12 agents.
  • scan — one invocation of this skill. A scan holds 1 or more runs.

The ledger's scans column counts scans. The report's seen in k/N runs counts runs inside one scan and is never written to the ledger. Two numbers, two names — do not mix them.

The pass count the scan runs is {passes} — settled in Turn 1b, 1 or more. The loop body is Turn 2 step 2c, Turn 2 step 3, Turn 3a, Turn 3b and Turn 4, and it runs once per pass. Everything before it runs once per scan, and Turn 5 closes the scan once, whatever {passes} is.

HARD RULE — the plain path must stay plain. A 1-pass scan with no flags reaches none of the memory steps: not Turn 1c, not Turn 2 step 2, not Turn 4 step 4, not Turn 4 step 6. It reads no ledger, writes no file anywhere, and prints exactly the report it printed before memory existed. Memory is on only when --memory was passed or the pass count is above 1. A later editor who is tempted to make any memory step unconditional is breaking this on purpose, not tidying up.

A 1-pass answer reaches none of the loop or memory machinery. No ledger is read or written. No Passes or Memory row in the Scope table. No seen in k/N runs. No KNOWN / NEW tag. No per-pass summary lines. The printed report is the report this skill printed before loop mode existed. The only trace of the picker is the question itself.

What it does write. Every scan writes .solidity-auditor/runs/{stamp}/ — one run-1.md, one scope.tsv, one full-report.md — because the report is assembled from those files at every pass count, and name, mode, files and threshold are needed by every report. This paragraph used to say a 1-pass answer creates no .solidity-auditor/ directory and no runs/ files; that is now false, and it was never what the rule was protecting.

What the rule protects, stated exactly: on the plain path the scan reads no ledger, writes no mem_ key, and prints a Scope table of exactly three rows — Mode, Files reviewed, Confidence threshold (1-100) — and no Passes row, no Memory row. Disk is not printed output. A later editor who makes a memory step unconditional is breaking this on purpose; writing the runs directory is not one of those steps.

--memory on a 1-pass run is the single exception: memory turns on, the loop machinery stays off.

Orchestration

Turn 1 — Discover. Print the banner, then make these parallel tool calls in one message:

a. Bash find for in-scope .sol files per mode selection b. Glob for **/references/hacking-agents/shared-rules.md — extract the references/ directory (two levels up) as {resolved_path} c. ToolSearch select:Agent d. Read the local VERSION file from the same directory as this skill e. Bash curl -sf https://raw.githubusercontent.com/pashov/skills/main/solidity-auditor/VERSION f. Bash mktemp -d ./.audit-XXXXXX → store as {bundle_dir} g. Bash date +%Y%m%d-%H%M%S → store as {stamp}, the scan-time stamp. One stamp per scan, computed once, here. The runs directory, the run files and any --file-output copy all carry it, so a report and the runs that produced it are tied together by eye.

Turn 1a — Open the scan directory. After the find returns, in one Bash command:

mkdir -p .solidity-auditor/runs/{stamp}
: > .solidity-auditor/runs/{stamp}/scope.tsv

Then write the three scope keys this turn knows. A scope key is written with one printf and never any other way:

printf '%s\t%s\n' name  "{project-name}" >> .solidity-auditor/runs/{stamp}/scope.tsv
printf '%s\t%s\n' mode  "{mode}"         >> .solidity-auditor/runs/{stamp}/scope.tsv
printf '%s\t%s\n' files "{file list}"    >> .solidity-auditor/runs/{stamp}/scope.tsv
  • {project-name} — the repo root basename, the same one report-formatting.md names.
  • {mode} — default or filename, as Mode Selection settled it.
  • {file list} — every in-scope path the find returned, space separated on one line, in find order. The assembler wraps them 3 per row; the order it prints is the order written here.

scope.tsv is key<TAB>value, append-only, last line per key wins. An absent key gives an absent table row — that is what keeps the plain scan's Scope table at three rows with no special case. A tab or a newline in a value breaks the row, so values are stripped, not escaped: no key here has any use for either character. Six writers across four turns append to this one file, and none of them ever rewrites or deletes a line.

name and mode are the only two keys a model types. Everything else is either shell knowledge or read by the assembler for itself: the threshold from the constant, N in seen in k/N runs from counting run files, the stamp from the directory's own name.

If the remote VERSION fetch succeeds, compare the two as numbers and warn only when the local one is lower: print ⚠️ You are not using the latest version. Please upgrade for best security coverage. See https://github.com/pashov/skills. If it fails, skip silently.

Lower, not different. A plain "differs" test warns the wrong person: somebody working on an unreleased version has a local VERSION above the published one, and gets told to upgrade to the version they are writing. Local equal to remote, or local above it, prints nothing.

Turn 1b — Model and pass count. This turn asks two questions in one AskUserQuestion call: which model the 12 agents use, and how many passes the scan runs. The runner is interrupted once, before any work starts.

The two questions do not fail the same way. On a runtime without AskUserQuestion and an Agent tool that takes a model parameter — Codex, Gemini, Cursor's native agent — the model question is skipped silently, {agent_model} is left unset, and no prose replaces it. The pass question is not skipped: it falls through to the printed block in Turn 1b-ii, which stops and waits. This turn as a whole is never skipped. A later editor must not restore a blanket "SKIP this turn entirely" rule: it was true when this turn asked one question, and it is false now.

Turn 1b-i — the AskUserQuestion call (Claude Code). Ask both questions in one call. Where the Agent tool takes no model parameter, ask the pass question alone.

Question 1 — model:

  1. Read your system prompt to detect your own model family (Opus, Sonnet, or Haiku). Ignore the version digits — the Agent tool's model parameter takes the family name ("opus" / "sonnet" / "haiku"), and the runtime resolves to the latest version in that family.

  2. Put this question in the call:

    • Question: "Which Claude model should the 12 audit agents use?"
    • Three single-select options. Mark the orchestrator's own family as (Recommended) and place it first.
    • On each option, set the description field to latest.
    • On each option, set the preview field verbatim (preserve all whitespace exactly — the box widths must stay equal across all three):

    Opus preview:

    ┌──────────────────────────────────────────────────────────┐
    │  opus  ·  highest reasoning  ·  most expensive           │
    └──────────────────────────────────────────────────────────┘
    

    Sonnet preview:

    ┌──────────────────────────────────────────────────────────┐
    │  sonnet  ·  balanced reasoning  ·  mid cost              │
    └──────────────────────────────────────────────────────────┘
    

    Haiku preview:

    ┌──────────────────────────────────────────────────────────┐
    │  haiku  ·  lowest reasoning  ·  cheapest                 │
    └──────────────────────────────────────────────────────────┘
    
  3. Store the runner's choice as {agent_model}. If no answer, default to the orchestrator's own model.

Question 2 — pass count. It goes in the same call, second:

  1. Question: "How many passes should this audit run? Each pass is a full 12-agent audit, and every pass after the first is told what the earlier ones found, so it hunts new ground. You get one combined report at the end."

    Three single-select options, 3 passes first and marked (Recommended). Each carries a preview box in this turn's style — the boxes are 60 characters wide, equal to the model picker's, so two questions in one prompt look like one thing. Set preview verbatim, whitespace preserved:

    Labeldescription
    3 passes (Recommended)~45 min
    1 pass~15 min
    5 passes~75 min

    3 passes preview:

    ┌──────────────────────────────────────────────────────────┐
    │  3 passes  ·  each pass hunts new ground  ·  ~45 min     │
    └──────────────────────────────────────────────────────────┘
    

    1 pass preview:

    ┌──────────────────────────────────────────────────────────┐
    │  1 pass  ·  today's audit  ·  ~15 min, nothing written   │
    └──────────────────────────────────────────────────────────┘
    

    5 passes preview:

    ┌──────────────────────────────────────────────────────────┐
    │  5 passes  ·  deepest sweep  ·  ~75 min                  │
    └──────────────────────────────────────────────────────────┘
    

    These are measured, not guessed — and they are a floor. A real 3-pass scan of 2,228 lines of Solidity across 10 files, 12 agents per pass on Opus, took 43 minutes wall clock: 11 minutes for pass 1 and about 16 for each of passes 2 and 3, which run slower because the growing known-findings.md is appended to all twelve bundles. Individual agents ran 3.5–11 minutes. A larger codebase takes longer; a smaller model is faster. Quote minutes rather than multipliers — "~3x time" told the runner nothing about whether to wait or come back after lunch. If these numbers are ever re-measured, correct them here rather than adding a second estimate somewhere else.

  2. Any other number needs no option of its own. AskUserQuestion always adds an Other choice with a free-text box, and the runner types their number there. Do NOT add a fourth option reading "your own number" — options are fixed choices, so it could not collect the number and would dead-end.

    Parse the Other answer for the first integer. Below 1 or above 10 → ask once more. A second unusable answer → 1 pass.

  3. Store the answer as {passes}. No answer at all → 1 pass.

Turn 1b-ii — the printed fallback (every runtime without AskUserQuestion). Print this exactly:

How many passes should this audit run?

Each pass is a full 12-agent audit. Every pass after the first is told what the
earlier passes found, so it hunts new ground. You get one combined report at the end.

  1) 1 pass    — today's audit, about 15 minutes. Nothing is written to disk.
  2) 3 passes  — recommended. About 45 minutes.
  3) 5 passes  — deepest sweep. About 75 minutes.

Answer with 1, 2, 3, or any pass count you want.

STOP here and wait for the runner's answer. Do NOT choose for them. Do NOT continue to Turn 2 with an assumed pass count. Do NOT start the scan and ask later.

This is the one place in this skill where a question is emitted as prose. Turn 1b forbids prose questions because the model picker has a safe default — the orchestrator's own model. A pass count has no safe default: 1 and 5 differ by 5x in time and in cost, and that is the runner's money. A later editor must not "fix" this by deleting the prose block or by picking a default. If you are reading this and it looks like an inconsistency, it is deliberate.

Answers 1, 2 and 3 are the three listed choices; any other integer is that many passes. Apply the same bounds as Turn 1b-i step 5 — below 1 or above 10, ask once more, then 1 pass.

Turn 1b-iii — when the runner already said. Ask nothing that has already been answered:

What arrivedWhat the picker does
--loop 5Skip the pass question, silently. 5 passes.
--loop with no numberSkip the pass question, silently. 3 passes.
"run 4 passes", "audit this four times"Skip the pass question, silently. 4 passes.
"loop mode", "run it a few times" — a request with no numberAsk. They asked for the feature, not for a count.
NothingAsk.

Skipping is silent in the first three rows — printing using 5 passes back at somebody who just typed --loop 5 is noise. Skipping the pass question never skips the model question, and the reverse holds too.

Turn 1b-iv — record the pass count. However {passes} was settled — asked, typed in the prose fallback, or read off a flag — write it once, here:

printf '%s\t%s\n' passes_planned "{passes}" >> .solidity-auditor/runs/{stamp}/scope.tsv

This is the P in the Scope table's Passes row. The R — how many passes actually produced a run file — is counted by the assembler from the run files themselves, so a pass that dies before it can record anything still lowers the count. The row is printed only when P is above 1, so writing the key on a 1-pass scan is harmless: passes_planned 1 prints no row.

Turn 1c — Memory read. SKIP this turn entirely when memory is off. It is a turn of its own, and not a step of Turn 2, because two of its outcomes stop or downgrade the whole scan — they have to be reached before any expensive work starts.

  1. Shell check. Memory is merged by awk. Run command -v awk >/dev/null once. If it fails, print memory needs a bash shell — on Windows install Git for Windows, turn memory off for this scan, scan as a plain 1-run scan, and touch no file. There is no second implementation of the merge; the rule that must never break lives in one language only.

  2. Stale temporary file. If .solidity-auditor/memory.tsv.tmp exists, a previous scan did not finish. Print warning: .solidity-auditor/memory.tsv.tmp left by an unfinished scan — overwriting, then carry on. It is overwritten by this scan's merge.

  3. Read the ledger. If .solidity-auditor/memory.tsv does not exist, this is a first-ever scan: memory is empty, every finding is NEW, the file is written at the end. Otherwise read it and validate it before using it:

    • the first line's first tab-separated field is exactly #solidity-auditor-memory v1
    • every later line has exactly 6 tab-separated fields

    On any failure print the path and the problem and STOP the scan. No agents, no report, no write. The user fixes or deletes the file. Do NOT start fresh and do NOT continue with a warning — a single bad row must never destroy real memory.

  4. Photocopy it. mkdir -p {bundle_dir} is already done; copy the ledger as it was before this scan started:

    cp .solidity-auditor/memory.tsv {bundle_dir}/memory-before.tsv 2>/dev/null || : > {bundle_dir}/memory-before.tsv
    

    Every merge this scan performs reads this photocopy, never the live file. Create it empty when no ledger exists, so the merge command below needs no special case.

  5. Create the scan's row file, empty: : > {bundle_dir}/scan-rows.tsv. Each run appends its gated rows to it.

  6. Hold the key list — column 1 of every row of the photocopy (after the prune in Turn 2, if one runs).

    The per-function bug-class vocabulary — for each contract|function prefix, the bug-class labels the ledger already holds — is not held in the orchestrator's head. Turn 2 step 2c writes it to {bundle_dir}/known-findings.md, where the agents read it inside their bundles and Turn 4 reads it back from disk. It is a file and not a memory because the two readers are a long scan apart, and because the same labels must reach both.

    The vocabulary is what keeps exact key matching honest: a bug class is a label written in words, so the same bug re-labelled is a second record and memory fails silently. Handing the existing labels back is how the same bug keeps the same key.

Turn 2 — Prepare. In one message, make parallel tool calls: (a) Read {resolved_path}/report-formatting.md, (b) Read {resolved_path}/judging.md, (c) Read {resolved_path}/agent-prompts.md, (d) Read {resolved_path}/report-language.md.

Why report-formatting.md is still read, now that nothing here composes a report. It is the shape Turn 4 step 5a writes each finding in — title line, location line, Description, diff Fix block — and the assembler pastes those bytes straight through. Read it as the contract the finding blocks meet, not as a template to imitate at the end.

report-language.md is read for the same reason, one level down. report-formatting.md settles the shape of a finding block; report-language.md settles the words inside it. Turn 4 step 5a writes the title and the Description, the assembler pastes them through unchanged, so those bytes are the last chance to write a sentence a developer can act on. It is Simplified Technical English (ASD-STE100), and it is not optional styling: a finding the developer cannot read is a finding they do not fix. The same file is in every agent bundle, so the sentence the pass writes and the sentence the agent handed it obey one rule.

Then build source.md, run the memory step, and only then cat the bundles — in that order, because the bundles carry a file the memory step writes:

Turn 2 is split across the loop. Step 1 runs once per scan: source.md provably cannot change between passes — one git SHA for the whole loop, no pruning between them — and it is the expensive half of the build. Step 2c and step 3 run once per pass, because only they carry the knowns, and re-catting twelve bundles from files already on disk is one Bash command. Steps 2a and 2b run once per scan with step 1, since both read that frozen source.

One {bundle_dir}, reused, everything overwritten. source.md is written once and never touched again; known-findings.md and the twelve agent-N-bundle.md files are overwritten each pass. Disk stays flat whether the runner picked 1 pass or 10 — a directory per pass would hold 12 × N copies of the whole repo. This is safe because Turn 3b is a hard barrier: no pass-K agent is still reading a bundle when pass K+1 overwrites it. The cost, accepted: after the loop you cannot see what pass 2 told its agents. The durable record is the run files and the ledger.

  1. Once per scan. {bundle_dir}/source.md — ALL in-scope .sol files, each with a ### path header and fenced code block.
  2. Turn 2 step 2 — Name map, prune and known findings (below). SKIPPED whole when memory is off. Parts a and b run once per scan; part c runs every pass, because the ledger it reads grows as the loop learns.
  3. Every pass. Agent bundles, in a single Bash command using cat (not shell variables or heredocs) = source.md + agent-specific files:

Shortened here. Read the whole file on GitHub.

Signals

GitHub stars
1k
Forks
224
Last commit
Sep 2026
Advanced
Item type
skill
Key
solidity-auditor
Source
github.com/pashov/skills