Wado performance
SkillAI & modelsAnalyze and improve the runtime speed of a Wado program's compiled guest Wasm — profile hot functions, read the generated WIR for allocations and copies, reason about the WasmGC cost model, and A/B-measure a fix. Use for any guest-side speed question, whatever the program does. For host-side native compiler profiling see profiling-wado-compiler; for wrong code out of an optimizer pass see optimizer-debug.
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the Wado performance skill
What this skill tells your AI
The instructions your AI receives, as published by wado-lang/wado in .claude/skills/wado-performance/SKILL.md and read by ahel’s review.
Speed of the compiled guest Wasm (what wasmtime runs), not the native
compiler — that is profiling-wado-compiler.
Loop: profile the hot function → read its WIR for what it allocates/copies per iteration → change one thing → A/B both arms in one session, plus the WIR diff of the hot function → keep or revert (§5 says which evidence decides).
1. Profile
wado run --profile guest,profile.json,1 prog.wado # interval 1 short runs, 0 exhaustive
wado test --profile guest,profile.json,1 file.wado # same, over one file's test blocks
node .claude/skills/wado-performance/scripts/analyze_guest_profile.ts profile.json [--top N]
The script reports self (leaf) and inclusive counts per function; names keep
monomorphization detail, so each instantiation is separate. Loop a one-shot hot
phase N times so it clears the fixed setup (aim ≥ ~200 samples). Firefox
Profiler (profiler.firefox.com) gives a flame graph; perf + --profile jitdump gives instruction-level (store- vs compute-bound), see
docs/jitdump-profiling.md.
Dev-profile inflation: a cargo run wado JITs guest code near-release but
runs the wasmtime runtime / GC / allocator at dev speed (~4–5× slower), so
profiles over-weight allocation/GC frames — read percentages as relative and
size any GC or allocation win by its release number, not the dev multiple. A
flat-CST rewrite that cut a benchmark ~3× on dev gained ~1.47× on release,
because release GC was only ~⅓ of wall-clock to begin with. Pure compute does not
inflate, so a compute-bound win carries over intact.
A sample lands at the next epoch check, not where the time went. The guest
profiler samples on an epoch deadline and wasmtime checks the epoch at function
entries and loop headers, so straight-line code is charged to whichever it
reaches next. A derived deserializer's field-dispatch chain reported as 73%
self in deserialize_i32 on an 80-field struct, and 19% in deserialize_bool
on cbor-twitter — neither function is more than a bounds check and two compares.
A hot small leaf is telling you how often it is entered, so go read its caller.
The profile ranks candidates; it does not locate them.
A low self-percentage rules out one dataset, not the function.
FieldSchema::lookup read 0.71% on json-catalog, whose widest struct is 16
fields. Rewriting that same function cut 44% off cbor-twitter's decode, where
User has 40. Profile a per-item cost on the input whose items are widest.
Rule out a super-linear pass before blaming GC — that same inflation makes an algorithmic blow-up read as GC-bound; sweep input size to tell them apart. Faster-than-linear growth is a hypothesis and not a verdict, since a live set can grow that way too, so the WIR is what settles which.
Sweeping a shape dimension is sharper than size, and the one to vary is the one the suspect is indexed by. Decoding 1000 CBOR records, holding that count fixed so no per-record term is left for the growth to be:
i32 fields per record | 5 | 10 | 20 | 40 | 80 |
|---|---|---|---|---|---|
| ns per field | 85 | 80 | 87 | 129 | 221 |
Hold everything but the dimension under test — a sweep that varies two answers about neither.
2. Read the WIR — allocations and copies first
wado dump -O2 prog.wado # final WIR
wado dump --tir-monomorphized prog.wado # how `?`, for-of, … desugar
Three villains, each a heap alloc or deep copy, and a bug when one lands per element in a loop:
struct.new/Box<…>— a heap object.for x of &listboxes every element (WasmGC has no interior references, so a by-ref iterator materializes&Tas a box). A tuple is a GC struct too, so retyping a two-field struct as[u64, u64]allocates the same. Two things remove the allocation: multivalue on a return the caller destructures, and SROA on a literal whose fields the optimizer can split into locals. A merge point defeats both preconditions, which is a pass to fix rather than a call site to rewrite (below).array.new/array.new_default— a fresh GC array (_defaultzero-fills); watch for one per call where a buffer could be reused.$value_copy$T…— a value-semantics deep copy of a value-typed binding/arg unless the source is fresh (a call / literal / variant result, or a fresh value's payload).x?desugars tomatch f() {…}, so freshness must see through thematch; a missed copy shows up here and is removable.
Also: a Trait::method(…) call left in a hot loop (the inliner declined it), and
array_set_u8 / array_get_value (bounds-checked; one per element is the store floor
for Array<T>-backed String / List).
A stdlib workaround is a bug report about a pass
An optimizer fix reaches every Wado program. The stdlib edit that routes around one reaches a single call site, and hides the gap that produced it. So when the shape you are about to rewrite by hand is one a pass exists for, name the pass and read its precondition first. That is where the fix belongs.
Three turned up this way while cutting fts. All three are also live in
short, the path every ${x} on a float takes, which is why fixing the first
paid on a benchmark fts never touched.
sroamatches only a direct literal binding. An inlinedget_pow10leftlet pm = <block with two exits>, one building the struct and one calling out for it, withpm.hi/pm.lothe only uses. ThePmHiLowas a heap object per conversion. Fixed by extendingslot_temp_sroa, which already scalarized the[tag, slots…]shape, to a struct literal and to an exit that hands over the aggregate: json-canada de +12.7%, ser +7.2%, anduscale_pow10deleted.multi_value_returnis all-or-nothing per callee. It wants every call site to belet $tmp = Call(f)whose only uses are field accesses.mul_pow10had seven sites and six were exactly that; the one yielding the call as a block value disqualified all seven. Fixing the first gap retired this one, since the offending site is now alet. The precondition is unchanged, so the next callee to hit it pays the same way.cold_outlinerefuses a region containing areturn(control_escapes), which is every rare slow path there is. Leavingfixed_width_for_prec's out-of-range tail inline cost json-canada ser 6.5% against a byte-identical serialize path: growing a hot function moves everything downstream of it in the module. Hand-splitting it into a function restored the row, andfixed_width_out_of_rangeinfpfmt.wadois that split.
3. WasmGC cost facts
- The live set is the cost, not the allocation count. The
copyingcollector traces what survives a cycle; an object that dies before the next one is never copied, however many there were. Cutting transient allocations therefore moves nothing — a compiler pass that removed thousands of per-tokenBox<i32>allocs measured within noise undercopying(and −0.7 ms/iter undernull). Chase the footprint, not the volume. The same rule retires "iterate by index to stopfor x of &listboxing": the boxes die immediately. - Module-lifetime GC data is a tax on every collection. A decoded table held
in a global as one
List<i32>per state (~7.4K permanently live objects) made identical hot-function wasm run 3–6× slower purely from the resident graph; flattening it to offset/count columns fixed it. A resident 160 KB flatList<i32>costs ~+0.9 ms/parse, the 7,400-list shape ~+2.4 ms. Prefer flat columns over nested lists, and don't build what nothing reads. Measure the GC share with--collector null(it leaks, so drive a fixed iteration count) vs--collector copying. with_capacityzero-fills.List::with_capacity(n)is anarray.new_default, so an over-sized arena pays for every slot it never uses — once badly enough to turn a 2× faster build into a 4× slower one. Growing from[]by doubling is not the fix either: it zero-fills ~2.4× more than a reasonable pre-size. Size it about right, or grow.- GC-array access is bounds-checked, no unchecked variant. A lookup table in a GC array adds a checked load per access — it lost to plain arithmetic.
- A lone
array.getcosts ~20 machine instructions; gets sharing a block cost ~8. wasmtime re-derives the object's null check, its length load and the overflow-checked element address per get, and neither hoists them out of a loop nor shares them across blocks — only across gets in one block. Read the actual sequence withwasmtime explore -W gc,function-references f.wat; a byte-at-a-time loop is 22 instructions and 6 branches per byte, four gets in one block are 18 + 4×8. So a scan reads several bytes per bounds check and then tests them:peek_after_whitespace_runincore:jsonis that shape, worth 12.6% on json-catalog deserialize. It pays in proportion to the run it covers, against the one partial block it always wastes — under ~16 bytes per run it is a loss. Four is where the sharing stops: a wider block only adds lone gets, and measures worse the wider it gets (dead-ends.md).array.setshares nothing: a store may write the header as far as Cranelift knows, so four adjacent sets reload the length four times. Onlyarray.copy/array.fillamortise a write. - SROA is priced by the aggregate's width, not by the allocation it removes.
Splitting a 40-slot tuple into locals deletes one
struct.newper struct and costs 6.5% on cbor-twitter: past the register file, fortyreflocals live across a call-heavy loop are forty spill slots reloaded at every call boundary, plus aref.nullinit apiece at entry. "The allocation is gone" says nothing about which side won (dead-ends.md). array.copyis fast; leave it alone. It has a fast path that does not call out to the runtime, and it beats a hand-written loop from a couple of bytes on — the loop pays the bounds check above on both the get and the set of every byte. Neither hand-roll it nor contort an algorithm to avoid it (dead-ends.md).- Constant
/and%are cheap (Cranelift magic-multiply,x/kandx%kfused) — don't trade a divide for extra multiplies. - A short compare cascade is not a dispatch problem. Cranelift lowers a short
else ifchain competitively, and amatchover it (abr_table) adds an indirect branch: two separate rewrites to jump tables measured flat and slightly slower. Such a frame is usually call-frequency-bound, not dispatch-bound — cut the calls, not the branch. What does answer to dispatch is a cascade long enough to pay for that branch, or one that is not a cascade at all: independentifs no arm leaves test every key whatever matched, whichnir/if_chain_to_matchis what fixes. A set membership test,k matches { A | B | 'x'..='z' | … }, is neither.match_to_bitsetlowers it to a branch-free mask test when the scrutinee is 32 bits or narrower and the members span at most 256 values, so write it as the set rather than hand-rolling a range compare or a table. Past that span it is abr_table. - Write into the caller's buffer, not a temp.
`{v}`allocates a throwawayStringand copies it in, per value;buf.push_display(&v)skips both. A run of adjacentpush/push_strcalls is fused into one capacity check bynir/string_push, so write the appends plainly and let it batch them. internal_raw_data()/ returningArray<T>by value is a copy API — for a single read useget_unchecked/set_byte_unchecked.- An
assertof a caller-guaranteed precondition is free; the same test as a guard is not.is_json_wsincore:jsonneedsb < 64(Wasm masks a shift count mod 64) and every call site short-circuits onb > b' 'first. Writing the precondition asassert b < 64measures flat on json-catalog deserialize; writing it asb < 64 &&in the returned expression costs 6%, and dropping the bitset for four compares costs 19%. So a hot leaf whose precondition the callers establish keeps both the assert and the fast body. Measure the assert and the guard as separate arms: folded into one they read as a single cost, and the assert takes the blame for what the guard spent.
4. Inlining is usually not the lever
wasmtime/Cranelift call small Wasm functions cheaply, so forcing inlining rarely moves wall-time and raising the threshold bloats hot loops (measured slower). The exception is a tight iteration-bound loop with a trivial body.
A rare heavy sub-case behind a cold_path() marker usually needs no
hand-splitting: nir/cold_outline moves what the marker opens into a function of
its own, so the leaf inlines at its hot-path size. Its region runs to the end of
the enclosing block, so a marker mid-loop-body is one it cannot take (see that
pass's module doc) and that shape still needs the split written out. Split by
hand also when the slow path is not rare: a width > 0 branch that runs every
time a width is set is hot when taken, which no marker should claim otherwise.
Write the stdlib to be fast without inline hints. A #[inline] /
#[inline(never)] in wado-compiler/lib/ is a claim the cost model got it
wrong, and it silently outlives whatever measurement justified it — the split
that #[inline(never)] was added for is usually one the inliner already
declines on size. Prefer changing the shape (a separate function, a smaller hot
path) and leave the decision to the threshold; reach for a hint only after
measuring that the shape alone does not get there, and say in a comment what it
buys.
5. Measurement
Only relative numbers carry signal. A/B both arms in the same session, best of
three or four, alternating and with the order swapped once — the first run of a
session reads high, so a fixed order silently taxes whichever arm goes second.
Run on an idle host, nothing else building: an A/B taken beside a compiling
test suite has put both arms inside each other's spread and flipped their
ranking. Check ps and free as well as uptime — a load average lags a
session that just started and says nothing about memory, and another agent's on
this box put the same test target at 145 s and at 2499 s before OOM-killing the
command after it. Nothing in a number says whether its host was idle, so
benchmark/README.md is a sanity check on the arm you just built, never the
control for it — even on the machine that produced it; a HEAD build has measured
615 MB/s against its own recorded 656 in the same afternoon. Isolate the phase —
A/B a float-format change on fts, not on a serialize benchmark that dilutes it.
A dev-build A/B is only valid where the dev build is. The inflation §1 describes flips A/B verdicts, not only profile weights. Dev runs the wasmtime runtime, GC and allocator at dev speed, so a row bound by allocation reads a different winner. json-canada is store- and compute-bound, so it matched release to under 2% and made an 8-second stdlib loop possible. On the same change dev called syntax-highlight -1.2% where release said +1.4%, and cbor-canada and cbor-twitter deserialize -2.9% and -1.2% where release said +0.3%. Every row that moved is a deserialize or a CST build, which is what allocates. Iterate on dev, then settle any row whose work is building an object graph on release.
A/B-ing a compiler change
A change to the compiler needs two compilers. benchmark-baseline builds
origin/main's once and caches it under that commit; WADO_BIN then runs it
through this tree's harness, so only the compiler differs — the baseline's own
benchmark/ would put the branch's harness changes inside the comparison too.
base=$(mise run benchmark-baseline) # ~5 min the first time, 2 s after
# alternate, so neither arm always goes second
WADO_BIN=$base mise run benchmark-all > b1.log 2>&1
mise run benchmark-all > h1.log 2>&1 # …and so on, 3 each
node benchmark/ab.ts --base b1.log b2.log b3.log --head h1.log h2.log h3.log
Time the Wado rows alone, with mise run all-wado. The reference arms (C,
Rust, JavaScript, the Java ones) run the same binary whatever the compiler does,
so re-timing them buys nothing and stretches a round from seconds to minutes.
That is the gap the host drifts across: a three-arm benchmark-all comparison
came back with ANTLR4 (Java) at -2.2%, count-prime / JavaScript at +1.4% and
a prime sieve 4.3% "faster" from a string-append change, all unreadable. The
same arms over all-wado, six rounds seconds apart, settled every row. Keep
sieve in the selection as the in-band control, and run benchmark-all once at
the end for the record.
mise run all-wado # every Wado row
mise run all-wado json_catalog sieve # those, by name
Its log feeds ab.ts and pick.ts like any other, one row per benchmark file.
Hash the wasm before you time anything. Compile every benchmark under both
compilers and compare. A row whose bytes are identical cannot have moved, so
whatever the suite says about it is the host. That is a stronger check than
timing, and it leaves only the few rows that differ to measure. A
field_scalarize fix came out byte-identical on all but three benchmarks. The
suite had meanwhile called fts 6.9% SLOWER with non-overlapping ranges; the
identical SHA-256 retired that reading outright.
Gate the compare on each compiler's exit status, not on its output file. A
compile that failed leaves the previous round's file in place, and comparing
those reads as "identical". That is the one answer this check must never give by
accident. Drop the schema modules, which are no world entry point, and give
http_routing the world it targets. Every remaining benchmark must compile, so
a FAILED row is one to go and read, and it carries the diagnostic explaining
it. wado compile reports on stderr even when it succeeds, so its output is
held back and printed with the failure, which keeps the sweep's own lines
readable.
for f in benchmark/*/*.wado; do
case "$f" in *_schema.wado) continue ;; esac
world=""
case "$f" in */http_routing/*) world="--world wasi:http/service" ;; esac
"$base" compile -O2 $world -o /tmp/b.wasm "$f" > /tmp/cc.log 2>&1 \
|| { echo "FAILED $f"; cat /tmp/cc.log; continue; }
target/release/wado compile -O2 $world -o /tmp/h.wasm "$f" > /tmp/cc.log 2>&1 \
|| { echo "FAILED $f"; cat /tmp/cc.log; continue; }
cmp -s /tmp/b.wasm /tmp/h.wasm || echo "DIFFERS $f"
done
Then time only the rows that differ, back to back, and read the rest as unmoved.
ab.ts decides each row by whether the arms' [min, max] overlap, not by the
delta: on a 5 ms benchmark a 6% gap between bests sits inside one arm's own
spread. Read the reference rows first — C, Rust and JavaScript run the same
binary in both arms, so a SLOWER among them is the host drifting and no Wado
row can be read either.
Confirm a surviving row before believing it: the whole-suite arms are minutes apart, and the reference rows only catch drift big enough to cross a range. Loop that one benchmark back to back and check the ranking holds pair by pair.
for i in 1 2 3 4 5; do
"$base" run -O2 benchmark/sieve/sieve.wado
target/release/wado run -O2 benchmark/sieve/sieve.wado
done
A/B-ing a stdlib change
A change to lib/core/*.wado needs two compilers as well, which the source tree
hides: a release build embeds the stdlib where a dev build reads it from disk.
benchmark/wado.sh falls back to cargo run --release whenever WADO_BIN is
unset, so swapping an arm's .wado files into the tree invalidates
wado-compiler and rebuilds it under lto=thin / codegen-units=1. That is
two full rebuilds per alternating round, and they are the wall clock rather than
the benchmark.
Build one binary per arm first, replacing the whole of lib/ for each. Copying
only the files that differ leaves behind any file the other arm deletes or
renames, and the binary then embeds a stdlib belonging to neither.
for arm in base head; do
rm -rf wado-compiler/lib
git checkout $arm -- wado-compiler/lib
cargo build --release --bin wado --quiet
cp target/release/wado /tmp/ab/wado-$arm
done
git checkout HEAD -- wado-compiler/lib
for r in 1 2 3; do
for arm in base head; do
WADO_BIN=/tmp/ab/wado-$arm mise run benchmark-json-catalog
done
done
The tree's sources stop mattering once the binaries exist, so a round costs what the benchmark costs. Rounds are cheap enough then to run six or ten of them, which is what it takes to resolve a delta near 1% out of this row's spread.
WADO_SKIP_PASS=<pass> is a third arm off the same binary, which is how a
regression is attributed to one pass without a third build. WADO_BENCH_FLAGS
sweeps a knob the same way, but the harness spends it on wado run, so a knob
compile alone accepts is one no sweep reaches.
Give a threshold a temporary env override and sweep it, rather than
rebuilding per value — and reach for it the moment a change looks like it only
pays above some size, because that shape usually means two rewrites are riding
one knob. if_chain_to_match appeared to need a 12-arm floor; overriding its
threshold and match_to_switch's separately showed the fusion was never the
cost at any width and the br_table past it was the whole of it on the row that
regressed, turning 3.6% down on cbor-catalog into 2.1% up. Delete the overrides before
committing: read per node visit, std::env::var is itself a compile-time cost.
What decides adoption, in priority order:
- The benchmark moves → keep it.
- The WIR A/B diff shows fewer instructions → keep it, benchmark flat or not. The benchmark simply does not reach them.
- The benchmark is flat and the diff is qualitative — a different sequence, with no reading of it that says which is faster → keep whichever emits the smaller wasm. This is the only question wasm size answers.
- Anything else → drop it, and write it up in
dead-ends.md.
Only a WIR diff decides case 2. Diff the two wado dump -O2 outputs and read
what the hot function issues per iteration — a run of N capacity checks collapsed
to one, a call gone from a loop body. Nothing else establishes "fewer
instructions": not the dump's line count, and not the wasm byte count.
Neither wasm size nor dump size correlates with speed. Smaller output is
routinely slower and larger output routinely faster — the bytes are mostly code
that never runs, and what does run is priced by what the loop executes. The three
quantities move independently: the append fusion grew wado dump -O2 on
syntax-highlight 8.3% (a fused write unparses its offset as an expression) and
shrank the -Os binary 1.5%, while the thing that justified it was a +8%
benchmark and a diff showing one less capacity check per key. Size is its own
budget (mise run report-wasm-size); as evidence about speed it is only the
tiebreaker at rank 3.
6. Lessons
dead-ends.md (next to this file) is the record: every optimization measured
and dropped, with the A/B that killed it and what it generalizes to. Read it
before starting, and add an entry whenever an A/B comes back flat or negative
— a dead end nobody wrote down is one somebody re-measures.
Stop when the floor is the representation — a store-bound loop on an
Array<T>-backed String is near-optimal short of leaving GC arrays.
See also
dead-ends.md— what has already been measured and dropped.profiling-wado-compiler— the nativewadobinary (host side).benchmark— run the suite / wasm-size report.optimizer-debug— a NIR/WIR pass producing wrong code, not just slow.
Signals
- GitHub stars
- 113
- Forks
- 2
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Gateway key
wado-performance- Source
- github.com/wado-lang/wado