Performance Optimization
SkillDev toolsLets your agent analyze and improve performance of the Static Web Server, covering profiling, caching, and resource usage.
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the Performance Optimization skill
About this capability
Optimize or review performance for the Static Web Server (SWS) project — profiling, bottlenecks, resource usage, compression, and caching
What this skill tells your AI
The instructions your AI receives, as published by static-web-server/static-web-server in .agents/skills/performance/SKILL.md and read by ahel’s review.
Load this skill when profiling, optimizing, or reviewing code for performance — latency, throughput, memory, or CPU.
When to load: a change touches the request hot path (handler.rs, static_files/, compression.rs, fs/stream.rs), a benchmark is being added or interpreted, a regression in requests/sec or latency is suspected, or a new dependency may affect binary size or runtime cost.
General Approach
- Measure before optimizing: Use profiling tools (perf, flamegraph) to identify bottlenecks. Never optimize based on intuition
- Set a target: Define acceptable latency/throughput before starting. Stop optimizing when the target is met
- Optimize the hot path: Focus on code that runs on every request. Startup code and config parsing are low priority
Rust Performance
- Profile with
perfandflamegraph:cargo flamegraph --bin static-web-serverfor CPU profiles. Profile under load (e.g.,wrkorbombardier) - Profile heap allocations with DHAT: Use DHAT or dhat-rs to find hot allocation sites. Reducing 10 allocations per million instructions can have measurable impact
- Avoid unnecessary allocations: Prefer
&Pathover&PathBuf,&[u8]overVec<u8>, pass by reference where ownership is not needed - Pre-compute at startup: Canonicalize paths, parse config, compile regex patterns, build Aho-Corasick automata once. Never on the request path
- Pre-allocate collections: Use
Vec::with_capacity,String::with_capacitywhen the size is known - Stream large responses: Use
tokio::fs::File+tokio::io::copyfor file serving. Never buffer the full file in memory - Inline small hot functions: Use
#[inline]on small functions called on every request (e.g., header name normalization, MIME lookups). Use#[cold]on error-path functions to guide branch prediction away from the hot path - Prefer
filter_mapoverfilter().map(): Avoids an intermediate layer in hot iterator chains - Use
iter().copied()for small types: When iterating over&u8,&u32, etc.,.copied()lets LLVM generate better code than receiving references - Use
chunks_exactwhen chunk size evenly divides length: Faster thanchunksbecause it eliminates a remainder check per iteration - Prefer
ok_or_elseoverok_or:ok_or(expensive())always evaluates its argument.ok_or_else(|| expensive())is lazy and only evaluates onNone - Eliminate bounds checks in hot loops: Use iteration instead of index-based access, or add an upfront assertion on the range to let the compiler prove bounds are safe
HTTP Performance
Connection Handling
- HTTP/1.1 keep-alive: Enabled by default via Hyper. Reduces connection setup overhead for subsequent requests
- HTTP/2 multiplexing: Enable with
--http2 --tls. Multiple concurrent streams over a single TCP connection - Worker threads: Default is
num_cpus * 1. Increase--threads-multiplierfor workloads with mixed CPU and I/O blocking (e.g., many concurrent clients with dynamic compression enabled — compression per-request is CPU-bound but high concurrency adds I/O wait interleaving). For pure CPU-bound workloads with minimal I/O, increasing threads beyond CPU count rarely helps. - Max blocking threads: Default 512. For I/O-heavy patterns (large file serving), this is sufficient
- Graceful shutdown: Use
--grace-periodto allow in-flight requests to complete before shutdown
Compression Tradeoffs
- Static compression is free: Pre-compressed
.br/.gz/.zstfiles are served with zero CPU. Always prefer this for production - Dynamic compression overhead: On-the-fly compression trades CPU for bandwidth. Use
--compression-level fastestfor high-traffic sites - Minimum size threshold: Responses below 860 bytes skip dynamic compression entirely — the overhead exceeds any bandwidth savings
- Compression algorithm priority (by compression ratio × speed): zstd > brotli > gzip > deflate. zstd offers the best ratio-speed tradeoff
Caching Headers
- Cache-Control is enabled by default: SWS sets
max-agebased on file extension:- 1 year for static assets (
.css,.js,.png,.woff2, etc.) - 1 hour for feeds/API (
.json,.xml,.rss,.atom) - 1 hour fallback for unknown extensions
- 1 year for static assets (
You can override these defaults per file or extension using the configuration file. The above values are defaults, not hardcoded limits.
- Conditional requests: SWS supports
If-Modified-SinceandIf-Unmodified-SinceviaConditionalHeaders. Returns 304 when the file hasn't changed - ETag not implemented: SWS uses
Last-Modifiedinstead. For byte-level cache validation, put SWS behind a CDN or reverse proxy
File I/O Performance
Buffering
- Optimal buffer size:
optimal_buf_size()selects the best buffer size based on file metadata (usesstd::fs::Metadata::blksize()when available) BufReaderwithtake(): For byte-range requests, aBufReaderwraps the file handle and limits bytes read to the requested range- Streaming avoids full-file buffering:
FileStreamreads in chunks. The response body is a stream, not a byte buffer
Path Operations
- Canonicalize once at startup: The root directory is canonicalized in
server/opts.rs. Per-request path resolution reuses this try_metadata()caches nothing: Each call (insrc/fs/meta.rs) is a filesystem syscall. The experimental memory cache feature (mini-moka, insrc/mem_cache/) caches file metadata and content- Avoid
clone()in the hot path:static_files.rsavoids cloning file paths for non-directory requests
Pre-compressed Static Files
- Zero-CPU serving: SWS detects
.br/.gz/.zstvariants viaAccept-Encodingand serves them directly. No compression step runs - Build-time pre-compression: Generate variants with maximum quality:
brotli -q 11,gzip -9,zstd -19. SWS serves them as-is - Vary header:
Vary: Accept-Encodingis appended so caches know to store multiple variants
Memory
- Minimal per-connection state: SWS stores only the remote address and handler opts (shared via
Arc). No per-connection buffers - Response body is a stream: File contents are streamed, not buffered. Exception: small generated responses (health endpoint, error pages, directory listing HTML)
- Experimental in-memory cache:
mini-moka(insrc/mem_cache/) caches hot files in memory with LFU admission and LRU eviction. Configurablecapacity(default 100 entries),ttl(default 1800s),tti, andmax_file_size. Keys useCompactStringto reduce allocation
Allocation Patterns
- Prefer
clone_fromover reassign-and-clone:a.clone_from(&b)reusesa's existing heap allocation when possible, avoiding an extra alloc/free. Especially valuable forVecandStringin hot loops - Reuse collections across iterations: Declare the collection outside the loop, call
.clear()at the end of each iteration. Avoids repeated alloc/free while keeping the heap allocation alive - Use
Cow<'_, str>/Cow<'_, Path>for mixed borrowed/owned data: Avoids allocating aString/PathBufwhen the data is already a static literal or an existing slice that won't be modified - Use
SmallVec<[T; N]>for short, stack-like sequences: When most allocations hold ≤ N elements (e.g., header value lists, index file candidates),smallvecavoids heap allocation entirely for the common case - Convert finalized
VectoBox<[T]>withinto_boxed_slice(): Drops the unused capacity word, shrinking the type from 3 words to 2. Good for config-time data that is built once and never grown - Return
impl Iterator<Item=T>instead ofVec<T>from helpers: Avoids an allocation when the caller only needs to iterate - Avoid
format!when a literal orwrite!suffices: Everyformat!call allocates aString. Write directly to a&mut Stringor usestd::fmt::Writeinstead
Type Sizes
- Keep hot types under 128 bytes: The compiler emits
memcpyfor values larger than 128 bytes. If a hot type exceeds this, check its layout withRUSTFLAGS=-Zprint-type-sizes cargo +nightly build --release - Box large enum variants: If one variant is much larger than the others, box its fields to bring all variants to a similar small size. Reduces stack pressure and cache churn
- Use smaller integer types for index/count fields: Prefer
u32overusizefor counts and offsets stored in frequently instantiated structs (e.g., header tables, path segments). Cast tousizeat use sites - Assert type sizes in tests: Add
static_assertions::assert_eq_size!(HotType, [u8; N]);for performance-critical types so that accidental size regressions cause a compile error
Release Build Configuration
The default cargo build --release profile is a good starting point, but the following options can improve throughput for production SWS builds:
| Option | Effect | Cargo.toml |
|---|---|---|
lto = "thin" | Cross-crate inlining, 5–15% speedup, moderate compile cost | [profile.release] |
codegen-units = 1 | Single codegen unit, enables more optimizations, slower compile | [profile.release] |
panic = "abort" | Removes unwinding machinery, smaller binary, slight speedup | [profile.release] |
For a custom server build where broad CPU compatibility is not required:
RUSTFLAGS="-C target-cpu=native" cargo build --release
This emits AVX/SSE instructions optimal for the build machine, which can improve compression throughput.
Note:
target-cpu=nativeproduces a non-portable binary. Do not use for distributed release artifacts.
Benchmarking
Micro-benchmarks (CodSpeed)
The benches/ directory is a standalone crate with Criterion benchmarks for hot-path
functions (path sanitization, header handling, redirects, basic auth). They run on every
pull request via the codspeed workflow and are tracked on
CodSpeed.
cd benches
# Plain Criterion run (local timings only)
cargo bench
# Same suite measured with CodSpeed CPU simulation (deterministic, flamegraphs)
cargo codspeed build
codspeed run --mode simulation -- cargo codspeed run
Add a bench when a change touches the request hot path so a future regression is caught by CI instead of by users.
Tools
- HTTP load generators:
wrk,bombardier,oha,hey - CPU profiling:
perf record+flamegraph,cargo flamegraph,cargo instruments(macOS),samply(cross-platform) - Allocation profiling:
dhat-rs(all platforms), DHAT via Valgrind (Linux) — identifies hot allocation sites - Monitoring: Prometheus metrics via
--metrics+/metricsendpoint
What to Measure
- Requests per second at concurrency levels: 1, 10, 100, 1000
- Latency percentiles: p50, p95, p99
- Memory usage: RSS before and under load
- CPU utilization: Per-core usage during load test
- Allocation rate: Use DHAT to confirm per-request allocations are not growing unexpectedly
Checklist
- Is there a benchmark or profile showing the bottleneck?
- Are paths canonicalized once, not per-request?
- Is static compression used where possible (zero CPU)?
- Is dynamic compression size-threshold applied (860 bytes)?
- Are large files streamed, not buffered?
- Are Cache-Control headers set appropriately for the content type?
- Are worker threads configured for the workload?
- Is keep-alive or HTTP/2 enabled for connection reuse?
- Are hot allocation sites identified with DHAT or similar?
- Are collections pre-allocated or reused rather than recreated per request?
- Are
clone()calls on the hot path justified — or replaceable withclone_from,Cow, or a reference? - Are hot types under 128 bytes (no unintended
memcpy)? - Are
#[inline]/#[cold]attributes applied where profiling shows they help? - Are
ok_or_else/ lazy combinators used instead of eager alternatives in hot paths?
Signals
- GitHub stars
- 2k
- Forks
- 130
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Gateway key
performance-static-web-server- Source
- github.com/static-web-server/static-web-server