Deploy on Fly.io

SkillCloud & infra

Use when deploying or operating an app on Fly.io — writing fly.toml, placing Machines in regions near users, attaching Volumes, managing secrets, or picking a scaling lever (autostop/autostart, scale count, fly-replay). NOT choosing which host to deploy on (that is `deployment`), NOT a git-push PaaS with no regions model (that is `railway`).

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the Deploy on Fly.io skill

What this skill tells your AI

The instructions your AI receives, as published by ericrisco/rsc-harness in skills/fly-io/SKILL.md and read by ahel’s review.

You are deploying an app to Fly.io: a fly.toml, Machines (Firecracker microVMs) placed in regions close to users, optional region-pinned Volumes, secrets, and the right scaling lever. Get the mental model right first, then the config follows. If none of that placement control matters, ../railway/SKILL.md is the git-push PaaS with no Machines/regions model.

Mental model

  • App is the logical unit. It owns a name, a primary_region, and config in fly.toml. Why: every command targets an app.
  • Process groups ([processes], e.g. web, worker) split one image into roles. Why: a web group takes traffic, a worker group does not — they bind services and VMs separately.
  • Machines are Firecracker microVMs running your image. Each runs in exactly one region. Why: latency and volumes are per-Machine, so placement is the whole game.
  • Fly Proxy is the anycast front door. It routes a request to the nearest running Machine, can start a stopped one, and obeys fly-replay headers. Why: it is what makes "global" cheap — you do not run a load balancer.
  • Volumes are local NVMe disks pinned to one Machine in one region. No replication. Why: this single fact dictates every stateful architecture decision below.

Deploy fast (4 commands)

fly launch                              # detects framework, generates fly.toml + Dockerfile, creates the app
fly secrets set DATABASE_URL=postgres://...  # restarts every Machine; never put this in [env]
fly deploy                              # builds image, runs release_command, rolls out Machines
fly scale count 2 --region iad,ams      # place Machines in Virginia + Amsterdam

fly launch is interactive and writes a starter fly.toml. Treat that file as a draft — review it against the next section before the first real deploy. Run fly status and fly logs after any deploy.

A fly.toml that works

app            = "my-api"
primary_region = "iad"            # 3-letter region code: iad, ord, ams, syd, gru, nrt...

[build]
  # dockerfile = "Dockerfile"     # Fly builds from your Dockerfile; see ../docker/SKILL.md

[deploy]
  release_command = "npm run migrate"   # one-shot Machine that runs BEFORE the new version goes live
  strategy        = "rolling"            # rolling | bluegreen | canary | immediate

[processes]
  web    = "node server.js"
  worker = "node worker.js"

[http_service]
  internal_port       = 8080
  force_https         = true
  auto_stop_machines  = "stop"   # "off" | "stop" | "suspend" — set WITH auto_start_machines
  auto_start_machines = true
  min_machines_running = 0       # 0 = scale to zero; honored only in primary_region
  processes           = ["web"]

  [http_service.concurrency]
    type       = "requests"
    soft_limit = 200             # Proxy starts spreading load past this
    hard_limit = 250             # Proxy stops sending past this

[[vm]]                            # formerly [[compute]]
  size       = "shared-cpu-1x"
  memory     = "512mb"
  cpu_kind   = "shared"          # "shared" | "performance"
  processes  = ["web"]

[[mounts]]
  source       = "data"          # volume NAME, created with `fly volumes create data`
  destination  = "/data"
  processes    = ["web"]
  initial_size = "1gb"

Full field surface ([[services]] vs [http_service], health checks, [[statics]], [[files]], all VM sizes, [restart], [metrics]) lives in references/fly-toml.md — read it when you need a key that is not above. Custom domains, certs and registrar-level DNS are ../domains-dns/SKILL.md.

Regions: place Machines near users

Pick the branch first, then run the commands.

Your app is...StrategyHow
Stateless (no local disk; DB elsewhere)Replicate the Machine into more regionsfly scale count 2 --region iad,ams,syd
Stateful with a VolumeKeep writes in primary_region, add read replicas + fly-replaysee references/multi-region.md
Needs one extra box nowClone a single Machine (gets a fresh volume)fly machine clone <id> --region syd
fly platform regions              # list region codes + names
fly scale count web=2 --region ams   # per-process, per-region count
fly scale show                    # what runs where, right now

Rules:

  • A request with no pinned region goes to the fastest Machine for that caller via anycast — multi-region is mostly "run Machines in more places."
  • fly scale count N --region a,b is the per-region count, not a total. Why: count 2 --region iad,ams means 2 in each, i.e. 4 Machines.
  • If any target region is out of capacity, the whole scale op fails — no partial placement. Retry with fewer regions or a different code.
  • "Slow for users in Sydney, app runs in iad" => add syd, not a bigger VM. Latency is distance, not CPU.

Volumes

A Fly Volume is a local NVMe disk pinned to one Machine in one region. There is no automatic replication between volumes. Encrypted at rest by default (--no-encryption to opt out — almost never do).

fly volumes create data --region iad --size 3
fly volumes list
  • One volume attaches to one Machine. Two Machines cannot share a volume. Why: it is block storage on one host, not a network filesystem.
  • fly scale count on a group with a [[mounts]] creates a new empty volume per new Machine — it does not copy your data. This is the #1 stateful gotcha.
# Bad: expecting two web Machines to "share" /data — they each get their own empty disk
[[mounts]]
  source      = "data"
  destination = "/data"
  processes   = ["web"]   # then `fly scale count web=3` => 3 separate, unsynced disks
# Good: one writer with the volume; replicas are stateless and read via the DB/fly-replay
[[mounts]]
  source      = "data"
  destination = "/data"
  processes   = ["writer"]   # a single-Machine process group; scale `web` separately, stateless

Replication is your app's job (LiteFS, app-level streaming, or a managed DB), never the volume's. See references/multi-region.md.

Secrets

fly secrets set STRIPE_KEY=sk_live_... SESSION_SECRET=...   # one rollout
fly secrets list           # shows NAME + digest + timestamp — never the value
fly secrets unset OLD_KEY
  • fly secrets set updates every Machine and restarts them — it resets the ephemeral filesystem. Why: batch your sets into one command so you trigger one rollout, not five.
  • Secrets arrive as environment variables in the guest. Read process.env.STRIPE_KEY.
  • Need a secret as a file on disk (a cert, a service-account JSON)? Use [[files]] with secret_name — see references/fly-toml.md.
# Bad: secret baked into the image / committed config
[env]
  STRIPE_KEY = "sk_live_51H..."   # in git, in the image layers, leaked
# Good: out of the repo, out of the image, encrypted in Fly's vault
fly secrets set STRIPE_KEY=sk_live_51H...

Treat secret hygiene as non-negotiable — see ../secure-coding/SKILL.md.

Scaling: pick the right lever

LeverWhat it doesReach for it when
auto_stop_machines / auto_start_machinesFly Proxy stops/starts a pre-created pool by load; never creates/destroysBursty or idle traffic; cut cost on quiet hours
fly scale countYou set how many Machines exist per region/processSteady baseline capacity; geographic spread
fly-autoscaler (superfly/fly-autoscaler)Scales Machine count off any Prometheus metricQueue depth / custom-metric driven autoscaling
fly-replay headerApp returns fly-replay so Proxy replays the request elsewhereForward writes to primary region; route by tenant

Key distinction: autostop ≠ autoscaler. Autostop only toggles Machines that already exist; it never changes the count. The metrics autoscaler is what actually adds/removes Machines. Set auto_stop_machines and auto_start_machines together — configuring one without the other is undefined behavior.

fly-replay is the multi-region write-forwarding pattern: read-replicas serve local reads, a write replies with fly-replay: region=<primary> and the Proxy re-runs the request there. Full header forms and the primary/replica split are in references/multi-region.md. These are the Fly levers only; platform-agnostic scaling theory (queues, sharding, load shedding) is ../scaling/SKILL.md.

Cost & HA

  • min 2 Machines for HA. A single Machine = a single point of failure; Fly recommends ≥2 per group in production.
  • Stopped Machines are cheap — you pay for rootfs/volume storage, not running compute. So a warm pool with auto_stop_machines = "stop" is the default cost play.
  • Scale to zero (min_machines_running = 0) trades cost for a cold start on the next request. If the first-request latency hurts, set min_machines_running = 1 to keep one warm. Note: min_machines_running is honored only in the primary region.
  • "suspend" resumes faster than "stop" (keeps memory snapshot) but is supported on fewer setups — verify before relying on it.

Verify

After writing or editing a fly.toml, run the checker:

scripts/verify.sh path/to/fly.toml   # defaults to ./fly.toml

It prefers fly config validate when flyctl is on PATH, else does structural checks (app, primary_region, an internal_port, and the autostop-pair lint). Read-only; exits nonzero on any FAIL.

Anti-patterns

Anti-patternWhy it bitesDo instead
Secrets in [env] or the DockerfileCommitted to git, baked into image layersfly secrets set
fly scale count on a [[mounts]] group expecting shared dataEach new Machine gets an empty volumeSingle writer + stateless replicas + DB/fly-replay
Setting only auto_stop_machinesHalf-configured autostop is undefinedSet start + stop keys together
Assuming autostop changes Machine countIt only toggles a fixed poolUse fly scale count or fly-autoscaler
One Machine in productionNo HA; a host blip = downtime≥2 Machines per group
Bigger VM to fix far-away latencyDistance, not CPU, is the costAdd a Machine in the user's region
Volume in a different region than its MachineCannot attach across regionsCreate the volume in the Machine's region
Treating Fly Postgres as managedFly Postgres is unmanaged; you operate itRoute to it here; operate it via ../postgresdb/SKILL.md
min_machines_running in a non-primary regionIgnored outside primaryKeep warm capacity via scale count there
fly deploy with no release_command for a schema changeNew code hits an old schema mid-rolloutrelease_command runs the migration first

Signals

GitHub stars
82
Forks
3
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
fly-io
Source
github.com/ericrisco/rsc-harness