Modal — serverless Python & GPU as decorators
SkillCloud & infraUse when running Python or GPU workloads serverlessly on Modal — modal.App, inline container Images, gpu= on @app.function, Volumes for weight caching, Cron schedules, ASGI endpoints, modal run vs serve vs deploy. NOT managed prediction APIs with no container of your own (that is replicate); NOT SSH-able GPU boxes rented by the hour (that is runpod).
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the Modal — serverless Python & GPU as decorators skill
What this skill tells your AI
The instructions your AI receives, as published by ericrisco/rsc-harness in skills/modal/SKILL.md and read by ahel’s review.
Modal runs your Python on remote containers without you ever writing a Dockerfile or a YAML
file. The mental model: infrastructure is declared inline as Python decorators. A
modal.App is the deployable unit; each @app.function runs in its own container built from
a modal.Image you describe in code; you attach a GPU, a Volume, or a Secret as keyword
arguments and the platform provisions, scales to zero, and tears down for you. There is no
control plane to babysit — the source file is the infra.
Pinned stack: modal 1.4.3 (released 2026-05-18), Python 3.10–3.14 (>=3.10,<3.15).
Install with pip install modal then modal setup to authenticate. Everything below uses
the Modal 1.0+ API; several pre-1.0 forms were removed and are called out as Bad→Good.
Not this skill
Modal owns the serverless-container-as-decorators surface and its CLI lifecycle; the contents of your function belong elsewhere.
| The job | Goes to |
|---|---|
| Calling a managed prediction API with no container of your own | replicate / together-fireworks / fal |
| Renting a persistent, SSH-able GPU box by the hour/week | runpod |
| FastAPI design (routing, Pydantic, deps) independent of host | fastapi |
| Writing a Dockerfile for a registry / k8s / Compose | docker |
| General Python language/runtime questions | python |
| RAG / LLM pipeline orchestration logic itself | llm-pipeline |
Decision: which entrypoint?
| You want… | Use | Persists after exit? |
|---|---|---|
| Run a function once and exit (script, batch) | modal run app.py + @app.local_entrypoint() | No (ephemeral) |
| Hot-reload dev loop for a web endpoint | modal serve app.py | No (dies on Ctrl-C) |
| A persistent named deployment (prod, schedules, endpoints) | modal deploy app.py | Yes |
| Fan out work across many containers | .map() / .starmap() / .spawn() inside an entrypoint | n/a |
Rule: schedules and live web endpoints require modal deploy. modal run exits when the
entrypoint returns, so a Cron defined under modal run never fires. modal serve is for the
dev loop only — it watches your files and redeploys on save, but the app vanishes when you
stop it.
The minimal app skeleton
import modal
app = modal.App("hello-modal")
# The image is the container spec. Build it once, reuse across functions.
image = modal.Image.debian_slim(python_version="3.12").uv_pip_install("requests")
@app.function(image=image)
def fetch(url: str) -> int:
import requests # imported INSIDE the function: it lives in the remote image, not locally
return len(requests.get(url).content)
@app.local_entrypoint()
def main() -> None:
# Runs on your laptop; .remote() ships the call to a Modal container.
print(fetch.remote("https://modal.com"))
Run it: modal run app.py. Bad = wiring infra with argparse + a bash launcher + a
hand-rolled Dockerfile. Good = the decorators above; the app, image, and scaling are all
declared in the one file. Note the in-function import: dependencies you uv_pip_install exist
in the remote image, so import them inside the function (or guard top-level imports), not at
module top where your laptop would need them too.
Images — pin, layer, cache
Build images by chaining methods on modal.Image. Rules, each with its why:
- Prefer
.uv_pip_install(...)over.pip_install(...)— it resolves and installs withuv, materially faster image builds. - Pin versions —
.uv_pip_install("torch==2.5.1", "transformers==4.46.0"). Unpinned deps make builds non-reproducible and silently drift on rebuild. - Order layers stable→volatile — system packages and big wheels first, your fast-changing code last. Modal caches each layer; a change busts that layer and everything after it.
- Add your own code with
.add_local_dir(...)/.add_local_python_source(...), not by pip-installing your repo. These are applied last so editing your source doesn't rebuild torch. .from_registry("...")when you need a specific base image;.apt_install("ffmpeg")for system binaries;.run_commands(...)for arbitrary build steps.
image = (
modal.Image.debian_slim(python_version="3.12")
.apt_install("ffmpeg") # stable: rarely changes
.uv_pip_install("torch==2.5.1", "transformers==4.46.0") # heavy wheels, pinned
.add_local_python_source("my_pkg") # volatile: your code, applied last
)
→ references/images-gpu-cookbook.md for vLLM / torch+CUDA
/ diffusers recipes and the download-once weight-cache pattern.
GPU — it's a string now
In Modal 1.0+ the GPU is a string on the decorator. The old modal.gpu.H100() objects
were removed.
- Single GPU:
gpu="H100". - Count via colon:
gpu="A100:2"(two A100s in one container). - Memory variant:
gpu="A100-80GB"(alsoA100-40GB). - Fallback list (first available wins):
gpu=["H100", "A100", "any"]. - Supported types:
T4, L4, A10, L40S, A100(-40GB/-80GB), RTX-PRO-6000, H100, H200, B200.
# Bad — removed API, raises at import.
# @app.function(gpu=modal.gpu.A100())
# Good — string form.
@app.function(image=image, gpu="A100-80GB", timeout=600)
def embed(texts: list[str]) -> list[list[float]]: ...
Pick the smallest GPU that fits: T4/L4 for cheap inference and small models, A10/L40S
mid-range, A100/H100 for training and large-model serving, H200/B200 for frontier-scale.
GPU time is billed per second a container is alive — never attach a GPU to a CPU-only job, and
keep scaledown_window tight so idle GPU containers don't burn money.
Scaling & lifecycle
Tune these keyword args on @app.function, each with its why:
| Param | Effect | Why |
|---|---|---|
min_containers=N | Keep N warm instances always running | Kills cold starts for latency-sensitive endpoints (costs idle compute) |
buffer_containers=N | Pre-warm N extra beyond current load | Smooths bursty traffic |
scaledown_window=300 | Seconds an idle container lingers before shutdown | Reuse hot containers across nearby calls; lower = cheaper, higher = warmer |
timeout=600 | Max seconds a single call may run | Caps runaway jobs |
retries=3 | Auto-retry failed inputs | Survives transient failures in .map() fan-outs |
Migration note: keep_warm → min_containers and container_idle_timeout →
scaledown_window in the 1.0 migration. The old names are gone.
Concurrency within a container is now its own decorator: @modal.concurrent(max_inputs=N)
stacked under @app.function (it replaces the old allow_concurrent_inputs= argument). Use it
so one container handles N simultaneous requests instead of one-per-container.
Volumes & Secrets
A Volume is a distributed filesystem you mount into containers to persist data across runs —
the canonical use is caching downloaded model weights so cold starts skip the re-download.
weights = modal.Volume.from_name("hf-cache", create_if_missing=True)
@app.function(image=image, gpu="H100", volumes={"/cache": weights})
def serve_model():
# Reader: refresh the view so you see writes from other containers.
weights.reload()
# ... load model from /cache ...
@app.function(image=image, volumes={"/cache": weights})
def download_weights():
# ... write files into /cache ...
weights.commit() # WITHOUT this, writes are NOT durable across containers
Gotcha: writers must call vol.commit() to persist; readers call vol.reload() to see
another container's committed writes. Forgetting commit() is the #1 "my cache is empty"
bug — the files existed in that container and vanished with it.
Secrets land as environment variables in the container:
@app.function(image=image, secrets=[modal.Secret.from_name("hf-token")])
def pull():
import os
token = os.environ["HF_TOKEN"] # value injected from the named Modal Secret
Never bake a token into the image (.run_commands("export TOKEN=...")) — it's recorded in
layer history. Use a Secret. The cookbook above also carries the HF/OpenAI secret patterns.
Web endpoints
Stack a web decorator under @app.function. Pick by surface:
| Decorator | Use for | Needs |
|---|---|---|
@modal.fastapi_endpoint() | A single GET/POST function-as-URL | fastapi[standard] in image |
@modal.asgi_app() | A full FastAPI/Starlette app you return | fastapi[standard] |
@modal.wsgi_app() | A Flask/Django WSGI app | the framework |
@modal.web_server(port=8000) | Your own server process (e.g. vLLM) on a port | the server |
Decorator stack order matters: @app.function is outermost (top), then optional
@modal.concurrent, then the web decorator innermost (bottom, closest to def).
@app.function(image=image, gpu="H100", min_containers=1, scaledown_window=300)
@modal.concurrent(max_inputs=10) # middle
@modal.asgi_app() # innermost
def web():
from fastapi import FastAPI
api = FastAPI()
@api.get("/health")
def health():
return {"ok": True}
return api
Develop with modal serve app.py (hot-reload); ship with modal deploy app.py (stable URL).
For custom domains, proxy-auth tokens, batching (@modal.batched), and concurrency tuning →
references/web-and-scaling.md. For the FastAPI app's own
design (routes, Pydantic, deps), that's fastapi — this skill only mounts it.
Scheduled jobs
# Fixed wall-clock time, with timezone — survives redeploys at the same clock time.
@app.function(schedule=modal.Cron("0 6 * * *", timezone="America/New_York"))
def nightly_report(): ...
# Interval relative to deploy time.
@app.function(schedule=modal.Period(hours=5))
def every_five_hours(): ...
Gotcha: Period is measured from deploy time and resets on every redeploy — redeploy
at 4:59 and your "every 5 hours" clock restarts. Cron is wall-clock stable; prefer it for
"run at 6am" semantics. Either way you must modal deploy (not modal run) for the
schedule to live on the platform.
Parallelism
Fan a function out across containers without managing a pool:
@app.local_entrypoint()
def main():
urls = ["https://a.com", "https://b.com", "https://c.com"]
# .map: one arg per call, results in input order.
sizes = list(fetch.map(urls))
# .starmap: each item is an argument tuple. .spawn: fire-and-forget -> handle.get() later.
handle = fetch.spawn("https://slow.com")
print(sizes, handle.get())
.map(iterable) returns results in order by default; pass order_outputs=False to yield as
they complete (faster when latencies vary). Combine with retries= on the function so a
single bad input doesn't sink the batch.
Anti-patterns
| Anti-pattern | Do instead |
|---|---|
"I'll use gpu=modal.gpu.A100() like the old docs" | Removed in 1.0. Use the string gpu="A100-80GB". |
| "Attach a GPU, it might speed up this CPU job" | GPU is billed per second alive. CPU-only job → no gpu=. |
| "My files are written, the Volume will keep them" | Not without vol.commit() (writer) / vol.reload() (reader). |
"Pin later — uv_pip_install('torch') is fine for now" | Unpinned deps drift; builds aren't reproducible. Pin every version. |
"modal run it, the endpoint/schedule will stay up" | run is ephemeral; it exits. Use modal deploy for anything persistent. |
| "Order the decorators however — Modal figures it out" | @app.function outermost, web decorator innermost. Wrong order errors. |
"Bake the HF token into the image with run_commands" | Leaks into layer history. Use modal.Secret.from_name(...). |
| "Just call the model via a managed API through Modal" | If you write no container, that's a managed-API job → replicate. |
| "I need a box to SSH into for a week" | That's a persistent rental → runpod, not Modal's scale-to-zero. |
"Set min_containers high so it's always fast" | Idle warm containers cost money 24/7. Tune scaledown_window first. |
Verify
scripts/verify.sh [TARGET] statically lints the nearest emitted Modal
*.py: it requires a modal.App(, fails if the removed modal.gpu. object form appears,
checks that any web decorator sits under an @app.function, and that any Volume uses
from_name(..., create_if_missing=...). It runs python -c "import modal" only if modal is
installed (skip-pass otherwise), needs no Modal credentials, and exits 0 on an empty target.
Project grounding (02-DOCS)
In a project with a 02-DOCS/ layer (the harness wiki), read
02-DOCS/wiki/stack/modal.md first, then record this app's real Modal choices there — GPU types,
image base, Volume names, schedule, endpoint shape — and index it in 02-DOCS/wiki/index.md. No
02-DOCS/? Skip silently.
Signals
- GitHub stars
- 82
- Forks
- 3
- Last commit
- Sep 2026
- Hacker News mentions
- 20
Advanced
- Catalog kind
- skill
- Gateway key
modal-ericrisco- Source
- github.com/ericrisco/rsc-harness