Vercel Functions
SkillCloud & infraLets your agent apply Vercel Claude Skill guidance to configure, debug, and optimize serverless functions.
Available today. Use it from your connected AI after setup.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
Then ask your AI: use the Vercel Functions skill
About this skill
Vercel Functions expert guidance, Node.js/Bun/Python runtimes, Fluid Compute, long-duration (30 min) functions, large functions (5 GB bundles), Docker/OCI container images, plan limits, streaming, WebSockets, and Cron Jobs. Use when configuring, debugging, or optimizing server-side code running on
What this skill tells your AI
The instructions your AI receives, as published by vercel/vercel-plugin in skills/vercel-functions/SKILL.md and read by ahel’s review.
You are an expert in Vercel Functions — the compute layer of the Vercel platform.
Prefer Node.js over the Edge runtime
Strongly prefer Node.js. Don't reach for runtime = 'edge' in new code, and recommend migrating off it when you find it in existing code.
This follows Vercel's own documented guidance — "we recommend migrating from edge to Node.js for improved performance and reliability" — and one hard constraint: as of Next.js 16.3, runtime = 'edge' is no longer supported. Routes and pages there run on Node.js regardless of what you write, so on 16.3+ this stops being a recommendation and becomes a migration you have to do.
Everywhere else it is a strong default, not a prohibition. Both runtimes run on the same Fluid Compute infrastructure, in the same regions, under the same Active CPU pricing — so in nearly every case Edge gains you nothing while costing you most of the Node.js API surface. If you have a specific, tested reason to stay on Edge, that's a legitimate call; just make it deliberately rather than by habit.
The default to reach for
// app/api/hello/route.ts — no runtime export needed.
export async function GET() {
return Response.json({ message: 'Hello from Node.js on Fluid Compute' })
}
Node.js is the default. Omit export const runtime entirely rather than writing export const runtime = 'nodejs'.
Reasons people reach for Edge — and what to do instead
| "I need Edge because…" | Reality | Do this instead |
|---|---|---|
| "…I need to stream / SSE / AI tokens" | Streaming is zero-config on Node.js. This is the single most common false belief. | Return a ReadableStream from a normal Node.js function |
| "…I need low latency" | Both run on Fluid Compute. Fluid pre-warms instances and caches bytecode; the difference is noise next to your DB/API round trips | Stay on Node.js; pin regions near your data |
| "…auth checks / redirects / A-B tests at the edge" | That's Routing Middleware's job, and Routing Middleware supports full Node.js — it is not edge-only | Use Routing Middleware (routing-middleware skill) |
| "…it's cheaper" | Identical Active CPU pricing | Stay on Node.js |
| "…it has faster cold starts" | Fluid Compute reuses warm instances across concurrent invocations and bytecode-caches Node 20+ in production | Stay on Node.js |
| "…my function must run globally" | Edge's global execution usually hurts — every DB query crosses an ocean | Single region (iad1 default) next to your database |
What Edge actually costs you
- No
fs, no native modules, norequire()— ESM only, and most npm packages with Node.js dependencies simply will not load - No
eval/new Function/ dynamicWebAssembly.instantiate - Code size limit after gzip: 1 MB (Hobby), 2 MB (Pro), 4 MB (Enterprise) — versus 250 MB uncompressed (up to 5 GB) on Node.js
- Must begin sending a response within 25 seconds (it may then stream for up to 300s). The 300s/800s/1800s duration limits below apply to the Node.js, Bun, and Python runtimes — not to Edge
- No long-duration or large-function support of any kind
Migrating an existing Edge function
Worth doing when you're already touching the file, and required on Next.js 16.3+. An Edge function that works today isn't an emergency.
- Remove
export const runtime = 'edge'(orruntime: 'edge'invercel.json/ theconfigobject). - Replace
next/serverEdge-only imports where applicable; the WebRequest/Responsehandler signature is unchanged, so most route handlers need no other edit. - If you pinned execution with the Edge-only
preferredRegion, useregionsinvercel.jsoninstead. - Confirm Fluid Compute is on (default since April 23, 2025) and redeploy.
There is no rollback story to plan for: Node.js is a superset of what the function could do on Edge.
Function Types
Node.js (the default)
- Full Node.js runtime, all npm packages available
- Default for Next.js route handlers, Server Actions, Server Components, and any file in
/api - Node.js 24 LTS is GA for builds and functions (V8 13.6, global
URLPattern, Undici v7, npm v11). Node.js 20 is deprecated on October 1, 2026 — move offnodejs20.x - Duration: 300s default on every plan; 800s max on Pro/Enterprise; 1800s with the extended-duration beta
Bun
Add "bunVersion": "1.x" to vercel.json to run functions on Bun instead of Node.js. ~28% lower latency for CPU-bound workloads. Supports Next.js, Express, Hono, Nitro, and Bun.serve as an entrypoint. Bun supports both large functions and extended max duration.
Python
Python 3.12 / 3.13 / 3.14 on Fluid Compute. FastAPI, Flask, and Django build into a single function from the resolved entrypoint — key vercel.json config on that entrypoint file (app/main.py, myproject/wsgi.py), not on /api routes. Python gets a 500 MB standard bundle limit (vs. 250 MB) and supports large functions and extended duration.
Rust
Rust functions run on Fluid Compute with HTTP streaming and Active CPU pricing. Official runtime (Beta) built on the vercel_runtime crate. Supports environment variables up to 64 KB.
Container images (Docker)
Any OCI image via Dockerfile.vercel. See Docker and Container Images below.
Edge (legacy — not recommended)
V8 isolates with a subset of Web APIs. Fine to leave in place on existing deployments, but not the runtime to pick for new work. See Prefer Node.js over the Edge runtime.
Choosing a Runtime
| Need | Runtime | Why |
|---|---|---|
| Anything not listed below | nodejs | The default, and correct nearly always |
| Full Node.js APIs, npm packages | nodejs | Full compatibility |
| AI streaming, SSE, WebSockets | nodejs | Zero-config streaming, long durations |
| Lower latency, CPU-bound work | nodejs + Bun | ~28% latency reduction |
| Database connections, heavy deps | nodejs | Pin regions next to the database |
| Data/ML libraries, big model files | nodejs or python + large functions | Up to 5 GB bundles |
| Systems-level performance | rust | Native speed on Fluid Compute |
| Custom system libraries (FFmpeg, Chromium), Go/Ruby/PHP, unsupported frameworks | container image | Bring your own Dockerfile |
| Auth, redirects, A/B tests before the cache | Routing Middleware | Runs on Node.js, framework-agnostic |
| Hours-to-months of execution | Vercel Workflow | Durable steps, no duration limit |
edge is deliberately absent: there's no row here where it's the better answer for new code.
Fluid Compute
Fluid Compute is the execution model for Vercel Functions — enabled by default for new projects since April 23, 2025, and available for the Node.js, Python, Bun, Rust, and Edge runtimes. Enable it explicitly per-deployment with {"fluid": true} in vercel.json, or project-wide in Settings → Functions.
Long-duration, large-function, and container-image support all depend on it.
Key behaviors:
- Optimized concurrency: multiple invocations share one instance instead of one microVM per request. Vercel prioritizes idle existing resources before allocating new ones. Available on the Node.js and Python runtimes.
- Active CPU pricing: you are billed for CPU time your code actually consumes, plus provisioned memory while requests are in flight, plus invocations. Waiting on I/O (AI models, DB queries) does not accrue Active CPU — which is what makes 30-minute functions affordable.
- Automatic cold start optimization: function pre-warming plus bytecode caching on Node.js 20+. Bytecode caching applies to production only — not dev or preview, so don't benchmark cold starts in a preview deployment.
- Error isolation: an uncaught exception or unhandled rejection is logged and in-flight requests are allowed to finish; one broken request will not crash its neighbors on the same instance.
- Cross-AZ and cross-region failover: fails over to another availability zone in-region first, then to the next closest region.
- Graceful shutdown:
SIGTERMbefore termination (see below).
Instance Sizes (memory / CPU)
| Type | Memory / CPU | Use |
|---|---|---|
| Standard (default) | 2 GB / 1 vCPU | Predictable performance for production workloads |
| Performance | 4 GB / 2 vCPU | Latency-sensitive applications and SSR workloads |
- With Fluid Compute enabled, memory cannot be set in
vercel.json— setting it there produces a build-time warning. Set it in the dashboard instead: Settings → Functions → Advanced Settings → Function CPU, then redeploy. (Thememorykey still exists for legacy non-Fluid deployments, which is why you will find older examples using it.) - Pro/Enterprise only. Hobby always runs Standard (2 GB / 1 vCPU) and cannot configure it. The Basic instance has been removed.
- More memory also means more CPU, which can reduce Active CPU billing for CPU-bound work by finishing sooner — but it raises Provisioned Memory cost while requests are in flight.
- Projects created before 2019-11-08 may still sit on legacy sizes (1024 MB / 0.6 vCPU on Hobby, 3008 MB / 1.67 vCPU on Pro) until you pick a size in the dashboard.
Settings precedence
Function code (export const maxDuration) → vercel.json → dashboard → Fluid defaults. Later entries lose.
Background Processing with waitUntil
waitUntil takes a Promise, not a callback. Passing a function does nothing — a common and silent bug.
import { waitUntil } from '@vercel/functions'
export async function POST(req: Request) {
const data = await req.json()
// Correct: invoke the async work and hand over the promise.
waitUntil(processAnalytics(data))
// For several tasks, combine them:
waitUntil(Promise.all([sendNotification(data), updateCache(data)]))
return Response.json({ received: true })
}
Next.js after (equivalent)
import { after } from 'next/server'
export async function POST(req: Request) {
const data = await req.json()
after(async () => {
await logToAnalytics(data)
})
return Response.json({ ok: true })
}
Graceful shutdown and request cancellation
// Runs on scale-down. 500 ms to clean up (30 s for container images).
process.on('SIGTERM', () => {
// flush buffers, close pools
})
Request cancellation is opt-in, per path. With it enabled, a client disconnect aborts request.signal and terminates the function — anything not wrapped in waitUntil/after is lost, which is exactly why it is not on by default.
{
"functions": {
"api/*": { "supportsCancellation": true }
}
}
export async function GET(request: Request) {
// Pass the signal through so upstream work stops too.
const res = await fetch('https://upstream.example.com', { signal: request.signal })
return new Response(res.body, { status: res.status })
}
Duration and Long-Duration Functions
Duration limits
With Fluid Compute (default), per Vercel's limits:
| Plan | Default | Maximum | Extended maximum |
|---|---|---|---|
| Hobby | 300s (5 min) | 300s (5 min) | — |
| Pro | 300s (5 min) | 800s | 1800s (30 min) — Beta |
| Enterprise | 300s (5 min) | 800s | 1800s (30 min) — Beta |
The 800s maximum is generally available on Pro and Enterprise. The 1800s extended maximum is in beta. Exceeding the limit returns 504 FUNCTION_INVOCATION_TIMEOUT.
Hobby's default and maximum are the same 300s — there is no headroom to raise, and no extended duration. Setting maxDuration above 300s on Hobby does nothing; upgrade to Pro.
Setting maxDuration
// app/api/report/route.ts — Next.js App Router (and Node.js, SvelteKit, Astro,
// Nuxt, Remix via their own config). Value is in seconds.
export const maxDuration = 800
export async function POST(request: Request) {
return Response.json({ ok: true })
}
For other frameworks and runtimes — Next.js < 13.5, Rust, Go, Python, Ruby — use vercel.json:
{
"$schema": "https://openapi.vercel.sh/vercel.json",
"functions": {
"api/long-task.py": { "maxDuration": 1800 }
}
}
Glob order matters, and Next.js projects using src/ must prefix paths with src/. For Python frameworks, key on the resolved entrypoint (app/main.py), not an /api route.
To change the project-wide default: Settings → Functions → Function Max Duration.
Extended max duration (30 minutes) — Beta
Pro and Enterprise teams can run individual functions for up to 1800s. Requirements, all of which are load-bearing:
- Per-function configuration only. Durations above 800s must be set in code or in
vercel.jsonfor that function. Project-level defaults above 800s are not supported during the beta — raising the dashboard default will not get you to 1800s. - Supported runtimes only:
nodejs20.x,nodejs22.x,nodejs24.x, Bun1.xand1.4.x,python3.12,python3.13,python3.14. - Fluid Compute must be enabled (default for new projects).
- Secure Compute and Static IPs do not support durations above 800s during the beta. If the project uses either, you are capped at 800s.
// app/api/long-task/route.ts
export const maxDuration = 1800 // 30 minutes
export async function POST(request: Request) {
await doTheLongThing()
return Response.json({ ok: true })
}
Keeping a long request alive
Over HTTP/2, Vercel sends connection-level PING frames while the response is idle. HTTP/1.1 has no equivalent, so HTTP/1.1 clients and intermediate proxies may still close an idle connection long before 30 minutes elapse. For any long-running handler, stream progress or heartbeat data while the work runs rather than going silent and emitting one payload at the end.
Use getDeadline() to find out how much time is actually left and bail out cleanly:
import { getDeadline } from '@vercel/functions'
const msRemaining = getDeadline().getTime() - Date.now()
Cost of long functions
Active CPU pricing is what makes this viable: a 25-minute function that spends 24 minutes awaiting an LLM bills almost no Active CPU, only Provisioned Memory for the instance while the request is in flight.
When 30 minutes is not enough
Do not chain functions, self-invoke, or poll to fake durability. Use Vercel Workflow, which pauses, resumes, and keeps state for minutes to months with no duration limit, plus automatic retries and crash safety. Rough guide:
- ≤ 300s → any plan, no configuration needed
- 300–800s → Pro/Enterprise, set
maxDuration - 800–1800s → Pro/Enterprise, extended-duration beta, per-function config
- Beyond 30 min, or needs to survive a crash/deploy → Vercel Workflow (
workflowskill)
Workflow steps themselves support extended function durations, so a single step can also run up to 30 minutes.
Large Functions (bundle size)
Standard limits
| Runtime | Uncompressed bundle limit |
|---|---|
| Node.js, Bun, Rust, Go | 250 MB (includes runtime layers) |
| Python | 500 MB |
| Edge runtime | 1 MB Hobby / 2 MB Pro / 4 MB Enterprise, after gzip |
Blowing the limit fails the build with Serverless Function has exceeded the unzipped maximum size of 250 MB.
Large functions — Beta
Large functions raise the uncompressed bundle ceiling to 5 GB. This is what makes Python data/AI libraries, model weights, browser automation (Playwright/Puppeteer), image/video processing, and big backend apps deployable as Functions.
- Runtimes: Node.js, Bun, Python.
- Requires Fluid Compute with Active CPU enabled (default for new projects).
- New projects are eligible by default. Existing projects opt in with the
VERCEL_SUPPORT_LARGE_FUNCTIONSenvironment variable, then redeploy:
vercel env add VERCEL_SUPPORT_LARGE_FUNCTIONS # value: 1 (use 0 to disable)
The environment variable always takes precedence over the project default, in both directions.
- Only functions that exceed the standard limit use the large-function path — everything under 250 MB keeps the normal, faster path, so enabling it is not a global performance trade.
- Not supported with Secure Compute or Static IPs.
Shrinking a bundle first
A 5 GB function still costs you cold-start time. Trim before you opt in:
In vercel.json (not supported in Next.js — see below):
{
"functions": {
"api/**/*.py": {
"excludeFiles": "{tests/**,__tests__/**,**/*.test.py,fixtures/**,testdata/**}"
}
}
}
- Next.js ignores
includeFiles/excludeFiles— useoutputFileTracingIncludes/outputFileTracingExcludesinnext.config.jsinstead. - Audit heavy imports, prefer dynamic
import(), and check for a package independenciesthat belongs indevDependencies.
Request and response payloads
Bundle size is not payload size. The request or response body of a Function is capped at 4.5 MB; exceeding it returns 413 FUNCTION_PAYLOAD_TOO_LARGE. For larger data:
- Uploads → Vercel Blob client uploads, which send the file browser → Blob directly, bypassing the function
- Large responses → stream them; streamed responses are not subject to the limit
- Otherwise, chunk across multiple requests
Docker and Container Images
Vercel Functions run OCI-compatible container images. This is first-class Docker support: bring a Dockerfile, get an autoscaling function with scale-to-zero and Active CPU pricing. It is not a VM or a long-lived server.
Quick start
Create Dockerfile.vercel (or Containerfile.vercel) at the project root. Vercel detects it automatically and adds a rewrite routing all traffic to the image.
# Dockerfile.vercel
FROM node:26-alpine
RUN npm i -g srvx
WORKDIR /app
COPY server.ts .
# srvx listens on $PORT by default
CMD ["srvx", "--prod"]
// server.ts
export default {
fetch(req: Request) {
return Response.json({ ip: req.headers.get('x-forwarded-for') })
},
}
Deploy with vercel deploy or a Git push. During the build, the image is built and pushed to Vercel Container Registry (VCR).
The rules that actually bite
- Serve HTTP on port 80, or override with the
PORTenvironment variable in project settings. A container that doesn't listen gets no traffic. - Containers must be stateless. Each instance takes a request, returns a response, and keeps nothing between calls — that is what allows autoscaling and scale-to-zero. Persist to a Marketplace database, Redis, or Blob; never to the container filesystem.
- Scale to zero after 5 minutes without traffic in production, 30 seconds in preview. Cold starts are real; do not assume a warm process.
SIGTERMwith a 30-second grace period on scale-down (regular functions get 500 ms). Use it to drain.- Logs are not per-request.
stdout/stderrare broadcast to all inflight requests of the instance, so correlate with your own request IDs. - Same Function limits and Active CPU pricing apply for size, memory, and duration.
- Secure Compute and Static IPs are not supported with custom container images. If you need either, deploy that part without a container.
- Local dev:
vercel devruns the image and requires thedockerCLI plus a running daemon.
Multiple services in one project
Use Services to deploy several frontends/backends in one project, containerized or not. Set runtime: "container" on any service you want built as an image; entrypoint points at the Dockerfile relative to that service's root.
{
"services": {
"frontend": { "runtime": "container", "root": "frontend/", "entrypoint": "Dockerfile.vercel" },
"backend": { "runtime": "container", "root": "backend/", "entrypoint": "Dockerfile.vercel" }
},
"rewrites": [
{ "source": "/api/(.*)", "destination": { "service": "backend" } },
{ "source": "/(.*)", "destination": { "service": "frontend" } }
]
}
Services are internal by default — without a top-level rewrite, nothing is publicly routable. When services is present, build/runtime keys (functions, buildCommand, installCommand, outputDirectory, framework) move into the service and are no longer valid at the top level.
Vercel Container Registry (VCR)
vercel vcr login docker # authenticate Docker with a short-lived OIDC token
vercel vcr image ls my-app # list images
vercel vcr image inspect my-app <id>
vercel vcr image rm my-app <id>
Registry limits: 2 GB per compressed layer, 15 GB total image size, 4 MB manifest, 1 MB config blob. Layers must be gzip or zstd compressed — uncompressed OCI layers are rejected. Repositories per project: 10 (Hobby) / 1,000 (Pro) / 5,000 (Enterprise). Storage is billed at $0.10 per GB.
When to reach for a container
Good fits: Go, Rust, Ruby, PHP, or other backends; apps needing system libraries like FFmpeg or Chromium; frameworks outside Vercel's auto-detection; guaranteed build/runtime parity across environments.
Poor fits: anything that must hold state in-process, keep a daemon alive between requests, or run background work independent of a request. Reach for Workflow, Queues, or Cron for those.
If your framework is already auto-detected and you have no system-library needs, the standard build is simpler and faster — a Dockerfile is not an upgrade by default.
Plan Limits at a Glance
| Hobby | Pro | Enterprise | |
|---|---|---|---|
| Duration (default / max) | 300s / 300s | 300s / 800s | 300s / 800s |
| Extended duration (beta) | — | 1800s | 1800s |
| Memory / CPU | 2 GB / 1 vCPU, not configurable | Standard or Performance (4 GB / 2 vCPU) | Standard or Performance |
| Bundle size | 250 MB (500 MB Python), 5 GB with large functions beta | same | same |
| Concurrency | auto-scales to 30,000 | 30,000 | 100,000+ |
| Regions | single region | up to 3 | all |
| Edge code size (gzipped) | 1 MB | 2 MB | 4 MB |
| VCR repos per project | 10 | 1,000 | 5,000 |
| Request/response body | 4.5 MB | 4.5 MB | 4.5 MB |
What changed for Hobby
Hobby function limits went up substantially with Fluid Compute, and stale 10s/60s numbers are a common source of bad advice:
- Duration: 60s → 300s for both the default and the maximum — a 5× increase. Hobby functions can run a full five minutes.
- CPU: the Basic instance was removed; Hobby now runs Standard, 1 vCPU / 2 GB (up from 1 vCPU / 1.7 GB), managed by Vercel with a minimum of 1 vCPU.
- Hobby still cannot configure memory/CPU, use the extended 30-minute duration, or run in multiple regions — those remain Pro/Enterprise.
Streaming
Zero-config streaming on the default Node.js runtime, including Server-Sent Events (SSE). Essential for AI applications.
You do NOT need
runtime = 'edge'for streaming or SSE. Streaming responses (ReadableStream,text/event-stream) work on the default Node.js runtime — this is the single most common reason people wrongly reach for Edge. Stay on Node.js (Fluid Compute) so you keep full Node.js APIs, npm packages, and longer durations; Edge offers no streaming advantage and caps you at 25s to first byte.
export async function POST(req: Request) {
const encoder = new TextEncoder()
const stream = new ReadableStream({
async start(controller) {
for (const chunk of data) {
controller.enqueue(encoder.encode(chunk))
await new Promise(r => setTimeout(r, 100))
}
controller.close()
},
})
return new Response(stream, {
headers: { 'Content-Type': 'text/event-stream' },
})
}
Shortened here. Read the whole file on GitHub.
Signals
- GitHub stars
- 290
- Forks
- 60
- Last commit
- Sep 2026
Advanced
- Item type
- skill
- Key
vercel-functions- Source
- github.com/vercel/vercel-plugin
github.com/vercel/vercel-plugin