Docker Agent: Serving, Sharing, and Evaluating
SkillCloud & infraUse this skill when exposing a Docker Agent as a server (MCP, HTTP API, A2A, ACP, or OpenAI-compatible chat), distributing an agent via an OCI registry with `docker agent share`, or measuring agent quality with `docker agent eval`. Even if the user just says they want to "turn my agent into an MCP server", "let Claude Desktop use my agent", "publish my agent to Docker Hub", "push my agent like an image", or "test my agent in CI", this skill applies. Covers `serve mcp/api/a2a/acp/chat` listen addresses and auth flags, `share push/pull`, eval session JSON format, scoring metrics, and the `--baseline` regression gate.
Available today. Use it from your connected AI after setup.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
Then ask your AI: use the Docker Agent: Serving, Sharing, and Evaluating skill
What this skill tells your AI
The instructions your AI receives, as published by docker/skills in skills/docker-agent-deploy/SKILL.md and read by ahel’s review.
Overview
This skill owns the integration surface of Docker Agent: making an agent
reachable by other software (docker agent serve), distributing it through
an OCI registry the way container images are distributed (docker agent share), and proving it still behaves after a change (docker agent eval).
It assumes the agent config already exists — see docker-agent-config for
authoring it, and docker-agent-run for interactive/local invocation.
When to use this skill
Activate this skill when:
- The user wants an agent reachable over MCP, an OpenAI-compatible chat endpoint, a plain HTTP API, or A2A/ACP.
- The user wants to publish an agent to Docker Hub (or any OCI registry) or pull one someone else published.
- The user wants automated evaluations (regression tests) for an agent, or wants to gate CI on eval results.
Do not use this skill when
Do not use this skill when:
- The task is authoring the agent.yaml itself (models, toolsets, sub_agents) — use
docker-agent-config. - The task is running the agent interactively on a developer's machine, choosing
--safety/--sandbox, or aliases — usedocker-agent-run.
Core guidance
Serving an agent
-
Five server modes, each with its own default loopback listen address — never expose any of them beyond loopback without authentication:
Mode Default listen Auth flag Has --safety?serve mcp127.0.0.1:8081--auth-token(only with--http)Yes (only with --http)serve api127.0.0.1:8080--auth-tokenNo serve chat127.0.0.1:8083--api-key/--api-key-envYes serve a2a127.0.0.1:8082--auth-tokenYes serve acp(stdio only) n/a No docker agent serve mcp ./agent.yaml --http --listen 127.0.0.1:9090 --auth-token "$TOKEN" -
serve mcpdefaults to stdio transport (for local clients like Claude Desktop); pass--httponly when you need a network-reachable MCP endpoint, and set--auth-tokenwhenever you do. -
Binding any server flag to a non-loopback address without an auth token/key is refused;
--insecure-no-authexists to force it and must be treated as a deliberate, documented exception, never a default. -
serve mcp(with--http),serve chat, andserve a2aexpose--safety(strict/balanced/restricted/autonomous); Docker's docs state it defaults torestrictedfor these modes when unset.serve apiandserve acpexpose no--safetyflag at all. Never raise--safetytoautonomouson a network-reachable listener; if a served agent must approve more, preferbalancedand keep auth enabled. -
serve apiaccepts a directory instead of a single file: every.yaml/.yml/.hclin it is exposed under/api/agents. Use--session-workingdir-rootto confine session working directories when the server is reachable by more than one user.
Sharing agents via OCI registries
- Push and pull agent configs the same way you push and pull images — same
registry, same
docker loginauth:docker agent share push ./agent.yaml docker.io/username/my-agent:latest docker agent share pull docker.io/username/my-agent:latest instruction_filecontents are inlined into the pushed artifact automatically, so a published agent stays self-contained — you do not need to bundle the referenced files separately.- Pin
sub_agentsthat reference the pushed artifact to a digest (name@sha256:...) once published, to avoid a per-run registry lookup and to guarantee the exact config a consumer gets. - Use
--forceonshare pullonly when you intend to overwrite a local copy that already exists; without it, an existing local config is left untouched.
Evaluating agents
- Evals live in an
evals/directory next to the agent config by default; each eval is one JSON session file capturing a user message, the recorded tool calls, and anevalsobject with the scoring criteria. - Create eval sessions from real conversations rather than hand-writing
JSON: run the agent interactively, then use the
/evalslash command in the TUI to save the session, and edit inrelevance/size/assertionscriteria afterward. - Four scoring dimensions: Tool Calls (F1 against the recorded sequence),
Relevance (LLM-judge,
--judge-model, defaultanthropic/claude-opus-5), Size (S/M/L/XL response-length bucket), and Assertions (deterministic checks; see the complete assertion-type list inreferences/eval-format.md). Prefer assertions overrelevancewhen a check can be exact: they need no judge model and are deterministic, not approximation-prone. - Evaluations run inside containers for isolation; a Docker-compatible
runtime is required. Dedicated provider API keys
(
ANTHROPIC_API_KEY/OPENAI_API_KEY) are forwarded automatically.GITHUB_TOKEN/GH_TOKENare not forwarded automatically (they're broad host credentials, not model keys) — pass them explicitly with-e GITHUB_TOKENwhen an agent's provider needs one (e.g.github-copilot). - Gate CI on regressions, not on absolute scores, with
--baseline:
A previously-passing eval that now fails always gates regardless of tolerance; cost changes are reported but never gate. A baseline or run with zero evaluations (e.g. andocker agent eval ./agent.yaml --baseline results/2026-08-01-run.json --regression-tolerance 0.05--onlypattern matching nothing) is rejected rather than reported as passing. - Use
--keep-containersplus your runtime'sexecto inspect a failed eval's container; the eval's.dbsession file holds the full conversation for offline debugging.
Verify
- After changing a served agent's config, re-run its evals with the same
explicit
--safetyvalue used in the deployment before restarting the listener — this catches an approval-policy regression before it reaches traffic. If a rollout must be rolled back, restore the prior config and safety flag; never restore an unauthenticated listener as a rollback shortcut.
Related skills
- For writing or changing the underlying
agent.yaml, usedocker-agent-config. - For local/interactive runs, safety-mode choice, and sandboxing, use
docker-agent-run.
References
references/eval-format.md— full eval session JSON schema and CLI flag table.references/sources.md— provenance of every rule in this skill.
Assets
assets/eval-session-example.json— a minimal eval session file to copy and adapt.
Checks
checks/verification.md— Verification runbook for serving, sharing, and evaluating an agent.
Signals
- GitHub stars
- 410
- Forks
- 21
- Last commit
- Sep 2026
ahel review
K2info
exfiltration (in checks/verification.md)
Automated review, not a security audit. Ruleset v1+k2.
Advanced
- Item type
- skill
- Key
docker-agent-deploy- Source
- github.com/docker/skills