
TL;DR
Docker Agent (
docker-agent, invoked asdocker agent) is an open-source, Apache-2.0 runtime written in Go that turns a declarative YAML file into a working AI agent, or a whole team of them.It is provider-agnostic (OpenAI, Anthropic, Gemini, Bedrock, Mistral, xAI, local models through Docker Model Runner, and many more) and uses MCP for tools.
Agents are packaged, pushed, pulled, signed and run through OCI registries, the same plumbing your container images already use.
Sandbox mode runs an agent inside an isolated VM behind a default-deny network, which is the sane way to hand a language model a shell.
The agent mess nobody planned for
Every team that has built an agent recently knows the arc. Day one is a script, an API key and a prompt, and it feels like magic. By day thirty it is a heap of framework glue, hand-wired tools, a shell tool with the run of your home directory, and a configuration that only works on one laptop. Sharing it means a README and a prayer.
We have seen this movie before. Before Docker, shipping software meant shipping instructions: install this, set that, hope the versions line up. Containers fixed it with a deceptively simple trio: a declarative definition, a portable artifact, and an isolated runtime.
Docker Agent applies that same trio to AI agents. The project's own tagline says it plainly: Run AI agents like containers. The rest of this post is a tour of how seriously it takes that idea.
What it is, in 60 seconds
Docker Agent is a docker CLI plugin. You describe an agent in YAML (an HCL flavor exists too) and run it:
agents:
root:
model: openai/gpt-5-mini
description: A research assistant that can search the web
instruction: |
Answer questions concisely. Search the web when you need
fresh facts, and say where the information came from.
toolsets:
- type: mcp
ref: docker:duckduckgodocker agent run agent.yamlThat is the whole program: a model, a job description, and a toolset, where docker:duckduckgo pulls a web-search MCP server from Docker's catalog. No orchestration code, no SDK to learn, no glue to maintain.
Because the agent is a text file, it behaves like any other configuration. You can diff it, review it in a pull request, version it, and hand it to a colleague.
Bring your own model (or several)
Model choice is a config line, not an architecture decision. The docs list around thirty providers, including Docker Model Runner for local inference, so switching vendors means editing one field.
The more useful consequence is mixing them. Models are first-class config objects, so each agent can use the one that fits its job (excerpt):
models:
fast:
provider: openai
model: gpt-5-mini
local:
provider: dmr
model: ai/qwen3
agents:
analyst:
model: fast # cheap and precise
helper:
model: local # runs on your machineSpend on a premium model where reasoning matters, and use a small or local one where it doesn't.
Tools: batteries included, plus the whole MCP ecosystem
Docker Agent ships built-in tools for files, shell, git and HTTP fetching, plus think, todo, plan and memory for reasoning and state. Retrieval is built in too, with BM25, embeddings, hybrid search and reranking. Anything else comes through MCP, whether the server is local, remote or Docker-based.
The guiding principle, echoed in the docs' best practices, is least privilege. Give each agent only the tools its role requires. A reviewer reads files; it does not need a shell.
Multi-agent, with two distinct patterns
Docker Agent does not treat agents talking to agents as one thing. It separates two patterns, and the one you pick decides how context flows.
Delegation (sub_agents) is hierarchical. A coordinator hands a well-scoped task to a specialist through the built-in transfer_task tool. The specialist works in its own sub-session, starting from a clean task description, and the coordinator waits for the result before carrying on. Think tech lead and engineers.
Handoffs (handoffs) are peer-to-peer. The entire conversation moves to the next agent, which sees the full history, and the previous agent steps out of the loop. Think assembly line or support routing. Because the graph can loop, iterative workflows work too.
You can combine both in one config. Here is a small development team:
agents:
root:
model: anthropic/claude-sonnet-4-5
description: Technical lead coordinating the team
instruction: |
Break requests into tasks, delegate to the right specialist,
and review results before answering.
sub_agents: [developer, reviewer]
toolsets:
- type: think
developer:
model: anthropic/claude-sonnet-4-5
description: Writes and tests code
instruction: Write clean, well-tested code.
toolsets:
- type: filesystem
- type: shell
reviewer:
model: openai/gpt-5
description: Reviews code for quality and security
instruction: Give specific, actionable review feedback.
toolsets:
- type: filesystemTwo model vendors are on the team. The reviewer can read files but cannot execute anything. The lead picks specialists from their description fields, so those descriptions are effectively the routing API.
When you can't trust the model to remember
Handoffs depend on the model choosing to call a handoff tool, and models sometimes forget. For pipelines where order is non-negotiable, force_handoff takes the decision away from the LLM: when the agent finishes, the runtime itself routes the conversation to the named agent. The config loader rejects self-references and cycles up front, so you discover mistakes at load time instead of at 2 a.m.
agents:
root:
force_handoff: summarizer # runs after root's final responseThere is also hook-driven routing. The before_agent_run and after_agent_complete hooks can pick the next agent through a command or an evaluator, validated against an allow-list, so routing can be deterministic logic instead of yet another model call.
Fan out in parallel
transfer_task is sequential. When tasks are independent, add the background_agents toolset and the coordinator can dispatch several at once, monitor them, wait for results, or cancel the ones it no longer needs. Independent branches run side by side instead of waiting in line.
Orchestrate the coding CLIs you already use
A sub-agent can also be backed by an external coding CLI (Claude Code, Codex, opencode or pi) instead of a model API. Swap model: for a harness: block and the orchestrator delegates to it like any other sub-agent:
agents:
root:
model: anthropic/claude-sonnet-4-5
description: Plans work and delegates it
instruction: Break coding tasks down and delegate them to the coders.
sub_agents: [claude-coder, codex-coder]
claude-coder:
description: Claude Code specialist
harness:
type: claude-code
codex-coder:
description: Codex specialist
harness:
type: codexDocker Agent handles orchestration and hooks, while the external CLI runs its own coding loop. One planner can then route work across the tools your team already uses.
Ship agents like images
If the YAML is the Dockerfile, the registry is the Hub. Agents are pushed to any OCI-compatible registry and pulled or run from anywhere:
docker agent share push ./agent.yaml docker.io/yourorg/my-agent:v1
docker agent run yourorg/my-agent:v1Registry references compose. A published agent can be listed as a sub-agent beside local ones, and docker agent serve api can run straight from a registry and re-pull on an interval, so pushing a new tag becomes a rollout.
Trust, not just distribution
Sharing executable behavior demands provenance. Pushing with --key records a signature (or an HMAC when you use a shared secret) over an in-toto statement wrapped in a DSSE envelope. Pullers holding the matching key can verify that the YAML was published by you and has not been altered. The format is standard enough that existing in-toto tooling can read it without Docker Agent. An --encrypt option goes a step further.
The docs are candid about the limits. Only the YAML bytes and the published reference are authenticated. Surrounding manifest annotations are not, and serving an older signed version under the same tag goes undetected. Pin by digest (@sha256:...) when it matters, as you would with container images. Pinned references also load from cache, which saves a registry round-trip on every run.
Sandboxing a model that has a shell
Shell access is what makes coding agents useful and what makes security teams nervous. Docker Agent's answer is a single flag:
docker agent run --sandbox agent.yamlThe agent launches inside an isolated sandbox VM via Docker Sandboxes. Your working directory is mounted read-write, the agent's config is mounted read-only, and the rest of the host is invisible. Network egress passes through a default-deny proxy that permits the major model providers and the hosts your toolsets need, plus whatever you explicitly allow:
runtime:
sandbox: true
network_allowlist:
- api.example.comAgent authors can bake sandboxing into the YAML so every caller gets it by default. The same works at the alias level with docker agent alias add safe-coder myorg/coder --sandbox. An explicit --sandbox=false still wins when you need to debug on the host.
Host skills and prompt files (like AGENTS.md) are staged into a read-only kit and run through a secret-redaction pass, which the docs describe as best-effort. For heavier jobs, --cloud runs the agent in a fresh remote sandbox with a time limit (one hour by default, 24 at most) and deliberately does not upload your directory, config or API keys.
One honest note: Docker Agent orchestrates the sandbox CLI. It does not implement the isolation itself, and you need Docker Sandboxes installed to use the flag.
One definition, many entry points
Most of the remaining surface is about reuse, and it all hangs off the same YAML. An API server and a chat server put an agent behind an endpoint. MCP mode lets other MCP clients call your agents as tools. A2A is the route for agents built with other frameworks. A Go SDK covers embedding, and a headless guide covers CI.
You do not need any of that on day one. The piece worth setting up early is docker agent eval, which replays saved conversations in isolated containers and scores the results. Agents change whenever you edit a prompt or swap a model, and without a regression check you find out in production. The docs are clear that evals are for catching regressions, not for choosing the best configuration.
They build it with itself
The maintainers dogfood the project. Contributors run docker agent run ./golang_developer.yaml to bring up a Go-developer agent for the repository itself. At the time of writing, the repo has over ten thousand commits, around 3.9k stars and nearly 500 forks. Traces of its earlier name, cagent, still appear in a few config paths.
Try it in five minutes
# Docker Desktop 4.63+ already includes it. Otherwise:
brew install docker-agent # symlink into ~/.docker/cli-plugins/ to run it as docker agent
export ANTHROPIC_API_KEY=your-key-here # or any supported provider key
docker agent new # generate an agent interactively
docker agent run agent.yaml # run your config
docker agent share push ./agent.yaml docker.io/yourorg/first-agent:v1From there, the examples/ directory in the repo has runnable configs for handoffs, forced handoffs, hook routing and registry sub-agents.
An honest take
Docker Agent is not Docker for AI in the sense of solving agent engineering. It is a standardization layer around four things: how an agent is configured, how agents are orchestrated, how they are distributed, and where they execute. Prompt quality, tool design and evaluation are still your problem.
Where it shines: teams that want agents to be reviewable, reproducible and shareable; platform groups that need a story for isolation and provenance; anyone who wants to mix model vendors, or run locally, without rewriting code; and prototypers who would rather edit YAML than wire up a framework.
Worth keeping in mind:
The online docs track the
mainbranch and may describe unreleased features. Stable docs live at docs.docker.com/ai/docker-agent, and some details, such as how sandbox mode is described, differ between the two.Routing is deterministic, quality is not.
force_handoffguarantees the next agent runs. It says nothing about whether the previous agent did good work.A signature proves who published the YAML, not that the agent is safe to run. Read it, pin the digest and sandbox it.
Declarative config has a ceiling. For custom logic, reach for hooks and the Go SDK.
The tool collects anonymous usage data. The telemetry page in the docs explains the details.
Sandbox mode requires Docker Sandboxes to be set up first.
Containers did not make software correct either. They made it portable, repeatable and isolated, and that is the bar to hold Docker Agent to.
Links
Repository: https://github.com/docker/docker-agent
Documentation: https://docker.github.io/docker-agent/ (stable docs: https://docs.docker.com/ai/docker-agent/)
Multi-agent guide: https://docker.github.io/docker-agent/concepts/multi-agent/
Agent distribution: https://docker.github.io/docker-agent/concepts/distribution/
Sandbox mode: https://docker.github.io/docker-agent/configuration/sandbox/
Community: Docker Community Slack, #docker-agent channel
Comments (0)
Join the discussion by logging into your account.
No comments yet. Be the first to comment!