Productionizing agents is notoriously difficult, because of their long-running, stateful nature. This reference architecture shows how to build durable, stateful, steerable, concurrent agents and agentic systems.
The architecture fully runs on Restate and your favorite container platform or serverless provider. Restate is a durable runtime for agents that gives you all the building blocks you need to build advanced, large-scale agentic systems without managing a large infra stack.
agent-reference.mp4
Each agent execution is a durable async process that is:
- STEERABLE: Each execution has a handle that is stable across process restarts and can be used to steer it, interrupt it, or approve an action. Interrupts automatically propagate through subagents.
- STATEFUL: Execution is isolated per agent session and stateful. Transcripts and profile data (memories, instructions) are stored in Restate's embedded KV store.
- RECOVERABLE: Restate automatically keeps a journal per agent execution, to recover it automatically after a failure.
- CONCURRENT: Agents can spawn parallel subagents and tools. Tools execute as durable concurrent tasks within the process and can share resources like sandbox connections. Subagents run as separate durable invocations with their own state and resources.
- SCALABLE: Agents can scale up to thousands of concurrent executions, with protection against concurrency issues and race conditions.
- PAUSABLE: Restate can suspend an invocation waiting on a durable timer, approval signal, or child invocation, then resume it when ready. Suspension releases that invocation's execution on serverless platforms.
Implementing these characteristics in a production-grade manner is challenging and usually requires a lot of infra and coordination logic. In this reference architecture, the agent processes rely on Restate to handle this complexity:
This is a runnable TypeScript implementation with an agent runtime, a typed client, and an optional demo UI. You can use it as a starting point for your own application or take individual patterns into an existing agent. The reference focuses on agent execution and session management. User accounts, OAuth flows, and evals are not included.
| Feature | What you can do |
|---|---|
| Queueing, steering, and interruption | Queue a new request, add instructions to an ongoing execution, or interrupt it and optionally start a replacement. |
| Parallel tools | Run independent tool calls concurrently and return their results or errors to the model. |
| Background operations | Continue model steps while tool-requested timers or approvals are pending; wait for or cancel them later. |
| Subagents | Delegate work to persistent agents with their own history, memory, and sandbox, then send follow-up tasks to the same agents. |
| Guardrails | Define guardrails that guard against dangerous or unwanted tool behavior. |
| Human approval | Wait durably for a decision. The controller continues accepting steering and interruption; a guardrail wait holds its current proposal. |
| Programmatic tool calling | Let the model write JavaScript that combines tool calls and returns a result without putting all intermediate data in its context. |
| Tool search | Find tools from MCP servers and Restate services, loading their schemas into model context only when needed. |
| Memory | Store information across turns, search memory descriptions, and retrieve the entries relevant to the task. |
| Compaction | Summarize older conversation history in the background and compact long-running turns, while keeping recent context verbatim. |
| Schedules | Schedule one-off or recurring messages, with a policy to queue, steer, or interrupt when the agent is busy. |
| Sandboxes | Read and write files and run commands in a local workspace or Modal sandbox, keeping files between turns. |
| Client updates | Follow a running agent, reconnect after a connection failure, and read new history from an offset. |
| Extensible tools | Add built-in tools, connect MCP servers, or expose Restate handlers as tools. |
| Output recovery | Retry a truncated model response at most once with a larger output budget, if it is below the runtime's cap. |
Each agent has two Restate Virtual Objects (stateful entities addressed by key) keyed by the same agentId:
Agenthandles incoming messages and tracks the active turn, queued input, profile, memories, approvals, and schedules.AgentSessionstores the conversation and runs the model/tool loop. EachdoTurninvocation handles one turn, which may incorporate several queued messages or steering updates.
This architecture makes the following advanced features possible:
When an Agent starts a turn, it gets the turnId back, with which it can:
- Steer the turn: adds an instruction to the context for the next LLM call. Tools already running can finish.
- Interrupt the turn: stops unfinished work, including subagent tasks, and asks the model to summarize what it completed.
Restate provides durable signal delivery. The reference implements the routing, turn state machine, and cleanup rules on top of it.
See steering and interruption.
Each step in an agent loop is recorded in Restate's journal: LLM calls, guardrail checks, tool calls, state updates, approvals,... After a failure or a long wait, the agent process can recover to the exact step where it left off, by replaying the journal.
Restate only adds a few milliseconds of overhead to persist a journal entry, making fine-grained recovery feasible.
Tools run concurrently in the same process and share resources such as sandbox connections. Each tool's durable steps are recorded independently, so recovery can reuse completed work across the batch. The runtime handles deterministic replay during recovery.
See parallel tools and background operations.
Agent IDs run in parallel across service instances, while each session runs one turn at a time. Restate coordinates state access, routes calls, and recovers work after process failures, without the need for locks or coordination.
See state ownership.
This architecture uses Restate's embedded KV store for both conversation state and context. Restate gives each agent session its own isolated store, and ensures that only a single process can write to it at a time.
This architecture also implements:
- Selective loading: history is stored in chunks, and memories are retrieved on demand. Clients read only the history they need to update the UI, and the model sees only relevant memories.
- Compaction background jobs: Older messages are compacted in the background, to avoid exceeding the LLM's context limit. The model sees the latest messages verbatim, incl. the last summary. The full chat transcript is retained so the UI client can retrieve it, when needed.
See history and context.
When a guardrail needs a person, the turn suspends: no process, only stored state. It can wait weeks, through new versions of the service, and resumes where it stopped.
See approvals.
The reference UI uses a typed client to read session data and follow changes to history, approvals, configuration, and schedules. The UI uses HTTP long-polling and revision tags to fetch only what changed, while transcript sequence numbers let it catch up after disconnects.
See session updates.
- Architecture: state owners, one request, notifications
- Protocol: every handler, ordering rules, clients
- Turn runtime: steps, guardrails, pending work, recovery
- Tools: built-ins, programmatic tool calls, dynamic tools, MCP
- Schedules and sandboxes
- Configuration: models, tools, environment variables, MCP servers and the reference UI
- Development: tests, debugging, packaging and the repository layout; PROJECT.md gives a file-by-file reading order
- Agent guide: read before changing runtime semantics
- AG-UI: connect AG-UI frontends such as CopilotKit
The documentation index has the full list.
plugins/restate-agent has two skills:
restate-agent: how to extend this agent. It covers tools, handlers, configuration and testing, and how to grow the agent into a full application with users, sessions and credentials.restate-gen-sdk: how to write the generator-SDK code the agent is built from.
The plugin bundles both skills and the Restate docs MCP server. Codex and Claude Code share the same skill files and MCP configuration.
For Codex, install the plugin from a terminal:
codex plugin marketplace add restatedev/agent
codex plugin add restate-agent@restate-agentStart a new Codex chat after installation. Opening the repository alone does
not enable the plugin. To install from a local checkout, run
codex plugin marketplace add . from the repository root instead of the
first command above. Codex supports the existing
marketplace catalog and uses the
Codex manifest to load the
skills and MCP server.
Claude Code offers to install the plugin when you open this repository. You can also install it by hand:
/plugin marketplace add restatedev/agent
/plugin install restate-agent@restate-agentFor other coding agents, install just the skills with
npx skills add restatedev/agent. That command does not configure the
Restate docs MCP server.
You need Node.js 22+, pnpm, the Restate server and CLI, and an OpenAI API key. The agent can also run on Anthropic, Google, xAI or DeepSeek models, or on open models through Ollama, vLLM or any OpenAI-compatible server; see models.
pnpm install
# Terminal 1: Restate, with the SDK features this example uses
RESTATE_EXPERIMENTAL_ENABLE_PROTOCOL_V7=true restate-server
# Terminal 2: the agent service on port 9080, registered with Restate
export OPENAI_API_KEY=your-api-key
pnpm dev:service
restate deployments register http://localhost:9080Talk to an agent through Restate ingress on port 8080. Any agent ID works; the first message creates the agent.
# Start a turn (or queue the message if one is running)
curl localhost:8080/Agent/demo/ask --json '{"message":"What is the weather in Berlin?"}'
# Redirect the running turn without cancelling its tools
curl localhost:8080/Agent/demo/steer --json '{"message":"Use Fahrenheit"}'
# Stop it, optionally queueing a replacement request
curl localhost:8080/Agent/demo/interrupt --json '{"reason":"Changed my mind"}'
# Read the conversation log
curl localhost:8080/AgentSession/demo/history --json '{"fromSequence":1,"limit":100}'Or chat in the reference UI: run pnpm dev:ui and open
http://127.0.0.1:3000/?agent=demo.
Try this: ask it to sleep for four minutes, then kill pnpm dev:service
and start it again. The turn picks up where it was, and the Restate UI
(http://localhost:9070) shows every model call and tool result in the
turn's journal.
