Durable Go agents
Build an agent as a graph of nodes, hand it to a Runner, and every node runs as a checkpointed activity in a Catalyst durable workflow. If the process dies, the run resumes from the last completed node.
You will learn how to:
- Build an agent graph with the framework adapter of your choice
- Run it on Catalyst with the
diagridCLI - Kill the agent mid-run and watch the same run recover
Prerequisites
- Go 1.26.4 or later
- The Diagrid CLI, logged in to a Catalyst project with a workflow state store and agent infrastructure enabled (see Framework requirements). Cloud projects with the managed KV store have agent infrastructure enabled automatically; elsewhere, turn it on with
diagrid project update <your-project> --enable-agent-infrastructure - An
OPENAI_API_KEY(both LangChainGo and Eino call OpenAI)
1. Get the library
The core library and each adapter are separate Go modules, so a framework's dependencies stay out of your build unless you use it:
- LangChainGo
- Eino
go get github.com/diagridio/go-ai
go get github.com/diagridio/go-ai/adapters/langchaingo
go get github.com/diagridio/go-ai
go get github.com/diagridio/go-ai/adapters/eino
go get github.com/cloudwego/eino-ext/components/model/openai
2. Build the graph
A node reads a channel from the shared state, calls the model, and writes its reply to another channel. The adapter turns your framework's chat model into that node. This agent has two: diagnose writes its reply to the diagnosis channel, and reboot reads that channel as its prompt:
- LangChainGo
- Eino
model, err := openai.New(openai.WithModel("gpt-4o"))
if err != nil {
return err
}
diagnose := langchaingo.ModelNode(model,
langchaingo.WithSystemPrompt("You are the control room operator."),
langchaingo.WithOutputKey("diagnosis"),
)
reboot := langchaingo.ModelNode(model,
langchaingo.WithSystemPrompt("Reboot the grid and report the security code once systems are back online."),
langchaingo.WithInputKey("diagnosis"),
langchaingo.WithOutputKey("output"),
)
framework := langchaingo.Framework
model, err := openai.NewChatModel(ctx, &openai.ChatModelConfig{
APIKey: os.Getenv("OPENAI_API_KEY"),
Model: "gpt-4o",
})
if err != nil {
return err
}
diagnose := eino.ChatModelNode(model,
eino.WithSystemPrompt("You are the control room operator."),
eino.WithOutputKey("diagnosis"),
)
reboot := eino.ChatModelNode(model,
eino.WithSystemPrompt("Reboot the grid and report the security code once systems are back online."),
eino.WithInputKey("diagnosis"),
eino.WithOutputKey("output"),
)
framework := eino.Framework
Wire the nodes into a graph and hand it to the runner. This part is identical for every framework:
graph := agent.NewGraph("control-room").
AddNode("diagnose", diagnose).
AddNode("reboot", reboot).
SetEntry("diagnose").
AddEdge("diagnose", "reboot").
AddEdge("reboot", agent.END)
runner, err := goai.NewRunner(ctx, goai.Config{
Graph: graph,
Name: "control-room",
Framework: framework,
MaxSteps: 50,
})
if err != nil {
return err
}
defer runner.Close()
out, err := runner.Invoke(ctx, agent.State{"input": "The park grid just went offline."},
goai.InvokeOptions{InstanceID: "control-room-001"})
NewRunner connects to Catalyst, registers the agent in the agent registry, and compiles the graph into a workflow. There is no backend or state store to wire up yourself. The framework value comes from your adapter's Framework constant, set in the tab above (langchaingo.Framework or eino.Framework) — it is recorded in the registry and becomes part of the workflow name.
3. Run it on Catalyst
Create the Agent resource first — it appears under Agents in the Catalyst console and authorizes the runtime to register. Use the same name your app runs as:
diagrid agent create control-room --project <your-project> --wait
Then run your app with a managed sidecar:
diagrid dev run --project <your-project> --id control-room -- go run .
On startup the runner registers the agent, and each Invoke schedules the graph as a workflow: every node runs as a checkpointed activity, and routing decisions replay deterministically in the orchestrator.
Recover from a crash
The InstanceID you pass to Invoke is caller-owned. Reuse it and Catalyst attaches to the existing run instead of starting a new one:
- Start the agent with a fixed instance id and kill the process after a node has finished (watch the logs for its output, then Ctrl-C).
- Start it again with the same id. The workflow resumes from the last completed node: a node that finished and checkpointed before the crash returns its recorded result on replay instead of running again, so it makes no duplicate model call.
This guarantee covers nodes that had already completed. A node that was still running when you killed the process — its call in flight, not yet checkpointed — is not guaranteed exactly-once and may run again on resume. Keep node work idempotent (or side-effect-free beyond the LLM call) if that timing matters for your agent.
The runnable examples wire this to the GOAI_RUN environment variable so you can try it without writing code — one example per framework, each with the full walkthrough in its README.