# Durable Go agents

Build an agent as a graph of nodes, hand it to a `Runner`, and every node runs as a checkpointed activity in a Catalyst [durable workflow](https://docs.diagrid.io/concepts/durable-execution). If the process dies, the run resumes from the last completed node.

You will learn how to:

- Build an agent graph with the framework adapter of your choice
- Run it on Catalyst with the `diagrid` CLI
- Kill the agent mid-run and watch the same run recover

## Prerequisites

- Go 1.26.4 or later
- The [Diagrid CLI](https://docs.diagrid.io/references/catalyst/catalyst-cli-intro), logged in to a Catalyst project with a workflow state store and agent infrastructure enabled (see [Framework requirements](https://docs.diagrid.io/develop/agents#framework-requirements)). Cloud projects with the managed KV store have agent infrastructure enabled automatically; elsewhere, turn it on with `diagrid project update <your-project> --enable-agent-infrastructure`
- An `OPENAI_API_KEY` (both LangChainGo and Eino call OpenAI)

## 1. Get the library

The core library and each adapter are separate Go modules, so a framework's dependencies stay out of your build unless you use it:

**LangChainGo**

```bash
go get github.com/diagridio/go-ai
go get github.com/diagridio/go-ai/adapters/langchaingo
```

**Eino**

```bash
go get github.com/diagridio/go-ai
go get github.com/diagridio/go-ai/adapters/eino
go get github.com/cloudwego/eino-ext/components/model/openai
```

## 2. Build the graph

A node reads a channel from the shared state, calls the model, and writes its reply to another channel. The adapter turns your framework's chat model into that node. This agent has two: `diagnose` writes its reply to the `diagnosis` channel, and `reboot` reads that channel as its prompt:

**LangChainGo**

```go
model, err := openai.New(openai.WithModel("gpt-4o"))
if err != nil {
	return err
}

diagnose := langchaingo.ModelNode(model,
	langchaingo.WithSystemPrompt("You are the control room operator."),
	langchaingo.WithOutputKey("diagnosis"),
)

reboot := langchaingo.ModelNode(model,
	langchaingo.WithSystemPrompt("Reboot the grid and report the security code once systems are back online."),
	langchaingo.WithInputKey("diagnosis"),
	langchaingo.WithOutputKey("output"),
)

framework := langchaingo.Framework
```

**Eino**

```go
model, err := openai.NewChatModel(ctx, &openai.ChatModelConfig{
	APIKey: os.Getenv("OPENAI_API_KEY"),
	Model:  "gpt-4o",
})
if err != nil {
	return err
}

diagnose := eino.ChatModelNode(model,
	eino.WithSystemPrompt("You are the control room operator."),
	eino.WithOutputKey("diagnosis"),
)

reboot := eino.ChatModelNode(model,
	eino.WithSystemPrompt("Reboot the grid and report the security code once systems are back online."),
	eino.WithInputKey("diagnosis"),
	eino.WithOutputKey("output"),
)

framework := eino.Framework
```

Wire the nodes into a graph and hand it to the runner. This part is identical for every framework:

```go
graph := agent.NewGraph("control-room").
	AddNode("diagnose", diagnose).
	AddNode("reboot", reboot).
	SetEntry("diagnose").
	AddEdge("diagnose", "reboot").
	AddEdge("reboot", agent.END)

runner, err := goai.NewRunner(ctx, goai.Config{
	Graph:     graph,
	Name:      "control-room",
	Framework: framework,
	MaxSteps:  50,
})
if err != nil {
	return err
}
defer runner.Close()

out, err := runner.Invoke(ctx, agent.State{"input": "The park grid just went offline."},
	goai.InvokeOptions{InstanceID: "control-room-001"})
```

`NewRunner` connects to Catalyst, registers the agent in the agent registry, and compiles the graph into a workflow. There is no backend or state store to wire up yourself. The `framework` value comes from your adapter's `Framework` constant, set in the tab above (`langchaingo.Framework` or `eino.Framework`) — it is recorded in the registry and becomes part of the workflow name.

## 3. Run it on Catalyst

Create the Agent resource first — it appears under **Agents** in the Catalyst console and authorizes the runtime to register. Use the same name your app runs as:

```bash
diagrid agent create control-room --project <your-project> --wait
```

Then run your app with a managed sidecar:

```bash
diagrid dev run --project <your-project> --id control-room -- go run .
```

On startup the runner registers the agent, and each `Invoke` schedules the graph as a workflow: every node runs as a checkpointed activity, and routing decisions replay deterministically in the orchestrator.

## Recover from a crash

The `InstanceID` you pass to `Invoke` is caller-owned. Reuse it and Catalyst attaches to the existing run instead of starting a new one:

1. Start the agent with a fixed instance id and kill the process after a node has finished (watch the logs for its output, then Ctrl-C).
2. Start it again with the same id. The workflow resumes from the last completed node: a node that finished and checkpointed before the crash returns its recorded result on replay instead of running again, so it makes no duplicate model call.

This guarantee covers nodes that had already completed. A node that was still running when you killed the process — its call in flight, not yet checkpointed — is not guaranteed exactly-once and may run again on resume. Keep node work idempotent (or side-effect-free beyond the LLM call) if that timing matters for your agent.

The [runnable examples](https://github.com/diagridio/go-ai/tree/main/examples) wire this to the `GOAI_RUN` environment variable so you can try it without writing code — one example per framework, each with the full walkthrough in its README.

## Next steps

- [go-ai repository](https://github.com/diagridio/go-ai) — The full API — including conditional edges to route between nodes based on state.
- [Runnable examples](https://github.com/diagridio/go-ai/tree/main/examples) — One example per framework, each with a crash-recovery walkthrough in its README.
- [Develop AI agents](https://docs.diagrid.io/develop/agents) — See how other frameworks compare and what every integration requires.
