Skip to main content

HolmesGPT + Dapr Workflow

HolmesGPT is an open-source agent for SRE and on-call work — alert triage, log analysis, and incident investigation across your Kubernetes and observability stack. Make your HolmesGPT investigations production-ready by running them as durable Dapr workflows on Catalyst, so a long-running investigation survives process crashes, deploys, and tool-call timeouts.

HolmesGPT is supported through the same diagridio/python-ai durable-execution package as the other Python frameworks. Like every agent framework on Catalyst, it requires a workflow state store; the Conversation API is optional.


Why Dapr Workflow for HolmesGPT?​

ChallengeDapr Solution
Long investigations lost when a process crashesDurable workflows checkpoint every step
Chat memory lost on restartState store persists agent memory across restarts
Custom retry logic for flaky tool callsBuilt-in retries with exponential backoff
No visibility into what the agent didBuilt-in observability traces every tool call

Get started​

Install the HolmesGPT extra in its own environment. It ships strict dependency pins that conflict with the other agent extras, so it is intentionally not part of diagrid[all]:

pip install "diagrid[holmesgpt]"

Then follow the runnable HolmesGPT example in the python-ai repository, which wraps a HolmesGPT investigation in a Dapr workflow and runs it on Catalyst:

See also​

  • Develop AI agents — framework requirements table and the other supported frameworks
  • AI agents — the durable execution and identity model agents run under
  • Durable workflows — the underlying primitive agents are built on