Azure deployment
This guide deploys a Catalyst Enterprise Self-Hosted region on Azure. Terraform creates an AKS cluster, a PostgreSQL Flexible Server, and the public IP address and DNS zone for the region. Then you install the Catalyst Helm chart on the cluster.
This guide is intended as a reference walkthrough for demonstration and proof-of-concept deployments. It is not a production-ready deployment recipe — use the Production Planning guidance and adapt this guide to your organization's networking, security, and operational requirements before going to production.
Architecture overview
The Terraform creates one resource group with a virtual network, an AKS cluster, and a PostgreSQL Flexible Server. Clients reach the Catalyst gateway over TLS, through a static public IP address and a wildcard domain in Azure DNS. The Catalyst agent in the cluster opens an mTLS connection to the Catalyst control plane in Diagrid Cloud, and the control plane manages the region over that connection.
The Terraform in guides/azure/terraform creates:
- Resource group, named
<cluster_name>-rg(catalyst-rgby default). - Virtual network (
10.0.0.0/16by default) with one subnet for the AKS nodes and one subnet delegated to PostgreSQL Flexible Server. - AKS cluster with Azure CNI, a Standard load balancer, an autoscaling node pool spread across availability zones, Entra ID authentication with Azure RBAC, and the OIDC issuer and workload identity turned on.
- PostgreSQL Flexible Server with private networking only. It has no public endpoint, so only resources inside the virtual network can reach it. It holds the Catalyst state store, the workflow data shown in the Catalyst console, and the scheduler's jobs and reminders.
- Gateway public IP, a static Standard IP address for the Catalyst gateway. It stays the same when you reinstall the chart.
- Azure DNS zone for the region's domain, with a wildcard record that points at the gateway public IP.
- cert-manager identity, a user-assigned managed identity federated to cert-manager's service account. It has permission to write the DNS-01 challenge records into the zone.
The same Terraform builds the regions of the Azure multi-region deployment. With its default settings, it builds the single standalone region this guide describes.
Prerequisites
- Azure CLI, signed in with
az loginto the subscription you deploy to - Terraform 1.9 or later, and
make - Diagrid CLI
- Helm and kubectl
kubelogin. Install it withaz aks install-cli. Don't usebrew install kubelogin: it installs a different tool with the same name.- jq
- Owner on the subscription, or Contributor plus User Access Administrator. The Terraform creates role assignments, and Contributor alone can't do that.
- Enough vCPU quota for the AKS nodes and the PostgreSQL server in the Azure region you deploy to. By default, that's two to five
Standard_D4s_v5nodes and oneGP_Standard_D2ds_v5server. Some subscriptions have no quota at all for theStandard DSv5 Family. In that case, setnode_instance_typeto a size your subscription can use. See Troubleshooting. - A domain you control, such as
catalyst.example.com. The Terraform creates an Azure DNS zone for it, and you delegate the domain to that zone at your registrar.
The shell commands in this guide assume a bash-compatible shell. On Windows, run them in WSL.
Clone the repository that holds the Terraform. Run every make command in this guide from charts/guides/azure:
git clone https://github.com/diagridio/charts.git
cd charts/guides/azure
1. Create a Catalyst region
Choose the region's domain and Azure region. Then register the region in Diagrid Cloud, with that domain as its ingress:
export INGRESS_DOMAIN="catalyst.example.com"
export LOCATION="westus2"
export PRIVATE_REGION="azure-region"
diagrid login
export JOIN_TOKEN=$(diagrid region create $PRIVATE_REGION \
--ingress "https://*.$INGRESS_DOMAIN:443" -o json | jq -r .joinToken)
diagrid region create fails if a region with that name already exists.
2. Configure the Terraform
Create the Terraform variables file. The command reads your subscription ID, tenant ID, and Entra ID object ID from the Azure CLI:
cat > terraform/terraform.tfvars << EOF
subscription_id = "$(az account show --query id -o tsv)"
tenant_id = "$(az account show --query tenantId -o tsv)"
location = "$LOCATION"
region_ingress_endpoint = "$INGRESS_DOMAIN"
# Who can use the cluster. Terraform grants no other access, so an empty list
# gives you a cluster nobody can use.
aks_admin_principal_ids = ["$(az ad signed-in-user show --query id -o tsv)"]
EOF
Choose a password for the PostgreSQL administrator. Pass it to Terraform as an environment variable, not in the variables file:
export PGPASSWORD="<a strong password>"
export TF_VAR_postgresql_password="$PGPASSWORD"
The password needs at least eight characters, from three of these four groups: uppercase letters, lowercase letters, digits, and symbols.
To match one of Catalyst's standard region sizes, copy the values from one of the tier files into terraform/terraform.tfvars. For other settings, see Configuration options.
3. Deploy the Azure resources
make init
make plan
make apply
make apply shows the plan and asks you to confirm it. Type yes. The apply takes 10 to 20 minutes, mostly for the AKS cluster and the PostgreSQL server.
When the apply finishes, delegate your domain to Azure DNS. First, get the zone's name servers:
make output dns_zone_name_servers
At your registrar, or in the parent zone's DNS, add NS records for $INGRESS_DOMAIN that point at those four name servers. Before you continue, check that the delegation works. The certificate in step 5 can't be issued until it does:
dig NS $INGRESS_DOMAIN +short
When the delegation is live, the command prints the Azure name servers.
4. Connect to the cluster
az aks get-credentials \
--resource-group $(make output resource_group_name) \
--name $(make output aks_cluster_name)
kubelogin convert-kubeconfig -l azurecli
kubectl get nodes
In this guide the AKS API server is public, so you can run kubectl and helm from your own machine. To limit who can reach it, set api_server_authorized_ip_ranges. See Configuration options.
5. Issue the wildcard certificate
The gateway serves TLS for *.$INGRESS_DOMAIN from a secret named cert-wildcard, and it isn't ready until that secret exists. cert-manager issues the certificate with a DNS-01 challenge, using the Azure identity the Terraform created for it.
Install cert-manager with workload identity turned on:
export CERT_MANAGER_CLIENT_ID=$(make output cert_manager_identity_client_id)
export DNS_ZONE_RG=$(make output cert_manager_dns_zone_resource_group_name)
helm repo add jetstack https://charts.jetstack.io --force-update
helm upgrade --install cert-manager jetstack/cert-manager \
--namespace cert-manager --create-namespace \
--set crds.enabled=true \
--set "serviceAccount.annotations.azure\.workload\.identity/client-id=$CERT_MANAGER_CLIENT_ID" \
--set-string "podLabels.azure\.workload\.identity/use=true"
Create a Let's Encrypt ClusterIssuer that solves the challenge in your Azure DNS zone. Replace <your-email-address> with your own email address:
cat <<EOF | kubectl apply -f -
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
name: letsencrypt
spec:
acme:
server: https://acme-v02.api.letsencrypt.org/directory
email: <your-email-address>
privateKeySecretRef:
name: letsencrypt-account-key
solvers:
- dns01:
azureDNS:
hostedZoneName: $INGRESS_DOMAIN
resourceGroupName: $DNS_ZONE_RG
subscriptionID: $(az account show --query id -o tsv)
environment: AzurePublicCloud
managedIdentity:
clientID: $CERT_MANAGER_CLIENT_ID
EOF
Request the certificate in the cra-agent namespace, where you install Catalyst in step 7. Create the namespace first. cert-manager never issues a Certificate in a namespace that doesn't exist:
kubectl create namespace cra-agent
cat <<EOF | kubectl apply -f -
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: cert-wildcard
namespace: cra-agent
spec:
secretName: cert-wildcard
issuerRef:
name: letsencrypt
kind: ClusterIssuer
dnsNames:
- "*.$INGRESS_DOMAIN"
EOF
kubectl wait --for=condition=Ready certificate/cert-wildcard -n cra-agent --timeout=5m
The first certificate usually takes a few minutes. If it stays False, run kubectl describe certificaterequest -n cra-agent to see why. A challenge stuck in presenting almost always means the DNS delegation from step 3 isn't live yet.
6. Let the database user stream changes
The Catalyst scheduler stores its jobs and reminders in PostgreSQL and reads changes over a logical replication connection. For that, the database administrator needs the REPLICATION attribute. Flexible Server doesn't grant it by default, and Terraform can't set it.
Only resources inside the virtual network can reach the server, so run the grant from a pod in the cluster:
kubectl run grant-replication --rm -i --restart=Never -n default \
--image=postgres:17-alpine --env="PGPASSWORD=$PGPASSWORD" --command -- \
psql "host=$(make output postgresql_endpoint) user=postgres dbname=catalyst sslmode=require" \
-c "ALTER ROLE postgres WITH REPLICATION;"
The command prints ALTER ROLE.
7. Configure and install Catalyst
Create a Helm values file. It connects Catalyst to the PostgreSQL server and points the gateway at the public IP the Terraform created:
export PG_ENDPOINT=$(make output postgresql_endpoint)
export RESOURCE_GROUP=$(make output resource_group_name)
export GATEWAY_PIP_NAME=$(make output gateway_public_ip_name)
cat > catalyst-values.yaml << EOF
agent:
config:
project:
default_managed_state_store_type: postgresql-shared-external
external_postgresql:
enabled: true
auth_type: connectionString
namespace: postgresql
connection_string_host: $PG_ENDPOINT
connection_string_port: 5432
connection_string_username: postgres
connection_string_password: "$PGPASSWORD"
gateway:
tls:
enabled: true
secretName: "cert-wildcard"
envoy:
service:
type: LoadBalancer
httpsPort: 443
httpsTargetPort: 8443
annotations:
# Use the static public IP the Terraform created for this region.
service.beta.kubernetes.io/azure-load-balancer-resource-group: "$RESOURCE_GROUP"
service.beta.kubernetes.io/azure-pip-name: "$GATEWAY_PIP_NAME"
EOF
This example stores Catalyst secrets in Kubernetes Secrets, the default. For other options, see the Helm chart reference.
Install Catalyst:
helm install catalyst oci://public.ecr.aws/diagrid/catalyst \
-n cra-agent \
-f catalyst-values.yaml \
--set join_token="${JOIN_TOKEN}"
Wait for all pods to be ready. This can take a few minutes:
kubectl -n cra-agent wait --for=condition=ready pod --all --timeout=5m
kubectl wait checks pods one at a time. If one pod isn't ready, it prints timed out for every pod after it, even pods that are ready. Run kubectl -n cra-agent get pods to see the real state of each pod.
When the agent connects, diagrid region list shows the region as online.
8. Create a project and test it
Create a project in the new region, and two apps to test the connection:
export PROJECT_NAME="azure-project"
diagrid project create $PROJECT_NAME --region $PRIVATE_REGION --use
diagrid app create app1
diagrid app create app2
# Wait until the apps are ready
diagrid app list
The gateway is public, so you can test from your own machine. In a second terminal, start a listener for app1:
diagrid project use azure-project
# Wait until you see:
# ✅ Connected App ID "app1" to http://localhost:<port> ⚡️
diagrid listen -a app1
In the first terminal, send a service invocation request from app2 to app1:
diagrid call invoke get app1.hello -a app2
The listener prints the requests it received. The first one, to /dapr/config, comes from the sidecar when it connects. The second one is your call:
{
"method": "GET",
"url": "/dapr/config"
}
{
"method": "GET",
"url": "/hello"
}
The request went over TLS through the region's gateway at the wildcard domain. This shows that apps in the region can call each other with Dapr's service invocation API. In this test, the Diagrid CLI plays both apps. For more, see Test Catalyst APIs using the Diagrid CLI.
To open the Catalyst console for the project, run diagrid web. To build and deploy your own apps, continue with the Connect to Catalyst guide.
Tear the region down
If you followed Managed identities or Catalyst identity, remove the identities first. Then remove the region.
Remove the identities
Skip this step if you followed neither section. The identities, and the roles they hold on your own Azure resources, live outside the Terraform, so make destroy doesn't remove them. This applies whether you ran the steps by hand or with the helper scripts.
Remove each identity's role assignments before you delete the identity. If you delete the identity first, its role assignments stay on your Key Vault, storage account, and other resources as Identity not found entries.
If you followed Managed identities, remove the catalyst identity. Run this from charts/guides/azure, before make destroy:
RESOURCE_GROUP=$(make output resource_group_name)
PRINCIPAL_ID=$(az identity show --name catalyst --resource-group "$RESOURCE_GROUP" \
--query principalId --output tsv)
# Every Azure role the identity holds in the subscription, such as the
# Key Vault and storage roles from step 5
az role assignment list --assignee "$PRINCIPAL_ID" --all --query "[].id" --output tsv \
| xargs -r az role assignment delete --ids
# Also removes the identity's federated credentials
az identity delete --name catalyst --resource-group "$RESOURCE_GROUP"
If you followed Catalyst identity, remove the project's app registration:
APP_ID=$(az ad app list --display-name "catalyst-azure-project" --query "[0].appId" --output tsv)
SP_OBJECT_ID=$(az ad sp show --id "$APP_ID" --query id --output tsv)
az role assignment list --assignee "$SP_OBJECT_ID" --all --query "[].id" --output tsv \
| xargs -r az role assignment delete --ids
# Also removes the service principal and the federated credentials
az ad app delete --id "$APP_ID"
az role assignment list --all only covers the current subscription. If you granted roles in another subscription, run it there too. If you ran a helper script with --cosmosdb, the script also granted a Cosmos DB data-plane role, which az role assignment doesn't list. Remove it for each principal, with PRINCIPAL_ID or SP_OBJECT_ID from above:
az cosmosdb sql role assignment list \
--account-name "<cosmosdb-account>" --resource-group "<cosmosdb-resource-group>" \
--query "[?principalId=='$PRINCIPAL_ID'].name" --output tsv \
| xargs -r -n1 az cosmosdb sql role assignment delete --yes \
--account-name "<cosmosdb-account>" --resource-group "<cosmosdb-resource-group>" \
--role-assignment-id
Remove the region
Delete the projects first, then uninstall the chart. This way the gateway releases the public IP while the cluster still exists. Then destroy the infrastructure and remove the region from Diagrid Cloud:
diagrid project delete azure-project
helm uninstall catalyst -n cra-agent
kubectl -n cra-agent wait --for=delete svc/gateway-envoy --timeout=5m
make destroy
diagrid region delete $PRIVATE_REGION
diagrid project delete and make destroy ask you to confirm, and diagrid region delete asks you to type the region's name. Like make apply, make destroy needs TF_VAR_postgresql_password to be set.
make destroy fails if the catalyst identity from Managed identities still exists. Terraform won't delete a resource group that holds resources it doesn't manage. Remove the identity as shown in Remove the identities, then run make destroy again.
When the destroy is done, remove the domain's NS records at your registrar.
Troubleshooting
| Symptom | Cause and fix |
|---|---|
An apply fails with MissingSubscriptionRegistration | The subscription is new and has no resource providers registered. Run az provider register -n <namespace> for Microsoft.Compute, Microsoft.ContainerService, Microsoft.DBforPostgreSQL, Microsoft.Network, Microsoft.OperationalInsights, and Microsoft.ManagedIdentity. Wait until az provider show -n <namespace> --query registrationState reports Registered. |
| The apply fails on vCPU quota | Each region has a quota per VM family and a total quota. az vm list-usage -l <region> -o table shows both. If your node size's family has no row, your subscription can't use that size, and you need a different one. |
| The cluster create fails on zone placement | Your subscription can't use the node size in the region's zones. If az vm list-skus -l <region> --size <size> --zone -o table returns nothing, that's the cause. Set availability_zones = [], or set it to the zones the command lists. |
HA is disabled for region <region> | The region doesn't offer zone-redundant PostgreSQL high availability to your subscription. Set postgresql_high_availability = false. |
K8sVersionNotSupported | The default cluster_version is now available only with Long-Term Support. Set cluster_version to a version shown as KubernetesOfficial in az aks get-versions -l <region> -o table. |
kubectl fails with executable kubelogin not found | Install kubelogin with az aks install-cli, not Homebrew, and run kubelogin convert-kubeconfig -l azurecli. |
kubectl fails with User does not have access to the resource in Azure | Your object ID isn't in aks_admin_principal_ids. Add it and apply again. If it was already there, kubelogin has cached an old token. Run rm -rf ~/.kube/cache/kubelogin. |
helm install cert-manager fails with cannot unmarshal bool | The pod label was passed with --set. Use --set-string instead. |
The catalyst-gateway pod stays in Init with FailedMount warnings for cert-wildcard | The wildcard certificate from step 5 isn't issued yet. Check it with kubectl get certificate -n cra-agent. |
The scheduler never becomes ready, and the agent restarts every five minutes with permission denied to start WAL sender | The REPLICATION grant is missing. |
Managed identities
Catalyst apps on AKS can use Azure Workload Identity to authenticate to Azure services, such as Azure Key Vault or Azure Storage, without storing long-lived secrets. The whole region shares one user-assigned managed identity, which is federated to each app's sidecar service account. At runtime, the sidecar and your app exchange the service account token for an Azure AD token.
The Terraform already turns on the OIDC issuer and workload identity on the AKS cluster.
Entra ID only accepts sidecar tokens that have the api://AzureADTokenExchange audience. The Catalyst certificate authority issues those tokens, so configure it to add that audience. Add the following to catalyst-values.yaml from step 7 and upgrade the release:
global:
sentry:
jwt_audiences: ["api://AzureADTokenExchange"]
How it works
- All apps in the region share one user-assigned managed identity, named
catalyst. - Each app's sidecar runs under its own Kubernetes service account, named
sidecar-<project-uid>-<appid-uid>in theprj-<project-uid>namespace. - A federated credential links that service account to the managed identity through the cluster's OIDC issuer. Azure then trusts tokens issued for the service account.
- You run
diagrid appid updateto give each app a label and an annotation. The label turns on Azure token injection for the sidecar pod, and the annotation holds the identity's client ID.
The setup-user-managed-catalyst-identity.sh helper script in guides/azure runs the steps below for one app. Run it with --project and --app, and optionally --keyvault or --storage-account. Its --resource-group, --cluster-name, and --location defaults match the Terraform defaults. Or follow the steps below by hand.
Configure a managed identity manually
These steps do what the helper script does, for one app. You create the catalyst identity (step 3) and its role assignments (step 5) once per region. For each extra app, repeat steps 2, 4, and 6, because each app needs its own federated credential, label, and annotation.
1. Set environment variables
Run these from charts/guides/azure, so make output can read the region's Terraform outputs:
export RESOURCE_GROUP=$(make output resource_group_name)
export LOCATION=$(make output aks_cluster_region)
export CLUSTER_NAME=$(make output aks_cluster_name)
export CATALYST_PROJECT="azure-project"
export CATALYST_APP="app1"
2. Resolve the sidecar service account
The sidecar's service account name and namespace come from the project and app UIDs. Look them up with the Diagrid CLI:
PROJECT_UID=$(diagrid project get "$CATALYST_PROJECT" --output json | jq -r '.metadata.uid')
APPID_UID=$(diagrid appid get "$CATALYST_APP" --project "$CATALYST_PROJECT" --output json | jq -r '.metadata.uid')
SERVICE_ACCOUNT_NAMESPACE="prj-${PROJECT_UID}"
SERVICE_ACCOUNT_NAME="sidecar-${PROJECT_UID}-${APPID_UID}"
3. Create the managed identity
Create the shared managed identity, unless it already exists, and save its client ID and principal ID:
az identity create \
--name catalyst \
--resource-group "$RESOURCE_GROUP" \
--location "$LOCATION"
USER_ASSIGNED_CLIENT_ID=$(az identity show \
--name catalyst --resource-group "$RESOURCE_GROUP" \
--query clientId --output tsv)
PRINCIPAL_ID=$(az identity show \
--name catalyst --resource-group "$RESOURCE_GROUP" \
--query principalId --output tsv)
4. Create the federated credential
Get the cluster's OIDC issuer URL, and federate the sidecar service account to the identity:
AKS_OIDC_ISSUER=$(make output aks_oidc_issuer_url)
az identity federated-credential create \
--name "catalyst-${PROJECT_UID}-${APPID_UID}" \
--identity-name catalyst \
--resource-group "$RESOURCE_GROUP" \
--issuer "$AKS_OIDC_ISSUER" \
--subject "system:serviceaccount:${SERVICE_ACCOUNT_NAMESPACE}:${SERVICE_ACCOUNT_NAME}" \
--audience api://AzureADTokenExchange
5. (Optional) Grant access to Azure resources
Give the identity the data-plane roles it needs. For example, read access to a Key Vault and read/write access to a storage account:
# Key Vault Secrets User
KEYVAULT_SCOPE=$(az keyvault show --name "<keyvault-name>" --query id --output tsv)
az role assignment create \
--assignee-object-id "$PRINCIPAL_ID" \
--assignee-principal-type ServicePrincipal \
--role "Key Vault Secrets User" \
--scope "$KEYVAULT_SCOPE"
# Storage Blob Data Contributor
STORAGE_SCOPE=$(az storage account show --name "<storage-account-name>" --query id --output tsv)
az role assignment create \
--assignee-object-id "$PRINCIPAL_ID" \
--assignee-principal-type ServicePrincipal \
--role "Storage Blob Data Contributor" \
--scope "$STORAGE_SCOPE"
6. Enable workload identity on the app
Add the workload identity label and annotation to the app with the Diagrid CLI, using the client ID from step 3. Diagrid applies them to the app's sidecar pod and service account. If diagrid listen from step 8 is still running for the app, stop it first. The update fails while a local connection holds the app:
diagrid appid update "$CATALYST_APP" \
--project "$CATALYST_PROJECT" \
--label azure.workload.identity/use=true \
--annotation azure.workload.identity/client-id="$USER_ASSIGNED_CLIENT_ID"
The azure.workload.identity/use=true label turns on Azure token injection for the sidecar pod. The azure.workload.identity/client-id annotation tells the workload identity webhook which managed identity to use.
--label and --annotation replace the app's existing workload labels and annotations, so pass every label and annotation the app needs in one command. Both flags work only in self-hosted regions.
All apps share the catalyst identity, so every app uses the same client ID. Only the federated credential from step 4 is different for each app. You still need to label and annotate each app. After the sidecar restarts, it runs as the managed identity and can access any Azure resource the identity has permission for.
Grant the management service access
A workflow or agent state store can reference a secret in your own secret store, such as an Azure Key Vault credential in a secretKeyRef. The regional Catalyst management service reads that secret itself when it serves workflow and agent data. It authenticates with its own Kubernetes service account, not an app's, so the shared catalyst identity must trust that service account too.
Create a federated credential for the management service account. For the installation in this guide, that's catalyst-management-sa in the cra-agent namespace:
az identity federated-credential create \
--name catalyst-management \
--identity-name catalyst \
--resource-group "$RESOURCE_GROUP" \
--issuer "$AKS_OIDC_ISSUER" \
--subject "system:serviceaccount:cra-agent:catalyst-management-sa" \
--audience api://AzureADTokenExchange
Management is part of the Catalyst Helm release, not a Catalyst app. So you turn on workload identity for it in the Helm values, not with diagrid appid update. Add the following to catalyst-values.yaml from step 7, using the client ID from step 3, and upgrade the release. The merge.selectorLabels entry adds the label to the management pods. It only changes the pod template and leaves the deployment's selector alone, so the upgrade is safe:
management:
merge:
selectorLabels:
azure.workload.identity/use: "true"
podAnnotations:
azure.workload.identity/client-id: "<client-id>"
serviceAccount:
annotations:
azure.workload.identity/client-id: "<client-id>"
The federated credential lets management sign in, but the identity still needs permission on the resources management reads. That includes the secret store the secretKeyRef points to, for example Key Vault Secrets User on the vault. Management uses the same catalyst identity as the apps, so the role assignments from step 5 cover both. The federated credential covers the whole region, not one app, so you set it up once. The setup-user-managed-catalyst-identity.sh helper script creates the federated credential and prints the Helm values to apply.
Catalyst identity
Catalyst apps can also authenticate to Azure services with their own Catalyst workload identity: the SPIFFE identity Catalyst gives each app. This way they don't borrow an AKS service account. The Azure federated credential trusts the Catalyst region's OIDC issuer and the app's SPIFFE ID, not the AKS cluster's OIDC issuer. At runtime, the sidecar exchanges its JWT-SVID for an Azure AD token, on whichever cluster it runs. Use this approach when you want the Azure identity tied to the Catalyst app itself, not to the AKS service account its sidecar runs under.
Entra ID only accepts federated tokens whose audience matches the federated credential's audience. So the JWT-SVIDs the sidecar presents must have the api://AzureADTokenExchange audience. Add the same global.sentry.jwt_audiences setting shown in Managed identities to catalyst-values.yaml and upgrade the release.
How it works
- An Azure AD application (app registration) for the project, named
catalyst-<project>, backs the identity. Your Azure component authenticates as its application (client) ID. - Each app has its own Catalyst workload identity: one SPIFFE ID for each region host its sidecar can run on. You find them in the app's status (
.status.spiffeIds). - A federated credential on the app registration trusts the Catalyst region's OIDC issuer and one of the app's SPIFFE IDs. Azure then accepts the JWT-SVID the sidecar creates for that identity, with the audience
api://AzureADTokenExchange. - The app registration's service principal gets the Azure roles the workload needs.
- You configure the Azure component with the app's client ID and tenant ID. It doesn't need a client secret.
The setup-federated-catalyst-identity.sh helper script in guides/azure runs the steps below for one app. Run it with --project and --app, and optionally --keyvault or --storage-account. It looks up the OIDC issuer from the project's region. Or follow the steps below by hand.
Configure a federated identity manually
These steps do what the helper script does, for one app. You create the project's app registration (step 3) and its role assignments (step 5) once per project, and all its apps share them. For each extra app, repeat steps 2 and 4, because each app needs its own federated credential.
1. Set environment variables
export CATALYST_PROJECT="azure-project"
export CATALYST_APP="app1"
# One app registration per project
export APP_DISPLAY_NAME="catalyst-${CATALYST_PROJECT}"
# The Catalyst OIDC issuer URL, resolved from the project's region
REGION_ID=$(diagrid project get "$CATALYST_PROJECT" --output json | jq -r '.spec.region')
export ISSUER=$(diagrid region get "$REGION_ID" --output json | jq -r '.status.endpoints.oidc')
2. Resolve the app's SPIFFE identities
A sidecar can run on more than one region host, and it has a different SPIFFE ID on each. Read them from the app's status with the Diagrid CLI, one per line:
SPIFFE_IDS=$(diagrid appid get "$CATALYST_APP" --project "$CATALYST_PROJECT" \
--output json | jq -r '.status.spiffeIds // [.status.spiffeId] | .[]')
3. Create the app registration
Create the project's app registration and its service principal, unless they already exist, and save the application ID:
APP_ID=$(az ad app list --display-name "$APP_DISPLAY_NAME" --query "[0].appId" --output tsv)
[ -z "$APP_ID" ] && APP_ID=$(az ad app create --display-name "$APP_DISPLAY_NAME" --query appId --output tsv)
SP_OBJECT_ID=$(az ad sp show --id "$APP_ID" --query id --output tsv 2>/dev/null \
|| az ad sp create --id "$APP_ID" --query id --output tsv)
4. Create a federated credential per SPIFFE ID
Federate each of the app's SPIFFE IDs to the app registration, trusting the region's OIDC issuer:
echo "$SPIFFE_IDS" | while IFS= read -r SUBJECT; do
AUTHORITY=${SUBJECT#spiffe://}; HOST_LABEL=${AUTHORITY%%.*}
cat > fic.json <<EOF
{
"name": "${CATALYST_APP}-${HOST_LABEL}",
"issuer": "${ISSUER}",
"subject": "${SUBJECT}",
"audiences": ["api://AzureADTokenExchange"]
}
EOF
az ad app federated-credential create --id "$APP_ID" --parameters @fic.json
done
5. (Optional) Grant access to Azure resources
Give the app registration's service principal the data-plane roles it needs. For example, read access to a Key Vault and read/write access to a storage account:
# Key Vault Secrets User
KEYVAULT_SCOPE=$(az keyvault show --name "<keyvault-name>" --query id --output tsv)
az role assignment create \
--assignee-object-id "$SP_OBJECT_ID" \
--assignee-principal-type ServicePrincipal \
--role "Key Vault Secrets User" \
--scope "$KEYVAULT_SCOPE"
# Storage Blob Data Contributor
STORAGE_SCOPE=$(az storage account show --name "<storage-account-name>" --query id --output tsv)
az role assignment create \
--assignee-object-id "$SP_OBJECT_ID" \
--assignee-principal-type ServicePrincipal \
--role "Storage Blob Data Contributor" \
--scope "$STORAGE_SCOPE"
6. Configure the Azure component
Configure the Azure component the app uses with the application's client ID and tenant ID. The federated credential supplies the token, so you don't need azureClientSecret:
TENANT_ID=$(az account show --query tenantId --output tsv)
diagrid component create my-keyvault \
--type secretstores.azure.keyvault \
--project "$CATALYST_PROJECT" \
--metadata vaultName="<keyvault-name>" \
--metadata azureClientId="$APP_ID" \
--metadata azureTenantId="$TENANT_ID" \
--scopes "$CATALYST_APP"
Federated identities depend on the app's SPIFFE workload identity, which only exists in self-hosted regions. Run step 4 again whenever the app's region hosts change, so each SPIFFE ID has a matching federated credential.
Once the component is scoped to the app, the sidecar exchanges the app's SPIFFE ID for an Azure AD token. It can then access any Azure resource the app registration's service principal has permission for. The credential belongs to the Catalyst app, not to an AKS service account, so the same Azure identity works wherever the sidecar runs.
Federate the management service identity
A workflow or agent state store can reference a secret in your own secret store, such as an Azure Key Vault credential in a secretKeyRef. The regional Catalyst management service reads that secret itself when it serves workflow and agent data. It authenticates with its own Catalyst identity, not an app's. That identity is a SPIFFE ID in the same trust domain as the app's region host, with a different path: spiffe://<region-trust-domain>/ns/cra-agent/management. Federate it on the project's app registration so management can read the store.
Build the management subjects from the app's SPIFFE IDs in step 2, one per region host, and create a federated credential for each:
echo "$SPIFFE_IDS" | while IFS= read -r SUBJECT; do
AUTHORITY=${SUBJECT#spiffe://}; AUTHORITY=${AUTHORITY%%/*}; HOST_LABEL=${AUTHORITY%%.*}
cat > fic.json <<EOF
{
"name": "management-${HOST_LABEL}",
"issuer": "${ISSUER}",
"subject": "spiffe://${AUTHORITY}/ns/cra-agent/management",
"audiences": ["api://AzureADTokenExchange"]
}
EOF
az ad app federated-credential create --id "$APP_ID" --parameters @fic.json
done
The federated credentials let management sign in, but its service principal still needs permission on the resources management reads. That includes the secret store the secretKeyRef points to, for example Key Vault Secrets User on the vault. Management uses the same application (client) ID as the apps, so the role assignments from step 5 cover both. There's one management credential per region host, shared by all apps on the project's app registration, so you create them once per project. As with step 4, run the loop again whenever the region hosts change. The setup-federated-catalyst-identity.sh helper script federates both the app and management.
Configuration options
Set these in terraform/terraform.tfvars and run make apply again. variables.tf describes every variable.
| Area | Variables |
|---|---|
| Naming | cluster_name (default catalyst) names the cluster and is the prefix for every other resource. resource_group_name overrides the resource group's name. |
| Network | vnet_cidr, aks_subnet_cidr, database_subnet_cidr, service_cidr, and dns_service_ip. The service range must not overlap the virtual network. |
| Cluster | cluster_version, node_instance_type, node_min_capacity, node_max_capacity, node_desired_capacity, and availability_zones. |
| Cluster access | api_server_authorized_ip_ranges limits who can reach the AKS API server. aks_admin_principal_ids and aks_readonly_principal_ids give access to Entra ID users, groups, or service principals. |
| PostgreSQL | postgresql_version, postgresql_instance_class, postgresql_allocated_storage, postgresql_high_availability, postgresql_backup_retention_period, and postgresql_geo_redundant_backup_enabled. |
| Bastion host | enable_bastion, bastion_ssh_public_key, and bastion_allowed_cidr add a jumpbox in the virtual network. It's the only way to open a psql session to the PostgreSQL server from outside the cluster. |
| Peering | enable_peering and peer_vnet_id peer the region's virtual network with one of yours. |
Limitations
- Secrets backends. Catalyst Enterprise Self-Hosted supports AWS Secrets Manager and Kubernetes Secrets. Azure Key Vault support is on the roadmap.
- Public endpoints. The gateway and the AKS API server both have public endpoints. The PostgreSQL server has none. To limit access to the API server, set
api_server_authorized_ip_ranges. - Region groups. This guide builds a standalone region. A region that already has projects can't join a region group, and group members need settings that this guide leaves at their defaults. If you plan to run a group, start with the Azure multi-region deployment guide. It uses the same Terraform.