# Metrics

The **Metrics** page in the Catalyst console gives platform engineers and business operators a live view of request volume, error rate, and tail latency across a project — answering "is traffic healthy right now, and which app needs attention?" without needing to install Prometheus or write a single query. For the conceptual model (metrics, traces, logs), see [Observability](https://docs.diagrid.io/concepts/observability).

![Catalyst console Metrics page showing the four project KPI cards (Total Requests, Avg Throughput, Error Rate, Avg P95 Latency), the app comparison table, and the aggregate HTTP error-rate chart](https://docs.diagrid.io/img/catalyst/metrics-dashboard-light.png)

## Dashboard panels

The Metrics page opens scoped to the **whole project**. It groups data into three areas.

### Project KPIs

Four cards summarise the selected time window:

| Card | What it measures | Question it answers |
| --- | --- | --- |
| **Total Requests** | HTTP + gRPC requests received across every app | "How much traffic did my project handle?" |
| **Avg Throughput** | Requests per second (weighted across protocols) | "What's the steady-state load?" |
| **Error Rate** | Weighted percentage of failed HTTP + gRPC calls | "Are calls failing project-wide?" |
| **Avg P95 Latency** | Worst-of HTTP / gRPC 95th-percentile latency, in milliseconds | "Are tail latencies acceptable?" |

### App comparison table

A row per app with **Status**, **Total Requests**, **Throughput (RPS)**, **Error Rate (%)**, and **P95 Latency (ms)**. Sort or filter by app name and toggle **Active Only** to hide idle workloads. This is where you answer "**is my app near its rate limit?**" — compare the RPS column against the per-app rate limit for your plan in [plan limits](https://docs.diagrid.io/operate/plans-and-support#per-catalyst-cloud-region) (100 requests per second per App ID on the free plan and 500 on Cloud Plus, counting HTTP and gRPC together).

Selecting a row opens the [App detail view](#app-detail-view).

### Aggregate charts

Two time-series charts for the project, split by protocol (HTTP / gRPC):

- **Error Rate over time** — pinpoints when failures started.
- **P95 Latency over time** — pinpoints when tail latency regressed. Hover any point to read the HTTP and gRPC values at that moment.

![Avg P95 Latency time-series chart in the Catalyst console, with separate HTTP and gRPC series and a tooltip showing per-protocol values at a single timestamp](https://docs.diagrid.io/img/catalyst/metrics-p95-latency-light.png)

## Filter and zoom

The toolbar applies to every panel on the page:

- **Protocol** — `All`, `HTTP`, or `gRPC`.
- **Time range** — `1h`, `6h`, `24h`, `7d`, or a custom range within the last seven days. The 7-day cap matches the default Cloud retention (see [Retention](#retention)).
- **App search** and **Active Only** — narrow the comparison table.

Dashboards refresh roughly once a minute while the page is open.

## App detail view

Click any app in the comparison table to drill into a per-app page with the same KPIs and charts scoped to that single workload. Use this view to confirm whether a project-level spike came from one noisy app or is spread across the fleet, and to see that app's headroom against the rate limit.

## Investigating a spike

A Metrics chart tells you *when* something changed, not *what* changed. To go from a spike to a root cause, keep the same app and time window and jump to:

- [**API Logs**](https://docs.diagrid.io/operate/project-operations/observability/api-logs) — filter by app and the time window from the chart to read the actual failing requests, methods, status codes, and latencies.
- [**Topology**](https://docs.diagrid.io/operate/project-operations/observability/topology) — select the same app to see which upstream callers or downstream components were involved during the window.

This pattern — spike on Metrics → request detail in API Logs → connections in Topology — covers most incident triage in Catalyst.

## Retention

On Catalyst Cloud, metrics are retained for **7 days**; logs for **3 days**. Both limits are listed in [Plans & support](https://docs.diagrid.io/operate/plans-and-support#limits-in-every-region). Paid plans and Catalyst Enterprise increase retention — contact Diagrid if your investigation horizon is longer than the default.

## Export to external monitoring

If you already run Datadog, Grafana, New Relic, or any OTLP-compatible backend, you can ship Catalyst telemetry to it instead of (or in addition to) the in-console dashboards.

- **Catalyst Cloud** — metric dashboards are in-console only today. Distributed traces *can* be exported by attaching a Dapr `Configuration` resource to your apps; see [Observability settings](https://docs.diagrid.io/operate/project-operations/observability) for the tracing-export workflow.
- **Catalyst Enterprise Self-Hosted** — the platform ships an OpenTelemetry Collector that supports OTLP receivers and exporters for Prometheus Remote Write and OTLP HTTP. See [Enterprise Self-Hosted Observability](https://docs.diagrid.io/operate/hosting/enterprise-self-hosted/observability) for the collector configuration.

## Related

- [Observability](https://docs.diagrid.io/concepts/observability) — the concept model behind metrics, traces, and logs.
- [API Logs](https://docs.diagrid.io/operate/project-operations/observability/api-logs) — per-call detail to pair with a metric spike.
- [Topology](https://docs.diagrid.io/operate/project-operations/observability/topology) — connection context for a noisy app.
- [Plan limits](https://docs.diagrid.io/operate/plans-and-support#plan-limits) — rate limits and retention.
