Skip to content

Runbook ​

One section per alert in the example rules. Each alert links here by name. A section says what the alert means, what to check first and how to fix it. The commands assume a Kubernetes deployment (kubectl) and psql access to the database; translate them for Docker or a bare binary.

Three things help with almost every alert:

  • Logs. The process logs structured JSON. Failed jobs log river job errored or river job panicked with kind, queue, job_id and attempt; provider hooks log hooks/ses: ... lines with the workspace_id.
  • /readyz. If it returns 503, the database is the first suspect.
  • Rollouts. Check whether the alert started with a deploy or a migration (Upgrading) before digging further.

SphericonDown ​

Means. Prometheus cannot scrape a sphericon metrics target for 5 minutes. The process is down, unreachable, or the metrics listener is not enabled.

Check.

  1. Are the pods running and ready? kubectl get pods -l app.kubernetes.io/name=sphericon
  2. Why did they stop? kubectl describe pod ... (OOMKilled, failed probe) and kubectl logs --previous ....
  3. Is it only the scrape? Hit /healthz on the public port. If the app is fine, check METRICS_ADDR (or metrics.enabled), the metrics Service and any network policy between Prometheus and the pod.
  4. If /readyz returns 503, the database is unreachable: see SphericonDBPoolSaturated and check PostgreSQL itself.

Fix. Restart on a crash loop caused by configuration (a missing secret, an invalid ENCRYPTION_KEY, a port clash between PORT and METRICS_ADDR all fail startup with a clear log line). After a bad deploy, roll back the image. Pending migrations are not applied automatically with several replicas: run them as described in Upgrading.

SphericonCollectErrorRatioHigh ​

Means. More than 1% of /collect requests (tracker events and identifies from customer sites) answered 5xx. Those events are lost unless the client retries, so this is the most urgent request alert.

Check.

  1. Which operation? sum by (operation_id, http_response_status_code) (rate(ogen_server_request_count_total{operation_id=~"Collect.*"}[5m])).
  2. The collect path writes to the domain-event outbox in PostgreSQL, so look for database trouble first: /readyz, SphericonDBPoolSaturated, SphericonDBPoolWaits, disk full, a failover in progress.
  3. Read the error logs of the failing pods for the underlying error.

Fix. Restore database health (free disk, resolve locks, fail over). If a deploy caused it, roll back. If load is the cause, add replicas, but check the pool sizes first, because each replica adds connections.

SphericonSiteErrorRatioHigh ​

Means. More than 5% of dashboard API (/site) requests answered 5xx for 10 minutes. People see errors in the UI.

Check.

  1. Narrow it down: topk(10, sum by (operation_id) (rate(ogen_server_request_count_total{operation_id=~"Site.*", http_response_status_code=~"5.."}[5m]))). One operation points at a bug or a slow query, all of them point at the database or a deploy.
  2. Compare with SphericonRequestLatencyHigh and the pool alerts.
  3. Look at the logs for the failing operation and at what changed (deploy, migration, a large import or broadcast running at the same time).

Fix. Roll back a bad release, relieve the database, or fix the failing query. A single operation failing for a single workspace is usually a data problem, not an outage.

SphericonApiErrorRatioHigh ​

Means. More than 5% of external API (/api, also used by MCP clients) requests answered 5xx for 10 minutes. Integrations calling your instance are failing.

Check. The same as SphericonSiteErrorRatioHigh, with operation_id!~"Site.*|Collect.*". Also look for one noisy client: a large batch import can saturate the database for everyone.

Fix. As for /site. If one client is the cause, ask it to slow down or lower its batch size.

SphericonRequestLatencyHigh ​

Means. The 95th percentile of /site and /api request duration has been above 2 seconds for 15 minutes. /collect is excluded because it should always be fast.

Check.

  1. Which operations are slow? histogram_quantile(0.95, sum by (le, operation_id) (rate(ogen_server_duration_milliseconds_bucket{operation_id!~"Collect.*"}[5m])))
  2. Is the database the bottleneck? Check pool usage and waits (SphericonDBPoolSaturated) and PostgreSQL itself (CPU, IO, long queries in pg_stat_activity).
  3. Large events tables make segment and analytics queries slow; confirm Event retention is running (ADR 0019).

Fix. Kill or fix the slow query, add a missing index, or scale the database. Adding replicas does not help if the database is the limit.

SphericonOutboxLagHigh ​

Means. A domain-event consumer group is more than 5 minutes behind. The groups are persist (writes the Event rows that analytics and segments read), automations (enrolls contacts), webhooks (dispatches outgoing webhooks) and suppression (adds bounced and complaining addresses to the suppression list). The consumer_group label tells you which one.

Impact. Lag in suppression delays blocking of hard-bouncing addresses, which hurts reputation. Lag in persist makes dashboards stale. The outbox is also only pruned below the slowest consumer's cursor, so a stuck group makes it grow.

Check.

  1. Is only one group behind? That group's handler is failing or stuck. Look at SphericonEventHandlerErrors for its handler and read the logs.
  2. Are all groups behind? The router is not running (process restarting, database trouble) or the database is slow. Check SphericonDown and the pool alerts.
  3. Lag drops by itself after a spike (a big import, an SES hook burst) once consumers catch up: watch the trend before acting.

Fix. A message that keeps failing is redelivered and holds its group back: fix the cause (a bug, a bad payload, an unreachable dependency) and ship a release; the group resumes from its cursor with no data loss. If a deploy introduced it, roll back.

SphericonEventHandlerErrors ​

Means. More than 5% of messages handled by one event handler returned an error for 15 minutes. The handler label is one of persist_event, enroll_automations, dispatch_webhooks or update_suppression.

Check.

  1. The logs of the process around the failures, for the underlying error.
  2. update_suppression and persist_event fail mostly on database errors. enroll_automations fails on a broken automation definition. dispatch_webhooks should not fail on a customer's unreachable endpoint (delivery is a river job), so a failure there points at sphericon itself.
  3. Correlate with SphericonOutboxLagHigh: failures that keep a message from being acknowledged show up as growing lag for that handler's group.

Fix. Fix the cause and deploy. Failed messages are redelivered and the handlers are written to tolerate redelivery, so there is nothing to replay by hand.

SphericonJobQueueBacklog ​

Means. In one river queue (default, broadcasts or webhooks) the oldest job that is ready to run has been waiting more than 15 minutes.

Check.

  1. Is it just load? A large broadcast legitimately queues many jobs. If river_queue_depth is high but falling, wait.
  2. Is anything being worked? sum by (queue) (rate(river_work_count_total[5m])). Zero with a growing backlog means the workers are not running or are stuck.
  3. Are the jobs failing and retrying? See SphericonJobFailureRatioHigh.
  4. Workers need database connections from the pgx pool: check db_pool_in_use{pool="pgx"} against PGX_MAX_CONNS (SphericonDBPoolSaturated).
  5. Check that the river schema is applied. It is created by sphericon migrate; an instance that skipped migration has no job tables and the job metrics and workers fail.

Fix. Restart stuck pods, raise PGX_MAX_CONNS (and the database max_connections accordingly) if the pool is the limit, or add replicas. Worker concurrency per queue is fixed in the binary.

SphericonJobFailureRatioHigh ​

Means. More than 20% of job attempts in a queue ended in error or panic for 30 minutes. Because retries count, this overstates how many jobs end up discarded, but a sustained high ratio always needs a look.

Check.

  1. Which kind? topk(10, sum by (kind) (rate(river_work_count_total{status!="ok"}[15m]))).

  2. Read the river job errored and river job panicked log lines for that kind; they carry the error, attempt and max_attempts.

  3. Common causes: the email provider rejecting sends (see SphericonEmailSendErrorsHigh), a customer's webhook endpoint down (the webhooks queue retries up to 10 times), DNS for sending-domain checks, or a bug.

  4. Jobs that used up their attempts are in state discarded and are kept for 14 days:

    sql
    SELECT kind, count(*), max(finalized_at)
    FROM river_job WHERE state = 'discarded'
    GROUP BY kind ORDER BY 2 DESC;

Fix. Fix the cause. Failing webhook deliveries to one customer endpoint are that customer's problem; tell them. Discarded jobs are not retried automatically: after the fix, re-run the underlying action (for example, re-send the broadcast to the missed recipients).

SphericonDBPoolSaturated ​

Means. One connection pool has been at or above 90% of its maximum for 10 minutes. pool is sql (requests, ent, event consumers; DB_MAX_OPEN_CONNS, default 15) or pgx (the job queue; PGX_MAX_CONNS, default 25).

Check.

  1. Slow queries holding connections: SELECT pid, state, wait_event, now() - query_start AS age, left(query, 80) FROM pg_stat_activity WHERE datname = current_database() ORDER BY age DESC;
  2. Idle-in-transaction sessions or locks blocking others.
  3. Does the total across replicas fit max_connections? Each replica opens both pools; see Sizing.

Fix. Remove the slow query or the lock first. If the load is real, raise DB_MAX_OPEN_CONNS or PGX_MAX_CONNS and make sure the sum over all replicas stays below the server's max_connections; beyond that, add a connection pooler. DB_MAX_IDLE_CONNS must not exceed DB_MAX_OPEN_CONNS.

SphericonDBPoolWaits ​

Means. More than 100 times in 15 minutes a request or worker had to wait for a free connection (db_pool_wait_count for the sql pool, db_pool_empty_acquire_count for the pgx pool). Unlike saturation this catches short, repeated stalls.

Check and fix. The same as SphericonDBPoolSaturated. Waiting that continues while db_pool_in_use is well below db_pool_max points at slow connection establishment (network, TLS, a pooler) instead of a too-small pool. On the pgx pool an acquire that only waited for a new connection to open also counts, so a cold or restarted pool can trip it.

SphericonEmailSendErrorsHigh ​

Means. More than 10% of send attempts through a provider (smtp or ses) failed for 15 minutes. These are failures to hand a message to the provider, not bounces.

Check.

  1. The send error in the logs of the river job errored lines for the send jobs: credentials revoked, a provider outage, throttling, an unverified sender or sending domain.
  2. For SES, check the AWS account status (sandbox, suspended sending, quota) and region. For SMTP, check the relay's status page and your credentials.
  3. Is it one Workspace? Provider credentials are per Workspace, so a single wrong credential fails only that Workspace's sends.

Fix. Rotate or correct the credentials in the Workspace's integration settings, wait out a provider incident, or request a higher quota. Sends run as jobs and are retried within their attempt limit; check SphericonJobFailureRatioHigh for anything that was discarded.

SphericonEmailBounceRateHigh ​

Means. More than 5% of the messages a provider accepted in the last hour bounced, with at least 100 accepted. High bounce rates damage sender reputation and can get an SES account placed under review.

Check.

  1. Which Workspace? The metric is instance-wide; the per-Workspace bounce rate is on the Workspace dashboard (ADR 0011).
  2. A recent import of an old or purchased list is the usual cause. Permanent bounces are added to the suppression list automatically, but the damage is done by the first send.
  3. Confirm the suppression consumer is keeping up (SphericonOutboxLagHigh); if it lags, bouncing addresses stay sendable.

Fix. Pause the offending broadcast, clean the list, and re-send only to verified addresses. Talk to the Workspace owner if the source of the list is the problem.

SphericonEmailComplaintRateHigh ​

Means. More than 0.1% of the messages a provider accepted in the last hour drew a spam complaint, with at least 1,000 accepted. This is the most damaging deliverability signal: SES reviews accounts above roughly this rate, and mailbox providers start filtering.

Check.

  1. Which Workspace and which broadcast? Use the Workspace dashboard (ADR 0011).
  2. Is the unsubscribe link working? The one-click unsubscribe headers (ADR 0012) must be present and the link reachable, or people use the spam button instead.
  3. Is the list opted in? Mail to people who never asked draws complaints.

Fix. Stop the offending broadcast now. Complaining addresses are suppressed automatically; confirm that the suppression consumer is current. Fix the source of the audience or the unsubscribe flow before sending again, and respond to the provider if they open a review.