Skip to content

Monitoring ​

sphericon is one process on one PostgreSQL database, so monitoring comes down to four questions: is the process up and reachable, are requests succeeding, is background work keeping up, and is the database and the email provider healthy. This page lists the probes and metrics that answer them, proposes service level objectives, and ships example alerting rules. Each alert links to a section of the runbook.

Probes ​

EndpointPortUse it for
GET /healthzPORTLiveness. Always 200 while the process serves; touches nothing.
GET /readyzPORTReadiness. Pings the database; 503 when it is unreachable.
GET /metricsMETRICS_ADDRPrometheus exposition. Opt-in, served on its own internal listener.

/metrics is never served on the public port and has no authentication: the network boundary is the control. Enable it with METRICS_ADDR (see Prometheus metrics) or, with the Helm chart, metrics.enabled=true, and keep the port off your public ingress.

With the Prometheus Operator, the chart can also create the scrape configuration and the rules below:

sh
helm upgrade --install sphericon charts/sphericon \
  --set existingSecret=sphericon-secrets \
  --set metrics.enabled=true \
  --set metrics.serviceMonitor.enabled=true \
  --set metrics.prometheusRule.enabled=true

Metrics reference ​

Names below are the ones Prometheus sees on /metrics. They come from OpenTelemetry, so counters end in _total and the unit becomes a suffix (_seconds, _milliseconds). Every series also carries otel_scope_* labels, which this page leaves out. Labels are bounded technical dimensions (handler, route template, queue, status class); tenant or recipient identifiers are never used as labels (ADR 0018).

HTTP ​

One set of metrics covers the three API surfaces (/site, /api and /collect); the MCP endpoint dispatches through the external API and is counted with it. The tracker, tracking links, provider hooks, probes and the SPA are not instrumented.

MetricTypeLabels
ogen_server_request_count_totalcounteroperation_id, http_route, http_response_status_code
ogen_server_errors_count_totalcountersame
ogen_server_duration_milliseconds_*histogramsame; _bucket, _sum, _count, milliseconds

The surface is not a label of its own. Tell them apart by operation_id: Collect* is /collect, Site* is /site, and everything else is the external /api. Status class is http_response_status_code=~"5..". ogen_server_errors_count_total counts requests that ended in an error of any class, 4xx included, so use the request counter for 5xx ratios.

Domain events (outbox) ​

MetricTypeLabelsMeaning
outbox_lag_secondsgaugeconsumer_groupAge of the oldest outbox row the group has not processed yet; 0 when caught up.
watermill_messages_processed_totalcounterhandler, errorMessages handled, by handler and whether it returned an error.
watermill_handler_duration_seconds_*histogramhandler, errorHandler run time.

Consumer groups are persist, automations, webhooks and suppression; their handlers are persist_event, enroll_automations, dispatch_webhooks and update_suppression. The lag gauge runs one query per scrape.

Background jobs (river) ​

MetricTypeLabelsMeaning
river_queue_depthgaugequeueJobs ready to run now. A queue with none reports no sample.
river_queue_oldest_available_age_secondsgaugequeueAge of the oldest ready job.
river_work_count_totalcounterqueue, kind, status, attempt, ...Job attempts; status is ok, error or panic.
river_work_duration_histogram_seconds_*histogramsameAttempt run time.

Queues are default, broadcasts and webhooks. Snoozed jobs count as ok.

Database pools ​

MetricTypeLabelsMeaning
db_pool_opengaugepoolConnections open (in use plus idle).
db_pool_in_usegaugepoolConnections in use.
db_pool_idlegaugepoolIdle connections.
db_pool_maxgaugepoolConfigured maximum (DB_MAX_OPEN_CONNS, PGX_MAX_CONNS).
db_pool_wait_countgaugepoolsql pool only: cumulative requests that waited for a connection (sql.DBStats.WaitCount); use increase().
db_pool_empty_acquire_countgaugepoolpgx pool only: cumulative acquires that found the pool empty and waited for a connection to be released or opened (pgxpool.Stat.EmptyAcquireCount); use increase().

pool is sql (requests, ent, event consumers) or pgx (the job queue). The two wait counters are driver-native and not identical: pgx also counts an acquire that waited only while a new connection was opened, so it can rise while the pool is below its maximum.

Email ​

MetricTypeLabelsMeaning
email_send_outcomes_totalcounterprovider, statusstatus is accepted, error, bounce or complaint; provider is smtp or ses.

bounce and complaint count provider reports about messages the provider had accepted, so they arrive later than the send. Per-workspace complaint and bounce rates are on the workspace dashboard (ADR 0011); the metric above is the instance-wide view.

Runtime ​

Go runtime metrics (go_goroutine_count, memory and GC series) and target_info (service name, version and commit) are exported alongside. The duration histograms with a _seconds unit use the OpenTelemetry default bucket boundaries, which are coarse for seconds (the first buckets are 5 and 10), so this page alerts on the HTTP histogram in milliseconds and not on those.

Service level objectives (proposed) ​

These are proposed starting points, not commitments. Nothing in the product enforces them; adopt, tighten or drop them to match what you promise your users. Objectives are measured over 30 days, and every query below is plain PromQL over the metrics above.

SLIProposed SLOMeasured as
/collect availability99.9% of requests not 5xx1 - 5xx / all for operation_id=~"Collect.*"
/site 5xx ratiobelow 1%5xx / all for operation_id=~"Site.*"
/api 5xx ratiobelow 1%5xx / all for the remaining operations
Outbox consumer lagunder 5 minutes, 99% of the timeshare of time max(outbox_lag_seconds) is at most 300
Discarded jobsbelow 1% of jobsno direct metric today; see below

/collect gets the strictest objective because a failed request is an event that is lost, while a failed /site request is retried by a person. Request latency and deliverability (send errors, bounces, complaints) have alerts below but no SLO: the right numbers depend on your traffic and your sending reputation.

txt
# /collect availability over 30 days
1 - (
  sum(increase(ogen_server_request_count_total{operation_id=~"Collect.*", http_response_status_code=~"5.."}[30d]))
  / sum(increase(ogen_server_request_count_total{operation_id=~"Collect.*"}[30d]))
)

# Outbox lag SLI: share of the last 30 days with every consumer within 5 minutes
avg_over_time((max(outbox_lag_seconds) <= bool 300)[30d:1m])

# Job failure ratio (upper bound for the discarded-job ratio)
sum(increase(river_work_count_total{status!="ok"}[30d])) / sum(increase(river_work_count_total[30d]))

River does not export a discarded-job counter. A job is discarded after its last attempt fails, so the failed-attempt ratio above is an upper bound (retried attempts count too). The exact figure is the number of discarded rows in river_job (kept 14 days) and the river job errored log lines with attempt equal to max_attempts. The example alerts watch the failure ratio and the queue age instead of a discard ratio for this reason.

Alert thresholds are deliberately looser than the objectives: an alert says "act now", an objective says "this is what we promise over a month". Burn-rate alerting is a good next step once you have run the objectives for a while.

Example Prometheus rules ​

The rules below are the same file the Helm chart renders into a PrometheusRule (charts/sphericon/files/alerts.yaml). Load them with rule_files: in a plain Prometheus, or let the chart manage them (metrics.prometheusRule.enabled=true; set metrics.prometheusRule.labels to match your Prometheus ruleSelector). SphericonDown matches job=~".*sphericon.*": adjust it to your scrape job name.

yaml
# Example Prometheus alerting rules for sphericon. This file is the single source: the Helm chart
# renders it into a PrometheusRule, and docs/operations/monitoring.md includes it verbatim.
# Every alert has a section with the same name in docs/operations/runbook.md (mise run
# check:helm enforces it). Thresholds are starting points; tune them to your traffic.
#
# The three API surfaces share one set of HTTP metrics and are told apart by operation_id:
# Collect* is /collect, Site* is /site, everything else is the external /api.
groups:
  - name: sphericon.availability
    rules:
      - alert: SphericonDown
        expr: up{job=~".*sphericon.*"} == 0
        for: 5m
        labels:
          severity: critical
        annotations:
          summary: sphericon metrics target {{ $labels.instance }} is down
          description: Prometheus has not been able to scrape {{ $labels.instance }} for 5 minutes.
          runbook_url: https://docs.getsphericon.app/operations/runbook#sphericondown

      - alert: SphericonCollectErrorRatioHigh
        expr: |
          sum(rate(ogen_server_request_count_total{operation_id=~"Collect.*", http_response_status_code=~"5.."}[5m]))
            / sum(rate(ogen_server_request_count_total{operation_id=~"Collect.*"}[5m])) > 0.01
        for: 10m
        labels:
          severity: critical
        annotations:
          summary: /collect is failing {{ $value | humanizePercentage }} of requests
          description: More than 1% of tracking requests answered 5xx for 10 minutes; events are being lost.
          runbook_url: https://docs.getsphericon.app/operations/runbook#sphericoncollecterrorratiohigh

      - alert: SphericonSiteErrorRatioHigh
        expr: |
          sum(rate(ogen_server_request_count_total{operation_id=~"Site.*", http_response_status_code=~"5.."}[5m]))
            / sum(rate(ogen_server_request_count_total{operation_id=~"Site.*"}[5m])) > 0.05
        for: 10m
        labels:
          severity: warning
        annotations:
          summary: /site is failing {{ $value | humanizePercentage }} of requests
          description: More than 5% of dashboard API requests answered 5xx for 10 minutes.
          runbook_url: https://docs.getsphericon.app/operations/runbook#sphericonsiteerrorratiohigh

      - alert: SphericonApiErrorRatioHigh
        expr: |
          sum(rate(ogen_server_request_count_total{operation_id!~"Site.*|Collect.*", http_response_status_code=~"5.."}[5m]))
            / sum(rate(ogen_server_request_count_total{operation_id!~"Site.*|Collect.*"}[5m])) > 0.05
        for: 10m
        labels:
          severity: warning
        annotations:
          summary: /api is failing {{ $value | humanizePercentage }} of requests
          description: More than 5% of external API requests answered 5xx for 10 minutes.
          runbook_url: https://docs.getsphericon.app/operations/runbook#sphericonapierrorratiohigh

      - alert: SphericonRequestLatencyHigh
        expr: |
          histogram_quantile(0.95,
            sum by (le) (rate(ogen_server_duration_milliseconds_bucket{operation_id!~"Collect.*"}[5m]))) > 2000
        for: 15m
        labels:
          severity: warning
        annotations:
          summary: p95 /site and /api latency is {{ $value | humanize }} ms
          description: The 95th percentile of request duration has been above 2 seconds for 15 minutes.
          runbook_url: https://docs.getsphericon.app/operations/runbook#sphericonrequestlatencyhigh

  - name: sphericon.events
    rules:
      - alert: SphericonOutboxLagHigh
        expr: max by (consumer_group) (outbox_lag_seconds) > 300
        for: 10m
        labels:
          severity: warning
        annotations:
          summary: Event consumer {{ $labels.consumer_group }} is {{ $value | humanizeDuration }} behind
          description: The oldest unprocessed domain event for {{ $labels.consumer_group }} is older than 5 minutes.
          runbook_url: https://docs.getsphericon.app/operations/runbook#sphericonoutboxlaghigh

      - alert: SphericonEventHandlerErrors
        expr: |
          sum by (handler) (rate(watermill_messages_processed_total{error="true"}[10m]))
            / sum by (handler) (rate(watermill_messages_processed_total[10m])) > 0.05
        for: 15m
        labels:
          severity: warning
        annotations:
          summary: Event handler {{ $labels.handler }} fails {{ $value | humanizePercentage }} of messages
          description: More than 5% of messages handled by {{ $labels.handler }} returned an error for 15 minutes.
          runbook_url: https://docs.getsphericon.app/operations/runbook#sphericoneventhandlererrors

  - name: sphericon.jobs
    rules:
      - alert: SphericonJobQueueBacklog
        expr: max by (queue) (river_queue_oldest_available_age_seconds) > 900
        for: 15m
        labels:
          severity: warning
        annotations:
          summary: Job queue {{ $labels.queue }} has had a ready job waiting {{ $value | humanizeDuration }}
          description: The oldest job ready to run in queue {{ $labels.queue }} has waited more than 15 minutes.
          runbook_url: https://docs.getsphericon.app/operations/runbook#sphericonjobqueuebacklog

      - alert: SphericonJobFailureRatioHigh
        expr: |
          sum by (queue) (rate(river_work_count_total{status!="ok"}[15m]))
            / sum by (queue) (rate(river_work_count_total[15m])) > 0.2
        for: 30m
        labels:
          severity: warning
        annotations:
          summary: Jobs in queue {{ $labels.queue }} fail {{ $value | humanizePercentage }} of attempts
          description: More than 20% of job attempts in queue {{ $labels.queue }} errored or panicked for 30 minutes.
          runbook_url: https://docs.getsphericon.app/operations/runbook#sphericonjobfailureratiohigh

  - name: sphericon.database
    rules:
      - alert: SphericonDBPoolSaturated
        expr: max by (pool) (db_pool_in_use / db_pool_max) >= 0.9
        for: 10m
        labels:
          severity: warning
        annotations:
          summary: Database pool {{ $labels.pool }} is {{ $value | humanizePercentage }} in use
          description: Connections in pool {{ $labels.pool }} have been at 90% of the configured maximum for 10 minutes.
          runbook_url: https://docs.getsphericon.app/operations/runbook#sphericondbpoolsaturated

      - alert: SphericonDBPoolWaits
        expr: sum by (pool) (increase(db_pool_wait_count[15m]) or increase(db_pool_empty_acquire_count[15m])) > 100
        for: 15m
        labels:
          severity: warning
        annotations:
          summary: Requests are waiting for a connection from pool {{ $labels.pool }}
          description: More than 100 connection waits in 15 minutes on pool {{ $labels.pool }} (database/sql waits for pool sql, pgx empty-pool acquires for pool pgx).
          runbook_url: https://docs.getsphericon.app/operations/runbook#sphericondbpoolwaits

  - name: sphericon.deliverability
    rules:
      - alert: SphericonEmailSendErrorsHigh
        expr: |
          sum by (provider) (rate(email_send_outcomes_total{status="error"}[15m]))
            / sum by (provider) (rate(email_send_outcomes_total{status=~"accepted|error"}[15m])) > 0.1
        for: 15m
        labels:
          severity: warning
        annotations:
          summary: Provider {{ $labels.provider }} rejects {{ $value | humanizePercentage }} of sends
          description: More than 10% of send attempts through {{ $labels.provider }} failed for 15 minutes.
          runbook_url: https://docs.getsphericon.app/operations/runbook#sphericonemailsenderrorshigh

      - alert: SphericonEmailBounceRateHigh
        expr: |
          sum by (provider) (increase(email_send_outcomes_total{status="bounce"}[1h]))
            / sum by (provider) (increase(email_send_outcomes_total{status="accepted"}[1h])) > 0.05
          and sum by (provider) (increase(email_send_outcomes_total{status="accepted"}[1h])) > 100
        for: 30m
        labels:
          severity: warning
        annotations:
          summary: Bounce rate through {{ $labels.provider }} is {{ $value | humanizePercentage }}
          description: More than 5% of messages accepted by {{ $labels.provider }} bounced in the last hour.
          runbook_url: https://docs.getsphericon.app/operations/runbook#sphericonemailbounceratehigh

      - alert: SphericonEmailComplaintRateHigh
        expr: |
          sum by (provider) (increase(email_send_outcomes_total{status="complaint"}[1h]))
            / sum by (provider) (increase(email_send_outcomes_total{status="accepted"}[1h])) > 0.001
          and sum by (provider) (increase(email_send_outcomes_total{status="accepted"}[1h])) > 1000
        for: 30m
        labels:
          severity: critical
        annotations:
          summary: Complaint rate through {{ $labels.provider }} is {{ $value | humanizePercentage }}
          description: More than 0.1% of messages accepted by {{ $labels.provider }} drew a spam complaint in the last hour.
          runbook_url: https://docs.getsphericon.app/operations/runbook#sphericonemailcomplaintratehigh
AlertSeverityFires when
SphericonDowncriticalA scrape target has been down for 5 minutes.
SphericonCollectErrorRatioHighcritical/collect 5xx ratio above 1% for 10 minutes.
SphericonSiteErrorRatioHighwarning/site 5xx ratio above 5% for 10 minutes.
SphericonApiErrorRatioHighwarning/api 5xx ratio above 5% for 10 minutes.
SphericonRequestLatencyHighwarning/site and /api p95 above 2 seconds for 15 minutes.
SphericonOutboxLagHighwarningA consumer group is over 5 minutes behind for 10 minutes.
SphericonEventHandlerErrorswarningA handler fails over 5% of messages for 15 minutes.
SphericonJobQueueBacklogwarningA ready job has waited over 15 minutes.
SphericonJobFailureRatioHighwarningOver 20% of job attempts fail for 30 minutes.
SphericonDBPoolSaturatedwarningA pool is at 90% of its maximum for 10 minutes.
SphericonDBPoolWaitswarningOver 100 connection waits in 15 minutes.
SphericonEmailSendErrorsHighwarningOver 10% of send attempts fail for 15 minutes.
SphericonEmailBounceRateHighwarningOver 5% of accepted messages bounce in an hour.
SphericonEmailComplaintRateHighcriticalOver 0.1% of accepted messages draw a complaint.

The bounce and complaint alerts only evaluate once the provider has accepted at least 100 and 1,000 messages in the hour, so a quiet instance does not page on one bad address.