PostgreSQL Cloud Monitoring: RDS, Azure and Cloud SQL

PostgreSQL is portable; managed-service telemetry is not. AWS RDS, Azure Database for PostgreSQL Flexible Server and Cloud SQL all expose CPU, storage, connection and network signals, but their definitions, sampling and included service activity differ. An alert copied between providers can be misleading even when its label looks familiar.

Know what the provider chart represents

RDS CloudWatch publishes instance-level measures such as CPU utilisation, connections, IOPS, latency and replication-slot lag. Azure Flexible Server exposes CPU, memory, storage, IOPS, queue depth, transaction-log storage and PostgreSQL-aware measures, with some metrics processed in batches. Cloud SQL provides reserved CPU utilisation, disk and connection metrics, and offers query-insight metrics only when the relevant feature is enabled. These are useful views, but none replaces PostgreSQL session state, locks or query plans.

Map a common operational question to each service

  • For CPU pressure, record whether the value is host, allocated or reserved compute utilisation.
  • For storage, distinguish data use, WAL or transaction-log use, and service-managed logs or backups.
  • For connections, verify whether idle sessions are included and how managed internal sessions are treated.
  • For replication, distinguish physical replica delay, logical-slot retention and HA behaviour.

Alert on a condition, not a name

A useful alert joins a sustained service metric change with business impact or database evidence. For example, high CPU plus rising query latency and active backends is more actionable than a short CPU peak. Account for provider retention, aggregation and delayed visibility when choosing windows and thresholds.

Review these mappings whenever you change service tier, storage configuration or provider generation. The available metrics and limits are part of the platform contract, not permanent universal database facts.

Map the failure to the evidence you can actually collect

ProblemCloud signalPostgreSQL follow-up
Storage rising quicklyUsed or free storage for the serviceSeparate table growth, WAL retention and temporary work.
Connection pressureProvider connection countBreak down pg_stat_activity by state and client.
Slow requests with low CPUCPU and storage latency where exposedInspect waits, locks and old transactions.

Azure's active_connections metric includes multiple session states despite its name. On any provider, confirm whether your chart means executing queries or connected sessions. Query-analysis features are another separate layer: check whether the required extension or provider feature is enabled before expecting historical statement detail.

Example: WAL retention threatens the storage budget

Suppose the platform sends a low-space alert while application table growth is modest. A stopped replication consumer is a plausible hypothesis, but not a conclusion. Inspect slot retention and consumer health through the database alongside the service's transaction-log or storage measures. Compare timestamps and units: a retained-byte quantity and a replica delay in seconds are different dimensions of the problem.

Estimate time to exhaustion from recent growth and available space, using several intervals because bursty workloads invalidate a straight-line forecast quickly. A storage expansion may buy investigation time; it does not repair the consumer. Agree whether recovery can resume from the existing position before taking any action that discards retained WAL.

Include visibility delay in the response design

Do not select a one-minute alert response target when a critical metric is delivered in larger batches. Use an independent reachability or application symptom check for urgent incidents and the delayed metric for explanation. Document missing-data behaviour so a failed collection pipeline cannot produce a misleading all-clear.

When changing tiers or migrating providers, test query-statistics access, supported extensions, session visibility and log export using the actual monitoring identity. Save a representative incident bundle before migration and compare the new evidence with it. Successful connectivity only proves a login works; it does not prove the team can investigate blocking, vacuum debt or a plan regression when production is busy.

Keep the evidence for the next incident

Mini DBA PostgreSQL monitoring provides query and metric history alongside session, lock, vacuum and WAL diagnostics, helping teams compare an incident with normal operation. Available evidence depends on the monitored version, permissions and configured collection.

References and further reading

Add comment