MySQL Cloud Monitoring: RDS vs Azure vs Cloud SQL Metrics

Managed MySQL services expose valuable platform metrics, but the labels do not all mean the same thing. A CPU percentage, network total or connection count may include provider activity as well as customer workload. Treat provider monitoring as an additional evidence layer, not a substitute for MySQL status, query and transaction diagnostics.

Start with the metric boundary

Amazon RDS publishes instance-level CloudWatch metrics such as CPU utilisation, connections, IOPS, throughput and latency. AWS documents that some traffic metrics include monitoring and replication activity, and its connection metric is not a complete count of every engine-created session. Azure Database for MySQL Flexible Server exposes host CPU, memory and storage I/O consumption percentages plus engine counters such as running threads and buffer-pool reads. Cloud SQL presents reserved CPU utilisation, disk operations and database metrics, with sampling and visibility delay documented per metric.

What this changes during an investigation

  • Do not compare an Azure I/O percentage directly with AWS IOPS or Cloud SQL disk operations.
  • Check the provider's sampling interval, aggregation and metric delay before aligning it with an incident.
  • Use MySQL query, lock and InnoDB evidence to explain a platform spike.
  • Include burstable CPU-credit metrics where the selected service tier uses them.

Build a portable runbook

Record the service, tier, region, storage model and enabled diagnostics for every production instance. Define a common vocabulary—CPU pressure, read latency, write pressure, connection growth and replication lag—then map it to each provider's actual metric. This prevents a dashboard from implying precision that the service does not expose.

Review provider documentation when you change tier or service generation. Metrics, limits and available diagnostics evolve, so alert thresholds should be tested against the exact service being operated.

A metric comparison you can use during an outage

ServiceUseful starting signalsInterpretation check
RDS for MySQLCPUUtilization, ReadLatency, WriteLatencyStorage latency is not query latency.
Azure MySQL Flexible Servercpu_percent, io_consumption_percentA percentage needs its configured capacity denominator.
Cloud SQL for MySQLdatabase/cpu/utilization, disk operation countsRead the metric kind before deriving a rate.

These examples are entry points, not a complete parity matrix. Advanced metrics, query diagnostics and operating-system visibility depend on the service configuration. Keep Aurora MySQL separate from ordinary RDS for MySQL in the inventory; its architecture and metrics should not be assumed identical because both accept MySQL connections.

Example: the cloud chart seems to contradict the application

An application records a short latency burst, but its cloud dashboard looks flat. First align time zones, bucket boundaries and aggregation. A one-minute average can hide a much shorter incident, while a delayed sample may arrive after the application has recovered. Ask whether the chart had actually received data for that interval before concluding that MySQL was healthy.

Check database evidence in parallel: slow statements, a live blocking chain, running threads and connection errors. Record the request identifier or job name, because a brief lock queue can vanish before a broad resource dashboard changes noticeably. A lack of CPU movement is consistent with waiting on a transaction.

Test monitoring permissions before you need them

Cloud IAM access to a metrics dashboard does not grant SQL privileges to inspect database sessions. Conversely, a database login cannot necessarily inspect storage configuration or provider events. Verify both access paths with the operational account and record unavailable evidence explicitly. An empty chart must not be represented as a measured zero.

When migrating providers, replay a representative workload and record latency, completed transactions and engine-side work on both services. Translate each metric by meaning and unit before comparing its values. Keep storage type, instance size and enabled diagnostics with the results. Otherwise, an apparently improved graph can reflect a changed denominator or collection setting rather than better customer performance.

Keep the evidence for the next incident

Mini DBA MySQL monitoring brings query and metric history together with InnoDB and connection diagnostics for investigating recurring incidents. Available evidence depends on the monitored version, permissions and configured collection.

References and further reading

Add comment