MySQL Cloud Alerts: CPU, Storage and Connection Runbooks

A managed MySQL alert should describe a condition that needs an operator decision, not merely a chart crossing a round number. CPU, storage, connections and I/O metrics are valuable early warnings, but they are service-layer signals. The action becomes clear only when they are correlated with MySQL running threads, query load, transactions, buffer-pool behaviour and replication state.

Design alerts as evidence chains

Instead of alerting on every CPU peak, look for sustained CPU with rising application latency or running threads. Instead of a generic storage alert, distinguish data growth, binary-log growth, provider backup use and temporary workload bursts. For replica issues, identify whether the metric represents delay, bytes retained, a thread state or an application-visible freshness risk.

Account for service-specific limits

Burstable tiers may have CPU-credit behaviour that changes the meaning of a high CPU period. Storage offerings can limit IOPS or throughput differently. Some provider network metrics include monitoring, replication or service traffic. Sampling intervals and aggregation can delay a visible spike. Record these characteristics in the alert runbook so on-call staff do not assume a metric is an instantaneous engine counter.

A practical alert checklist

  • Name the exact cloud service, tier and metric definition.
  • Use a duration and baseline that fit the workload cycle.
  • Link to the database query, lock and session checks needed next.
  • Define who owns scaling, query change and application-pool remediation.
  • Test the notification during a controlled load or game day.

Review thresholds after major releases, tier changes and seasonal traffic shifts. An alert is successful when it helps someone make the next correct decision, not when it creates the most messages.

Write the response before choosing the threshold

For CPU, the response might begin with query digests and running threads. For rapidly falling storage headroom, it begins with growth rate and the source of new bytes. For failed connections, it begins with the exact error and whether clients can reach the endpoint. These are different runbooks even if the provider sends all three notifications through the same channel.

ConditionFirst evidence to captureDecision
Sustained CPU with slow requestsQuery mix, concurrency and request latencyReduce expensive work or add justified capacity.
Fast storage growthFree space, recent rate and file categoryProtect headroom and investigate retention.
Connection failuresClient error, reachability and pool stateRoute to network, authentication or capacity owner.

Example: a nightly load causes an alert storm

An import predictably uses substantial CPU for fifteen minutes while customer latency remains acceptable. A low fixed threshold pages every night. Suppressing the entire maintenance period would also hide an import that never completes or consumes all storage. Use the known job schedule as context and monitor its duration and impact, with an escalation for behaviour outside the expected envelope.

Keep thresholds tied to observed service objectives. A workload with strict response limits may need investigation before CPU reaches a conventional round number. A batch service may tolerate sustained utilisation provided its completion deadline and downstream freshness remain safe. One global threshold cannot encode both priorities.

Handle absence and recovery explicitly

Missing metrics can mean a stopped instance, a collection fault, changed permissions or delayed delivery. Decide how each alert treats missing data and test it. Do not fill every gap with zero. Recovery should require evidence that the original condition cleared for a meaningful period; one low sample between repeated bursts should not produce a reassuring recovery message.

Exercise the runbook with a controlled scenario before enabling broad notifications. Confirm the recipient, working links, timestamps and access to the diagnostic view. Measure how long an operator takes to identify the owner and next action. After a real incident, adjust the alert using the observed failure pattern, and retain the previous rule so the reasoning behind the change remains reviewable.

Keep the evidence for the next incident

Mini DBA MySQL monitoring brings query and metric history together with InnoDB and connection diagnostics for investigating recurring incidents. Available evidence depends on the monitored version, permissions and configured collection.

References and further reading

Add comment