A cloud database estate often puts MySQL, PostgreSQL and Oracle behind dashboards that all display CPU, memory, storage, network and connections. That similarity is useful for triage, but it can be dangerous when it hides what the metric actually measures. A reliable investigation map has a common first step and platform-specific second step.
Use a common triage layer
At estate level, ask which service changed, when it changed, who owns it and whether the impact is isolated. Compare availability, provider capacity signals, connection growth, storage trend and application latency. This gives an incident commander a consistent way to prioritise without pretending that Oracle DB time, PostgreSQL wait events and MySQL InnoDB activity are equivalent.
Then switch to the engine's own evidence
- MySQL: inspect running threads, InnoDB locking, buffer-pool behaviour, query access and replication.
- PostgreSQL: inspect backend state, wait events, locks, transaction age, WAL and vacuum activity.
- Oracle: inspect DB time, wait classes, sessions, SQL, redo, I/O and tablespace behaviour.
Document metric semantics
For each provider and service tier, record the metric source, unit, aggregation, retention, delay and included service activity. A storage percentage may include logs and provider-managed files. A network total may include replication. A connection count may include idle or service sessions. These details determine whether two graphs can be compared and whether a threshold is meaningful.
The result is a disciplined hand-off: triage with a shared operational vocabulary, investigate with native evidence, and record the exact metric meaning in the post-incident review. That approach scales better than trying to force every database into one generic performance score.
Start with the customer operation
Write a concrete symptom: checkout p95 increased, an import missed its deadline, or a report timed out. Record the affected endpoint, tenant or job, the first observed time and whether failures are widespread. This makes it possible to distinguish a database problem from a slow downstream service, a connection-pool queue or a network path that the database never sees.
Create a small incident timeline containing the application symptom, resource signals, changed SQL or sessions and recent deployments. Preserve source timestamps and time zones. A dashboard screenshot without a time range is much less useful during the next shift or a post-incident review.
A worked triage scenario
Imagine all three engines are healthy by reachability checks, but one business service becomes slow after an application release. Its database CPU is low, connections rise and requests time out. Start by identifying the target database and active sessions. If the sessions wait on locks, follow transaction ownership. If the database sees little new work, inspect the client pool and application dependencies instead.
The engine determines the detailed branch. MySQL may expose an InnoDB transaction holding rows; PostgreSQL may show an old idle transaction; Oracle may identify a blocking session and enqueue event. These observations can support a similar application explanation without being interchangeable counters or identical remedies.
Build an evidence handoff between teams
A useful handoff to the application owner includes the business operation, session identity, statement or query fingerprint, transaction age and the interval of impact. A handoff to infrastructure includes service identity, storage or compute configuration, measured latency and the matching workload period. Avoid forwarding only a graph with the instruction to add capacity.
Mark unknowns explicitly. If historical SQL was not collected, say that the current snapshot cannot reconstruct the earlier query. If host access is unavailable on a managed service, identify which provider metric or support route covers that layer. An honest gap is more useful than an invented explanation.
Close the incident with a reproducible comparison
Record the intervention and compare the same customer operation under a similar workload. Validate successful throughput, tail latency and errors, then look for shifted pressure on replicas, storage or application queues. Keep this bundle with the runbook. The next incident should begin with a known method and a usable baseline rather than the same argument over which dashboard to trust.
Keep the evidence for the next incident
Mini DBA cross-platform monitoring provides a common route into database-specific diagnostics and retained query and metric history. Cloud-provider metrics remain a separate source for service capacity and infrastructure evidence.
References and further reading