Oracle data guard lag alert


The Data Guard Lag alert notifies you when its configured condition is met on Oracle instances so you can investigate and respond.

Screenshot pending: Mini DBA Oracle Data Guard Lag alert screenshot placeholder

Alert summary

  • Platform: Oracle
  • Alert category: Availability
  • Default enabled: true
  • Default evaluation frequency: Minute
  • Threshold label: Seconds lag
  • Unit: seconds

What Mini DBA checks

Mini DBA describes this alert as: Oracle Data Guard transport/apply lag from V$DATAGUARD_STATS Sustained transport lag points at redo shipping, listener, network, or archive destination problems. Sustained apply lag points at standby recovery, I/O, CPU, or workload pressure. Investigate before recovery point or failover readiness is threatened. Mini DBA evaluates this alert once a minute so changes are detected quickly. Because duration is used, Mini DBA can avoid treating a single short spike as a full incident when the condition clears quickly.

How this alert helps

This alert protects high availability and disaster recovery by warning when replicas, replication slots, or apply processes fall behind. Lag can turn a planned failover into data loss risk or leave reporting users looking at stale data.

When to enable it

Enable it where replication, read replicas, failover, or reporting copies are part of the service design. Disable it on standalone systems that do not use replication so the alert list stays focused.

Threshold guidance

Threshold meaning: Seconds lag. Major threshold: 900 seconds. Minor threshold: 300 seconds. Comparison direction: over. Duration is used, so prefer requiring the condition to persist before paging people for transient spikes. For an over-threshold alert, decreasing a threshold makes that severity fire sooner; increasing it tolerates more load or pressure. Set the minor threshold as an early-warning level and the major threshold at the highest acceptable value for the service.

Remediation for an active alert

Check the sender, receiver, apply process, network path, and oldest retained log position. Restart failed replication workers only after capturing the error, then address the root cause such as disk pressure, long transactions, schema drift, network latency, or insufficient replica capacity.

Investigation workflow

  1. Confirm the alert is still active and note the first seen time, affected instance, and severity.
  2. Review the primary, replica, transport path, apply process, retained log position, and oldest open transaction in Mini DBA before changing configuration or ending sessions.
  3. Compare the current value with the normal baseline for the same time of day or maintenance window.
  4. Record the cause, corrective action, and whether thresholds or routing should be adjusted after the incident.

Avoiding alert noise

For recovery and replication alerts, route notifications to the people who own recovery objectives. Higher thresholds can be reasonable for reporting replicas, but production failover paths normally need tighter settings and explicit escalation.

Related pages