PostgreSQL WAL archiver failures alert


The WAL Archiver Failures alert notifies you when its configured condition is met on PostgreSQL instances so you can investigate and respond.

Screenshot pending: Mini DBA PostgreSQL WAL Archiver Failures alert screenshot placeholder

Alert summary

  • Platform: PostgreSQL
  • Alert category: Disk
  • Default enabled: true
  • Default evaluation frequency: Minute
  • Threshold label: Failures per minute
  • Unit: failures/min

What Mini DBA checks

Mini DBA describes this alert as: WAL archiving failures can break point-in-time recovery and let WAL accumulate on disk. Mini DBA evaluates this alert once a minute so changes are detected quickly.

How this alert helps

This alert protects storage and recovery capacity. It helps you find growth, retention, and backup conditions that can stop writes, break recovery objectives, or leave the server without enough working space for normal database activity.

When to enable it

Enable it once the instance has a normal workload baseline. It is especially useful on busy OLTP, reporting, and integration systems where resource pressure can quickly become user-visible.

Threshold guidance

Threshold meaning: Failures per minute. Major threshold: 1 failures/min. Minor threshold: 0 failures/min. Comparison direction: over. For an over-threshold alert, decreasing a threshold makes that severity fire sooner; increasing it tolerates more load or pressure. Set the minor threshold as an early-warning level and the major threshold at the highest acceptable value for the service.

Remediation for an active alert

Free space by removing safe-to-delete files, expanding the volume or tablespace, moving growth-heavy objects, correcting retention settings, or shrinking only after a documented one-off event. For recovery-related areas, verify that backups and log shipping or archiving are healthy before deleting anything.

Investigation workflow

  1. Confirm the alert is still active and note the first seen time, affected instance, and severity.
  2. Review the affected disk, tablespace, log stream, backup job, retention setting, and recent growth pattern in Mini DBA before changing configuration or ending sessions.
  3. Compare the current value with the normal baseline for the same time of day or maintenance window.
  4. Record the cause, corrective action, and whether thresholds or routing should be adjusted after the incident.

Avoiding alert noise

Storage alerts should usually stay enabled even on quiet systems because the impact of missing them is high. Tune warning thresholds to leave enough time for approval, provisioning, and validation, especially when storage changes are handled by another infrastructure team.

Related pages