PostgreSQL Inactive Replication Slots Alert


Use the Inactive Replication Slots alert in Mini DBA to monitor PostgreSQL instances and make this condition visible before it becomes a wider database incident.

Screenshot pending: Mini DBA PostgreSQL Inactive Replication Slots alert screenshot placeholder

Alert Summary

  • Platform: PostgreSQL
  • Alert category: Availability
  • Default enabled: true
  • Default evaluation frequency: Minute
  • Threshold label: Inactive slots
  • Unit: slots

What Mini DBA Checks

Mini DBA describes this alert as: Inactive replication slots can retain WAL indefinitely. Review whether the slot is expected to be inactive or should be dropped. The default evaluation frequency is Minute, so the alert is intended to be close enough to operational reality for live triage. Because duration is used, Mini DBA can avoid treating a single short spike as a full incident when the condition clears quickly.

Why This Alert Is Helpful

This alert protects high availability and disaster recovery by warning when replicas, replication slots, or apply processes fall behind. Lag can turn a planned failover into data loss risk or leave reporting users looking at stale data.

When To Enable It

Enable it where replication, read replicas, failover, or reporting copies are part of the service design. Disable it on standalone systems that do not use replication so the alert list stays focused.

Threshold Guidance

Threshold meaning: Inactive slots. Major threshold: 0 slots. Minor threshold: 0 slots. Comparison direction: "over". Duration is used, so prefer requiring the condition to persist before paging people for transient spikes. Use higher thresholds on batch-heavy, development, or intentionally bursty systems where brief pressure is expected. Use lower thresholds on latency-sensitive production systems, small instances with little headroom, and services with strict recovery or availability commitments.

Remediation For An Active Alert

Check the sender, receiver, apply process, network path, and oldest retained log position. Restart failed replication workers only after capturing the error, then address the root cause such as disk pressure, long transactions, schema drift, network latency, or insufficient replica capacity.

Investigation Workflow

  1. Confirm the alert is still active and note the first seen time, affected instance, and severity.
  2. Review the top consumers, host counters, service tier or instance size, and recent workload changes in Mini DBA before changing configuration or ending sessions.
  3. Compare the current value with the normal baseline for the same time of day or maintenance window.
  4. Record the cause, corrective action, and whether thresholds or routing should be adjusted after the incident.

Avoiding Alert Noise

If this alert is noisy, check whether the workload naturally has short bursts. Raise the duration or threshold for expected batch windows, but keep a lower route for production systems where user-facing latency matters. Avoid simply disabling the alert until you know whether the noise is caused by threshold choice or by a real capacity trend.

Related Pages