> ## Documentation Index
> Fetch the complete documentation index at: https://www.paradedb.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Monitoring and Troubleshooting

> Monitor replication health, resolve apply failures, and recover from WAL buildup

Monitor both the publisher and ParadeDB to catch replication failures before they affect query freshness or fill the publisher’s disk.

## Monitoring the Publisher

Permanent logical replication is operationally safe only if you watch the
publisher, not just the subscriber. The most important signal is how much WAL a
logical slot is retaining.

```sql theme={null}
SELECT
  slot_name,
  active,
  restart_lsn,
  confirmed_flush_lsn,
  pg_size_pretty(pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn)) AS retained_wal,
  wal_status,
  safe_wal_size,
  inactive_since
FROM pg_replication_slots
WHERE slot_type = 'logical';
```

Watch for:

* `retained_wal` growing steadily because the subscriber is not acknowledging
  WAL quickly enough
* `inactive_since` becoming non-`NULL` for longer than expected
* `wal_status` showing that the slot is under pressure
* Filesystem usage on the volume that contains `pg_wal`

To reduce blast radius, configure `max_slot_wal_keep_size` on the publisher.
This caps how much WAL a slot may retain, but it can also invalidate a lagging
subscriber, so it should be paired with alerting and a reseed plan.

## Monitoring the Subscriber

Use the subscriber to confirm that apply workers are healthy and that errors are
not accumulating:

```sql theme={null}
SELECT subname, worker_type, received_lsn, latest_end_lsn, latest_end_time
FROM pg_stat_subscription;

SELECT subname, apply_error_count, sync_error_count
FROM pg_stat_subscription_stats;
```

If `latest_end_time` stops advancing or `apply_error_count` increases, inspect
the subscriber logs immediately.

## Troubleshooting Apply Failures

One common cause of apply-worker failures is schema drift between the publisher
and subscriber.

Two common log patterns for schema drift are:

```text theme={null}
logical replication target relation "public.doctor" is missing replicated columns: "personnel_id", "role_function_id"
```

```text theme={null}
logical replication apply worker for subscription "paradedb_subscription" has started
background worker "logical replication apply worker" (PID 2570238) exited with exit code 1
```

The first message is the root cause. The second means the apply worker crashed
after hitting that error and Postgres will try to restart it.

When you see these messages:

1. Inspect the subscriber logs for the first schema-mismatch error, not just the
   worker restart message
2. Compare the affected table definition on the publisher and ParadeDB
3. Apply the missing DDL on ParadeDB
4. Re-enable or refresh the subscription if needed
5. Rebuild any ParadeDB indexes affected by the schema change

Another common cause of apply-worker failures is a logical replication conflict.
For example, a duplicate key, a permissions failure on the target table, or
row-level security on the subscriber can stop replication even when the schemas
match.

```text theme={null}
ERROR: duplicate key value violates unique constraint ...
CONTEXT: processing remote data during INSERT for replication target relation ...
```

When you suspect a replication conflict:

1. Inspect the subscriber logs for the first conflict error and note the finish
   LSN and replication origin if Postgres logged them
2. Resolve the underlying issue on the subscriber, such as conflicting local
   data, missing privileges, or row-level security policy interference
3. Resume replication normally once the conflict is removed
4. Only if you intentionally want to discard that remote transaction, use
   `ALTER SUBSCRIPTION ... SKIP` with care

Skipping a conflicting transaction can leave the subscriber inconsistent, so it
should be treated as a last resort rather than the default fix. For conflict
types and the Postgres recovery workflow, see the [Postgres logical
replication conflicts
documentation](https://www.postgresql.org/docs/current/logical-replication-conflicts.html).

## Emergency: WAL Keeps Accumulating on the Publisher

If the logical slot on the publisher is filling disk and ParadeDB cannot catch
up quickly enough, the priority is protecting the publisher.

1. First, fix the subscriber if the issue is simple and recent, such as a
   schema mismatch or networking issue
2. If the publisher is running out of disk and the subscriber can be rebuilt,
   remove the subscription or drop the logical slot so the publisher can recycle
   WAL again
3. Recreate the subscription and reseed ParadeDB once the publisher is safe

<Warning>
  Disabling the subscription is not an emergency fix for WAL buildup. A disabled
  subscription still leaves the logical slot behind on the publisher, and that
  slot can continue retaining WAL.
</Warning>

If the subscriber is reachable and healthy enough to cleanly tear down, dropping
the subscription is the cleanest path:

```sql theme={null}
DROP SUBSCRIPTION paradedb_subscription;
```

To protect the publisher from continued `pg_wal` growth when you are
intentionally giving up the current replica state, drop the slot on the
publisher:

```sql theme={null}
SELECT pg_drop_replication_slot('paradedb_subscription');
```

After either step, ParadeDB must be reinitialized from a fresh schema and data
copy before it can resume as a logical subscriber.

## Common Pitfalls

* Starting with pre-populated subscriber tables while using `copy_data = true`
* Applying DDL on only one side of the replication link
* Forgetting that new tables must be added to the publication and refreshed on
  the subscription
* Writing directly to subscribed tables on ParadeDB, which can create conflicts
  with incoming replicated changes
* Leaving a broken logical slot unattended on the publisher until `pg_wal`
  fills disk
* Assuming `ALTER SUBSCRIPTION ... DISABLE` relieves publisher-side WAL pressure
