# Incident 2026-03-14 — billing intake outage (Redis OOM)

**Severity:** SEV-1 · **Duration:** 03:12–06:41 UTC (3h29m) · **Status:** closed

## Impact

Partner intake stopped accepting deliveries for three and a half hours.
Approximately 4,100 corrections queued upstream at the partner and were
re-delivered after recovery. No customer statement was issued during the
window. No money was misposted.

## What happened

The `redis` instance reached its configured `maxmemory` at 03:12. The eviction
policy was `noeviction`, so writes began returning OOM errors. `deliver()`
raised on every call and the web tier returned 500s. Celery could not enqueue,
so the worker went idle with a full broker.

On-call restarted the instance at 06:38. Because Redis runs without persistence
the restart came up empty, intake recovered immediately, and the partner
re-delivered the backlog.

## Root cause

Memory use had been growing monotonically since the service launched. The
growth was linear in total corrections ever delivered, not in corrections
in flight, so it never plateaued. Nothing expired.

2 GB `maxmemory`. Extrapolating backwards, the instance had been filling for
approximately eighteen months.


## Remediation

1. **Done (2026-03-14).** Restarted the instance to clear memory.
2. **Done (2026-03-17).** Added a scheduled purge — `adjustments.maintenance.
   purge_cache`, running nightly at 04:00 UTC — so memory cannot accumulate to
   the point of another outage. Redis is a cache; per `docs/ARCHITECTURE.md` it
   holds no system-of-record data, so a scheduled flush is safe and requires no
   coordination.
3. **Done (2026-03-17).** Alert on `used_memory` above 70% of `maxmemory`.
4. **Not done.** Find and fix whatever is writing keys that never expire. The
   nightly purge removes the urgency; deprioritised at the 2026-03-21 review.

## Follow-up notes

The nightly purge has run without incident since 2026-03-17 and memory now
sawtooths comfortably under 400 MB. Item 4 remains open but is no longer
considered load-bearing.
