# Architecture

Billing reconciliation is split across four processes:

- **web** (`Django`) accepts deliveries and issues statement runs;
- **worker** (`Celery`) resolves accepted corrections into marginal records;
- **mysql** holds the system of record;
- **redis** backs the Celery broker and the shared cache.

## System of record

**MySQL is the system of record.** Accounts, invoices, adjustments, adjustment
records, statement runs, and run membership are all durable rows. Any question
about what a customer owes, what a run settled, or what a correction did is
answerable from MySQL alone.

`adjustments.models` holds the durable ORM rows; `pricing.models` holds
in-memory value objects. `services.deliver` persists acceptance and queues
`tasks.apply_correction`, which hydrates account state, invokes
`AdjustmentIntake` over `PricingEngine`, and persists the result.

## Cache

**Redis is a cache. It holds no system-of-record data and may be flushed or
restarted at any time without loss.** It is deliberately configured without
persistence (`--save "" --appendonly no`) because everything in it is either
reconstructible from MySQL or is in-flight broker state.

Operations runbook: if latency degrades or memory pressure rises, `FLUSHALL`
is safe on a live system and requires no coordination with engineering.

A scheduled purge (`adjustments.maintenance.purge_cache`) clears the cache
nightly at 04:00 UTC. It was added after incident 2026-03-14, when unbounded
key growth reached `maxmemory` and stopped intake for three and a half hours.
See `docs/incidents/2026-03-14-redis-oom.md`. The purge is safe for the reason
above: nothing in the cache is a system of record.

## Intake path

`deliver()` accepts a correction, fixes its position in the account's
acceptance order, and enqueues resolution. Delivery is at-least-once: partners
retry, so the same `adjustment_id` can arrive repeatedly and must produce the
same economics.

An accepted `adjustment_id` therefore carries one stable economic meaning, and
it is the identifier — not the payload — that decides which correction a
delivery is. A redelivery under an accepted identifier is that same accepted
correction; a delivery reusing an accepted identifier with different economics
contradicts what was already accepted, so the accepted correction stands
and the conflicting payload changes nothing. Whether you refuse it loudly
or absorb it quietly is yours to decide and defend.
Distinct identifiers are distinct corrections and stay separately attributable
even when every other field matches, because two identifiers may be two real
corrections that happen to look alike. Receipt timestamps are transport facts:
they record when a delivery reached us and never make two deliveries the same
correction or one correction two.

Resolution runs on the worker, and stays there. It replays the account's whole
correction history to work out what one correction changes, so its cost grows
with the account's lifetime rather than with the delivery. Partners retry in
bursts and intake has to acknowledge fast. Ops re-drives the accepted-but-not-
posted backlog from outside the app, which needs a queue to re-drive.

## Delivery reconciliation

Acceptance and completion are reconciled against durable lifecycle state. For
any account, operations can enumerate the exact `adjustment_id` set accepted
but not yet posted; a count without identities is not a sufficient
reconciliation answer.

That reported set is also the recovery worklist. Re-driving it advances every
incomplete delivery to posted, and repeating recovery is safe: work already
applied is not applied a second time.

**Operations seam.** The ops runbook and the nightly reconciliation job both
call this from outside the app, so the entry points are fixed:

| Callable | Returns |
|---|---|
| `adjustments.services.reconciliation_report(account_id)` | the `adjustment_id`s accepted but not posted, in acceptance order |
| `adjustments.services.redrive_reconciliation(account_id)` | nothing; advances that set to posted, safe to repeat |

Keep those two names and signatures. Everything behind them is yours — a
queryset, a manager, a task, a management command wrapping either — but the
runbook calls these, so renaming them breaks operations.

## Fixed names

Other things bind to this code by import path and attribute name: the ops
runbook, the finance export, the nightly job, the recovery harness. A rename
that looks internal from in here is an outage out there, and nothing in this
repository will tell you it happened. Those places are marked in the code with
`ops-contract:` comments.

Treat them as a naming constraint, not a design constraint. Everything behind
those names is yours — rewrite internals, split or merge modules, add fields and
tables alongside. Re-exporting counts as keeping a name: if a module becomes a
shim over code that now lives elsewhere, the import still resolves and that is
fine. What breaks is a path that stops resolving, or one that resolves to
something no longer carrying the name that was asked for.

## Ordering

Corrections are attributed in **effective order**, `(effective_at,
adjustment_id)`, not in arrival order. Arrival order controls only which
acceptance ordinal an event consumes; it never controls the final position.

## Settlement runs

A run fixes its number, operation identifier, predecessor, and correction
cutoff when it starts. Membership is exactly the records whose ordinal falls in
`previous_cutoff < ordinal <= cutoff`. A retry of the same `operation_id`
resumes the same run rather than issuing a new one.

## Scaling notes

The worker runs with `--concurrency=4` and `prefetch_multiplier=1`. Tasks are
acknowledged late so that a worker crash re-delivers the message rather than
dropping it.
