# Invoicing challenge

**CONFIDENTIAL CHALLENGE MATERIAL — do not share or post any part of this
package or your solution, ever.** Doing so ends your candidacy and permanently
disqualifies you from all current and future roles with the hiring company.

This package is challenge version `inv-2026.10`.

---

## Read this first: what is real and what is fiction

**This file and `tools/` are the only real-world artifacts in this package.**
They are the hiring company talking to you directly: what to do, how long you
have, and how to package what you produce.

**Everything else is in-world.** `adjustments/`, `billing/`, `docs/`,
`docs/chat/`, `docs/incidents/`, and `tests/` are the fictional company's
codebase, documentation, team conversation, and incident history. Treat them as
you would inherited production material at a new job: written by real people,
at different times, under pressure, with the errors that implies. Nothing in
them is guaranteed to be accurate, current, or consistent with the code.

### Which source governs when they disagree

They will disagree. Resolving that is part of the work, so these are the rules:

- **The standing contracts in `docs/` govern.** Chat messages and attachments
  are input, not contract.
- **An attachment or a chat message changes a contract only where Marcus
  explicitly adopts a specific part**, and the change extends only to the part
  he named. An unadopted spec is a proposal, however detailed it looks.
- **The documents describe the system as it was last reviewed.** They are not
  generated from the code and are not verified against it. Where a document and
  the running system disagree, you decide which one is wrong, say so, and
  defend the call in your release note.

---

## Your brief

You are Gordon, the engineer who has just taken over billing reconciliation.
Read `docs/chat/billing-recon.md` — the team's channel, in date order, plus the
one attachment posted in it. That is where the work comes from.

You have **4 hours** from submitting the Start Form.

**Reserve the final 30 minutes for export, packaging, and submission.** That
is not padding. `make submit` runs checks that can fail, the archive has to be
uploaded somewhere that serves it without a login, and the Results Form takes
a few minutes to fill in honestly. Candidates who code until the last moment
lose work that was finished but never delivered.

**There is more work here than fits in 4 hours, on purpose. You are not
expected to finish everything.** There is roughly a day of work in the
channel. What we are reading is which parts you chose, how well you did them,
and whether you can say plainly what you left alone.

Play to your strengths. The work divides
about three ways, and a strong submission will go deep on one or two:

- **Data and durability** — what survives a restart, a flush, a crash
  mid-operation, a task delivered twice. `docs/ARCHITECTURE.md`.
- **Business logic** — what the money should be, in what order corrections
  apply, what an issued invoice is allowed to say afterwards.
  `docs/PRICING_MATH.md` and `docs/SETTLEMENT_RUNS.md`.
- **The customer-facing screen** — whether someone who is not us can look at a
  corrected bill and understand it. `docs/STATEMENT.md` and `mobile/`.

Depth beats breadth. Four areas touched shallowly reads worse here than one
area finished and three declared untouched — and the confidence table is where
you declare that. **An honest `untested` costs you very little. A `verified`
on something that turns out to be broken costs you a great deal.** Guessing
high is the single most expensive thing you can do in this exercise.

**Do a test export in your first five minutes.** See "Approved agent surfaces
and required exports" below. If you discover at the end that you cannot export
your transcripts, your submission will not be reviewed — and by then it is too
late to switch tools.

A reading order that avoids meeting terms before they are defined:
`docs/GLOSSARY.md`, `docs/ARCHITECTURE.md`, `docs/PRICING_MATH.md`,
`docs/SETTLEMENT_RUNS.md`, `docs/STATEMENT.md`, the incident report, then the
channel and its attachment, then the code.

## The stack

Django · MySQL · Celery · Redis. One command brings it all up:

```
make up          # build and start web, worker, mysql, redis
make test        # run the visible tests
make down        # tear down
```

**Why nothing here checks who is asking.** This is the internal billing
service. In production it sits behind the customer app's gateway, which
authenticates the customer and scopes the request to their own account before
it ever reaches Django, so the views in `api/` take the account id as given.
Locally you talk to the service directly, which is why
`GET /api/accounts/<id>/statement/<period>` will answer for any account id you
type. Authentication, authorization and rate limiting are the gateway's, and
they are not in this repository. Don't add them here. If you think that
boundary is wrong, or you find something that makes it unsafe, that belongs in
your release note — noticing it is worth more to us than building it.

`make up` only migrates the database -- it creates no account and no invoice.
`GET /api/accounts/acct-demo/statement/2026-06` 404s against a fresh stack.
Seed the deterministic demo account once the stack is up:

```
make demo        # creates acct-demo (period 2026-06) with a few corrections
```

`make demo` is idempotent: running it again against the same database is a
no-op, not an error. It creates one issued invoice and a handful of
corrections chosen to be the interesting cases in `docs/STATEMENT.md`'s
audit-history section -- among them a correction withdrawn to a 0.00 net
effect, and a version later superseded by a replacement -- so the screen
below has something real, and non-trivial, to render. Confirm it worked
without the app at all:

```
curl http://127.0.0.1:8000/api/accounts/acct-demo/statement/2026-06
```

The customer-facing statement screen is a React Native app in `mobile/`. It
talks to the service over the API in `api/` and renders what
`docs/STATEMENT.md` describes. It runs on the host rather than in docker:

```
cd mobile && npm install
make test-mobile   # the visible mobile tests, from the repository root
```

**Scope: this repository has the statement screen, not the whole app.** The
screen was moved into the customer app about a year ago; sign-in, the account
switcher and the customer's list of invoices belong to that app and are not
billing's code, so none of it is checked out here. There is nothing to sign
into — `App.js` mounts `StatementScreen` straight onto the seeded demo account
so you can run it. Don't build a login or an invoice list, and don't add a
navigation library. If the lack of functionality is genuinely in your way, say so
in your release note.

### Running the app against the stack above

The screen's API base URL is `mobile/src/api.js`'s `BASE_URL`: the
`EXPO_PUBLIC_STATEMENT_API` environment variable if set, else
`http://localhost:8000`. Which value is correct depends on where the app
itself is running, not on where Django is running (Django is always in
Docker, published on host port 8000) -- copy `mobile/.env.example` to
`mobile/.env` and uncomment the line for your target:

| Where the app runs | Correct API base URL | Why |
|---|---|---|
| Web (`npm run build:web`, or press `w` after `npm start`) | `http://localhost:8000` (default -- no `.env` needed) | The browser and Docker's published port are on the same host. |
| iOS simulator (`npm run ios`) | `http://localhost:8000` (default -- no `.env` needed) | The simulator shares the host Mac's network stack, so "localhost" already means this machine. |
| Android emulator (`npm run android`) | `http://10.0.2.2:8000` | The emulator has its own virtual network; "localhost" there means the emulator itself, not your host. `10.0.2.2` is the emulator's fixed alias for the host machine. |
| A physical device on the same Wi-Fi/LAN | `http://<your-host-LAN-IP>:8000` | Neither `localhost` nor `10.0.2.2` reaches your machine from a separate device. |

A complete path from nothing to seeing the screen with real data:

```
make up && make demo             # stack up, demo account seeded
cd mobile && npm install
npm start                        # Metro/Expo dev server; press w for web
```

**Some visible tests fail on the package as supplied. That is the work, not a
broken checkout.** They are ordinary failures with ordinary messages; read them.

**The first `make up` takes about a minute after MySQL reports healthy**, while
migrations run. `curl` will refuse the connection until it finishes — `docker
compose logs -f web` shows what it is doing.

The visible tests in `tests/` and `mobile-tests/` cover the basic supported
workflow only. They are not the contract. A larger private suite will run against
your submission, derived from the documents in `docs/`. **Both the service and
the mobile UI are graded**, and the UI is 30 of the 100 points.

**The runtime is fixed.** Your submission is checked in an environment with no
network access that installs exactly the four pinned packages in
`requirements.txt`. Anything you import beyond those and the standard library
will fail there even if it works on your machine.

Both settlement policies described under "Closed periods" in
`docs/PRICING_MATH.md` are exercised. Behavior is credited only where it holds
for both, so handling one family and not the other scores as handling neither.

## What to deliver

1. **Working code.** Fix what is broken, build what is asked for, and keep the
   documented contracts. Take ownership and think like an actual engineer planning
   to be at the company long-term.
2. **Document corrections.** If you conclude a standing contract is wrong,
   update it for the next engineer and defend why in `RELEASE_NOTE.md`. A human
   judges corrections; automated grading stays against the contracts as
   shipped, so edits never move the goalposts.
3. **`RELEASE_NOTE.md`** at the repository root, covering:
   - what this release includes and what it is gated on — anything you are
     deliberately not shipping, and any precondition someone else has to
     satisfy first;
   - what changed and why;
   - verification performed versus assumptions;
   - residual risk;
   - the disposition of **every** ask in the channel, with a counterexample or
     a contract citation for anything you reject.
4. **A confidence table** in `RELEASE_NOTE.md`, under the heading
   `## Confidence`. Declare exactly one confidence value for each of the four
   contract documents below. See "How the confidence table is scored".
5. **Transcripts.** Every agent session, exported natively, in
   `transcripts/native/`, listed in `transcripts/INDEX.md`.

You are expected to work agentically. Use an approved agent surface, keep the
complete native transcript, and declare it. Don't edit files by hand.

### How the confidence table is scored

Declare one confidence value for each of the four contract documents. Each row
means **the behavior that document specifies** — if you changed something, the
document you were reading is the row it belongs to. Each row asks how much of
what was *broken* there you repaired; behavior that already worked is not
counted, so a document you never needed to touch cannot count against you.

Copy this table into `RELEASE_NOTE.md` and fill in every row:

| Contract document | Confidence |
|---|---|
| `docs/PRICING_MATH.md` | |
| `docs/ARCHITECTURE.md` | |
| `docs/SETTLEMENT_RUNS.md` | |
| `docs/STATEMENT.md` | |

| Value | Means |
|---|---|
| `verified` | I fixed what was broken here and checked that it works |
| `confident` | I believe I fixed it; my checking was partial |
| `unsure` | I cannot vouch for it either way |
| `untested` | I did not get to this |

Scoring rewards an honest, accurate reading of your own work: over-claiming
costs more than under-claiming, so the truthful answer is also the
best-scoring one.

Fill in every row; a row you leave out is scored as if you had declared
`verified`, and omitting the table entirely scores zero. Explain your reasoning
in prose after the table — write it however you like, the table is what we
read.

### Approved agent surfaces and required exports

| Approved surface | Required export |
|---|---|
| Claude Code CLI | Run `/export transcripts/native/claude/cli-<id>.txt` in every session; also copy the matching raw JSONL and any `subagents/` records from `~/.claude/projects/`. |
| Claude Desktop's local Code tab | Local top-level **Code tab** sessions only. Run `/export transcripts/native/claude/desktop-<id>.txt`; also copy the matching raw JSONL and `subagents/` records. |
| Codex CLI | Copy session JSONL from `${CODEX_HOME:-~/.codex}/sessions/` including dated subfolders. Do not use `--ephemeral`. |
| Codex in the ChatGPT desktop app (local) | Copy local session JSONL from `${CODEX_HOME:-~/.codex}/sessions/`. |
| Cursor's local IDE Agent | Use the Agent chat history's Export action; save the native Markdown. |
| Cursor Agent CLI in captured print mode | `--print --output-format stream-json`, captured through `tee`. |

Use this exact table in `transcripts/INDEX.md`:

| Approved surface | Exact model identifier | Reasoning/thinking setting | Purpose | Session/thread ID | Parent ID | Native export path |
|---|---|---|---|---|---|---|
| Codex CLI | gpt-5.6-sol | xhigh | Example only | session-id | not delegated | transcripts/native/codex/session.jsonl |

State the model identifier exactly as the surface reports it. Use `not exposed`
for the reasoning setting **only** if the surface genuinely does not display
one. Export transcripts unmodified; you may redact secrets, but note every
redaction in `INDEX.md`. Summaries, screenshots, or rewritten logs do not
replace native exports.

**Do a test export before you start.** If you discover at the end that you
cannot export your transcripts, your submission will not be reviewed.

## Packaging and submission

Do not build the archive by hand, and do not use your file manager's
"Compress" — it adds metadata that makes the archive unreadable to us.

```
make submit                     # builds submission.zip, prints the SHA-256
make verify                     # runs the same checks we run
make check-link URL=<your url>  # proves your link serves the zip anonymously
```

Then upload the ZIP to a **direct download link that needs no login or
cookies**, keep it unchanged and downloadable for seven days, and submit the
Results Form with the URL and the lowercase SHA-256 that `make submit` printed.

`make check-link` exists because a share link that shows *you* a preview page
in your logged-in browser very often serves that same preview page — not your
file — to us. Use it to check that we'll be able to process your submission.
