# Financial Document Intelligence Platform

Architecture principle:

- Structured systems extract, calculate, search, and verify.
- The LLM explains grounded evidence — it must not invent totals or IDs.

## Package layout (`app/fdi/`)

| Module | Role |
|--------|------|
| `types.py` | Document / identifier / intent taxonomies |
| `schema.py` | Durable SQL SoR (`fdi_documents`, line items, identifiers, entities) |
| `classify.py` | Multi-type document classifier |
| `identifiers.py` | Exact GSTIN/PAN/UTR/IFSC/… extract + lookup (never embeddings) |
| `extract.py` | Typed extraction → lines, assertions, parties |
| `extractors.py` | Type enrichers (invoice, GST, payslip, CAS, insurance, loan, BS/P&L, receipt, PO) |
| `store.py` | Persist structured index |
| `entities.py` | Merchant alias seed + counterparty relink |
| `compute.py` | Deterministic SUM/COUNT/GROUP/ID tools |
| `multidoc.py` | Compare, reconcile invoice↔bank, timeline, duplicates |
| `hybrid.py` | BM25 + dense RRF fusion (+ optional cross-encoder rerank) |
| `planner.py` | Intent → tool plan → execute |
| `verify.py` | Confidence + grounded answer rendering |
| `pipeline.py` | Dual-write hook after vector ingest |

## Chat modes (`KB_CHAT_MODE`)

| Mode | Behavior |
|------|----------|
| `fdi` | Planner SQL first, hybrid RAG fallback (**default**) |
| `rag` | Vector knowledge-base only (hybrid retrieval when enabled) |
| `hybrid` | Legacy statement insight cards |

## Env flags

```
KB_CHAT_MODE=fdi
FDI_ENABLED=true
FDI_HYBRID_SEARCH=true
FDI_RERANK=false
```

## Ops

```bash
python scripts/backfill_fdi.py   # re-index ready PDFs into FDI tables
```

Hierarchical chunk metadata (`parent_chunk_id`, `section`) is written on ingest for page→child retrieval.
