# Accountants SaaS Chatbot Agent

Multi-tenant RAG chatbot for accounting and finance. Each organization gets its own knowledge base; the shared finance dataset remains available as a platform library.

## Stack

- **LLM:** Ollama (default: `phi3:mini`)
- **Embeddings:** BGE-small-en-v1.5 (sentence-transformers)
- **Vector DB:** ChromaDB (shared `finance_data` + per-org tenant stores)
- **API:** FastAPI
- **Database:** SQLite (users, orgs, KB docs, API keys)
- **Auth:** JWT (workspace users) + API keys

## Prerequisites

1. [Ollama](https://ollama.com) installed and running
2. Python 3.10+
3. Pull a model:

```bash
ollama pull phi3:mini
```

## Setup

```bash
cd accountants_chatbot
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
cp .env.example .env   # edit credentials and model if needed
```

## Index the shared finance dataset

Full dataset (~105K records, 30–60 min on CPU):

```bash
python scripts/ingest.py
```

Quick test with 500 records:

```bash
python scripts/ingest.py --limit 500
```

## Run

### Development (backend + frontend separately)

**Terminal 1 — API**

```bash
uvicorn app.main:app --host 127.0.0.1 --port 8000 --reload
```

**Terminal 2 — React UI (hot reload)**

```bash
cd frontend
npm install   # first time only
npm run dev
```

- **Dev UI:** http://127.0.0.1:5173/ (Vite proxies `/saas` → uvicorn `:8000`)
- **API docs:** http://127.0.0.1:8000/docs

### Production-style (uvicorn serves built UI)

```bash
cd frontend && npm run build && cd ..
uvicorn app.main:app --host 0.0.0.0 --port 8000 --reload
```

- **SaaS agent UI:** http://localhost:8000  
- **Classic anonymous chat:** http://localhost:8000/chat-classic  
- **Admin UI:** http://localhost:8000/admin  
- **API docs:** http://localhost:8000/docs  

### SaaS quick start

1. Open http://localhost:8000 and **Create workspace**
2. Go to **Knowledge base** → upload PDF / TXT / MD / CSV  
3. Ask questions in **Agent chat** — answers come from your org KB (`tenant` mode)  
4. In **Settings**, switch mode to `platform` or `combined` to include the shared finance dataset  

### SaaS API

| Method | Path | Description |
|--------|------|-------------|
| POST | `/saas/auth/register` | Create user + organization |
| POST | `/saas/auth/login` | Login, get JWT |
| GET | `/saas/me` | Current user, org, KB stats |
| PATCH | `/saas/org/kb-mode` | `platform` / `tenant` / `combined` |
| POST | `/saas/kb/documents` | Upload & index KB file |
| GET | `/saas/kb/documents` | List KB documents |
| DELETE | `/saas/kb/documents/{id}` | Remove document |
| POST | `/saas/agent/chat` | Tenant-aware agent chat |
| POST | `/saas/api-keys` | Create API key (`X-API-Key`) |

## Index the dataset

Full dataset (~105K records, 30–60 min on CPU):

```bash
python scripts/ingest.py
```

Quick test with 500 records:

```bash
python scripts/ingest.py --limit 500
```

Resume a previous interrupted run:

```bash
python scripts/ingest.py --resume
```

Re-index from scratch:

```bash
python scripts/ingest.py --reset
```

## Ingest PDF documents

```bash
python scripts/ingest_pdf.py path/to/policy.pdf
python scripts/ingest_pdf.py docs/               # all PDFs in a folder
```

## Run the API

```bash
uvicorn app.main:app --host 0.0.0.0 --port 8000 --reload
```

- **Chat UI:** http://localhost:8000
- **Admin UI:** http://localhost:8000/admin  (default: `admin` / `changeme`)
- **API docs:** http://localhost:8000/docs

## API endpoints

| Method | Path | Description |
|--------|------|-------------|
| GET | `/health` | System status |
| POST | `/chat` | Ask a question |
| POST | `/feedback` | Submit thumbs up/down + correction |
| POST | `/admin/login` | Get JWT token |
| GET | `/admin/feedback` | List feedback (auth required) |
| PATCH | `/admin/feedback/{id}` | Approve / dismiss feedback |
| GET | `/admin/audit` | Audit log of all queries |

## Configuration

Edit `.env`:

| Variable | Default | Description |
|----------|---------|-------------|
| `OLLAMA_BASE_URL` | `http://localhost:11434` | Ollama API URL |
| `OLLAMA_MODEL` | `phi3:mini` | Model name in Ollama |
| `RETRIEVAL_TOP_K` | `5` | Chunks retrieved per query |
| `SIMILARITY_THRESHOLD` | `0.3` | Minimum relevance filter |
| `ADMIN_USERNAME` | `admin` | Admin panel username |
| `ADMIN_PASSWORD` | `changeme` | **Change this in production!** |
| `JWT_SECRET` | random | JWT signing secret |
| `CORS_ORIGINS` | `*` | Allowed origins for CORS |

## WordPress / avatar integration

Point your avatar/chat page to this API:

```bash
curl -X POST http://YOUR_SERVER:8000/chat \
  -H "Content-Type: application/json" \
  -d '{"question": "Which companies are in oil drilling?", "session_id": "wp-user-1"}'
```

Set `CORS_ORIGINS` in `.env` to your WordPress domain (not `*` in production).

## Structured queries (Phase 3)

For filter-style questions, the bot uses pandas on the CSV (not only vector search), e.g.:

- `Show AP transactions over 500 in Q3 2023`
- `Which companies are in the oil drilling industry?`
- `What is the average debt-to-equity ratio?`

## PDF calculation (high-accuracy path)

Open **http://localhost:8000/pdf**

Flow:
1. Upload PDF
2. Digital text PDFs → tables extracted with **pdfplumber**, validated, may calculate immediately
3. Scanned PDFs → **Tesseract OCR draft only** — you must confirm/edit amounts
4. **Calculate** runs deterministic totals (not the LLM)

Install OCR system packages (optional, for scans):

```bash
sudo apt install -y poppler-utils tesseract-ocr
pip install pdf2image pytesseract python-multipart
```

API:
- `POST /pdf/extract` — upload
- `POST /pdf/{job_id}/confirm` — confirm/edit rows
- `POST /pdf/{job_id}/calculate` — compute (blocked until confirmed for OCR)

## Next steps

1. Ingest policy PDFs: `python scripts/ingest_pdf.py path/to/policies/`
2. Connect WordPress avatar page to `POST /chat`
3. Change `ADMIN_PASSWORD` / `JWT_SECRET` before any shared deploy
4. Review thumbs-down feedback in `/admin` and improve prompts/data as needed

## Project structure

```
accountants_chatbot/
├── app/
│   ├── main.py          ← FastAPI app + all routes
│   ├── config.py        ← Settings (pydantic-settings)
│   ├── database.py      ← SQLite (audit, feedback, sessions)
│   ├── auth.py          ← JWT authentication
│   ├── models/
│   │   └── schemas.py   ← Pydantic request/response models
│   ├── pdf/
│   │   ├── extract.py   ← digital tables + OCR draft
│   │   ├── calc.py      ← validation + deterministic totals
│   │   └── store.py     ← confirm-before-calc jobs
│   └── rag/
│       ├── chain.py     ← RAG service (embeddings + Ollama)
│       ├── structured.py← pandas filters for transactions/companies
│       └── prompts.py   ← System + user prompt templates
├── scripts/
│   ├── ingest.py        ← CSV → ChromaDB ingestion
│   └── ingest_pdf.py    ← PDF → ChromaDB ingestion
├── static/
│   ├── chat.html        ← Chat UI
│   ├── pdf.html         ← PDF calculator UI
│   └── admin.html       ← Admin feedback review panel
├── data/
│   ├── chroma/          ← Vector store (auto-created)
│   ├── pdf_uploads/     ← Uploaded PDFs + job JSON
│   └── chatbot.db       ← SQLite database (auto-created)
├── final_finance_dataset.csv
├── requirements.txt
└── .env
```
