# Accountants Chatbot — How It Works (Non-Technical Guide)

This guide explains how the accountants chatbot works, which AI models it uses, how to make answers faster, and which models are better when you care more about quality than speed.

For install commands and API details, see `README.md`.

---

## What the Accountants Chatbot Does

The accountants chatbot is a local accounting assistant. Everything runs on your machine — no ChatGPT cloud account is required for the main chat.

The accountants chatbot can:

1. **Answer finance questions** from a large combined finance dataset (about 4–5 source datasets merged into ~105,000 records).
2. **Explain accounting concepts** in plain language (e.g. current ratio, working capital).
3. **Do exact math** (e.g. `12 + 5 × 3`) without guessing.
4. **Read balance-sheet PDFs**, pull out numbers, check if the sheet balances, and answer follow-up questions about that document.
5. **Remember the current chat** so you can ask follow-ups.
6. **Save feedback** (thumbs up/down) for later review in the admin panel.

---

## How a Question Is Handled (Simple Flow)

When you send a message, the accountants chatbot does **not** always “think with AI.” It picks the cheapest accurate path first:

```
Your message
    │
    ├─ Pure math?  → Calculator (instant, exact)
    │
    ├─ PDF uploaded in this chat?  → Read that PDF + ask the language model
    │
    ├─ Greeting / help / thanks / bye?  → Short friendly reply
    │
    ├─ Filter-style data question?  → Search the spreadsheet directly
    │                              (e.g. “companies in oil drilling”)
    │
    └─ Everything else  → Find similar records in the knowledge base
                           then ask the language model to answer
```

**Why this matters:** Math and spreadsheet filters are fast and precise. The language model is used when you need explanation, PDF reading, or open-ended answers — and that is usually the slower step.

---

## Models Used Today

The accountants chatbot uses **two different kinds of models**, for two different jobs.

### 1. Reasoning / chat model (the “brain”)

| Item | Current setting |
|------|-----------------|
| Product | **Ollama** (runs models on your PC) |
| Model name | **`phi3:mini`** (Microsoft Phi-3 Mini) |
| Where to change | `.env` → `OLLAMA_MODEL` |
| Used for | Explaining concepts, dataset answers, PDF Q&A, greetings |
| Not used for | Pure arithmetic; PDF total/ratio math (those are calculated separately) |

**In plain terms:** Phi-3 Mini is a small, efficient model. It is a good default on a normal laptop or CPU-only machine because it answers quickly enough for demos while still handling accounting language reasonably well.

### 2. Search model (the “librarian”)

| Item | Current setting |
|------|-----------------|
| Model | **`BAAI/bge-small-en-v1.5`** |
| Job | Turn questions and documents into numbers so similar text can be found quickly |
| Used when | Looking up relevant rows from the finance dataset |

This model does **not** write answers. It only helps find the right pieces of data to show the reasoning model.

### 3. What is not an AI model

- **Calculator** — normal math rules (exact).
- **PDF number totals / ratios** — fixed accounting formulas (totals, current ratio, debt-to-equity, etc.).
- **Spreadsheet filters** — direct lookup/filter on the CSV.

So when someone asks “which model calculates the current ratio?” — the answer is: **none**. The ratio is calculated by rules; the language model only explains or discusses it when asked in chat.

---

## The Finance Dataset

The knowledge base is one large combined file built from **about 4–5 finance/accounting datasets** (investment survey data, accounting transactions, general accounting data, and company financial statements such as P&L, balance sheet, cash flow, ratios, and prices).

- **File:** `final_finance_dataset.csv` (the merged result)
- **Size:** about **105,000** records after combining
- **Content:** companies, industries, transactions, financial statements, ratios, and related finance fields from those source datasets
- **How the accountants chatbot uses it:** After a one-time “indexing” step, the accountants chatbot can search this combined data and answer questions about it.

Without indexing, dataset Q&A will not work. PDF upload and calculator can still work if the rest of the app is running.

---

## How PDF Understanding Works (Non-Technical)

When you attach a balance sheet PDF:

1. The accountants chatbot **extracts text and tables** from the file (digital PDFs are best).
2. If the PDF is a **scan/image**, optional OCR can try to read it (less reliable — review numbers if accuracy matters).
3. The accountants chatbot **adds up** assets, liabilities, and equity and computes simple ratios.
4. The **language model** reads the extracted text and answers your question in normal English.

The accountants chatbot does **not** “look at the page like a human with eyes” using a vision model for digital PDFs. It reads the text inside the PDF (or OCR text for scans).

**Tip for demos:** Use clear, text-based PDFs (e.g. samples in `samples/balance_sheets/`).

---

## Response Time — What Makes Answers Slow or Fast

Typical time is spent on:

| Step | Usually |
|------|---------|
| Math-only questions | Very fast |
| Spreadsheet filter questions | Fast |
| Short greetings | Fast–medium |
| Dataset + AI explanation | Medium–slow |
| PDF upload + first question | Slowest (file processing + AI) |
| First message after starting the app | Extra slow (models loading into memory) |

### How to decrease response time

Do these in order of impact:

#### A. Use a smaller / faster chat model (biggest lever)

Current default **`phi3:mini`** is already chosen for speed on CPU.

To go even faster (with some quality loss):

- Try smaller / faster Ollama models such as **`qwen2.5:0.5b`**, **`qwen2.5:1.5b`**, or **`llama3.2:1b`**
- Avoid large models on CPU (e.g. 8B–70B) unless you have a strong GPU

Change in `.env`:

```text
OLLAMA_MODEL=phi3:mini
```

Then pull the model in Ollama and **restart** the app.

#### B. Shorten how much the model is allowed to write

In `.env`:

| Setting | What it does | Faster tip |
|---------|--------------|------------|
| `OLLAMA_NUM_PREDICT` | Max length of normal/data answers | Lower (e.g. `120`–`160`) |
| `OLLAMA_NUM_PREDICT_CONVERSATIONAL` | Max length of greetings/help | Lower (e.g. `60`–`80`) |
| `OLLAMA_NUM_CTX` | How much context the model keeps | Keep modest (e.g. `2048`; higher = slower) |

Shorter answers = faster replies. Too low = cut-off or thin answers.

#### C. Retrieve less data per question

| Setting | Faster tip |
|---------|------------|
| `RETRIEVAL_TOP_K` | Keep at `3` or try `2` |
| `MAX_CONTEXT_CHARS_PER_DOC` | Keep around `600` or slightly lower |
| `SIMILARITY_THRESHOLD` | Slightly higher (e.g. `0.5`) to skip weak matches |

Less text sent to the model → faster generation.

#### D. Hardware

- **GPU** with Ollama makes a large difference for any model above ~3B parameters.
- On **CPU only**, prefer Mini / 1B–3.8B-class models.
- Close other heavy apps during demos.
- Keep Ollama running so the model stays warm.

#### E. Product habits that feel faster

- Prefer **math** and **clear data questions** when you want snappy demos.
- For PDFs: ask a **specific question** (“What is total equity?”) instead of only “analyze everything.”
- First message after startup will be slower — send a quick “hi” once before the real demo.

---

## Which Models Are Best for Reasoning?

There is a trade-off: **better reasoning usually means slower answers** (especially without a GPU).

### Recommended choices for this project

| Goal | Suggested Ollama model | Notes |
|------|------------------------|--------|
| **Best balance (current default)** | `phi3:mini` | Good speed on CPU; solid for demos and accounting Q&A |
| **Better reasoning / writing** | `llama3.1:8b` or `qwen2.5:7b` | Clearer, more careful answers; noticeably slower on CPU |
| **Stronger reasoning (if you have GPU)** | `llama3.1:8b`, `qwen2.5:14b`, or similar mid-size models | Better at multi-step PDF and finance explanations |
| **Maximum speed** | `llama3.2:1b` / `qwen2.5:1.5b` | Fast; more mistakes on hard accounting questions |
| **Avoid for local CPU demos** | 70B-class models | Too slow without serious hardware |

### Practical recommendation

- **Demos on a laptop / CPU:** keep **`phi3:mini`**.
- **Need smarter explanations and you can wait 5–20+ seconds:** try **`llama3.1:8b`** or **`qwen2.5:7b`**.
- **Need both speed and quality:** use a **GPU**, then a mid-size model (7B–8B) is usually the sweet spot.

Changing the model does **not** change the spreadsheet math or PDF totals — only how well the accountants chatbot **explains and answers in language**.

### How to switch the chat model

1. Install/pull in Ollama, e.g. `ollama pull llama3.1:8b`
2. Set `OLLAMA_MODEL=llama3.1:8b` in `.env`
3. Restart the chatbot server
4. Check the health page / admin readiness that the new model is listed

---

## Quality vs Speed Cheat Sheet

| If you want… | Do this |
|--------------|---------|
| Faster replies | Smaller model + lower `OLLAMA_NUM_PREDICT` + keep `RETRIEVAL_TOP_K=2` or `3` |
| Better reasoning | Larger model (7B–8B+) + GPU if possible |
| More accurate numbers from PDFs | Use clean digital PDFs; rely on calculated totals, not free-form AI guesses |
| Exact arithmetic | Type the expression as math; the accountants chatbot uses the calculator |
| Better dataset answers | Make sure the full CSV was indexed; ask specific questions (company, industry, period) |

---

## Day-to-Day Usage Tips

1. **New chat** when switching topics or PDFs — keeps context clean.
2. **Upload PDF + question together** for the clearest first answer.
3. Use **thumbs down + a short note** when an answer is wrong — useful for improving prompts later.
4. Admin panel reviews feedback and past queries (login required).
5. After changing `.env` or code, **restart** the server so settings apply.

---

## Quick FAQ

**Does the accountants chatbot always use AI?**  
No. Math and many filter-style data questions skip the language model or use it lightly.

**Which model “reasons”?**  
The Ollama chat model (default `phi3:mini`).

**Which model finds records in the dataset?**  
The embedding model `BAAI/bge-small-en-v1.5`.

**Why is the first answer slow?**  
Models load into memory on first use.

**Will a bigger model improve PDF totals?**  
No. Totals/ratios are calculated separately. A bigger model mainly improves explanations and Q&A wording.

**Can we use ChatGPT / cloud models instead?**  
This project is built for local Ollama. Cloud models would need a separate integration and would leave the “fully local” design.

---

## Summary

- **The accountants chatbot** answers accounting questions using a large combined dataset (about 4–5 source datasets), optional PDFs, a calculator, and a language model.
- **Reasoning model today:** `phi3:mini` via Ollama.
- **Search model:** `BAAI/bge-small-en-v1.5`.
- **To go faster:** smaller model, shorter max answer length, less retrieved context, GPU if available, warm start before demos.
- **For better reasoning:** move up to ~7B–8B models (ideally with a GPU); accept slower replies on CPU.

*Keep this guide updated when the default model or speed settings in `.env` change.*
