# Sergio Backend - FastAPI & Playwright Automation Server

The **Sergio Backend** is a FastAPI REST service that provides API endpoints for user authentication and dashboard operations, and manages a Playwright automation wrapper for interacting with the Hudl web portal.

---

## 1. Tech Stack

* **Python 3.10+ / FastAPI**: Async web framework.
* **Playwright**: Automates browser tasks on Hudl using a single persistent context.
* **SQLAlchemy ORM**: Database object-relational mapping.
* **PostgreSQL**: Database for users, roles, and OTPs.
* **SMTP (smtplib)**: Gmail SMTP relay for sending OTP codes.
* **Bcrypt & Python-Jose**: Hashing passwords and JWT token management.

---

## 2. Prerequisites

Before installing the backend, make sure you have the following installed on your local system:
1. **Python**: Python 3.10 or higher.
2. **PostgreSQL**: A running instance of PostgreSQL database server.
3. **SMTP Credentials**: Access to an SMTP mail server (like Gmail with an App Password) to send verification OTPs.
4. **Hudl Credentials**: A valid Hudl account email and password to test the automation flows.

---

## 3. Installation & Setup

Follow these steps to set up the backend locally:

### Step 1: Navigate to the Backend Directory
Open your terminal and enter the `backend/` directory:
```bash
cd backend
```

### Step 2: Create a Python Virtual Environment
Initialize a virtual environment to manage dependencies:
```bash
python3 -m venv venv
```

### Step 3: Activate the Virtual Environment
Activate the environment to isolate dependencies:
* **Linux/macOS**:
  ```bash
  source venv/bin/activate
  ```
* **Windows (Command Prompt)**:
  ```cmd
  venv\Scripts\activate
  ```
* **Windows (PowerShell)**:
  ```powershell
  .\venv\Scripts\activate
  ```

### Step 4: Install Python Dependencies
Install all required python packages specified in the requirements file:
```bash
pip install -r requirements.txt
```

### Step 5: Install Playwright Browsers
Install the Playwright browser binaries required for execution. We only need Chromium for this application:
```bash
playwright install chromium
```

### Step 6: Database Setup
1. Log in to your local PostgreSQL server.
2. Create a new database named `sergio`:
   ```sql
   CREATE DATABASE sergio;
   ```
   *(Note: Table creation and default roles seeding (`user` and `admin`) happen automatically when the FastAPI app starts).*

### Step 7: Environment Variables Setup
1. Copy the `.env_example` template file to a new file named `.env`:
   ```bash
   cp .env_example .env
   ```
2. Open `.env` and fill in the values corresponding to your local configuration:
   - Set database credentials (`DB_HOST`, `DB_PORT`, `DB_NAME`, `DB_USER`, `DB_PASSWORD`).
   - Set SMTP config (`SMTP_USER`, `SMTP_PASSWORD`, `SMTP_FROM`, etc.) for email dispatch.
   - Set `CREDENTIAL_ENCRYPTION_KEY` (required for account ingestion). Generate with:
     ```bash
     python -c "from cryptography.fernet import Fernet; print(Fernet.generate_key().decode())"
     ```
   - Optionally enter Hudl credentials (`HUDL_EMAIL`, `HUDL_PASSWORD`) as a fallback when no account has been ingested via the API.
   - Set `HEADLESS=False` to view the automated browser during local development/debugging.

---

## 4. How to Run the Server

With the virtual environment active and the database running, launch the backend application:

```bash
uvicorn app.main:app --reload --host 127.0.0.1 --port 8000
```

* **Interactive API Documentation**: Once running, open [http://127.0.0.1:8000/docs](http://127.0.0.1:8000/docs) in your browser to view the interactive Swagger UI and test APIs directly.
* **Auto-migration & Seeding**: The app will automatically build the tables (`users`, `roles`, `otp_verifications`, `accounts`) on startup and seed the roles database.

### Account Details Ingestion

Operators can submit Hudl credentials through the API instead of editing server `.env` files:

1. Authenticate via `POST /api/v1/auth/login` to obtain a JWT.
2. Submit account details to `POST /api/v1/accounts` with the Bearer token.
3. The controller validates the payload, performs a live Hudl login via Playwright, then encrypts and stores the credentials in PostgreSQL.
4. All Hudl automation (`POST /api/v1/hudl/login`, exchanges, uploads, library, etc.) reads credentials from the active stored account. If none is stored, it falls back to `HUDL_*` environment variables.

Example request body:

```json
{
  "account_id": "school-girls-bball",
  "name": "Example High School Girls Basketball",
  "credentials": {
    "hudl_email": "coach@school.edu",
    "hudl_password": "your-password",
    "hudl_login_url": "https://identity.hudl.com/u/login/identifier"
  },
  "metadata": {
    "team_id": "65072"
  }
}
```

Responses: `201 Created` on success, `400` if Hudl login fails, `409` if `account_id` already exists.

---

## 5. Backend Architecture & Key Details

### 📂 Directory Layout
```
backend/
├── app/
│   ├── automation/          # Playwright tasks & browser control
│   │   ├── tasks/           # Individual automation scripts (login, library, upload, etc.)
│   │   ├── browser_manager.py
│   │   ├── runner.py        # Task execution queue & locks
│   │   └── selectors.py     # XPath selectors for Hudl pages
│   ├── core/                # Database engine and config files
│   ├── models/              # SQLAlchemy database tables
│   ├── routes/              # FastAPI APIRouters (Auth, Accounts & Hudl endpoints)
│   ├── schemas/             # Pydantic models for validation
│   ├── services/            # Auth business logic
│   └── main.py              # Application entrypoint
├── browser_data/            # Persistent browser profile state (cookies, localstorage)
├── requirements.txt         # Dependencies list
└── .env                     # Local environment settings
```

### 🧠 Core Automation Systems

#### 1. Browser Lifecycle Manager (`browser_manager.py`)
To minimize server resource consumption, the `BrowserManager` employs several optimization strategies:
* **Persistent Session Context**: Cookies and session caches are persisted to `backend/browser_data/`. This avoids having to authenticate with Hudl on every execution.
* **Resource Optimization**: Blocks the loading of heavy resources (images, media, web fonts) and tracking scripts (Google Analytics, Mixpanel, Sentry) to maximize navigation speeds.
* **Idle Timeout**: Automatically shuts down the Chromium browser after **5 minutes (300 seconds)** of inactivity.
* **Daily Refresh**: Clears all profile cookies and starts a clean browser context every 24 hours to prevent memory leaks and handle stale login states.

#### 2. Serialized Task execution (`runner.py`)
To prevent concurrent requests from interfering with each other (e.g. multiple users trying to click different buttons on the same page), the backend uses a sequential queue:
* **Global Async Lock**: All routes coordinate through `run_task()`, which uses an `asyncio.Lock()` to process automation requests one at a time.
* **Busy Flag**: Keeps the idle timeout checker from closing the browser while a transaction is actively being performed.

#### 3. HLS Stream Proxy (`hudl_routes.py`)
Hudl streams library videos using HLS (HTTP Live Streaming). These streams contain `.m3u8` playlists and encrypted `.ts` segments protected by Hudl cookies. Standard HTML5 players cannot access these segments.
* **Proxying**: The `/api/v1/hudl/library/stream-proxy` endpoint intercepts client video requests.
* **Cookie Injection**: It attaches active Hudl session cookies collected from the Playwright browser.
* **Manifest Rewriting**: It parses the `.m3u8` manifest, converts all relative segment and key paths to absolute URLs, and rewrites them so that the client player fetches all components back through the Sergio stream proxy with a valid JWT token.
