--- title: Champ Chatbot Demo emoji: 🤖 colorFrom: yellow colorTo: gray sdk: docker pinned: false app_port: 8000 python_version: "3.11" --- # MARVIN WebUI Demo A lightweight chat interface powered by the MARVIN model, designed for easy deployment and testing. This project can run locally with Docker Compose or be deployed to platforms like Hugging Face Spaces. ## Features - Simple and fast chat interface - Python backend with async endpoints - Frontend served through a minimal web app - Containerized so can start everything with a single command ## Requirements - Docker - Docker Compose ## Architecture The chat path is four layers, one-way dependencies (top → bottom): ``` HTTP route (main.py) ↓ SessionRouter (classes/session_router.py) — registry of (session, conversation) → AgentClient ↓ AgentClient (agent/agent_client.py + subclasses) — one per conversation; owns ConversationHistory ↓ Agent (agent/agent.py) — chat loop, tool dispatch via SkillsManager ↓ ChatProvider (providers/) — wire-level adapter per backend SDK ``` **`model_type`s supported** by `SessionRouter._create_agent_client`: | `model_type` | AgentClient | Provider | Notes | |---|---|---|---| | `skills_wiki` | `SkillsAgentClient` | `HFChatProvider` (Groq) | gpt-oss-20b with skills + tool use; language-leak correction. Pediatric answers grounded on the built wiki (`pediatry_wiki` skill). This is the deployed pediatric agent. (The manual-chunks `pediatry` skill / `skills` route exists for local A/B eval only and is not deployed.) | | `champ` | `ChampAgentClient` | `HFChatProvider` (Groq) | gpt-oss-20b with FAISS RAG + safety triage | | `openai` | `GptAgentClient` | `OpenAIProvider` | single-shot chat completion | | `google-conservative` | `GeminiAgentClient` | `GeminiProvider` | temperature 0.2 | | `google-creative` | `GeminiAgentClient` | `GeminiProvider` | temperature 1.0 | | `fake` | `FakeAgentClient` | `FakeProvider` | echo, for tests/demos — no API key needed | Provider skeletons exist; only HF and Gemini/OpenAI are implemented. Adding a new backend means writing a `providers/.py` implementing the `ChatProvider` protocol from `providers/protocol.py`, plus a branch in `SessionRouter._create_agent_client`. ## Local Development ### Environment variables The app uses several environment variables. Create a `.env` file at the project root: ``` # Local DynamoDB (docker-compose.dev.yml exposes it on port 3000) DYNAMODB_ENDPOINT=http://localhost:3000 ENV=dev # Model providers — required for the model_types you actually use HF_TOKEN= # HuggingFace, used by the skills + champ paths (via Groq) HF_ACCESS_TOKEN= # used by the pediatry / pediatry_wiki skill subagents (generate_grounded_response, generate_wiki_response) OPENAI_API_KEY= # required for model_type="openai" GEMINI_API_KEY= # required for model_type="google-conservative" / "google-creative" ``` Production (HF Spaces) additionally needs `AWS_ACCESS_KEY`, `AWS_SECRET_ACCESS_KEY`, `AWS_REGION`, `DDB_TABLE`, `DDB_ENVIRONMENT_IMPACT_TABLE`. ### Start the database service To run the local DynamoDB service: To run the database service: ``` docker-compose -f docker-compose.dev.yml up -d ``` ### Start the backend and frontend service From the project root: ``` docker compose up --build ``` Once everything is ready, open: ``` http://localhost:8000 ``` ### Stopping the project Use: ``` docker compose down ``` ### Rebuilding after code changes Use: ``` docker compose up --build ``` ### Running without Docker Before installing the dependencies, install `uv`: ``` pip install uv ``` `uv` is a python package manager similar to `pip`. However, it permits overriding package version conflicts. This allows installing packages that *theorically* incomptatible but are necessary to run the app. After installing `uv`, create your virtual environment, then run: ``` uv pip install --no-cache-dir -r requirements.txt ``` #### Installation problems with Windows ##### libmagic (Failed to find libmagic) When running the app with uvicorn for the first time on Windows, you might get the error `Failed to find libmagic`. Do these steps to fix the issue: 1. Go [here](https://pypi.org/project/python-magic-bin/0.4.14/#files), then download the wheel that matches the number of bits of your CPU - For 64 bits: python_magic_bin-0.4.14-py2.py3-none-win_amd64.whl - For 32 bits: python_magic_bin-0.4.14-py2.py3-none-win32.whl 2. Then run `pip install python_magic_bin-0.4.14-py2.py3-none-win_amd64.whl --force-reinstall` or `pip install python_magic_bin-0.4.14-py2.py3-none-win32.whl --force-reinstall` depending on the downloaded wheel #### Installation problems with Mac (Apple Silicon) ##### libmagic Installing `libmagic` on Mac is often problematic. If it fails, do these steps: 1. Run `brew install libmagic` 2. Copy the magic directory from [this repository](https://github.com/SHi-ON/libmagic-apple-silicon) to the directory where your Python environment libraries are located. Run `$ pip list -v` to be able to locate the path to your libraries directory. As an explanation on the origin of the magic directory, it has been derived from an Intel-based Mac with python-magic installed via pip. 3. Copy ``libmagic.1.dylib`` from the lib directory in the libmagic that Homebrew has installed to the ``magic/libmagic`` directory in Step 2 to replace the ``YOUR_libmagic.dylib``. Please note that you need to copy the original file, not the alias (symbolic link). Run ``$ brew list -v`` to help you locate the path to the library installed by Homebrew. A typical path looks like ``/usr/local/Cellar/libmagic/5.44/lib``. 4. Rename the copied file `libmagic.1.dylib` to `libmagic.dylib` ##### SSL: CERTIFICATE_VERIFY_FAILED If this happens when trying to run the app, in `Applications/Python3.11`, execute the file `Install Certificates.command`. --- ## Deployment on HuggingFace Spaces A HuggingFace Space uses Git to store the code. It requires a YAML block at the beginning of the `README.md` file that defines metadata variables. Notably, this file specifies the app title (`Champ Chatbot Demo`), the SDK (`docker`), and the app port (`8000`). Since `docker` is the selected SDK, the space builds the app using the `Dockerfile` at the project root. ### Update code To update the code in the space, click on `+ Contribute` button in the upper-right corner of the Files page, then click `Upload Files`. Follow the instructions to commit/upload your local files to the space. You could add the Git repo as a remote to your local Git repository, but it would add unnecessary complexity. HuggingFace is stricter than Gitlab concerning best Git practices. You would have to configure `git-xet` and delete the `.env` file and the binary file in `rag_data` from the Git history to be able to push your changes. The `.env` file has not been added to the space. The environment variables are stored in the settings page. ## Testing The repo has four distinct test surfaces: - **Unit tests** under `tests/` — fast, run with `pytest`. - **Scenario evaluation** (`experiments/behavior_eval/run_scenarios.py`) — drives the agent through scenarios and asks a judge LLM to verdict each criterion. - **Groundedness / language evaluation** (`experiments/groundedness_eval/run_groundedness_eval.py`) — per-question check that the agent stays grounded in the reference material, and that French questions don't leak English (and vice-versa). - **Judge evaluation** (`experiments/judge_eval/evaluate_judge.py`) — measures the judge model's own quality against human-labeled gold answers (e.g. Cohen's kappa). `pytest` is reserved for unit tests. Everything in `experiments/` is run as a standalone Python script. ### Unit tests Install dev deps once, then run: ```bash pip install -r requirements-dev.txt pytest ``` Some tests are marked as `resource_intensive` and skipped by default. To include them: ```bash pytest -m resource_intensive ``` To run every test (including skipped): ```bash pytest -m "" ``` ### Code coverage `coverage` is a Python library that measures code coverage. To use it, run: ```bash coverage run -m pytest ``` To see a short summary of the results, run: ```bash coverage report ``` For a more detailed presentation, run: ```bash coverage html ``` To run `pytest` with additionnal arguments, you can run, for example: ```bash coverage run -m pytest -m resource_intensive ``` ## Scenario evaluation Scenario evaluation runs the agent against (scenario, criterion) pairs and asks one or more judge LLMs to verdict each criterion. Driven by `experiments/behavior_eval/run_scenarios.py`. ```bash python -m experiments.behavior_eval.run_scenarios [--skill NAME] [--judge LABELS] \ [--wandb-project PROJECT] [--wandb-experiment NAME] ``` - `--skill NAME` — restrict to one skill's scenarios (e.g. `--skill pediatry`). - `--judge LABELS` — comma-separated judge labels (`oss`, `qwen`). Default: all. - `--wandb-experiment NAME` — enable W&B logging under this run name. Omit to disable W&B. - `--wandb-project NAME` — defaults to `marvin-agent`. Exits 0 only if every criterion verdict is `PASS`. `FAIL` / `NOT_TRIGGERED` / agent-crash all cause a non-zero exit. Each criterion line is logged with its full test id. Scenarios live in `experiments/behavior_eval/definitions//` as `human.json` (and optionally `generated.json`). Criteria live alongside as `criteria.json`. Failed JSON outputs and agent-crash transcripts are dumped to `evaluation_results/`. ## Groundedness / language evaluation Per-question evaluation of the agent: did the answer stay grounded in the reference material, and did the response language match the question language? Driven by `experiments/groundedness_eval/run_groundedness_eval.py`. ```bash python -m experiments.groundedness_eval.run_groundedness_eval [--questions-set {easy,hard}] \ [--split {val,test,all}] [--skill {pediatry,pediatry_wiki}] \ [--checks {both,groundedness,language}] [--limit N] \ [--out REPORT.md] [--provider NAME] [--no-language-directive] \ [--no-save-transcripts] \ [--wandb-project PROJECT] [--wandb-experiment NAME] ``` - `--questions-set` — `easy` (default) runs the templated single-illness set and honours `--split`; `hard` runs the hand-authored adversarial set (English-only, `--split` ignored). Both, with their per-skill gold references and language, live in `experiments/groundedness_eval/illness_questions.py`. - `--split` — `val` (tuning), `test` (held-out), or `all` (default). Only applies to `--questions-set easy`. Membership is the per-question `split` field in `experiments/groundedness_eval/illness_questions.py` (by illness, so no illness leaks between tuning and held-out). - `--skill` — `pediatry` (default, manual chunks) or `pediatry_wiki` (Karpathy-style wiki). Selects which pediatric retrieval surface the agent uses. Each skill is judged against its own gold namespace (manual chunks vs. built wiki pages), defined per question in `illness_questions.py`. - `--checks` — `groundedness` (judge only), `language` (French-leakage only), or `both` (default). - `--limit` — first N questions only. - `--no-language-directive` — skip telling the agent which language to use (measures baseline leakage). - `--dry-mapping` — print `(question → reference file)` pairs and exit; no LLM calls. A/B comparison example — same questions, both retrieval surfaces: ```bash python -m experiments.groundedness_eval.run_groundedness_eval --questions-set hard --skill pediatry --wandb-experiment manual_chunks python -m experiments.groundedness_eval.run_groundedness_eval --questions-set hard --skill pediatry_wiki --wandb-experiment wiki ``` A markdown report lands in `reports/` by default. Per-question transcripts go to `evaluation_results/` unless `--no-save-transcripts` is set. ## Judge evaluation Measures the judge model itself against gold-labeled answers — Cohen's kappa, precision/recall on each verdict, etc. Driven by `experiments/judge_eval/evaluate_judge.py`. ```bash python -m experiments.judge_eval.evaluate_judge [--split {val,test,all}] \ [--gold-dir DIR] [--val-topics PATH] [--judge-model ID] [--provider NAME] \ [--out REPORT.md] [--wandb-project PROJECT] [--wandb-experiment NAME] ``` - `--split` — `val` (default), `test`, or `all`. Splits gold cases by illness (`topic_file`): a case is val if any of its topic files is listed in `--val-topics`, test otherwise. - `--val-topics` — file listing the val-split `topic_file` names, one per line (`#` comments allowed). Defaults to `experiments/groundedness_eval/val_topics.txt`. - `--judge-model` — judge model id (pair with `--provider`); defaults to the standard agent model. - `--gold-dir` — gold-case directory (default `experiments/judge_eval/cases_gpt_original/`). - `--out` — markdown report path; defaults to `reports/judge_eval_{split}_{ts}.md`. Gold labels live in `experiments/judge_eval/cases_gpt_original/` (GPT-era) and `experiments/judge_eval/cases_gemma/` (Gemma-era) — pick which via `--set`. Reports go to `experiments/judge_eval/reports/`. Use this when you change the judge prompt or swap judge models — it tells you whether the new judge agrees with humans more or less than the old one. ## Load testing [k6](https://k6.io/open-source/) is an open-source tool for performing load testing. Test cases are defined in JavaScript files and can be run using the command `k6 run .js`. ### k6 installation On Debian/Ubuntu: ``` sudo apt-get update sudo apt-get install k6 ``` On Windows: ``` winget install k6 --source winget ``` On Docker: ``` docker pull grafana/k6 ``` For more options, see [Install k6](https://grafana.com/docs/k6/latest/set-up/install-k6/). ### Test scenarios The test cases are defined in the folder `/tests/stress_tests/`: - `chat_session.js` simulates 150 users sending three messages to one specific model. - `file_upload.js` simulates 150 users sending three PDF files. - `chat_session_with_file.js` simulates 150 users sending one PDF file followed by three messages to one specific model. - `website_spike.js` simulates 150 users connecting to the application home web page. #### Chat session test scenario The chat session scenario must be run by specifying the model type and the URL of the server. For example, the following command simulates 150 users making three requests at `https://-champ-chatbot.hf.space` to the model `champ`: ``` k6 run chat_session.js -e MODEL_TYPE=champ -e URL=https://-champ-chatbot.hf.space ``` The possible values for `MODEL_TYPE` are `skills_wiki`, `champ`, `openai`, `google-conservative`, `google-creative`, and `fake`. To find your HuggingFace Space backend URL, follow these steps: 1. Go to your space 2. Click on the **three dots** in the top right corner 3. Select **Embed this Space** 4. Look for the **Direct URL** in the code snippet. Typically, the URL follows this format: `https://-.hf.space`. To test locally, simply use `http://localhost:8000` The file `message_examples.txt` contains 450 pediatric medical prompts (generated by Gemini and Sonnet). `chat_session.js` uses this file to simulate real user messages. #### File upload test scenario The file upload scenario must be run by specifying the file to send and the URL of the server. Each virtual user will upload the file 3 times to the server. ``` k6 run file_uploads.js -e FILE=my_pdf_file.pdf -e URL=https://-champ-chatbot.hf.space ``` Make sure the file is at the same directory level as the test file. #### Chat with file test scenario The file upload scenario must be run by specifying the PDF file, the model type and the URL of the server. Each virtual user will upload the file once then send three messages to the server. ``` k6 run chat_session_with_file.js -e FILE=my_pdf_file.pdf -e MODEL_TYPE=champ -e URL=https://-champ-chatbot.hf.space ``` The possible values for `MODEL_TYPE` are `skills_wiki`, `champ`, `openai`, `google-conservative`, `google-creative`, and `fake`. Make sure the file is at the same directory level as the test file. #### Website spike test scenario The website spike scenario must be run by specifying the website URL which is simply the HuggingFace Space URL: ``` k6 run website_spike.js -e URL=https://huggingface.co/spaces//champ-chatbot ```