# Infrastructure Plan — CHAMP Pediatric Chatbot Status: **draft for review**. Written 2026-08-05. Nothing here is implemented. Scope: production deployment of the pediatric chatbot web app — UI hosting, Firebase authentication, persistence of accounts / conversations / environmental impact, and a two-backend inference plane (self-hosted Gemma 4 26B-E4B during high traffic, Vertex AI during low traffic). --- ## 1. Starting constraints Given: 1. The web UI must be deployed. 2. Firebase Authentication for identity. 3. Persist user accounts, their conversations, and their environmental impact. 4. High traffic → a self-hosted instance serving Gemma 4 26B-E4B on an A10G GPU, used by all four LLM roles (main agent, wiki subagent, judge, attribution check). 5. Low traffic → the same model behind a serverless API. Vertex AI (GCP) is the chosen endpoint, because Bedrock does not host Gemma in a Canadian region. 6. The backend server must also be hosted. Established during planning: - **Data residency: Canada.** Confirmed workable for every component including Firebase Auth, so the whole topology sits in Montreal (`ca-central-1` / `northamerica-northeast1`). - **Data classification (PHI vs. Law 25 personal data vs. de-identified) is still undetermined** — the full policy is not yet available. It gates retention, audit, and vendor-agreement requirements rather than the topology. - **Traffic volume and concurrency are unknown** and will stay unknown until there are real users. The plan is therefore sequenced to *measure before sizing* (§8). - **Vertex bills per token with no idle cost**, and the established crossover against a self-hosted A10G is **~78 requests/hour**. This is what justifies the two-backend design. - Judge rounds per turn are expected to drop below 2 with recent pipeline improvements, which lowers per-turn token cost and raises that crossover (§7). --- ## 2. What exists today | Piece | Today | Production-ready? | |---|---|---| | Web UI | Expo Router app in `client/`, web target, `EXPO_PUBLIC_BACKEND_URL`. Legacy Jinja templates + `static/` also served from `/` by FastAPI | Needs a host; needs the two UIs disentangled | | Backend | FastAPI monolith, Docker, deployed to **HF Spaces** (`README.md` front-matter, `app_port: 8000`) | No | | Auth | **None.** `user_id` / `participant_id` are client-supplied fields on `ChatRequest` | No | | Session state | In-RAM: `SessionRouter`, `SessionTracker`, `SessionDocumentStore` (up to 30 MB uploads/session) | **Blocks horizontal scaling** | | Conversations | DynamoDB single-table, `PK=SESSION#`, `SK=TS##`, `GSI1=USER#` | Schema is fine; region and access path are not | | Environmental impact | DynamoDB `environmental-impact`, `PK="SERVER#HF-Space-01"` (hardcoded), ecologits token estimates + codecarbon infra accumulation | Schema needs work (§6.3) | | Inference | HF `InferenceClient` → Groq, gpt-oss-20b, elaborate 429 backoff in `providers/hf.py` | Being replaced by the two-backend plane | | Tables | Created at *import time* in `helpers/dynamodb_helper.py` | Must move to IaC | | Telemetry | `telemetry.py` returns early unless `ENV=dev` — **tracing is off in production** | Backwards; fix | | Rate limiting | `slowapi`, in-process, keyed by client IP | Breaks across replicas | --- ## 3. Topology decision The tension: Firebase (auth) and Vertex (serverless inference) are GCP-only; the team prefers AWS; DynamoDB is already on AWS. One fact narrows this: **the A10G is effectively an AWS-only SKU** (the `g5` family — NVIDIA built it for AWS). GCP's nearest 24 GB equivalent is the **L4** (`g2` family), and the L4 is **materially worse for this workload**: ~300 GB/s of memory bandwidth against the A10G's ~600 GB/s. Token generation is memory-bound, so decode on an L4 runs at roughly half the A10G's rate — and this pipeline's latency is dominated by judge *output* tokens, i.e. decode. The L4's advantages (native FP8, 72 W vs 150 W) do not compensate: 26B at FP8 is ~26 GB and does not fit in 24 GB anyway, so INT4 is required on either card, and the power saving mostly cancels against the longer runtime. See §6.4. Getting A10G-class bandwidth on GCP means jumping to an A100, which is a large cost step. **So the GPU node wants to be on AWS.** ### Option A — All-GCP, `northamerica-northeast1` (Montreal) — **recommended** Firebase Hosting + Firebase Auth + Cloud Run (backend) + Firestore + Vertex AI + GCE L4 for the self-hosted node. - One cloud → one IAM model, one audit trail, one vendor in the privacy impact assessment, one residency argument to make. - The inference hot path never leaves the provider boundary. - Cost: migrating the DynamoDB tables and `helpers/dynamodb_helper.py`, plus the analysis notebook (`analysis/chat_log/dynamodb_chat_log_analysis.ipynb`), and going against the team's AWS preference. ### Option B — AWS `ca-central-1` for compute + data, GCP for auth and serverless inference ECS Fargate or App Runner + DynamoDB `ca-central-1` + `g5.xlarge` (A10G), with Firebase Auth used only as a token issuer and Vertex called over the internet. - **Technically fine.** Both regions are physically in Montreal, so cross-cloud latency is single-digit milliseconds against multi-second inference, and egress volume is small (~200 KB/call × a few calls per turn ⇒ cents per month). The usual "don't split clouds" performance argument does not really apply here. - The real cost is **governance, not engineering**: two vendor assessments, two sets of credentials and rotation policies, two audit trails, and a documented cross-provider data flow carrying user text — for a workload whose classification is still undetermined. - Firebase Auth on its own does *not* force GCP; verifying a Firebase ID token server-side needs only Google's public JWKS. It is Vertex that pulls the hot path across. ### Option C — Split by plane (UI+identity GCP, data+GPU AWS) Most seams, no benefit over A or B. Not recommended. ### Recommendation **Option B.** The GPU node needs A10G-class memory bandwidth (above), which is AWS-only at this price point; the team already prefers AWS; and DynamoDB is already there. Firebase Auth does not pull the backend onto GCP — verifying an ID token needs only Google's public JWKS — and while Vertex does put the low-traffic inference path on GCP, both regions are physically in Montreal, so the cross-provider hop costs milliseconds and negligible egress. The remaining argument for all-GCP is **governance**: one vendor instead of two in the privacy impact assessment, one audit trail, one residency argument. That is a real cost, but it does not outweigh halving decode throughput on a latency-critical clinical pipeline. Budget for the second vendor review and document the AWS→Vertex flow as a disclosure rather than an internal call. *(Earlier drafts of this document recommended Option A on the mistaken premise that the L4 was the better card. Corrected 2026-08-05.)* The rest of this document names AWS services first, with GCP equivalents inline. --- ## 4. Target architecture Option B: compute and data on AWS `ca-central-1`, identity and serverless inference on GCP. Both regions are in Montreal. ```mermaid flowchart TB subgraph client["Browser"] UI["Expo web bundle"] end subgraph gcp["GCP — northamerica-northeast1"] AUTH["Firebase Auth
ID tokens, custom claims"] HOST["Firebase Hosting
static bundle, CDN"] VERTEX["Vertex AI endpoint
Gemma 4 26B-E4B
per-token billing"] end subgraph aws["AWS — ca-central-1"] API["ECS Fargate / App Runner
FastAPI, scale 1..N"] REDIS["ElastiCache Redis
sessions, rate limits"] S3["S3
spilled step payloads"] DB[("DynamoDB
users / conversations / impact")] SM["Secrets Manager"] VLLM["EC2 g5 (A10G) + vLLM
same model, private VPC
scaled 0..1..N"] end ROUTER{{"ModelRouter
in backend"}} UI -->|"static"| HOST UI -->|"sign in"| AUTH UI -->|"Bearer ID token"| API API -->|"verify token via JWKS"| AUTH API --> REDIS API --> S3 API --> DB API --> SM API --> ROUTER ROUTER -->|"above ~78 req/hr"| VLLM ROUTER -->|"below threshold / failover"| VERTEX ``` The only steady-state cross-provider flows are JWKS token verification (no user data) and the Vertex inference calls (user text — document this one in the privacy assessment). --- ## 5. Prerequisites in the existing codebase These are blockers. None of them are about the new infrastructure — they are things the current code does that prevent it from being deployed behind an autoscaler at all. **P1 — Externalize session state.** `SessionRouter` holds `(session, conversation) → AgentClient` (each owning its `ConversationHistory`) in a process-local dict; `SessionTracker` and `SessionDocumentStore` likewise. With more than one replica, a user's second message can land on a replica that has never seen their first. `main.py` already flags this in a comment. Fix: conversation history rehydrated from the database (or Redis) per request. Uploaded documents stop being session state entirely — under §6.3.9 an attachment is committed with the message that carries it, so there is nothing to hold between requests and `SessionDocumentStore` goes away. **P2 — Authentication.** `user_id` and `participant_id` arrive in the request body and are never verified — today anyone can write chat events under anyone's ID. Derive the identity from the verified Firebase token instead, including for guests (anonymous auth still yields a verifiable token). `participant_id` is being retired into `user_id` — see §6.3.2. **P3 — `model_type` is client-chosen.** `ChatRequest.model_type` lets the caller pick `openai`, `google-creative`, `champ`, etc. In production that is both a cost exposure (a scripted client can bill your OpenAI key) and a safety one (a health question answered by a non-grounded path that skips the judge). Restrict to a server-side allow-list; gate the research model types behind an admin/researcher custom claim. **P4 — Table creation at import.** `helpers/dynamodb_helper.py` calls `create_*_table_if_not_exists` at module import, so every cold start issues `list_tables` / `create_table`. Move schema to IaC (Terraform); the app should only ever read/write. **P5 — Rate limiting is per-process and per-IP.** `slowapi` with `get_remote_address` under N replicas means the effective limit is N× the configured one, and IP is the wrong key once there are accounts. Move to a Redis-backed limiter keyed by `user_id`. **P6 — Turn telemetry on in production.** `setup_telemetry` returns immediately unless `ENV=dev`. Export OTLP to Cloud Trace (or your APM) in staging and prod; a distributed judge loop is exactly the thing you cannot debug from logs. **P7 — Deletion path.** There is no way to delete a user's data. Under Law 25 (and any PHI reading) you need erasure spanning the conversation table, the impact table, uploaded documents, and any derived analysis exports. Design the keys for it now — see §6.3. **P8 — PII redaction is not wired up.** `classes/pii_filter.py` defines `PIIFilter` and has unit tests, but **it has no call site anywhere outside its own test file**. Nothing on the chat path, the logging path, or the upload path calls it. Everything §6.3.5 specifies therefore describes a system that does not exist yet, and the gap is widest on the document path: uploaded text is rendered straight into the system prompt by `agent/documents.py`, so a photographed vaccination record or discharge summary reaches the inference provider verbatim — and OCR'd clinical paperwork is far denser in identifiers than anything a parent types. Two channels need covering, not one: `messages` and the uploaded-document block. --- ## 6. Component design ### 6.1 Web UI Static export (`npx expo export -p web`) → **Firebase Hosting**. It is a CDN-fronted static bundle; the artifacts contain no personal data, so residency does not bind here. Firebase Hosting also gives you the auth-domain and preview channels for staging. Decide what happens to the legacy Jinja UI in `templates/` + `static/`. Recommendation: stop serving it from the production backend (it predates the Expo client and duplicates it); keep it only in the dev image if it is still useful for internal demos. `EXPO_PUBLIC_BACKEND_URL` becomes a build-time variable per environment. AWS equivalent: S3 + CloudFront. ### 6.2 Authentication - Client signs in with Firebase Auth, sends `Authorization: Bearer `. - Backend verifies via `firebase-admin` (`verify_id_token`) with cached JWKS. A FastAPI dependency resolves `firebase_uid` → internal opaque `user_id`; every handler takes the authenticated identity, never the payload field. - **Custom claims** carry whatever flag the P3 model-type gate ends up needing. Out of scope for the data model — the schema in §6.3 does not model access control. - **Never write conversation content or health data into Firebase.** Firebase holds credentials and claims only. The `firebase_uid` → `user_id` mapping lives in the Canadian database. - ✅ **Residency: resolved** (2026-08-05). Firebase Auth data can be held in Canada, so Identity Platform is usable under a Canada-only rule. Keep the "nothing but credentials and claims in Firebase" discipline anyway — it keeps the erasure path (P7) in one place. ### 6.3 Data layer Decisions recorded from the data-modelling discussion of 2026-08-05. A concrete schema sketch lives in [`schema.sql`](schema.sql). It is **documentation, not a migration** — it has never been executed against a PostgreSQL server, so treat it as a statement of the model rather than as runnable DDL. Its header lists the choices it had to make that the prose below leaves open. #### 6.3.1 Store choice — PostgreSQL recommended (not yet decided) **Recommendation: one PostgreSQL instance (RDS, `ca-central-1`) for everything**, replacing DynamoDB. Reasons, in order of weight: 1. **The data is relational now.** Versioned consent history with referential integrity, transactions (record consent + write its audit row atomically), temporal demographics, a many-to-many on occupation, and research queries that join across all of it. The existing code already resorts to full-table `scan`s for analytics — the tell that the access patterns outgrew the store. 2. **DynamoDB's 400 KB item limit is a real cliff.** Today an answer's whole pipeline record (tool calls, judge verdicts, reasoning) rides in one item, and judge reasoning alone is ~84% of judge output by the project's own measurements. A long multi-round answer can approach the limit, and the write simply fails when it does. Decomposing into per-step rows (§6.3.6b) removes the cliff on its own, and `JSONB` has no practical ceiling for what remains; oversized payloads spill to object storage with a pointer. 3. **Volume does not justify NoSQL.** The self-hosting crossover is 78 requests/hour. 4. **One store to back up, restore, and erase from** — which matters a great deal for P7. Postgres specifically over MySQL for `JSONB`: indexable, queryable JSON lets traces live beside relational data instead of forcing a second store. *Defensible alternative if the conversation pipeline shouldn't be disturbed:* Postgres for accounts and consent, DynamoDB retained for transcripts. That split follows a real line (OLTP vs. append-only event log) rather than an arbitrary one. Migration cost either way touches `helpers/dynamodb_helper.py`, `analysis/chat_log/`, and the current `AWS_REGION` default of `us-east-1`. #### 6.3.2 The identity boundary Direct identifiers (email, phone, password) live **only in Firebase Auth**. Everything else is keyed on an opaque internal `user_id`. This keeps direct identifiers out of the conversation store. It does **not** make transcripts anonymous — see §6.3.5. - `users` — `user_id` (UUID, PK), `created_at`, `locale`, `account_type` (`registered` | `guest`). **No lifecycle status column** (decided 2026-08-06): account lifecycle is not specified, so row absence means deleted. States such as `suspended` get added when a requirement for them appears, not on speculation. - `auth_identities` — one user may hold several bindings: Firebase (including *anonymous* uids for guests) and, if WhatsApp ships, a phone number. `(provider, provider_uid)` unique; FK to `users`. **`user_id` must not require a `firebase_uid`.** - **`participant_id` is retired** and merged into `user_id` (decided 2026-08-05). It currently keys the environmental-impact records and `analysis/metadata/event_metadata_helper.py`, so the merge is a migration touchpoint. If a research dataset is later published, generate the study identifier **at export time** rather than reintroducing a stored column — that keeps identifiers unjoinable across releases. - **Username: undecided.** If it becomes a login credential, Firebase cannot help — it supports email/password, phone, and OAuth, not usernames. That requires a server-side username→email resolve step (returning a uniform error either way, or it is a user-enumeration oracle) and a **unique constraint in Postgres**, since Firebase does not enforce uniqueness on `displayName`. Cheap hedge: reserve the column with a unique constraint now; retrofitting uniqueness after accounts exist means resolving collisions by hand. Real name is a separate field and direct PII — do not collect it without a concrete use. #### 6.3.3 Consent — an append-only event log Not three date columns. Those cannot express grant → withdraw → re-grant, and cannot record *what* was consented to. ``` consent_events (consent_event_id, user_id, purpose, action, notice_version, occurred_at, source) ``` **Purposes (decided 2026-08-06, Malik — consent handled outside this document):** | Code | Covers | |---|---| | `service_and_storage` | Using the service, and storing conversations so users can see their history when they log back in | | `research_improvement` | Using conversations to improve the system and the model, for research | More purposes are expected. `purpose` is therefore a **lookup table, not an ENUM** — adding a row must not require a migration, and retiring one must not break historical rows. - `action` ∈ { granted, withdrawn }. Append-only: never updated, never deleted. - `notice_version` is required — revising the privacy notice does not carry prior consent forward. - Current state is a view over the log, not a stored column. - There is an existing `user_consents/` directory and a `consent: bool` on `ChatRequest`; both fold into this. #### 6.3.4 Demographics — temporal, not snapshotted `ProfileBase` today carries `age_group` (adult brackets `18-24` … `65+`), `gender`, and `roles` (a **set** of `patient | clinician | computer-scientist | researcher | other`). None of them reach the model — they are logged and consumed only by `analysis/metadata/event_metadata_helper.py` for cohort distributions. They are research demographics about the user, nothing more. `roles` is the "occupation" field, and being multi-valued it is a junction table, not a column. Three ways to answer "what was this user's occupation when this message was sent": | | Duplication | Correct history | Verdict | |---|---|---|---| | Join to current profile | none | ✗ retroactively rewrites past messages | reject | | Snapshot onto each message | heavy | ✓ | rejected — see below | | **Temporal history table** | none | ✓ | **chosen** | ``` user_demographics_history (user_id, valid_from, valid_to, age_group, gender) user_roles_history (user_id, role, valid_from, valid_to) ``` Research queries join on `message.created_at BETWEEN valid_from AND valid_to`; hide it behind a view. This is not a normalization compromise — *occupation-at-time-of-message* depends on `(user_id, time)`, not on `user_id` alone, so the temporal table **is** the normalized form. Snapshotting was rejected on a second count: these attributes are consent-conditional, so withdrawal must erase them. That is a few row deletes against a history table versus updating every message the user ever sent. The current DynamoDB logging copies `age_group`/`gender`/ `roles` onto every event — the migration is the moment to stop. Exception, for later: a value that was an actual *input to inference* belongs on the message, because it helped produce the answer. None of these currently qualify. #### 6.3.5 Transcripts and the PII vault A teammate owns the PII pipeline: identifiers are replaced before storage and restored on output. **The vault is durable** (decided 2026-08-05) — users must see their own history as they typed it, not placeholders, however long afterwards. Consequences that follow from durability: - **Tokenized *and* encrypted at rest.** Tokenization's coverage is bounded by detector recall, and recall will not be 100% on bilingual free text with indirect identifiers ("my son at St-Joseph in Verdun"). A missed identifier sits in plaintext in the supposedly de-identified transcript; encryption at rest is what covers that gap. The research corpus is therefore **sensitive, not anonymous**. - **The vault lives in its own schema under a distinct KMS key**, separate from the transcript store. - **The substitute is a plausible value, not a placeholder** — `Paris → Marseille`, chosen by the redactor so the surrounding text still reads naturally to the model. The redaction policy (which entities are replaced at all, and with what) is the teammate's; the schema's job is to make the swap reversible and the reversal unambiguous. - Vault rows store ciphertext, not plaintext: `(conversation_id, surrogate, value_hmac, entity_type, ciphertext, created_at)`. - **Constrained in both directions, because the mapping has two callers.** Restoration looks up `surrogate → real value`; redaction looks up `real value → surrogate` to stay consistent across turns. The primary key `(conversation_id, surrogate)` covers the first, a unique constraint on `(conversation_id, value_hmac)` covers the second, and together they keep the mapping bijective inside a conversation. - **`value_hmac` is an index, not a second copy of the secret.** It is `HMAC(per-conversation salt, normalized value)`, kept because ciphertext cannot be constrained or searched — encryption is randomized, so the same value encrypts to different bytes each time. The hash is deterministic, so it can carry the unique constraint and answer "have I already substituted this value here?" with no decryption on the write path. Salting per conversation keeps it non-comparable across conversations. - **One data key per conversation** (`pii.conversation_keys`), generated locally, used for every row in that conversation, and stored only in wrapped form under the KMS key. Reading a conversation is one KMS call to unwrap plus local decryption of each row, rather than one KMS call per value. This also makes the audit trail legible — CloudTrail records one key use per conversation opened instead of one per identifier — though only if the unwrapped key is not cached across requests. That caching decision is worth making deliberately: it trades the audit trail for speed. - **Restoration fails closed.** On a vault miss, render the surrogate or an error; never fall back to a lookup that could surface another conversation's value. - Batch lookups when restoring long conversations. - ⚠️ **Coordinate point-in-time recovery across both stores.** Restoring the transcript database to a different point than the vault yields dangling tokens and unrenderable history. This is normally discovered during the first real restore. - **Possible intermediate, undecided (§6.3.9):** deleting a user's vault rows would leave their transcripts intact but permanently unresolvable — the research record survives, the link back to real values does not. Whether that state is ever wanted is an open question, and it is the only thing `pii.conversation_keys` exists for. Note also that it does **not** survive a restore: the wrapped key sits in the same database as the ciphertext, so point-in-time recovery brings both back. Destroying re-identification for good depends on backup retention as much as on the delete. #### 6.3.6 Conversations and messages ``` conversations (conversation_id, user_id, session_id, started_at, last_activity_at) messages (message_id, conversation_id, seq, role, content, lang, platform, created_at) ``` How an assistant message was produced lives in `pipeline_runs` and `pipeline_steps` — see §6.3.6b. **Object storage is provider-neutral, and now only carries spilled step payloads.** Attachments no longer use it at all — only extracted text is kept, and that lives in Postgres (§6.3.9). The remaining pointer, `pipeline_steps.payload_object_key`, is named that rather than `s3_key` because S3 versus GCS is a deployment decision (§3 is not final) and the bucket is configuration rather than data. **One row per message, not per exchange** (decided 2026-08-06). A row is *either* a user message *or* an assistant message, distinguished by `role`. The alternative — one row holding a query and its response — matches the intuitive meaning of "turn" and matches today's `log_chat_event`, which writes `human_message` and `reply` in a single item. It was rejected because: - **WhatsApp breaks the pairing assumption.** People send several messages in a row before anything answers. A row that must hold exactly one query and one response cannot represent that. - **Messages are the superset.** An exchange is derivable by pairing a user message with the assistant message that follows it; the reverse is not. - **It maps onto the LLM messages array.** P1 requires rehydrating history from the database on every request; `ORDER BY seq` gives that almost verbatim. - **Failure is cleaner.** A request that dies before answering is a row with no successor, rather than a half-populated row. The table is named `messages` rather than `turns` precisely because "turn" reads as *exchange* to most people. Note that elsewhere in this document — §6.5 and §7 — "turn" still means an exchange (one user question and the pipeline's answer to it, which is 4+ LLM calls). That usage is intentional and unrelated to the table. *Migration note:* existing DynamoDB items are exchange-shaped, so backfilling them means splitting each item into two rows and synthesising `seq`. **`platform` ∈ { web, mobile, whatsapp } lives on the message, not the conversation** (decided 2026-08-06). A conversation can start on mobile and continue on desktop, so a conversation-level column would be wrong as soon as anyone switches device. On assistant messages it records the client of the request that produced the answer, which keeps it meaningful on every row. A conversation's platform (or platforms) is a query over its messages. #### 6.3.6b Pipeline runs and steps **An assistant answer is decomposed into a run and its steps** (decided 2026-08-06). It is not an atomic thing: it is produced by an ordered sequence of steps, some LLM calls (draft, judge, redraft, attribution, translation), some deterministic (PII redaction, language detection, tool execution). ``` pipeline_runs (run_id, trigger_message_id, assistant_message_id, skill_variant, wiki_version, outcome, started_at, ended_at, diagnostics_expires_at) pipeline_steps (step_id, run_id, seq, parent_step_id, step_type, model_code, provider_code, prompt_id, prompt_tokens, completion_tokens, started_at, ended_at, status, error_detail, reasoning_tokenized, payload_tokenized, …) ``` **There is no `assistant_message_details` table.** An earlier draft put the assistant-only fields in a 1:1 subtype of `messages`. Introducing runs supersedes it: everything that table held — outcome, skill variant, wiki and prompt version, timing, retention — is a property of the *run*. The message is the run's output, not its owner. Four things this gets right that per-message columns could not: 1. **`model_code` and `provider_code` belong to the step.** The judge may run on different weights than the drafter (§6.5 records that as a live plan), and during a backend cutover one step can be self-hosted while the next hits Vertex. One model per answer cannot express either. 2. **Token counts are per call.** Which is what the impact coefficients actually need (§6.3.7). 3. **A run can exist with no message.** Requests that die before replying are exactly the ones worth investigating — the timed-out redraft that shipped ungrounded content, the turn that bypassed the judge entirely. A message-anchored model cannot record them. 4. **Latency analysis becomes SQL.** `experiments/timing_probe/scaffold_rounds.py` exists to reconstruct judge and attribution round sequences from raw events; with steps as rows, that reconstruction is a query. **Deterministic steps are recorded too**, not just LLM calls. They burn no tokens and their rows are mostly NULL, but the *absence* of a `pii_redaction` step is itself evidence — being able to show that redaction ran on a given answer is worth more than the rows cost. `step_types.is_llm_call` keeps queries from hardcoding which types consume tokens. Two distinctions the schema enforces: - **`judge_rounds` is not stored.** It is `COUNT(*)` over steps of type `judge`. A counter that can disagree with the steps it counts is worse than a query. - **`status` is not a verdict.** A judge returning FAIL executed perfectly: `status = 'ok'`, with the FAIL in `payload_tokenized`. `error` and `timeout` mean the step did not complete. Conflating the two would make "how often does the judge fail" unanswerable. `reasoning_tokenized` and `payload_tokenized` live on the **step**, so the judge's reasoning is attributable to the judge rather than concatenated into one per-answer blob. Both are tokenised like message content: reasoning quotes the user, and tool arguments derive from user text, so both carry PII. `diagnostics_expires_at` on the run gives all of it a retention clock separate from the conversation's. Reference tables — `step_types`, `models`, `inference_providers`, `skill_variants`, `wiki_builds`, `prompts` — are described below. **System prompts are not messages** (decided 2026-08-06). `messages.role` stays `user | assistant`. Four reasons, the first decisive: there is no single system prompt — the main agent, wiki subagent, judge and attribution check each have their own, so a conversation-level system message cannot express "the judge was told X while the drafter was told Y". They are also not conversation events (the user never sees them, so every read for display would need to filter them out), they would duplicate enormous text on every call, and a prompt is a versioned artifact rather than an event. The same reasoning excludes a `tool` role: tool results are step payloads. What a step was told is `pipeline_steps.prompt_id`. **All six reference tables are lookups, not enums or free text** (decided 2026-08-06), because each list changes on its own clock and none should need a migration to extend: | Table | Shape | Notes | |---|---|---| | `step_types` | `step_type` PK, `is_llm_call` | `triage`, `clarify`, `tool_call`, `draft`, `judge`, `redraft`, `attribution`, `translation`, `language_check`, `pii_redaction`, `pii_restoration`. Two of those are not implemented yet, which is precisely why this is a table | | `models` | `model_code` PK, `family` | The version is part of the code (`gemma-4-26b-e4b`), since that is how models are identified in practice | | `inference_providers` | `provider_code` PK, `is_self_hosted` | **Replaces the old `inference_backend` enum.** Self-hosting and the serverless providers (vertex, groq, scaleway, …) are one axis, so one table covers both; `is_self_hosted` keeps the metered-vs-estimated split queryable without hardcoding names | | `skill_variants` | `skill_variant` PK | Mirrors the directories under `agent/skills/`; grows every time a variant is trialled. Not to be confused with `ChatRequest.model_type`, which is the routing choice and maps here rather than to `models` | | `wiki_builds` | `wiki_version` PK (content hash), `name`, `git_sha`, `article_count`, `built_at` | Hash is the identity, `name` is what a human says out loud, `git_sha` ties it to the source revision | | `prompts` | `prompt_id` PK (content hash of the template), `name`, **`text`**, `git_sha` | Replaces the earlier `prompt_versions`, which versioned "the prompt set" as one thing. The useful granularity is per step — the judge and the drafter are told different things — so `prompt_id` sits on `pipeline_steps`. Stores the template **text**, not just a hash, so it can actually be read later | ⚠️ **Foreign keys need the referenced row to exist at write time.** If the app writes an answer citing a wiki build or prompt version nobody registered, the `INSERT` fails and the run is lost to protect its own metadata — a bad trade. Register both at application startup, since the image already builds the wiki. The alternative is to leave those two as plain text with a best-effort catalogue and no constraint. #### 6.3.7 Environmental impact **Per pipeline step**, rolled up to the user through a view (requirement 3 says *their* environmental impact). ``` environmental_impact (impact_id, source_type, step_id, provider_code, node_id, region, energy_kwh, gwp_kgco2eq, water_l, adpe_kgsbeq, pe_mj, measured, created_at) ``` **Impact attaches to a step, not a message** (decided 2026-08-06). Each LLM call has its own model, its own provider and its own completion count — exactly what the coefficients need. One figure per message was always an approximation, and it breaks outright the moment the judge runs on different weights than the drafter. - `source_type` ∈ { inference, infrastructure } — the existing distinction, kept. Infrastructure events (idle node hours) belong to no step, so `step_id` is NULL for them. - `provider_code` / `node_id` / `region` replace the hardcoded `PK="SERVER#HF-Space-01"`, which would otherwise collapse self-hosted and Vertex impact into one indistinguishable pile. - `measured` — true when metered on our own GPU, false when estimated by ecologits. The two are not comparable and should never be summed without the distinction (§6.6). - **No `user_id` column.** It is derivable through steps → runs → messages → conversations, and the `user_impact` view hides that join so the carbon-footprint screen doesn't hand-write it. - **`step_id` is `ON DELETE SET NULL`, not `CASCADE`.** Erasing a user must remove the attribution without silently reducing total reported emissions, so erased rows survive as unattributed aggregate. That is also why no CHECK ties `source_type` to `step_id` — an inference row can legitimately end up with a NULL step. **This is what makes the coefficient question recoverable.** §6.6 flags that `agent.py` feeds `total_tokens` into per-token coefficients that ecologits derives from *output* tokens. With per-step prompt and completion counts stored separately, a corrected figure can be recomputed from history. With one summed total per message, it could not. #### 6.3.8 Guest accounts Requirement added 2026-08-05. **Mechanism is straightforward:** Firebase anonymous authentication issues a real uid with no credentials, and supports *linking* credentials later to upgrade the same uid into a permanent account. A guest who signs up keeps their `user_id` and their history carries over — so do not mint a new `user_id` on upgrade. The `auth_identities` table absorbs this as just another binding. **The hard part is consent, not identity.** A guest who cannot authenticate later cannot exercise erasure, and you cannot verify them if they try. So informed, revocable consent is not really obtainable from them. Therefore, by default: - Minimal collection: **no demographics** (they are consent-conditional). - Short retention on transcripts. - **Excluded from the research corpus.** If they upgrade, ask for research consent then and bring their history in. - Aggregate-only environmental impact. - Keep **IP-based rate limiting** on the guest path — per-user limits are weak when a user can mint a new identity freely (interacts with P5). - Anonymous uids live in browser local storage, so guest continuity is best-effort and clearing site data orphans the account. Schedule cleanup of orphaned anonymous accounts. #### 6.3.9 Open items in this section - Store choice: Postgres vs. keeping DynamoDB (§6.3.1). - Username: display name, credential, or dropped entirely (§6.3.2). - Whether real name is collected at all. - Retention periods per record family — accounts, transcripts, step payloads, vault entries, impact rows. They differ. - How surrogate consistency within a conversation is achieved — see below. - The full `pii_entity_t` taxonomy. The current list is provisional and Malik will supply the complete set once the redactor's entity list is settled. Known gap: a city has no category — `address` is a different granularity and would draw from a different surrogate pool. - Whether erasure ever needs to keep the transcripts while destroying re-identification — see below. **Decided 2026-08-06 — uploaded documents are stored as extracted text only.** The original bytes are discarded after extraction. The text is redacted once at upload and stored like any other user content, sharing the vault, the restore path and the erasure cascade. It is small (`MAX_DOCUMENT_TOKENS` caps it near 120 KB), so it lives in Postgres and **no object storage is involved**. What this rules out, and why the alternative lost: keeping the original would mean redacting on *every read*, so the content is never at rest in de-identified form, and each re-extraction risks producing surrogates that differ from those already baked into the stored transcript. The vault limits that drift — `value_hmac` reuses the surrogate already assigned for a value in that conversation — but a newly-detected identifier still gets a fresh one. More decisively, **an original largely defeats partial erasure**: a stored PDF is un-redacted by definition, so deleting vault rows leaves the child's name still printed on the file. The costs accepted in exchange are that the user sees a filename they cannot reopen, and that OCR mistakes are permanent. Note this closed a real divergence: `attachments` had modelled object-storage pointers while the code stored only extracted text in process RAM (`classes/session_document_store.py`), discarding it at session end. Neither half matched, and extracted text had no column at all. **Open item: is partial erasure ever wanted?** (Malik, 2026-08-06 — undecided.) Two shapes. *Full erasure*: delete the `users` row, everything cascades, nothing survives. This behaves identically under any key design. *Partial erasure*: keep the redacted transcripts as research material but destroy the mapping back to the real values, by deleting that user's vault rows. The schema currently supports the second, and that is the **only** reason `pii.conversation_keys` exists. If erasure is always full erasure, the table is unnecessary: `token_map` should reference `conversations` directly, and values can be encrypted under the app's KMS key with no intermediate data key. The arguments that would survive dropping it are audit granularity — one key use logged per conversation opened, rather than one per identifier — and blast radius. Both real, both modest. Relevant to deciding: partial erasure is a live-data operation, not a guarantee against backups. The wrapped key lives alongside the ciphertext, so a point-in-time restore brings back both. **Open item: how is the surrogate kept consistent within a conversation?** The team may use a different strategy than the one the schema currently assumes (Malik, 2026-08-06 — flagged deliberately, mechanism not yet settled). What §6.3.5 assumes today: the redactor looks up `value_hmac` before substituting, reuses the surrogate already recorded for that value, and the `UNIQUE (conversation_id, value_hmac)` constraint makes a second surrogate for the same value impossible to insert. Consistency is enforced by the database, on the write path. A different mechanism — deriving the surrogate from the value, carrying the assignment in application or session state, or fixing it per user rather than per conversation — would change which parts of the vault are load-bearing. Specifically at risk: whether `value_hmac` is needed at all, whether the unique constraint is the right enforcement point or becomes redundant, and whether `conversation_id` remains the correct scope for both keys. **The restore direction is unaffected** — `(conversation_id, surrogate) → ciphertext` holds regardless of how the surrogate was chosen — so this is contained to the write path. **Open item: store the rendered prompt, or reconstruct it?** `prompts` stores each *template*. The prompt actually sent to a model is the template plus the category index, recalled article text, and conversation history. Whether to persist that too is undecided (Malik, 2026-08-06 — he is not convinced reconstruction is sufficient, and the reasoning below is why that scepticism is well founded). *The case for reconstructing:* the ingredients are all already recorded — `prompt_id`, `wiki_version`, the recall step's payload, and the conversation itself. Storing the rendered text duplicates the index and article text on every call, and it is another copy of user content to tokenise, retain and erase. *The case for storing:* **reconstruction is a derivation, and derivations rot.** Reproducing a rendered prompt needs not just the ingredients but the assembly logic *as it was at the time* — history truncation at `MAX_HISTORY`, uploaded-document blocks, language-correction passes. Re- running that a year later against a changed codebase is fragile in a way a stored artifact is not. For a pipeline whose answers may be questioned clinically, "what exactly did the model see" is the strongest evidence available, and a reconstruction that is subtly wrong is worse than no reconstruction because it looks authoritative. *Rough volume, so the decision is not made on vibes:* ~10–20k tokens per call × ~5 calls per answer ≈ 200–400 KB of rendered prompt per answer. At the §7 crossover rate that is roughly 0.4–0.75 GB/day, or 150–270 GB/year uncompressed — appreciably less after compression, since the text is highly repetitive. **Storage cost is not the deciding factor**; the duplicated copy of user content, and its retention and erasure obligations, is the real cost. *Middle grounds worth considering before choosing:* 1. **Always store a hash of the rendered prompt**, full text only sometimes. Cheap, and it lets you *verify* later that a reconstruction is faithful — which removes most of the risk from the reconstruct option. 2. **Full text to object storage with a lifecycle rule**, not into Postgres: recent answers fully reproducible, older ones reconstructable. Debugging needs are recent; audit needs are older but coarser. 3. **Sample it** — every call for a debugging window or an eval cohort, a fraction otherwise. 4. **Record the assembly as data** rather than deriving it from code: an ordered list of component references (template, article ids, message ids) per step. Reconstruction then depends on data rather than on code that may have changed. The schema already accommodates any of these: `pipeline_steps.payload_object_key` exists for spilling large content, and a `rendered_prompt_sha256` column would be a one-line addition. **Open item: can a user actually be deleted?** As `schema.sql` stands, no — and this needs a decision before erasure is implemented. This matters more now that there is no lifecycle status column: with soft delete gone, a hard delete is the only representation of a deleted account, and it currently cannot succeed. Two things block it. `consent_events.user_id` is `ON DELETE RESTRICT`, so deleting a user whose consent events exist raises a foreign-key violation — and every user has at least one, from signup. Separately, the append-only trigger on `consent_events` fires on `DELETE` as well as `UPDATE` and raises unconditionally; cascading deletes fire child triggers, so switching the foreign key to `ON DELETE CASCADE` would not help on its own. That is deliberate to the extent that it forces the question rather than silently destroying records, because two obligations pull opposite ways: erasure says remove the user's data, while accountability generally requires being able to show that consent was obtained and that its withdrawal was honoured. Deleting the consent log destroys that evidence. The options, once the retention question is answered: 1. **`ON DELETE CASCADE`** — consent events die with the user. Simplest and maximal erasure; no proof of consent survives. 2. **Keep the events, sever the identity link** — retain `occurred_at`, `purpose_code`, `notice_version`; drop or replace `user_id`. Keeps an auditable record that consent activity occurred without it being personal data. Limitation: once severed you can no longer demonstrate anything about a *specific* person. This is an `UPDATE`, so it also needs the trigger scoped to `UPDATE` only — or a documented exception for the erasure path. 3. **Retain in full** — only viable if retention of the consent log is separately justified. Whichever is chosen, the append-only trigger needs narrowing so that the erasure path is possible at all. Left unchanged for now on purpose: the trigger and the foreign key should be adjusted together, once the legal answer determines which of the three above applies. ### 6.4 Self-hosted inference node **Serving stack**: vLLM, OpenAI-compatible HTTP API. That matters because it drops straight into the existing `providers/` protocol — a new `providers/vllm.py` (or the OpenAI provider pointed at a different `base_url`) rather than a new integration. **Card choice.** Approximate figures — confirm against current spec sheets: | Card | VRAM | Mem. bandwidth | FP8 | TDP | Where | |---|---|---|---|---|---| | **A10G** | 24 GB | ~600 GB/s | no | 150 W | AWS `g5` only | | **L4** | 24 GB | ~300 GB/s | yes | 72 W | GCP `g2`, AWS `g6`, Cloud Run GPU | | **L40S** | 48 GB | ~864 GB/s | yes | 350 W | AWS `g6e` | | A100 40 GB | 40 GB | ~1555 GB/s | no | 400 W | both, much pricier | Decode is memory-bandwidth-bound — roughly, each generated token requires reading the active weights out of VRAM — so bandwidth, not FLOPS, sets token throughput. Given that this pipeline's measured latency is dominated by judge output tokens, **bandwidth is the number to optimize.** The A10G and L4 have similar tensor FLOPS and very different bandwidth; comparing them on FLOPS is the easy mistake. **"E4B" = 26B total parameters, ~4B activated per token** (confirmed 2026-08-05). This is a sparse MoE, and it has two consequences that pull in opposite directions: - *Capacity is set by the full 26B.* All experts must be resident. At bf16 that is ~52 GB; INT4 (AWQ/GPTQ) lands around ~13 GB. **Quantization is not optional on a 24 GB card.** FP8 (~26 GB) does not fit, so the L4's FP8 support buys nothing for weights here — only for KV cache and activations. - *Bandwidth per token is set by the ~4B active.* At batch size 1 that is only ~2 GB read per token, so decode is cheap on either card. **But this does not narrow the A10G/L4 gap where it matters:** as concurrency rises, the union of experts activated across a batch grows toward the full 26B, so the node behaves closer to a dense model under load — and the node only runs during high traffic, i.e. at high batch. The bandwidth advantage holds. **KV cache, not weights, is the tight resource.** ~13 GB of weights plus ~1–2 GB of framework and CUDA-graph overhead leaves roughly 8–9 GB of a 24 GB card for KV cache. That is workable, but vLLM's prefix cache draws from the same block pool — less headroom means a lower hit rate, and prefix caching is the biggest throughput lever available here (below). Tune `--gpu-memory-utilization`, measure the cache hit rate rather than assuming it, and check whether FP8 KV cache is supported for this model on Ampere if you need more room. ⚠️ **MoE-specific quantization risk.** Experts carry less redundancy than dense layers, so MoE models are frequently more sensitive to 4-bit quantization than a dense model of the same parameter count. Confirm the AWQ/GPTQ config keeps the **router** and shared layers at higher precision — a degraded router silently selects the wrong experts, which the eval harness would catch but a smoke test would not. **Worth pricing: `g6e` / L40S.** 48 GB and ~864 GB/s means FP8 weights fit comfortably and decode is faster than A10G. It costs more per hour, which raises the serverless break-even (§7) — but it removes INT4 quality risk entirely. On a pipeline where a judge adjudicates individual factual claims, that trade may be worth buying. If you stay on INT4 — which the current plan does — **the quantized model must be re-run through the eval harness** (`experiments/run_simple_eval.py`) before it is trusted. INT4 quality loss on a clinical grounding pipeline is not something to assume away, and the MoE sensitivity above raises the stakes. ⚠️ Confirm `g5` and `g6e` availability in `ca-central-1` — not every GPU family ships in every region. ⚠️ Also verify the "E4B" naming: if Gemma 4 follows the Gemma 3n convention, *E4B* means ~4B **effective** active parameters out of 26B total. That improves throughput substantially but does **not** reduce resident weight memory unless the serving stack supports offloading the inactive parameters — confirm what vLLM actually does with this architecture before sizing the card. **Prefix caching is the highest-leverage setting you have.** Every judge and redraft call resends a large shared prefix (system prompt + the category index + recalled article text). vLLM's automatic prefix caching turns that from recomputation into a cache hit. Given that one user turn is 4+ LLM calls over largely overlapping context, enable it and measure — this is likely worth more than any other single tuning knob. **Warm-up is the operational hazard.** Loading 13 GB of weights plus CUDA graph capture takes minutes. The node must not receive traffic until `/health` reports the model ready, and the autoscaler must not treat "instance running" as "instance serving." **Observability**: vLLM exports Prometheus metrics — `vllm:num_requests_waiting`, `vllm:gpu_cache_usage_perc`, time-to-first-token. These are the inputs to the routing decision in §6.5, so wire them up before the router, not after. ### 6.5 The router between the two backends This is the novel part of the design and deserves explicit rules. **Placement.** One `ModelRouter` in the backend chooses a provider per *call*, not per session. Call sites (main agent, subagent, judge, attribution) ask the router for a provider; they do not know which backend they got. Per-call granularity is what makes mid-conversation failover possible. **Modes.** `INFERENCE_MODE = serverless | self_hosted | auto`, read from config, not baked into the image. Ship with the manual modes first; `auto` comes later (§8, Phase 3). **Scale up on leading indicators, not lagging ones.** The obvious trigger — "we started getting 429s from Vertex, bring up the GPU" — is the wrong one, because the GPU takes minutes to warm and you would be down for all of it. Trigger on arrival rate (requests/min over a rolling window) crossing a threshold for D minutes. Note the existing 429-backoff machinery in `providers/hf.py` (up to 8 retries, waits up to 65 s): under Vertex the same pattern would silently absorb a load spike as latency instead of surfacing it as pressure. Make rate-limit waits a first-class metric feeding the scale-up signal. **Scale down with drain and hysteresis.** Stop routing new work to the node, let in-flight calls finish, then terminate. Hysteresis (a long quiet period before scale-down, longer than the scale-up threshold) prevents thrashing on a workload that is bursty by nature. **Failover is one-directional and per-call.** If the self-hosted node errors or times out, that call retries on Vertex. The reverse (Vertex fails → self-hosted) only works if the node is already warm; do not build it as a recovery mechanism. **⚠️ Prompt/quality parity is a genuine risk, not a formality.** Your own notes already record that `reasoning_effort` is passed through `extra_body` because it is specific to gpt-oss's harmony format. The same class of divergence applies here: the same weights served by vLLM and by Vertex can differ in chat template application, default sampling parameters, tool-call serialization, and stop-token handling. A pipeline that depends on structured JSON from a judge is exactly where that bites. **Gate the router behind an eval-parity check**: run `experiments/run_simple_eval.py` against both backends and compare pass rates before allowing `auto` mode. Per your existing methodology note, stratify by clarify outcome rather than pooling into one number. **Billing model: confirmed per-token, no idle cost** (Malik, 2026-08-05). This is what makes the two-backend design work — the low-traffic path genuinely costs nothing when idle. The established crossover is **~78 requests/hour** against a self-hosted A10G; see §7 for what moves it. **Cloud Run GPU: investigated, low priority.** Cloud Run offers scale-to-zero request-billed GPUs, but they are **L4s** — the same bandwidth ceiling described in §6.4 — plus a multi-minute cold start to load multi-GB weights. Against a per-token Vertex endpoint with no cold start, it does not win the low-traffic path, and against an A10G it does not win the high-traffic one. Keep it on the list, but it is not the subsystem-eliminating shortcut it first appeared to be. ### 6.6 Environmental impact accounting Self-hosting changes this qualitatively and mostly for the better. - **On your own GPU you can measure instead of estimate.** codecarbon (or direct `nvidia-smi` power sampling) on the inference node gives real energy draw, replacing ecologits' token-based approximation. Keep ecologits for the Vertex path, where you have no hardware visibility. - **Update the electricity mix constants.** `helpers/impacts_tracker_helper.py` currently carries US-grid factors (`QWEN_ELECTRICITY_MIX_GWP = 0.38355 kgCO2eq/kWh`). Quebec's grid is overwhelmingly hydro and roughly an order of magnitude cleaner. If you host in Montreal, using US factors understates your advantage substantially — and for a project that *publishes* impact numbers, using the wrong grid is a correctness bug. - **Attribute idle cost honestly.** A GPU node held warm during low traffic burns energy and amortizes embodied carbon whether or not it serves requests. If per-request impact only counts active inference, the published numbers are wrong in a way that flatters the self-hosted path. Log an `infrastructure`-type event per node-hour (the schema already distinguishes `inference` from `infrastructure`) and decide explicitly how idle time is allocated across users — or report it separately rather than amortizing it. - Emit `provider_code`, `node_id`, and `region` on every impact event so the paths can be compared. Right now they would be indistinguishable. ### 6.7 Cross-cutting - **Secrets**: Secret Manager (or AWS Secrets Manager), injected at runtime. The current `.env`-file pattern and `constants.py`'s import-time `raise` on a missing `HF_TOKEN` need to become a startup health check instead of an import crash. - **CI/CD**: build the API image and the web bundle on push; deploy to staging automatically, to prod on tag. Note the Dockerfile runs `build_articles.py` at build time — keep the wiki build reproducible and pinned, since it determines what the model can ground on. The split index (`data/wiki/built/index_by_category/`) must be generated in the same step. - **Environments**: `dev` (local + DynamoDB Local / Firestore emulator), `staging`, `prod`, with separate projects and separate data stores. Staging must not read production conversations. - **Logging**: `classes/pii_filter.py` exists — make sure it is applied on the logging path, not only the model path. Log lines currently include full chat items (`logger.info(f"Logging conversation: {item}")` in `dynamodb_helper.py`), which will put user text into Cloud Logging. Fix before any real user touches this. - **Retention & backup**: point-in-time recovery on the database; a documented retention period per record family; the erasure path from P7. - **Network**: the GPU node should not be internet-facing — private VPC, reachable only from the backend service. --- ## 7. Cost model **Established: the crossover is ~78 requests/hour** — below that, the per-token Vertex endpoint is cheaper than a self-hosted A10G; above it, the node wins. That figure is the premise of the whole two-backend design. Two things will move it, and both should be re-derived before the router is tuned: **The threshold rises when judge rounds drop below 2.** Node cost is fixed per hour; fewer judge rounds means fewer tokens per request, so the serverless side gets cheaper per request while the node does not. Expect the crossover to land somewhat above 78 once the improvement ships — meaning the node is worth running *less* often than currently planned. Pull `tokens_per_turn` from the timing-probe data (it already records per-stage token counts) rather than re-estimating. **A bigger card raises it too.** If §6.4's `g6e` / L40S option is chosen to escape INT4, the higher hourly rate pushes the crossover up proportionally. ### The missing number: saturation 78 req/hr is the **lower** bound of the node's useful window. The upper bound — the arrival rate at which one A10G saturates and latency degrades — is not yet known, and it determines whether "one node during peaks" is a viable design or whether peaks need a scaling group with different economics. This requires a load test against the real four-call pipeline. vLLM **prefix caching will move the ceiling substantially**, because the shared prefix (system prompt + category index + recalled article text) is resent on every judge round. Note that caching raises *capacity*, not the break-even — the node's hourly cost is the same whether or not it is cached. ### Switching policy Do not switch at exactly 78. With a multi-minute warm-up, traffic oscillating around the threshold will thrash the node, and running it at 79 req/hr saves nothing while costing warm-up time and risk. Trigger scale-up at a comfortable multiple of the crossover sustained over several minutes, and scale down only after a longer quiet period (§6.5). Costs easy to forget: idle hours of a warm node, ElastiCache/Memorystore, the load balancer, egress, and the second cloud's baseline. --- ## 8. Phasing **Phase 0 — unblock (no new infrastructure).** P1–P7 from §5. Externalize state, add auth, lock down `model_type`, move table creation to IaC, shared rate limiting, telemetry on, erasure path. Independently: get the compliance policy and classify the data. **Phase 1 — deploy with one inference backend.** Static UI on Firebase Hosting, backend on ECS Fargate / App Runner in `ca-central-1`, DynamoDB moved to `ca-central-1`, **Vertex as the only inference path**. No GPU node, no router. This ordering is deliberate given that traffic is unknown: you cannot size a GPU node without arrival-rate data, and a serverless-only deployment produces exactly that data at zero idle cost. It also means the first production deployment has one fewer novel component to debug. **Phase 2 — add the self-hosted node, manually switched.** vLLM image, quantized weights, warm-up gate, private networking, `providers/vllm.py`, metrics. Flip `INFERENCE_MODE` by hand and run the eval harness against both backends. Do not automate routing until parity is demonstrated. **Phase 3 — automatic routing.** Arrival-rate scale-up with hysteresis, drain-on-scale-down, per-call failover, cost dashboard showing actual vs. break-even. **Phase 4 — hardening.** DR and restore rehearsal, retention enforcement, audit logging, completed privacy impact assessment, load test against the real topology. --- ## 9. Open questions and verification list Still blocking: 1. **Data classification** — PHI, Law 25 personal data, or de-identified? Malik does not yet have the full policy. Gates the vendor list and the retention/audit requirements, though §3's topology holds either way now that residency is settled. Parked until closer to implementation (Malik, 2026-08-05 — none of these block Phase 0): 2. **`g6e` / L40S pricing** (§6.4, §7) — buys FP8 and escapes INT4 quality risk, at a higher crossover point. 3. **`g5` / `g6e` availability in `ca-central-1`.** 4. **Single-node saturation rate** (§7) — the upper bound of the node's useful window; needs a load test against the real four-call pipeline. 5. **vLLM's handling of this MoE architecture** (§6.4) — expert placement, and whether FP8 KV cache is available on Ampere for it. Resolved: - ~~Vertex billing model~~ — per-token, no idle cost; crossover ~78 req/hr. - ~~Firebase Auth residency~~ — data can reside in Canada; Identity Platform is usable (§6.2). - ~~Gemma 4 "E4B" semantics~~ — 26B total, ~4B activated per token. Fits a 24 GB card at INT4; KV cache is the constraint (§6.4). - ~~Cloud Run GPU as a two-backend replacement~~ — L4-based with multi-minute cold starts; deprioritized (§6.5). - ~~Cloud choice~~ — Option B (§3), driven by the A10G bandwidth advantage. Team decisions with no external dependency: 6. Retention periods per record family. 7. Whether the legacy Jinja UI survives. 8. Whether idle GPU energy is amortized into per-user impact or reported separately (§6.6).